← All notes · 10 min read

Synthetic Data's Greater Potential

Synthetic data's potential marketing goes well beyond completing surveys

Photo by Hunter Harritt on Unsplash
Photo by Hunter Harritt on Unsplash
Synthetic Data as a Demand Intelligence Layer

In my last post, I discussed some of the limitations of synthetic data, particularly the digital twin approaches that run the risk of committing a homunculus fallacy. I want to change my tone for this one to be more optimistic. There is immense potential in synthetic data for marketing intelligence. In fact, I think this potential is currently undersold. Synthetic data methods could solve the data integration challenge marketers face today. They could be the single source of truth for understanding customer demand.

The Underwhelming Offers Dominating Synthetic Data

The most frustrating part of the synthetic data discussion in marketing right now is how much of it centers on pretty low-stakes applications. Researchers out there are touting synthetic respondents as a means of accessing “hard to reach sample” and as “concept screeners” to replace some concept testing. Reading between the lines, the impact seems relatively marginal.

Consider the hard-to-reach sample use case. In most situations where I would even consider a synthetic respondent to fill quota, it’s because the client just has a checklist of quota cells they need to hit. Without synthetic respondents at my disposal, I would recommend using findings from the rest of the research to infer what is happening in the hard-to-reach quota. If we know what men think, and we know what Gen Z thinks, and we have a handful of male Gen Zs to check against, we can normally infer what a more robust read of Gen Z males would say. And if that doesn’t work, because we think Gen Z males are very different for some reason, or because those handful of examples are giving wildly different signal from what we expect, then the answer is more original sample, not synthetic sample.

The main innovation here is quantifying these qualitative inferences and putting uncertainty measures around them, which is valuable. But it isn’t game-changing for practitioners like me. I see this being useful primarily in large programs where hitting every single quota cell is a non-negotiable requirement, and where the client-side insights manager you’re working with does not have the decision-making authority to change the requirements (e.g. a global tracker). In this case, synthetic respondents are mostly a means for large survey research companies to check more boxes on an RFQ. They will probably become standard in the industry. But not game-changing.

The concept screener offerings are probably more of a dead-end. These are cases where we want to run a concept, creative, or some other innovation through a model to score its market potential instead of running a concept test study. This is pretty clearly a response to market pressures for faster and cheaper results, always at the expense of quality. I can’t tell you the number of requests I’ve seen to take rich and immersive behavioural methods and compromise them to be ever faster and ever cheaper, until their original value proposition was nearly unrecognizable.

So I get it. That’s what the market is asking for. What’s faster and cheaper than your fastest and cheapest possible concept test? Not doing the test at all, and just running it through the algorithm.

The problem is the confidence people have in these approaches is embarrassingly poor. The consistent recommendation right now is to use them for “early stage concepts” to make real concept testing later more efficient. Results are positioned as “directional” and suitable for “low stakes” situations. Which to me reads as, “we don’t have confidence that this works but outright calling it wrong would be rude.”

I think there are only two directions this goes. Either these tools evolve into more bespoke solutions, grounded in a rigorous enough foundational research program to be useful for the specific client sponsor. Or everyone decides that the even faster and cheaper solution than an unreliable algorithm is to just not test concepts at all, and be done with the whole farce.

The amount of industry hype around low-stakes scoring algorithms and RFQ checkboxing tools is frustrating because the underlying technology could do so much more. I think this is mostly an artefact of trying to find product-market fit in a market that is fairly regimented in what and how it buys. You have to sell what people are buying. But since I’m not selling either of these things, let’s explore what the bigger potential could be.

The Problem Synthetic Data Could Solve

I mentioned before that synthetic data has the potential to be a single integrated data layer that could have a far-reaching impact in marketing. To understand this, let’s drill down on the problem a bit further.

The need for a single layer comes from the proliferation of distinct data sources that influence marketing decision making. Each of these sources has three characteristics that make it valuable. First, gathering: a method for collecting information. Second, transformation: a process for converting that data into another form. And third, action: a way to act on the data to get value from it. Data is collected, analyzed in some way to make it meaningful, and then acted upon to extract value.

As marketing operations have become populated by a growing variety of toolsets and vendors, data sources have grown up alongside them. This is because of the obvious utility of adding market sensing capabilities to each tool set. If you partner with a credit card company, for example, to reach their customers, they’ll set up tools to monitor their customer base’s buying behaviour (gather), identify buying signals (transform), and then launch a campaign to target those signals (action). The tighter that loop between market sensing and acting, the more value the marketing partner provides.

But those market signals are tailor-built for their tools, and they don’t tell the general story. In the credit card example, the signal received doesn’t tell you about the whole market. It only tells you about the credit card company’s customers. And worse, really only the subset of spending those customers actually do through that credit card. These insights are exceedingly valuable for marketing through the credit card partner. But if you try to generalize the findings you receive from the partner to other tools and other contexts, you’re going to find they fall flat.

So now you have many tools, and many partners, and many narratives, grounded in many different data sets.

You could try going to “objective third parties” whose job is just market sensing divorced from execution. But these have their own challenges.

Looking at industry press, trend hunting, or other journalistic sources, for example, can surface interesting narratives. But you’ll be hindered by their own gather-transform-action loops. A journalist gathers information by interviewing people, reading press releases, or digging through social media. They take that information and transform it into an article. The action happens when the article is read, shared, and repeated in boardrooms, generating attention and influence for the journalist. The result is a gather-transform-action loop optimized for novelty and attention seeking rather than critical perspective.

Or you could commission some market research. Market research often presents itself as a neutral third party, which seems like a good option. The problem is that neutrality usually comes from a lack of integration with tools to action the results. So you commission a study, the supplier gathers some data, they transform it into a compelling report, and… nothing happens. There is no mechanism built into the process to do anything about it. This normally results in ballooning investment into the transformation step. The industry does more and more work to analyze data into “actionable results” so that when it lands in a marketing manager’s hands, they know exactly what to do with it.

Unfortunately, this investment often comes at the expense of data gathering, which inevitably leads to the kind of situation I described earlier, where the industry tries to find ways around the data gathering step altogether. And even if research firms made stronger investments in data gathering, it’s dubious that any standalone data gathering effort could match the practical utility of data gathered natively in the marketing activation workflow.

Some folks have also tried to tighten that research-to-action loop by conducting research with their creative agency or a management consulting firm. This is fraught with the same misaligned incentives as the other examples. The gather-transform-action loop remains tightly coupled with the other activities of the firm, and the results inevitably grow to serve the business of the agency or consultancy doing the research.

So what do we do? None of the data gathered by anyone in these roles is wrong, presumably. It’s just biased to serve the function of the tools and organizations that execute it. And if you try to remove that execution function, the utility of the results drops off precipitously.

This is where I think synthetic data could play a powerful role.

The Opportunity for Synthetic Data

Synthetic data, at its heart, is an algorithm. A model of the world we build in a machine to understand how it works. It is not at all new to the world of business. The discipline of management science is effectively built on it. Conjoint analysis is synthetic data. Sales forecasts are synthetic data. Marketing mix modelling is synthetic data. We have been doing this for decades.

What is new is the scale and sophistication of these models. We can now reasonably simulate markets and consumer populations as a whole.

So what does this have to do with the data integration problem?

If we have many data sources, each with their own gather-transform-action loop, each producing a partial and biased view of the market, then what we need is a model that can take these partial views as inputs and produce a more general output. A model that doesn’t replace the data sources but ingests them. A model that we can query with new questions and get reasonable answers from, without commissioning a new study.

This is conceptually very different from what most synthetic data vendors are selling today. They are selling outputs. Answers to specific surveys. Scores for specific concepts. The bigger opportunity is to sell the model itself. Or rather, the analytical layer the model enables. A demand intelligence layer.

What would this look like in practice?

Imagine taking your brand health tracker and integrating it with your media mix model. Not stapling reports together. Actually integrating the underlying data, such that when your brand health score moves in a region, you can immediately see how the media spend in that region might have contributed, and project forward what spend adjustments would do to the brand metric. Now add your retail panel data, your social listening, your loyalty program. Each source feeds the model. The model resolves contradictions where it can, and quantifies uncertainty where it can’t. And critically, when a new question comes up, you don’t commission a study. You query the model.

This is what an integrated demand intelligence layer looks like. Market sensing infrastructure sitting underneath your marketing operations, doing the integration work that no individual tool or partner is incentivized to do.

How, exactly, this would be achieved is still TBD. But, there are a lot of different bets being placed in this space.

The first is the digital twin or synthetic respondent model. Companies like Aaru, Simile, Evidenza, Synthetic Users, and AskReplicas use LLM-based agents trained on interviews, behavioural data, or personas to simulate how consumers might respond to products, messaging, pricing, or experiences. I covered this approach in the last post, where I warned against committing the homunculus fallacy. But that doesn’t mean any of these specific companies are committing this error, so they are worth evaluating cautiously.

The second category is emerging from traditional market research and panel providers. Qualtrics, Toluna, YouGov (following its acquisition of Yabble), NielsenIQ, and Ipsos are layering AI-generated insights, synthetic sample augmentation, and automated analysis onto their existing human respondent ecosystems. This is a hybrid play, and it has the advantage of being grounded in real data these firms already own. But the underlying product is still mostly about filling out survey responses faster. Which loops us back to the low-stakes applications I started this post complaining about.

The third approach is the most structurally different, and the one I find most promising. The synthetic population or market simulation model, exemplified by Arima, builds census- and demographic-grounded synthetic populations to support applications like media planning, audience intelligence, marketing mix modelling, retail strategy, and macro demand forecasting. The advantage of starting from a population, rather than from agents or survey responses, is that the population grounds the model in real-world distributions and real-world geography.

Which approach will win is anyone’s guess. Methodologically, I’m most partial to population simulations like Arima. But, digital twin companies appear to have attracted the most investment. And hybrid solution are support by industry incumbents that can get these solutions in the hands of users right away. So it’s a competitive race.

Thinking Bigger About Synthetic Data

The real potential for synthetic data is to be that integrated demand intelligence layer. A single layer that pulls the proliferation of marketing data sources together into one place. It’s hard to overstate how significant that would be.

Which is why the current discourse is so frustrating. It’s all cost cutting. Faster concept tests, cheaper hard-to-reach sample, fewer surveys. Synthetic data positioned as a way to spend less on the things you were already doing.

But I think that’s the wrong question. The more interesting question is what becomes possible that wasn’t possible before. Asking questions and getting answers in minutes instead of months. Testing strategies across dozens of markets before committing. Connecting brand health to media spend to retail behaviour in a single model. These aren’t just

cheaper versions of things you already do. They’re things you couldn’t do at all before, at any price.

First published in The Knowledge Stack.

The notes, by email

Roughly monthly. No sequence.

More notes

All notes →
July 2026 · 14 min

Which Model Should Answer Your Survey?

July 2026 · 10 min

Roleplay or Reasoning?

June 2026 · 15 min

Experimentalist Innovation in Distributed Retail