← All notes · 15 min read

The Homunculus Fallacy

How digital twins can mislead in an era of synthetic data

There’s a particular mode of AI deployment in marketing that’s proving quite popular in certain circles - Digital Twins. They are attracting more than their fair share of attention and dollars at the moment. The total digital-twin AI market is expected to reach $29 billion by 2029, with over 900 start-ups globally competing for a piece of it. Applications range from enhancing clinical research to marketing intelligence - the latter being the obvious focus of this blog. Not without controversy, of course. But through the din of investor and marketing hype, I think there’s a critical fallacy being committed here. Maybe not in all cases but certainly in some. And, if we could better understand this fallacy, we’d be able to better engage with these new approaches to consumer intelligence with clarity. It’s a modern version of the homunculus fallacy.

The Promise of Digital Twins

First definitions, because these terms are not used very consistently. A ‘Digital Twin’ is a computer simulation of something that exists in the real world. It has its roots in mechanical engineering and has found additional success in urban planning and transportation. By creating a digital model of a thing in the real world you gain the ability to plug it into a simulation and see what it does. It can help with planning in very complex scenarios where results could be chaotic. These digital twins improve our understanding of objects and systems by formalizing our assumptions about how they work and testing those assumptions against the real world.

The cause of recent excitement over digital twins is the introduction of AI tools, particularly LLMs, to these simulation models promising to better simulate humans. The basic idea is that if you condition an LLM on enough information about a real person, the model can generate responses that approximate what that person would actually say or do — and you could use this for market intelligence in place of direct research. These AI tools could answer surveys, evaluate new concepts, and even be placed in large-scale market simulations to explore campaign outcomes.

Excitement for this approach really took off after a 2024 paper from Stanford by Joon Sung Park and colleagues titled "Generative Agent Simulations of 1,000 People." The team conducted two-hour structured interviews with 1,052 Americans and used the transcripts to construct generative agents intended to twin each participant. The agents replicated participants' answers on the General Social Survey at roughly 85% of the rate at which the participants themselves replicated their own answers two weeks later. The researchers went on to commercialize this approach by founding a company called Simile and raising $100 million in funding.

Others soon followed the Simile example. PyMC Labs and Colgate-Palmolive published a paper in 2025 on what they call Semantic Similarity Rating, which has formed the basis of PyMC’s digital twin business. YouGov acquired Yabble, adopting its Gen-AI Personas (a digital twin variation). Ipsos struck a research partnership with Stanford to commercialize academic work in this space. NielsenIQ launched BASES AI, and Qualtrics rolled out Edge Audiences. The speed and enthusiasm for adopting this technology is virtually unprecedented.

The appeal of this technology is very clear. Give an AI some context and instruction about who it should roleplay as and how it should act, then present that AI with a scenario, and record the response. The ability to incorporate real intelligence in our models of consumer behaviour promises to make them truly generalist, letting us explore all different markets and contexts. And we can do this infinitely faster and more flexibly than running surveys, A/B tests, test market studies, or virtually any other type of market intelligence activity. This would be a huge win if it works.

Pushback

Others in the industry share my skepticism. This enthusiasm is misplaced."

I think there is a very specific type of error happening here that will cause significant problems down the road when initial investor enthusiasm wears off and the real commercial impact of these technologies becomes apparent. And this is specifically the idea that an AI tool, instantiated with the right context and prompting, can provide insight into how people would think, behave, and react to novel situations. There is a fundamental error at the heart of this proposition. The homunculus fallacy.

The “homunculus fallacy” is a term first coined by Anthony Kenny in 1971 to describe a logical error in theory of mind explaining mental processes by attributing human-like cognitive functions (like seeing, thinking, or remembering) to individual parts of the brain or internal "little men" (homunculi). An example would be explaining vision in terms of the eye casting light upon the retina in a particular pattern, which the brain then “sees” as if there were some little person in the brain doing the seeing for it. It builds on Wittgenstein's view that only a human being as a whole can be said to see, hear, or feel, not a body part. Anil K. Seth has recently made a related argument about philosophy of mind and AI consciousness.

Applied to the concept of digital twins, I believe a very similar category error is happening. When a vendor tells you a digital twin in their synthetic panel "thinks like" a 45-year-old grocery shopper from Cleveland, or that their digital twin "prefers" certain product attributes, or that their AI persona "represents the views of" a demographic segment, they are doing something philosophically suspicious. They are taking properties that meaningfully apply to whole, embedded, biographical persons and they are attributing those properties to, essentially, a token-prediction system. There is no little 45-year-old grocery shopper inside the model. There is a mathematical function that produces likely next words given a context. And if we remove this word ‘intelligence’ from our artificial systems, as it seems to pre-suppose an entity existing within it, and focus on its token-prediction functionality, it’s actually extremely surprising that a string of text describing the demographic and psychographic characteristics of a person would somehow accurately predict that real-world person. People are not strings of text, after all. Treating them as such is the error.

And, aside from the headline-grabbing positive results that helped raise millions of dollars in venture funding, the rest of the academic record on digital twins is less impressive.

What the Empirical Record Shows

The most damning empirical evidence on this question is recent. In 2025, a research team that included Olivier Toubia, a marketing professor at Columbia, published “A Mega-Study of Digital Twins Reveals Strengths, Weaknesses and Opportunities for Further Improvement.” It is the largest and most rigorous test of the digital twin approach to date. It consists of nineteen pre-registered new studies, run in parallel, comparing real human samples to synthetic digital twins. Each twin was constructed using its corresponding human's answers to around five hundred prior survey questions.

The results were not encouraging. Across the nineteen pre-registered studies:

  • The average correlation between twin responses and human responses was approximately 0.2, roughly the predictive lift you might get from knowing someone’s age and gender.

  • Responses were systematically under-dispersed, meaning the spread of opinion was much narrower than in the human sample.

  • Outputs were biased toward higher-education, higher-income, and moderate respondents.

These failures showed up despite the rich personal grounding each twin had been given.

The mega-study findings line up with a body of independent work that has been accumulating since 2023. There’s a lot now and the list is growing but a greatest hits list includes:

  • James Bisbee and colleagues found that 48% of the regression coefficients estimated from synthetic data were significantly different from the real-data coefficients, and among those, the sign of the relationship flipped 32% of the time

  • Bisbee also found silent model updates by OpenAI resulted in the same prompt run in April, June, and July of 2023 producing different response distributions

  • Mohammad Atari and colleagues found GPT values clustered closely to the United States and English-speaking, wealthy nations, and that the correlation between GPT-human similarity and a country’s cultural distance from the United States was negative 0.7

  • Malik Stromberg and colleagues tested LLM responses against YouGov BrandIndex data on 114 brands found systematic overprediction at the consideration and purchase-conversion stages

  • Shibani Santurkar and colleagues at Stanford found that the gap between LLM responses and the average American on opinion questions drawn from the Pew American Trends Panel was on par with the gap between Democrats and Republicans on climate change, and that persona steering did not close it

  • Tiancheng Hu and Nigel Collier found that persona variables, the demographic prompts vendors lean on so heavily, accounted for less than 10% of the variance in LLM annotations

  • Lindia Tjuatja and colleagues at CMU found that LLMs respond strongly to manipulations that don’t affect humans, like minor typos, while shrugging off manipulations that produce well-known biases in real human respondents

  • Gati Aher and colleagues, replicating classic economic and psychological experiments, found a “hyper-accuracy distortion” where LLM crowds outperformed humans on factual subtasks

Why It Was Never Going to Work

You might respond to all of this by saying the empirical failures are temporary. Today’s models are flawed, sure, but tomorrow’s models will be better. It’s a data or engineering problem. This is the worst they will ever be. And so on. But there are some deeper conceptual reasons to think this category of approach is headed in the wrong direction.

Consider SimpleBench, a popular benchmark designed to test the kinds of everyday reasoning an average person handles without much effort. Things like spatial-temporal reasoning, basic social intelligence, and recognizing when a question is phrased as a trick. Non-specialized humans score around 84%. The same frontier LLMs being marketed as digital twins of consumers score well below that, and you can check the live leaderboard for yourself. This is a systematic gap between how a model handles language about the world and how a person who actually lives in the world handles the same question. By extension, a model that has read every word ever written about consumer behaviour has not, by virtue of that reading, become a consumer. It has become very good at producing text about consumer behaviour. The outputs may look identical, but they are produced by structurally different processes, and the difference shows up the moment you ask the system to do something its training data does not cover.

This compounds with what is actually in the training data. Marketers tend to think of LLMs as having read everything, and in some sense they have. But everything here means the slice of the world that exists as scrapeable text on the public internet, and that slice is narrow. Consider just the demographics: roughly two-thirds of Reddit users are men, mostly aged eighteen to twenty-nine. Wikipedia editors are overwhelmingly male, with female participation hovering somewhere between nine and fifteen percent. The most demographically representative discussions of consumer behaviour, the ones held in proprietary panels and behind publisher paywalls, are largely absent. What remains is fan fiction, hobbyist forums, marketing copy, opinion blogs, social media, and Wikipedia. The web is not the world.

How LLMs Are Engineered Away from Being Human

Beyond the empirical and the ontological arguments, which can be a bit abstract for some, there’s also a very practical challenge: LLMs are not designed to be human and these differences are not an accident. They are designed, very deliberately, to behave in ways that are unlike how an arbitrary human would behave. An LLM that produced the median output you would get from a random sample of internet text would be unhelpful, frequently offensive, and commercially unviable. So, the labs intervene, with intent, to shape the model’s outputs in particular directions. Once you understand what those interventions actually do, the idea that the resulting system can serve as a neutral proxy for “what humans think” becomes very hard to maintain.

The most consequential intervention is reinforcement learning from human feedback, or RLHF. The basic idea is that after the model has been pre-trained on a large corpus of text, a much smaller group of human labelers is asked to rate model outputs against various criteria. The model is then fine-tuned to produce more outputs that the labelers liked and fewer that they disliked. This is the step that turns a pre-trained model into a “helpful, harmless, and honest” assistant, and it is also where most of the model’s “personality” is set. It’s worth asking who is conducting this RLHF work and with which framework, as it is these values that are ultimately reflected in a model’s preferences instead of “humans” in any general sense.

In addition, a more recent intervention is something called Constitutional AI, pioneered by Anthropic. The idea is that rather than relying purely on labeler ratings, you give the model an explicit document of principles, a “constitution,” and use it to guide both labelling and model self-correction. Anthropic has published its constitution, and the contents are very illuminating for our purposes. One of the explicit principles reads: “Choose the response that is least likely to imply that you have preferences, feelings, opinions, or religious beliefs, or a human identity or life history, such as having a place of birth, relationships, family, memories, gender, age.” Note how antagonistic this version of a constitution is to the idea of digital twins in market research.

Going deeper, there is also the question of training data filtering. The widely-used Colossal Clean Crawled Corpus (C4), which underlies many large models, was filtered with a profanity blocklist intended to remove offensive content. Researchers at the Allen Institute showed that this blocklist removed 42% of African American English documents and 32% of Hispanic-aligned English documents, compared with about 6% of White-aligned English documents. Minority dialects were filtered out at roughly seven times the rate of the dominant dialect, before any RLHF, before any constitutional principles, just at the data preparation stage. If you then fine-tune that model and ask it to “respond as a 35-year-old African American mother in Atlanta,” the model will do its best, but its best is a reconstruction from an already-impoverished signal, and the reconstruction will lean heavily on whatever stereotypes were over-represented in the surviving corpus.

All of these concepts are important to the issue of alignment in AI research, which is a central and evolving concern for AI researchers. The concept that LLMs would naturally reflect the latent preferences of humans by default is simply false. And this is well understood, acknowledged, and actively intervened on by leaders in the industry.

AI Still Has a Place

I don’t want this essay to come across as a blanket objection to the use of AI in market intelligence. I am very bullish on the technology and think it should be part of all future workflows. It’s specifically the attribution of human-like attributes to AI, and the methodological implications of doing so, that I believe is problematic. AI tools are remarkably good at processing, summarizing, and reasoning over text, which has huge implications for the understanding and application of research. Adjacent innovations in semantic embedding technologies will change how we handle qualitative data. That is to say, we should treat AI as an analytic tool in the research process, first and foremost, and not the subject of the research itself.

Key Take-Aways

The modern homunculus fallacy can be summarized thusly:

  • Understanding human characteristics requires reference to actual humans because of their embodied nature as real physical, psychological, and biographical entities that exist in the world. Prompting an AI to roleplay as an embodied human supposes the existence of a little proto-human in the model that does not exist.

  • Contemporary LLMs are constrained by their training data that is often limited, systematically biased, and impoverished of the relevant information we want to represent in our models. (Reading about consumer behaviour does not make one a consumer).

  • AI alignment research shows that the values and behaviours of frontier models are set intentionally by AI engineers and should not be expected to be a neutral and unbiased reflection of the training data (even if that data were neutral and unbiased).

As a result, we’ve the impacts of this fallacy borne out empirically. Attempts at replicating positive digital twin experiments are met with little to no success. Correlations with real human behaviour are very weak. Digital twins overwhelmingly look like highly educated, high-income Americans - the exact group of people who built them. Responses are highly sensitive to variations in input text and model set-up. Results are not even directionally accurate a third of the time.

This is frustrating because the academic story and the commercial enthusiasm for digital twins appear to be telling two different stories. Simile is a typical example. Joon Sung Park and colleagues' Stanford paper was an influential first step. But future contributions from Park will obviously be constrained by Simile's investors and the need to protect proprietary IP. The founding of Aaru tells a similar story with a famous EY Wealth Management case study showing the success of their approach - and the remainder of their efforts remaining stubbornly hidden from view. The pattern on the commercial front is consistent: one or two powerful case studies for the pitch deck, and everything else locked behind IP protections. I'm not claiming nefarious intent. This is just the reality of capitalism. But it does create a situation where positive evidence stays hidden while negative evidence continues to pile up.

And because of this information asymmetry, I don’t want to make any specific accusations about any specific vendors. I don’t know how Simile and Aaru or any other digital twin vendors manage the homunculus fallacy. I only mention them for their profile and impressive valuations. However, as a buyer, understanding this fallacy should help you vet potential vendors. Here are some approaches I would suggest:

  • Understand if AI is even used in this digital twin product. Digital twins pre-date AI. They are not synonymous with it. It might not be an issue at all.

  • If AI is used, ask what role it is playing. And be specific. A watchout is if the AI is being instructed to roleplay as a person. We’ve seen time and again that this does not work and is the homunculus fallacy in a nutshell. However, AI could be used in other ways, such as research assistance to search for and reference information, which may be more defensible.

  • Get clarity on which models are used and their alignment training. Absolutely do not let vendors handwave this away. It is common for frontier labs to publish system cards with critical information about their model development. Even if the vendor claims to have trained their own model, they can (should) be able to provide something similar to you.

  • Ask for third-party evaluation. It’s very easy for a vendor to curate a single working case example. But if these models are as high-quality, reliable, and affordable as everyone claims, then having an independent third party validate the results should be trivial. I’m not aware of anyone doing this yet, but you won’t get it unless you ask.

Closing Thought

The little person inside the model does not exist. What exists is a mathematical function trained to produce text that humans find useful, polite, and occasionally compelling. This function has no preferences of its own, no demographic identity, no biographical history, and no place in any real population. When you ask it to act like a 45-year-old grocery shopper, you do not summon a 45-year-old grocery shopper. You summon the function, conditioned on the words “45-year-old grocery shopper,” and the function produces text. The text may sound persuasive, but it is not the report of a 45-year-old grocery shopper, and treating it as one is the crux of the homunculus fallacy. Being mindful of this error can help us be better buyers of new and emerging market intelligence services. The promise of digital twins and the other capabilities made possible by AI is exciting, but this is new territory, and it can be easy to get swept up in flawed thinking.

First published in The Knowledge Stack.

The notes, by email

Roughly monthly. No sequence.

More notes

All notes →
July 2026 · 14 min

Which Model Should Answer Your Survey?

July 2026 · 10 min

Roleplay or Reasoning?

June 2026 · 15 min

Experimentalist Innovation in Distributed Retail