Jason Duke, Founder, Kronaxis
Tag: Research
Here is the most dangerous result a synthetic panel can produce: the right answer. Not because a right answer is bad, but because a single headline number that matches reality is the easiest thing in the world to hit by accident, and it tells you almost nothing about whether the panel underneath it is real. Two careful benchmarks from 2026 make this concrete, and they should change how anyone buys or sells synthetic research.
A panel that passed the test, for the wrong reason
A group led by Yuxuan Cai and Yequan Hu took a real behavioural experiment, 843 people judging an affordable housing development sited near their own homes, and asked eight language models to reproduce it. Support fell as the project moved closer, and fell more sharply for owners than renters. That owner versus renter gap was the headline estimand, and they set a fair equivalence bar around it in advance.
One model, and only one, passed. And then they looked underneath the pass. It cleared the bar only because its errors cancelled: it matched Democrats, badly understated Republicans, and overstated Independents, and those opposite mistakes averaged out to something that looked correct. It reproduced roughly a tenth of the within group variation real residents showed, so every subgroup came out far more uniform than real people are. And when the researchers simply reordered the questions, the model's headline estimate shifted by 0.367, against a human shift near zero. A real population does not change its mind because you asked the questions in a different order.
Agreement on one aggregate estimand can coexist with substantial population mismatch. Cai, Hu and colleagues, 2026.
Read that twice, because it is the whole point. The average matched. Everything that generates the average did not.
The same wall, from a different direction
The other benchmark, from Zihan Chen, Di Zhu and Lei Zheng, comes at it across two large surveys, the US General Social Survey and the World Values Survey, and four models including a frontier one. Their verdict on the individual level is blunt: a demographic prompted model never beats a plain demographic lookup table, and on cross cultural values every model they tested fell well below it. Where it really breaks is the same place: the subgroup. The models bind who you are to what you think far more tightly than reality does, and the cost lands on the exact question a research buyer is paying for. Asked which group holds the most extreme view, the models picked the wrong group half the time on US questions and 72 per cent of the time across cultures. And, usefully for anyone still waiting on the next model to fix this, the larger models stereotyped more, not less.
Why this is a buyer's problem, not a lab curiosity
Almost nobody buys synthetic research to learn the national average. They buy it to learn which segment differs, where the resistance sits, which group to design for. That is the subgroup layer, and the subgroup layer is precisely where both papers show the failure concentrates. A vendor can show you a chart where the synthetic total lines up beautifully with a real benchmark, and that chart can be true, and the segment level story you actually act on can still be invented. A matching average is not evidence that the parts are right. It can be the accident that hides the parts being wrong.
So the test that matters is not "does the aggregate match". It is "does it hold at the subgroup, does it carry the real spread within each group, and does it stay put when you perturb the prompt". We wrote about the wider version of this in Scale Will Not Fix Synthetic Research, and about the single call version of the same collapse in the KNOWS/DOES split.
Being fair to the evidence
Two honest notes, because overstating this would be its own kind of dishonesty. First, neither paper says synthetic panels are useless. Both are explicit that the population average is often reproduced well. The failure is at the individual and subgroup levels, which is where the value is, but it is not "everything is broken". Second, neither paper tested grounding as a fix. They evaluated thin personas, a party label and a housing status, and they document the problem rather than solve it. So no honest reading of them lets us claim they prove any particular remedy works. They prove the failure. The remedy has to be shown separately, on real people.
What we do about it, and how we prove it
Our answer starts from the assumption these papers establish. We ground each persona in a measured disposition, our DYNAMICS-8 profile, and in cultural priors drawn from established social science, so that the persona carries real signal rather than a demographic caricature. But the part that matters here is not the grounding, it is the standard of proof. We do not validate a panel by showing you a matching total. We test it against real people in a trial whose analysis is fixed in advance, so a passing aggregate cannot be mistaken for a working panel, and so a lucky average cannot be dressed up as a result. We set out why that discipline is the real bar in what preregistration actually proves.
So the next time a synthetic panel shows you a perfect headline number, ask the quiet question these two papers were built to answer: does it survive at the subgroup, or is this the average that lied?
See the evidence
Read the study, then build a grounded panel and judge the reasoning behind every response for yourself.
Get Your API Key