← Back to Blog

The State of Synthetic Populations, 2026

Jason Duke, Founder, Kronaxis

Tag: Research

2026 is the year synthetic population research stopped selling the dream and started measuring the reality. For two years the pitch was simple: give a language model a persona, ask it questions, and read the answers as a crowd. This year a run of careful papers, from machine learning, from the study of society, and from economics, tested that pitch properly. This is a map of where the field now stands, what it agrees is broken, what it has stopped hoping scale will fix, and where the answer is actually being found.

What the field now agrees is broken

The central finding of the year is that running one independent model call per person, and aggregating, does not reproduce a real population. This is no longer a suspicion. Jang, Lee and Kim showed the mechanism at the level of a single model: an instruction tuned model does not sample from a distribution, it collapses to one answer, returning the same response on more than half of a public opinion benchmark, and the alignment training itself is the cause. Ozkan showed the same thing at the level of a whole population, where thousands of independent agents pile onto a modal default and the distribution never appears. Two results, from opposite ends of the problem, describing one wall. We wrote about what this means for buyers in why most synthetic research does not reproduce real people.

What the field has stopped hoping scale will fix

For a while the answer to any weakness was patience: the next model will be larger, so the gap will close. A Stanford group put that to the test. Ziems and colleagues found that scale helps in easy settings but stalls on the three that matter most: underrepresented populations, forecasting over time, and calibration to how people actually behave. Fine tuning did not rescue those either. The quiet consensus forming now is that these gaps are not about model size at all. They are about grounding. We took that argument apart in Scale Will Not Fix Synthetic Research.

What the field has learned personas can and cannot do

The year also brought a sharper sense of scope. Aslam found that a compact statistical model can out predict interview grounded language personas at the individual level, and located the real value of a language model in understanding a question and carrying influence rather than in being a person. Hu and colleagues had already shown that a persona label moves a model's output only modestly, and only in predictable cases. Read together, these draw a boundary: a persona is not a substitute for a person, and a thin persona is barely a nudge. The useful question is what grounding turns a label into.

Where the answer is being found

Put the year's findings side by side and they point one way. If independent draws collapse, you have to elicit the distribution rather than stack up single answers. If scale will not close the gaps on minorities and behaviour, you have to ground the persona in something real rather than wait for a bigger model. And if a bare label barely moves the needle, the grounding has to carry actual signal about the person. Collapse, scale, and scope are three descriptions of the same missing ingredient, and the ingredient is grounding.

This is the part we find worth saying plainly. We did not arrive here by reading this year's papers. We built our system around these three failures before the literature named them, because we had already run into all three ourselves. Our personas carry a measured disposition, our DYNAMICS-8 profile, and cultural grounding drawn from established social science rather than from whatever a web corpus happened to hold. We do not run one collapsing call per respondent, and we never bet on the next model to save the method.

The evidence that separates a position from a claim

Agreeing with the field's diagnosis is easy, and by now everyone does. The hard part is showing that your answer works on real people. In a choice experiment registered in advance, with 117 real people, matching the character of content to a reader's own disposition raised conversion from nineteen to thirty one per cent, a lift of more than sixty per cent, carried by the deeper part of disposition that no reader consciously reports. The analysis was fixed before a single response was collected. That is a randomised controlled trial on humans, not a synthetic benchmark grading its own homework, and it is the kind of evidence the field will increasingly be asked for.

We are not going to set out how we solved the underlying problem. That is the asset, and we intend to keep it. But the direction of the field is no longer in doubt, and the result is ours to show.

What to watch in the rest of 2026

Three things are worth watching. First, whether calibration becomes the standard demand: elicited distributions can be mis dispersed, and the teams that measure and correct that will pull ahead of the teams that do not. Second, whether the field moves from stated attitudes to behaviour as the test that counts, because that is where the softest claims fall apart. Third, whether buyers start asking vendors the one question that ends most conversations, which is whether there is evidence on real people, gathered before the data was collected. If you only take one habit into the second half of the year, we set it out in five questions to ask before you trust a synthetic panel.

See the evidence

Read the study, then build a grounded panel and judge the reasoning behind every response for yourself.

Get Your API Key