Jason Duke, Founder, Kronaxis
Tag: Research
The default way to build a synthetic population is simple to describe. Take a language model, give it a persona, ask your question, and treat the answer as one draw from that person. Do it a thousand times and you have a thousand respondents. Aggregate, and you have a poll.
Through 2026 a run of independent research groups took that method apart, from several different angles, and reached the same verdict: it does not reproduce a real population. If the synthetic research you are buying runs one model call per respondent, this is about the tool you are using. We reached that same conclusion earlier, and it is the reason we never built our system that way.
The problem is real, and it is not shallow
This is not one team with a grudge. It is separate groups, in machine learning, in the physics of society, and in economics, arriving independently at the same wall.
The mechanism is now understood. In Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe (Jang, Lee and Kim, 2026), the authors show that an instruction tuned model does not sample from a distribution at all: it collapses to a single output. The same persona, asked the same question, returns the same answer on more than half the items in a public opinion benchmark. And the cause is the alignment training itself, so the very thing that makes a model a capable assistant is what breaks it as a population.
At the population level the picture is the same. In Distribution-First Population Simulation (Ozkan, 2026), thousands of independent agents grounded on real survey respondents pile onto a modal default and fail to reproduce the population's distribution, and the collapse is large and systematic. Two papers, from opposite directions, describing the same failure.
And scale will not quietly fix it
The comforting assumption is that this is a small model problem that the next generation grows out of. A Stanford group tested exactly that. In Will Scaling Improve Social Simulation with LLMs? (Ziems and colleagues, 2026), they find that scale helps in many settings, but not on the ones that matter most: underrepresented populations, forecasting over time, and calibration to how people actually behave. On those, even much larger and fine tuned models barely move. Waiting for a bigger model is not a plan.
Nor is this a story told by only three papers. Mirror (Aslam, 2026) finds that a compact statistical model can out predict interview grounded language model personas at the individual level, and that the genuine value of a model lies in understanding a question and carrying influence, not in being a person. Quantifying the Persona Effect (Hu and colleagues) had already shown, earlier, that a persona label moves a model's output only modestly and only in predictable cases. The field is converging, and it is converging on a hard truth.
We reached this conclusion first, and acted on it
Here is the part that matters most. We did not read these papers and change course. We built our system around exactly these failure modes before this literature existed, because we had already seen them for ourselves.
Our published work and our human study predate this run of papers. We agreed, early, with what the field is now formally proving: that a thin persona over an aligned model collapses, that aggregating independent draws does not recover a population, and that grounding is not optional but the whole game. Agreeing with the diagnosis is the easy part, and plenty of teams now nod along to it. The difference is that we agreed early enough to go and solve it, while a great deal of what is sold today as synthetic research is still built on the very approach these papers take apart.
We are not going to set out how we solved it. That is the asset, and we intend to keep it. What we will do is show you the result, because the result is the part a buyer can actually check.
The proof, on real humans
Our personas are not demographic labels with a prompt wrapped around them. Each carries a measured disposition, our DYNAMICS-8 profile, and is grounded so that it behaves like a specific person rather than defaulting to whatever a web corpus already contained. That is a claim, and claims are cheap in this field. So here is the test that is not.
In a choice experiment registered in advance, with 117 real people, matching the character of content to a reader's own disposition raised conversion from nineteen to thirty one per cent. That is a lift of more than sixty per cent, and the loud, obvious, easily copied signals carried none of it. It was the deeper half, the part a reader never consciously reports, that moved the decision. This was a randomised controlled trial on humans, its analysis fixed before a single response was collected, not a synthetic benchmark grading its own homework. It is direct evidence that grounded disposition changes what people do, not merely what they say.
Stated plainly
We do not sell magic. Synthetic populations are, in our view, tools for hypothesis generation and rapid screening: they narrow the field, surface the dynamics behind a preference, and let a team test twenty ideas in a morning. On the highest stakes decisions they belong alongside real people, not instead of them. Anyone promising a synthetic panel that perfectly replaces human research is selling you the brochure, not the science. What we promise is narrower and real: a grounded system, an honest account of what it is for, and evidence gathered the way evidence is supposed to be gathered.
The question to ask
This year the field spent a great deal of ink naming a problem. Naming it is now easy; every vendor can quote the same papers back to you. Solving it is the hard part, and having solved it early, before it was fashionable to talk about, is harder still.
So when you evaluate a synthetic research tool, do not ask whether it can describe the problem. Ask whether it saw the problem coming, and whether it can show you, on real people and with evidence collected the way science collects it, that its answer works. We can. That is the difference between a demonstration and a product.
See the evidence
Read the study, then build a grounded panel and judge the reasoning behind every response for yourself.
Get Your API Key