Jason Duke, Founder, Kronaxis
Tag: Research
There is a bet buried inside most synthetic research, and it is rarely said out loud. When a synthetic panel gets a population wrong today, the answer on offer is patience. The next model will be larger, so the gap will close on its own. Buy now, and the tool you own will quietly get better while you sleep.
It is a comfortable bet, and through 2026 it was tested directly. It lost on exactly the questions that make synthetic research worth doing. This is the companion to our piece on why most synthetic research does not reproduce real people, and it answers the obvious follow up: if the method is broken, does the next generation of model simply fix it? No. And it is worth being precise about why not.
The bet the industry is making
The bet rests on a real pattern. Larger models are better at almost everything, and for a decade the safe prediction has been that whatever a model does badly this year it will do adequately next year. Applied to synthetic populations, the reasoning runs like this: today's model collapses a persona to a single modal answer, so tomorrow's, being larger, will collapse less, and eventually the collapse will be small enough to ignore. Scale as a cure. It is the assumption that lets a vendor sell you a method that does not yet work, on the promise that it soon will.
A Stanford group tested the bet, and it lost
In Will Scaling Improve Social Simulation with LLMs? (Ziems and colleagues, 2026), a Stanford group did the unglamorous work of measuring it. They scaled models up and asked where the improvement actually landed. In many easy settings it did land: bigger models were better. But on three things it barely moved, and those three things are the whole reason a serious buyer commissions social research in the first place.
The first is underrepresented populations, the very groups a headline sample tends to miss and a client most needs to understand. The second is forecasting over time, saying not what people think today but what they will do next. The third is calibration to how people actually behave, as opposed to what they will tell a survey. On all three, scale stalled, and fine tuning did not rescue it. The gaps that matter are not the gaps that scale closes.
Why the open gaps are the ones that matter
This is not bad luck, and it is worth seeing why. Scale improves what is well represented in the training data, because more parameters and more text let a model absorb more of what is common. The minority voice, the shift that has not happened yet, the quiet gap between stated attitude and real behaviour: these are precisely the things that are rare, or absent, or actively misreported in the text a model learns from. No amount of the same kind of data teaches them. A larger model trained the same way is a sharper version of the same blind spot. The failure is not in the size of the model. It is in the assumption that a general web trained model, scaled, becomes a specific person.
We never took that bet
We did not build a system and hope the next model would fix its weakest part. We started from the assumption these papers have now confirmed: that scale would not close these gaps, because the gaps are about grounding and disposition, not raw capability. So we grounded. Every persona we build carries a measured disposition, our DYNAMICS-8 profile, and cultural grounding drawn from established social science rather than from whatever a web corpus happened to contain. That work does not get cheaper or more optional as models grow. It is the part that scale was never going to do for us, and we did it deliberately, early, and by hand.
We are not going to set out how we did it. That is the asset, and we intend to keep it. What we can do is show the result, because the result is what a buyer can check without taking our word for anything.
The proof, on real humans
In a choice experiment registered in advance, with 117 real people, matching the character of content to a reader's own disposition raised conversion from nineteen to thirty one per cent. That is a lift of more than sixty per cent, and it came from the deeper half of disposition, the part a reader never consciously reports, not from the loud signals anyone can copy. The analysis was fixed before a single response was collected. This is a randomised controlled trial on humans, not a synthetic benchmark marking its own work. It is evidence that grounded disposition changes what people do, and it did not require a larger model to produce. It required the right model of a person.
What to do instead of waiting
If a synthetic research tool is weak today on minorities, on forecasting, or on real behaviour, do not accept the next model as the plan. The one paper that measured it says the next model will not fix those. Ask instead what the tool does that does not depend on scale at all: how it grounds a persona, whether its disposition is measured or merely labelled, and whether it can show you, on real people, that its answer moves behaviour. Waiting is not a strategy. Grounding is, and it was available all along to anyone willing to do the harder thing first.
See the evidence
Read the study, then build a grounded panel and judge the reasoning behind every response for yourself.
Get Your API Key