Jason Duke, Founder, Kronaxis
Tag: Research
Ask a language model to pick a number from one to a hundred. Do it a few hundred times. You will get 42 on roughly four goes in five. Ask it to flip a coin and it comes up heads about five times in six. This is not a party trick, it is the whole synthetic research industry's problem in one sentence, and a group at KAIST has just named it.
The name is the KNOWS/DOES split, and the paper is Jang, Lee and Kim (2026). Their finding is precise and slightly unsettling: a model can describe a distribution accurately, and in the same breath be completely unable to sample from it. It knows the shape. It cannot roll the dice.
What collapse looks like
Their headline result is stark. Asked to role play survey respondents, the model gave the identical answer on every single call for 57 per cent of persona and question pairs, and repeated itself more than nine times in ten for 80 per cent of them. On five synthetic target distributions it emitted a single answer on over 94 per cent of calls. The die does not come up random. It comes up 42.
This matters because most synthetic research quietly rests on the opposite assumption. The method the field calls silicon sampling treats each model call as an independent draw from a persona's response distribution: run the call many times, aggregate, read the spread as a crowd. The KAIST result says the draw does not exist. What you get back is not a distribution, it is the same modal answer wearing a hundred different name badges. We described the population level version of this failure in why most synthetic research does not reproduce real people.
The cause is the alignment, not the size
The tempting explanation is that the next, larger model will fix it. The paper closes that door. When they compared base models against their instruction tuned versions, every instruction tuned model was worse than its base on every task. On one family the error more than doubled, from 0.26 to 0.54. The descriptive ability was already sitting in the base model; instruction tuning damaged the sampling primitive while leaving the knowledge intact. And because one of the models they tested was fine tuned without reinforcement learning from human feedback and still failed, they attribute the collapse to instruction tuning broadly, not to any one training recipe.
They can even see it in the logits. The gap between the top two options runs to more than fourteen units of information on some targets, far wider than any temperature setting reachable through a normal interface could flatten. You cannot turn a dial and get your distribution back. This is the same wall we wrote about in Scale Will Not Fix Synthetic Research, now visible at the level of a single token.
Instruction tuned models do not sample from distributions, they collapse to a single output. The same model that cannot sample from a distribution can describe it accurately in a single call. Jang, Lee and Kim, 2026.
The honest part, because it cuts both ways
Here is where we have to be fair to the paper, and to ourselves. The KAIST group found a cheap fix. If you simply ask the model to describe the distribution in one call, as a set of probabilities, the error roughly halves. So for a population the model already knows well, you do not need a hundred collapsing role plays, you need one honest question. That is a genuinely useful trick, and it partly competes with the idea that you must build something elaborate to get a distribution out.
But read their own caveat closely. The describe trick has a ceiling, and they say so plainly: on populations outside the model's training distribution, the describer's accuracy degrades. The model can recite the shape of a crowd it has seen. It has far less to say about a crowd it has not, a specific region, a niche segment, a group its training data underrepresented, the very slices where a real research question usually lives. That ceiling is not a footnote to us. It is the whole reason grounding exists.
What we do about it
Our answer starts from the assumption the KAIST paper proves: do not stack up single answers hoping a distribution appears. We elicit the distribution rather than sample it by brute force, and then we ground each persona in something real, a measured disposition from our DYNAMICS-8 profile and cultural priors drawn from established social science, precisely so the model is not left guessing about the groups it never learned. Describe beats collapse on the easy, well known crowd. Grounding is what you reach for when the crowd is the one the model cannot recite from memory, which is most of the time in real work.
We did not arrive at this by reading the paper. We built the system around the collapse before the field had a name for it, because we had run into it ourselves. What we can now do is point at an independent group that measured the wall and marked exactly where it is.
The evidence that separates a position from a claim
Agreeing with a diagnosis is easy. Showing your answer works on real people is not. In a choice experiment registered in advance, with 117 real people, matching the character of content to a reader's own disposition raised conversion from nineteen to thirty one per cent, a lift of more than sixty per cent, carried by the deeper part of disposition that no reader consciously reports. The analysis was fixed before a single response was collected. That is a randomised controlled trial on humans, not a synthetic benchmark grading its own homework. We set out why that distinction matters in what preregistration actually proves.
So the next time a vendor shows you a synthetic crowd, ask the quiet question: is this a distribution, or is it 42 in a hundred costumes? The honest ones will know exactly what you mean.
See the evidence
Read the study, then build a grounded panel and judge the reasoning behind every response for yourself.
Get Your API Key