Jason Duke, Founder, Kronaxis
Tag: Research
Something changed in synthetic research this year. A run of independent papers named, in careful detail, why the standard approach does not reproduce a real population. That was good for the field. It was also, quietly, good for anyone selling a weak product, because now every vendor can quote the same papers back to you and sound as though they have solved the very thing the papers describe.
Naming a problem and solving it are not the same act, and the gap between them is exactly where a buyer gets hurt. So here are five questions to put to any synthetic research tool, ours included. Each one is drawn straight from the 2026 literature, and each is hard to answer well unless the work is genuinely done. We tell you what a good answer sounds like, and at the end, plainly, why we can give all five. If you would like the background first, start with why most synthetic research does not reproduce real people.
1. Does it run one model call per respondent?
This is the first question because it is the most revealing. The default design gives a model a persona, asks the question once, and treats that single answer as one person's view. Jang, Lee and Kim (2026) showed that an aligned model does not sample from a distribution when you do this; it collapses to one answer, returning the same response on more than half of a public opinion benchmark. Ozkan (2026) showed the same collapse at the level of a whole population. If a vendor's method is one call per respondent, aggregated, you are buying the exact design these papers take apart. A good answer describes how the tool recovers a real distribution rather than stacking up identical draws.
2. Is the persona grounded, or is it a label over a web corpus?
A persona that is a few demographic words in front of a general model tends to answer as the internet's average of those words, not as a person. Hu and colleagues found that a persona label moves a model's output only modestly, and only in predictable cases. So ask what the grounding actually is. A good answer points to something specific and external: measured disposition, established social science, a real basis for why this persona differs from the default. A weak answer is a longer list of adjectives.
3. Does the pitch depend on the next model being bigger?
If a tool is thin today but the vendor promises scale will fix it, check that promise against the one group that measured it. Ziems and colleagues at Stanford (2026) found that scale does not close the gaps that matter: underrepresented populations, forecasting over time, and calibration to real behaviour. A good answer does not lean on a future model at all. It explains what the tool does now that does not depend on scale, because those gaps are about grounding, not size. We wrote about this in Scale Will Not Fix Synthetic Research.
4. Can it show behaviour, not just agreement?
It is easy to build a persona that nods along to a survey question. It is much harder to build one whose disposition changes what it would actually do. Aslam (2026) found that a compact statistical model can out predict interview grounded language personas at the individual level, which is a warning about where the real value sits. Ask whether the tool has been tested on a behavioural outcome, a choice or an action, not only on stated attitudes. A good answer distinguishes what a persona says from what it does, and can show the second.
5. Is there evidence on real humans, gathered the way science gathers it?
This is the one that ends most conversations. A synthetic benchmark that grades its own homework is not evidence; it is the same model marking itself. Ask whether the tool has ever been checked against real people, in a study designed before the data was collected, so the result could have come out against the vendor and did not. A good answer is a preregistered trial on humans with a clear, reported effect. A weak answer is a demo, a case study written after the fact, or a chart with no way to have failed.
Why we can answer all five
We did not assemble these questions after reading the papers. We built our system around these failure modes before this literature existed, because we had already run into them ourselves. We do not run one collapsing call per respondent. Our personas are grounded on measured disposition, our DYNAMICS-8 profile, and on established social science rather than on whatever a web corpus contained. We do not need the next model to be larger. And we have the fifth answer, the one that matters most.
In a choice experiment registered in advance, with 117 real people, matching the character of content to a reader's own disposition raised conversion from nineteen to thirty one per cent, a lift of more than sixty per cent, carried by the deeper part of disposition that no reader consciously reports. The analysis was fixed before a single response was collected. It is a randomised controlled trial on humans, not a synthetic score marking itself.
We are not going to set out how we solved the problem behind these questions. That is the asset, and we intend to keep it. But the questions are yours to use on anyone, and the fifth answer is the one to hold out for. Ask for evidence on real people, collected the way science collects it. If a vendor cannot give it, the rest of the conversation is a brochure.
See the evidence
Read the study, then build a grounded panel and judge the reasoning behind every response for yourself.
Get Your API Key