Jason Duke, Founder, Kronaxis
Tag: Methods
The word doing the most work in synthetic research marketing this year is "validated". It appears on nearly every product page, and almost none of them mean the same thing by it. For most, validation is a chart the vendor produced after building the tool, showing the tool agreeing with something. That is not evidence. It is a demonstration, and the difference is the whole subject of this piece.
The problem with a benchmark that grades itself
Here is the trap that catches most synthetic validation. You build a system, you run it against a benchmark, you look at the results, and if they are not good enough you adjust the system and run it again. By the time you publish the number, the benchmark has quietly become the thing you optimised against. The model is, in effect, marking its own homework, and it has seen the answer key. A result reached this way cannot fail, because any run that failed was tuned away before anyone saw it. A test that cannot fail proves nothing.
This matters more for synthetic populations than almost anywhere else, because the output is plausible by construction. A language model will always produce a fluent, reasonable looking answer. Fluency is not accuracy, and a chart of the model agreeing with a benchmark it was tuned on measures neither.
What preregistration is, and why it is a high bar
Preregistration is a simple, unforgiving discipline borrowed from clinical science. Before you collect a single piece of data, you write down exactly what you will measure, how you will measure it, and what result would count as success or failure. You lodge that plan where it cannot be quietly changed later. Then you run the study.
The power of it is what it takes away. You can no longer move the goalposts after seeing the data, because the goalposts are already fixed and on the record. You cannot try twenty analyses and report the one that worked, because you named the analysis in advance. You cannot claim a prediction after the fact, because the prediction is timestamped before the outcome. Preregistration is how a result earns the right to be called evidence: it made a claim in public that could have come out against you, and it did not.
What our registered trial showed
We hold our own work to this bar, which is why we talk about one study more than any other. In a choice experiment registered in advance, with 117 real people, we tested whether matching the character of content to a reader's own disposition would change what they did. The plan was fixed before the first response arrived. Matching to disposition raised conversion from nineteen to thirty one per cent, a lift of more than sixty per cent. The signal did not come from the loud, obvious, easily copied cues. It came from the deeper half of disposition, the part a reader never consciously reports. We tell the fuller story in the half of your web page you cannot see is the half that sells.
The number matters, but the design matters more. Because the analysis was fixed beforehand, that result could have shown no effect, or a smaller one, and we would have had to report it. It showed a large one. That is the difference between a finding and a brochure.
Why this is the question that ends the conversation
This is the fifth of the five questions we suggest putting to any synthetic research vendor, and it is the one that settles things. Ask whether the tool has ever been checked against real people, in a study whose analysis was fixed before the data existed. A good answer is a preregistered trial with a clear, reported effect. A weak answer is a demo, a case study written after the fact, or a chart with no way to have failed. The gap between those two is the gap between science and salesmanship, and it is worth insisting on.
The honest limits
One trial is one trial. Preregistration makes a result trustworthy; it does not make it universal. Our study measured a specific behaviour, in a specific setting, and the honest reading is that it is strong evidence for a claim, not proof of every claim. That is exactly why we keep running the discipline rather than resting on one number, and why we would rather show you a single result that could have failed than a wall of charts that never could. Evidence that can fail and did not is the only kind worth trusting, ours included.
See the evidence
Read the study, then build a grounded panel and judge the reasoning behind every response for yourself.
Get Your API Key