← All posts

The Long Idea: what we can prove, and the claim we cannot make

Almost everything we build is an answer to one problem, so start there. More and more of what you read was written by a machine, and the machines are built to sound right, not to be right. A language model is trained on a single reward: produce text a person would plausibly have written. It is superb at that. What it is never doing is checking whether the words are true, or fair, or trying to move you. It was not asked to.

So the world is filling with writing that is fluent by design and truthful only by accident, and the old instinct we all use, trusting writing that reads well, is now the wrong instinct, because reading well is the one thing the machine is guaranteed to do whether or not any of it is real.

Plausible is not the same as true. That sentence is the whole programme.

Once you take it seriously you cannot trust the smooth surface of a piece of writing. You have to measure the things the surface hides. Is there real substance here, or just performance? Is it true? Is it trying to move me? And if a machine acts on my behalf, can I prove it behaved? Each of those questions became an instrument, and the instruments turned out to share a shape. That shared shape is the long idea.

The instruments, briefly

Substance against performance. Two pieces of writing can say the same thing and land completely differently, and almost all of that difference comes down to two directions. Matter is how much substance is in the writing, rigorous and backed up, or thin. Manner is how it carries itself, plain and institutional, or warm and vivid. A dense legal opinion is high matter, low manner; brilliant clickbait is the reverse. A machine optimising for reads well is optimising for manner, so being able to read matter separately is being able to see past the very thing the machine is good at.

People have a character too, the settled kind psychology has measured for a century, and the striking early finding is that the character of a person and the character of writing are joined by a map. Not a crude one. A person's openness tends to surface as originality in what they write, their steadiness as substance, and the links run at an angle rather than straight across.

Truth as a lookup. For most of the practical web, truth is not an opinion to weigh but a fact to check. Whether a company exists, filed the accounts it claims, is who it says. So our truth engine does not ponder whether something feels true; it checks against the public record and grounds the verdict in what it retrieves.

Manipulation has a shape, certain moves in how a message is built to get past your judgement rather than through it, and because we can read those moves we can measure them. And provable compliance: if a machine acts for you in a regulated setting, we make each action carry its own checkable proof that it was allowed, rather than a promise you have to trust. The opposite of "trust the AI": here is the proof, check it yourself.

These look like separate tools. They are one idea seen from different sides: every one is a refusal to trust the smooth surface.

The surprising part: it keeps turning up

If the matter and manner structure were only true of, say, English social media, it would be a curiosity. It is not. The same two directions kept appearing across millions of websites in dozens of languages, in the speeches of many national parliaments, across centuries of persuasion, and in the slow drift of the web itself. Different substrates, different tongues, the same simple shape underneath. That is the long idea of the title: beneath a great many things that look separate, personality, persuasion, truth, attention, there may be one small shared structure. We are careful with "may", and the rest of this piece is why.

What the experiments have actually found

It is one thing for our measurements to resemble established theory. It is another to ask whether they carry information those theories missed, and whether the structure is something you can act on rather than only describe. So we froze the instruments, changed nothing, and put a series of hard questions to them. Here is the honest ledger.

It predicts, not just describes

We gave the frozen system something it was never built to predict, how good an argument is as judged by real people, on thousands of short arguments. The ordinary tools went first: length, reading difficulty, how emotional the language is, an established measure from linguistics. Then we added our character reading on top, and because it was frozen we could not tune it to win. It improved the prediction by a clear margin, and the improvement was information the ordinary tools did not hold. We then let a large modern machine reading of the same text compete; on its own it did better than our small map, no surprise, but when we added our frozen map on top the prediction improved again. The powerful reading could not recover what our small, readable map adds. The two see partly different things.

Measuring and controlling are not the same

A sceptic could say our scores are just what our own model thinks. So we edited texts on purpose, adding a citation to raise rigour, adding urgency to raise emotion, and had two unrelated models from different independent labs score them blind. Both saw exactly the changes we made. The qualities are real, not one model's private opinion. Then a harder test: asked to change one quality and leave the rest alone, a model could not do it cleanly. Changing one dragged others with it, in a structured, repeatable way. The lesson is that describing something and controlling it are different. Like a car with eight gauges but only three controls, character has about eight qualities you can measure but only two or three you can actually move by rewriting, and that small control structure holds up when you change the generator, the language, the tool doing the rewriting and the tool doing the reading.

Does it read a person the way people do?

The sharpest test of an instrument is whether it recovers judgements independent humans already made, in public research where people labelled a text long before we existed. We locked our pass marks in advance, deliberately demanding, and checked. The result deserves to be stated in two halves that must not be blurred.

Evidence positive

Our reading of rigour tracks human judgements of argument quality; our reading of emotion recovers human emotion labels across two separate corpora and in the right direction; the shape of our emotion space lines up with the human one far above what chance allows, about a hundred and forty six times over. There is even a small but real behavioural signal: inside randomised tests where a publisher showed different headlines to comparable audiences, our attention reading predicts which one people actually clicked. These are inconsistent with the instruments being an arbitrary code that models merely agree on among themselves.

Preregistered gate not passed

Every one of those effects is moderate to weak, and none clears the strict bar we set ourselves in advance. We did not move the bar to meet them. That is the honest state: real, consistent, human grounded, and short of the demanding threshold. One correspondence failed outright, and we banked the failure rather than quietly dropping it.

Can we move a real person? Not shown.

This is the line we will not soften. There is already good evidence that the properties of a message affect what a crowd does on average, which is why the headline result holds that manner earns attention and matter earns conviction. But that is a statement about averages, not about steering an individual. The cleanest causal test we have run in real people, with randomised messages, a frozen prediction made in advance, and a working control to prove the pipeline could detect a signal, returned a genuine null: our geometry added essentially nothing to predicting how a person's mood then moved. That is negative evidence for the strong claim, not weakly encouraging evidence, and we record it as such.

Three questions, kept apart

Holding those results apart matters, because three very different claims are easy to let bleed into one.

One, can these fuzzy things be made measurable? Increasingly yes. Emotion, disposition and the character of writing can be represented as positions in measurable spaces, and the relations between those spaces can themselves be measured. This makes them operational, not physical. This is not the reading of minds. It is estimating a limited set of defined variables from observable evidence, imperfectly and with measurable error. We treat psychological states as real variables that can be located, related and perturbed, and we claim nothing more mystical than that. Somebody can accept every qualification in this piece and still misread the phrase "read a person"; the plain answer is that we estimate defined quantities, we do not see inside anyone.

Two, can we deliberately change those represented states, on the machine and text side, and predict the knock on effects? Substantially yes. Push one axis in generated text and the others move in sparse, signed, predictable ways, and that control structure travels across generators, languages and tools. This is the beginning of a control theory over text, not merely a good classifier.

Three, can we measure a real person and then move them from where they are toward a chosen state, in a way tailored to them? Not shown. One meaningful null, one unusable dataset.

We have built much of the instrumentation a science of human state and its movement would need. We have not shown that such a transition rule exists in people at all.

The programme is strong at both ends, measurement in front, verification and proof behind, and the gap is in the middle: the causal law connecting a measured person, a controllable message and a moved outcome. Having all the parts is not the same as having that law. The parts do not need wiring together; they need the law, and the failures we have run did the useful thing, they told us exactly where it is missing and eliminated the weak ways of looking for it.

The experiment that would put the central claim to a real test

So we have named the test that could prove the deep claim, and could just as easily sink it. Take a single message with fixed facts. Rewrite it to sit at chosen places on the map, its substance and its manner up or down, keeping the facts, the length and the author the same. For a particular person we have already measured, write down beforehand exactly how they should respond, and how a different person should respond differently. Then show it to real people and see. If moving the message on the map moves people the way we predicted in advance, and moves different people differently as their own measurement said it should, the map is part of the machinery of how words work on people. If the message moves and the people do not, the claim fails and we keep an excellent description rather than a mechanism. Either way, we learn something true. Move the map, predict the person.

One thing that experiment would not do, even if it succeeded, is turn itself into a product. Showing that a change can be predicted and deciding that it should be deployed are separate questions, and we treat them as separate. The reason verification, policy and provable restriction sit in the same architecture as the measurement is precisely this: a capability to move a state, if it ever existed, would be unacceptable without the constraint that proves each use stayed within the rules. We are building the accountability half deliberately alongside the measurement half, not as an afterthought to it.

Why we tell it this way

Because the honesty is the product. Anyone can publish a confident story; the machines are built to. What is scarce is a reading you can check, a claim with a stated boundary, and a failure reported in full. That is the posture the whole programme takes, and it is the same posture we bring to the work we do for clients: measure what actually changed, show where the evidence stops, and prove it rather than assert it.

Read it, cite it, argue with it

If measuring what is real rather than what is plausible is the standard your work needs, we should talk.

Get in touch