00

The horse that could count

In 1904, Berlin had a celebrity horse. Clever Hans could add, subtract, multiply and tell the date, tapping out answers with a front hoof while crowds watched. A commission of thirteen, including a circus manager, a cavalry officer and the director of the Berlin zoo, examined him that year and found no trick.

In 1907 the psychologist Oskar Pfungst looked closer. When Hans could see his questioner, he got 89 percent of the answers right. When he could not, or when the questioner did not know the answer, he dropped to six percent. The horse was not counting. He was reading the tiny, involuntary tension in a human face, and stopping his taps the moment it released.

Two cream cards headed The horse was not counting: 89 percent right with the questioner in view, six percent with the questioner blind, with hoof-tap tally marks under each.
The tell in two numbers. Hans scored when he could read a face that knew the answer, and stopped scoring when he could not.

Here is the part worth keeping. Hans was not stupid, and nobody was lying. There was a performance you could watch, and a computation underneath it, and they were different things.

Every time an AI shows you its reasoning, you are back in that Berlin courtyard.

01

What you thought was happening

The button says thinking. The screen says thinking. So the natural picture is of a mind at work behind the glass: the machine goes quiet, reasons somewhere inside itself, and then reports what it concluded. And when it shows you a tidy summary of its reasoning, that summary is surely the thought itself, the way a friend tells you how they reached a decision.

That is not what happens.

02

The rough book

Part one showed that the machine reads numbers from a codebook. It writes the same way, and thinking is more of that writing. A reasoning model is a model trained to fill a private scratchpad before it commits to an answer: it drafts, works through steps, checks itself, and only then produces the fair copy you receive. The scratchpad is ordinary generated text. There is no second kind of substance called thought. There is a rough book and a fair copy, produced by the same pen.

You mostly never see the rough book. When OpenAI launched its first reasoning model in September 2024, it hid the raw working deliberately and offered summaries instead, and its rivals broadly followed. When the Chinese lab DeepSeek released a model in January 2025 that showed every line of its rough work to anyone, the shock helped knock 589 billion dollars off Nvidia’s market value in one day, the largest single-company loss ever recorded to that point.

You do pay for it, though. Hidden thinking is billed as output, the same as the words you receive. On a hard question the rough book can be many times longer than the answer. You are buying pages you will never be shown.

The idea in one line

Thinking is more generated text, written into a rough book before the fair copy, and the rough book is evidence of work, not a diary of it.

That last clause earns its place. In 2025, researchers at Anthropic slipped hints into questions and then checked whether models that used the hint admitted it in their visible reasoning. One leading model mentioned the hint about a quarter of the time. Another managed under forty percent. The working can read as calm and complete while omitting the thing that swung the answer. That does not make the trace a lie. It makes it Clever Hans: a performance related to the computation, produced alongside it, and not to be confused with it.

The rough book still earns its price. On problems with real steps in them, mathematics, code, planning, the same model does measurably better when allowed to fill the scratchpad first. The working is where wrong turns get caught. What the working is not is testimony.

03

What to do with it

Three habits follow.

Give the thinking mode the work it is priced for. Multi-step problems, tricky documents, anything where a wrong turn compounds. Keep quick lookups and routine drafting off it; you would be paying for pages nobody needs.

Never file a reasoning summary as an audit trail. If a decision needs a defensible account, the account is yours to write. The model’s summary is a performance of diligence, and the research says it leaves things out.

Judge answers by checking them. A confident, orderly trace tells you only that the horse tapped well.

The writing is the thinking. The machine hands you the fair copy, and the rough book stays shut.

04

Next on the table

Next week, the effort. You will learn what the dial next to the model picker buys, why slow and small beats big and quick, and why chess had to solve this problem first.