00

The horse that could count

In 1904, Berlin had a celebrity horse. Clever Hans could add, subtract, multiply and tell the date, tapping out answers with a front hoof while crowds watched. A commission of thirteen, including a circus manager, a cavalry officer and the director of the Berlin zoo, examined him that year and found no trick.

In 1907 the psychologist Oskar Pfungst looked closer. When Hans could see his questioner, he got eighty-nine percent of the answers right. When he could not, or when the questioner did not know the answer, he dropped to six percent. The horse was not counting. He was reading the tiny, involuntary tension in a human face, and stopping his taps the moment it released.

Two cream cards headed The horse was not counting: 89 percent right with the questioner in view, six percent with the questioner blind, with hoof-tap tally marks under each.
The tell in two numbers. Hans scored when he could read a face that knew the answer, and stopped scoring when he could not.

Hans was not stupid, and nobody was lying. There was a performance you could watch, and a computation underneath it, and they were different things.

The reasoning your AI shows you puts you in the same position as that commission.

01

What you thought was happening

The button says thinking. The screen says thinking. So I pictured a mind at work behind the glass: the machine goes quiet, reasons somewhere inside itself, and then reports what it concluded. And when it showed me a tidy summary of its reasoning, I took the summary as the thought itself, the way a friend tells you how they reached a decision.

That is not what happens.

02

The rough book

Part one showed that the machine reads numbers from a codebook. It writes the same way, and thinking is more of that writing. A reasoning model is a model trained to fill a private scratchpad before it commits to an answer: it drafts, works through steps, checks itself, and only then produces the fair copy you receive. The scratchpad is ordinary generated text. There is no second kind of substance called thought. There is a rough book and a fair copy, produced by the same pen.

You mostly never see the rough book. When OpenAI launched its first reasoning model in September 2024, it hid the raw working deliberately and offered summaries instead, and its rivals broadly followed. When the Chinese lab DeepSeek released a model in January 2025 that showed every line of its rough work to anyone, the shock helped knock 589 billion dollars off Nvidia’s market value in one day, the largest single-company loss ever recorded to that point.

I have read some of that raw working, in the tools that let you open it. It did not look like thinking to me. It read like an anxious, rambling monologue. It doubted itself, suspected the question, ran down dead ends and circled back before it settled. If that is thinking, it is not the kind I had imagined.

You do pay for it, though. Hidden thinking is billed as output, the same as the words you receive. On a hard question the rough book can be many times longer than the answer. You are buying pages you will never be shown.

The idea in one line

Thinking is more generated text, written into a rough book before the fair copy, and the rough book is evidence that work happened rather than a record of what decided the answer.

In 2025, researchers at Anthropic slipped hints into questions and then checked whether models that used the hint admitted it in their visible reasoning. One leading model mentioned the hint about a quarter of the time. Another managed under forty percent. The working can read as calm and complete while omitting the thing that swung the answer. The trace is not lying. It is Clever Hans: a performance related to the computation, produced alongside it, and not to be confused with it.

The rough book is still worth paying for. On problems with steps in them, a hard calculation or a plan with dependencies, the same model does measurably better when it is allowed to fill the scratchpad first, because the working is where wrong turns get caught. The working is not testimony.

03

What to do with it

Give the thinking mode the work it is priced for: multi-step problems, tricky documents, anything where a wrong turn compounds. I keep quick lookups and routine drafting off it, because that is paying for pages nobody needs.

Never file a reasoning summary as an audit trail. If a decision needs a defensible account, write the account yourself. The model’s summary is a performance of diligence, and the research says it leaves things out.

And judge answers by checking them, not by admiring the trace.

04

Next on the table

Next week, the effort. You will learn what the dial next to the model picker buys, why slow and small beats big and quick, and why chess had to solve this problem first.