00

Day nine of the trial

Picture a courtroom on day nine of a long trial. At the front sits the stenographer, whose record is perfect: ask for any sentence from day one and it can be produced, word for word. In the box sits the jury, whose record is nothing of the kind. Ask a juror about day one and you get the gist, plus the occasional confident error. They remember this morning’s testimony well.

The jurors are the same twelve people they were on day one, and they are far less reliable.

When the trial ends, the court keeps the transcript.

Your longest chat with an AI is day nine of that trial.

01

What you thought was happening

The context window is sold by size, hundreds of pages, whole books, so the natural picture is a warehouse: whatever you paste is stored, and stored means attended to, so more context can only help. I pasted whole reports to be safe, and I kept one mega-thread running for months because the machine had it all. I have watched what happens to threads like that. Somewhere past the middle the answers start to drift. The machine contradicts something it said an hour earlier, forgets a rule I gave it at the start, and gets stranger with every reply while still sounding perfectly confident.

It reminded me of a school chess tournament in grade eleven where someone sneaked in a flask of vodka. The strongest player in the room got sillier with every round.

I also assumed that closing the tab ended it. The conversation was over, so it was gone. In fact the machine keeps every word and reads the middle badly, and the provider keeps the chat after you close it.

02

The jury and the stenographer

Everything from the earlier parts ends up in the window. Your words arrive there as tokens and the model’s rough working is written there. The memory file, the skill and every tool result an agent collects are pasted in too. They all share one space.

The idea in one line

Eleven of thirteen models fell below half their short-context performance by 32,000 tokens, a small fraction of the window they advertise.

Windows are vast now, and the storage half of the courtroom holds up. The stenographer has every word. What fails is the jury. In July 2025 the research team at Chroma tested eighteen leading models and found performance degrading as input grew, even on trivially simple tasks. They called it “context rot”. The NoLiMa benchmark found that when the question does not share words with the passage that answers it, eleven of the thirteen models tested fell below half their short-context performance by 32,000 tokens, which is roughly fifty pages and a small fraction of their advertised windows. And the earliest of these results, from 2023, mapped where attention goes: models retrieve best from the start and end of the window, and worst from the middle, like a juror who recalls opening statements and this morning.

I suspect the drift I saw in my long threads was mostly this, the middle of the trial going out of focus, though I never measured it and the vodka comparison is mine rather than the researchers’.

One long chat drawn as three zones: the start and end filled in and labelled read closely, the wider middle left empty and labelled most often missed.
Models retrieve best from the start and end of the window, and worst from the middle.

Andrej Karpathy argued in 2025 that the craft worth naming is “context engineering”: deciding what deserves a place in the window, rather than phrasing the question cleverly. Every technique in this season is that craft from a different angle. The skill from part six stays folded until it is needed. The memory file from part five carries three lines instead of thirty conversations. Both keep the window from filling with things the task does not need.

Closing the tab clears your screen and leaves the record where it was. The transcript stays on the provider’s servers, governed by retention periods and training policies that differ sharply between personal accounts and business accounts, and that move too often to print here. So pasting the whole report to be safe makes the answers worse and also puts more of your material on the provider’s servers. The settings page decides what the provider keeps, and this site’s guide to the training switches walks through them. For words that should never reach a provider at all, the answer has not changed since part five: run a model on your own machine.

03

What to do with it

Start fresh ruthlessly: one task per chat. I gave up the mega-thread after the drift got too obvious to ignore, and I now open a new chat the moment the machine contradicts itself or loses the thread of what I asked for.

Put the key document where the model reads best, at the start or the end of what you paste and never in the middle, and restate the question at the end. Paste the ten pages the task needs, not the report they came from (I still catch myself pasting the report).

Once this week, open the settings page and read what your provider keeps: retention, training, and the business terms if you use a work account. Closing the tab does not touch any of it.

04

The table is cleared

The machine still works exactly as it did eight weeks ago, and you now know what the parts are called. It reads tokens, and it drafts in a rough book before it answers. You turn the effort dial when the problem is hard, and you can read the model picker as a spec sheet rather than a leaderboard. The memory file is yours to open and correct, a skill is a file you can hand to someone else, and an agent gets only the permissions you grant it. All of it rests on the window. The whole season stays on the shelf at Anatomy of an Answer, and the field guides pick up where the essays leave off. Pass the parts on.