00

The badge on the boot

A 2.0 badge on the boot of a car used to give you a reasonable idea of what was under the bonnet. As manufacturers grew more creative with their numbering, choosing a car started to need rather more homework.

Hardware shops took the plainer route. For decades the big catalogues sold the same tool at three prices, labelled good, better and best, and most buyers correctly took the middle one.

AI companies have arrived at much the same place as the car makers. Open the model menu and you get names, numbers and various suggestions of speed or intelligence. You can usually tell which one the company considers its premium model. Nothing on the menu tells you whether you need it.

01

What you thought was happening

Faced with a menu you cannot read, the obvious rule is that the list is a ranking, and I assumed for a while that I should use the most capable model I could get. Then I tried Fable. In its early days I often struggled to follow the answers. Sometimes I could not work out whether it was talking nonsense or whether I was not clever enough to understand it. Neither was a particularly comfortable possibility.

I kept going back to Sonnet. I understood its answers and had come to trust it. Whatever Fable could do, Sonnet was more useful to me for that work.

02

Four axes

Strip the names off and there are four things I would look at.

The axisIn one line
What it can handleTry it on the kind of work you do
How recently it learnedEvery model stopped reading at a cutoff date
How long it can thinkThe rough book and the dial from parts two and three
What it costs to runMoney and seconds, and your own time

What it can handle. Start with the kind of work you want done. Rewording an email is a different task from finding a fault in software or comparing several complicated proposals, and a model that does one well may be less impressive at another. Try a few examples of your own work. Does it follow the instructions? Does it miss details? How much correcting do you have to do before the answer is useful?

Include something whose answer you already know. My preference for Sonnet’s writing told me something about readability. Checking its work tells me something about accuracy, which deserves a separate judgement.

How recently it learned. Every model stopped reading at a cutoff date, and the providers publish them. Anthropic’s comparison table lists a cutoff for each model next to its price, and on the day I checked, the cheapest and the dearest were sixteen months apart. Google’s model pages do the same. The date is no guarantee that the model knows every fact from before it. A brilliant model with an old cutoff is a brilliant colleague back from two years on a desert island.

An assistant can also search the web or read a document you give it, which takes it past the cutoff. If the subject has changed recently, ask it to look the answer up and show its sources, then open the important links and check the dates. For a summary of a document you supplied, the question is whether it understood the document, and the cutoff does not come into it.

How long it can think. Parts two and three opened the rough book and the dial that rations it. Where the dial is available, a difficult problem can be given more processing time, at the cost of a longer wait. That pays when a task has several interacting parts: a plan with conflicting constraints, or a calculation where one assumption affects everything after it.

The model and the effort setting both count, and a smaller tier on a long think beats a bigger one answering off the cuff more often than the menu order suggests. If the answer is missing information, supply it first. More thinking cannot reliably fill a gap in the facts.

What it costs to run. Money and seconds. The same Anthropic table puts the price next to the cutoff, and its dearest model costs ten times its cheapest for the same length of text. If you pay by subscription the money is invisible, but you still notice how long an answer takes and how much work it leaves you with.

A more demanding task may justify the wait. Compare how long it takes to get a usable result, including the corrections and retries.

The idea in one line

Models differ in what they can handle, how recently they learned, how long they can think and what they cost, so the question to ask is which one is enough for this job.

A card headed Choosing an AI model, listing four things to compare: what it can handle, how recently it learned, how long it can think, and what it costs to run, each with a one-sentence explanation.
Four things to compare before you pick a model from the menu.

None of the four comes with a number on the menu.

If the menu is this unreadable, why not let the shop choose? That was tried, at scale, in August 2025. OpenAI launched GPT-5 with an automatic router meant to retire the picker: the system would read each question and choose a model for you. On launch day the router broke. In the chief executive’s own words, “the autoswitcher was out of commission for a chunk of the day”, and the new system “seemed way dumber” because every question was being sent to the wrong kitchen. Users revolted, not least because the old models had been switched off overnight, and within days the company restored the previous menu for paying customers. Routers have improved since, but the lesson, I think, stands: the shopkeeper chooses by guessing your job from one sentence. You know the job.

03

What to do with it

Pick a model that works well for your usual tasks and make it your default. I keep routine drafting, summaries and everyday questions on a small or middle tier, and I cannot see enough improvement from the dearer option to justify waiting for it. You do not need to reconsider the whole menu every time you write an email.

When an answer disappoints, work out what went wrong. If the instructions were unclear, clarify them. If it used old information, ask for a search. If it lost track of the problem despite having what it needed, turn up the effort or move up a tier.

For work with consequences, start with a capable model and make time to check the result. Waiting for an obvious failure is a poor test when you might not recognise the mistake.

If your assistant offers to choose the model for you, let it, and tell it what the answer is for and how carefully it needs to work. Find out where the override lives for the times the result falls short. When you are unsure whether an upgrade is worth it, give both models the same task with the same instructions and material, something typical of your week, and compare the answers for accuracy, omissions and the editing they need.

04

Next on the table

Next week, the memory. You will learn what the machine keeps about you between conversations, where it keeps it, and who is allowed to read the book.