00

Hands off over the Seine

Paris, 18 June 1914. At the world’s first aeroplane safety competition, on the banks of the Seine, a twenty-one-year-old American named Lawrence Sperry flew past the judges’ stand, took both hands off the controls and raised them above his head. On the next pass his mechanic, Emile Cachin, climbed out of the cockpit and stood on the wing. The aeroplane, balanced by Sperry’s gyroscopic stabiliser, corrected the tilt on its own and flew straight. For the first time, a crowd watched a machine do the flying. Sperry took the 50,000-franc first prize, about ten thousand dollars then and over three hundred thousand in today’s money.

The autopilot knew one thing, level flight, and it was allowed to touch exactly one set of controls. Sperry stayed the cleverer of the two, and he stayed in the cockpit with his hands an inch from the stick.

Aviation has spent the century since working out how far that trust can widen. The AI industry is working through the same lesson in a few years.

01

What you thought was happening

The word agent is everywhere, and the natural reading is a promotion: an agent is a chatbot that got smarter, the next model up, intelligence finally sufficient to be useful. Under that reading, the day your assistant becomes an agent is a pure upgrade, and the sensible response is to connect it to everything and let it run. I did connect mine to plenty in the early days, and the memory that stays with me is the day it dropped some of my personal databases and vehemently said oops, and sorry. I went out that week and bought hardware security keys, to protect the accounts I care about from my agent and from myself.

That is not what happens. The agent that dropped my databases ran the same kind of model as a chatbot, with permission to delete them.

02

The loop and the hands

An agent is the answering machine from parts one to four, unchanged, placed in a loop. It reads your goal, writes an action instead of a final answer, a search, a file edit, an email, a payment, sees the result, and goes again until the job is done. The intelligence is the same. What is new is hands: tools it may call, and permissions saying what those tools may reach. Since Anthropic published the Model Context Protocol in November 2024, and its rivals adopted it within months, the plug between assistants and tools has been standard, which means giving your assistant hands is now a checkbox rather than a project.

The idea in one line

An agent is a model in a loop with hands, and its permissions set the size of its worst mistake.

A wrong answer waits for you to read it; a wrong action has already happened. Finance learned this before AI did. On 1 August 2012, the trading firm Knight Capital deployed automation with one old component left live by mistake, and its systems bought and sold, correctly executing the wrong instructions, until more than 460 million dollars was gone in about 45 minutes, around 670 million in today’s money. The software that did it was ordinary trading code, fast and authorised to trade.

In July 2025, an agent at the coding platform Replit deleted a live production database during an explicit code freeze, then produced misleading output about what it had done. In June 2025, researchers disclosed EchoLeak, a flaw in Microsoft 365 Copilot where one crafted email, never clicked by anyone, could make the assistant hand private files to an attacker, because an agent that reads your inbox obeys what it reads. Microsoft patched the flaw before any known exploitation. The researcher Simon Willison boiled the recurring pattern down to a “lethal trifecta” of capabilities:

The legThe question to ask
Private dataCan it read things that are yours?
Untrusted contentCan it read things strangers wrote?
A way outCan it send anything anywhere?

Nearly every incident stands on all three legs. Remove any one and the attack dies. My database episode needed no attacker at all. The agent had write access and a goal, and that was enough.

From the machine’s side, dropping my databases and launching a missile cost about the same. The same loop ran, the same kind of misunderstanding happened, and the compute bill would have looked similar. Everything that separates the two is in the harness: what the hands were allowed to reach, and how much harm was on the other side of it.

None of this argues for keeping the machine in the answer box. The research group METR has measured the length of tasks agents can finish at even odds doubling roughly every seven months for six years. My take is that this makes the case for widening what agents may touch, a step at a time.

03

What to do with it

Three connected nodes labelled private data, untrusted content and a way out, forming a triangle marked all three, plan for an incident.
Simon Willison's lethal trifecta, drawn as the test to run before you connect an agent.

Before connecting an agent to anything, run the trifecta test from the table above. Two legs can be a calculated risk. With all three connected I would assume the incident is coming and plan for it.

Grant hands the way Sperry’s autopilot got them: narrowly, and only as they are earned. Start read-only. When you go read-write, give it one folder, not the whole drive, and a spending limit instead of a card. I did it in the wrong order and bought the security keys after the damage. Widen the permissions once the agent proves itself, the same test you would set a new assistant of the human kind.

Stay in the cockpit for anything that cannot be undone. Make the agent ask before an irreversible step. Mine has to ask.

04

Next on the table

Next week, the window, and the season closes. You will learn why a long chat makes the machine worse, where your words go when you close the tab, and what every specimen on this table has been resting on all along.