The routine that lied to me

7 August 2026

Every Thursday morning a routine wakes up, reads the week’s AI news, checks the claims against primary sources and writes me a draft. Nobody starts it. It has been running for weeks.

It has also been wrong twice, in two different ways, and those are the interesting part.

The memory that could not exist

The routine needs to remember what it covered last week, or it pitches the same story three times running. The memory was already in my mailbox: its own previous drafts.

Except it kept reporting that it found none. Six drafts, plainly visible in Gmail, and the routine insisted there was nothing there.

The Gmail search tool excludes drafts by design. It says so in its own documentation. The routine was searching a mailbox that structurally could not contain its own memory, and it was reporting that honestly: no memory found.

What impressed me, on reflection, is what it did not do. It did not invent a plausible list of last week’s topics. It said the thing was empty. The fix was one line, using the tool that lists drafts instead of the one that searches threads.

The claim that was not true

The second failure was worse, because it was confident.

In one issue the routine reported that Anthropic’s own benchmarks showed coding performance peaking at medium reasoning effort. Specific, checkable, and the kind of counterintuitive detail that makes a newsletter feel sharp.

I opened the source. It showed the opposite.

The claim never went out, because the routine’s own instructions say to open at least one primary source for every item and confirm the specific number before writing. That step is slow and it is the only reason this is a story about a process working rather than a correction.

What this actually taught me

The instinct is to conclude that agents cannot be trusted with research. That is the wrong lesson, and it is expensive, because it throws away the part that works.

The right one is narrower. An agent deserves exactly as much trust as its work can be checked. The routine reads forty sources I would not read. It also produces a claim that sounds right and is not, roughly once a month. Both of those are true at the same time, and the second is only dangerous if nothing sits between it and a reader.

A failure caught before sending costs nothing. One that gets through costs the thing the newsletter is for.

So the routine drafts. It never sends. And every factual claim gets opened against its source before it reaches anyone, including the ones that look obviously fine. Especially those.


I write What Actually Works every week: one real task handed to AI, with the prompts, the cost, and where it broke.