Somebody on your team is fixing the same AI mistake every week

Edition 200 - Here's how to make the correction stick.

Hereโ€™s what weโ€™re reading and thinking about this week:

Anthropic just published what happened inside their company once AI was writing most of their code. They got about 8 times more of it done per quarter than they used to.

So what broke?

Not the writing. The checking.

The automatic checks that run before any of it goes out the door grew 25 times in six months. The team barely grew at all. They tried to fix the pileup three times, and each fix held for less time than the one before it.

Now, most people will file that under engineering news and move on.

But take the code out of it and it's a story a lot of people are living right now. The accountant using AI to categorize a month of transactions still checks them before the close. The recruiter using AI to write job posts still reads every one. The support team using AI to draft replies still reads every reply.

The making got fast. The checking didn't.

The bottleneck doesn't disappear, it moves onto a person

Here's the version we run into most.

A team we work with uses AI to draft replies to customer support tickets. Drafting used to take the bulk of the time. Now it's basically instant. And every single reply still sits there waiting for a human to read it before it can go out.

So the work didn't really get faster. It got faster in one place and then piled up in front of one person... which is not the same thing at all.

The part that's actually broken

Watch what happens in that queue.

The support person opens the draft. The tone is a little off, so they fix it. It's missing the policy, so they add it. It's twice as long as it needs to be, so they trim it. Then they hit send, and all three of those corrections leave with the ticket.

Next week, same three fixes. Week after that, same three fixes.

So the person gets very good at correcting the same mistakes, and the AI never gets any better. That's the trap. It looks like review, and it's really just rewriting with extra steps.

Our take: if your correction never makes it back into the system, you're not reviewing the AI. You're working for it.

What good looks like has to be written down

Anthropic didn't solve their problem by adding more reviewers. They got specific about which checks actually needed to run.

Everywhere else, that starts with success criteria. Most teams are working off "make it sound like us," which nobody can check quickly and nobody can hand back to the AI.

Here's what that support team wrote down instead:

  • Encouraging of the learner, never condescending

  • Only uses material approved in our methodology

  • Brief, 8th grade reading level

Two things happen once that list exists.

Review gets fast, because the reviewer is checking against a list instead of arguing with their own taste. And the corrections finally have somewhere to go. "Missing the policy" stops being a thing a human fixes by hand every week and turns into a line the AI reads before it drafts.

That's progressive delegation doing its real work. You earn your way to less review by turning corrections into criteria, not by lowering the bar.

And honestly, none of this is hard... it's just the step everyone skips, because fixing the draft in front of you always feels faster than writing down the rule.

What to do this week

  1. Pull last week's AI drafts that somebody edited before they went out. Read 10 of them.

  2. Write down the edits that keep showing up. Three is plenty.

  3. Turn those three into success criteria, and put them in two places: where the AI reads them, and where the reviewer checks them.

What's the edit you keep making by hand? Hit reply and tell us. We're guessing tone or a missing policy, and we're collecting these.

LINKS

For your reading list ๐Ÿ“š

That's all!

We'll see you again soon. Thoughts, feedback and questions are much appreciated - respond here or shoot us a note at [email protected]

Cheers,

๐Ÿช„ The AMP Team (formerly: the AI Exchange Team)