Doctors just built a test for whether AI knows it's wrong

Edition 194 - Here's how to apply the thinking to your work

Here’s what we’re reading and thinking about this week:

The most useful thing an AI can tell you is that it wasn't sure.

Almost none of them do it on their own, and almost nobody thinks to ask them to.

The news this week: Doctors built a test for AI uncertainty

We were very interested to see this one, because it comes from radiology, where a confident wrong answer can kill someone.

It's called RadLE 2.0. Older tests scored one thing, which was whether the AI got the right answer, and this one adds a score for something else entirely: can the AI tell when it's out of its depth and hand the case back to a human?

The people who built it explain the point in a single line. A confident wrong answer is more dangerous than an honest "I don't know."

That lands differently once you look at your own work, because your AI answers everything in the same steady voice, and whether it knows the answer cold or it's mostly guessing, what shows up on your screen looks about the same. Clean, organized, and sure of itself.

That isn't the AI being dishonest with you. Nobody ever told it that doubt was allowed.

Solve #1: Give it permission to say "I'm not sure"

So smart operators tell it, and they write that permission into the playbook itself rather than hoping it shows up, because the AI is never going to take it on its own.

The simple version is a single line asking it to tell you when it isn't sure and where. The better version, and the one that changes how the whole thing runs, is asking to see its second and third idea too.

Solve #2: Ask for three ideas, not one

We use a format called the 1-3-1, which is just here's the situation, here are three options, and here's the one I'd pick and why.

Why does this work?

Well, it's actually inspired by a management technique - it's what a good employee does when they bring you something hard. Instead of hiding the options they didn't choose, they show you what they weighed and then commit to a pick, and that combination is what makes it useful to you.

When AI hands you one confident answer, you're getting the end of its thinking without any of the middle, and the middle is where it made all the real choices.

The impact: You'll spend less time checking, not more

You'd think asking for more would mean reading more, but it works the other way around.

Checking a single answer means quietly redoing the work in your head to see whether you'd land in the same place. Checking a 1-3-1 takes about fifteen seconds, because you're only judging a choice that somebody already laid out for you, and you can usually tell right away if it weighed the wrong thing.

That's the difference between a review step you keep and one you quietly stop doing after a couple of weeks.

Food for thought :)

Do this this week

Take one playbook you already run and add this to the end of it:

"Before you give me your answer, give me three options and tell me which one you'd pick and why. Then list anything you weren't sure about, anything in my instructions that was unclear, and anything you had to assume."

Run it once and look at what it picked, then spend a little longer on what it almost picked.

Then hit reply and tell us whether it chose the one you would have. We read every one.

LINKS

For your reading list 📚

That's all!

We'll see you again soon. Thoughts, feedback and questions are much appreciated - respond here or shoot us a note at [email protected]

Cheers,

🪄 The AMP Team (formerly: the AI Exchange Team)