- AMP (Formerly: The AI Exchange)
- Posts
- Doctors just built a test for whether AI knows it's wrong
Doctors just built a test for whether AI knows it's wrong
Edition 194 - Here's how to apply the thinking to your work
Here’s what we’re reading and thinking about this week:
The most useful thing an AI can tell you is that it wasn't sure.
Almost none of them do it on their own, and almost nobody thinks to ask them to.
The news this week: Doctors built a test for AI uncertainty
We were very interested to see this one, because it comes from radiology, where a confident wrong answer can kill someone.
It's called RadLE 2.0. Older tests scored one thing, which was whether the AI got the right answer, and this one adds a score for something else entirely: can the AI tell when it's out of its depth and hand the case back to a human?
The people who built it explain the point in a single line. A confident wrong answer is more dangerous than an honest "I don't know."
That lands differently once you look at your own work, because your AI answers everything in the same steady voice, and whether it knows the answer cold or it's mostly guessing, what shows up on your screen looks about the same. Clean, organized, and sure of itself.
That isn't the AI being dishonest with you. Nobody ever told it that doubt was allowed.
Solve #1: Give it permission to say "I'm not sure"
So smart operators tell it, and they write that permission into the playbook itself rather than hoping it shows up, because the AI is never going to take it on its own.
The simple version is a single line asking it to tell you when it isn't sure and where. The better version, and the one that changes how the whole thing runs, is asking to see its second and third idea too.
Solve #2: Ask for three ideas, not one
We use a format called the 1-3-1, which is just here's the situation, here are three options, and here's the one I'd pick and why.
Why does this work?
Well, it's actually inspired by a management technique - it's what a good employee does when they bring you something hard. Instead of hiding the options they didn't choose, they show you what they weighed and then commit to a pick, and that combination is what makes it useful to you.
When AI hands you one confident answer, you're getting the end of its thinking without any of the middle, and the middle is where it made all the real choices.
The impact: You'll spend less time checking, not more
You'd think asking for more would mean reading more, but it works the other way around.
Checking a single answer means quietly redoing the work in your head to see whether you'd land in the same place. Checking a 1-3-1 takes about fifteen seconds, because you're only judging a choice that somebody already laid out for you, and you can usually tell right away if it weighed the wrong thing.
That's the difference between a review step you keep and one you quietly stop doing after a couple of weeks.
Food for thought :)
Do this this week
Take one playbook you already run and add this to the end of it:
"Before you give me your answer, give me three options and tell me which one you'd pick and why. Then list anything you weren't sure about, anything in my instructions that was unclear, and anything you had to assume."
Run it once and look at what it picked, then spend a little longer on what it almost picked.
Then hit reply and tell us whether it chose the one you would have. We read every one.
LINKS
For your reading list 📚
An autonomous AI agent broke into Hugging Face over a weekend. Then the defenders' own AI refused to help investigate, because it couldn't tell them apart from the attacker.
Everyone assumes the AI-heavy companies cut headcount first. JLL asked 2,200 leaders and got the opposite: the ones furthest along are hiring hardest, entry-level included.
A Chinese open model just took the top coding spot from Fable 5 and GPT-5.6. The full model goes public July 27, which is the part worth watching.
The District 9 director released a 13-minute film made entirely with AI, and the reviews were brutal.
That's all!
We'll see you again soon. Thoughts, feedback and questions are much appreciated - respond here or shoot us a note at [email protected]
Cheers,
🪄 The AMP Team (formerly: the AI Exchange Team)