- AMP (Formerly: The AI Exchange)
- Posts
- GPT-6 can do the work. Can your instructions survive it?
GPT-6 can do the work. Can your instructions survive it?
Edition 199 - More capable AI makes vague instructions more expensive.
Here’s what we’re reading and thinking about this week:
OpenAI just released GPT-6 Astra, its most capable model yet, built to carry complex work across apps and files.
That sounds like hiring an extremely capable new employee... except this one has a habit of going rogue and almost never asks for feedback.
The smarter the model gets, the more expensive vague instructions become.
OpenAI's announcement focuses on what the model can do. For operators, that is only half the story. The other half is whether your instructions can survive all that capability.
Because intelligence was never the main constraint. Reliability and control were.
Smart was never the problem
Think about a sales-proposal AI agent. It reads a call transcript, pulls in feedback from email, and rebuilds the proposal. The output looks polished, the work happens fast, and the team saves hours.
Then one day it goes entirely off script and proposes work the company cannot deliver.
That is not a writing problem. The AI did a very good job writing the wrong proposal.
The process never made clear what the agent could promise, what it had to treat as a hard constraint, or when it needed a person to approve the work. More capability would not fix that. It would help the agent make the same mistake faster, across more files, with more confidence.
Giving agents a job description
We talk a lot about playbooks here. You can, of course, run your playbooks in general-purpose agents, but if you are going to build more specialized agents, then this is where building job descriptions really helps.
For example:
Its job. What outcome does it own?
Its capabilities. What actions may it take to produce that outcome?
Its constraints. What may it never infer, change, approve, or promise?
Its operating context. What tools, files, systems, and information may it access?
This is what you would do for an extremely capable new hire. You would not say, "Handle sales," give them access to every system, and hope they develop good judgment by Thursday.
But teams do the AI version of that every day.
Design the moment when it asks
Here's the thing though... AI rarely volunteers that it needs feedback.
It does not feel uncertain in a way you can see. It does not stop by your desk. And it has no instinct for the moment when a small assumption becomes a large business risk.
You have to design that moment into the playbook.
Name when the AI must stop. Name what it must escalate. Name what a human reviews before the work moves forward. If a proposal includes a new service, changes pricing, or makes a promise outside the approved scope, the agent does not keep going. It asks.
Capability makes playbooks more valuable
The argument against detailed playbooks usually sounds like this: the models are smart enough now. We should not need to explain so much.
But smart was never the problem. In fact, smart makes AI better at following a strong playbook.
Better models do not make playbooks obsolete. They turn vague instructions into risk at scale.
The upside is just as big. Once the job, boundaries, context, and review points are clear, you can build systems that handle much more complex work while you keep reliability and control.
That's the shift. The question is no longer whether the AI can do the work. It is whether you have defined the work well enough to let it.
LINKS
For your reading list 📚
Nvidia just bought Hugging Face for $13 billion, and promised to keep it open. That's one enormous ecosystem bet.
The data and the panic are a mismatch… AI adoption is climbing fast, while layoffs remain rare… what is going on??
We thought this was cool: Gemini can now search video for exact moments instead of watching every frame.
Five major launches in three days? Enterprise buyers have officially entered "model fatigue".
Meta is teaching robot arms to plug in servers, and yes, the robots spend a lot of time charging.
That's all!
We'll see you again soon. Thoughts, feedback and questions are much appreciated - respond here or shoot us a note at [email protected]
Cheers,
🪄 The AMP Team (formerly: the AI Exchange Team)