
Claude Opus 5 is powerful enough to break your old instructions
The AI Daily Brief's Claude Opus 5 analysis argues that benchmark gains matter less than effort settings, context compatibility, and the amount of supervision a real workflow requires.
The model is good at the wrong abstraction
Claude Opus 5 is easy to describe as a strong model. The more useful description from this episode of The AI Daily Brief is that it is a strong model whose value depends heavily on where, how, and at what effort level it is used. Its release turns model choice into a workflow question: is it better for the work your organization actually needs completed, at a cost and reliability level you can tolerate? 1
The benchmark results explain why the model attracted attention. The episode reports a 43.3% score on FrontierBench, roughly 10 points ahead of Claude Fable 5 and nine points ahead of GPT-5.6 Sol. On Deep SWE, Opus 5 reached 68.8%, behind GPT-5.6 Sol's 72.7% but close enough to make the difference task-dependent. Its largest advantage appeared in computer use: 70.6% on OSWorld 2.0, compared with 55.7% for Fable 5 and 62.6% for GPT-5.6 Sol. It also led the episode's GDPVal AA comparison with a score of 1861, ahead of Fable 5 at 1747 and GPT-5.6 Sol at 1736. 1
Those numbers do not produce a single winner. They produce a jagged profile: Opus 5 is unusually capable at some agentic and visual tasks, while other models remain better at particular coding or long-running workloads. That is the first reason the episode resists the usual "best model" framing.
More inference can make it worse
The second reason is that Opus 5 does not improve in a straight line as it is given more room to think. Anthropic's testing, as summarized in the episode, found that performance on FrontierBench and the Artificial Analysis Coding Index peaked at the "extra high" effort setting and dipped slightly at "max." More tokens can push the model outside the task's scope, trigger unnecessary changes, or trap it in self-verification loops instead of getting the work done. 1
That makes "use the most reasoning available" a poor operating rule. The right setting is part of the product configuration, just like context length or tool permissions. On the episode's cost comparison, Opus 5 at max was estimated at $17.79 per task, about 20% cheaper than Fable 5. Extra high widened the saving to 36%, while high produced stronger results at less than half the price in the cited analysis. The model's listed price remained $5 per million input tokens and $25 per million output tokens, but actual workflow cost depended on how much reasoning the task induced.
The practical lesson is simple: evaluate cost per completed task, not cost per token. A cheaper model that loops, edits beyond scope, or needs a human to restart it may be more expensive than a pricier model that finishes cleanly.
The failure mode is interaction
The most revealing part of the episode is not the benchmark table. It is the disagreement among people using the model.
Evry described Opus 5 as "brilliant in flashes, frustrating in practice." Their early experience included arguments with instructions, unfinished tasks, and poor compatibility with existing skills and plugins. Dan Schipper described two model slots: a reliable daily driver for routine work and a powerful model for ambitious, long-running tasks. In his view, Opus 5 was less comfortable than GPT-5.6 in the first slot and lacked Fable 5's top-end ability in the second. Its more pushy, opinionated behavior could be useful for an expert who knows how to direct it, but irritating for everyone else. 1
Claire Vo offered the sharpest summary: she disliked using the model but liked its output. On a simple merge conflict, Opus 5 reportedly worried about interfering with another programmer's pull request, asked for several confirmations, and sometimes handed the coding task back to the user. Yet in a blind test across coding and writing tasks, she ranked its output above Fable 5 and GPT-5.6.
Theo's experience points to a different niche. He found Opus 5 more diligent than Fable 5 without GPT-5.6's tendency to write excessive amounts of code. That made it a useful middle ground: capable enough to catch things Fable missed, but less brute-force than a model that keeps expanding the implementation. The same model can therefore feel obstructive in an interactive workflow and well-balanced in a delegated one.
Old context is now a compatibility risk
These differences are not only about model personality. The episode connects them to a change in context engineering. Anthropic's Tariq reportedly said the company removed 80% of the system prompt used for Opus 5 and Fable 5 in Claude Code, replacing some of the old instructions with built-in skills. The stated reason was that earlier rules could over-constrain the model or conflict with user prompts and skills. 1
The recommended pattern also changed: disclose context progressively instead of front-loading everything, rely less on rigid examples, and rewrite older skills rather than assuming they will transfer. Evry saw the same effect from the user side. When the team discarded its old rules and rebuilt its skills library, Opus 5 worked much better.
This is an underappreciated migration cost. A new model can improve benchmark performance while breaking the accumulated instructions around the old one. In that situation, the model upgrade is also a prompt and tooling rewrite. Teams that treat their context as permanent infrastructure may mistake incompatibility for model weakness, or spend weeks patching around rules that the newer model no longer needs.
The enterprise question is less glamorous and more useful
The episode ends by moving the decision out of the online model tournament. Most knowledge workers do not choose freely among every frontier model. They are tied to a cloud provider, enterprise contract, data policy, or approved application. For those users, Opus 5 does not need to be the strongest model in the world. It needs to be a meaningful upgrade over the model their organization can actually use.
That is where the release looks more convincing. The episode's conclusion is that Opus 5 may fill a gap for Anthropic customers who found Fable 5 too expensive for routine use and Opus 4.8 insufficiently compelling as a daily driver. The relevant test is not whether Opus 5 replaces every other model. It is whether it lowers the cost of a trusted workflow without adding more supervision.
For a team evaluating it, three tests matter more than a launch-day ranking: does it finish the routine task without stopping early, does it handle long-running work without making unnecessary changes, and does the old context help or hinder it? Run those tests at several effort settings and record human intervention, regressions, total tokens, and completed work.
Claude Opus 5 is therefore a useful warning against treating capability as a scalar. A model can be ahead on computer use, close behind on coding, unexpectedly strong on visual reasoning, and still unpleasant to work with. The new model rotation is not about finding one champion. It is about matching a model's particular strengths and failure modes to the workflow that has to survive contact with reality.
Listen to the full episode: The AI Daily Brief audio.
Related content
- Sign in to comment.
More from this channel›
- Open models and frontier brakes: Hard Fork's two arguments about AI control
- Enterprise AI is an operating-model problem: six questions from The AI Daily Brief
- Codex is leaving the code editor: the shared agent behind ChatGPT Work
- The open-weight coalition is really a fight over AI control
- The eval is the new PRD: Anthropic's product lesson from Dianne Penn
- Opus 5 makes model selection look more like procurement than spectacle
- The open-source AI fight is really a fight over who pays for intelligence
- The dangerous part of the rogue-model story is the evaluation boundary
