Why working with Opus 5 feels worse, despite better benchmarks
A developer's blog post argues that Opus 5 is more capable than Opus 4.7 and 4.8 on benchmarks, yet worse to work with — because benchmark-driven training rewards confident assumptions over stopping to ask questions.
An opinion piece published on August 14, 2026, under the title “Why does Opus 5 feel worse to work with?”, has struck a chord with developers: in the author's view, and that of colleagues he says he has spoken with, working with Opus 5 feels like a downgrade compared to Opus 4.7, Opus 4.8 and Fable.
More capable, yet harder to work with
The author is careful to say he is not claiming a step backwards in capability: by his account, Opus 5 is a more capable model than Opus 4.7 and Opus 4.8, and even rivals Fable on benchmarks. The difference, he argues, is behavioural. The older models stop and ask questions when the intent is unclear, avoid assumptions without checking, and do not reinterpret or update a user's plan without asking. Opus 5, in his description, does not — and therefore needs careful babysitting.
Two forces behind the shift
His explanation is explicitly speculative. He points to two compounding pressures at Anthropic, and at frontier labs in general: the desire to build a self-improving AI capable of recursively bootstrapping itself towards AGI/ASI, and the pressure to score highly on benchmarks. Many benchmark tasks, he notes, are an open secret of being ill-defined, unfair or hackable — but a good coding task is self-contained: it can be solved without hints, without reading the task creator's mind, and without outside information. Training on such tasks, in his view, selects for models that make bold, usually-correct assumptions when things are ambiguous, and penalises models that stop to ask.
Real life is not a benchmark
That tendency is the opposite of what many users want from a coding agent. The author argues it is nearly impossible to write down all the context, intentions, business implications and budget constraints an agent could need, so ambiguity is inevitable — and being able to stop and ask questions is valuable. Real life, he writes, does not have a guaranteed right answer to every question, and with real consequences on the line he does not want an agent simply taking its best guess.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.