"The new model got worse" is now a monthly ritual. Opus 5 is the...

@bally_kehal
Bally_AgenticAI@bally_kehal
1 views Aug 19, 2026 ~2 min read
Advertisement
1
"The new model got worse" is now a monthly ritual. Opus 5 is the latest target - too verbose, over-engineers, treats every bug like a Sev-1.

Half of it is your context. Half is real behavior drift. The problem: your evals can't tell which.

A framework for behavioral evals:
2
Your eval suite measures capability: did it pass, hit the benchmark, solve the task.

It says nothing about HOW. Verbosity, scope creep, severity inflation, sycophancy - all invisible to a pass/fail grade.

Capability can hold flat while behavior regresses hard.
3
Behavior is a product surface. A model that solves the task but rewrites 800 lines you never asked for is expensive - in tokens, in review time, in trust.

"Correct but annoying" still churns users. Benchmarks score the correct. Nobody scores the annoying.
4
Make behavior measurable. Pick axes that map to real cost:
- Scope: lines changed vs lines requested
- Verbosity: output tokens per unit of task
- Severity: does its triage match yours
- Deference: does it push back or steamroll
5
Build a golden set of 30-50 prompts from your actual workload. Not benchmark tasks - YOUR tasks.

Run every model version against it. Log the behavioral axes, not just pass/fail. Now "it feels worse" becomes a number you can diff across releases.
6
Gate releases on behavior, not just scores. A model that gains 2pts on a benchmark but doubles output length and scope is a regression for most production agents.

The vendor optimizes for the leaderboard. You optimize for your workload. They diverge more every release.
7
This is why version-pinning matters. "Latest" silently changes behavior under you. If you can't measure the drift, you can't defend it - you just wake up to angry users and a Reddit thread.

Pin the version, then upgrade deliberately against your golden set.
8
The teams that stay calm through every "model got worse" cycle aren't lucky. They have behavioral evals and can point to exactly what changed.

Capability is table stakes. Behavior is the moat. Measure it.

@AnthropicAI #AIeval #AgenticAI
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial

Advertisement