"The new model got worse" is now a monthly ritual. Opus 5 is the...

Half of it is your context. Half is real behavior drift. The problem: your evals can't tell which.
A framework for behavioral evals:
It says nothing about HOW. Verbosity, scope creep, severity inflation, sycophancy - all invisible to a pass/fail grade.
Capability can hold flat while behavior regresses hard.
"Correct but annoying" still churns users. Benchmarks score the correct. Nobody scores the annoying.
- Scope: lines changed vs lines requested
- Verbosity: output tokens per unit of task
- Severity: does its triage match yours
- Deference: does it push back or steamroll
Run every model version against it. Log the behavioral axes, not just pass/fail. Now "it feels worse" becomes a number you can diff across releases.
The vendor optimizes for the leaderboard. You optimize for your workload. They diverge more every release.
Pin the version, then upgrade deliberately against your golden set.
Capability is table stakes. Behavior is the moat. Measure it.
@AnthropicAI #AIeval #AgenticAI