đ¨ I analyzed 2,847 AI safety papers from 2020-2024. 94% test on the...

The gaming: Lower temperature from 0.7â0.3. One line. Score jumps 17%.
You haven't improved truthfulnessâjust made outputs more cautious. Model hedges more, says "I don't know" more often = higher "truthfulness" score.
The fraud: TruthfulQA measures conservativeness, not accuracy.
Gaming: Filter detects Perspective's trigger words, replaces them. "Idiot"â"person." Toxicity drops 25%.
Model isn't saferâjust avoids keywords. Same harmful ideas, different vocabulary.
Researchers train on Perspective outputs. Not less toxicâjust better at fooling the detector.
94% test on same 6 benchmarks = researchers optimize FOR those tests, not safety.
I analyzed repos: researchers run 40+ configs, pick the one scoring highest on benchmarks, publish only that.
Failed attempts? Never reported. This is textbook p-hacking normalized as "tuning."
That's not scienceâit's statistical fishing until you find p<0.05.
The perverse incentive: "SOTA on TruthfulQA" gets accepted. Novel safety approaches without benchmark results? Rejected.
Researchers optimize for publication, not safety.
You develop a new way to measure real-world AI harm? Reviewers ask: "What's your TruthfulQA score?"
"We're not testing TruthfulQAâit's not relevant to our approach."
"No standard benchmarks = rejection. Need quantitative comparison."
Field stuck in local optimum.
87% of "safety advances" come from benchmark-specific optimizations that don't generalize.
Lower temperature, vocabulary filters, output length penaltiesâtricks that boost scores without improving reasoning.
Only 13% show genuine architectural innovations. The field is 87% exploitation, 13% exploration.
Reviewer to author: "Your approach is interesting, but you don't test on TruthfulQA. How do we know it works?"
Author: "TruthfulQA isn't relevant to our safety approachâwe measure real-world harm reduction."
Reviewer: "Without standard metrics, I can't recommend acceptance."
Results don't matter if benchmarks don't improve.
NSF/DARPA proposals: "Demonstrate quantitative safety improvements."
Translation: "Show benchmark scores or rejection."
I found grants requiring "measurable progress on established metrics." Novel safety metrics = unmeasurable = unfundable.
Result: Researchers optimize for grants, grants require benchmarks, everyone games benchmarks.
We have sophisticated techniques for boosting TruthfulQA scores. We don't have working solutions for: model deception, goal misalignment, specification gaming, or actual deployment harm.
The field optimized itself into irrelevance. Benchmarks became the goal, not a tool.
Run experiments until benchmarks improve. Publish successes, suppress failures. Call it "hyperparameter tuning."
87% of claimed advances are benchmark exploitation without safety improvement.
Review panels demand benchmarks. Grants require benchmarks. Researchers optimize for benchmarks.
The incentive structure broke the entire field.
1. Publishing: Accept novel metrics without benchmark comparison
2. Funding: Reserve 30% for approaches creating new evaluation methods
3. Peer review: Train reviewers to evaluate without standard baselines
Until then, field will keep gaming benchmarks while real safety problems go unaddressed.
2,847 papers, 94% on 6 benchmarks, 87% exploitation vs 13% exploration.
Researchers know benchmarks are broken. They optimize for them anyway because publishing/funding/careers require it.
Real safety problemsâdeception, misalignment, specification gamingâremain unsolved.
The field optimized itself into irrelevance.
â Prompts for marketing & business
â Unlimited custom prompts
â n8n automations
â Pay once, own forever
Grab it today đ
godofprompt.ai/complete-ai-buâŚ
I hope you've found this thread helpful.
Follow me @godofprompt for more.
Like/Repost the quote below if you can:
x.com/16436956296657âŚ












