An agent that beats Claude Mythos on Terminal Bench and SWE-bench...

@lihanc02
Hanchen Li@lihanc02
20 views Apr 10, 2026 ~1 min read
Advertisement
1
An agent that beats Claude Mythos on Terminal Bench and SWE-bench Verified?

πŸŽ‰We are excited to share Terminator-1, our newest agent that achieved 95+% on SWE-bench Verified and Terminal-Bench with @MogicianTony!

We show that besides model capabilities, well-designed harness could actually boost the accuracy by 3x in coding tasks.

Well if you really wanted you could get 100% accuracy without solving a single task.

The actual finding is that most AI benchmarks can be easily reward-hacked with simple exploits. Read more about the same 7 design flaws that almost every evaluation has ⬇️
x.com/MogicianTony/s…
Media image
2
It is just a hack
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial

Advertisement