💻 Tech & Development🤖 AI & Machine Learning
We ran 5 coding-agent harnesses on one model, DeepSeek V4.1 Flash, across the 30 SWE-bench Lite tasks.
The harness alone moved success from 61% to 75% and cost per attempt by 3.2×.
Two agents even tried to cheat their way out of the sandbox, see how 🧵 ...
Sep 30, 2026