🤖 AI & Machine Learning We benchmarked DeepSeek V4.1 Flash by @deepseek_ai . It reached 98%...OpenDesign@OpenDesignHQ 13 views Sep 10, 2026 ~1 min read Advertisement 1 We benchmarked DeepSeek V4.1 Flash by @deepseek_ai .It reached 98% of GPT-6 Astra’s score at 1.4% of the cost on everyday design tasks based on user requests.Every model except Astra scored lower AND cost more.Are open models overtaking closed ones?Full results below ↘️ 2 1) Most benchmarks test models’ limits. But we noticed that most users just want to know: which model should I use for everyday design work?We built OpenDesign Arena to help you choose.open-design.ai/zh/llm-arena-f… 3 2) Are closed models doomed?11 models scored lower AND cost more than DeepSeek V4.1 Flash across our real-world design benchmark.Only GPT-6 Astra scored higher. 4 3) Cheaper. Faster. Nearly the same score.Average time & cost per artifact:DeepSeek V4.1 Flash: 5.3 min / $0.023GPT-6 Astra: 11.1 min / $1.61Claude Fable 5.1: 12.8 min / $3.6698% of Astra’s score. 1/70th the cost. 5 4) Top-2 score. Fastest among the top 3.DeepSeek V4.1 Flash delivers in 5.3 minutes—less than half the time of Astra or Claude. 6 5)DeepSeek V4.1 Flash across app design tasks:Web apps: #3 — 82.4Mobile apps: #5 — 82.5Desktop apps: tied #3 — 84.4 7 6) DeepSeek V4.1 Flash on websites and dashboards:Websites & landing pages: #4 — 79.3Dashboards & admin panels: #10 — 76.1Ahead of Astra on landing pages. Behind on dashboards.Choose for your task, not just the overall ranking. 8 Explore the full benchmark results and comparisons here:open-design.ai/llm-arena-for-…Save this thread — create a free accountSave this thread Sign Up