We benchmarked DeepSeek V4.1 Flash by @deepseek_ai . It reached 98%...

@OpenDesignHQ
OpenDesign@OpenDesignHQ
13 views Sep 10, 2026 ~1 min read
Advertisement
1
We benchmarked DeepSeek V4.1 Flash by @deepseek_ai .

It reached 98% of GPT-6 Astra’s score at 1.4% of the cost on everyday design tasks based on user requests.

Every model except Astra scored lower AND cost more.

Are open models overtaking closed ones?
Full results below ↘️
Media image
2
1) Most benchmarks test models’ limits. But we noticed that most users just want to know: which model should I use for everyday design work?

We built OpenDesign Arena to help you choose.
open-design.ai/zh/llm-arena-f…
Media image
3
2) Are closed models doomed?

11 models scored lower AND cost more than DeepSeek V4.1 Flash across our real-world design benchmark.

Only GPT-6 Astra scored higher.
Media image
4
3) Cheaper. Faster. Nearly the same score.

Average time & cost per artifact:
DeepSeek V4.1 Flash: 5.3 min / $0.023
GPT-6 Astra: 11.1 min / $1.61
Claude Fable 5.1: 12.8 min / $3.66

98% of Astra’s score. 1/70th the cost.
Media image
5
4) Top-2 score. Fastest among the top 3.
DeepSeek V4.1 Flash delivers in 5.3 minutes—less than half the time of Astra or Claude.
Media image
6
5)DeepSeek V4.1 Flash across app design tasks:

Web apps: #3 — 82.4
Mobile apps: #5 — 82.5
Desktop apps: tied #3 — 84.4
Media image
Media image
Media image
7
6) DeepSeek V4.1 Flash on websites and dashboards:

Websites & landing pages: #4 — 79.3
Dashboards & admin panels: #10 — 76.1
Ahead of Astra on landing pages. Behind on dashboards.

Choose for your task, not just the overall ranking.
Media image
Media image
8
Explore the full benchmark results and comparisons here:
open-design.ai/llm-arena-for-…
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial

Advertisement