1:09

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models...

@thehypedotnews
thehype.@thehypedotnews
3 views Aug 30, 2026 ~3 min read
glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash

four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief

the setup: our own agent loop on @OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked.

tasks:

1. house – a plot and a palette, no plan. shape, height and material are the model's call
2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it
3. bridge – a river with one islet and banks at different heights. cross it however you want

models: @Zai_org glm 5.3 flash, @Alibaba_Qwen qwen 3.8 flash, @GoogleDeepMind gemini 3.7 flash, @deepseek_ai v4 flash vision

all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied

- total cost, three builds
#1 glm 5.3 flash – $0.201
#2 gemini 3.7 flash – $0.871
#3 qwen 3.8 flash – $1.058
#4 deepseek v4 flash – $1.567

- wall clock, three builds
#1 gemini 3.7 flash – 91m
#2 glm 5.3 flash – 228m
#3 deepseek v4 flash – 502m
#4 qwen 3.8 flash – 912m

- total tokens
#1 gemini 3.7 flash – 3,567,052
#2 glm 5.3 flash – 4,732,748
#3 qwen 3.8 flash – 13,469,333
#4 deepseek v4 flash – 18,230,076

- defects logged by the site
#1 deepseek v4 flash – 59
#2 gemini 3.7 flash – 132
#3 glm 5.3 flash – 221
#4 qwen 3.8 flash – 350

- material tallied across three builds
#1 gemini 3.7 flash – $359,884
#2 glm 5.3 flash – $583,358
#3 deepseek v4 flash – $1,327,484
#4 qwen 3.8 flash – $2,188,625

observations:

• glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate

• what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set

• gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock

• gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3

• qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there

conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest!

follow @thehypedotnews for 24/7 ai news, analysis and breakdowns
Actions
Advertisement