I took Qwen3.8-27B apart to see how it works inside. The plan was...

Every ffn neuron tapped, all 1,114,112 of them. 3.92M tokens pushed through. 87 minutes on the 4x3090.
The result surprised me. Neurons that fire on more than 90% of tokens: 2.
Thats what youre looking at, one frame per token. Full numbers below 🧵

The recipe is 3 papers deep: moefication (arxiv 2110.01786), cmoe (2502.04416), expertweaver (2602.15521, icml 2026). Wiretap which ffn neurons fire on which tokens, glue the always-on ones into a shared expert, cluster the rest by who fires together, build the router from the same stats. cmoe turns a 7B dense into a usable MoE in 5 minutes. Their best score so far is Qwen2.5-7B keeping 67.0 avg out of 75.2 dense, thats 89%, after you skip 25% of the ffn. The aggressive mode with 25% active needs 200B tokens of retraining. Im not doing that part.
Nobody tried this on anything newer than 2024 or bigger than 8B. Qwen3.8-27B has 64 layers with 17408 ffn neurons each, 1.11M neurons total, and 17.1B of its 27.8B params are ffn, 62% of the model. Im cutting at 75, 50 and 25 percent active, rerunning the same evals against dense, plotting what survives.
My prediction: nothing will work out. The trick feeds on lazy neurons and the 27B is a frontier distill, probably packed too dense to cut. If quality craters, the crater is the number. Nobody has measured how much redundancy a 2026 dense model still carries, win or wreck I get it first.
Perfect router at 50% budget: keeps 99.9% of the fired neurons, but only 57% of the energy. Picking 54% of neurons at random keeps 54%.
So the clever selection barely beats a coin flip on the thing that probably matters.
My bet: 50% survives ugly, 25% is dead on arrival, and layer 63 with its 3 experts breaks first.



