I took Qwen3.8-27B apart to see how it works inside. The plan was...

@superalesha
Alexey Fateev@superalesha
5 views Aug 23, 2026 ~2 min read
Advertisement
1
I took Qwen3.8-27B apart to see how it works inside. The plan was to carve a MoE out of it.

Every ffn neuron tapped, all 1,114,112 of them. 3.92M tokens pushed through. 87 minutes on the 4x3090.

The result surprised me. Neurons that fire on more than 90% of tokens: 2.

Thats what youre looking at, one frame per token. Full numbers below 🧵
0:55
@superalesha
Alexey Fateev@superalesha
Qwen didnt give us a small MoE in the 3.8 drop, so im cutting one out of Qwen3.8-27B myself. Zero training, pure surgery on the weights.

The recipe is 3 papers deep: moefication (arxiv 2110.01786), cmoe (2502.04416), expertweaver (2602.15521, icml 2026). Wiretap which ffn neurons fire on which tokens, glue the always-on ones into a shared expert, cluster the rest by who fires together, build the router from the same stats. cmoe turns a 7B dense into a usable MoE in 5 minutes. Their best score so far is Qwen2.5-7B keeping 67.0 avg out of 75.2 dense, thats 89%, after you skip 25% of the ffn. The aggressive mode with 25% active needs 200B tokens of retraining. Im not doing that part.

Nobody tried this on anything newer than 2024 or bigger than 8B. Qwen3.8-27B has 64 layers with 17408 ffn neurons each, 1.11M neurons total, and 17.1B of its 27.8B params are ffn, 62% of the model. Im cutting at 75, 50 and 25 percent active, rerunning the same evals against dense, plotting what survives.

My prediction: nothing will work out. The trick feeds on lazy neurons and the 27B is a frontier distill, probably packed too dense to cut. If quality craters, the crater is the number. Nobody has measured how much redundancy a 2026 dense model still carries, win or wreck I get it first.
Media image
2
The recipe from the papers starts by gluing the always-on neurons into a shared expert. On this model that population is two neurons wide.

One smooth tail, no split to cut along.

The quiet ones are real though: 19.6% of neurons almost never fire, and together they carry 2.9% of all firing.
Media image
3
There is slack, its just not where you would put it.

Last layer: 87% of its neurons almost never fire. Middle of the model, layers 16 to 31: only 9%.

So cutting every layer by the same amount is the wrong move. You undercut the tail and you butcher the middle, and the middle is doing the work.
Media image
4
Qwen3.8-27B alternates 48 cheap linear attention layers with 16 full attention ones. I assumed the ffn sitting above each type would pick up different habits.

It doesnt. Medians 0.00229 against 0.00223. Same shape, same everything.

Null result, but its a new one.
Media image
5
The top-64 neurons I count as fired carry only 5 to 14% of the layers actual activation mass. Rank by that and you keep the soloists while you mute the choir.

Perfect router at 50% budget: keeps 99.9% of the fired neurons, but only 57% of the energy. Picking 54% of neurons at random keeps 54%.

So the clever selection barely beats a coin flip on the thing that probably matters.
6
The experts are carved. Stage 3 is the eval that settles it.

My bet: 50% survives ugly, 25% is dead on arrival, and layer 63 with its 3 experts breaks first.
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial

Advertisement