đ¨ This is wild. A new paper from the Ling team just dropped "Every...

A new paper from the Ling team just dropped "Every Attention Matters" and it quietly rewrites how long-context reasoning works in LLMs.
Their new Ring-linear architecture mixes Softmax and Linear Attention, cutting inference cost by 10x while keeping SOTA accuracy up to 128K tokens.
Even crazier:
⢠Training efficiency +50%
⢠Inference speed +90%
⢠Stable RL optimization over ultra-long sequences
Basically, they solved long-context scaling without trillion-parameter overkill.
The future isnât bigger models. Itâs smarter attention.








