If you ask a frontier LLM a multi-hop reasoning question, e.g.,...

We investigated how + what the models are doing and showed the computation is readable from the hidden states over the "meaningless" filler tokens: arxiv.org/abs/2607.03502
blog.redwoodresearch.org/p/recent-llms-…

