Finally broke the 3k token per second input/prompt processing...

Results and steps to reproduce up on @LottoLabs LocalMaxxing here: localmaxxing.com/runs/cmouqgx9q…
3130t/s pp2048 is close to 4x faster than the fastest M5 Max number I could find on Reddit.
For long running agents, input token processing can be at least as important as output token processing and Spark shines for that!
localmaxxing.com/runs/cmomgvsoo…
(Though I hope that ultimately a lot of these sorts of optimizations become vLLM defaults in the future)
github.com/my-other-githu…
GB10 has a lot of potential if the software can catch up!