@_avichawla

Avi Chawla (@_avichawla)

View on X 21 Unrolled Threads
Thread Archive
190
🤖 AI & Machine Learning💻 Tech & Development

Everything you need to turn an open-source LLM into a fast, local decision engine without retraining it. It covers next-token scoring, fixed choices with probability distributions, SGLang, and a hands-on benchmark against standard text generation on the same model....

Sep 21, 2026
Thread Archive
185
🤖 AI & Machine Learning💻 Tech & Development

Everything you need to understand where your input tokens are being recomputed and what to do about it. It covers the four cache layers from first principles, their trade-offs, what happens when they interact, and the five most common problems that inhibit cache reuse....

Aug 28, 2026
Thread Archive
164
🤖 AI & Machine Learning💻 Tech & Development

Everything you need to understand where your agent's tokens actually go and what to do about it. It covers what a production harness owns beyond the execution loop, the four strategies that keep context flat, how credentials stay out of the sandbox, and how three harnesses compare on the same 14 tas...

Aug 24, 2026
Thread Archive
148
🤖 AI & Machine Learning🔬 Science & Research📰 News & Politics

8 techniques to make an LLM reason better at inference time, covered with tradeoffs and practical notes. Every one of them is running in production at a frontier lab today, and the research behind them comes from Google, OpenAI, and Anthropic....

Aug 15, 2026
Thread Archive
187
🤖 AI & Machine Learning

The technique behind vLLM's 23x throughput jump and the default scheduler in every serving engine you use. Covered with internals, the token budget, KV allocation, and the preemption that makes you pay for the same prefill twice....

Aug 13, 2026
Thread Archive
155
🤖 AI & Machine Learning

Every generate() call to an LLM runs two distinct computational phases on the same GPU:...

Jun 29, 2026
Thread Archive
76
🤖 AI & Machine Learning💻 Tech & Development

- Google Maps uses graph ML to predict ETA - Netflix uses graph ML in recommendation - Spotify uses graph ML in recommendation - Pinterest uses graph ML in recommendation Here are 6 must-know ways for graph feature engineering (with code):...

Dec 12, 2025
Thread Archive
103
🤖 AI & Machine Learning🔬 Science & Research💻 Tech & Development

Fine-tuning LLM Agents without Fine-tuning LLMs! Imagine improving your AI agent's performance from experience without ever touching the model weights. It's just like how humans remember past episodes and learn from them. That's precisely what Memento does. The core concept: Instead of updating...

Oct 24, 2025
Thread Archive
68
🤖 AI & Machine Learning

KV caching in LLMs, clearly explained (with visuals):...

Oct 07, 2025