Inference Engineering the new hot space in LLM π₯ A book Inference...

A book Inference Engineering by @philipkiely & @baseten gaining popularity in tech world.
I am on the way to complete it and as a reader with ML background, I've written an article on how to start learning Inference Engineering:
- It builds on ML engineering, neural networks and LLMs. That ground comes first.
- Training happens once. Inference happens every time a user types.
- A trained model is fixed. You cannot make it smarter, only cheaper and faster, and that gap is multiples.
- Inference runs in two phases: prefill and decode.
- Decode is where the cost lives.
- The KV cache is the trick that stops decode redoing work.
- To start you need LLM basics and a little GPU. Not CUDA.
- Measure before you optimise, then ask the only question that matters. Why is the GPU idle?
I added all mandatory things you need to start with resources and survey about the future of inference engineering.
Personally I come from MLOps background and nowadays using LLM calls also impacts both user experience and business. While cost is always headache.
This field allows us to understand inference and build your own dedicated architecture about infernece.
Eventually improves LLM calling and business.
Read the full article.