6 Topics
LLM Inference
Understand how large language models generate text efficiently — from KV caching to batching strategies and serving infrastructure.
Help Make This Better
This guide is open source. Got an idea for a new topic? Found a bug? Want to improve an explanation? Every contribution helps.
01
KV CacheStore computed keys and values to avoid redundant workSep 13, 2026
E02
Mega-KernelsFuse operations and move whole transformer steps into a single GPU kernel to cut launch overhead and HBM trafficSep 24, 2026
E03
Prompt CachingReuse computed KV caches across API requests to save cost and latencySep 13, 2026
I04
Batching & ThroughputProcess multiple requests simultaneously for higher throughputSep 13, 2026
I05
Running Models LocallyRun LLMs on your own hardware for privacy, speed, and zero API costsSep 13, 2026
B06
VRAM CalculatorCalculate a local LLM memory budget from explicit weight, cache and runtime assumptionsSep 13, 2026
B