What is Batching in LLM Inference?
Batching processes several requests together. During decoding, this can share the cost of reading model weights across sequences. The benefit depends on context length, memory traffic, compute capacity and scheduling.
Static Batching
A static batch starts a fixed set of requests together. Completed requests leave unused capacity until the longest request finishes and the next batch starts.
Dynamic / Continuous Batching
Continuous batching can admit waiting requests between decoding steps as capacity becomes available. The scheduler must respect memory limits, request lengths and service goals.
Throughput vs Batch Size
Batching shares weight reads across requests. Throughput can approach a compute or memory limit. Longer contexts add KV traffic and memory use. The point where a batch stops fitting depends on the model, context, precision and device capacity.
A hypothetical dense decoder
7B FP16 weights (14 GB), 32 layers, 8 KV heads × 128, FP16 KV, 1 TB/s memory bandwidth and 100 TFLOP/s compute. The lower bound is max(weight + KV read time, matrix compute time). It omits attention FLOPs, communication and kernel overhead; it is not a benchmark.
Prefill vs Decode Phase
Prefill processes prompt positions together within each layer. Decode adds tokens sequentially per request. Their relative speed depends on model, hardware, context and batching; there is no universal multiplier.
Dependency example, not a speed race: prompt positions can be processed together within a layer, while each new output token depends on preceding outputs. The animation clock does not represent hardware time.
Prefill
Decode
Continuous Batching
Continuous batching admits new requests when slots become free. This example shows occupancy for four slots with supplied request lengths. Occupied slots are not a hardware utilization measurement.
💡 Continuous scheduling reuses free slots. Real servers must also respect memory and scheduling limits.