Local memory, not a token shortcut
An n-gram table maps short token sequences to learned vectors. A retrieved vector is added to the current token representation near the embedding or an early hidden layer; the backbone still computes the hidden state and next-token logits.
The table and backbone learn together. Frequent local regularities can occupy table capacity and add little dense compute, but a lookup neither outputs nor guarantees the next token.
Generate with vs. without n-gram memory
Both paths start with the same tokens, base representation, and pre-lookup logits. Switch examples to see when a local lookup helps—and when it cannot.
The capital of France is▌Same base token representation
xₜ = token_embed(t)Same illustrative starting scores
Paris4.2the3.8located2.9Without n-gram embeddings
Normal model path
xₜ → backbone → hₜ → LM headThe backbone must recover the familiar continuation from its ordinary hidden state.
Illustrative downstream logits
Paris6.1the4.0located3.1With n-gram embeddings
Bigram / trigram lookup
[France · is] + [of · France · is] → learned rows foundAdd retrieved vector
x′ₜ = xₜ + e_phraseNormal model path
x′ₜ → backbone → h′ₜ → LM headThe matching local patterns retrieve a trained vector. After addition, downstream layers put more mass on the fitting continuation.
Illustrative downstream logits
Paris7.8the3.5located2.7The vector changes the hidden representation, not the answer directly. The backbone and LM head still produce and sample the next-token distribution; a lookup is a learned bias, never a guaranteed continuation.
Where gains come from
- Frequent phrases and syntax patterns become reusable local vectors instead of being reconstructed entirely by dense layers.
- Large lookup tables add parameter capacity while each token reads only a few rows, so added FLOPs can stay small.
- Deterministic addressing is cheap and can be prefetched; useful vectors reshape later logits through the normal backbone.
Where the lookup is useless
- Novel combinations and rare or unseen n-grams have no well-trained local row.
- Long-range dependencies, semantic reasoning, and facts outside the local window still require the backbone.
- Hash collisions can mix unrelated patterns; bandwidth, cache locality, host transfers, and batching can erase latency gains.
- Extra capacity does not guarantee downstream accuracy: an unhelpful vector may do nothing or add noise.
Inspect keys and values
A compact view of how local token sequences address learned rows.
The hash and vector values here are illustrative; production systems choose their own addressing and collision strategy.
| Position | Local key | Illustrative slot | Learned value vector |
|---|---|---|---|
| t=1 | [the · quick] | 746,393 | [1.0, 1.3, 0.9, 0.7, …] |
| t=2 | [quick · brown] | 1,712,346 | [1.1, -0.9, -0.7, 0.5, …] |
| t=3 | [brown · fox] | 18,478,650 | [1.1, 0.2, 1.4, 0.4, …] |
Real-world case study: Qwen3.8-Flash-Next
Qwen uses bigram and trigram tables as one component of a published experimental sparse model. System-level benchmark gains cannot be assigned to this component alone.
These are lookup-capacity parameters, not 51B dense parameters evaluated for every token. Only addressed rows are fetched.
Do not confuse the mechanism
Not n-gram speculative decoding
Speculative decoding drafts and verifies future tokens. N-gram embeddings retrieve a vector for the current representation; they draft nothing.
Not Multi-Token Prediction (MTP)
MTP trains several future offsets. N-gram embeddings are keyed memory for already observed local sequences.