N-gram Embeddings

Expert

See how learned local-pattern vectors can steer generation without replacing the language model.

Last updated: Aug 27, 2026

Local memory, not a token shortcut

An n-gram table maps short token sequences to learned vectors. A retrieved vector is added to the current token representation near the embedding or an early hidden layer; the backbone still computes the hidden state and next-token logits.

The table and backbone learn together. Frequent local regularities can occupy table capacity and add little dense compute, but a lookup neither outputs nor guarantees the next token.

Generate with vs. without n-gram memory

Both paths start with the same tokens, base representation, and pre-lookup logits. Switch examples to see when a local lookup helps—and when it cannot.

Scenario
Same context on both pathsFrequent local match
The capital of France is▌

Same base token representation

xₜ = token_embed(t)

Same illustrative starting scores

Paris4.2
the3.8
located2.9

Without n-gram embeddings

Normal model path

xₜ → backbone → hₜ → LM head

The backbone must recover the familiar continuation from its ordinary hidden state.

Illustrative downstream logits

Paris6.1
the4.0
located3.1

With n-gram embeddings

Bigram / trigram lookup

[France · is] + [of · France · is] → learned rows found

Add retrieved vector

x′ₜ = xₜ + e_phrase

Normal model path

x′ₜ → backbone → h′ₜ → LM head

The matching local patterns retrieve a trained vector. After addition, downstream layers put more mass on the fitting continuation.

Illustrative downstream logits

Paris7.8
the3.5
located2.7

The vector changes the hidden representation, not the answer directly. The backbone and LM head still produce and sample the next-token distribution; a lookup is a learned bias, never a guaranteed continuation.

Where gains come from

  • Frequent phrases and syntax patterns become reusable local vectors instead of being reconstructed entirely by dense layers.
  • Large lookup tables add parameter capacity while each token reads only a few rows, so added FLOPs can stay small.
  • Deterministic addressing is cheap and can be prefetched; useful vectors reshape later logits through the normal backbone.

Where the lookup is useless

  • Novel combinations and rare or unseen n-grams have no well-trained local row.
  • Long-range dependencies, semantic reasoning, and facts outside the local window still require the backbone.
  • Hash collisions can mix unrelated patterns; bandwidth, cache locality, host transfers, and batching can erase latency gains.
  • Extra capacity does not guarantee downstream accuracy: an unhelpful vector may do nothing or add noise.

Inspect keys and values

A compact view of how local token sequences address learned rows.

N-gram order

The hash and vector values here are illustrative; production systems choose their own addressing and collision strategy.

PositionLocal keyIllustrative slotLearned value vector
t=1[the · quick]746,393[1.0, 1.3, 0.9, 0.7, …]
t=2[quick · brown]1,712,346[1.1, -0.9, -0.7, 0.5, …]
t=3[brown · fox]18,478,650[1.1, 0.2, 1.4, 0.4, …]

Real-world case study: Qwen3.8-Flash-Next

Qwen uses bigram and trigram tables as one component of a published experimental sparse model. System-level benchmark gains cannot be assigned to this component alone.

+51B lookup parameters20M bigram + trigram entriesInjected at Layer 2Host-memory prefetch

These are lookup-capacity parameters, not 51B dense parameters evaluated for every token. Only addressed rows are fetched.

Do not confuse the mechanism

Not n-gram speculative decoding

Speculative decoding drafts and verifies future tokens. N-gram embeddings retrieve a vector for the current representation; they draft nothing.

Not Multi-Token Prediction (MTP)

MTP trains several future offsets. N-gram embeddings are keyed memory for already observed local sequences.