31 Topics
Large Language Models
Understand the inner workings of large language models, from tokenization to attention mechanisms.
Help Make This Better
This guide is open source. Got an idea for a new topic? Found a bug? Want to improve an explanation? Every contribution helps.
Fundamentals
01
TokenizationHow text gets broken into pieces the model understandsSep 13, 2026
B02
Subtoken BlindnessWhy models miss letters, digits, and exact counts hidden inside tokensSep 13, 2026
I03
EmbeddingsTurning words into numbers that capture meaningSep 13, 2026
B04
Positional Encoding & RoPEHow LLMs represent token order with positional encodings, RoPE, and long-context tradeoffs.Jul 12, 2026
E05
Attention MechanismHow transformers decide which tokens matter for each predictionSep 13, 2026
I06
Multi-Head Attention, MQA & GQAHow attention heads run learned projections in parallel, and how MQA/GQA share key-value heads.Jun 12, 2026
IBehavior
01
TemperatureHow logits become sampling probabilities.Sep 13, 2026
B02
Reasoning Models & Inference-Time ComputeHow models spend extra runtime compute on search, candidate generation, and verification — and when it pays off.Jul 12, 2026
E03
Jagged FrontierWhy frontier models solve hard benchmarks and still fail simple-looking tasksSep 13, 2026
I04
Context RotEvidence on context length, position and distractors, with clear experimental limits.Sep 13, 2026
ICapabilities
01
RAGAugment LLM answers with retrieved external knowledgeSep 13, 2026
I02
Vision & ImagesHow LLMs process and understand images alongside textSep 13, 2026
B03
Visual ChallengesWhere vision models fail: counting, spatial reasoning, OCRSep 13, 2026
I04
Agentic VisionActive visual investigation through zoom, crop, and code executionSep 13, 2026
E05
MultimodalityProcessing text, images, audio, and video in one modelSep 13, 2026
BArchitecture
01
Transformer ArchitectureThe foundational architecture behind GPT, BERT, and all modern LLMsSep 13, 2026
I02
Looped TransformersExplore how recurrent depth updates a hidden state, reuses model weights, and changes the tradeoffs between memory, compute, latency, and monitorability.Sep 7, 2026
E03
Feed-Forward NetworksHow the same MLP transforms each token position; routing is covered in Mixture of Experts.Sep 13, 2026
I04
Residual Stream & LayerNormHow residual streams and normalization keep deep transformer stacks trainable.Sep 13, 2026
I05
Next-Token PredictionHow logits, softmax, and decoding turn model outputs into text one token at a time.Sep 13, 2026
B06
LLM TrainingFrom pretraining on raw text to fine-tuning with human feedbackSep 13, 2026
I07
Training DataWhere AI models get their knowledge: legitimate sources, controversies, and synthetic dataSep 13, 2026
I08
Mixture of ExpertsRun only a fraction of model parameters per tokenSep 13, 2026
E09
QuantizationShrink model size by reducing numerical precisionSep 13, 2026
E10
Nested LearningLearning algorithms that operate at multiple levelsSep 13, 2026
E11
Multi-Token Prediction (MTP)Train models to predict several future tokens at once instead of only the next tokenSep 13, 2026
E12
N-gram EmbeddingsLearned lookup tables for short local token patternsAug 27, 2026
E13
DistillationTransfer knowledge from a large model to a smaller oneSep 13, 2026
I14
Fine-Tuning & LoRAAdapt large models efficiently by training only tiny low-rank matricesSep 13, 2026
I15
AbliterationRemove a model's refusal behavior by ablating a single direction in its residual streamSep 13, 2026
E16
Speculative DecodingSpeed up inference by drafting tokens with a smaller modelSep 13, 2026
E