Master AI Concepts
Through Experience

Explore artificial intelligence and large language models through beautiful, interactive demonstrations. Learn by doing, not just reading.

Topics
68
Interactive guides to understand AI concepts
Recently Updated
Mega-Kernels
Sep 24, 2026
LLM
Architecture
MTP · MoE · KV cache · inference

LMF Blog

Latest AI news, model analysis, and deep dives

Interactive Demos

Visual Learning

Build Intuition

Recently Updated

What's New

Help Make This Better

This guide is open source. Got an idea for a new topic? Found a bug? Want to improve an explanation? Every contribution helps.

Explore Topics

Dive into interactive lessons

Artificial Intelligence

Large Language Models
Architecture
01
Transformer ArchitectureThe foundational architecture behind GPT, BERT, and all modern LLMsSep 13, 2026
I
02
Looped TransformersExplore how recurrent depth updates a hidden state, reuses model weights, and changes the tradeoffs between memory, compute, latency, and monitorability.Sep 7, 2026
E
03
Feed-Forward NetworksHow the same MLP transforms each token position; routing is covered in Mixture of Experts.Sep 13, 2026
I
04
Residual Stream & LayerNormHow residual streams and normalization keep deep transformer stacks trainable.Sep 13, 2026
I
05
Next-Token PredictionHow logits, softmax, and decoding turn model outputs into text one token at a time.Sep 13, 2026
B
06
LLM TrainingFrom pretraining on raw text to fine-tuning with human feedbackSep 13, 2026
I
07
Training DataWhere AI models get their knowledge: legitimate sources, controversies, and synthetic dataSep 13, 2026
I
08
Mixture of ExpertsRun only a fraction of model parameters per tokenSep 13, 2026
E
09
QuantizationShrink model size by reducing numerical precisionSep 13, 2026
E
10
Nested LearningLearning algorithms that operate at multiple levelsSep 13, 2026
E
11
Multi-Token Prediction (MTP)Train models to predict several future tokens at once instead of only the next tokenSep 13, 2026
E
12
N-gram EmbeddingsLearned lookup tables for short local token patternsAug 27, 2026
E
13
DistillationTransfer knowledge from a large model to a smaller oneSep 13, 2026
I
14
Fine-Tuning & LoRAAdapt large models efficiently by training only tiny low-rank matricesSep 13, 2026
I
15
AbliterationRemove a model's refusal behavior by ablating a single direction in its residual streamSep 13, 2026
E
16
Speculative DecodingSpeed up inference by drafting tokens with a smaller modelSep 13, 2026
E
Pro tip: PressCtrl+Kto search topics