Research Preview
Nested Learning was presented by Google Research at NeurIPS 2025. The article discusses that research and the HOPE experiments. The illustrations are conceptual; they do not establish universal retention or deployment readiness.
The Forgetting Problem
Learning a new skill can interfere with an old one when both rely on shared mechanisms. In neural networks, sequential optimization can similarly change parameters useful for earlier tasks. How much interference occurs depends on the tasks and training setup.
“Can a system adapt to new information while preserving useful earlier behavior?”
Nested Learning studies interacting optimization processes with different update frequencies. The human-learning analogy is an intuition, not evidence that its mechanisms reproduce a biological brain.
See It Happen: Catastrophic Forgetting
Changing shared parameters can improve a new objective while worsening an earlier one. The scalar example below makes that conflict explicit and computes both losses. It is not a simulation of HOPE or a claim that every model must forget.
Two conflicting tasks, one parameter
Actual gradient updates in a toy model: task A wants w = 1, task B wants w = −1. Loss = ½(w − target)², learning rate = 0.25. This explains interference in one shared parameter, but models neither an LLM nor HOPE.
w = 0.000 · Step: 0
Loss A: 0.5000
Loss B: 0.5000
The proposed framework
Describe a model and its optimizers as interacting learning processes that update at different frequencies. Those levels can have their own context and objectives. Choosing frequencies alone does not guarantee that interference disappears.
🐢 Slow
Parameters updated less frequently
🚶 Medium
An intermediate update schedule
⚡ Fast
State updated more frequently
Different update schedules can isolate some changes in different state or parameter groups. The schedule alone does not prove that those groups cannot interfere; that requires a defined objective and evaluation.
Watch the Loops in Action
Conceptual animation of three update frequencies. Rotation and timing are chosen for visibility; they do not measure training speed, knowledge or a specific HOPE configuration.
Nested Learning Loops
Evaluate continual learning
A performance comparison needs a defined benchmark and matched conditions. The following describes an evaluation protocol, not fabricated percentages for traditional and nested learning.
What a useful comparison must measure
The HOPE work reports experiments on language modeling, long context and continual learning. It does not establish a universal percentage of retained knowledge. A comparison needs matched data, task sequences and compute budgets, with earlier tasks evaluated again after each update.
- Measure A before and after training on B.
- Report performance on the new task and its cost as well.
- Document update frequencies and memory budget.
- Compare multiple task orders and random seeds.
The Hope Architecture
HOPE is the authors’ proof-of-concept architecture, combining self-modifying recurrent memory with a continuum of memory update frequencies. The reported experiments support particular configurations and tasks.
How Hope Works
Why This Matters
The framework raises research questions about adaptation and memory:
Always Learning
Can a model incorporate new information while retaining earlier task performance?
More Efficient
Which update schedules work within a measured compute and memory budget?
Brain-Like
Different timescales offer a modeling perspective; similarity to brains needs separate evidence.
Key Takeaways
- 1Sequential training can interfere with earlier tasks; measure that effect rather than assuming it.
- 2Nested Learning describes interacting learning processes with different update frequencies.
- 3HOPE is a research architecture evaluated on specified tasks; its results do not establish universal retention.
- 4Keep conceptual diagrams, numerical toy models and experimental benchmark results visibly distinct.