Reinforcement Learning

Intermediate

How agents learn by acting, receiving rewards, and improving a policy over repeated experience.

Last updated: Sep 13, 2026

Learning from consequences

Reinforcement learning trains an agent through interaction. Instead of being shown the correct answer for every example, the agent tries actions, receives rewards, and gradually learns a policy that produces better long-term outcomes.

The hard part is credit assignment: if a reward arrives after 200 steps, which earlier decisions deserve credit or blame? That is why RL is powerful, unstable, and often expensive.

Q-learning gridworld

The agent updates Q-values from rewards. Arrows show the currently preferred action. Changing exploration affects the data collected, so compare several runs.

🤖0.0
↑0.0
↑0.0
↑0.0
↑0.0
↑0.0
■
■
⚠
↑0.0
↑0.0
↑0.0
↑0.0
↑0.0
↑0.0
↑0.0
⚠
↑0.0
■
↑0.0
↑0.0
↑0.0
↑0.0
↑0.0
🏁
Robot: current state · flag: +10 terminal · warning: −8 terminal · square: wall · arrow: best learned action
Exploration ε30%

High ε samples more actions. Low ε follows the learned Q-values more often.

Episodes
0
Steps
0
Return
0
Learning trace
Q-values start at zero. Train episodes and watch the policy change.

The RL loop

1

Agent

The learner or decision-maker.

2

Environment

Receives actions and returns observations and rewards.

3

State / observation

What the agent can observe; it may be only part of the underlying state.

4

Action

A choice available to the agent.

5

Reward

A scalar training signal, not a complete specification of the intended task.

6

Policy

A strategy mapping observations to actions.

Major method families

Q-learning

Updates estimates of expected return for state-action pairs.

Policy gradients

Adjust policy parameters using sampled returns or advantage estimates.

Actor-critic

Combines a policy with a learned value estimate.

PPO / GRPO

Policy-optimization families with different estimators and update constraints. A gridworld Q-learning demo is not their implementation.

Why RL matters for modern AI

RL is the conceptual bridge between passive prediction and agentic behavior. Robots, game-playing systems, recommender systems, tool-using agents, and LLM post-training all use the same idea: optimize choices based on feedback from the world or a reward model.

In LLMs, RLHF and newer verifiable-reward training methods use RL-style optimization to make models more helpful, safer, or better at reasoning. The model is not merely imitating text; it is being pushed toward behavior that earns higher reward.