Learning from consequences
Reinforcement learning trains an agent through interaction. Instead of being shown the correct answer for every example, the agent tries actions, receives rewards, and gradually learns a policy that produces better long-term outcomes.
Q-learning gridworld
The agent updates Q-values from rewards. Arrows show the currently preferred action. Changing exploration affects the data collected, so compare several runs.
High ε samples more actions. Low ε follows the learned Q-values more often.
The RL loop
Agent
The learner or decision-maker.
Environment
Receives actions and returns observations and rewards.
State / observation
What the agent can observe; it may be only part of the underlying state.
Action
A choice available to the agent.
Reward
A scalar training signal, not a complete specification of the intended task.
Policy
A strategy mapping observations to actions.
Major method families
Q-learning
Updates estimates of expected return for state-action pairs.
Policy gradients
Adjust policy parameters using sampled returns or advantage estimates.
Actor-critic
Combines a policy with a learned value estimate.
PPO / GRPO
Policy-optimization families with different estimators and update constraints. A gridworld Q-learning demo is not their implementation.
Why RL matters for modern AI
RL is the conceptual bridge between passive prediction and agentic behavior. Robots, game-playing systems, recommender systems, tool-using agents, and LLM post-training all use the same idea: optimize choices based on feedback from the world or a reward model.
In LLMs, RLHF and newer verifiable-reward training methods use RL-style optimization to make models more helpful, safer, or better at reasoning. The model is not merely imitating text; it is being pushed toward behavior that earns higher reward.