Reinforcement Learning
Acting well when outcomes are uncertain: the Bellman equation and how to solve it, then what changes when the environment is unknown and the agent has to learn from experience alone.
Sign in to take quizzes, earn XP, and unlock stages as you reach 90% mastery.
MDPs and the Bellman Equation
30 min · 130 XPThe five parts of a Markov decision process, why a policy beats a plan under uncertainty, and the self-consistency condition every utility must satisfy.
Open lesson →Sign in to take the 3-question quiz.
Value and Policy Iteration
35 min · 150 XPTurning the Bellman equation into an assignment and sweeping until it settles, then the alternating evaluate-and-improve loop that usually finds the policy first.
Open lesson →Sign in to take the 3-question quiz.
Q-Learning and Exploration
35 min · 160 XPLearning to act well with no model of the environment: the temporal-difference update, value propagating backwards, and why a purely greedy agent gets stuck.
Open lesson →Sign in to take the 3-question quiz.