Skip to content
Kudos AI

Reinforcement Learning

Acting well when outcomes are uncertain: the Bellman equation and how to solve it, then what changes when the environment is unknown and the agent has to learn from experience alone.

Advanced440 XP~2 h90% to advance

Sign in to take quizzes, earn XP, and unlock stages as you reach 90% mastery.

  1. MDPs and the Bellman Equation

    30 min · 130 XP

    The five parts of a Markov decision process, why a policy beats a plan under uncertainty, and the self-consistency condition every utility must satisfy.

    Open lesson →

    Sign in to take the 3-question quiz.

  2. Value and Policy Iteration

    35 min · 150 XP

    Turning the Bellman equation into an assignment and sweeping until it settles, then the alternating evaluate-and-improve loop that usually finds the policy first.

    Open lesson →

    Sign in to take the 3-question quiz.

  3. Q-Learning and Exploration

    35 min · 160 XP

    Learning to act well with no model of the environment: the temporal-difference update, value propagating backwards, and why a purely greedy agent gets stuck.

    Open lesson →

    Sign in to take the 3-question quiz.

Complete all modules → earn the Reinforcement Learning badge