Markov Decision Processes
How to plan when actions do not reliably do what you intend: states, transition models, rewards and discounting, the Bellman equation, and value iteration worked numerically to its fixed point.
Value iteration, policy iteration, and Q-learning on the same gridworld, so a planner that knows the model can be compared directly against a learner that does not.
A small Markov decision process used as a controlled setting for comparing planning against learning. Value iteration and policy iteration are run first, with full knowledge of the transition model, to compute the optimal value function and policy exactly. Q-learning is then run on the same world with no model at all, learning only from sampled transitions, and its estimates are compared against the planned ground truth - which is what makes the comparison informative: the learner recovers the same optimal policy while its value estimates remain slightly off, and the residual gap is sampling error rather than a bug. The lab also exposes the parameters that actually govern behaviour: the discount factor, the learning rate, and the exploration schedule, with the sweep over epsilon showing why a purely greedy learner can settle on a worse policy. Implementation is in progress and no source repository has been published yet.
How to plan when actions do not reliably do what you intend: states, transition models, rewards and discounting, the Bellman equation, and value iteration worked numerically to its fixed point.
Learning to act well without a model of the world: temporal-difference updates, the Q-learning rule, exploration versus exploitation, and a run that recovers the planned optimum from experience alone.