Decisions Under Partial Observability
An agent that cannot see which state it is in has to act on a distribution instead. That distribution is itself always observable, which turns the problem back into an MDP - over a continuous space, on which the exact algorithms do not close.
Sign in to take quizzes, earn XP, and unlock stages as you reach 90% mastery.
Belief States and the Update
25 min · 100 XPReplacing the unknown state with a distribution over states, the filtering step that maintains it, and the reason a noisy sensor cannot drive that distribution to certainty however long you watch.
Open lesson →Sign in to take the 3-question quiz.
POMDPs as Belief-State MDPs
30 min · 120 XPThe reduction that turns a partially observable problem into a fully observable one over beliefs, conditional plans as hyperplanes, and the value function that is piecewise linear and convex because it is a maximum over them.
Open lesson →Sign in to take the 3-question quiz.
Why Exact Solution Does Not Scale
30 min · 120 XPThe count of conditional plans, the measured growth of the undominated set under exact value iteration, and the two things practice does instead: discretise the belief space, or sample it online.
Open lesson →Sign in to take the 3-question quiz.