Skip to content
Kudos AI

Belief State

The probability distribution an agent holds over the states it might be in, given everything it has done and perceived - the thing it can act on when the state itself is hidden.

Also known as: Belief distribution, Information state

Understanding Belief State

When an agent cannot observe which state it occupies, the natural object to reason about is not the state but the distribution over states consistent with its history. That distribution is the belief state. It is a sufficient statistic for the entire past: given the belief, no earlier action or percept adds anything to the prediction of the future, which is what allows planning to proceed from the belief alone rather than from an ever-growing history.

Maintaining it is one recursion. Given a belief, an action and a resulting percept, the new belief is obtained by pushing the old one through the transition model, multiplying pointwise by the probability of the percept in each candidate successor state, and renormalising. This is exactly the forward step of hidden Markov model filtering; the only addition is that the transition model now depends on the action the agent chose.

The crucial structural fact is that the belief state is observable to the agent by definition - it is a summary of the agent's own history, not of the world. That is what makes it usable as a state variable. A policy that says "if you are in state 3, do this" cannot be executed by an agent that does not know whether it is in state 3; a policy over beliefs always can be.

Belief states appear well beyond probabilistic planning. In sensorless and contingency search the belief is a set of possible states rather than a distribution, and search proceeds over those sets. In temporal models the filtered distribution over the current state is a belief state maintained by the same recursion. And in partially observable MDPs the belief space becomes the state space of an equivalent, fully observable problem.

How to Calculate

b'(s') = α · P(e | s') Σ_s P(s' | s, a) b(s)

where

b'(s')
the updated probability of being in state s′ after acting and perceiving
P(e | s′)
the sensor model: how likely this percept is in the candidate state
P(s' | s, a)
the transition model for the action actually taken
α
the normalising constant that makes the new belief sum to one

Example of Belief State

Take a two-state world in which one action persists with probability 0.9, the other switches with probability 0.9, and the sensor reports the correct state with probability 0.6. From an even belief, either action leaves the prediction even - one persists and the other switches, which from a 50/50 split is the same thing - so the percept does all the work, and a reading of "state 1" moves the belief to exactly 3/5.

Confidence is harder to accumulate than to lose. A belief of 0.99 taking the persist action and then a single contradicting percept collapses to 446/527 = 0.846300, most of that the transition leak, which alone takes 0.99 to a prediction of 0.892; eight consecutive confirming percepts starting from 0.5 raise it only to 0.817218.

It never reaches one. Repeating the persist action with the same percept forever converges to (3 + √105)/16 = 0.827934, the point at which the evidence pulling the belief outward exactly balances the transition noise pushing it back toward the middle. The agent must therefore plan while still uncertain, which is the whole difficulty of partial observability.

Frequently Asked Questions

How does it differ from the state of a hidden Markov model?

It is the same object - the filtered distribution over the hidden state - with actions added. In an HMM the transitions happen to the agent; in a belief-state formulation the agent chooses which transition model applies, and so influences what it will learn as well as where it will go.

Does the belief always converge?

Not to certainty, and not necessarily to a single point. With a noisy sensor and stochastic transitions it settles into a region determined by the balance between the two, and with informative-enough sensors and slow-enough dynamics that region can be tight. Assuming convergence to the truth is a common and expensive mistake.

Why is it a sufficient statistic?

Because the transition and sensor models are Markov: the next state depends only on the current state and action, and the percept only on the current state. Everything the history has to say about the future is therefore already contained in the current distribution over states.

The Bottom Line

A belief state is the agent's distribution over where it might be, maintained by the filtering recursion and always available to it. Planning over beliefs rather than states is what makes partially observable problems tractable in principle - and the fact that beliefs form a continuum is what makes them hard in practice.