Reinforcement Learning
Learning from consequences rather than from labels. Markov decision processes, value and policy methods, and the exploration problem at the heart of it.
Learning paths (1)
Encyclopedia (2)
Markov Decision Process
A formal model of sequential decision-making in which outcomes are partly random, defined by states, actions, transition probabilities, and rewards.
Q-Learning
A reinforcement learning algorithm that learns the value of taking each action in each state directly from experience, without a model of the environment.
Articles (2)
Reinforcement Learning and Q-Learning
Learning to act well without a model of the world: temporal-difference updates, the Q-learning rule, exploration versus exploitation, and a run that recovers the planned optimum from experience alone.
Markov Decision Processes
How to plan when actions do not reliably do what you intend: states, transition models, rewards and discounting, the Bellman equation, and value iteration worked numerically to its fixed point.