Skip to content
Kudos AI

Tagged “reward-design”

1 article.

5 min readReinforcement Learning

The Parameter Nobody Chooses

The living reward in a grid world is written down once and never discussed, and the optimal policy is a step function of it: eight thresholds between -3 and 0, each flipping exactly one square. The textbook value of -0.04 sits 0.0048 away from the one that decides whether the agent takes the shortcut past the pit, and above -0.0221, when steps are nearly free, the optimal move in one corner is to walk into a wall on purpose.

Artificial Intelligence