Shortcut or shelter?
The policy boundary in a noisy maze
An exposed route is shorter. A protected route is slower. This experiment solves for the reliability at which the preference changes, then tests whether three TD-control methods recover the boundary implied by the objective they learn.
Q-learning · SARSA · Expected SARSA