IRL LAB / FIELD NOTE 001EST. 2026
Small worlds.
Hard questions.
A personal laboratory for rebuilding reinforcement learning from first principles—one controlled environment, exact baseline, and honest uncertainty statement at a time.
Study 01 is runningExact analysis complete · empirical conclusions withheld
01 / CURRENT STUDY
Shortcut
or shelter?
A noisy corridor is fast. A southern route is longer but protected. Where does the optimal route switch—and does a learning algorithm target the same boundary when exploration continues?
Q-learningSARSAExpected SARSA
Read the live research scaffold Known dynamicsTwo consequence lawsPaired seed panel
THE LAB CONTRACT
Mechanisms before leaderboards.
- 01Solve what can be solved.
Known finite environments get exact dynamic-programming baselines before training begins.
- 02Match the comparison.
Methods share interaction budgets, seed panels, and held-out evaluation conditions.
- 03Separate design from evidence.
Calibration shapes the protocol. Full-run artifacts support the published result.
- 04Show what would change the conclusion.
Unresolved actions, instability, and inconvenient outcomes remain visible.
REPRODUCIBILITY LEDGER
Evidence has a paper trail.
StudyShortcut or Shelter?
Exact modelComplete
Full artifactsPending validation
Public resultWithheld