Course material for a deep reinforcement learning module, kept with the work I did on top of the provided harness.
drl_sample_project_python/drl_lib/to_do/— the graded part: dynamic programming, Monte-Carlo methods and temporal-difference learning over grid-world and line-world MDPs, driven frommain.py.drl_sample_project_python/drl_lib/do_not_touch/— the instructor-provided environment wrappers and result structures.drl_contracts/— the Rust trait definitions describing the environment API the harness expects.drl_sample_project/— the reference Rust sample project shipped with the course.
The short version: implement the three methods under to_do/, then run
python drl_sample_project_python/main.py to see each demo print its learning curves.