This project trains a bipedal robot to walk using the Truncated Quartile Critics (TQC) reinforcement learning method. It includes training runs for both the standard BipedalWalker environment and the more difficult BipedalWalker Hardcore variant.
- Code and training notebook: ./agent-code.ipynb
- Normal environment logs and video: ./agent-log.txt, ./agent-video,episode=1500,score=333.mp4
- Hardcore environment logs and video: ./agent-hardcore-log.txt, ./agent-hardcore-video,episode=1500,score=313.mp4
- Paper: ./learning-to-walk-paper.pdf