RL training environments with verifiable rewards for coding agents. Works with TRL, Unsloth, verl, OpenRLHF.
-
Updated
Apr 24, 2026 - Python
RL training environments with verifiable rewards for coding agents. Works with TRL, Unsloth, verl, OpenRLHF.
SpaceMining: a novel RL environment beyond LLM priors
Reinforcement learning strategies for AWS DeepRacer — from stable baseline to sub-9 second laps on the re:Invent 2018 track.
Check the checker: an open standard for testing the graders of RL environments, with proof before every accusation.
AWS Deep Racer workflow: reward functions, log analysis, and references for model tuning.
To associate your repository with the reward-function topic, visit your repo's landing page and select "manage topics."