Framework for evaluating and improving agents
-
Updated
Oct 7, 2026 - Python
Framework for evaluating and improving agents
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
Turn any repository into verifiable RL environments for coding agents - Harbor tasks you can train on, evaluate and share on the Hugging Face Hub
FineEnvs — RL Environments 101: building and scaling RL environments in the age of LLMs
A Universal Platform for Training and Evaluation of Mobile Interaction
A graphical interface for reinforcement learning and gym-based environments.
Interoperating between (Deep) Reiforcement Learning libraries
Gymnasium-style API standard for RL environment creation in JAX
Workspace manager for coding agents. Interactively solve and develop Harbor tasks.
Awesome Environment Scaling for AI Agents — a periodically updated survey and curated list of papers, projects, RL environments, world models, agent sandboxes, and open research problems
An agentic auditor for RL environments and eval suites. Emits an Environment Card: every claim tied to a probe result, and every probe that could not run, with the reason.
Create new gridworld gym environments easily
Check the checker: an open standard for testing the graders of RL environments, with proof before every accusation.
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
A configurable 3D jump-chess framework with A/B geometries, multi-player rule switches, and RL/self-play training support.
One contract for every RL environment: load any format, judge it your way, improve any model with any learner.
Independent exact certification of machine-generated mathematics — exact arithmetic, no code shared with the claimant, refusal as a verdict.
A lightweight, open-source framework that turns historical GitHub pull requests into reproducible, verifiable software-engineering tasks for training and evaluating coding agents.
To associate your repository with the rl-environments topic, visit your repo's landing page and select "manage topics."