BioAgent Gym is a repository of independently installable biomedical agents and native Harbor leaderboards. Harbor 0.23.0 owns Docker environments, trials, timeouts, verification, and result storage. This repository supplies agents, dataset tasks, scoring code, and native job YAML.
The supported catalog is agents/README.md. Install only the shared item runner and the agent you intend to use:
python3.12 -m venv .venv-harbor
.venv-harbor/bin/pip install -e . -e runner -e 'agents/dswizard[harbor]'
.venv-harbor/bin/pip install -e 'agents/deepevidence[harbor]'Coder, DSWizard, and DeepEvidence are separate distributions. Their wheels own
only their agent namespace; bioagent-harbor-runtime is a separate lightweight
distribution.
The current leaderboard catalog is benchmarks/README.md:
Each benchmark directory contains its tasks, jobs, scoring source, preparation script, and tests. Prepare pinned data once:
.venv-harbor/bin/python benchmarks/biodsbench/prepare.py
.venv-harbor/bin/python benchmarks/biomedicine-deep-research/prepare.pyPreparation writes public items, private verifier references, large external
inputs, and grader deployment copies. It never rewrites committed
task.toml, Dockerfiles, instructions, or job YAML.
Run deterministic smoke jobs:
ROOT="$(pwd)"
.venv-harbor/bin/harbor run \
-c benchmarks/biodsbench/tests/jobs/dswizard-mock.yaml \
--jobs-dir "$ROOT/.harbor/jobs" -y
.venv-harbor/bin/harbor run \
-c benchmarks/biomedicine-deep-research/tests/jobs/deepevidence-mock.yaml \
--jobs-dir "$ROOT/.harbor/jobs" -yRun the formal leaderboard configurations with the required provider secret:
.venv-harbor/bin/harbor run \
-c benchmarks/biodsbench/jobs/dswizard.yaml --env-file .env \
--jobs-dir "$ROOT/.harbor/jobs" -y
.venv-harbor/bin/harbor run \
-c benchmarks/biomedicine-deep-research/jobs/deepevidence.yaml --env-file .env \
--jobs-dir "$ROOT/.harbor/jobs" -y
.venv-harbor/bin/python benchmarks/biomedicine-deep-research/summarize.py \
"$ROOT/.harbor/jobs/<deepevidence-job-directory>"BioDSBench ranks the 118-item Python task; R remains inventoried but unsupported.
Biomedical Deep Research ranks verifier-split choice items by exact match and
evidence-gap retrieval by recall@30. To experiment on fit or tune, copy the
formal YAML into the gitignored local-jobs/ directory and change both the
agent kwargs.split and verifier BIOAGENT_SPLIT; these values must agree.
Use harbor view .harbor/jobs to inspect submissions, item artifacts, verifier
logs, and rewards. See architecture.md,
adding-agents.md, and
adding-benchmarks.md. Historical code is isolated
under legacy/.