Hello ResearchClawBench maintainers,
We are preparing an official Pass@1 submission for ScholarX, a cloud scientific-research agent, and want to follow the current canonical import process before running the frozen 40-task campaign.
Our planned evaluation baseline is:
- ResearchClawBench commit:
01bc2371f698f755892935ef2965a95a790ff0db
- ResearchClawBench-Home commit used for leaderboard comparison:
5d63ab64b6d22e5669f58b7aa49be433c853b629
- Coverage: all 40 base tasks, one clean run per task for Pass@1
- Agent and judge are isolated: the agent cannot access
target_study, checklist.json, sibling tasks, earlier scored runs, or the scorer
- Each run preserves
_meta.json, _score.json, code, outputs, report, figures, duration, timestamps, model identity, and immutable hashes
- Infrastructure-only reruns are retained and disclosed separately; scientific low scores are not rerun
Could you please confirm the current submission protocol?
- Should a new agent submit results through a PR to
ResearchClawBench-Home, an issue/email handoff to maintainers, or another import channel?
- Is a 40-task leaderboard-only result set sufficient, or do you require complete run directories and reports as well?
- Which exact fields and meanings are required for
_meta.json, _score.json, duration, cost, and timestamp?
- Which judge model/configuration and retry rule should be used for a comparable official Pass@1 submission?
- How should deterministic infrastructure failures and judge failures be represented, and when is one clean rerun permitted?
- Is the agent preset/logo/project metadata submitted in the same change as results, in a separate PR, or added by maintainers during import?
- May we disclose earlier diagnostic runs and then submit one separately frozen 40-task campaign as the official result?
We will not claim an official rank until the result appears in the official ResearchClawBench Home data. Thank you for helping us make the submission directly compatible with your current process.
Hello ResearchClawBench maintainers,
We are preparing an official Pass@1 submission for ScholarX, a cloud scientific-research agent, and want to follow the current canonical import process before running the frozen 40-task campaign.
Our planned evaluation baseline is:
01bc2371f698f755892935ef2965a95a790ff0db5d63ab64b6d22e5669f58b7aa49be433c853b629target_study,checklist.json, sibling tasks, earlier scored runs, or the scorer_meta.json,_score.json, code, outputs, report, figures, duration, timestamps, model identity, and immutable hashesCould you please confirm the current submission protocol?
ResearchClawBench-Home, an issue/email handoff to maintainers, or another import channel?_meta.json,_score.json, duration, cost, and timestamp?We will not claim an official rank until the result appears in the official ResearchClawBench Home data. Thank you for helping us make the submission directly compatible with your current process.