GASBench evaluates AI-generated content detectors on image, video, and audio datasets for Bittensor Subnet 34. It handles dataset downloads, media preprocessing, inference, and scoring.
Requires Python 3.10 or newer. From a checkout of this repository:
pip install -e '.[gpu]'For dataset definitions alone, use pip install -e .; the base package does not
include the benchmark dependencies.
Prepare a model directory containing model_config.yaml, model.py, and
safetensors weights. See the model specification for
configuration examples and the input/output contract.
Start with a debug run:
gasbench run --image-model ./my_model --debug --cache-dir ./cacheUse --video-model or --audio-model for other modalities. Replace --debug
with --full for a full evaluation.
The CLI saves each run under results/ with:
results.json: scores, timing, and dataset breakdowns.records.parquet: per-sample predictions.summary.txt: a readable report.
See scoring to interpret the results, or
gasbench run --help for all options. The running guide
covers dataset selection, caching, resuming runs, and robustness evaluation.
import asyncio
from gasbench import run_benchmark, print_benchmark_summary, save_results_to_json
results = asyncio.run(run_benchmark(
model_path="./my_model",
modality="image", # "image", "video", or "audio"
mode="debug",
cache_dir="./cache",
))
print_benchmark_summary(results)
save_results_to_json(results, output_dir="./results")- Model specification: package a model for evaluation.
- Classification and scoring: class labels, metrics, and score weighting.
- Running benchmarks: configure and manage evaluations.
- Discriminative mining guide: submit a model to Subnet 34.
- Releases: changes between versions.