Skip to content

About

cheatbench.ai

Resources

Stars

16 stars

Watchers

6 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

CheatBench

Website | Paper

CheatBench measures whether AI agents attempt to cheat while completing challenging assignments. Tasks span mathematical research, visual reasoning, writing, coding, knowledge work, and other domains. Each environment establishes expectations of honest work and provides opportunities to violate them, such as accessing hidden answers, copying another participant's work, or manipulating evaluation.

This repository contains the task environments, agent adapters, evaluation runner, and behavioral judges. Task success and cheating are measured separately.

Agent harness support

  • Claude Agent SDK (claude-sdk)
  • Codex (codex)
  • Gemini CLI (gemini-cli)
  • Grok Build (grok-build)
  • Kimi Code (kimi-code)
  • Muse Code (muse-code)
  • DeepSeek Harness (deepseek-harness)
  • Terminus-2 (terminus-2)

Each model pins its harness and version in configs/models.yaml. An explicit --harness uses the latest release; add --harness-version to pin it. Terminus-2 uses a fixed Harbor revision and calls the model API directly. See configuration for overrides and limits.

Container support

  • Enroot — requires Linux and Python 3.11+.
  • Docker — Software Engineering through Harbor; shared-runner support is planned.
  • Apptainer — planned.

Quick start

Use Linux with Python 3.11 and the Enroot/FUSE/user-namespace prerequisites. Install the repository and task-build dependencies in a fresh virtual environment:

python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[build,svgs]'
python -m core.sandbox --check --private-net

Add your provider API credentials to a local .env file, then run:

python run.py TASK --model MODEL --limit 1

Choose TASK from Tasks, MODEL from configs/models.yaml. The default judge is gpt-6-sol. Missing task inputs and container images are built automatically, including the Office image and the host-only, pinned Stockfish engine for chess. Omit --limit 1 to run all examples.

Outputs and judging

Runs write to outputs/: run.json records settings, episodes.jsonl records execution outcomes, trajectories/ contains agent interactions, and judge.jsonl contains behavioral labels and explanations. The common judge combines the shared rubric with each task's judge_schema.py.

To rejudge a saved run:

python judge.py outputs/<run-directory> --redo

Tasks

Category Tasks
Mathematical Research OpenMath, OpenMath Agent
Multimodal GeoGuessr, Visual Puzzles
Creative Writing Creative Writing
SVG Competition SVG Competition
Menial Computation Prime Factorization, Subset Sum
Biology and Bioinformatics Protein Design
Knowledge Work Knowledge Work
Board Games Chess, Go
Sycophancy Sycophancy
Software Engineering Software Engineering

Sycophancy uses its direct API runner; Software Engineering uses Harbor. Their task READMEs document the commands. To add a task, follow the task guide.

Citation

@misc{phan2026cheatbenchmeasuringrewardgaming,
      title={CheatBench: Measuring Reward Gaming in AI Agents}, 
      author={Long Phan and Stephen K. Yang and Jason J. Lim and Mantas Mazeika and Wenyu Zhang and Zheyuan Liu and Richard Ren and Jingxiang Meng and Yaoteng Tan and Weiliang Zhao and Addison Wu and Matei Anghel and Dan Hendrycks},
      year={2026},
      eprint={2609.36308},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.36308}, 
}

License

See LICENSE and NOTICE for licensing and third-party attribution.

About

cheatbench.ai

Resources

Stars

16 stars

Watchers

6 watching

Forks

Releases

Packages

Contributors

Languages