CheatBench measures whether AI agents attempt to cheat while completing challenging assignments. Tasks span mathematical research, visual reasoning, writing, coding, knowledge work, and other domains. Each environment establishes expectations of honest work and provides opportunities to violate them, such as accessing hidden answers, copying another participant's work, or manipulating evaluation.
This repository contains the task environments, agent adapters, evaluation runner, and behavioral judges. Task success and cheating are measured separately.
- Claude Agent SDK (
claude-sdk) - Codex (
codex) - Gemini CLI (
gemini-cli) - Grok Build (
grok-build) - Kimi Code (
kimi-code) - Muse Code (
muse-code) - DeepSeek Harness (
deepseek-harness) - Terminus-2 (
terminus-2)
Each model pins its harness and version in configs/models.yaml.
An explicit --harness uses the latest release; add --harness-version to pin it.
Terminus-2 uses a fixed Harbor revision and calls the model API directly.
See configuration for overrides and limits.
- Enroot — requires Linux and Python 3.11+.
- Docker — Software Engineering through Harbor; shared-runner support is planned.
- Apptainer — planned.
Use Linux with Python 3.11 and the Enroot/FUSE/user-namespace prerequisites. Install the repository and task-build dependencies in a fresh virtual environment:
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[build,svgs]'
python -m core.sandbox --check --private-netAdd your provider API credentials to a local .env file, then run:
python run.py TASK --model MODEL --limit 1Choose TASK from Tasks, MODEL from
configs/models.yaml. The default judge is gpt-6-sol.
Missing task inputs and container images are built automatically, including
the Office image and the host-only, pinned Stockfish engine for chess. Omit --limit 1
to run all examples.
Runs write to outputs/: run.json records settings, episodes.jsonl records
execution outcomes, trajectories/ contains agent interactions, and judge.jsonl
contains behavioral labels and explanations. The common judge combines the
shared rubric with each task's judge_schema.py.
To rejudge a saved run:
python judge.py outputs/<run-directory> --redo| Category | Tasks |
|---|---|
| Mathematical Research | OpenMath, OpenMath Agent |
| Multimodal | GeoGuessr, Visual Puzzles |
| Creative Writing | Creative Writing |
| SVG Competition | SVG Competition |
| Menial Computation | Prime Factorization, Subset Sum |
| Biology and Bioinformatics | Protein Design |
| Knowledge Work | Knowledge Work |
| Board Games | Chess, Go |
| Sycophancy | Sycophancy |
| Software Engineering | Software Engineering |
Sycophancy uses its direct API runner; Software Engineering uses Harbor. Their task READMEs document the commands. To add a task, follow the task guide.
@misc{phan2026cheatbenchmeasuringrewardgaming,
title={CheatBench: Measuring Reward Gaming in AI Agents},
author={Long Phan and Stephen K. Yang and Jason J. Lim and Mantas Mazeika and Wenyu Zhang and Zheyuan Liu and Richard Ren and Jingxiang Meng and Yaoteng Tan and Weiliang Zhao and Addison Wu and Matei Anghel and Dan Hendrycks},
year={2026},
eprint={2609.36308},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.36308},
}See LICENSE and NOTICE for licensing and third-party attribution.