BioProBench is the first large-scale, integrated multi-task benchmark designed specifically for Large Language Models (LLMs) in the life sciences. It moves beyond simple declarative QA to encompass a comprehensive suite of tasks critical for true procedural text comprehension and execution.
Biological protocols are the fundamental bedrock of reproducible and safe life science research. While LLMs have shown remarkable capabilities on general tasks, their systematic evaluation on highly specialized, accuracy-critical, and inherently procedural texts remains limited.
BioProBench fills this gap by providing a robust framework to evaluate LLMs on diverse aspects of protocol understanding and reasoning. It serves as the foundational evaluation environment for downstream execution agents like our BioProAgent.
- 📚 Unprecedented Scale: Grounded in a highly-curated, open-source corpus of 22,413 original biological protocols, yielding 523,784 high-quality structured instances.
- 🎯 Comprehensive Tasks: A suite of 5 core tasks challenging LLMs on temporal dependencies, conditional logic, and generation:
[PQA]Protocol Question Answering[ORD]Step Ordering[ERR]Error Correction[GEN]Protocol Generation[REA]Protocol Reasoning
- 🧬 Broad Domain Coverage: Data sourced from 5 authoritative repositories spanning 16 biological subdomains.
- 🔬 Standardized Evaluation: A robust framework combining standard NLP metrics with novel domain-specific measures (e.g., Step Recall, Step Precision).
💡 Build your own Scientific AI? We've got you covered.
We have encapsulated our evaluation framework into a self-contained, easily deployable agent skill.
Located in skills/evaluate-protocol-outputs/, this module can be directly imported into other AI agent environments (e.g., AutoGen, LangChain, or custom frameworks) independently of this repository, while retaining full BioProBench compatibility.
BioProBench provides a layered data design to support various model development stages. All data files are rigorously formatted in JSON and split into Train/Test sets.
| Task | Description | File Names |
|---|---|---|
| PQA | Question Answering | PQA_train.json, PQA_test.json |
| ORD | Step Ordering | ORD_train.json, ORD_test.json |
| ERR | Error Correction | ERR_train.json, ERR_test.json |
| GEN | Protocol Generation | GEN_train.json, GEN_test.json |
| REA | Protocol Reasoning | REA_train.json, REA_test.json |
| Corpus | Full Raw Corpus | protocols-io.json, Nature-Protocols.json, etc. |
- 📥 Download Access: Hugging Face Repository
To keep the repository clean, we've organized our inference and evaluation scripts logically. Click to expand the instructions below.
1. Running Inference (API or Local Models)
For researchers who wish to reproduce our results or benchmark new models, we provide easy-to-use inference scripts in the Scripts/ directory.
Using an API (e.g., OpenAI, Anthropic, Gemini):
cd Scripts
python generate_response.py
Configuration in generate_response.py:
API_KEY = 'YOUR_API_KEY'
BASE_URL = '[https://api.openai.com/v1](https://api.openai.com/v1)'
MODEL_NAME = 'o3-mini'
TASK_NAME = 'PQA' # Options: 'PQA', 'ORD', 'ERR', 'REA-ERR', 'GEN', 'REA-GEN'Using Local Models (Huggingface):
cd Scripts
python generate_response_local.py
Configuration in generate_response_local.py:
MODEL_NAME = 'meta-llama/Meta-Llama-3-8B-Instruct'
TASK_NAME = 'PQA'
TEST_FILE_PATH = f"../Data/{TASK_NAME.split('-')[-1]}_test.json"2. Evaluation
Each task has a standalone evaluation script in the Metrics/ directory.
| Task | Script | Output Metrics |
|---|---|---|
| GEN | ./Metrics/GEN.py |
BLEU, Keyword-based, Step Recall/Precision |
| PQA | ./Metrics/PQA.py |
Accuracy, Brier Score |
| ERR | ./Metrics/ERR.py |
Accuracy, Precision, Recall, F1 |
| ORD | ./Metrics/ORD.py |
Exact Match, Kendall's tau |
| REA | ./Metrics/REA-ERR.py |
Accuracy, Precision, Recall, Consistency |
Usage Example:
- Open the script (e.g.,
ERR.py) and set your model's response path:
output_file_path = "/absolute/path/to/model_response.json" - Execute the evaluation:
cd Metrics
python ERR.py
After evaluating 12 mainstream open-source and closed-source LLMs (including frontier models), we uncovered critical insights:
- Surface vs. Deep Understanding: Top models perform well on qualitative tasks (e.g., ~74% PQA-Acc.), but struggle drastically with quantitative precision and safety awareness.
- The Generation Bottleneck: Performance plummets on Step Ordering (ORD-EM ~50%) and Protocol Generation (GEN-BLEU <15%), highlighting a profound difficulty in managing temporal dependencies and generating coherent procedures.
- Domain Models Fall Short: Interestingly, smaller bio-specific models often lag behind general frontier LLMs on complex procedural content, suggesting structural reasoning capacity is as vital as domain vocabulary.
While BioProBench diagnoses the cognitive gaps of LLMs, wet-lab environments demand zero-defect physical execution. To bridge the gap from computer simulation to in vitro experiments, we introduce BioProAgent.
By grounding probabilistic LLM reasoning within a deterministic Finite State Machine (FSM) and enforcing a strict "Design-Verify-Rectify" workflow, BioProAgent achieves 95.6% physical compliance and an 88.7% success rate in error recovery.
Explore our related initiatives pushing the boundaries of AI in science:
- 🧪 ChemCoTBench: A step-by-step, application-oriented benchmark evaluating LLM reasoning in chemical applications.
- 🧬 ProLLaMA: A multitask protein language model enhanced by the Evolutionary Protein Generation Framework (EPGF).
We welcome contributions! Whether it's adding new protocol sources, creating novel tasks, or improving annotations, your pull requests are highly appreciated.
For dataset access, collaboration inquiries, or support, please reach out to: 📧 sunshineliuyuyang@gmail.com
If you find our benchmark, datasets, or the overarching BioProProject useful in your research, please consider citing our work:
@inproceedings{liu2026bioprobench,
title={BioProBench: A Corpus and Benchmark for Biological Protocol Reasoning in Autonomous Science},
author={Liu, Yuyang and Lv, Liuzhenghao and Zhang, Xiancheng and Wang, Jingya and Yuan, Li and Tian, Yonghong},
booktitle={Proceedings of the 43rd International Conference on Machine Learning (ICML)},
year={2026}
}
@inproceedings{liu2026bioproagent,
title={BioProAgent: Neuro-Symbolic Grounding for Constrained Scientific Planning},
author={Liu, Yuyang and Wang, Jingya and Lv, Liuzhenghao and Tian, Yonghong},
booktitle={ACL 2026 Oral},
year={2026}
}


