| Rank | Model | Score |
|---|---|---|
| 1 | gpt-5-5 | 0.805 |
1 phaseActive
Real-world bioinformatics benchmark evaluating AI agents on multi-step computational biology workflows requiring code execution and domain knowledge.
Quick answer: BixBench (Mitchener et al., 2025) evaluates AI models on multi-step bioinformatics and computational biology data analysis, requiring code execution, statistical reasoning, and biological domain knowledge to interpret experimental data.
| Property | Value |
|---|---|
| Domain | Bioinformatics & computational biology |
| Metric | Accuracy on multi-step scientific workflows |
| Evaluation | Code execution + domain knowledge scoring |
| Categories | Science, Agents |
Source: Mitchener et al. 2025. Last updated 2026-07-24.