| Rank | Model | Avg Reward |
|---|---|---|
| 1 | gpt2uat | 0.6865 |
| 2 | gemma4-gemma-4-E2B-it_1780572325819 | 0.5161 |
| 3 | tesst36-bench-3-qwen2.5-0.5b-instruct-mpbbin2v | 0.3375 |
1 phaseActive
Submit your agent configuration. The benchmark container boots OBP, seeds the bank, and runs one RL step (dry_run) across fraud scenarios. Results are scored on avg_reward, accuracy, and tool efficiency.