| Rank | Model | Score |
|---|---|---|
| 1 | grok-4-6 | 1577 |
| 2 | kimi-k3 | 1548 |
1 phaseActive
Artificial Analysis's Elo-rated benchmark for AI agent performance on professional knowledge-work deliverables. Metric: Elo rating.
Quick answer: AA-Briefcase is an Elo-rated benchmark from Artificial Analysis that evaluates AI agents on professional knowledge-work deliverables — the kind of documents, analyses, and reports a professional would carry in a briefcase. Human or model raters compare pairwise outputs to derive a relative Elo rating. Kimi K3 scores 1548 Elo as of July 2026.
What it tests: The quality of AI-generated professional knowledge-work outputs (reports, analyses, documents), rated via pairwise comparison rather than binary pass/fail.
Why it matters: Many knowledge-work tasks don't have a single "correct" answer, making Elo-style pairwise comparison a better fit than accuracy scoring. AA-Briefcase, from independent evaluator Artificial Analysis, provides a comparable cross-model quality signal.
Known limitations: Elo scores are relative to the pool of models being compared and the specific rating methodology, so they should not be interpreted as absolute accuracy percentages.
AA-Briefcase evaluates the quality of professional knowledge-work deliverables generated by AI models — such as written reports, analyses, and business documents — using pairwise comparison judged by human or model raters. Results are aggregated into an Elo rating, similar to the methodology Artificial Analysis uses across its broader intelligence benchmarking suite (e.g., GDPval-AA).
| Field | Value |
|---|---|
| Task category | Agent / knowledge work |
| Metric | Elo rating |
| Saturation | Low |
| Created by | Artificial Analysis |
| Website | artificialanalysis.ai |
Model outputs on professional knowledge-work tasks are compared pairwise against outputs from other models, and results are aggregated into a relative Elo rating — higher Elo indicates outputs that are more consistently preferred in head-to-head comparison.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 1548 Elo | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run AA-Briefcase.
| Benchmark | What it tests | Metric |
|---|---|---|
| AA-Briefcase | Professional knowledge-work deliverable quality | Elo |
| GDPval-AA v2 | Economically valuable task performance | Elo |
| Agents' Last Exam | Economically valuable long-horizon agent tasks | Pass rate |
| Harvey Lab-AA | Legal knowledge-work quality | % accuracy |
Benchgen lets you track AA-Briefcase-style Elo comparisons for your own model's knowledge-work outputs over time, alongside your other benchmark results.