Benchgen

AA-Briefcase — Results

RankModelScore
1grok-4-61577
2kimi-k31548
A

AA-Briefcase

1 phaseActive

Artificial Analysis's Elo-rated benchmark for AI agent performance on professional knowledge-work deliverables. Metric: Elo rating.

Overview

AA-Briefcase

Category Metric Saturation Created

Website

Quick answer: AA-Briefcase is an Elo-rated benchmark from Artificial Analysis that evaluates AI agents on professional knowledge-work deliverables — the kind of documents, analyses, and reports a professional would carry in a briefcase. Human or model raters compare pairwise outputs to derive a relative Elo rating. Kimi K3 scores 1548 Elo as of July 2026.

At a Glance

What it tests: The quality of AI-generated professional knowledge-work outputs (reports, analyses, documents), rated via pairwise comparison rather than binary pass/fail.

Why it matters: Many knowledge-work tasks don't have a single "correct" answer, making Elo-style pairwise comparison a better fit than accuracy scoring. AA-Briefcase, from independent evaluator Artificial Analysis, provides a comparable cross-model quality signal.

Known limitations: Elo scores are relative to the pool of models being compared and the specific rating methodology, so they should not be interpreted as absolute accuracy percentages.

What AA-Briefcase Measures

AA-Briefcase evaluates the quality of professional knowledge-work deliverables generated by AI models — such as written reports, analyses, and business documents — using pairwise comparison judged by human or model raters. Results are aggregated into an Elo rating, similar to the methodology Artificial Analysis uses across its broader intelligence benchmarking suite (e.g., GDPval-AA).

Benchmark Specifications

FieldValue
Task categoryAgent / knowledge work
MetricElo rating
SaturationLow
Created byArtificial Analysis
Websiteartificialanalysis.ai

How AA-Briefcase Is Scored

Model outputs on professional knowledge-work tasks are compared pairwise against outputs from other models, and results are aggregated into a relative Elo rating — higher Elo indicates outputs that are more consistently preferred in head-to-head comparison.

State-of-the-Art Results

RankModelScoreSourceDate
1Kimi K31548 EloKimi K3 technical report2026-07

Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.

AA-Briefcase on Benchgen

No Benchgen results yet — be the first to run AA-Briefcase.

AA-Briefcase vs Other Benchmarks

BenchmarkWhat it testsMetric
AA-BriefcaseProfessional knowledge-work deliverable qualityElo
GDPval-AA v2Economically valuable task performanceElo
Agents' Last ExamEconomically valuable long-horizon agent tasksPass rate
Harvey Lab-AALegal knowledge-work quality% accuracy

Run AA-Briefcase on Your Model

Benchgen lets you track AA-Briefcase-style Elo comparisons for your own model's knowledge-work outputs over time, alongside your other benchmark results.

Frequently Asked Questions

What is AA-Briefcase? AA-Briefcase is an Elo-rated benchmark from Artificial Analysis that evaluates the quality of AI-generated professional knowledge-work deliverables via pairwise comparison.
What does a good score look like on AA-Briefcase? Kimi K3 reports 1548 Elo as of July 2026. As with any Elo system, scores are meaningful relative to other models in the same rating pool rather than as an absolute percentage.
Who created AA-Briefcase? AA-Briefcase is created and maintained by Artificial Analysis, an independent AI model evaluation and analysis platform.