| Rank | Model | Score |
|---|---|---|
| 1 | gpt-6-astra | 0.84 |
1 phaseActive
Optical music recognition benchmark on the OpenScore String Quartet Corpus — scored as 1 minus normalized edit distance.
Quick answer: OpenScore String Quartets is a benchmark built on the OpenScore String Quartet Corpus (Gotham, Redbond, Bower & Jonas, 2023) that tests optical music recognition (OMR) — a model's ability to read a musical score image and transcribe it accurately. Scores are reported as 1 minus the OMR normalized edit distance (OMR-NED), so higher is better. GPT-6 Astra scores 0.84, versus 0.19 for GPT-5.6 Sol.
What it tests: Whether a multimodal model can visually parse a string quartet score (an image of sheet music) and transcribe its notation accurately — pitches, rhythms, and structure across four instrumental parts.
Why it matters: Optical music recognition is a specialized visual-reasoning task combining fine-grained image understanding with structured symbolic output. Strong performance signals genuine visual precision, not just pattern-matching on common image types — a distinct capability from general OCR or document understanding.
Known limitations: Narrow domain (classical string quartet scores specifically); the corpus and evaluation methodology were originally built for musicology research, not AI benchmarking, so task framing for LLM evaluation is comparatively new.
The OpenScore String Quartet Corpus (Gotham et al., 2023) is a large, high-quality digital corpus of string quartet scores originally built for computational musicology research. Repurposed as an AI benchmark, it tests whether a model can take an image of a musical score and produce an accurate machine-readable transcription — a task requiring precise visual parsing of dense, structured notation across multiple simultaneous instrumental parts.
Performance is measured as 1 minus the OMR normalized edit distance (OMR-NED) between the model's transcription and the ground-truth score — a standard optical-music-recognition metric that penalizes insertions, deletions, and substitutions relative to the correct transcription, normalized by score length. A score of 1.0 would represent a perfect transcription; scores near 0 indicate the transcription bears little resemblance to the source.
| Field | Value |
|---|---|
| Task category | Reasoning / Multimodal |
| Metric | 1 - OMR-NED (normalized edit distance, higher = better) |
| Saturation | Medium — meaningful spread between frontier models |
| Created by | Mark R. H. Gotham, Maureen Redbond, Bruno Bower, Peter Jonas |
| Source paper | Gotham et al., ACM DLfM 2023 |
The score is computed as 1 minus the normalized edit distance between the model's transcription output and the ground-truth score encoding. This means the metric behaves like an accuracy score on a 0-1 scale, where 1.0 is a perfect transcription and lower values reflect proportionally more transcription errors relative to the total content of the score.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | GPT-6 Astra | 0.84 | OpenAI: GPT-6 Astra | 2026-09 |
| 2 | GPT-5.6 Sol | 0.19 | OpenAI: GPT-6 Astra | 2026-09 |
Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026), citing the OpenScore String Quartet Corpus (Gotham et al. 2023).
No Benchgen results yet — be the first to run OpenScore String Quartets.
| Benchmark | What it tests | Saturation |
|---|---|---|
| OpenScore String Quartets | Optical music recognition (score transcription) | Medium |
| ScreenSpot-Pro | Visual grounding for GUI elements | Low |
OpenScore String Quartets is a niche but genuinely differentiating visual-precision test — unlike most multimodal benchmarks, it demands exact structured transcription rather than natural-language description.
Benchgen lets teams benchmark visual transcription accuracy across model versions, tracking regressions in fine-grained multimodal parsing that a single vendor snapshot would miss.
Benchmark definition paraphrased from Gotham et al. 2023. State-of-the-art scores sourced from OpenAI's GPT-6 Astra announcement. Last updated 2026-09-07.