Benchgen

OpenScore String Quartets — Results

RankModelScore
1gpt-6-astra0.84
O

OpenScore String Quartets

1 phaseActive

Optical music recognition benchmark on the OpenScore String Quartet Corpus — scored as 1 minus normalized edit distance.

Overview

OpenScore String Quartets

Category Metric Saturation Created

Paper

Quick answer: OpenScore String Quartets is a benchmark built on the OpenScore String Quartet Corpus (Gotham, Redbond, Bower & Jonas, 2023) that tests optical music recognition (OMR) — a model's ability to read a musical score image and transcribe it accurately. Scores are reported as 1 minus the OMR normalized edit distance (OMR-NED), so higher is better. GPT-6 Astra scores 0.84, versus 0.19 for GPT-5.6 Sol.

At a Glance

What it tests: Whether a multimodal model can visually parse a string quartet score (an image of sheet music) and transcribe its notation accurately — pitches, rhythms, and structure across four instrumental parts.

Why it matters: Optical music recognition is a specialized visual-reasoning task combining fine-grained image understanding with structured symbolic output. Strong performance signals genuine visual precision, not just pattern-matching on common image types — a distinct capability from general OCR or document understanding.

Known limitations: Narrow domain (classical string quartet scores specifically); the corpus and evaluation methodology were originally built for musicology research, not AI benchmarking, so task framing for LLM evaluation is comparatively new.

What OpenScore String Quartets Measures

The OpenScore String Quartet Corpus (Gotham et al., 2023) is a large, high-quality digital corpus of string quartet scores originally built for computational musicology research. Repurposed as an AI benchmark, it tests whether a model can take an image of a musical score and produce an accurate machine-readable transcription — a task requiring precise visual parsing of dense, structured notation across multiple simultaneous instrumental parts.

Performance is measured as 1 minus the OMR normalized edit distance (OMR-NED) between the model's transcription and the ground-truth score — a standard optical-music-recognition metric that penalizes insertions, deletions, and substitutions relative to the correct transcription, normalized by score length. A score of 1.0 would represent a perfect transcription; scores near 0 indicate the transcription bears little resemblance to the source.

Benchmark Specifications

FieldValue
Task categoryReasoning / Multimodal
Metric1 - OMR-NED (normalized edit distance, higher = better)
SaturationMedium — meaningful spread between frontier models
Created byMark R. H. Gotham, Maureen Redbond, Bruno Bower, Peter Jonas
Source paperGotham et al., ACM DLfM 2023

How OpenScore String Quartets Is Scored

The score is computed as 1 minus the normalized edit distance between the model's transcription output and the ground-truth score encoding. This means the metric behaves like an accuracy score on a 0-1 scale, where 1.0 is a perfect transcription and lower values reflect proportionally more transcription errors relative to the total content of the score.

State-of-the-Art Results

RankModelScoreSourceDate
1GPT-6 Astra0.84OpenAI: GPT-6 Astra2026-09
2GPT-5.6 Sol0.19OpenAI: GPT-6 Astra2026-09

Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026), citing the OpenScore String Quartet Corpus (Gotham et al. 2023).

OpenScore String Quartets on Benchgen

No Benchgen results yet — be the first to run OpenScore String Quartets.

OpenScore String Quartets vs Other Benchmarks

BenchmarkWhat it testsSaturation
OpenScore String QuartetsOptical music recognition (score transcription)Medium
ScreenSpot-ProVisual grounding for GUI elementsLow

OpenScore String Quartets is a niche but genuinely differentiating visual-precision test — unlike most multimodal benchmarks, it demands exact structured transcription rather than natural-language description.

Run OpenScore String Quartets on Your Model

Benchgen lets teams benchmark visual transcription accuracy across model versions, tracking regressions in fine-grained multimodal parsing that a single vendor snapshot would miss.

Frequently Asked Questions

What is OpenScore String Quartets? A benchmark testing AI optical music recognition on string quartet scores from the OpenScore String Quartet Corpus, scored as 1 minus normalized edit distance (OMR-NED).
What does a good score look like? Scores close to 1.0 indicate near-perfect transcription. As of September 2026, GPT-6 Astra leads at 0.84, a large jump from GPT-5.6 Sol's 0.19.
Who created OpenScore String Quartets? The underlying corpus was created by Mark R. H. Gotham, Maureen Redbond, Bruno Bower, and Peter Jonas, published at ACM DLfM 2023.

Benchmark definition paraphrased from Gotham et al. 2023. State-of-the-art scores sourced from OpenAI's GPT-6 Astra announcement. Last updated 2026-09-07.