Benchgen

ScreenSpot-Pro — Results

RankModelScore
1muse-glimmer75.4
S

ScreenSpot-Pro

1 phaseActive

High-resolution professional GUI grounding benchmark spanning 23 apps across 5 industries and 3 OSes. CC BY 4.0. Metric: % click-region accuracy.

Overview

ScreenSpot-Pro

Category Metric Saturation Created

Paper Leaderboard

Quick answer: ScreenSpot-Pro tests multimodal models' ability to ground GUI instructions (click the right element) on authentic high-resolution screenshots from 23 professional applications across 5 industries and 3 operating systems. It's substantially harder than earlier GUI-grounding benchmarks — the best model in the original paper reached only 18.9% accuracy.

What ScreenSpot-Pro Measures

ScreenSpot-Pro evaluates the GUI grounding capability of multimodal models: given a natural-language instruction and a screenshot, can the model correctly localize the target UI element? Unlike general-purpose GUI benchmarks built on consumer web/mobile apps, ScreenSpot-Pro uses authentic high-resolution screenshots from specialized professional software — spanning CAD, creative, scientific, and other domains — where small target sizes, dense layouts, and unfamiliar interfaces make grounding meaningfully harder.

The paper found existing GUI grounding models perform poorly on this dataset (best baseline: 18.9%), and introduced ScreenSeekeR, a visual search method that narrows the search area using a planner's GUI knowledge, achieving a new state of the art of 48.1% without additional training — evidence that professional, high-resolution GUI grounding remains an open problem even for strong general-purpose models.

Benchmark Specifications

FieldValue
Task categoryReasoning (multimodal GUI grounding)
Metric% accuracy (correct click-region localization)
Coverage23 applications across 5 industries and 3 operating systems
SaturationLow
Created byLi et al.
Source paperLi et al. 2025
Leaderboardgui-agent.github.io/grounding-leaderboard

How ScreenSpot-Pro Is Scored

Models are scored on the percentage of instructions where the predicted click location falls within the correct UI element's bounding region, evaluated across authentic high-resolution screenshots. Because targets in professional software are often small and visually dense, scores in the 40-75% range represent strong performance for current frontier multimodal models, well below saturation.

ScreenSpot-Pro on Benchgen

No Benchgen results yet — be the first to run ScreenSpot-Pro.

ScreenSpot-Pro vs Other Benchmarks

BenchmarkWhat it testsSaturation
ScreenSpot-ProHigh-resolution professional GUI element groundingLow
OSWorld-VerifiedEnd-to-end real desktop/web computer-use tasksLow
OmniDocBenchDocument parsing (PDFs, scanned text)Low

ScreenSpot-Pro isolates the grounding step (where to click) rather than full task completion, complementing end-to-end computer-use benchmarks like OSWorld-Verified.

Frequently Asked Questions

What is ScreenSpot-Pro? ScreenSpot-Pro is a benchmark testing multimodal models' ability to ground natural-language GUI instructions to the correct on-screen element, using authentic high-resolution screenshots from 23 professional applications.
What does a good ScreenSpot-Pro score look like? The original paper's best baseline reached only 18.9% accuracy; a specialized visual-search method (ScreenSeekeR) reached 48.1% — scores above 40% represent strong performance on this still-unsaturated benchmark.
Who created ScreenSpot-Pro? ScreenSpot-Pro was created by Kaixin Li, Ziyang Meng, and collaborators, published in April 2025.
Is ScreenSpot-Pro saturated? No — even specialized methods reach under 50% accuracy, leaving substantial headroom for improvement in professional GUI grounding.

Benchmark definition based on Li et al., "ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use" (arXiv:2504.07981). Last updated 2026-08-10.