Zoom in, Click out:
Unlocking and Evaluating the Potential of Zooming for GUI Grounding

* Equal contribution † Corresponding authors

1 Xi’an Jiaotong University · 2 Princeton University · 3 Peking University · 4 University of Chinese Academy of Sciences · 5 The University of Hong Kong · 6 Michigan State University

ZoomClick performance on ScreenSpot-Pro and accuracy-efficiency trade-offs across model scales
Left: performance of existing GUI grounding methods on ScreenSpot-Pro. Right: accuracy-efficiency trade-offs without ZoomClick across model scales.
01 / The idea

Abstract

Grounding is a fundamental capability for building graphical user interface (GUI) agents. Although existing approaches rely on large-scale bounding box supervision, they still face challenges such as cross-platform generalization, complex layout analysis, and fine-grained element localization.

We investigate zoom as a strong yet underexplored prior for GUI grounding, and propose a training-free method, ZoomClick. By characterizing four key properties of zoom—pre-zoom, depth, shrink size, and minimal crop size—we unlock its full capabilities for dynamic spatial focusing and adaptive context switching.

73.1%UI-Venus-72B + ZoomClick
on ScreenSpot-Pro
42.5%UI-Venus-72B + ZoomClick
on UI-Vision
+34.4%relative gain for
Qwen3-VL-32B
02 / Method

Method

ZoomClick integrates four critical properties of zoom: pre-zoom, depth, shrink size, and minimal crop size. The pipeline specifies when to zoom in, how to zoom, and when to click out.

ZoomClick framework showing pre-zoom, iterative narrowing, and final prediction stages
Model Framework of ZoomClick. Min_crop_size represents the lower bound of the viewport during Iterative Narrowing.
01

When to zoom in

ZoomClick begins with a Pre-Zoom that compares a global prediction with four local predictions over a predefined grid. A local candidate becomes the starting point when its distance to the global candidate falls below a threshold.

02

How to zoom

Each iteration crops a region of the specified shrink size directly in the original coordinate system, preventing iterative drift from relative cropping and avoiding boundary overflow that introduces irrelevant context.

03

When to click out

Termination occurs when no target is detected, the zoom depth is reached, or the crop hits the minimal size, enabling resolution-adaptive stopping and preventing over-zooming into unrecognizable context.

03 / Benchmark

GUIZoom-Bench

Existing GUI grounding benchmarks fail to reveal essential shortcomings or model-specific limitations under different zoom conditions. GUIZoom-Bench systematically dissects zoom behavior, exposing failure modes and the adaptability boundaries of different model families.

Easy-normal: the target is visually salient and contextually clear; excessive zooming can hurt by unnecessarily discarding useful context.

Easy-mislead: the model starts correct but becomes wrong after zooming, as distractors grow more salient once global cues vanish.

Hard-normal: the target is small or visually subtle, initially hard to locate but gradually revealed through zooming.

Hard-mislead: the scene contains distractors visually similar to the target.

Hard-est: even repeated zooming fails to clarify the target.

GUIZoom-Bench curation process and number of samples in five behavior categories
Overview of benchmark design and zoom-depth evaluation. Data organization of GUIZoom-Bench and model accuracy as zoom depth increases.
Examples of GUIZoom-Bench zoom behaviors
Data examples of each category in GUIZoom-Bench.
04 / Results

Results

Accuracy and efficiency comparison showing ZoomClick gains across model scales
Accuracy-efficiency trade-offs without ZoomClick across model scales.
Accuracy comparison with and without ZoomClick across grounding backbones
Accuracy vs. zoom depth across models.
Qualitative GUI grounding results comparing baseline and ZoomClick predictions
Qualitative results of ZoomClick on ScreenSpot-Pro: deeper zoom resolves errors from earlier depths.
05 / Cite

BibTeX

@misc{jiang2025zoominclickout,
      title={Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding},
      author={Zhiyuan Jiang and Shenghao Xie and Wenyi Li and Wenqiang Zu and Peihang Li and Jiahao Qiu and Siqi Pei and Lei Ma and Tiejun Huang and Mengdi Wang and Shilong Liu},
      year={2025},
      eprint={2512.05941},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2512.05941},
}