When to zoom in
ZoomClick begins with a Pre-Zoom that compares a global prediction with four local predictions over a predefined grid. A local candidate becomes the starting point when its distance to the global candidate falls below a threshold.
Grounding is a fundamental capability for building graphical user interface (GUI) agents. Although existing approaches rely on large-scale bounding box supervision, they still face challenges such as cross-platform generalization, complex layout analysis, and fine-grained element localization.
We investigate zoom as a strong yet underexplored prior for GUI grounding, and propose a training-free method, ZoomClick. By characterizing four key properties of zoom—pre-zoom, depth, shrink size, and minimal crop size—we unlock its full capabilities for dynamic spatial focusing and adaptive context switching.
ZoomClick integrates four critical properties of zoom: pre-zoom, depth, shrink size, and minimal crop size. The pipeline specifies when to zoom in, how to zoom, and when to click out.
ZoomClick begins with a Pre-Zoom that compares a global prediction with four local predictions over a predefined grid. A local candidate becomes the starting point when its distance to the global candidate falls below a threshold.
Each iteration crops a region of the specified shrink size directly in the original coordinate system, preventing iterative drift from relative cropping and avoiding boundary overflow that introduces irrelevant context.
Termination occurs when no target is detected, the zoom depth is reached, or the crop hits the minimal size, enabling resolution-adaptive stopping and preventing over-zooming into unrecognizable context.
Existing GUI grounding benchmarks fail to reveal essential shortcomings or model-specific limitations under different zoom conditions. GUIZoom-Bench systematically dissects zoom behavior, exposing failure modes and the adaptability boundaries of different model families.
Easy-normal: the target is visually salient and contextually clear; excessive zooming can hurt by unnecessarily discarding useful context.
Easy-mislead: the model starts correct but becomes wrong after zooming, as distractors grow more salient once global cues vanish.
Hard-normal: the target is small or visually subtle, initially hard to locate but gradually revealed through zooming.
Hard-mislead: the scene contains distractors visually similar to the target.
Hard-est: even repeated zooming fails to clarify the target.
@misc{jiang2025zoominclickout,
title={Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding},
author={Zhiyuan Jiang and Shenghao Xie and Wenyi Li and Wenqiang Zu and Peihang Li and Jiahao Qiu and Siqi Pei and Lei Ma and Tiejun Huang and Mengdi Wang and Shilong Liu},
year={2025},
eprint={2512.05941},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2512.05941},
}