Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
Minxing Luo, Linlong Fan, Wang Qiushi, Ge Wu, Yiyan Luo, Yuhang Yu, Jinwei Chen, Yaxing Wang, Qingnan Fan, Jian Yang
TL;DR
TIGER addresses the persistent trade-off between readability and image quality in scene-text super-resolution by decoupling glyph restoration from full-image enhancement. It introduces a two-stage, text-first paradigm where a region-level text restoration module reconstructs accurate glyph structures that guide a ControlNet-like diffusion-based image restoration stage, all trained with a two-phase strategy on synthetic and real data. The authors also contribute UZ-ST, an extreme-zoom Chinese scene-text dataset with precise LR–HR alignment, enabling robust evaluation under challenging degradations. Across Real-CE and UZ-ST benchmarks, TIGER achieves state-of-the-art performance in both image quality metrics and text readability (OCR-A), demonstrating the effectiveness of explicit glyph-structure guidance for coherent, artifact-free text in high-resolution scenes. The work advances practical scene-text SR with broad implications for document restoration, accessibility, and navigation tasks.
Abstract
Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text-Image Guided supEr-Resolution), a novel two-stage framework that breaks this trade-off through a "text-first, image-later" paradigm. TIGER explicitly decouples glyph restoration from image enhancement: it first reconstructs precise text structures and uses them to guide full-image super-resolution. This ensures high fidelity and readability. To support comprehensive training and evaluation, we present the UZ-ST (UltraZoom-Scene Text) dataset, the first Chinese scene text dataset with extreme zoom. Extensive experiments show TIGER achieves state-of-the-art performance, enhancing readability and image quality.
