Table of Contents
Fetching ...

Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

Minxing Luo, Linlong Fan, Wang Qiushi, Ge Wu, Yiyan Luo, Yuhang Yu, Jinwei Chen, Yaxing Wang, Qingnan Fan, Jian Yang

TL;DR

TIGER addresses the persistent trade-off between readability and image quality in scene-text super-resolution by decoupling glyph restoration from full-image enhancement. It introduces a two-stage, text-first paradigm where a region-level text restoration module reconstructs accurate glyph structures that guide a ControlNet-like diffusion-based image restoration stage, all trained with a two-phase strategy on synthetic and real data. The authors also contribute UZ-ST, an extreme-zoom Chinese scene-text dataset with precise LR–HR alignment, enabling robust evaluation under challenging degradations. Across Real-CE and UZ-ST benchmarks, TIGER achieves state-of-the-art performance in both image quality metrics and text readability (OCR-A), demonstrating the effectiveness of explicit glyph-structure guidance for coherent, artifact-free text in high-resolution scenes. The work advances practical scene-text SR with broad implications for document restoration, accessibility, and navigation tasks.

Abstract

Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text-Image Guided supEr-Resolution), a novel two-stage framework that breaks this trade-off through a "text-first, image-later" paradigm. TIGER explicitly decouples glyph restoration from image enhancement: it first reconstructs precise text structures and uses them to guide full-image super-resolution. This ensures high fidelity and readability. To support comprehensive training and evaluation, we present the UZ-ST (UltraZoom-Scene Text) dataset, the first Chinese scene text dataset with extreme zoom. Extensive experiments show TIGER achieves state-of-the-art performance, enhancing readability and image quality.

Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

TL;DR

TIGER addresses the persistent trade-off between readability and image quality in scene-text super-resolution by decoupling glyph restoration from full-image enhancement. It introduces a two-stage, text-first paradigm where a region-level text restoration module reconstructs accurate glyph structures that guide a ControlNet-like diffusion-based image restoration stage, all trained with a two-phase strategy on synthetic and real data. The authors also contribute UZ-ST, an extreme-zoom Chinese scene-text dataset with precise LR–HR alignment, enabling robust evaluation under challenging degradations. Across Real-CE and UZ-ST benchmarks, TIGER achieves state-of-the-art performance in both image quality metrics and text readability (OCR-A), demonstrating the effectiveness of explicit glyph-structure guidance for coherent, artifact-free text in high-resolution scenes. The work advances practical scene-text SR with broad implications for document restoration, accessibility, and navigation tasks.

Abstract

Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text-Image Guided supEr-Resolution), a novel two-stage framework that breaks this trade-off through a "text-first, image-later" paradigm. TIGER explicitly decouples glyph restoration from image enhancement: it first reconstructs precise text structures and uses them to guide full-image super-resolution. This ensures high fidelity and readability. To support comprehensive training and evaluation, we present the UZ-ST (UltraZoom-Scene Text) dataset, the first Chinese scene text dataset with extreme zoom. Extensive experiments show TIGER achieves state-of-the-art performance, enhancing readability and image quality.
Paper Structure (23 sections, 8 equations, 10 figures, 11 tables)

This paper contains 23 sections, 8 equations, 10 figures, 11 tables.

Figures (10)

  • Figure 1: We present TIGER (Text–Image Guided supEr-Resolution), a novel framework for scene text super-resolution. Its ‘text-first, image-later’ paradigm ensures accurate glyph restoration and consistently high overall image fidelity and visual quality.
  • Figure 2: The framework of TIGER, which includes the Text Restoration stage (stage 1) and the Image Enhancement stage (stage 2). Stage 1 recover accurate glyph structures from text regions. Stage 2 uses them to guide full-image restoration for coherent text and background.
  • Figure 3: Overview of UZ-ST (UltraZoom-Scene Text). (a) Real-CE LRs show only mild degradation (red box), while UZ-ST LRs exhibit stronger degradation (red box), enabling a more comprehensive evaluation. (b) Coarse-to-fine alignment: images are sorted by focal length, each warped to the next higher-focal neighbor using an estimated homography matrix, then refined to the 200 mm GT.
  • Figure 4: Qualitative Evaluation on Real-CE and UZ-ST.
  • Figure 5: Qualitative Results of Ablation Study with stage 2 fixed as the baseline.
  • ...and 5 more figures