Cobra: Efficient Line Art COlorization with BRoAder References

Junhao Zhuang; Lingen Li; Xuan Ju; Zhaoyang Zhang; Chun Yuan; Ying Shan

Cobra: Efficient Line Art COlorization with BRoAder References

Junhao Zhuang, Lingen Li, Xuan Ju, Zhaoyang Zhang, Chun Yuan, Ying Shan

TL;DR

Cobra tackles the challenge of colorizing line art with extensive reference guidance by designing a long-context diffusion framework that scales to over 200 reference images while maintaining low latency. It introduces a Causal Sparse DiT with KV-Cache, and Localized Reusable Position Encoding to efficiently fuse many references without altering pre-trained 2D encodings. A Line Art Guider, along with a Self-Attention-Only block, line-art style augmentation, and a hint-point sampling strategy, enables precise color ID preservation and flexible color hints. Empirical results on Cobra-Bench show superior image quality, color ID accuracy, and speed compared to ColorFlow and other baselines, with a clear industrial impact for multi-reference comic colorization.

Abstract

The comic production industry requires reference-based line art colorization with high accuracy, efficiency, contextual consistency, and flexible control. A comic page often involves diverse characters, objects, and backgrounds, which complicates the coloring process. Despite advancements in diffusion models for image generation, their application in line art colorization remains limited, facing challenges related to handling extensive reference images, time-consuming inference, and flexible control. We investigate the necessity of extensive contextual image guidance on the quality of line art colorization. To address these challenges, we introduce Cobra, an efficient and versatile method that supports color hints and utilizes over 200 reference images while maintaining low latency. Central to Cobra is a Causal Sparse DiT architecture, which leverages specially designed positional encodings, causal sparse attention, and Key-Value Cache to effectively manage long-context references and ensure color identity consistency. Results demonstrate that Cobra achieves accurate line art colorization through extensive contextual reference, significantly enhancing inference speed and interactivity, thereby meeting critical industrial demands. We release our codes and models on our project page: https://zhuang2002.github.io/Cobra/.

Cobra: Efficient Line Art COlorization with BRoAder References

TL;DR

Abstract

Cobra: Efficient Line Art COlorization with BRoAder References

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (12)