Match-and-Fuse: Consistent Generation from Unstructured Image Sets

Kate Feingold; Omri Kaduri; Tali Dekel

Match-and-Fuse: Consistent Generation from Unstructured Image Sets

Kate Feingold, Omri Kaduri, Tali Dekel

TL;DR

The paper addresses generating coherent edits from unstructured image sets by introducing Match-and-Fuse, a zero-shot, training-free framework that preserves cross-image consistency for shared content. It models the set as a Pairwise Consistency Graph and employs Multiview Feature Fusion guided by dense 2D correspondences to enforce global coherence across all pairwise grids, extending the grid prior without masks. Per-image prompts are automatically composed from set-level prompts, and a lightweight Feature Guidance term further aligns features across views. Extensive experiments demonstrate superior cross-image consistency and visual fidelity compared to strong baselines, with diverse applications including storyboard-style edits and flow-based localized adjustments. The work advances set-to-set generation and lays groundwork for scalable, consistent editing across image collections and potentially video collections.

Abstract

We present Match-and-Fuse - a zero-shot, training-free method for consistent controlled generation of unstructured image sets - collections that share a common visual element, yet differ in viewpoint, time of capture, and surrounding content. Unlike existing methods that operate on individual images or densely sampled videos, our framework performs set-to-set generation: given a source set and user prompts, it produces a new set that preserves cross-image consistency of shared content. Our key idea is to model the task as a graph, where each node corresponds to an image and each edge triggers a joint generation of image pairs. This formulation consolidates all pairwise generations into a unified framework, enforcing their local consistency while ensuring global coherence across the entire set. This is achieved by fusing internal features across image pairs, guided by dense input correspondences, without requiring masks or manual supervision. It also allows us to leverage an emergent prior in text-to-image models that encourages coherent generation when multiple views share a single canvas. Match-and-Fuse achieves state-of-the-art consistency and visual quality, and unlocks new capabilities for content creation from image collections.

Match-and-Fuse: Consistent Generation from Unstructured Image Sets

TL;DR

Abstract

Match-and-Fuse: Consistent Generation from Unstructured Image Sets

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (18)