Table of Contents
Fetching ...

GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer

Sayan Deb Sarkar, Sinisa Stekovic, Vincent Lepetit, Iro Armeni

TL;DR

GuideFlow3D tackles 3D appearance transfer under large geometry differences by using a training-free framework that combines a pretrained rectified flow with inference-time guided latent optimization. It introduces two differentiable guidance losses—part-aware appearance loss and self-similarity loss—to enforce local semantic correspondences and structure during sampling, conditioning on images, meshes, or text. The method builds on structured latents (SLat) from Trellis and applies a Bayesian-inspired objective within a universal guidance setup, updating the latent trajectory with hat z_t = z_t + Δt v_θ(z_t,t|c) + ∇_{z_t} ℒ(z|c) to achieve realistic, geometry-aware transfers without retraining. Experiments on synthetic and in-the-wild datasets show clear gains over baselines in appearance fidelity and structural coherence, validated through GPT-based rankings and user studies. The approach is general, extensible to other diffusion models and guidance functions, and has practical impact for 3D asset stylization in AR/VR and gaming pipelines.

Abstract

Transferring appearance to 3D assets using different representations of the appearance object - such as images or text - has garnered interest due to its wide range of applications in industries like gaming, augmented reality, and digital content creation. However, state-of-the-art methods still fail when the geometry between the input and appearance objects is significantly different. A straightforward approach is to directly apply a 3D generative model, but we show that this ultimately fails to produce appealing results. Instead, we propose a principled approach inspired by universal guidance. Given a pretrained rectified flow model conditioned on image or text, our training-free method interacts with the sampling process by periodically adding guidance. This guidance can be modeled as a differentiable loss function, and we experiment with two different types of guidance including part-aware losses for appearance and self-similarity. Our experiments show that our approach successfully transfers texture and geometric details to the input 3D asset, outperforming baselines both qualitatively and quantitatively. We also show that traditional metrics are not suitable for evaluating the task due to their inability of focusing on local details and comparing dissimilar inputs, in absence of ground truth data. We thus evaluate appearance transfer quality with a GPT-based system objectively ranking outputs, ensuring robust and human-like assessment, as further confirmed by our user study. Beyond showcased scenarios, our method is general and could be extended to different types of diffusion models and guidance functions.

GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer

TL;DR

GuideFlow3D tackles 3D appearance transfer under large geometry differences by using a training-free framework that combines a pretrained rectified flow with inference-time guided latent optimization. It introduces two differentiable guidance losses—part-aware appearance loss and self-similarity loss—to enforce local semantic correspondences and structure during sampling, conditioning on images, meshes, or text. The method builds on structured latents (SLat) from Trellis and applies a Bayesian-inspired objective within a universal guidance setup, updating the latent trajectory with hat z_t = z_t + Δt v_θ(z_t,t|c) + ∇_{z_t} ℒ(z|c) to achieve realistic, geometry-aware transfers without retraining. Experiments on synthetic and in-the-wild datasets show clear gains over baselines in appearance fidelity and structural coherence, validated through GPT-based rankings and user studies. The approach is general, extensible to other diffusion models and guidance functions, and has practical impact for 3D asset stylization in AR/VR and gaming pipelines.

Abstract

Transferring appearance to 3D assets using different representations of the appearance object - such as images or text - has garnered interest due to its wide range of applications in industries like gaming, augmented reality, and digital content creation. However, state-of-the-art methods still fail when the geometry between the input and appearance objects is significantly different. A straightforward approach is to directly apply a 3D generative model, but we show that this ultimately fails to produce appealing results. Instead, we propose a principled approach inspired by universal guidance. Given a pretrained rectified flow model conditioned on image or text, our training-free method interacts with the sampling process by periodically adding guidance. This guidance can be modeled as a differentiable loss function, and we experiment with two different types of guidance including part-aware losses for appearance and self-similarity. Our experiments show that our approach successfully transfers texture and geometric details to the input 3D asset, outperforming baselines both qualitatively and quantitatively. We also show that traditional metrics are not suitable for evaluating the task due to their inability of focusing on local details and comparing dissimilar inputs, in absence of ground truth data. We thus evaluate appearance transfer quality with a GPT-based system objectively ranking outputs, ensuring robust and human-like assessment, as further confirmed by our user study. Beyond showcased scenarios, our method is general and could be extended to different types of diffusion models and guidance functions.
Paper Structure (20 sections, 10 equations, 14 figures, 6 tables)

This paper contains 20 sections, 10 equations, 14 figures, 6 tables.

Figures (14)

  • Figure 1: GuideFlow3D is a method for 3D appearance transfer robust to strong geometric variations between objects. Given an input 3D mesh, e.g., designed using simple 3D primitives, it transfers the texture and fine geometric details of an appearance object (e.g., the rounded edges of the table on the top left and the base and mattress distinction of the bed on the top right) but preserves the geometric form of the input mesh. Its flexibility across appearance modalities like meshes or text makes GuideFlow3D efficient for generating diverse 3D assets.
  • Figure 2: GuideFlow3D introduces guided rectified flow for appearance transfer between input object $\mathcal{O}^q$ and an appearance object. We extend the denoising process of structured latents $\tilde{z}^q$, conditioned by $\mathbf{c}$, by introducing an objective function $\mathcal{L}$ that enforces strong geometric and semantic priors during the process. We show denoised structured latents at different stages of the process, along with corresponding meshes decoded using a pretrained decoder $\mathcal{D}_{\textit{3D}}$. The output 3D asset displays robustness to strong geometric variations between input and appearance objects.
  • Figure 3: GuideFlow3D with different guidance objectives. (a) When a textured mesh is available for appearance object $\mathcal{O}^a$, we use our co-segmentation based objective $\mathcal{L}_{appearance}$ to guide appearance transfer. It encourages consistency between structured latents $z^a$ and noisy latents $\tilde{z}^q$. In such case, we use an image of object $\mathcal{O}^a$ to condition the generative model $\mathcal{R}$. (b) When textured mesh is not available, we use our geometric clustering based objective $\mathcal{L}_{structure}$ for guidance. It encourages intra-cluster similarity and inter-cluster disparity when denoising $\tilde{z}^q$. We use text or image to condition $\mathcal{R}$ in such case. For both cases, we use decoder $\mathcal{D_{\textit{3D}}}$ to obtain the output 3D asset.
  • Figure 4: Effect of different modules. (a) Using only optimization with our objective functions to transfer appearance is insufficient as it does not enforce realistic distribution over the structured latent space. (b) The rectified flow model from trellis fails to transfer appearance when appearance and input objects have significantly different geometries. (c) We obtain appealing 3D assets when using rectified flow guidance of our GuideFlow3D.
  • Figure 5: Qualitative Comparisons showing quality of appearance transfer. Top and bottom rows show intra-class (chair to chair) and inter-class (cabinet to bunk bed) results respectively. In both examples, MambaST mambast blends the textures from both input and appearance objects, giving a grey hue to the final result. EasiTex easitex generates non-smooth repetitions of texture and fails to generate textures for the entire object (e.g., the handles of the chair or the bottom part of the bunk bed). Cross Image Attention crossimageattention performs better but omits texture details (cabinet's wood texture) and fails to preserve good local geometry. Trellis trellis preserves better the texture of the appearance object and does better on matching it to the input object, however, it fails at providing a uniform texture on the arms of the chair and does not preserve the overall geometry of the bunk bed by closing the side hole. Ours performs the best by adhering to the appearance object's texture and matching it on the input object, while preserving the overall geometry of the input object.
  • ...and 9 more figures