Table of Contents
Fetching ...

CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models

Denis Rychkovskiy

TL;DR

CADE 2.5 targets CFG-induced degradation in SD/SDXL latent diffusion by introducing ZeResFDG, a sampler-side stack that combines Frequency-Decoupled Guidance, Energy Rescaling, and Zero-Projection, all steered by a spectral EMA with hysteresis. The approach reshapes guidance rather than altering the model, using per-sample energy matching and parallel/orthogonal decomposition to preserve global tone while enhancing micro-detail, complemented by a lightweight inference-time stabilizer, QSilk Micrograin. The method is training-free, compatible with velocity parameterizations, and integrates into SD/SDXL pipelines with a four-pass preset workflow to stabilize tone early, grow detail mid, balance and finish late, and polish at the end. Empirical visuals and qualitative assessments indicate improved sharpness, prompt adherence, and artifact control at moderate CFG scales, with robust high-frequency texture at high resolutions and minimal overhead. Together, these contributions offer a practical, plug-and-play improvement for real-world latent-diffusion image synthesis without retraining.

Abstract

We introduce CADE 2.5 (Comfy Adaptive Detail Enhancer), a sampler-level guidance stack for SD/SDXL latent diffusion models. The central module, ZeResFDG, unifies (i) frequency-decoupled guidance that reweights low- and high-frequency components of the guidance signal, (ii) energy rescaling that matches the per-sample magnitude of the guided prediction to the positive branch, and (iii) zero-projection that removes the component parallel to the unconditional direction. A lightweight spectral EMA with hysteresis switches between a conservative and a detail-seeking mode as structure crystallizes during sampling. Across SD/SDXL samplers, ZeResFDG improves sharpness, prompt adherence, and artifact control at moderate guidance scales without any retraining. In addition, we employ a training-free inference-time stabilizer, QSilk Micrograin Stabilizer (quantile clamp + depth/edge-gated micro-detail injection), which improves robustness and yields natural high-frequency micro-texture at high resolutions with negligible overhead. For completeness we note that the same rule is compatible with alternative parameterizations (e.g., velocity), which we briefly discuss in the Appendix; however, this paper focuses on SD/SDXL latent diffusion models.

CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models

TL;DR

CADE 2.5 targets CFG-induced degradation in SD/SDXL latent diffusion by introducing ZeResFDG, a sampler-side stack that combines Frequency-Decoupled Guidance, Energy Rescaling, and Zero-Projection, all steered by a spectral EMA with hysteresis. The approach reshapes guidance rather than altering the model, using per-sample energy matching and parallel/orthogonal decomposition to preserve global tone while enhancing micro-detail, complemented by a lightweight inference-time stabilizer, QSilk Micrograin. The method is training-free, compatible with velocity parameterizations, and integrates into SD/SDXL pipelines with a four-pass preset workflow to stabilize tone early, grow detail mid, balance and finish late, and polish at the end. Empirical visuals and qualitative assessments indicate improved sharpness, prompt adherence, and artifact control at moderate CFG scales, with robust high-frequency texture at high resolutions and minimal overhead. Together, these contributions offer a practical, plug-and-play improvement for real-world latent-diffusion image synthesis without retraining.

Abstract

We introduce CADE 2.5 (Comfy Adaptive Detail Enhancer), a sampler-level guidance stack for SD/SDXL latent diffusion models. The central module, ZeResFDG, unifies (i) frequency-decoupled guidance that reweights low- and high-frequency components of the guidance signal, (ii) energy rescaling that matches the per-sample magnitude of the guided prediction to the positive branch, and (iii) zero-projection that removes the component parallel to the unconditional direction. A lightweight spectral EMA with hysteresis switches between a conservative and a detail-seeking mode as structure crystallizes during sampling. Across SD/SDXL samplers, ZeResFDG improves sharpness, prompt adherence, and artifact control at moderate guidance scales without any retraining. In addition, we employ a training-free inference-time stabilizer, QSilk Micrograin Stabilizer (quantile clamp + depth/edge-gated micro-detail injection), which improves robustness and yields natural high-frequency micro-texture at high resolutions with negligible overhead. For completeness we note that the same rule is compatible with alternative parameterizations (e.g., velocity), which we briefly discuss in the Appendix; however, this paper focuses on SD/SDXL latent diffusion models.
Paper Structure (22 sections, 2 equations, 3 figures)

This paper contains 22 sections, 2 equations, 3 figures.

Figures (3)

  • Figure 1: Qualitative samples "Anime style" produced with CADE 2.5 (ZeResFDG pipe (SDXL)).
  • Figure 2: Qualitative samples "Photo style" produced with CADE 2.5 (ZeResFDG pipe (SDXL)).
  • Figure 3: Qualitative samples "Photo style" produced with CADE 2.5 (ZeResFDG pipe (SDXL)).