Table of Contents
Fetching ...

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

Yuancheng Xu, Wenqi Xian, Li Ma, Julien Philip, Ahmet Levent Taşel, Yiwei Zhao, Ryan Burgert, Mingming He, Oliver Hermann, Oliver Pilarski, Rahul Garg, Paul Debevec, Ning Yu

TL;DR

The paper tackles the challenge of producing identity-consistent, photorealistic videos of customized subjects under 3D camera motion. It introduces a customization data pipeline that couples professional volumetric captures with 4D Gaussian Splatting reconstructions and relighting to train a camera-conditioned video diffusion model in two stages: general camera-conditioned pretraining and subject-specific fine-tuning. Key contributions include methods for multi-subject generation via joint training and noise blending, scene and real-life customization, and control over motion and spatial layout, all validated through extensive benchmarks, ablations, and user studies. The results demonstrate improved multi-view identity preservation, accurate camera control, and realistic lighting, enabling practical applications in virtual production and beyond.

Abstract

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric capture performances re-rendered with diverse camera trajectories via 4D Gaussian Splatting (4DGS), lighting variability obtained with a video relighting model. We fine-tune state-of-the-art open-source video diffusion models on this data to provide strong multi-view identity preservation, precise camera control, and lighting adaptability. Our framework also supports core capabilities for virtual production, including multi-subject generation using two approaches: joint training and noise blending, the latter enabling efficient composition of independently customized models at inference time; it also achieves scene and real-life video customization as well as control over motion and spatial layout during customization. Extensive experiments show improved video quality, higher personalization accuracy, and enhanced camera control and lighting adaptability, advancing the integration of video generation into virtual production. Our project page is available at: https://eyeline-labs.github.io/Virtually-Being.

Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures

TL;DR

The paper tackles the challenge of producing identity-consistent, photorealistic videos of customized subjects under 3D camera motion. It introduces a customization data pipeline that couples professional volumetric captures with 4D Gaussian Splatting reconstructions and relighting to train a camera-conditioned video diffusion model in two stages: general camera-conditioned pretraining and subject-specific fine-tuning. Key contributions include methods for multi-subject generation via joint training and noise blending, scene and real-life customization, and control over motion and spatial layout, all validated through extensive benchmarks, ablations, and user studies. The results demonstrate improved multi-view identity preservation, accurate camera control, and realistic lighting, enabling practical applications in virtual production and beyond.

Abstract

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric capture performances re-rendered with diverse camera trajectories via 4D Gaussian Splatting (4DGS), lighting variability obtained with a video relighting model. We fine-tune state-of-the-art open-source video diffusion models on this data to provide strong multi-view identity preservation, precise camera control, and lighting adaptability. Our framework also supports core capabilities for virtual production, including multi-subject generation using two approaches: joint training and noise blending, the latter enabling efficient composition of independently customized models at inference time; it also achieves scene and real-life video customization as well as control over motion and spatial layout during customization. Extensive experiments show improved video quality, higher personalization accuracy, and enhanced camera control and lighting adaptability, advancing the integration of video generation into virtual production. Our project page is available at: https://eyeline-labs.github.io/Virtually-Being.
Paper Structure (53 sections, 18 figures, 2 tables)

This paper contains 53 sections, 18 figures, 2 tables.

Figures (18)

  • Figure 1: Overview of our training pipeline (right) for camera-controlled customized video generation, consisting of a camera pretraining stage and a customization stage. The data pipeline (left) generates customization data by capturing multi-view performances, applying 4D Gaussian Splatting, and rendering videos with diverse viewpoints, camera motions, and lighting.
  • Figure 2: Customization results of T2V baselines and our method, demonstrating superior multi-view identity preservation by our approach.
  • Figure 3: Generated videos with multi-view data (top) and frontal-view-only data (bottom). Multi-view training yields markedly better identity preservation across viewpoints.
  • Figure 4: Generated videos with additional relit data (top) and without (bottom). Relit data significantly enhances lighting realism and diversity.
  • Figure 5: Generated videos with (top) and without (bottom) joint-subject data. Including joint-subject data improves multi-subject interaction quality.
  • ...and 13 more figures