Pith. sign in

REVIEW 3 major objections 7 minor 243 references

ID-V2V: Identity-Preserving Video Restylization

T0 review · 3 major / 7 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read A video-to-video framework that propagates scene, lighting, and style edits from an edited keyframe across a video while preserving facial identity and fine-grained facial performance, by decoupling edit-driven synthesis from source-grounde

desk verdict A credible, practically useful video restylization system with a clever single-video training recipe; the main unresolved risk is the unmeasured fidelity of the relighting model that generates the training conditions. read the letter →

arxiv 2607.22830 v1 pith:6Y5DFVDH submitted 2026-07-24 cs.CV

classification cs.CV
keywords identity-preservingvideorestylizationrelightingvideo-to-videogenerationdiffusioneditingkeyframe-basedfacialperformancepreservationmulti-subjecttrainingdatasynthesis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces identity-preserving video restylization: given a source video of a person performing and an edited first frame that changes background, lighting, or style, generate a full video that carries the edit everywhere while keeping facial likeness and performance (expressions, gaze, lip sync) intact. The central claim is that this is achievable by separating the problem into two halves: flexible edit-driven synthesis, guided by the edited keyframe and scene depth, and strict source-grounded identity preservation, guided by pixel-level face information. Because real paired before/after restylized videos are scarce, the paper constructs training pairs from a single video: a relighting model converts each cinematic training face to a fixed 'regular' top-lit condition, and the model learns to transform such regular lighting into the stylized lighting of the edited keyframe. On single- and two-subject benchmarks, ID-V2V reports substantially higher identity and expression-preservation scores than existing video-to-video baselines, and wins 70–86% in user preference. If correct, this enables a capture-first, restyle-later production workflow in which actors are filmed under simple lighting and the visual look is designed in post-production.

What carries the argument

The key machinery is the two-channel conditioning scheme. 'Relit face video' — the source face with illumination normalized to a fixed top-lit configuration by a dedicated image-based relighting model — supplies pixel-level appearance cues so the generator learns to transform regular lighting into the stylized lighting of the edited keyframe. Facial normal maps add geometric constraints that resolve lighting/shape ambiguity and prevent identity drift. On the synthesis side, the edited first frame and a source depth sequence specify the target scene and motion. All controls are injected through a shared multi-control diffusion branch whose features are summed before entering the main video tr

What would settle it

Run ID-V2V on a set of source videos with deliberately non-top-lit illumination (strong colored light, hard shadows, underlighting), edit the first frame to a neutral look, and measure AdaFace and expression-preservation scores in later frames. If residual color casts or shadows reappear and identity/expression scores drop systematically, the claim that illumination is the primary permissible variation holds only inside a narrow training-lighting regime.

Watch

Extended reading notes

Core claim

The paper's core discovery is that identity preservation and edit propagation can be decoupled, and that identity preservation can be treated as a video relighting problem. Under identity-preserving restylization, facial structure and expression are invariant; illumination is the primary permissible variation. ID-V2V therefore conditions the video generator on relit face regions (the source face with its lighting normalized) and facial normal maps to anchor likeness and performance, while the edited keyframe plus a depth sequence drives the synthesis of the new scene and lighting across time. Training pairs are synthesized from a single stylized video by relighting faces to a regular top-lit

Load-bearing premise

The method assumes the source video was captured under roughly regular, top-lit illumination matching the training relighting target, and that the relighting model used to build training pairs preserves identity and expressions; if either gives way, the identity-preservation channel is corrupted.

Editorial extensions

If this is right

  • A capture-first workflow becomes practical: shoot under simple lighting, edit one keyframe to set the look, and generate the whole restyled video with identity and performance intact.
  • Paired training data is no longer a bottleneck: synthetic pairs derived from single videos via relighting are sufficient to train the model.
  • Fine-grained facial performance cues such as lip sync, gaze, and micro-expressions are retained at levels that landmark- or embedding-based video generators do not reach.
  • Multi-subject sequences remain coherent: each person's identity and their facial interactions are preserved while the scene is restyled.
  • The relighting-based control generalizes beyond the face at inference, extending to full-scene relighting of body, objects, and background without explicit training for it.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The fidelity ceiling of the whole pipeline is set by the relighting model: any identity or expression distortion introduced when synthesizing the relit training faces will propagate into the restyled video, so improving relighting quality is a direct lever on the method's core claim.
  • The fixed top-lit training target quietly restricts the method to source footage with ordinary illumination; training with a distribution of canonical lighting conditions would be a natural extension to cover the colored-light and hard-shadow failure documented in the paper.
  • The decoupling principle generalizes: any edit that leaves facial structure and performance invariant, such as makeup or hairstyle changes, could reuse the same identity-preservation channel, with new synthesis-side controls for the edit type.
  • The single-video synthetic-pair construction is a transferable recipe for video-editing tasks that suffer from scarce paired corpora, not just for lighting-driven restylization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper formalizes identity-preserving video restylization, in which an edited first frame specifies scene/lighting/style edits that must propagate across a source video while preserving facial identity and fine-grained facial performance (expressions, gaze, lip sync). The proposed system, ID-V2V, decouples edit-driven synthesis (conditioned on the edited keyframe and a depth sequence) from source-grounded identity preservation (conditioned on relit face regions and facial normal maps). Training pairs are synthesized from single videos: a learned image-based relighting model converts cinematic training faces to a fixed 'regular' top-lit condition, and the video model is trained to map those relit faces back to the original cinematic target. At inference, the original source-face pixels replace the relit-face condition. Experiments on 100 single-subject and 60 two-subject videos report large gains over baselines in AdaFace, Exp-AU, Exp-EMOCA, and user-study win rates, with ablations showing both relit-face and normal controls contribute. The paper also qualitatively demonstrates multi-subject (>2) restylization and full-scene relighting.

Significance. If the central claim holds, ID-V2V would be a practically valuable contribution: it enables a capture-first, restyle-later workflow for human-centric video editing, with an elegant solution to the paired-data scarcity problem via single-video training-pair construction. The decoupling insight and the use of relit faces as a proxy for source video lighting are conceptually clean. The reported quantitative gains are large and consistent, and the user-study preferences are strong. The paper releases code, which supports reproducibility. However, the significance is contingent on the validity of the relighting-model-based training pipeline and on the statistical reliability of the evaluations; both currently have gaps.

major comments (3)
  1. [§3.2, §3.3] The training and inference protocols create a distributional gap that is load-bearing for the central claim. The model is trained on relit face regions produced by a learned relighting model, but at inference it is conditioned on the original source-face pixels. The paper asserts (Fig. 2, §3.2) that relit faces 'simulate the lighting conditions of the source video at inference,' but the relighting model's fidelity is never quantitatively evaluated. If the relighting model distorts identity, expression, gaze, or lip shape, the video model learns a mapping from a corrupted condition, and at inference the undistorted source face is off-distribution. This would degrade identity/performance preservation even under the intended regular-lighting regime, a failure more fundamental than the acknowledged extreme-lighting limitation. The paper reports only qualitative examples (Fig. 6). Please repo
  2. [Table 1, §4.1, §4.2] The claim that ID-V2V 'significantly outperforms' baselines is not statistically substantiated. Table 1 reports point estimates without error bars, confidence intervals, or significance tests. The evaluation datasets are self-curated (100+60 videos) with manual filtering of edited keyframes (§4.1), and the face-related metrics are computed only on frames with detected faces (§4.2). The user study uses 26 participants and 20 samples per setting, but win rates are reported without variance or significance. Please provide per-metric uncertainty (e.g., bootstrap CIs) and paired significance tests (e.g., Wilcoxon signed-rank) for the facial metrics and user-study results, and describe the manual filtering criteria to assess potential selection bias.
  3. [§4.1, Fig. 8] The paper lists multi-subject support as a key contribution, but the quantitative evaluation is limited to two subjects; the more-than-two-subject results are only qualitative (Fig. 8). Given that existing baselines are 'unable to reliably handle more subjects' and thus are not compared, the evidence for the claimed multi-subject advantage is thin. Please either provide quantitative results for datasets with three or more subjects or explicitly scope the claim to two subjects, with the >2-subject case presented as a qualitative demonstration.
minor comments (7)
  1. [Author affiliations] The affiliations contain duplicated 'and' (e.g., 'United States of America and and Eyeline Labs'). Please correct.
  2. [§3.2] The citation 'LuxPostFacto [Debevec et al. 2000]' is incorrect: LuxPostFacto is Mei et al. 2025, while Debevec et al. 2000 is the OLAT data source. Please fix the reference.
  3. [§3.2] In the Relighting Model Training paragraph, the OLAT data description cites '[Debevec et al. 2000]' twice; the first mention should likely cite Mei et al. 2025 for the LuxPostFacto hybrid dataset.
  4. [Fig. 1] The text says 'illustrated in 1' — should be 'Fig. 1'.
  5. [§3.2] The sentence 'The model takes the edited keyframe together with a depth sequence extracted from the source video using DepthAnything 2 [Yang et al. 2024a].' is missing a period after 'DepthAnything 2' and the citation placement is awkward.
  6. [§4.3] The user study description says 'each containing 20 samples, with 26 participants,' but Table 1 reports win rates as percentages. Clarify the denominator (e.g., total pairwise responses) and report exact counts or confidence intervals.
  7. [Table 1] The VBench metrics (Subject Consistency, Background Consistency, Temporal Flickering, Human Anatomy) are reported to three decimals without any uncertainty; consider consolidating these or reporting with appropriate precision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: training pairs are self-supervised, evaluation is on held-out videos with independently edited keyframes, and source-pixel conditioning is the intended mechanism rather than a postdicted fit.

full rationale

ID-V2V's derivation chain is not circular. The missing paired data are addressed by constructing training pairs from a single video via a relighting model: cinematic training faces are relit to a regular top-lit condition, and the video model is trained to map that relit face plus keyframe/depth controls back to the original cinematic target (Sec. 3.2). This is self-supervised pair construction, not a postdiction of the evaluation metrics. Evaluation (Sec. 4) uses 100 single-subject and 60 two-subject held-out videos with keyframes edited by an independent image-editing model (Qwen-Image-Edit), and identity/performance metrics compare generated videos against the source. No parameter is fitted to the test set. Using original source face pixels at inference (Sec. 3.3) is an explicit design for the task, not a hidden equivalence: the edited keyframe still must be propagated, and ablations (Table 1) show identity metrics degrade when face-video or normal signals are removed, so scores are not forced by construction. Author-overlapping citations (LuxPostFacto, DifFRelight, OLAT) are implementation/data references, not load-bearing uniqueness claims. The acknowledged failure under extreme source lighting (Sec. 4.3, Fig. 5, Limitations) is an honest boundary on the claim, and the unquantified fidelity of the relighting model is a validation gap, not a circular step. Therefore no circularity is warranted.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The paper is an empirical ML system, so there are no theory-level fitted constants. The loaded assumptions are domain assumptions about illumination invariance, relighting fidelity, and the validity of self-supervised training pairs. No new physical or conceptual entities are introduced.

free parameters (2)
  • Top-lighting HDRI configuration = vertically half-white/half-black HDRI (hand-chosen)
    All training faces are relit to this fixed regular-lighting condition; it defines the inference-time source lighting distribution and is chosen by design rather than learned or justified by data.
  • Face curation thresholds = up to 5 faces/frame; discard videos with >40% frames lacking faces
    These hand-set thresholds shape the identity-control supervision and the multi-subject evaluation; they affect which videos the model can handle.
assumptions (4)
  • domain assumption Facial structure and expression remain invariant under restylization; illumination is the primary permissible facial variation.
    Core premise of §3.2 'Identity Preservation as Relighting'; if legitimate style edits alter facial appearance, the training supervision and task definition are incorrect.
  • domain assumption The image relighting model converts arbitrary cinematic lighting to regular top lighting without changing identity or performance.
    Training pairs rely on this relighting model (§3.2); its failures or identity distortions are inherited by ID-V2V.
  • domain assumption At inference, source videos are captured under approximately regular lighting that matches the top-lit training condition.
    Required for the inference procedure in §3.3; the paper's own §4.3 failure mode (Fig. 5) shows colored/hard-shadow lighting breaks the method.
  • domain assumption Pretrained DepthAnythingV2, DAViD, SCRFD, AdaFace, and EMOCA provide adequate depth, normals, face detection, and expression estimates.
    These models generate the control signals and the evaluation metrics (§3.2, §4.2); errors in them propagate to both training and measurement.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ID-V2V: Identity-Preserving Video Restylization." pith.science (2026). https://pith.science/paper/6Y5DFVDH

@misc{pith2026260722830,
  author       = {Pith},
  title        = {Pith review of: ID-V2V: Identity-Preserving Video Restylization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6Y5DFVDH}},
  note         = {Machine review of arXiv:2607.22830}
}
read the original abstract

In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and lip synchronization. A key obstacle is the absence of paired training data, as identity-preserving restylized video pairs are rare in real-world settings. To address this, we propose a decoupling of source-grounded identity preservation and edit-driven video synthesis. Our key insight is that facial appearance and expression should remain invariant, with illumination being the primary permissible variation. We therefore cast identity preservation as a video relighting problem, while modeling visual edit propagation as controlled video synthesis guided by the edited keyframe. Building on this formulation, we introduce ID-V2V, a video-to-video generative framework integrating complementary control signals: relit facial regions and facial normal maps tightly constrain facial likeness and performance, while edited keyframes and depth sequences enable flexible and temporally coherent generation. This design enables constructing training pairs from a single video, eliminating the need for scarce paired data. Extensive experiments demonstrate that ID-V2V significantly outperforms existing methods in preserving facial likeness and fine-grained facial performance, supports both single- and multi-subject scenarios, and delivers high visual quality, highlighting its potential as a human-centric tool for real-world content production. The code is available at: https://github.com/Eyeline-Labs/ID-V2V.

Figures

Figures reproduced from arXiv: 2607.22830 by the authors.

Figure 1
Figure 1. ID-V2V overview. ID-V2V performs identity-preserving video restylization by taking an edited first frame from a source video that specifies new visual [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Training data construction. Identity-preserving video restylization is [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. ID-V2V architecture. The model conditions on relit facial regions and [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Single-subject comparison. The first row shows the edited first frame and the source video, while the following rows present the generated videos. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Failure case under irregular source lighting. Although the edited first [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Top: Training video with cinematic lighting. Bottom: Relit face video produced by our image based relighting model. The original facial illumination is removed and replaced with regular lighting to simulate the lighting conditions encountered during ID-V2V inference, w…
Figure 7
Figure 7. Figure 7: Two-subject comparison. The first row shows the edited first frame and the source video, while the following rows present the generated videos. ID-V2V [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: More than two subjects. In each pair, the first row is the source video and the second row is generated by ID-V2V, which reliably handles video [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Relighting. For each pair, the first row shows the source video and the second row shows the output generated by ID-V2V. The edited keyframe modifies [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

243 extracted references · 1 canonical work pages

  1. [1]

    ICLR , year=

    Animatediff: Animate your personalized text-to-image diffusion models without specific tuning , author=. ICLR , year=

  2. [2]

    CVPR , year=

    Align your latents: High-resolution video synthesis with latent diffusion models , author=. CVPR , year=

  3. [3]

    arXiv , year=

    Cogvideox: Text-to-video diffusion models with an expert transformer , author=. arXiv , year=

  4. [4]

    ECCV , year=

    Sparsectrl: Adding sparse controls to text-to-video diffusion models , author=. ECCV , year=

  5. [5]

    arXiv , year=

    Stable video diffusion: Scaling latent video diffusion models to large datasets , author=. arXiv , year=

  6. [6]

    2024 , author=

    Video generation models as world simulators. 2024 , author=. URL https://openai. com/research/video-generation-models-as-world-simulators , year=

  7. [7]

    ICLR , year=

    Language Model Beats Diffusion--Tokenizer is Key to Visual Generation , author=. ICLR , year=

  8. [8]

    arXiv , year=

    Latte: Latent diffusion transformer for video generation , author=. arXiv , year=

Show all 243 references
  1. [9]

    CVPR , year=

    Peekaboo: Interactive video generation via masked-diffusion , author=. CVPR , year=

  2. [10]

    SIGGRAPH , year=

    Direct-a-video: Customized video generation with user-directed camera movement and object motion , author=. SIGGRAPH , year=

  3. [11]

    SIGGRAPH , year=

    Motionctrl: A unified and flexible motion controller for video generation , author=. SIGGRAPH , year=

  4. [12]

    SIGGRAPH , year=

    Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling , author=. SIGGRAPH , year=

  5. [13]

    ECCV , year=

    Draganything: Motion control for anything using entity representation , author=. ECCV , year=

  6. [14]

    arXiv , year=

    MotionClone: Training-Free Motion Cloning for Controllable Video Generation , author=. arXiv , year=

  7. [15]

    NeurIPS , year=

    Videocomposer: Compositional video synthesis with motion controllability , author=. NeurIPS , year=

  8. [16]

    ICLR , year=

    Tokenflow: Consistent diffusion features for consistent video editing , author=. ICLR , year=

  9. [17]

    AAAI , year=

    Scalable Motion Style Transfer with Constrained Diffusion Generation , author=. AAAI , year=

  10. [18]

    arXiv , year=

    Anyv2v: A plug-and-play framework for any video-to-video editing tasks , author=. arXiv , year=

  11. [19]

    arXiv , year=

    Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control , author=. arXiv , year=

  12. [20]

    arXiv , year=

    CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation , author=. arXiv , year=

  13. [21]

    ICLR , year=

    How I Warped Your Noise: a Temporally-Correlated Noise Prior for Diffusion Models , author=. ICLR , year=

  14. [22]

    arXiv , year=

    Continuous 3D Perception Model with Persistent State , author=. arXiv , year=

  15. [23]

    arXiv , year=

    SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation , author=. arXiv , year=

  16. [24]

    2022 , booktitle=

    LoRA: Low-Rank Adaptation of Large Language Models , author=. 2022 , booktitle=

  17. [25]

    CVPR , year=

    Space-time diffusion features for zero-shot text-driven motion transfer , author=. CVPR , year=

  18. [26]

    ICCV , year=

    Preserve your own correlation: A noise prior for video diffusion models , author=. ICCV , year=

  19. [27]

    arXiv , year=

    Control-a-video: Controllable text-to-video generation with diffusion models , author=. arXiv , year=

  20. [28]

    arXiv , year=

    Infinite-Resolution Integral Noise Warping for Diffusion Models , author=. arXiv , year=

  21. [29]

    SIGGRAPH Asia , year=

    DifFRelight: Diffusion-Based Facial Performance Relighting , author=. SIGGRAPH Asia , year=

  22. [30]

    URL https://github.com/deep-floyd/IF?tab=readme-ov-file , year=

    DeepFloyd IF , author=. URL https://github.com/deep-floyd/IF?tab=readme-ov-file , year=

  23. [31]

    CVPR , year=

    Vbench: Comprehensive benchmark suite for video generative models , author=. CVPR , year=

  24. [32]

    arXiv , year=

    The 2017 davis challenge on video object segmentation , author=. arXiv , year=

  25. [33]

    CVPR , year=

    Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision , author=. CVPR , year=

  26. [34]

    CVPR , year=

    Wonderjourney: Going from anywhere to everywhere , author=. CVPR , year=

  27. [35]

    ECCV , year=

    Raft: Recurrent all-pairs field transforms for optical flow , author=. ECCV , year=

  28. [36]

    2021 , publisher=

    Learning Blender , author=. 2021 , publisher=

  29. [37]

    NeurIPS , year=

    MotionCraft: Physics-based Zero-Shot Video Generation , author=. NeurIPS , year=

  30. [38]

    International Conference on Learning Representations (ICLR) , year=

    Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting , author=. International Conference on Learning Representations (ICLR) , year=

  31. [39]

    ICLR , year=

    Score-based generative modeling through stochastic differential equations , author=. ICLR , year=

  32. [40]

    NeurIPS , year=

    Denoising diffusion probabilistic models , author=. NeurIPS , year=

  33. [41]

    ICLR , year=

    Denoising diffusion implicit models , author=. ICLR , year=

  34. [42]

    NeurIPS , year=

    Elucidating the design space of diffusion-based generative models , author=. NeurIPS , year=

  35. [43]

    arXiv , year=

    Glide: Towards photorealistic image generation and editing with text-guided diffusion models , author=. arXiv , year=

  36. [44]

    arXiv , year=

    Classifier-free diffusion guidance , author=. arXiv , year=

  37. [45]

    CVPR , year=

    High-resolution image synthesis with latent diffusion models , author=. CVPR , year=

  38. [46]

    arXiv , year=

    Sdxl: Improving latent diffusion models for high-resolution image synthesis , author=. arXiv , year=

  39. [47]

    ICML , year=

    Learning transferable visual models from natural language supervision , author=. ICML , year=

  40. [48]

    JMLR , year=

    Exploring the limits of transfer learning with a unified text-to-text transformer , author=. JMLR , year=

  41. [49]

    CVPR , year=

    Instructpix2pix: Learning to follow image editing instructions , author=. CVPR , year=

  42. [50]

    ICCV , year=

    Adding conditional control to text-to-image diffusion models , author=. ICCV , year=

  43. [51]

    CVPR , year=

    Repurposing diffusion-based image generators for monocular depth estimation , author=. CVPR , year=

  44. [52]

    ICLR , year=

    Sdedit: Guided image synthesis and editing with stochastic differential equations , author=. ICLR , year=

  45. [53]

    Liu, Guan-Horng and Vahdat, Arash and Huang, De-An and Theodorou, Evangelos A and Nie, Weili and Anandkumar, Anima , booktitle=. I \^

  46. [54]

    NeurIPS , year=

    Resshift: Efficient diffusion model for image super-resolution by residual shifting , author=. NeurIPS , year=

  47. [55]

    ICCV , year=

    Scalable diffusion models with transformers , author=. ICCV , year=

  48. [56]

    arXiv , year=

    Videocrafter1: Open diffusion models for high-quality video generation , author=. arXiv , year=

  49. [57]

    CVPR , year=

    Hierarchical spatio-temporal decoupling for text-to-video generation , author=. CVPR , year=

  50. [58]

    ECCV , year=

    Dynamicrafter: Animating open-domain images with video diffusion priors , author=. ECCV , year=

  51. [59]

    arXiv , year=

    StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation , author=. arXiv , year=

  52. [60]

    arXiv , year=

    Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory , author=. arXiv , year=

  53. [61]

    arXiv , year=

    Cameractrl: Enabling camera control for text-to-video generation , author=. arXiv , year=

  54. [62]

    arXiv , year=

    Viewcrafter: Taming video diffusion models for high-fidelity novel view synthesis , author=. arXiv , year=

  55. [63]

    arXiv , year=

    Training-free Camera Control for Video Generation , author=. arXiv , year=

  56. [64]

    arXiv , year=

    ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning , author=. arXiv , year=

  57. [65]

    CVPR , year=

    The unreasonable effectiveness of deep features as a perceptual metric , author=. CVPR , year=

  58. [66]

    SSIM , author=

    Image quality metrics: PSNR vs. SSIM , author=. ICPR , year=

  59. [67]

    ECCV , year=

    Learning blind video temporal consistency , author=. ECCV , year=

  60. [68]

    arXiv , year=

    Cotracker: It is better to track together , author=. arXiv , year=

  61. [69]

    arXiv , year=

    Towards accurate generative models of video: A new metric & challenges , author=. arXiv , year=

  62. [70]

    2024 , eprint=

    CogVLM: Visual Expert for Pretrained Language Models , author=. 2024 , eprint=

  63. [71]

    2022 , eprint=

    Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding , author=. 2022 , eprint=

  64. [72]

    2020 , booktitle=

    NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis , author=. 2020 , booktitle=

  65. [73]

    ACM Trans

    Thomas M\"uller and Alex Evans and Christoph Schied and Alexander Keller , title =. ACM Trans. Graph. , issue_date =. 2022 , pages =. doi:10.1145/3528223.3530127 , publisher =

  66. [74]

    ICCV , year=

    Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields , author=. ICCV , year=

  67. [75]

    ICCV , year=

    Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields , author=. ICCV , year=

  68. [76]

    and Bouaziz, Sofien and Goldman, Dan B and Martin-Brualla, Ricardo and Seitz, Steven M

    Park, Keunhong and Sinha, Utkarsh and Hedman, Peter and Barron, Jonathan T. and Bouaziz, Sofien and Goldman, Dan B and Martin-Brualla, Ricardo and Seitz, Steven M. , title =. ACM Trans. Graph. , issue_date =. 2021 , articleno =

  69. [77]

    Point-Based Neural Rendering with Per-View Optimization

    Kopanas, Georgios and Philip, Julien and Leimkühler, Thomas and Drettakis, George. Point-Based Neural Rendering with Per-View Optimization. Computer Graphics Forum (Proceedings of the Eurographics Symposium on Rendering). 2021

  70. [78]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    Point-nerf: Point-based neural radiance fields , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  71. [79]

    ACM Transactions on Graphics (ToG) , volume=

    Adop: Approximate differentiable one-pixel point rendering , author=. ACM Transactions on Graphics (ToG) , volume=. 2022 , publisher=

  72. [80]

    Computer Graphics Forum , volume=

    Cinematic Gaussians: Real-Time HDR Radiance Fields with Depth , author=. Computer Graphics Forum , volume=. 2024 , organization=

  73. [81]

    arXiv preprint arXiv:2111.14292 , year =

    Deblur-NeRF: Neural Radiance Fields from Blurry Images , author =. arXiv preprint arXiv:2111.14292 , year =

  74. [82]

    Srinivasan and Jonathan T

    Ben Mildenhall and Peter Hedman and Ricardo Martin-Brualla and Pratul P. Srinivasan and Jonathan T. Barron , journal=

  75. [83]

    3D Gaussian Splatting for Real-Time Radiance Field Rendering , journal =

    Kerbl, Bernhard and Kopanas, Georgios and Leimk. 3D Gaussian Splatting for Real-Time Radiance Field Rendering , journal =

  76. [84]

    Conference on Computer Vision and Pattern Recognition (CVPR) , year =

    Yu, Zehao and Chen, Anpei and Huang, Binbin and Sattler, Torsten and Geiger, Andreas , title =. Conference on Computer Vision and Pattern Recognition (CVPR) , year =

  77. [85]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    3D Gaussian Splatting as Markov Chain Monte Carlo , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  78. [86]

    CVPR , year =

    Wu, Guanjun and Yi, Taoran and Fang, Jiemin and Xie, Lingxi and Zhang, Xiaopeng and Wei, Wei and Liu, Wenyu and Tian, Qi and Wang, Xinggang , title =. CVPR , year =

  79. [87]

    Fast Dynamic Radiance Fields with Time-Aware Neural Voxels , year =

    Fang, Jiemin and Yi, Taoran and Wang, Xinggang and Xie, Lingxi and Zhang, Xiaopeng and Liu, Wenyu and Nie. Fast Dynamic Radiance Fields with Time-Aware Neural Voxels , year =

  80. [88]

    ACM Transactions on Graphics (TOG) , volume=

    Bilateral Guided Radiance Field Processing , author=. ACM Transactions on Graphics (TOG) , volume=. 2024 , publisher=

  81. [89]

    International Conference on Learning Representations (ICLR) , year =

    Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting , author =. International Conference on Learning Representations (ICLR) , year =

  82. [90]

    2024 , isbn =

    Duan, Yuanxing and Wei, Fangyin and Dai, Qiyu and He, Yuhang and Chen, Wenzheng and Chen, Baoquan , title =. 2024 , isbn =. doi:10.1145/3641519.3657463 , booktitle =

  83. [91]

    European Conference on Computer Vision (ECCV) , year =

    SuperGaussian: Repurposing Video Models for 3D Super Resolution , author =. European Conference on Computer Vision (ECCV) , year =

  84. [92]

    2023 , eprint=

    ResShift: Efficient Diffusion Model for Image Super-resolution by Residual Shifting , author=. 2023 , eprint=

  85. [93]

    Beardsley and Craig Gotsman and Robert W

    Thabo Beeler and Fabian Hahn and Derek Bradley and Bernd Bickel and Paul A. Beardsley and Craig Gotsman and Robert W. Sumner and Markus H. Gross , title =. 2011 , url =

  86. [94]

    Debevec and Jovan Popovic and Szymon Rusinkiewicz and Wojciech Matusik , title =

    Daniel Vlasic and Pieter Peers and Ilya Baran and Paul E. Debevec and Jovan Popovic and Szymon Rusinkiewicz and Wojciech Matusik , title =. 2009 , url =

  87. [95]

    Comprehensive Facial Performance Capture , journal =

    Graham Fyffe and Tim Hawkins and Chris Watts and Wan. Comprehensive Facial Performance Capture , journal =. 2011 , url =

  88. [96]

    Probabilistic Deformable Surface Tracking from Multiple Videos , booktitle =

    Cedric Cagniart and Edmond Boyer and Slobodan Ilic , editor =. Probabilistic Deformable Surface Tracking from Multiple Videos , booktitle =. 2010 , url =

  89. [97]

    Kirk and Steve Sullivan , title =

    Alvaro Collet and Ming Chuang and Pat Sweeney and Don Gillett and Dennis Evseev and David Calabrese and Hugues Hoppe and Adam G. Kirk and Steve Sullivan , title =. 2015 , url =

  90. [98]

    Performance capture from sparse multi-view video , journal =

    Edilson de Aguiar and Carsten Stoll and Christian Theobalt and Naveed Ahmed and Hans. Performance capture from sparse multi-view video , journal =. 2008 , url =

  91. [99]

    2017 , url =

    Mingsong Dou and Philip Davidson and Sean Ryan Fanello and Sameh Khamis and Adarsh Kowdle and Christoph Rhemann and Vladimir Tankovich and Shahram Izadi , title =. 2017 , url =

  92. [100]

    The relightables: volumetric performance capture of humans with realistic relighting , journal =

    Kaiwen Guo and Peter Lincoln and Philip Davidson and Jay Busch and Xueming Yu and Matt Whalen and Geoff Harvey and Sergio Orts. The relightables: volumetric performance capture of humans with realistic relighting , journal =. 2019 , url =

  93. [101]

    2008 , url =

    Daniel Vlasic and Ilya Baran and Wojciech Matusik and Jovan Popovic , title =. 2008 , url =

  94. [102]

    Deep blending for free-viewpoint image-based rendering , journal =

    Peter Hedman and Julien Philip and True Price and Jan. Deep blending for free-viewpoint image-based rendering , journal =. 2018 , url =

  95. [103]

    2019 , url =

    Zexiang Xu and Sai Bi and Kalyan Sunkavalli and Sunil Hadap and Hao Su and Ravi Ramamoorthi , title =. 2019 , url =

  96. [104]

    Saragih and Gabriel Schwartz and Andreas M

    Stephen Lombardi and Tomas Simon and Jason M. Saragih and Gabriel Schwartz and Andreas M. Lehrmann and Yaser Sheikh , title =. 2019 , url =

  97. [105]

    Deep relightable textures: volumetric performance capture with neural rendering , journal =

    Abhimitra Meka and Rohit Pandey and Christian H. Deep relightable textures: volumetric performance capture with neural rendering , journal =. 2020 , url =

  98. [106]

    Mixture of volumetric primitives for efficient neural rendering , journal =

    Stephen Lombardi and Tomas Simon and Gabriel Schwartz and Michael Zollh. Mixture of volumetric primitives for efficient neural rendering , journal =. 2021 , url =

  99. [107]

    2023 , url =

    Sida Peng and Yunzhi Yan and Qing Shuai and Hujun Bao and Xiaowei Zhou , title =. 2023 , url =

  100. [108]

    2022 , url =

    Fuqiang Zhao and Yuheng Jiang and Kaixin Yao and Jiakai Zhang and Liao Wang and Haizhao Dai and Yuhui Zhong and Yingliang Zhang and Minye Wu and Lan Xu and Jingyi Yu , title =. 2022 , url =

  101. [109]

    CoRR , volume =

    Haotong Lin and Sida Peng and Zhen Xu and Tao Xie and Xingyi He and Hujun Bao and Xiaowei Zhou , title =. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2310.08585 , eprinttype =

  102. [110]

    HumanRF: High-Fidelity Neural Radiance Fields for Humans in Motion , journal =

    Mustafa Isik and Martin R. HumanRF: High-Fidelity Neural Radiance Fields for Humans in Motion , journal =

  103. [111]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    Jiang, Yuheng and Shen, Zhehao and Wang, Penghao and Su, Zhuo and Hong, Yu and Zhang, Yingliang and Yu, Jingyi and Xu, Lan , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =. 2024 , pages =

  104. [112]

    Dense correspondence finding for parametrization-free animation reconstruction from video , booktitle =

    Naveed Ahmed and Christian Theobalt and Christian R. Dense correspondence finding for parametrization-free animation reconstruction from video , booktitle =. 2008 , url =

  105. [113]

    Barron and Sofien Bouaziz and Dan B

    Keunhong Park and Utkarsh Sinha and Jonathan T. Barron and Sofien Bouaziz and Dan B. Goldman and Steven M. Seitz and Ricardo Martin. Nerfies: Deformable Neural Radiance Fields , booktitle =. 2021 , url =

  106. [114]

    Neural Radiance Flow for 4D View Synthesis and Video Processing , booktitle =

    Yilun Du and Yinan Zhang and Hong. Neural Radiance Flow for 4D View Synthesis and Video Processing , booktitle =

  107. [115]

    K-Planes: Explicit Radiance Fields in Space, Time, and Appearance , booktitle =

    Sara Fridovich. K-Planes: Explicit Radiance Fields in Space, Time, and Appearance , booktitle =. 2023 , url =

  108. [116]

    2024 , url =

    Zhen Xu and Sida Peng and Haotong Lin and Guangzhao He and Jiaming Sun and Yujun Shen and Hujun Bao and Xiaowei Zhou , title =. 2024 , url =

  109. [117]

    Barron and Sofien Bouaziz and Dan B

    Keunhong Park and Utkarsh Sinha and Peter Hedman and Jonathan T. Barron and Sofien Bouaziz and Dan B. Goldman and Ricardo Martin. HyperNeRF: a higher-dimensional representation for topologically varying neural radiance fields , journal =. 2021 , url =

  110. [118]

    2024 , url =

    Zhen Xu and Yinghao Xu and Zhiyuan Yu and Sida Peng and Jiaming Sun and Hujun Bao and Xiaowei Zhou , title =. 2024 , url =

  111. [119]

    Debevec , editor =

    Mingming He and Pascal Clausen and Ahmet Levent Tasel and Li Ma and Oliver Pilarski and Wenqi Xian and Laszlo Rikker and Xueming Yu and Ryan Burgert and Ning Yu and Paul E. Debevec , editor =. DifFRelight: Diffusion-Based Facial Performance Relighting , booktitle =. 2024 , url =

  112. [120]

    2024 , date =

    Maya , url =. 2024 , date =

  113. [121]

    Gen-3 Alpha , year =

  114. [122]

    1998 , isbn =

    Guenter, Brian and Grimm, Cindy and Wood, Daniel and Malvar, Henrique and Pighin, Fredric , title =. 1998 , isbn =. doi:10.1145/280814.280822 , booktitle =

  115. [123]

    and Rander, P

    Kanade, T. and Rander, P. and Narayanan, P.J. , journal=. Virtualized reality: constructing virtual worlds from real scenes , year=

  116. [124]

    2006 , booktitle =

    Einarsson, Per and Chabert, Charles-Felix and Jones, Andrew and Ma, Wan-Chun and Lamond, Bruce and Hawkins, Tim and Bolas, Mark and Sylwan, Sebastian and Debevec, Paul , title =. 2006 , booktitle =

  117. [125]

    2019 , journal =

    Guo, Kaiwen and Lincoln, Peter and Davidson, Philip and Busch, Jay and Yu, Xueming and Whalen, Matt and Harvey, Geoff and Orts-Escolano, Sergio and Pandey, Rohit and Dourgarian, Jason and Tang, Danhang and Tkach, Anastasia and Kowdle, Adarsh and Cooper, Emily and Dou, Mingsong...

  118. [126]

    ICCV , year=

    Swin Transformer: Hierarchical Vision Transformer using Shifted Windows , author=. ICCV , year=

  119. [127]

    CVPR , year=

    SinSR: diffusion-based image super-resolution in a single step , author=. CVPR , year=

  120. [128]

    2024 , journal=

    HunyuanVideo: A Systematic Framework For Large Video Generative Models , author=. 2024 , journal=

  121. [129]

    , year =

    Wan: Open and Advanced Large-Scale Video Generative Models , author =. , year =

  122. [130]

    arXiv , year=

    Cat4d: Create anything in 4d with multi-view video diffusion models , author=. arXiv , year=

  123. [131]

    arXiv , year=

    Wonderland: Navigating 3D Scenes from a Single Image , author=. arXiv , year=

  124. [132]

    CVPR , year=

    Dreamvideo: Composing your dream videos with customized subject and motion , author=. CVPR , year=

  125. [133]

    ECCV , year=

    VideoStudio: Generating Consistent-Content and Multi-Scene Videos , author=. ECCV , year=

  126. [134]

    arXiv , year=

    Magic-me: Identity-specific video customized diffusion , author=. arXiv , year=

  127. [135]

    CVPR , year=

    Identity-Preserving Text-to-Video Generation by Frequency Decomposition , author=. CVPR , year=

  128. [136]

    arXiv , year=

    Multi-subject Open-set Personalization in Video Generation , author=. arXiv , year=

  129. [137]

    arXiv , year=

    Dynamic Concepts Personalization from Single Videos , author=. arXiv , year=

  130. [138]

    CVPR , year=

    Person image synthesis via denoising diffusion model , author=. CVPR , year=

  131. [139]

    CVPR , year=

    Dreampose: Fashion video synthesis with stable diffusion , author=. CVPR , year=

  132. [140]

    ACM Computing Surveys , year=

    Appearance and pose-guided human generation: A survey , author=. ACM Computing Surveys , year=

  133. [141]

    SIGGRAPH Asia , year=

    TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model , author=. SIGGRAPH Asia , year=

  134. [142]

    CVPR , year=

    One-shot free-view neural talking-head synthesis for video conferencing , author=. CVPR , year=

  135. [143]

    ECCV , year=

    Neural voice puppetry: Audio-driven facial reenactment , author=. ECCV , year=

  136. [144]

    arXiv , year=

    Audio2head: Audio-driven one-shot talking-head generation with natural head motion , author=. arXiv , year=

  137. [145]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    Progressive disentangled representation learning for fine-grained controllable talking head synthesis , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  138. [146]

    arXiv , year=

    Motionbooth: Motion-aware customized text-to-video generation , author=. arXiv , year=

  139. [147]

    CVPR , year=

    Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise , author=. CVPR , year=

  140. [148]

    arXiv preprint arXiv:2301.11280 , year=

    Text-to-4d dynamic scene generation , author=. arXiv preprint arXiv:2301.11280 , year=

  141. [149]

    CVPR , year=

    4d-fy: Text-to-4d generation using hybrid score distillation sampling , author=. CVPR , year=

  142. [150]

    ECCV , year=

    Tc4d: Trajectory-conditioned text-to-4d generation , author=. ECCV , year=

  143. [151]

    CVPR , year=

    A unified approach for text-and image-guided 4d scene generation , author=. CVPR , year=

  144. [152]

    arXiv , year=

    Diffusion4d: Fast spatial-temporal consistent 4d generation via video diffusion models , author=. arXiv , year=

  145. [153]

    NeurIPS , year=

    L4gm: Large 4d gaussian reconstruction model , author=. NeurIPS , year=

  146. [154]

    arXiv , year=

    Dimensionx: Create any 3d and 4d scenes from a single image with controllable video diffusion , author=. arXiv , year=

  147. [155]

    ICLR , year=

    Dreamfusion: Text-to-3d using 2d diffusion , author=. ICLR , year=

  148. [156]

    CVPR , year=

    Neural scene flow fields for space-time view synthesis of dynamic scenes , author=. CVPR , year=

  149. [157]

    CVPR , year=

    Dynibar: Neural dynamic image-based rendering , author=. CVPR , year=

  150. [158]

    arXiv , year=

    Shape of motion: 4d reconstruction from a single video , author=. arXiv , year=

  151. [159]

    arXiv , year=

    Mosca: Dynamic gaussian fusion from casual videos via 4d motion scaffolds , author=. arXiv , year=

  152. [160]

    arXiv , year=

    Lrm: Large reconstruction model for single image to 3d , author=. arXiv , year=

  153. [161]

    ECCV , year=

    Gs-lrm: Large reconstruction model for 3d gaussian splatting , author=. ECCV , year=

  154. [162]

    3DV , year=

    Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis , author=. 3DV , year=

  155. [163]

    arXiv , year=

    Reconx: Reconstruct any scene from sparse views with video diffusion model , author=. arXiv , year=

  156. [164]

    NeurIPS , year=

    Light field networks: Neural scene representations with single-evaluation rendering , author=. NeurIPS , year=

  157. [165]

    CVPR , year=

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation , author=. CVPR , year=

  158. [166]

    CVPR , year=

    Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset , author=. CVPR , year=

  159. [167]

    2025 , url =

    Poly Haven , title =. 2025 , url =

  160. [168]

    SIGGRAPH , year=

    Acquiring the reflectance field of a human face , author=. SIGGRAPH , year=

  161. [169]

    , author=

    Total relighting: learning to relight portraits for background replacement. , author=. TOG , year=

  162. [170]

    ECCV , year=

    Omg: Occlusion-friendly personalized multi-concept generation in diffusion models , author=. ECCV , year=

  163. [171]

    ECCV , year=

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection , author=. ECCV , year=

  164. [172]

    arXiv , year=

    Sam 2: Segment anything in images and videos , author=. arXiv , year=

  165. [173]

    and Tulyakov, Sergey , title =

    Bahmani, Sherwin and Skorokhodov, Ivan and Qian, Guocheng and Siarohin, Aliaksandr and Menapace, Willi and Tagliasacchi, Andrea and Lindell, David B. and Tulyakov, Sergey , title =. CVPR , year =

  166. [174]

    and Tulyakov, Sergey , title =

    Bahmani, Sherwin and Skorokhodov, Ivan and Siarohin, Aliaksandr and Menapace, Willi and Qian, Guocheng and Vasilkovsky, Michael and Lee, Hsin-Ying and Wang, Chaoyang and Zou, Jiaxu and Tagliasacchi, Andrea and Lindell, David B. and Tulyakov, Sergey , title =. ICLR , year =

  167. [175]

    NIPS , year=

    Video Diffusion Models are Training-free Motion Interpreter and Controller , author=. NIPS , year=

  168. [176]

    2024 , journal=

    Controlling Space and Time with Diffusion Models , author=. 2024 , journal=

  169. [177]

    ICLR , year=

    Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport , author=. ICLR , year=

  170. [178]

    2024 , journal=

    MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model , author=. 2024 , journal=

  171. [179]

    2024 , journal=

    Motion Prompting: Controlling Video Generation with Motion Trajectories , author=. 2024 , journal=

  172. [180]

    arXiv , year=

    Consistentid: Portrait generation with multimodal fine-grained identity preserving , author=. arXiv , year=

  173. [181]

    CVPR , year=

    Videobooth: Diffusion-based video generation with image prompts , author=. CVPR , year=

  174. [182]

    2021 , journal=

    Sample and Computation Redistribution for Efficient Face Detection , author=. 2021 , journal=

  175. [183]

    CVPR , year=

    Adaface: Quality adaptive margin for face recognition , author=. CVPR , year=

  176. [184]

    IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

    Implicit Neural Representations with Structured Latent Codes for Human Body Modeling , author=. IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

  177. [185]

    CVPR , year=

    Neural Body: Implicit Neural Representations with Structured Latent Codes for Novel View Synthesis of Dynamic Humans , author=. CVPR , year=

  178. [186]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    Gaussianavatars: Photorealistic head avatars with rigged 3d gaussians , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  179. [187]

    ACM SIGGRAPH Conference Proceedings, Denver, CO, United States, July 28 - August 1, 2024 , year =

    Shengjie Ma and Yanlin Weng and Tianjia Shao and Kun Zhou , title =. ACM SIGGRAPH Conference Proceedings, Denver, CO, United States, July 28 - August 1, 2024 , year =

  180. [188]

    arXiv preprint arXiv:2402.09368 , year=

    Magic-me: Identity-specific video customized diffusion , author=. arXiv preprint arXiv:2402.09368 , year=

  181. [189]

    The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

    Identity Decoupling for Multi-Subject Personalization of Text-to-Image Models , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

  182. [190]

    arXiv preprint arXiv:2503.10592 , year=

    CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models , author=. arXiv preprint arXiv:2503.10592 , year=

  183. [191]

    arXiv preprint arXiv:2502.11079 , year=

    Phantom: Subject-consistent video generation via cross-modal alignment , author=. arXiv preprint arXiv:2502.11079 , year=

  184. [192]

    VideoAlchemy: Open-set Personalization in Video Generation , author=

  185. [193]

    arXiv preprint arXiv:2501.04698 , year=

    ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning , author=. arXiv preprint arXiv:2501.04698 , year=

  186. [194]

    arXiv preprint arXiv:1805.09817 , year=

    Stereo magnification: Learning view synthesis using multiplane images , author=. arXiv preprint arXiv:1805.09817 , year=

  187. [195]

    The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

    Humanvid: Demystifying training data for camera-controllable human image animation , author=. The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

  188. [196]

    arXiv preprint arXiv:2501.04001 , year=

    Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos , author=. arXiv preprint arXiv:2501.04001 , year=

  189. [197]

    arXiv preprint arXiv:2503.11647 , year=

    ReCamMaster: Camera-Controlled Generative Rendering from A Single Video , author=. arXiv preprint arXiv:2503.11647 , year=

  190. [198]

    arXiv preprint arXiv:2503.05638 , year=

    TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models , author=. arXiv preprint arXiv:2503.05638 , year=

  191. [199]

    Environmental Psychology & Nonverbal Behavior , year=

    Facial action coding system , author=. Environmental Psychology & Nonverbal Behavior , year=

  192. [200]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    Emoca: Emotion driven monocular face capture and animation , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  193. [201]

    arXiv preprint arXiv:2601.03233 , year=

    LTX-2: Efficient Joint Audio-Visual Foundation Model , author=. arXiv preprint arXiv:2601.03233 , year=

  194. [202]

    Proceedings of the SIGGRAPH Asia 2025 Conference Papers , pages=

    Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Volumetric Performance Captures , author=. Proceedings of the SIGGRAPH Asia 2025 Conference Papers , pages=

  195. [203]

    arXiv preprint arXiv:2501.13452 , year=

    EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion , author=. arXiv preprint arXiv:2501.13452 , year=

  196. [204]

    arXiv preprint arXiv:2509.08519 , year=

    Humo: Human-centric video generation via collaborative multi-modal conditioning , author=. arXiv preprint arXiv:2509.08519 , year=

  197. [205]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    Magicmirror: Id-preserved video generation in video diffusion transformers , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  198. [206]

    arXiv preprint arXiv:2404.15275 , year=

    Id-animator: Zero-shot identity-preserving human video generation , author=. arXiv preprint arXiv:2404.15275 , year=

  199. [207]

    arXiv preprint arXiv:2509.14055 , year=

    Wan-animate: Unified character animation and replacement with holistic replication , author=. arXiv preprint arXiv:2509.14055 , year=

  200. [208]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    Animate anyone: Consistent and controllable image-to-video synthesis for character animation , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  201. [209]

    arXiv preprint arXiv:2503.07598 , year=

    Vace: All-in-one video creation and editing , author=. arXiv preprint arXiv:2503.07598 , year=

  202. [210]

    arXiv preprint arXiv:2507.12956 , year=

    Fantasyportrait: Enhancing multi-character portrait animation with expression-augmented diffusion transformers , author=. arXiv preprint arXiv:2507.12956 , year=

  203. [211]

    arXiv preprint arXiv:2511.19320 , year=

    SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation , author=. arXiv preprint arXiv:2511.19320 , year=

  204. [212]

    arXiv preprint arXiv:2508.02324 , year=

    Qwen-image technical report , author=. arXiv preprint arXiv:2508.02324 , year=

  205. [213]

    URL https://developers.googleblog.com/en/introducing-gemini-2-5-flash-image/ , year=

    Introducing gemini 2.5 flash image, our state-of-the-art image model , author=. URL https://developers.googleblog.com/en/introducing-gemini-2-5-flash-image/ , year=

  206. [214]

    Advances in Neural Information Processing Systems , volume=

    Depth anything v2 , author=. Advances in Neural Information Processing Systems , volume=

  207. [215]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    DAViD: Data-efficient and Accurate Vision Models from Synthetic Data , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  208. [216]

    5-vl technical report , author=

    Qwen2. 5-vl technical report , author=. arXiv preprint arXiv:2502.13923 , year=

  209. [217]

    arXiv preprint arXiv:2507.06261 , year=

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities , author=. arXiv preprint arXiv:2507.06261 , year=

  210. [218]

    arXiv preprint arXiv:2503.21755 , year=

    Vbench-2.0: Advancing video generation benchmark suite for intrinsic faithfulness , author=. arXiv preprint arXiv:2503.21755 , year=

  211. [219]

    Poly Haven , author =

  212. [220]

    arXiv preprint arXiv:2512.05115 , year=

    Light-X: Generative 4D Video Rendering with Camera and Illumination Control , author=. arXiv preprint arXiv:2512.05115 , year=

  213. [221]

    Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year =

    Light-A-Video: Training-free Video Relighting via Progressive Light Fusion , author =. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , year =

  214. [222]

    arXiv preprint arXiv:2501.16330 , year=

    RelightVid: Temporal-consistent diffusion model for video relighting , author=. arXiv preprint arXiv:2501.16330 , year=

  215. [223]

    arXiv preprint arXiv:2506.15673 , year=

    UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting , author=. arXiv preprint arXiv:2506.15673 , year=

  216. [224]

    arXiv preprint arXiv:2511.06271 , year=

    Relightmaster: Precise video relighting with multi-plane light images , author=. arXiv preprint arXiv:2511.06271 , year=

  217. [225]

    arXiv preprint arXiv:2508.12945 , year=

    Lumen: Consistent video relighting and harmonious background replacement with video generative models , author=. arXiv preprint arXiv:2508.12945 , year=

  218. [226]

    ACM SIGGRAPH Asia Conference Papers , year =

    3DPR: Single Image 3D Portrait Relighting with Generative Priors , author =. ACM SIGGRAPH Asia Conference Papers , year =

  219. [227]

    ACM Transactions on Graphics (TOG) , volume =

    Lite2Relight: 3D-aware Single Image Portrait Relighting , author =. ACM Transactions on Graphics (TOG) , volume =

  220. [228]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

    DiFaReli++: Diffusion Face Relighting with Consistent Cast Shadows , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

  221. [229]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

    Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , year =

  222. [230]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

    Image Style Transfer Using Convolutional Neural Networks , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

  223. [231]

    European Conference on Computer Vision , pages=

    Perceptual Losses for Real-Time Style Transfer and Super-Resolution , author=. European Conference on Computer Vision , pages=. 2016 , organization=

  224. [232]

    Proceedings of the IEEE International Conference on Computer Vision , pages=

    Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization , author=. Proceedings of the IEEE International Conference on Computer Vision , pages=

  225. [233]

    German Conference on Pattern Recognition , pages=

    Artistic Style Transfer for Videos , author=. German Conference on Pattern Recognition , pages=. 2016 , organization=

  226. [234]

    ACM Transactions on Graphics , volume=

    Stylizing Video by Example , author=. ACM Transactions on Graphics , volume=. 2019 , publisher=

  227. [235]

    Computer Graphics Forum , volume=

    StructuReiser: A Structure-preserving Video Stylization Method , author=. Computer Graphics Forum , volume=. 2025 , organization=

  228. [236]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  229. [237]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year=

    Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year=

  230. [238]

    Anand, Tharun and Garg, Aryan and Mitra, Kaushik , booktitle=

  231. [239]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    The Devil Is in the Details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  232. [240]

    Liu, Pengwei and Yuan, Hangjie and Dong, Bo and Xing, Jiazheng and Wang, Jinwang and Zhao, Rui and Chen, Weihua and Wang, Fan , booktitle=

  233. [241]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    Relightful Harmonization: Lighting-Aware Portrait Background Replacement , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  234. [242]

    ACM Transactions on Graphics , volume=

    Generative Portrait Shadow Removal , author=. ACM Transactions on Graphics , volume=. 2024 , publisher=

  235. [243]

    Kim, Hoon and Jang, Minje and Yoon, Wonjun and Lee, Jisoo and Na, Donghyun and Woo, Sanghyun , booktitle=

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.