REVIEW 4 major objections 5 minor 94 references
Seeing Voices: Generating A-Roll Video from Audio with Mirage
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Mirage claims a single diffusion transformer can synthesize photorealistic, expressive talking-person video directly from raw audio, without a reference image or domain-specific losses.
desk verdict Genuinely new setting and a clean architecture recipe, but the central comparative claim rests on selected stills and the authors' own inspection — no baselines, no user study, no metrics. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the asymmetric self-attention block inherited from Mochi, an open-source text-to-video Diffusion Transformer, and extended to multiple modalities: video, audio, text, and reference-image tokens are concatenated along the sequence dimension and attend to one another through a single self-attention operation, with each modality given its own learned rotary position encoding (RoPE) and MLP. This design makes cross-modal interaction a default property of the architecture rather than something added by specialized cross-attention, and it is trained with latent flow matching under a simple warm-up and stitching schedule. The same attention block learns to balance modalities, so conditioning on audio, text, or reference images requires no architecture or loss changes.
What would settle it
Feed Mirage a set of audio tracks that the filter was designed to remove — singing, speech with heavy background noise, or a low-motion monologue — and compare lip-sync accuracy and human-rated expressiveness against clips that pass the SyncNet and motion filters; a large drop on the filtered-out categories would show the performance depends on the filter rather than on general audio-to-video ability.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that the asymmetric self-attention layers of a Diffusion Transformer can be trained to balance attention among video, audio, text, and reference-image tokens, so that no cross-attention blocks or domain-specific components are required to synthesize video from audio. Mirage encodes video with a 3D spatiotemporal VAE, audio with wav2vec embeddings, text with T5, and generates 4-8 second 720p portrait videos by flow matching. When trained on a heavily filtered A-roll dataset of single-speaker scenes, the model produces outputs the paper describes as photorealistic and expressive: accurate plosive lip closures, fluid co-articulation in tongue twisters, natural blinks and gaze, emotion-matched facial expressions, semantically timed gestures, and — when text is omitted — plausible reconstructions of speaker appearance and indoor/outdoor environment from acoustic properties such as reverberation and background noise. The paper explicitly notes that audio-only outputs are visually degraded compared with text-audio conditioning.
Load-bearing premise
The model's ability to generate expressive A-roll from arbitrary audio depends on the filtered training data being representative of natural talking-person footage, yet the filtering removes short scenes, low-motion clips, text overlays, screen splits, and clips with low lip-sync confidence — and the paper notes that singing or background noise often lowers that confidence.
Editorial extensions
If this is right
- Audio-only conditioning is enough to infer speaker appearance, scene type, and even microphone-relative motion, so a text description is not required for plausible A-roll.
- Because conditioning modalities are just token streams concatenated into self-attention, the same recipe extends to reference-image conditioning and could absorb new modalities without architecture changes.
- Paired with a text-to-speech model, Mirage turns raw text into a full multimodal video: script, voice, and synchronized on-screen performance.
- When text and audio conflict, the model interpolates between them, tending to side with vocal characteristics for speaker appearance, while aligned text and audio give the most realistic outputs.
Reading between the lines
- If the generalization claim holds, audio-to-video could be treated as a general token-mixing problem, collapsing gesture synthesis, talking-head animation, and scene reconstruction into a single model rather than separate pipelines.
- A direct test of the paper's thesis would hold the architecture fixed and vary only the data filter: if expressive singing or noisy-location speech generate poor visuals, the claimed from-scratch capability is narrower than stated.
- The paper's qualitative evidence for gesture-semantic alignment is explicitly speculative about mechanism; a controlled study varying utterance content while holding prosody fixed could determine whether gestures track meaning or just vocal dynamics.
- Applying the same recipe to non-human subjects, such as animals or animated characters, would test whether the self-attention mixing generalizes beyond people or depends on the face-centric training distribution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Mirage, a 10-billion-parameter Diffusion Transformer (DiT) for generating A-roll video from audio, optionally conditioned on text and reference images. The architecture extends Mochi-style asymmetric self-attention to joint sequences of audio, text, and video tokens, trained with latent flow matching. The paper also describes a large-scale data processing and filtering pipeline, along with qualitative evaluations of plosive articulation, eye behavior, emotional expression, coarticulation, paralinguistic sounds, audio-only conditioning, mismatched text/audio, and gesture semantics. The central claims are that Mirage produces photorealistic, expressive video from raw audio alone and achieves 'superior subjective quality' relative to domain-specific talking-head and gesture-generation methods, while avoiding audio-specific architecture or loss components.
Significance. If substantiated, the unified asymmetric self-attention recipe and the scalable data pipeline would be practically significant, potentially simplifying audio-to-video and multimodal video generation. The paper is commendably explicit about the speculative nature of some claims (e.g., Sec. 5.8), and it provides a public video gallery and product links for readers to inspect outputs directly. However, the central comparative claim of superior subjective quality is not supported by any controlled evaluation: there is no user study, no baseline comparison, and no quantitative metric such as SyncNet confidence or offset on generated videos. As it stands, the significance of the method relative to existing talking-head and gesture-generation systems remains unestablished.
major comments (4)
- [Sec. 5 (Results, opening paragraph)] The paper defines four evaluation criteria (prompt adherence, subject facial details, subject body motion, background and prop fidelity) but never operationalizes them: no scores, no inter-rater agreement, no error bars, and no baseline systems are presented. The abstract and introduction claim 'superior subjective quality' to methods with audio-specific architectures or losses, yet the results section contains only selected still frames, qualitative statements such as 'Mirage demonstrates exceptional precision' (Sec. 5.1) and 'empirical observations indicate' (Sec. 5.2), and a link to a video gallery. This is the central claim of the paper, and it is unsupported without a human preference study or controlled comparison against representative baselines such as VASA-1, EMO, Hallo3, or SadTalker. I request that the authors add a properly controlled evaluation, including at minimum a human preference study and quantitative metrics.
- [Sec. 5.1 and Sec. 5.4 (Plosive/viseme dynamics and coarticulation)] The claims of precise plosive articulation and smooth coarticulation rest on author visual inspection of selected still frames. Although SyncNet is used to filter the training data (Sec. 3.7.3), no SyncNet confidence or offset score, or any other lip-sync metric, is reported on generated videos. Please provide quantitative lip-sync evaluation (e.g., SyncNet confidence/offset on a sample of generated videos) and, ideally, a blind rating study of articulation accuracy to support the claim of 'exceptional precision' at plosive articulation.
- [Sec. 3.7.3 (Data filtering)] The data filtering pipeline excludes scenes under 2 seconds, low P-frame packet size (low motion), text overlays, screen splits, and clips with low SyncNet confidence. The paper itself notes (Sec. 3.7.3) that background noise or singing yields lower SyncNet confidence scores. If these filters preferentially remove expressive, atypical, or noisy performances, the central generalization claim that Mirage 'excels at generating realistic, expressive output imagery from scratch given an audio input' is not established for such inputs. Please report dataset statistics after each filtering stage and evaluate Mirage on out-of-distribution audio (e.g., singing, heavy background noise, short utterances, rapidly cut scenes) to characterize the actual operating range of the model.
- [Sec. 1 and Sec. 3.4 (Warm-up and stitching recipe)] The paper's central technical contribution is described as a 'simple warm-up and stitching recipe' for training asymmetric self-attention to balance audio, text, and video modalities, either from scratch or from silent-video checkpoints. However, the recipe is never specified in enough detail to be reproduced or validated, and no ablation is presented to show that it is necessary or that it generalizes from silent-video checkpoints to audio-conditioned generation. Please provide the full recipe (warm-up schedule, stitching procedure, hyperparameters) and include ablations comparing training from scratch versus finetuning, and with versus without the proposed recipe.
minor comments (5)
- [Sec. 3.7.4 (Video captioning)] The prompts used for the two VLM passes (static keyframes and dense grid) are not included; providing them would improve reproducibility of the captioning pipeline.
- [Sec. 5.10 (Background and prop fidelity)] The statement that Mirage 'tends to struggle when excessive complexity is added to the prompt' is not supported by examples or metrics; please specify what failure modes were observed and, if possible, provide a systematic comparison across prompt complexity levels.
- [Sec. 5.6 (Audio-only conditioning)] The phrase 'The model most likely correlates correct environmental context' is presented as speculation; please either add evidence (e.g., classification accuracy or human judgments of environment matching) or rephrase as a hypothesis.
- [Sec. 4.2.2 (Quantization)] The claim that FP8 matmul quantization introduces 'minimal visual quality degradation' is not accompanied by any quantitative evaluation; a perceptual metric or small user study would make this assertion verifiable.
- [General] The manuscript contains duplicated sentences and formatting artifacts (e.g., repeated lines in the abstract and introduction, broken URLs in the text); a thorough proofread and cleanup is needed before publication.
Circularity Check
No significant circularity: the method is an empirical generative system trained with standard flow matching, and no claimed prediction reduces to its inputs by construction.
full rationale
Mirage is an empirical system, not a derivation. The training objective (Sec. 3.5) is latent flow matching: the velocity prediction is trained to map Gaussian noise to latent video, conditioned on audio and text tokens. The conditioning audio is encoded by a fixed pretrained wav2vec model (Sec. 3.2) and text by T5 (Sec. 3.3), both external to the paper. The generated video is not a re-display of a fitted parameter: the model is asked to synthesize video tokens from input audio/text, and the loss is computed against video latents in the training distribution. No equation in the paper defines an output in terms of the same quantity it claims to predict. The data pipeline (Sec. 3.7) filters training clips using SyncNet and motion/overlay detectors, but none of these filters are presented as evaluating the generation claim. The evaluation section (Sec. 5) is qualitative and relies on the authors' visual inspection, which weakens the evidence for subjective-quality claims, but this is an evidence gap rather than a circular reduction: the claims are not derived from the evaluation criteria by construction. Self-citations in related work (e.g., Jaegle et al.) are historical context and not load-bearing for Mirage's central claim. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (4)
- SyncNet confidence and offset acceptance thresholds =
not specified
- P-frame mean packet size threshold =
predefined threshold, value not specified
- Text overlay area threshold =
1% of any frame area
- Minimum scene duration =
2 seconds
assumptions (4)
- standard math Flow matching / rectified flow objective is a valid generative training target.
- domain assumption wav2vec 2.0 audio features contain enough identity, prosody, and environment information for visual generation.
- domain assumption The filtered single-speaker A-roll dataset is representative of natural talking performances.
- ad hoc to paper The warm-up and stitching recipe for asymmetric self-attention generalizes from silent-video checkpoints to audio-conditioned generation.
Cite this review
Pith. "Pith review of Seeing Voices: Generating A-Roll Video from Audio with Mirage." pith.science (2026). https://pith.science/paper/CSN6EL4R
@misc{pith2026250608279,
author = {Pith},
title = {Pith review of: Seeing Voices: Generating A-Roll Video from Audio with Mirage},
year = {2026},
howpublished = {\url{https://pith.science/paper/CSN6EL4R}},
note = {Machine review of arXiv:2506.08279}
}
read the original abstract
From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see (the video's image sequence). Current approaches to video generation either ignore sound to focus on general-purpose but silent image sequence generation or address both visual and audio elements but focus on restricted application domains such as re-dubbing. We introduce Mirage, an audio-to-video foundation model that excels at generating realistic, expressive output imagery from scratch given an audio input. When integrated with existing methods for speech synthesis (text-to-speech, or TTS), Mirage results in compelling multimodal video. When trained on audio-video footage of people talking (A-roll) and conditioned on audio containing speech, Mirage generates video of people delivering a believable interpretation of the performance implicit in input audio. Our central technical contribution is a unified method for training self-attention-based audio-to-video generation models, either from scratch or given existing weights. This methodology allows Mirage to retain generality as an approach to audio-to-video generation while producing outputs of superior subjective quality to methods that incorporate audio-specific architectures or loss components specific to people, speech, or details of how images or audio are captured. We encourage readers to watch and listen to the results of Mirage for themselves (see paper and comments for links).
Reference graph
Works this paper leans on
-
[1]
Botev, A
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 4343 A. Botev, A. Jaegle, P. Wirnsberger, D. Hennes, and I. Higgins. Which priors matter? benchmarking A. Botev, A. Jaegle, P. Wirnsberger, D. Hennes, and I. Higgins. Which priors matter? benchmarking models for learning latent...
2021
-
[9]
McGraw-Hill, 2009., chapter
work page 2009
-
[10]
Chhatre, R
K. Chhatre, R. Danˇeˇcek, N. Athanasiou, G. Becherini, C. Peters, M. J. Black, and T. Bolkart. K. Chhatre, R. Danˇeˇcek, N. Athanasiou, G. Becherini, C. Peters, M. J. Black, and T. Bolkart. Emotional speech-driven 3d body animation via disentangled latent diffusion. In Emotional speech-driven 3d body animation via disentangled latent diffusion. In Proceed...
2024
-
[13]
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 4444 G. Cloud. Cloud storage consistency. G. Cloud. Cloud storage consistency. https://cloud.google.com/storage/docs/consistencyhttps://cloud.google.com/storage/docs/consistency„ 2025.„
2025
-
[14]
M. M. Cohen, D. W. Massaro, and R. Clark. Training a talking head. In M. M. Cohen, D. W. Massaro, and R. Clark. Training a talking head. In Proceedings of the International Proceedings of the International Conference on Multimodal InterfacesConference on Multimodal Interfaces, 2002.,
2002
-
[16]
Denton and R
E. Denton and R. Fergus. Stochastic video generation with a learned prior. In E. Denton and R. Fergus. Stochastic video generation with a learned prior. In Proceedings of Proceedings of International Conference on Machine LearningInternational Conference on Machine Learning, 2018.,
2018
-
[17]
Dhariwal and A
P. Dhariwal and A. Nichol. Diffusion models beat GANs on image synthesis. In P. Dhariwal and A. Nichol. Diffusion models beat GANs on image synthesis. In Proceedings of Neural Proceedings of Neural Information Processing SystemsInformation Processing Systems, 2021.,
2021
-
[18]
Columbia University Press, 2009.2009. S. E. Eskimez, R. K. Maddox, C. Xu, and Z. Duan. End-to-end generation of talking faces from S. E. Eskimez, R. K. Maddox, C. Xu, and Z. Duan. End-to-end generation of talking faces from noisy speech. In noisy speech. In ICASSP 2020-2020 IEEE international conference on acoustics, speech and signal ICASSP 2020-2020 IEE...
arXiv 2009
Show all 94 references
-
[20]
Ferstl and R
Y . Ferstl and R. McDonnell. Investigating the use of recurrent motion modelling for speech gesture Y . Ferstl and R. McDonnell. Investigating the use of recurrent motion modelling for speech gesture generation. In generation. In Proceedings of the International Conference on ...
2018
-
[21]
Gafni, J
G. Gafni, J. Thies, M. Zollhöfer, and M. Nießner. Dynamic neural radiance fields for monocular 4d G. Gafni, J. Thies, M. Zollhöfer, and M. Nießner. Dynamic neural radiance fields for monocular 4d facial avatar reconstruction. In facial avatar reconstruction. In Proceedings of ...
2021
- [22]
-
[23]
URL https://github.com/genmoai/mochihttps://github.com/genmoai/mochi.. S. Ghorbani, Y . Ferstl, D. Holden, N. F. Troje, and M.-A. Carbonneau. ZeroEGGS: Zero-shot example-S. Ghorbani, Y . Ferstl, D. Holden, N. F. Troje, and M.-A. Carbonneau. ZeroEGGS: Zero-shot example- based g...
2022
-
[24]
Girdhar, M
R. Girdhar, M. Singh, A. Brown, Q. Duval, S. Azadi, S. S. Rambhatla, A. Shah, X. Yin, D. Parikh, R. Girdhar, M. Singh, A. Brown, Q. Duval, S. Azadi, S. S. Rambhatla, A. Shah, X. Yin, D. Parikh, and I. Misra. Factorizing text-to-video generation by explicit image conditioning. ...
2024
-
[25]
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio. Generative adversarial nets. In Bengio. Generative adversarial ne...
2014
-
[27]
Y . Guo, C. Yang, A. Rao, Z. Liang, Y . Wang, Y . Qiao, M. Agrawala, D. Lin, and B. Dai. AnimateDiff: Y . Guo, C. Yang, A. Rao, Z. Liang, Y . Wang, Y . Qiao, M. Agrawala, D. Lin, and B. Dai. AnimateDiff: Animate your personalized text-to-image diffusion models without specific...
2024
-
[29]
Hannun, C
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, and A. Y . Ng. Deep speech: Scaling up end-to-end speech r...
2014 arXiv
-
[30]
Hawthorne, A
C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, S. Dieleman, O. Vinyals, C. Hawthorne, A. Jaegle, C. Cangea, S. Borgeaud, C. Nash, M. Malinowski, S. Dieleman, O. Vinyals, M. Botvinick, I. Simon, H. Sheahan, N. Zeghidour, J.-B. Alayrac, J. Carreira, and...
2022
-
[31]
B. Heo, S. Park, D. Han, and S. Yun. Rotary position embedding for Vision Transformer. In B. Heo, S. Park, D. Han, and S. Yun. Rotary position embedding for Vision Transformer. In Proceedings Proceedings of European Conference on Computer Visionof European Conference on Comput...
2024
-
[32]
Ho and T
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 4646 J. Ho and T. Salimans. Classifier-free diffusion guidance. J. Ho and T. Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:220...
2022 arXiv
-
[33]
J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In Proceedings of Neural Proceedings of Neural Information Processing SystemsInformation Processing Systems, 2020.,
2020
-
[34]
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet. Video diffusion models. In J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet. Video diffusion models. In Proceedings of Neural Information Processing SystemsProceedings of Neural Infor...
2022
-
[35]
Hochreiter and J
S. Hochreiter and J. Schmidhuber. Long Short-Term Memory. S. Hochreiter and J. Schmidhuber. Long Short-Term Memory. Neural computationNeural computation, 1997.,
1997
-
[37]
URL https://arxiv.org/abs/2411.18664https://arxiv.org/abs/2411.18664.. S. A. Jacobs, M. Tanaka, C. Zhang, M. Zhang, S. L. Song, S. Rajbhandari, and Y . He. DeepSpeed S. A. Jacobs, M. Tanaka, C. Zhang, M. Zhang, S. L. Song, S. Rajbhandari, and Y . He. DeepSpeed Ulysses: System ...
2023 arXiv
-
[39]
Jaegle, S
A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, O. Hénaff, M. M. Botvinick, A. Zisserman, O. Vinyals, and J. C...
2022
-
[40]
Kalchbrenner, A
N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. N. Kalchbrenner, A. van den Oord, K. Simonyan, I. Danihelka, O. Vinyals, A. Graves, and K. Kavukcuoglu. Video pixel networks. Kavukcuoglu. Video pixel networks. arXiv:1610.00527arXiv:161...
-
[41]
Karras, T
T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen. Audio-driven facial animation by joint end-to-T. Karras, T. Aila, S. Laine, A. Herva, and J. Lehtinen. Audio-driven facial animation by joint end-to- end learning of pose and emotion. end learning of pose and emotion. AC...
2017
-
[44]
Kopp and I
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 4747 S. Kopp and I. Wachsmuth. Synthesizing multimodal utterances for conversational agents. S. Kopp and I. Wachsmuth. Synthesizing multimodal utte...
2004
-
[46]
J. Li, D. Kang, W. Pei, X. Zhe, Z. He, and L. Bao. Audio2Gestures: Generating diverse gestures J. Li, D. Kang, W. Pei, X. Zhe, Z. He, and L. Bao. Audio2Gestures: Generating diverse gestures from speech audio with conditional variational autoencoders. In from speech audio with ...
2021
-
[47]
T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero. T. Li, T. Bolkart, M. J. Black, H. Li, and J. Romero. Learning a model of facial shape and expression Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphicsfrom 4D scans. ACM Transaction...
2017
-
[48]
G. Lin, J. Jiang, J. Yang, Z. Zheng, and C. Liang. Omnihuman-1: Rethinking the scaling-up of one-G. Lin, J. Jiang, J. Yang, Z. Zheng, and C. Liang. Omnihuman-1: Rethinking the scaling-up of one- stage conditioned human animation models. stage conditioned human animation models...
2025 arXiv
-
[49]
Lipman, R
Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling. modeling. arXiv:2210.02747arXiv:2210.02747, 2023.,
-
[50]
H. Liu, M. Zaharia, and P. Abbeel. Ring attention with blockwise Transformers for near-infinite H. Liu, M. Zaharia, and P. Abbeel. Ring attention with blockwise Transformers for near-infinite context. context. arXiv:2310.0188arXiv:2310.01889, 2023a.9, 2023a. Q. Liu, J. He, Q. ...
-
[51]
X. Liu, C. Gong, and Q. Liu. Flow straight and fast: Learning to generate and transfer data with X. Liu, C. Gong, and Q. Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In rectified flow. In Proceedings of International Conference on Le...
2019
-
[52]
Lugaresi, J
C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. G. C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. G. Yong, J. Lee, et al. MediaPipe: A framework for building perception pipelines. Yong, ...
1906 arXiv
-
[53]
Mathieu, C
M. Mathieu, C. Couprie, and Y . LeCun. Deep multi-scale video prediction beyond mean square error. M. Mathieu, C. Couprie, and Y . LeCun. Deep multi-scale video prediction beyond mean square error. In In Proceedings of International Conference on Learning RepresentationsProcee...
2016
-
[54]
Mediainfo
MediaArea. Mediainfo. MediaArea. Mediainfo. https://mediaarea.net/en/MediaInfohttps://mediaarea.net/en/MediaInfo„ 2025.„
2025
-
[55]
Medina, D
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 4848 S. Medina, D. Tome, C. Stoll, M. Tiede, K. Munhall, A. G. Hauptmann, and I. Matthews. Speech S. Medina, D. Tome, C. Stoll, M. Tiede, K. Munhal...
2022
-
[56]
Mittal and B
G. Mittal and B. Wang. Animating face using disentangled audio representations. In G. Mittal and B. Wang. Animating face using disentangled audio representations. In Proceedings of Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Visionthe IEEE/CVF Win...
2020
-
[57]
C. Nash, J. Carreira, J. C. Walker, I. Barr, A. Jaegle, M. Malinowski, and P. Battaglia. Transframer: C. Nash, J. Carreira, J. C. Walker, I. Barr, A. Jaegle, M. Malinowski, and P. Battaglia. Transframer: Arbitrary frame prediction with generative models. Arbitrary frame predic...
2023
-
[59]
OpenAI, , A
URL https://openai.com/index/video-https://openai.com/index/video- generation-models-as-world-simulatorsgeneration-models-as-world-simulators.. OpenAI, , A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. OpenAI, , A. Hurst, A. Lerer, A. P. Gouc...
-
[60]
J. Pan, C. Wang, X. Jia, J. Shao, L. Sheng, J. Yan, and X. Wang. Video generation from single semantic J. Pan, C. Wang, X. Jia, J. Shao, L. Sheng, J. Yan, and X. Wang. Video generation from single semantic label map. In label map. In Proceedings of IEEE Conference on Computer ...
2019
-
[61]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S...
2019
-
[62]
Peebles and S
W. Peebles and S. Xie. Scalable diffusion models with Transformers. In W. Peebles and S. Xie. Scalable diffusion models with Transformers. In Proceedings of IEEE Proceedings of IEEE International Conference on Computer VisionInternational Conference on Computer Vision, 2023.,
2023
-
[63]
Pelachaud and M
C. Pelachaud and M. Bilvi. Computational model of believable conversational agents. C. Pelachaud and M. Bilvi. Computational model of believable conversational agents. Communication Communication in Multiagent Systemsin Multiagent Systems, 2003.,
2003
-
[65]
Radford, J
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 4949 A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever. Robust speech recognition A. Radford, J. W. Kim, T. Xu, G. Brockman,...
2022
-
[66]
Ronneberger, P
O. Ronneberger, P. Fischer, and T. Brox. U-Net: Convolutional networks for biomedical image O. Ronneberger, P. Fischer, and T. Brox. U-Net: Convolutional networks for biomedical image segmentation. In segmentation. In Proceedings of the International Conference on Medical Imag...
2015
-
[68]
Schneider, A
S. Schneider, A. Baevski, R. Collobert, and M. Auli. wav2vec: Unsupervised pre-training for speech S. Schneider, A. Baevski, R. Collobert, and M. Auli. wav2vec: Unsupervised pre-training for speech recognition. recognition. arXiv:1904.05862arXiv:1904.05862, 2019.,
1904 arXiv
-
[69]
J. Shah, G. Bikshandi, Y . Zhang, V . Thakkar, Pr. Ramani, and T. Dao. Flashattention-3: Fast and J. Shah, G. Bikshandi, Y . Zhang, V . Thakkar, Pr. Ramani, and T. Dao. Flashattention-3: Fast and accurate attention with asynchrony and low-precision. In accurate attention with ...
2024
-
[70]
S. Shen, W. Zhao, Z. Meng, W. Li, Z. Zhu, J. Zhou, and J. Lu. Difftalk: Crafting diffusion models S. Shen, W. Zhao, Z. Meng, W. Li, Z. Zhu, J. Zhou, and J. Lu. Difftalk: Crafting diffusion models for generalized audio-driven portraits animation. In for generalized audio-driven...
2023
-
[71]
Singer, A
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al. U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al. Make-a-Video: Text-to-video generation without text-video data. Make-a-Vide...
2022 arXiv
-
[72]
Sohl-Dickstein, E
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In nonequilibrium thermodynamics. In Proceedings o...
2015
-
[73]
Stypułkowski, K
M. Stypułkowski, K. V ougioukas, S. He, M. Zi˛eba, S. Petridis, and M. Pantic. Diffused heads: M. Stypułkowski, K. V ougioukas, S. He, M. Zi˛eba, S. Petridis, and M. Pantic. Diffused heads: Diffusion models beat GANs on talking-face generation. In Diffusion models beat GANs on...
2024
-
[74]
J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. RoFormer: Enhanced Transformer with Rotary J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. RoFormer: Enhanced Transformer with Rotary Position Embedding. Position Embedding. NeurocomputingNeurocomputing, 2024.,
2024
-
[75]
Taylor, T
S. Taylor, T. Kim, Y . Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews. A deep S. Taylor, T. Kim, Y . Yue, M. Mahler, J. Krahe, A. G. Rodriguez, J. Hodgins, and I. Matthews. A deep learning approach for generalized speech animation. learning approach for...
2017
-
[76]
Taylor, J
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 5050 S. Taylor, J. Windle, D. Greenwood, and I. Matthews. Speech-driven conversational agents using S. Taylor, J. Windle, D. Greenwood, and I. Matt...
2021
-
[77]
Team Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, Team Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, J. Zeng, J. Wang, J. Zhang, J. Zhou, J. Wang, J. Chen, K. Zhu, K. Zhao, K. Yan, L. Huang, M. F...
2025 arXiv
-
[80]
S. Tomar. Converting video formats with FFmpeg. S. Tomar. Converting video formats with FFmpeg. Linux JournalLinux Journal, 2006.,
2006
-
[82]
Tulyakov, M.-Y
S. Tulyakov, M.-Y . Liu, X. Yang, and J. Kautz. MoCoGAN: Decomposing motion and content for S. Tulyakov, M.-Y . Liu, X. Yang, and J. Kautz. MoCoGAN: Decomposing motion and content for video generation. In video generation. In Proceedings of IEEE Conference on Computer Vision a...
2018
-
[83]
V ondrick, H
C. V ondrick, H. Pirsiavash, and A. Torralba. Anticipating visual representations from unlabeled video. C. V ondrick, H. Pirsiavash, and A. Torralba. Anticipating visual representations from unlabeled video. In In Proceedings of IEEE Conference on Computer Vision and Pattern R...
2016
-
[84]
W. Wang, H. Yang, Z. Tuo, H. He, J. Zhu, J. Fu, and J. Liu. VideoFactory: Swap attention in W. Wang, H. Yang, Z. Tuo, H. He, J. Zhu, J. Fu, and J. Liu. VideoFactory: Swap attention in spatiotemporal diffusions for text-to-video generation. spatiotemporal diffusions for text-to...
2025
-
[85]
C. Wei, B. Sun, H. Ma, J. Hou, F. Juefei-Xu, Z. He, X. Dai, L. Zhang, K. Li, T. Hou, et al. MoCha: C. Wei, B. Sun, H. Ma, J. Hou, F. Juefei-Xu, Z. He, X. Dai, L. Zhang, K. Li, T. Hou, et al. MoCha: Towards movie-grade talking character synthesis. Towards movie-grade talking ch...
2025 arXiv
-
[86]
URL https://lilianweng.github.io/posts/2024-https://lilianweng.github.io/posts/2024- 04-12-diffusion-video/04-12-diffusion-video/.. J. Z. Wu, Y . Ge, X. Wang, S. W. Lei, Y . Gu, Y . Shi, W. Hsu, Y . Shan, X. Qie, and M. Z. Shou. Tune-J. Z. Wu, Y . Ge, X. Wang, S. W. Lei, Y . G...
2024
-
[87]
J. Xing, M. Xia, Y . Zhang, H. Chen, W. Yu, H. Liu, G. Liu, X. Wang, Y . Shan, and T.-T. Wong. J. Xing, M. Xia, Y . Zhang, H. Chen, W. Yu, H. Liu, G. Liu, X. Wang, Y . Shan, and T.-T. Wong. DynamiCrafter: Animating open-domain images with video diffusion priors. In DynamiCraft...
2024
-
[88]
S. Xu, G. Chen, Y .-X. Guo, J. Yang, C. Li, Z. Zang, Y . Zhang, X. Tong, and B. Guo. V ASA-1: Lifelike S. Xu, G. Chen, Y .-X. Guo, J. Yang, C. Li, Z. Zang, Y . Zhang, X. Tong, and B. Guo. V ASA-1: Lifelike audio-driven talking faces generated in real time. In audio-driven talk...
2024
-
[89]
W. Yan, Y . Zhang, P. Abbeel, and A. Srinivas. VideoGPT: Video generation using VQ-V AE and W. Yan, Y . Zhang, P. Abbeel, and A. Srinivas. VideoGPT: Video generation using VQ-V AE and Transformers. Transformers. arXiv:2104.10157arXiv:2104.10157, 2021.,
2021 arXiv
-
[90]
Yariv, Y
G. Yariv, Y . Kirstain, A. Zohar, S. Sheynin, Y . Taigman, Y . Adi, S. Benaim, and A. Polyak. Through- G. Yariv, Y . Kirstain, A. Zohar, S. Sheynin, Y . Taigman, Y . Adi, S. Benaim, and A. Polyak. Through- the-mask: Mask-based motion trajectories for image-to-video generation....
2025 arXiv
-
[91]
H. Yi, T. Ye, S. Shao, X. Yang, J. Zhao, H. Guo, T. Wang, Q. Yin, Z. Xie, L. Zhu, W. Li, M. H. Yi, T. Ye, S. Shao, X. Yang, J. Zhao, H. Guo, T. Wang, Q. Yin, Z. Xie, L. Zhu, W. Li, M. Lingelbach, and D. Zhou. MagicInfinite: Generating infinite talking videos with your words an...
2025 arXiv
-
[92]
Yoon, W.-R
Y . Yoon, W.-R. Ko, M. Jang, J. Lee, J. Kim, and G. Lee. Robots learn social skills: End-to-end Y . Yoon, W.-R. Ko, M. Jang, J. Lee, J. Kim, and G. Lee. Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots. In learning of co-speec...
2019
-
[93]
Zhang, Y
C. Zhang, Y . Zhao, Y . Huang, M. Zeng, S. Ni, M. Budagavi, and X. Guo. Facial: Synthesizing dynamic C. Zhang, Y . Zhao, Y . Huang, M. Zeng, S. Ni, M. Budagavi, and X. Guo. Facial: Synthesizing dynamic talking face with implicit attribute learning. In talking face with implici...
2023 arXiv
-
[94]
Zhang, L
Z. Zhang, L. Li, Y . Ding, and C. Fan. Flow-guided one-shot talking face generation with a high-Z. Zhang, L. Li, Y . Ding, and C. Fan. Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset. In resolution audio-visual dataset. In Proceedings ...
2022
-
[95]
Zhong, C
CAPTIONSCAPTIONS Seeing V oices: Generating A-Roll from Audio with MirageSeeing V oices: Generating A-Roll from Audio with Mirage 5252 W. Zhong, C. Fang, Y . Cai, P. Wei, G. Zhao, L. Lin, and G. Li. Identity-preserving talking face W. Zhong, C. Fang, Y . Cai, P. Wei, G. Zhao, ...
2023
-
[96]
Y . Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li. MakeltTalk: speaker-aware Y . Zhou, X. Han, E. Shechtman, J. Echevarria, E. Kalogerakis, and D. Li. MakeltTalk: speaker-aware talking-head animation. talking-head animation. ACM Transactions On Graphics ...
2020
-
[97]
S. Zhu, J. L. Chen, Z. Dai, Y . Xu, X. Cao, Y . Yao, H. Zhu, and S. Zhu. Champ: Controllable and S. Zhu, J. L. Chen, Z. Dai, Y . Xu, X. Cao, Y . Yao, H. Zhu, and S. Zhu. Champ: Controllable and consistent human image animation with 3d parametric guidance. consistent human imag...
2024 arXiv
-
[1986]
Cassell, H
J. Cassell, H. H. Vilhjálmsson, and T. Bickmore. BEAT: the behavior expression animation toolkit. J. Cassell, H. H. Vilhjálmsson, and T. Bickmore. BEAT: the behavior expression animation toolkit. In In Proceedings of the Annual Conference on Computer Graphics and Interactive T...
2001
-
[1997]
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed. HuBERT: Self-W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed. HuBERT: Self- supervised speech representation learning by masked prediction of hidden units. su...
2021
-
[2001]
Castellano
B. Castellano. PySceneDetect. B. Castellano. PySceneDetect. https://www.scenedetect.com/https://www.scenedetect.com/, 2025.,
2025
-
[2002]
Cudeiro, T
D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black. Capture, learning, and synthesis of D. Cudeiro, T. Bolkart, C. Laidlaw, A. Ranjan, and M. J. Black. Capture, learning, and synthesis of 3d speaking styles. In 3d speaking styles. In Proceedings of the IEEE/CVF Con...
2025 arXiv
-
[2003]
Prajwal, R
K. Prajwal, R. Mukhopadhyay, V . P. Namboodiri, and C. Jawahar. A lip sync expert is all you need K. Prajwal, R. Mukhopadhyay, V . P. Namboodiri, and C. Jawahar. A lip sync expert is all you need for speech to lip generation in the wild. In for speech to lip generation in the ...
2020
-
[2004]
Lahiri, V
A. Lahiri, V . Kwatra, C. Frueh, J. Lewis, and C. Bregler. Lipsync3d: Data-efficient learning of A. Lahiri, V . Kwatra, C. Frueh, J. Lewis, and C. Bregler. Lipsync3d: Data-efficient learning of personalized 3d talking faces from video using pose and lighting normalization. In ...
2021
-
[2005]
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. HunyuanVideo: A systematic framework for large video generative models. HunyuanVideo: A s...
2024 arXiv
-
[2006]
P. Toth, D. J. Rezende, A. Jaegle, S. Racanière, A. Botev, and I. Higgins. Hamiltonian generative P. Toth, D. J. Rezende, A. Jaegle, S. Racanière, A. Botev, and I. Higgins. Hamiltonian generative networks. networks. Proceedings of International Conference on Learning Represent...
2020
-
[2008]
J. Oh, X. Guo, H. Lee, R. Lewis, and S. Singh. Action-conditional video prediction using deep J. Oh, X. Guo, H. Lee, R. Lewis, and S. Singh. Action-conditional video prediction using deep networks in Atari games. In networks in Atari games. In Proceedings of Neural Information...
2015
-
[2009]
L. Tian, Q. Wang, B. Zhang, and L. Bo. EMO: Emote Portrait Alive – generating expressiveL. Tian, Q. Wang, B. Zhang, and L. Bo. EMO: Emote Portrait Alive – generating expressive portrait videos with audio2video diffusion model under weak conditions. In portrait videos with audi...
2024
-
[2014]
Grassal, M
P.-W. Grassal, M. Prinzler, T. Leistner, C. Rother, M. Nießner, and J. Thies. Neural head avatars P.-W. Grassal, M. Prinzler, T. Leistner, C. Rother, M. Nießner, and J. Thies. Neural head avatars from monocular RGB videos. In from monocular RGB videos. In Proceedings of IEEE C...
2022
-
[2015]
Rybkin, K
O. Rybkin, K. Pertsch, K. G. Derpanis, K. Daniilidis, and A. Jaegle. Learning what you can do before O. Rybkin, K. Pertsch, K. G. Derpanis, K. Daniilidis, and A. Jaegle. Learning what you can do before doing anything. In doing anything. In Proceedings of International Conferen...
2019
-
[2016]
Clark, J
A. Clark, J. Donahue, and K. Simonyan. Adversarial video generation on complex datasets. A. Clark, J. Donahue, and K. Simonyan. Adversarial video generation on complex datasets. arXiv:1907.06571arXiv:1907.06571, 2019.,
1907 arXiv
-
[2017]
M. Kipp. M. Kipp. Gesture generation by imitation: From human behavior to computer character animationGesture generation by imitation: From human behavior to computer character animation . . Universal-Publishers, 2005.Universal-Publishers,
2005
-
[2018]
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Yu, and F. Wei. WavLM: La...
2022
-
[2019]
J. Canny. A computational approach to edge detection. J. Canny. A computational approach to edge detection. IEEE Transactions on Pattern Analysis and IEEE Transactions on Pattern Analysis and Machine IntelligenceMachine Intelligence, 1986.,
1986
-
[2020]
Fang and S
J. Fang and S. Zhao. USP: A unified sequence parallelism approach for long context generative AI. J. Fang and S. Zhao. USP: A unified sequence parallelism approach for long context generative AI. arXiv:2405.07719arXiv:2405.07719, 2024.,
2024 arXiv
-
[2021]
G. Bradski. The OpenCV Library. G. Bradski. The OpenCV Library. Dr. Dobb’s Journal of Software ToolsDr. Dobb’s Journal of Software Tools, 2000.,
2000
-
[2022]
Y . Chen, J. Cao, A. Kag, V . Goel, S. Korolev, C. Jiang, S. Tulyakov, and J. Ren. Towards physical Y . Chen, J. Cao, A. Kag, V . Goel, S. Korolev, C. Jiang, S. Tulyakov, and J. Ren. Towards physical understanding in video generation: A 3d point regularization approach. unders...
2025
-
[2023]
Jaegle, O
A. Jaegle, O. Rybkin, K. G. Derpanis, and K. Daniilidis. Predicting the future with transformational A. Jaegle, O. Rybkin, K. G. Derpanis, and K. Daniilidis. Predicting the future with transformational states. states. arXiv:1803.09760arXiv:1803.09760, 2018.,
2018 arXiv
-
[2025]
L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu. Lip movements generation at a glance. In L. Chen, Z. Li, R. K. Maddox, Z. Duan, and C. Xu. Lip movements generation at a glance. In Proceedings of the European conference on computer vision (ECCV)Proceedings of the European con...
2018
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.