Pith. sign in

REVIEW 21 cited by

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.01061 v3 pith:ACKR4SU3 submitted 2025-02-03 cs.CV

classification cs.CV
keywords generationhumanomnihumanaudio-drivensupportsvideoanimationconditions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propose OmniHuman, a Diffusion Transformer-based framework that scales up data by mixing motion-related conditions into the training phase. To this end, we introduce two training principles for these mixed conditions, along with the corresponding model architecture and inference strategy. These designs enable OmniHuman to fully leverage data-driven motion generation, ultimately achieving highly realistic human video generation. More importantly, OmniHuman supports various portrait contents (face close-up, portrait, half-body, full-body), supports both talking and singing, handles human-object interactions and challenging body poses, and accommodates different image styles. Compared to existing end-to-end audio-driven methods, OmniHuman not only produces more realistic videos, but also offers greater flexibility in inputs. It also supports multiple driving modalities (audio-driven, video-driven and combined driving signals). Video samples are provided on the ttfamily project page (https://omnihuman-lab.github.io)

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AgentHOI generates human-object interaction videos from text plus one human image and one object image, using multi-agent action planning and implicit text-to-motion feature alignment inside a video diffusion model.

  2. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.

  3. Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation

    cs.LG 2026-01 conditional novelty 6.0 of 10

    A causal diffusion-forcing model generates interactive head-avatar motion with 500ms motion-generation latency and learns expressive reactions via DPO with synthetic negative samples.

  4. OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A single autoregressive diffusion model, trained on a new 286-hour SMPL-X dataset, generates whole-body motion from text, audio, and spatial-temporal control signals, with reference-motion conditioning.

  5. UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A unified audio-video generator built by stitching pre-trained video and music diffusion models, trained on 7,600 hours of data, with a new evaluation benchmark.

  6. FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A three-part system, Talking-Critic, Talking-NSQ, and TLPO, aligns diffusion portrait animation models to human preferences and improves lip-sync, motion naturalness, and visual quality.

  7. SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SpeakerVid-5M provides 5.2 million audio-visual human clips (8,743 hours) with rich annotations and a dyadic interaction benchmark for training interactive virtual humans.

  8. AnyAni: An Interactive System with Generative AI for Animation Effect Creation and Code Understanding in Web Development

    cs.HC 2025-06 conditional novelty 6.0 of 10

    AnyAni combines LLM generation, a version tree, and video-based checking to help front-end developers create and understand web animations; a nine-person study reports usability gains over a chatbot baseline.

  9. DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.

  10. HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    HunyuanVideo-HOMA generates human-object interaction videos from weak, sparse inputs: one arm pose, an object center dot, a human photo, and an object photo.

  11. Audio-Sync Video Generation with Multi-Stream Temporal Control

    cs.CV 2025-06 reject novelty 6.0 of 10

    MTV splits audio into speech, effects, and music to separately drive lip sync, event timing, and visual mood in video generation, trained on a new 392K-clip dataset.

  12. Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A UK-based mixed-methods study finds a wide gap between public trust in biometrics and expert concern about deepfake spoofing, and proposes a tri-layer mitigation framework.

  13. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  14. JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A dual-branch diffusion transformer with joint video-audio self-attention and a keypoint-based mouth-area loss reports top lip-sync and speech metrics on two benchmarks.

  15. InfinityHuman: Towards Long-Term Audio-Driven Human

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.

  16. MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation

    cs.CV 2025-08 reject novelty 5.0 of 10

    A new autoregressive video-generation framework for interactive digital humans, with a 64x compression autoencoder and a diffusion renderer, claims real-time multimodal control but is only demonstrated for audio input.

  17. InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Sparse-frame dubbing with adjacent-chunk keyframe sampling lets a streaming audio-video model produce full-body motion synchronized to new audio while preserving identity and camera motion.

  18. FramePrompt: In-context Controllable Animation with Zero Structural Changes

    cs.GR 2025-06 conditional novelty 5.0 of 10

    FramePrompt turns character animation into a video-continuation task by concatenating reference image, skeleton frames, and target frames into one sequence, then training the pretrained Wan-I2V model to generate only ...

  19. AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Timestep-segment preference optimization with separate motion and fidelity LoRAs improves audio-driven human animation quality and allows a 3.3x inference speedup.

  20. Seeing Voices: Generating A-Roll Video from Audio with Mirage

    cs.CV 2025-06 reject novelty 5.0 of 10

    Mirage generates photorealistic A-roll videos of people speaking directly from audio, using only joint self-attention over audio, text, and video tokens.

  21. Wan-S2V: Audio-Driven Cinematic Video Generation

    cs.CV 2025-08 reject novelty 4.0 of 10

    Wan-S2V is an audio-driven video generator built on Wan, claiming better cinematic character animation than prior systems, though the evaluation is limited.

Pith tools