Pith. sign in

REVIEW 5 cited by

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.19209 v1 pith:ISHHTLC4 submitted 2025-08-26 cs.CV

classification cs.CV
keywords charactermodelmultimodalsemanticanimationsaudiobeyondcoherent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm, lacking a deeper semantic understanding of emotion, intent, or context. To bridge this gap, \textbf{we propose a framework designed to generate character animations that are not only physically plausible but also semantically coherent and expressive.} Our model, \textbf{OmniHuman-1.5}, is built upon two key technical contributions. First, we leverage Multimodal Large Language Models to synthesize a structured textual representation of conditions that provides high-level semantic guidance. This guidance steers our motion generator beyond simplistic rhythmic synchronization, enabling the production of actions that are contextually and emotionally resonant. Second, to ensure the effective fusion of these multimodal inputs and mitigate inter-modality conflicts, we introduce a specialized Multimodal DiT architecture with a novel Pseudo Last Frame design. The synergy of these components allows our model to accurately interpret the joint semantics of audio, images, and text, thereby generating motions that are deeply coherent with the character, scene, and linguistic content. Extensive experiments demonstrate that our model achieves leading performance across a comprehensive set of metrics, including lip-sync accuracy, video quality, motion naturalness and semantic consistency with textual prompts. Furthermore, our approach shows remarkable extensibility to complex scenarios, such as those involving multi-person and non-human subjects. Homepage: \href{https://omnihuman-lab.github.io/v1_5/}

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ETHead: Generating Expressive 3D Facial Animation and Head Movement from Speech

    cs.GR 2026-08 conditional novelty 6.0 of 10

    ETHead pre-trains an emotion-aware speech encoder on 2D talking-head videos and uses it to guide a diffusion-based 3D talking-head generator, improving emotional expressiveness and head motion.

  2. AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AgentHOI generates human-object interaction videos from text plus one human image and one object image, using multi-agent action planning and implicit text-to-motion feature alignment inside a video diffusion model.

  3. Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical global-keyframe then local-refinement diffusion pipeline produces stable 720p/30fps music-to-dance videos longer than one minute across five genres.

  4. Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Live Avatar reports real-time streamable generation from a 14B audio-driven diffusion model at ~20 FPS on 5 H800s with stable identity over 10,000 seconds.

  5. Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model

    cs.CV 2026-08 conditional novelty 5.0 of 10

    Diff-VF extends short-video diffusion models to longer videos at inference time by combining noise mixing, weighted window fusion, and temporally dilated sampling, with no retraining.

Pith tools