Pith. sign in

REVIEW 9 cited by

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.15111 v1 pith:M362KP5L submitted 2025-01-25 cs.CV

classification cs.CV
keywords human-centricscenesunderstandinghumanomnimodelbranchesfeaturesindividuals
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they generally lack effectiveness in human-centric scenes due to the absence of large-scale, specialized datasets and non-targeted architectures. In this work, we developed HumanOmni, the industry's first human-centric Omni-multimodal large language model. We constructed a dataset containing over 2.4 million human-centric video clips with detailed captions and more than 14 million instructions, facilitating the understanding of diverse human-centric scenes. HumanOmni includes three specialized branches for understanding different types of scenes. It adaptively fuses features from these branches based on user instructions, significantly enhancing visual understanding in scenes centered around individuals. Moreover, HumanOmni integrates audio features to ensure a comprehensive understanding of environments and individuals. Our experiments validate HumanOmni's advanced capabilities in handling human-centric scenes across a variety of tasks, including emotion recognition, facial expression description, and action understanding. Our model will be open-sourced to facilitate further development and collaboration within both academia and industry.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

    cs.AI 2026-07 conditional novelty 7.0 of 10

    A new in-cabin dataset pairs RGB/IR video, audio, and Chinese dialogue text with emotion, fatigue, and distraction labels, and baselines show fusion beats single modalities on the Chinese partition.

  2. Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Benchmarks seven open-source audio-video-text MLLMs on six affective datasets and shows a generative-knowledge prompting step improves fine-tuned emotion recognition.

  3. LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free token compression method using semantic connected components in space and time keeps video understanding accuracy high even when retaining only 5-10% of visual tokens.

  4. HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Requiring omni-modal models to summarize context before reasoning, with LLM-judged context and logical rewards, improves human-intent reasoning benchmarks.

  5. Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A new large benchmark for AI-generated human-centric video quality with pairwise preferences, plus a Mixture-of-Experts MLLM that outperforms prior methods on rating, comparison, and Q&A.

  6. Advancing the Foundation Model for Music Understanding

    cs.SD 2025-08 unverdicted novelty 5.0 of 10

    MuFun is proposed as a unified music foundation model that jointly handles instrumental and lyrical content, and it is claimed to outperform existing audio language models on the authors' new MuCUE benchmark.

  7. FaceLLM: A Multimodal Large Language Model for Face Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fine-tuning InternVL3 on ChatGPT-generated face QA pairs yields a face-specialized MLLM with the highest reported accuracy among MLLMs on FaceXBench.

  8. From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Text models encode linguistic taxonomies early and densely; speech models develop them later and less prominently, with multimodal models showing intermediate patterns.

  9. Grounding Intelligence in Movement

    cs.AI 2025-07 conditional novelty 4.0 of 10

    Movement should be treated as a first-class AI modeling modality, and a unified, biomechanically grounded movement foundation model built from aggregated data across species and sensors is the proposed path forward.

Pith tools