Pith. sign in

Paper Citation Record · LEDGER

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2501.15111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15111 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:19.930713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.340081Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e9f303b-77b4-4360-9f91-55a7dd733807 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.563253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:7c9a494bbf84badac859ae9aae02f4428edd5663bf527d16367bd6bf651a3e07

Observation 6a2b2aea-b7d0-4516-80ab-a56c46060c39 · inbound

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models cites this paper.

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:19.930713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:19.930713Z digest=sha256:03a4a57f69f6fd78c771381e79d24e17ebe1ff837b7e5cd4ce78b6af555d5082

Observation b6abf601-e5a5-40b6-998a-9d42fbc6d159 · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.070930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.070930Z digest=sha256:11e452e32ed9e6231cd668d8484c4918caaff1e38f58ec446fa772145ebf0a84

Observation 1fb4ef1b-8bbc-4fc1-86c2-1b82f69ac1de · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.656003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.656003Z digest=sha256:a0b96d0d3e089628aec030dfde23653ef338ff8dbbd6f9d286fc80b0e513e063

Observation 2d8ab3f4-bcb6-4714-a34a-933d1bc3c8e0 · inbound

Grounding Intelligence in Movement cites this paper.

Grounding Intelligence in Movement HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:32.881282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:32.881282Z digest=sha256:0ed772c198fed36d00ca80a7fa227f1c9bea33d8f7c537f75d42533fe2f0f767

Observation e079e223-3733-428a-8926-77c04bf46313 · inbound

FaceLLM: A Multimodal Large Language Model for Face Understanding cites this paper.

FaceLLM: A Multimodal Large Language Model for Face Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:41.688863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:41.688863Z digest=sha256:b30f1134eb821f89d19683f53504b447a26b1d73b58581509a1f5778247c148d

Observation 100cee19-04d4-461f-b53b-523d9ca1c7ac · inbound

Advancing the Foundation Model for Music Understanding cites this paper.

Advancing the Foundation Model for Music Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:55.526816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:55.526816Z digest=sha256:151cbbe2b10e556de083e65afea7ae8bd4e843a179650324ea9539f50d5d2245

Observation fa56c0e6-2f13-4c61-8dc4-f259ad81b243 · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.540778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.540778Z digest=sha256:59dbb05103c4396ea465837bcca2d5dd0b7c4020487dc743839e35d4b7273910

Observation 4402c2e6-7863-41c9-b114-689858367479 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.107948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:c5a21779c9150bee84fd5c79b511c140113121c00f7f152836291490a09b936f

Observation 909bd532-5a03-4c10-b716-673949aa9b66 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.643788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:b52ef827b16fdbb966b5a785051d3ced8aa9e95d7a234078d15fd172e2100553

Observation 11cec45c-71e6-466f-97e3-23c5c3c0950d · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:55:53.282103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:2f97ba2ed36bf0f0c7b24fb1d65d038d78aac611a13c9fb11a174003ba1df854

Observation 4062d7e7-479f-4d9f-a260-29c62f8a83aa · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.970344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:2bbf56707f0193f8c14b6d32d849a572bd1114a023acdaf699a13b50af2d1f1f

Observation 95265503-a77e-4605-b5a6-f35c0d1754c7 · inbound

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions cites this paper.

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:53:04.276506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:49:33.107658Z digest=sha256:439978d667bf600122da29bf5bb465050db7f6dad8d145f0001a59623db2ce8d

Observation 3f0a8955-3999-4a65-b62d-a7cc1225d133 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.134430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:c4c241ac658e1d2ac8f2ec09ecccfc1709e71d08d86a6cc806b4559c4299c7d0

Observation a746a167-4cf3-491d-857e-f4a5e5deb2b7 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 283

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.220651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:8a54bcd599fe2531463377243392b0e74beab1244d03c5e29a101f2d80aa08e3

Observation 5bcc8013-624b-4b6c-b3b1-2f639fee0d58 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.047112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:96eb52b6cf651224ddba01a2677f6bec981767d07276311a504d65789b00ddd9

Observation 0c2ce232-45d3-46e7-92f1-58edd413481d · inbound

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions cites this paper.

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 91

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T18:53:39.002635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T18:49:18.815456Z digest=sha256:535629a263c846a490f47b3903e6217c2f24bea01c3430fa8a574312d564659d

Observation fa6cc457-39d6-4973-b9a0-195c6e5ec81d · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.342785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:abddb8aa47d3aebe6968b980c0cb4f26c3197c50d18de16108b362a5001d4ced

Observation 4ff8a9ec-fe22-4656-9a19-5af71dfaf5c6 · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:19:39.200505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T05:18:29.720630Z digest=sha256:68a8caccb32ce4cf7dc22b7fbe428d039de72782587d7e4317dd0f5b362d87ac

Observation 430b824a-7ecb-463d-b4af-de283cf666fe · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.327465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:06:26.561698Z digest=sha256:551e4dfea0aa94e240f58afa10eeed5a5689fb2815f64c958a0c00d58bbdf29d

Observation 85f127db-a05d-4fc2-868a-471b1e4a238a · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.204673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:d566cd9146a76324cb56d4b24f9d421c086e845d7e0b133ce67262804f18682c

Observation 2b0317bc-baa6-490e-815f-2d0073361aed · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.372188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:e7944401b9b78684af8680c6c0c2bdf1d030579dd2df990147882769188e27bb

Observation 69e6726e-d2c1-4fdf-91df-1ca81034a101 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.398553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:e58c03653e51cf1dceb231c946617fd8c80229d732e6eaed956992041067ca7d

Observation a89b4cd4-0ccc-4af1-9a1c-c3661f978015 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.283536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:9da478bfc01498332cba8e1a775d3f9fe17e4016f7404423cef0068997ca8969

Observation 8789fd0b-e0c6-4ee2-b437-9aeea8698f35 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.342275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:34a15b2c381bb1de44c661d103274622fcf01f6f33efd9b827c6616d8cd6eb27

Observation 9152cc76-dd74-4f33-b421-a8b6edea5666 · inbound

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring cites this paper.

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:09.040422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:09.040422Z digest=sha256:778dd28a0117ac333f9dbd0177e8da13fab52124e099078c6d0dac0c2d794230

Observation d47f81cc-cea3-4ad9-a62b-73d38c8f2d38 · inbound

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model cites this paper.

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T20:07:41.111549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:07:41.111549Z digest=sha256:7e5541cda7a990d5fcdbc403751e1f04d1677181f231350afe5582ff2237bdf5