Pith. sign in

Paper Citation Record · LEDGER

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2501.15111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15111 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:19.930713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.340081Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e9f303b-77b4-4360-9f91-55a7dd733807 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.563253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:9cdb871f4a303eb0989e56190672f747a28138b9a0a9d4f94302acdd05bb734e

Observation 6a2b2aea-b7d0-4516-80ab-a56c46060c39 · inbound

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models cites this paper.

From Words to Waves: Analyzing Concept Formation in Speech and Text-Based Foundation Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:19.930713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:19.930713Z digest=sha256:03a4a57f69f6fd78c771381e79d24e17ebe1ff837b7e5cd4ce78b6af555d5082

Observation b6abf601-e5a5-40b6-998a-9d42fbc6d159 · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.070930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.070930Z digest=sha256:11e452e32ed9e6231cd668d8484c4918caaff1e38f58ec446fa772145ebf0a84

Observation 1fb4ef1b-8bbc-4fc1-86c2-1b82f69ac1de · inbound

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs cites this paper.

LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:36.656003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:36.656003Z digest=sha256:830df917dccf67ccd36e8309a556af256cb9581619116788781a09e214d5c069

Observation 2d8ab3f4-bcb6-4714-a34a-933d1bc3c8e0 · inbound

Grounding Intelligence in Movement cites this paper.

Grounding Intelligence in Movement HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:32.881282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:32.881282Z digest=sha256:0ed772c198fed36d00ca80a7fa227f1c9bea33d8f7c537f75d42533fe2f0f767

Observation e079e223-3733-428a-8926-77c04bf46313 · inbound

FaceLLM: A Multimodal Large Language Model for Face Understanding cites this paper.

FaceLLM: A Multimodal Large Language Model for Face Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T17:39:41.688863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:39:41.688863Z digest=sha256:b30f1134eb821f89d19683f53504b447a26b1d73b58581509a1f5778247c148d

Observation 100cee19-04d4-461f-b53b-523d9ca1c7ac · inbound

Advancing the Foundation Model for Music Understanding cites this paper.

Advancing the Foundation Model for Music Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:55.526816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:55.526816Z digest=sha256:151cbbe2b10e556de083e65afea7ae8bd4e843a179650324ea9539f50d5d2245

Observation fa56c0e6-2f13-4c61-8dc4-f259ad81b243 · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.540778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.540778Z digest=sha256:59dbb05103c4396ea465837bcca2d5dd0b7c4020487dc743839e35d4b7273910

Observation 4402c2e6-7863-41c9-b114-689858367479 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.107948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:4171a692f0a91d59e9a7a0a0751e19550a8d1b86edc83d95c8f77f49ee3a4f46

Observation 909bd532-5a03-4c10-b716-673949aa9b66 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.643788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:0c834f3c483e6a3a9692ebcb7a9d871d6aef04853e1a24004eb540be64c2baef

Observation 11cec45c-71e6-466f-97e3-23c5c3c0950d · inbound

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis cites this paper.

C2F-Thinker: Coarse-to-Fine Reasoning with Hint-Guided Reinforcement Learning for Multimodal Sentiment Analysis HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:55:53.282103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T13:51:40.334057Z digest=sha256:d9c4ad9fac4a1e90f3852a2bdd7cb8412c96b01cd030107867caf1803b0ff210

Observation 4062d7e7-479f-4d9f-a260-29c62f8a83aa · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.970344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:0df9b15cd3acb0748cfe828b0223a7a5d40a7b75eb7397e2f867d3d7fd6bbae6

Observation 95265503-a77e-4605-b5a6-f35c0d1754c7 · inbound

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions cites this paper.

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:53:04.276506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:49:33.107658Z digest=sha256:87dee98141e21774b84409aff0849307576fa8dbbdee58731b5a4f1978ef42fd

Observation 3f0a8955-3999-4a65-b62d-a7cc1225d133 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.134430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:9b27a48a3b571a48a6ce01149d7f325c257616d9ff46b66f6cd872054869bcc2

Observation a746a167-4cf3-491d-857e-f4a5e5deb2b7 · inbound

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective cites this paper.

Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 283

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:51:03.220651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T04:32:29.428080Z digest=sha256:bfba2e9891326cdf5ec73197eff9cf9518cc42dc7b9cdbe469f755a5f0d80b00

Observation 5bcc8013-624b-4b6c-b3b1-2f639fee0d58 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.047112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:87e6f634da1e8eefa6fb0d8a4c29e4d2b30134a0d8ac621898c7ce39cc63a980

Observation 0c2ce232-45d3-46e7-92f1-58edd413481d · inbound

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions cites this paper.

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 91

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T18:53:39.002635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T18:49:18.815456Z digest=sha256:fa1e76c5ab7218a153ae0b4f12607f271db312049b7b39a5a73cf082d6f8fede

Observation fa6cc457-39d6-4973-b9a0-195c6e5ec81d · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.342785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:db3554997128aec11ab4c3fb0bbff0ef2b9a59bf71193f97cb0f5485238d0654

Observation 4ff8a9ec-fe22-4656-9a19-5af71dfaf5c6 · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:19:39.200505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T05:18:29.720630Z digest=sha256:dfc53199e601264960b8613b3a91c1d67d3ccbfb2832103543d026cd06467eb6

Observation 430b824a-7ecb-463d-b4af-de283cf666fe · inbound

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition cites this paper.

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.327465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:06:26.561698Z digest=sha256:30a15727c4c8e18cb2510eda15efcbe4d0b654bf39f84839c04eaa3ed76280b0

Observation 85f127db-a05d-4fc2-868a-471b1e4a238a · inbound

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind cites this paper.

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.204673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:04:39.753443Z digest=sha256:06b932f82fada35eabf88be2026b7af0f38d13afdc123271f289426f859fd25d

Observation 2b0317bc-baa6-490e-815f-2d0073361aed · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.372188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:11fe0f69bea0abb02eee40264fab5e0e69dd41c2f9c577c04c433d1b69d0c2e2

Observation 69e6726e-d2c1-4fdf-91df-1ca81034a101 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.398553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:752844653345f026b02a2c22c56d3204287669cda754605ae58a410cd6e799e6

Observation a89b4cd4-0ccc-4af1-9a1c-c3661f978015 · inbound

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression cites this paper.

AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.283536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:16:56.174638Z digest=sha256:3ab3f5795243c6819a931e6f73d78092343534177af0fc5ebb7e9a81d996e4b9

Observation 8789fd0b-e0c6-4ee2-b437-9aeea8698f35 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.342275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:e9018f461d26900541aeafc64f480b8721d5b46292c71787ea77cbe92c693f23

Observation 9152cc76-dd74-4f33-b421-a8b6edea5666 · inbound

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring cites this paper.

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:09.040422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:09.040422Z digest=sha256:778dd28a0117ac333f9dbd0177e8da13fab52124e099078c6d0dac0c2d794230

Observation d47f81cc-cea3-4ad9-a62b-73d38c8f2d38 · inbound

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model cites this paper.

Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T20:07:41.111549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:07:41.111549Z digest=sha256:7e5541cda7a990d5fcdbc403751e1f04d1677181f231350afe5582ff2237bdf5