Pith. sign in

Paper Citation Record · LEDGER

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 25 inbound Pith citation observations for arXiv:2506.21277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21277 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:16.371915Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T15:40:22.379357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.286106Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d9bc3846-ed5c-434e-b476-f42861c1cec7 · outbound

This paper cites Qwen2.5-Omni Technical Report.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Qwen2.5-Omni Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.100076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.100076Z digest=sha256:62249ccdb0558c5545d989d807d6cf4a36a74c90364b06c173445a8ebc9a7654

Observation 02c0e5a7-4ff0-45df-83de-54ac2fc6f8d4 · outbound

This paper cites Ocean-omni: To understand the world with omni-modality,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Ocean-omni: To understand the world with omni-modality,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.189958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.189958Z digest=sha256:daeb67afed48af9e4f87d0fc9e5e449615784bfd80e2d69eb7f27a7221fb7ab8

Observation b21fc665-8563-4dac-84d1-1b4f15f5c7a2 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.250680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.250680Z digest=sha256:114cf1863b7e2a51d433c02d5a2da6bf979a66757862d03b07265033f3307ee2

Observation d5cdd6a8-e1dd-4942-a1e9-e3d1f4c900d9 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.311095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.311095Z digest=sha256:d1da22783c4d81eea2e1a159e0cfca6d051e2271087416f2151512bc6bf1412a

Observation cc1a1c1b-483e-4412-9a66-b0bac0933d76 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.354539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.354539Z digest=sha256:8396e375830902b424dec7bbb6690cfef1b5433b68ec2149dde90550bc3ec1da

Observation fccec2a2-899d-4fb8-9ec7-dcb3dea65492 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.414790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.414790Z digest=sha256:b8a657fcb0324f664a076f7914ce51dabac17e450e06046e1171f6c68989acfd

Observation 2017c838-81a4-459c-89f9-da640f51f34e · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.542217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.542217Z digest=sha256:c694dc21f41c432c74a34e9bb43670026997cac201e1b9098b00059acf86a25e

Observation 0a69b6e7-73fe-43b2-a617-8f53fb280752 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.605014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.605014Z digest=sha256:11f3f09f50a7b4190f24f4a46a20146bfe7712dfe8dd8f1e227058f98e3af8d0

Observation 83c73b5c-953f-4b72-b537-44a7165a64ef · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.642332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.642332Z digest=sha256:2428c5b696069db32509dc7ae74be143d2c0873739ceacfc1a260a1f8b4734a7

Observation 2e8db14b-11e6-41f0-95b6-b26a32bac4a4 · outbound

This paper cites Let’s verify step by step,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Let’s verify step by step,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.716590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.716590Z digest=sha256:00b5e8f27b8e6bee922257fc9d97dc945b26213ed5814a64871b3c5f94a0ba80

Observation 00094bd8-5090-4b3d-a269-be2a6a281e40 · outbound

This paper cites Deep reinforcement learning from human preferences,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Deep reinforcement learning from human preferences,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.791285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.791285Z digest=sha256:96d755ffbbc2a78bb6e52e5cb740d43f461871f12fe6a2b306e263e3043b7c25

Observation c2c5ed5d-25b5-49d3-b8da-6b4715ff9d5c · outbound

This paper cites Training language models to follow instructions with human feedback,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Training language models to follow instructions with human feedback,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.854051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.854051Z digest=sha256:7d3556e574f5cde4ef6e60eef7c7f8b9fc2d82971815360e9ea5f73e7a48c2b8

Observation cadeb560-c2f2-400e-90c2-e440629f71fa · outbound

This paper cites Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Daily-omni: Towards audio-visual reasoning with temporal alignment across modalities,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:13.927417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:13.927417Z digest=sha256:d137835155300defafd97895462fc8c3e4b3164d291c8eb80cb915dab58926ef

Observation 4e3c89a2-e02d-4033-8cbf-0a45e2aebf3e · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.008552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.008552Z digest=sha256:00e247becf29a7f06467aa9e97ba7913c20ec75239185eb5023d4e0584094cf4

Observation b6abf601-e5a5-40b6-998a-9d42fbc6d159 · outbound

This paper cites HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.070930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.070930Z digest=sha256:8e9522e87081014bbcfeb46b58fe1849e75e3b5488ad3c8fca05fb36c15de06f

Observation 8fd7646a-2680-4dfc-ad4e-60c889f0c778 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.132504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.132504Z digest=sha256:af863bcacbcbcde53c22872e6b9c9e9f1dd0ff4d6ad698e9f8075fc90545f999

Observation 3ec31ea1-a90e-4493-a56b-c09be0b7b9bb · outbound

This paper cites InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.198672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.198672Z digest=sha256:cee77c95ed17b8f75e2abea14ea04b851233b51b8cc452dcc6da2e256fd397fd

Observation dac0beed-5947-41bd-acb0-9d2fb423d933 · outbound

This paper cites ViSpeak: Visual Instruction Feedback in Streaming Videos.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context ViSpeak: Visual Instruction Feedback in Streaming Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.252623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.252623Z digest=sha256:5f7dd28bfa9094bd79f6bc383a66b358fc03829b041e4673d36780185870264a

Observation eb3617e0-da3f-4432-a830-7d9fd7b958eb · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.298315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.298315Z digest=sha256:092e36a77e2b41d3600ed7359569af795b2ff68134ce17d13f3c9ae3d088d7eb

Observation e1e885b5-8e18-4d07-801f-3ce16ed229f2 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:18.497671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:36:14.371699Z digest=sha256:e8e5be0a7476de44435b3f531d4dce38ef513d15056d0207ff82a38b08efd476

Observation da4ef8b4-87b9-44b4-bb5c-c0544257c32a · outbound

This paper cites ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.442972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.442972Z digest=sha256:ce8103a076233ea7eecdc32fd6aee0c868ce2c9f4b28248fdcaffc1c061ee3af

Observation 868b216d-5934-4f50-84e1-78e86dcae62b · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:18.346314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:36:14.506749Z digest=sha256:0b70f62f0c00e419bc2b1a727871ebaba5d67dba2f6111bc3e678183004eab80

Observation 0ccc3c08-c875-420a-b814-3aef9656fe1c · outbound

This paper cites Omnibench: Towards the future of universal omni-language models,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Omnibench: Towards the future of universal omni-language models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.573744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.573744Z digest=sha256:91cd804affa2ffeb12e6444b057b73e2205f0b27dffe4443b979e39b191b53c4

Observation 7762b41f-f762-4dd8-8538-925bb04c6bbb · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.648069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.648069Z digest=sha256:bc21457ac981f6fb18e295433658df586f92f84b1a229d2ded3021bd3099bf58

Observation be188088-53f3-405c-8485-d3c128398b7b · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.709508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.709508Z digest=sha256:e1b8e1aad34626b493f97c0eb45e7aefc2284b38bb6c55da2dbeaea8893a9e07

Observation e3a3aa44-7bc6-4e3b-b410-5bbf810d6588 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.773345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.773345Z digest=sha256:30697ba9bf523941592490399347726aa1d9842d1b3c6d527b0f804d3c39a00a

Observation 7acb12a4-1f90-432c-a59d-a77bcbe0afa5 · outbound

This paper cites Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.841838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.841838Z digest=sha256:7cebc5b264a20567a869b58a69a6f5d1a90a359d1783a1bbf29916167c28976f

Observation ccc0b7ab-2e2b-4db4-ad07-1df3b15e635b · outbound

This paper cites Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.896841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.896841Z digest=sha256:01baa8eb6d25edbd63a99644b8019327ed10d2022fdcca757aaa3da2b32b83c0

Observation b5e06ef8-e5bd-4aa1-a082-be15c5a55d81 · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.965189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.965189Z digest=sha256:0f2b9603996e8dcce91710c5cd00e02f61d3622d703661549ef4beb5e8e4c019

Observation 2caf0f52-2c2f-4cb8-bcf5-c365bbbc7f9b · outbound

This paper cites EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.038514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.038514Z digest=sha256:c208b120cf8a27f77d09923cee70b23675ef668ff0cb68c865cb54fe17a23c15

Observation 8f50cf77-265a-42b5-a02d-2589dfdb4edc · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.103448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.103448Z digest=sha256:f23d0326ec6f3530cdf2d9882db1edd94c951a8d71420b998c71f1fe58daf277

Observation 41b50db4-1a75-499d-8957-a94752de91bb · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.170351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.170351Z digest=sha256:1024d235b9a08cffdcf1d138e8193516bd120887f81f5bd91f90d89fec87587d

Observation 976133f1-7b17-43e2-9abe-947893c2214f · outbound

This paper cites Social-iq: A question answering benchmark for artificial social intelligence,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Social-iq: A question answering benchmark for artificial social intelligence,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:18.184139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:36:15.224232Z digest=sha256:00a61432f87435e40e113c45ef8cabeee4c8f71af70660107a34f501c786c53a

Observation 6f381da0-0ce4-4880-bcce-f2d97caca2c4 · outbound

This paper cites Social-iq 2.0 challenge: Benchmarking multimodal social understanding,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Social-iq 2.0 challenge: Benchmarking multimodal social understanding,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:18.047248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:36:15.295001Z digest=sha256:32c31c90b148e996ffef8daaaecb335ff654b96b6c280f95d027bc1123938666

Observation 6fb568fc-47de-4fac-9f85-d4df837e52ca · outbound

This paper cites Explainable Multimodal Emotion Recognition.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Explainable Multimodal Emotion Recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.352601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.352601Z digest=sha256:71b084fa9d9efc03b7999ff04684315efd87e750d07c1b04b04c6ae1dfa5d333

Observation eff13216-134e-4527-abcc-8b7c9b4f2b5c · outbound

This paper cites MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context MDPE: A Multimodal Deception Dataset with Personality and Emotional Characteristics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.404271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.404271Z digest=sha256:1037e883038a97c2187964a39710482ed0c24ae878042663a624e49a22977211

Observation 0ac7f971-756b-4df0-a07a-adb9a148390e · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.530089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.530089Z digest=sha256:0f1f7a817983c0d2aca23790fde251d18707b513f8370861051b20af417ce145

Observation 6e4ade0f-7057-44cf-9251-b15672f059d8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Understanding R1-Zero-Like Training: A Critical Perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.612633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.612633Z digest=sha256:20b46c166f92a7d153a049f9ad95cc1a8249758a4b19d23fe78dae3445c61c6d

Observation 737c5beb-ed58-4d15-bef9-3c84af483d50 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Bleu: a method for automatic evaluation of machine translation,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.669533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.669533Z digest=sha256:0acf3cd934bc03751d9737c3566a9e90d611e4eeb957bb52e892bf4bdf544bce

Observation 65645993-1ed6-42cb-9e0d-07efdf6be16a · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Rouge: A package for automatic evaluation of summaries,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.737016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.737016Z digest=sha256:c9d4334e38f37996e9974d445d955281cd4094a167947dce97203983c9b38bab

Observation dab78a1e-8eb2-4f93-a1f6-d2595de50d6e · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.835979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.835979Z digest=sha256:faef8fe4531530a8e9e374456644e981b3d61f7fc143fda6045d923753a18fac

Observation cc07c2b5-094e-49ea-8069-d1d9f188f4d7 · outbound

This paper cites Gemini 2.5 pro,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Gemini 2.5 pro,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:17.887554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:36:15.883030Z digest=sha256:eca391de9bf509e7d3a5dbdaafeb111fd3e61566c5610f222a5e08de7a0d13aa

Observation 361cf6a6-743f-43ea-9829-93aa50d84b9a · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:15.960063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:15.960063Z digest=sha256:a6c5023e6a8998aad0f39c3f1de732e17797bd0a13af87c04ce9cbe66caf8589

Observation dd9cc03c-bc9a-42bf-9f5a-a2e6bf2ac43f · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:16.021781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:16.021781Z digest=sha256:b14d0e6d6c979cc66a75003c54b39d3f520e3b4716deaa56446e1f1c4c1532dd

Observation a0a0d3ee-7351-42c2-8670-9c7d5b0507d0 · outbound

This paper cites Introducing the next generation of Claude,.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Introducing the next generation of Claude,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:17.754014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:36:16.124771Z digest=sha256:0272aa8765906962365a914c4d905972365454b19d952571c48c840f086e0fa3

Observation 41a5225d-a5b5-4634-887a-413f92aba628 · outbound

This paper cites GPT-4o System Card.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context GPT-4o System Card

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:16.173667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:16.173667Z digest=sha256:8741b5f900c4ba3ffad9b40f1121d20eddcbd1e9bc9a1cbe0047e6aab0812f38

Observation 942a1024-803c-49a2-93fa-b668823c8534 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:16.270052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:16.270052Z digest=sha256:b80a5981483c1b5595a971bcd2352ee43f2dd58cbe84794f2886bb81ca5e5796

Observation 64396dc5-da23-4b78-9b50-770932dda3d7 · outbound

This paper cites OpenAI o1 System Card.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context OpenAI o1 System Card

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:16.328116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:16.328116Z digest=sha256:d7c69d8dfd6cf447f3649cadc894d1f539c25154081227584476e1baf9c4b28f

Observation c104b7a9-1c1f-41f6-97fa-cccf779cf544 · outbound

This paper cites AffectGPT: Dataset and Framework for Explainable Multimodal Emotion Recognition.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context AffectGPT: Dataset and Framework for Explainable Multimodal Emotion Recognition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:16.371915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:16.371915Z digest=sha256:40ebfe8d8a0be1fcb017fb981f3255a8c88b76d9d4ab44156893d9b0b8c74a6f

Pith citing papers

Observation 5254e4da-853c-4ac8-8643-6b2e024fa21e · inbound

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models cites this paper.

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.048217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:45:37.418493Z digest=sha256:589b64c88d92bea79f70181ace528583ab6f2bada39c8406365ebcb35fd3651e

Observation c9b2aa38-e488-4e17-83a9-43891b92c31f · inbound

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs cites this paper.

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T17:16:42.542703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:16:42.542703Z digest=sha256:63348a34fa78e9ace8b32a95e49fe5981fc419d0589fe081a63a91a93e754bb0

Observation 95218483-f9a3-4ebb-b2e0-62ea836414ee · inbound

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions cites this paper.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.377213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.377213Z digest=sha256:2818a5742e192bbe476d6a519a20298ee6943a6726494346b335b9f2fb0f69fc

Observation 48f64fe9-9131-45e2-8b4d-3a57bd8dfc5c · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:11:01.853846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:9b036b95bc6419384ab3cdaddedf51436367a9d3adb6aca2fe517fe186419f09

Observation ad84e811-f1ba-43c7-ba5f-a6cfadc8f22f · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.635645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:a10596677aab2190ba7d110a0d578e435ce2df2bbed0fc373db6da68d5d92583

Observation 2d070b2f-fe4b-43aa-8871-d9c741d48a2e · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.096457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:bd7b32778185b8d659bb45574b64f4aa7c2c00fad9ba815579817fac943168d3

Observation 9c31247b-557b-4232-b4a8-1d464eac4536 · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.195070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:f8d2cc43747f619b7905111475ed660b7d8fcd9d5c2d8f2f02302e61b0525405

Observation f0ea401e-48d2-42cb-97e1-5bb6bd5864e1 · inbound

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization cites this paper.

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:24.087743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:01:23.060562Z digest=sha256:b77d8633e582d3dd7d3c1ba717d26771364e133f89646d81afd6491868f35e3b

Observation 6fec6123-dd93-45a8-be08-fed71f6a3750 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.705531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:ec0c44819946de80840ac4826735ed2fe9f0f751e75fdd9da230bef2d80137d9

Observation 5e34008f-04ec-4570-9bde-2018ab456caa · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:09:50.263174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:b57fd28ce3c5d56e176a8bc13e0d29bb27352d28d2130b4b90ec1e26e44c3279

Observation f2b4c485-182d-433f-a46b-7206a4820b9c · inbound

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models cites this paper.

OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:16.258450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T04:52:03.076788Z digest=sha256:48cb6d0f26a5dc152671d131eb0ce96cb6644199f4f91629da349a782584ad2f

Observation f4bb1498-5f94-4c4e-9d23-241b5b30ffcb · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.164422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:eaeedac640f53f601120aa556d434f4c73eda1f834b458705b396a88bd3b5d30

Observation 98d9eaa6-52e2-4e35-ae19-f7b6495b6242 · inbound

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning cites this paper.

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:36:10.573414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T06:34:57.483234Z digest=sha256:3137127c33220de51a861655f55320077dee0fadd8642982da93e0438af27c05

Observation b44a9268-54eb-4b16-9f84-0587a19943c4 · inbound

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning cites this paper.

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:03.460392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T09:43:56.299302Z digest=sha256:62584c57992d2dc3f02382b1eac75810f46521c14d12a536bc5f0deb90c89680

Observation db542710-299f-4a10-8c70-18144ce331a2 · inbound

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models cites this paper.

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.332893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T17:37:11.371892Z digest=sha256:f8847cb4018ceae676b35fb3a71db161e7c70003c5be60317bbf91d847311e6b

Observation be526fec-bd7e-4d07-86f0-23097d16a4f7 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:20:06.288028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:5dd93310ff59c8b78578f37ed50661c6cd58013d236d92ad92988bbaf390521b

Observation db441dfa-501f-404a-b6d7-55eb6216d76b · inbound

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy cites this paper.

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:51.450720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T05:06:09.216428Z digest=sha256:64b2a2f64bbe0f99f2e1cf584ef94b3096421bdf92b563a4863801043bba45e4

Observation 4f5011fd-6082-4ad2-83be-cf2feec23b41 · inbound

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning cites this paper.

Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:38:39.614890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T16:37:06.384435Z digest=sha256:cc6ffc755bf813c5d772dd570838df5738795d1fd4d49b1ab2a2a9adc7a1a79e

Observation 8890fdbd-95f2-49ea-9cb8-00dcbccba527 · inbound

Empowering Long-form Omni-modal Understanding with Robust Audio Perception cites this paper.

Empowering Long-form Omni-modal Understanding with Robust Audio Perception HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T12:48:58.688011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:48:58.688011Z digest=sha256:79e0f663298e0798e345a35feb0e2f570e2588a224f45cf2e6f75bbcb3ed1ae0

Observation 6dedb69d-cfeb-400f-85af-d9d6873b62f4 · inbound

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning cites this paper.

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:40:15.585890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:40:15.585890Z digest=sha256:4116e4629b91c9d86adfd97b919cdbaaf67f1c9198895244c375fe8576368799

Observation 7602b7d7-0cd3-4c99-ad6b-1853d8bd6901 · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T03:25:31.638356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:25:31.638356Z digest=sha256:859be9de332bd4488db2b7c8f38a2366cea58994774ce613989a46b5920bd4f6

Observation 2bcccdc2-9feb-42fb-acce-f629ddb82464 · inbound

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models cites this paper.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.515743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.515743Z digest=sha256:cb7461e577b0397bf7dc7d2f7022beabc168b4b4e3f61d00f5b3c863ddc38bea

Observation 5ff0a44c-6bed-496d-a2cd-b373b08cf750 · inbound

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models cites this paper.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.926459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.926459Z digest=sha256:b12ed2c7af8f7856c41dfbcdf9d3d4b43b1d98e8b612549659344d6926dc235d

Observation 1074ca18-4c48-469d-a37d-db46b91920bc · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.385618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.385618Z digest=sha256:8857ca0f2b34b1e2ca84ba68e9878cb1f059a6b22f946d55d4ec5c56194a22ba

Observation 311bef0b-dcd4-4b30-829b-aeffeb57682e · inbound

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? cites this paper.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.379357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.379357Z digest=sha256:c1a71d4d80024e49c0b9be8f215eeb8b1e750f7c1eb78c3951708c8d77bd249b