Pith. sign in

Paper Citation Record · LEDGER

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

As of 17 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2505.20256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20256 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:55.925858Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:07:21.383338Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.771664Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2232d5e1-3447-49bc-a637-195515e6dddc · outbound

This paper cites Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:50.909522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:50.909522Z digest=sha256:01fdac7b07ede9a6aa7dd23c2aee2b671655793b32246b0c2e4ce3920ddf7cc4

Observation d14e2908-41ad-4ccf-8550-d93ac69167d0 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.005221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.005221Z digest=sha256:afd3436e477e08609cfcab49fa5c10ee3279ffd9f4f88b14f40248ad29df400f

Observation de29e52b-f9c2-48b2-9b1b-4832bbcb07e3 · outbound

This paper cites Attention-based multimodal fusion for video description.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Attention-based multimodal fusion for video description

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:00.035632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:51.127713Z digest=sha256:1e00dbc814e8672b0ceec5659052d1acb523f19c62a7df614f82d6b4f1981c88

Observation db73eb7c-7f3e-4be0-910c-c0d9f3f905d4 · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Merlot reserve: Neural script knowledge through vision and language and sound

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.778055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:51.236879Z digest=sha256:77628b37387ce07a2c46a8eb32e34aef6123f25cb97e98d32357944fad84d053

Observation 1a800756-ad6a-4554-a341-e925a6bc26bc · outbound

This paper cites GPT-4o System Card.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration GPT-4o System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.365763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.365763Z digest=sha256:0f249ed1bb1f7ee730643a28996ee6228dedae4971d7f047e1f6b366f169d863

Observation 5a787639-7038-44a8-9486-9a03bd435645 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.438063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.438063Z digest=sha256:3d68dc996c53d3dab87f4e831ec9819fd099ecab4796bb9fdd4cd08a221fd0ec

Observation be76c2b1-c408-4058-bf43-522ab05a4069 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.602440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.602440Z digest=sha256:8822f329abefc6c69f03c29f3cdfa520af683132e203f452bf5ac812f6874a44

Observation 2dba246a-4fb9-4b09-9aae-66b5f621b3a7 · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.arXiv preprint arXiv:2504.10068, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Mavors: Multi-granularity video representation for multimodal large language model.arXiv preprint arXiv:2504.10068, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.665984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.665984Z digest=sha256:33e7a0b60ca69647697e45c9ca9cda1d09a1d8fb6d5208d5dd8899b057c57db2

Observation ec614a37-190e-4dd2-9939-6843c541cbfd · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Lisa: Reasoning segmentation via large language model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.626240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:51.756900Z digest=sha256:d21c9956a8daaa5549c09ed77b0db272b0695c9221927cbbb7ad5ea26dcb10d2

Observation 8b47e923-4132-4fb2-b4e6-cd716d3c4cb8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.849583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.849583Z digest=sha256:5ec72fc50d050fda8a6bfa718a140013440076eaf07d9aed6c398db73ebe3a3f

Observation da2c69d7-9b43-48ea-8f86-2b0e4b4ecd34 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.955495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.955495Z digest=sha256:60dc7264d38b6fcb0f90e177fcfc66902bac674cdadf7ea7a1195c7f73f73c82

Observation 89728534-4ffa-413d-875c-af011ef21c29 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.067992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.067992Z digest=sha256:a82f3721f141f9d8f5bc3e1a4b53880b1053a452d6015f9dacd99e456d4b137d

Observation ab764237-0230-4fb1-8c10-aa1c07c08872 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.167671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.167671Z digest=sha256:6ae688946ee4d5a25f9bb2243f72af780e33f133ffc79cb2059af136b0eb439b

Observation 8e81f7b7-2473-478f-ba07-1d29a9a0dee1 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Glamm: Pixel grounding large multimodal model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.476621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:52.272410Z digest=sha256:744e1e5f53283ed26096cb5a5f6a9e56ccd7c19c5ad6174b4f6db2d43fa18118

Observation b90be1db-5ac2-4685-bb08-193e845dc734 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.274294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:52.368454Z digest=sha256:003516330d0b9857b03d58c592e830a58079cf846237518b6294e29be227d3f5

Observation e806b0ba-fb70-439e-bafa-89354ea7f852 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.459678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.459678Z digest=sha256:4a925a82a1a7407874f479bcdb38ee9644450998792f9dd63252cd35938af8b2

Observation 1768ab9d-764b-4ddd-92f2-d163ff7de1c4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.541895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.541895Z digest=sha256:88daed296def81332c8abbc3fceff4bb99df7a3fb756f7b506a71dc13a73876e

Observation 0c77d780-6a4c-4f63-8797-8eaa616d254f · outbound

This paper cites an unresolved cited work.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:59.154520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:52.652896Z digest=sha256:efe243b686f8c469895236844bccd208b25c2d3ef96642e8c0bba17ae6d2faf8

Observation 4127edd4-722e-45d1-831c-5edaeb2445eb · outbound

This paper cites OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.741949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.741949Z digest=sha256:9463ce1078850efabb866e5df9f3f6500af7eb7796ce328a138f629d1fc6b3be

Observation 4ea6870f-df15-4826-9250-2e6d29f1f584 · outbound

This paper cites Ref- avs: Refer and segment objects in audio-visual scenes.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Ref- avs: Refer and segment objects in audio-visual scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.977483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:52.854703Z digest=sha256:ca78a3545b762efb5b4f0ca4110e69765393f827884ab59499bfa8b1c789c10f

Observation 7dc7196d-1d7d-4a5b-9b90-e1fa8f024dc9 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.905262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.905262Z digest=sha256:2f366419d52c9f31eba67d84b30f4500bfacd77d9accaeec237314e6e3f93acd

Observation 0836b924-e8b9-452e-96a7-5763e205d746 · outbound

This paper cites GPT-4 Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.016804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.016804Z digest=sha256:ac4bfdf509164350f348b72fed978b930361059ba24cd3ca59ef8106b9b2c0d7

Observation 56cf3a5c-718e-402e-a6fa-5b117e9945c1 · outbound

This paper cites Qwen Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.127369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.127369Z digest=sha256:1e5afdf15564dc290228f8c1bafbf1dd093dcd2880f392a4bdd41cef31211f13

Observation 37f867b5-ad9a-4d51-9c38-7b886fb81541 · outbound

This paper cites DeepSeek-V3 Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeek-V3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.175517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.175517Z digest=sha256:d44bf834606ef172cf72e2cd4d15281de49341688c365450337d6a8cc4b0ad84

Observation fb06e6fa-3c67-441d-a998-be2940b55e26 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.295895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.295895Z digest=sha256:ffd709b2f8d60f4286c6ce2a9893373f06ffd6e972bca77dd672f9ae369a5841

Observation 373e72bc-1f8a-40de-bfaa-f65fa07922f0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.381768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.381768Z digest=sha256:6e1a1b13bc061db450dfe99d0310bd7a3d7a3f8c0d90fed977cd172606e89101

Observation e126d7be-81b3-4c64-988b-4074a4487086 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.820469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:53.464889Z digest=sha256:26f6dbe880bec6b2cd22b27970c277e1104d14a720a6b12afda545fa98a1e8f2

Observation 811d0ee5-5e8d-4291-b3e4-a069be9de93d · outbound

This paper cites Visual instruction tuning, 2023.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual instruction tuning, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.519220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.519220Z digest=sha256:b47e0050eea0bd198fdd6bdfcda5528afbe7b7e62c857ffb1bd5740ae1f03876

Observation 8a7f098b-4c37-4905-9a09-dca496a1a4fd · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.724157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:53.626176Z digest=sha256:a8f452408ac70615d8702cc3b48a4a673ab612f3cfdf277ccebe442e6acffaf9

Observation 0fb3a83b-ac11-422f-95b0-30d323ede06c · outbound

This paper cites Qwen2.5-Omni Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen2.5-Omni Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.705119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.705119Z digest=sha256:0463a465f764323c2703e58229a3bb7a324cfddde8b8cdceba24ac0d5e878fff

Observation 1596d63f-ea15-4435-9835-bba8ac980638 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models, 2024.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Omnibench: Towards the future of universal omni-language models, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.455156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:53.786386Z digest=sha256:373432e17e1c19ba8be6b00de4b42717c8cd7984a35c0392de6a738bcd4004bd

Observation 7acc720e-527d-4dd3-ae3e-2dbdd88aae6c · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.893277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.893277Z digest=sha256:158c9348e26f716a825ca71697a3014d2eca44efc8aca8da35544c2ee6d35cb6

Observation dc2370a3-20a0-475c-ab8b-8c05fdd5def9 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.986036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.986036Z digest=sha256:4a4d9223e2c37a98d30852193947c935b7f2a6e570fef13329b37f83d63bc0ef

Observation ed74b537-c464-4e8c-ae1c-f4a2f2fb4fba · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.059626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.059626Z digest=sha256:519c74bdee2d614549023487dab08130749e845b534af42a43d323c7b57dea4e

Observation 5f7eead2-aa70-42dd-8e9d-189cd6d4bd3a · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.164163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.164163Z digest=sha256:bf79c498fa703050bd520ffffff290b04ebb0c3cd914d9bb794f8859391c5d84

Observation 84479aa3-2f2e-43ac-82ad-7c4c856e9879 · outbound

This paper cites R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.301784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:54.253030Z digest=sha256:db88e87e643736e57c4478d4025f5e4cad6385ad04e35362bf81e36c5cd08b95

Observation 76ba6c11-fdad-4b64-8791-6417ff75765f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration SAM 2: Segment Anything in Images and Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.312129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.312129Z digest=sha256:0c6c0d689b66effc139a17d5a28d28e083b466a3fdf1037621ced3d53dfe148a

Observation 091839ed-40c2-4cb8-931b-5ad43caaca4f · outbound

This paper cites End-to-end object detection with transformers.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration End-to-end object detection with transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.165672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:54.373043Z digest=sha256:1dc1a160664af3fceaefe9f4fcecbb2e635fa261aa3deadc8cf271c888542760

Observation ee01d021-eb3e-42d2-a4e6-7f0ab97fa304 · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.046533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:54.490953Z digest=sha256:10b36d98be05ee8a91742cdde976e30ba64ae6423a30a5ec87c0d16fc87cd5e3

Observation 9a828aec-1762-4492-81f7-72aa9620e47a · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Generation and comprehension of unambiguous object descriptions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.852661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:54.574343Z digest=sha256:0a09be520f6c2fbb0d2db854faa0ff611b52ad3b74aaf32ebb34708dffda975b

Observation 2760fb1d-f276-4eb6-8b9e-6aae6a500311 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.636835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.636835Z digest=sha256:b90089b11337940852aa9e8d42e8827f21a07930c2b9f702e7d1914cb74bc820

Observation b1dfba77-6695-479a-82fd-bed20612b408 · outbound

This paper cites Avsbench: A pixel-level audio- visual segmentation benchmark.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Avsbench: A pixel-level audio- visual segmentation benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.620689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:54.754197Z digest=sha256:19f4f8aa19c8a1e2a3a12757eade3e47f2edcf3dbb13be1058eb30463dcadba3

Observation d024c242-252c-4501-ae1f-6bcb9173d017 · outbound

This paper cites Avsegformer: Audio-visual segmentation with transformer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Avsegformer: Audio-visual segmentation with transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.429546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:54.861182Z digest=sha256:96cffdb835d0b8df7a9806c064a961de2d891eecbe7bea35e36f1bc998bd8459

Observation 2121f94a-ed7e-481f-ac14-118ba329eea8 · outbound

This paper cites Prompting segmentation with sound is generalizable audio-visual source localizer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Prompting segmentation with sound is generalizable audio-visual source localizer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.234971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:54.935916Z digest=sha256:a1e73e826f73ee2418dd4094160b586b58ba7fd66338d3f398befb1714b119a1

Observation fff28671-19d9-4f53-95d3-f7cedab4b6f5 · outbound

This paper cites Language as queries for referring video object segmentation.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Language as queries for referring video object segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.080900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:55.054466Z digest=sha256:d6752a7890693352529f2457a166ac3da2c7c10267a9d61f2d370f99b6002d11

Observation b616ca86-1092-41be-9e2b-4c47089ceeb6 · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Robust referring video object segmentation with cyclic structural consensus

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.912927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:55.159388Z digest=sha256:8cfdccaa4facaac034da1fdaf111068d2da7b4521695a908191093c4a980ea7b

Observation 28b03452-ecb1-438a-bf2a-f2a1a82b46a0 · outbound

This paper cites TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.274394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.274394Z digest=sha256:edf28340d29e73586ac1065ec2777fa50aaae6c9a94a6fb162b973498c86e3eb

Observation 05bc131d-b4c1-4755-91d1-174dceae61f9 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.376208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.376208Z digest=sha256:e28e985ff051073e728d714cd63aa05b2859da9d7db0c928b12565872db417db

Observation ed5cd32a-4c1b-4cbd-b21c-9fe3843808f9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVA-OneVision: Easy Visual Task Transfer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.437771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.437771Z digest=sha256:2a3f394b39e92b5c73051001ebdf9f499ef180d832f9a89672034679d3448b5d

Observation a3fcd63b-ceaa-46f9-bd7f-2fcd422481f8 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.535106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.535106Z digest=sha256:1a6effcdd6d9302440fb247c4a02dcf4e8a80fb307dada24d9adcefc2261ceeb

Observation 41704680-dbff-4445-9e9d-9142addf6d12 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.605882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.605882Z digest=sha256:064f10e29ffc26191eda1a5f6e6c80a37bc6f11a4eebed5c2f11ec09a4315920

Observation 54bdf31a-b38e-46dd-8339-8ec5817828e8 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.662704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.662704Z digest=sha256:e59eb098710c8578b7f4d1d10977e3c7a87511bd8024afeaa432cd8f2793dacc

Observation 30c68b51-6c35-4cf8-b986-ded5388dec21 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.721329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.721329Z digest=sha256:bc6e9d70422a4a808f8835320be03f6cfb9ecc96767bb8aa2c661cdab4b032d0

Observation e8162af0-fae5-4271-94fb-4f446fa0a774 · outbound

This paper cites an unresolved cited work.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:56.762742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:02:55.829573Z digest=sha256:0353f6cab69deed64df770d2e440bc197a9f7bf8d1b9813e52a9257fded4027f

Observation ec44756a-9734-4793-b05a-a27563163f71 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.925858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.925858Z digest=sha256:cf0f9fa9c3500c68fa40f5ebd785847111ca21a097638b371d9ae98e534cafc9

Pith citing papers

Observation 748bcc57-ef43-4c6c-874a-623cafbcc7f9 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.383338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.383338Z digest=sha256:4c97ad4792c00dc167e8608d06bb1cc32abe175dbe4fee832b9719ccec72259d

Observation 11fa81bf-7cdf-4185-802d-a848fa6ccfb4 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.128382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:63fc5ac212d65ef0c195d6736c706727ccf253893986bc55876617830d5ffab3

Observation 9a349702-83c9-405d-ae4c-4abb8a12968e · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.120343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:f253f823d4df264f691c72c4d3ec62d9c5e847b5ee19c8a05e961dc2566e43fb

Observation 60a8a706-ac62-4312-9138-576e90d3bfca · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:59.982181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:eda3533917f897143d2de1677928c34b95329aebf3993961dc0cbb55d2df1207

Observation a35bda6a-7128-40f8-b307-a518e00bcfc2 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:11:01.973351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:4e33368b2f891898b6723ca9b36a9f65e7ca86c46943ea1b6081933e6dc0c77d

Observation 5cd62825-00f7-443b-b17b-2f24b32fc7a8 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.827672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:12c88bba093fab4f030a0bf6f8f11cd63ffc3dbfd9ff305ae09cb58d48e5f2e8

Observation 3eb7c43c-ccbd-4e9a-a055-c0581c6ba1b4 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.298818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:f06c3d921f66c6a17a39ac756919e065e917e1af7d95eb7a6b7bc4d3eea998af

Observation 6eb1a9ba-ff0e-40fe-a774-7737a268bd6a · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:5735ba5c8174d7c92d32d1e1c38a4e70b74fe8dbaba461b303f0123d2ddd7409

Observation 33213efd-ece5-4758-826a-c2693ae4e7c5 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.082628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:bdd16a2ecf0a8c274978e38956ba84275f680620e09c3317846ed9a7cb3bbd46

Observation 6baa4af0-3229-4b99-b6c1-dfea506983bf · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.211166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:16bd2e4590e0a45319aed959df65b877a927cd0f6e752493d76358b063f62984

Observation cf8a1670-c68e-4524-ad51-1587028d1389 · inbound

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition cites this paper.

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:53.325802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T02:30:44.486435Z digest=sha256:63f7c94193dac3384505f92a8124039e7568905d30aa448843522de4867c910c

Observation e46a2ae7-4b41-41ec-9c82-fcdb79ddfe19 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.904231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:54611c3fc557e343fbf8d80de1382c2d8cd4d10dade964ac24af9b6758e4fadb

Observation 59631231-82f1-48ef-91b3-68a6c3106ccb · inbound

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching cites this paper.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.106597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:a9add58da0ff7d068f23412c8b1e3be22e983c82f0fe5915ea03f66df239b446

Observation 6c937d94-f391-4895-86bf-b614ce2e5c02 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.773107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:474d1eec94d2ca8f344a940e46e8e2e728518b7d630b66699ea57e077f05fb2f