Pith. sign in

Paper Citation Record · LEDGER

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2505.20256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20256 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:55.925858Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:07:21.383338Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.771664Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2232d5e1-3447-49bc-a637-195515e6dddc · outbound

This paper cites Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:50.909522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:50.909522Z digest=sha256:6cead4621b0eeb16bcb257523b84d1a6f61ce5b54bfcc236ce737db4b6629fe1

Observation d14e2908-41ad-4ccf-8550-d93ac69167d0 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.005221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.005221Z digest=sha256:e495b70e79f55d967779c616ff30e0fd1ecb79a7bc05da65535aef96545e5979

Observation de29e52b-f9c2-48b2-9b1b-4832bbcb07e3 · outbound

This paper cites Attention-based multimodal fusion for video description.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Attention-based multimodal fusion for video description

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:00.035632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:51.127713Z digest=sha256:2671c3775657e1e1a66984098aa31868c89a58b394441300f6cc4dfd277ffecd

Observation db73eb7c-7f3e-4be0-910c-c0d9f3f905d4 · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Merlot reserve: Neural script knowledge through vision and language and sound

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.778055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:51.236879Z digest=sha256:9dd95945d423fe57c5447785dc87cf2f788fa9aca49ae09e13cb305298551a36

Observation 1a800756-ad6a-4554-a341-e925a6bc26bc · outbound

This paper cites GPT-4o System Card.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration GPT-4o System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.365763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.365763Z digest=sha256:01c3ab89269fbf4e946f6d54c7a01300f60399c148880ab52633bd7eb90b08ca

Observation 5a787639-7038-44a8-9486-9a03bd435645 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.438063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.438063Z digest=sha256:de484041db140845cb425663236914adc134a8726a51b2f65030ce2fd55aa10b

Observation be76c2b1-c408-4058-bf43-522ab05a4069 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.602440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.602440Z digest=sha256:d29695b70d2f2089a51c7c8b03fde2598c15e4e02a75f96acb5e0048af2e37cc

Observation 2dba246a-4fb9-4b09-9aae-66b5f621b3a7 · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.arXiv preprint arXiv:2504.10068, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Mavors: Multi-granularity video representation for multimodal large language model.arXiv preprint arXiv:2504.10068, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.665984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.665984Z digest=sha256:1bbbe9c47b0e009428f61e05e88e357e69675ab861c4ffea5d07ac6e903a7135

Observation ec614a37-190e-4dd2-9939-6843c541cbfd · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Lisa: Reasoning segmentation via large language model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.626240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:51.756900Z digest=sha256:e70ca542cd00982755e139a5fd93e6d68dada0f1a1a002417100436bcb492ded

Observation 8b47e923-4132-4fb2-b4e6-cd716d3c4cb8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.849583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.849583Z digest=sha256:8a5072dee3908b2ba5855563cd911927af0f1a9dd7d9bbe524a8260d85d0d731

Observation da2c69d7-9b43-48ea-8f86-2b0e4b4ecd34 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.955495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.955495Z digest=sha256:159b0ef1c566d642989884a1de283e8d581a856f5169462c970fc51776bd485e

Observation 89728534-4ffa-413d-875c-af011ef21c29 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.067992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.067992Z digest=sha256:f8fe46deccb1f792121c24642edd0e4ecac59cf74c71b71689bbe17e5edb4baf

Observation ab764237-0230-4fb1-8c10-aa1c07c08872 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.167671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.167671Z digest=sha256:d25e371e688a82e53d5807899401d773164b4856771c99c2b985adb9831c6de4

Observation 8e81f7b7-2473-478f-ba07-1d29a9a0dee1 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Glamm: Pixel grounding large multimodal model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.476621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:52.272410Z digest=sha256:d814c295be115434d0f1ef544c1713ca70e82002c6a41f344d699f67e456ee7e

Observation b90be1db-5ac2-4685-bb08-193e845dc734 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.274294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:52.368454Z digest=sha256:bff9f834f1ab2961cdc39b1431a52dd4bc42bc4236c64c84a1d7f96f862ee520

Observation e806b0ba-fb70-439e-bafa-89354ea7f852 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.459678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.459678Z digest=sha256:9e311d47b72db56475e244827232f105f94cfcf2b0738ba5bf9c4beb86297b69

Observation 1768ab9d-764b-4ddd-92f2-d163ff7de1c4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.541895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.541895Z digest=sha256:793daaf14d08d496fdea90d96f18ee04468392a03fd6bd4da43111e9c4dc6238

Observation 0c77d780-6a4c-4f63-8797-8eaa616d254f · outbound

This paper cites an unresolved cited work.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:59.154520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:52.652896Z digest=sha256:b12078a06088638dfbd9f9aaf96d87c48bba11a2ff187523c9b0f5b259133e0e

Observation 4127edd4-722e-45d1-831c-5edaeb2445eb · outbound

This paper cites OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.741949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.741949Z digest=sha256:4efa9500f5b6711117d7a073ac80604373d1c91ef40feb600cee1321e714f220

Observation 4ea6870f-df15-4826-9250-2e6d29f1f584 · outbound

This paper cites Ref- avs: Refer and segment objects in audio-visual scenes.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Ref- avs: Refer and segment objects in audio-visual scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.977483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:52.854703Z digest=sha256:eb9b8280383acabf45968d6743161e1230c8f4924c794d6fefc947ce23d022ae

Observation 7dc7196d-1d7d-4a5b-9b90-e1fa8f024dc9 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.905262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.905262Z digest=sha256:c36f285ffcf40b9def41f28a325a332f93573f58c844ff7ea40df53e4dbeacac

Observation 0836b924-e8b9-452e-96a7-5763e205d746 · outbound

This paper cites GPT-4 Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.016804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.016804Z digest=sha256:d2d31009e59e850d22d33e053491faab9e74b0d7080dfa35643eac9e53b5c4e0

Observation 56cf3a5c-718e-402e-a6fa-5b117e9945c1 · outbound

This paper cites Qwen Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.127369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.127369Z digest=sha256:f5f8856ac28a486cb1e62a7130a43012a387eb032482bc1a5c187a853071175d

Observation 37f867b5-ad9a-4d51-9c38-7b886fb81541 · outbound

This paper cites DeepSeek-V3 Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeek-V3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.175517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.175517Z digest=sha256:675364330b7141ca04f374b7646d889d319503e66d4111e0b5d343baca15eb0e

Observation fb06e6fa-3c67-441d-a998-be2940b55e26 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.295895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.295895Z digest=sha256:e8e9ab5d8a2a0b0f8bce95f0dfa4bdf2919d5e08d5ceaded992c18923d1f3b79

Observation 373e72bc-1f8a-40de-bfaa-f65fa07922f0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.381768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.381768Z digest=sha256:f9219570000fbdd613b5c0574aa197894928da183ebb64142768658724c0e381

Observation e126d7be-81b3-4c64-988b-4074a4487086 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.820469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:53.464889Z digest=sha256:c04a7e9521da4c0c16406a0e4ad02ff1cbc5fd465b628b79fa7e15c526b3d4ad

Observation 811d0ee5-5e8d-4291-b3e4-a069be9de93d · outbound

This paper cites Visual instruction tuning, 2023.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual instruction tuning, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.519220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.519220Z digest=sha256:4befee49d4cc19ef2088f181e6a61ba6efab5c8381fa3454da9bc2119c9299ad

Observation 8a7f098b-4c37-4905-9a09-dca496a1a4fd · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.724157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:53.626176Z digest=sha256:61f991c52d0275e9916a0b434816a8db63ae85aeca13cd325b5edc70fc3c2a37

Observation 0fb3a83b-ac11-422f-95b0-30d323ede06c · outbound

This paper cites Qwen2.5-Omni Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen2.5-Omni Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.705119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.705119Z digest=sha256:16627320ff25813140efc55fa48693c9e0d2304ef05e6fe9ffe2d5c1b88a108a

Observation 1596d63f-ea15-4435-9835-bba8ac980638 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models, 2024.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Omnibench: Towards the future of universal omni-language models, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.455156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:53.786386Z digest=sha256:001d0cf1e0f51eb3272949c0b2b1e6dfbfec5f5c476d2211c4b38ec54c9a0afb

Observation 7acc720e-527d-4dd3-ae3e-2dbdd88aae6c · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.893277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.893277Z digest=sha256:c0b708d8edc8fe275a6b56b3d010654c52dfe0df5fbb1f19e578d1eca9f811d4

Observation dc2370a3-20a0-475c-ab8b-8c05fdd5def9 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.986036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.986036Z digest=sha256:fa07c166e9501e3aa1397a4b4d60daaf0165c50a97ec07ab2656c8dba1ecfaf0

Observation ed74b537-c464-4e8c-ae1c-f4a2f2fb4fba · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.059626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.059626Z digest=sha256:b4423b12e39772f3fb42a40097e38c930f9a6a55ab64593042b8252104b75061

Observation 5f7eead2-aa70-42dd-8e9d-189cd6d4bd3a · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.164163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.164163Z digest=sha256:440443139e75cd224f21c8e22593490400c031c3127044815ffc867ac7d4d1ab

Observation 84479aa3-2f2e-43ac-82ad-7c4c856e9879 · outbound

This paper cites R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.301784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:54.253030Z digest=sha256:d7ca97f07fb45066c18a1eb1a44dbc9d490d550efa29153408eb1afe33a42a88

Observation 76ba6c11-fdad-4b64-8791-6417ff75765f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration SAM 2: Segment Anything in Images and Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.312129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.312129Z digest=sha256:64fdde18ea2d573830b190e86d5336c78eb8cb136684907e8a714968d39e2be3

Observation 091839ed-40c2-4cb8-931b-5ad43caaca4f · outbound

This paper cites End-to-end object detection with transformers.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration End-to-end object detection with transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.165672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:54.373043Z digest=sha256:e8e03ff2e470517b12a62df14cc76bbcee2159c60300df2612f5ef11f511ce80

Observation ee01d021-eb3e-42d2-a4e6-7f0ab97fa304 · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.046533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:54.490953Z digest=sha256:0f0fa9272dff755b0495bba41077e64475d1385dbd82e7cf4e24c505361f504e

Observation 9a828aec-1762-4492-81f7-72aa9620e47a · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Generation and comprehension of unambiguous object descriptions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.852661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:54.574343Z digest=sha256:9ee995942f61175df7199fe83c3e4343f636a6c183a6225b2453083e9d213f1d

Observation 2760fb1d-f276-4eb6-8b9e-6aae6a500311 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.636835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.636835Z digest=sha256:cc9562dc9e5d4a662ab4d5b65fe89a9937cd8440b9c6407d5bc63084ce7aab43

Observation b1dfba77-6695-479a-82fd-bed20612b408 · outbound

This paper cites Avsbench: A pixel-level audio- visual segmentation benchmark.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Avsbench: A pixel-level audio- visual segmentation benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.620689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:54.754197Z digest=sha256:ce1a603bceb432c5e4ec71b2e64f1e43416452f0ea7f057a68b3d99adb905467

Observation d024c242-252c-4501-ae1f-6bcb9173d017 · outbound

This paper cites Avsegformer: Audio-visual segmentation with transformer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Avsegformer: Audio-visual segmentation with transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.429546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:54.861182Z digest=sha256:01f359a267418289f044d39be42f3a908ee45d06e84a073270df9116ed6c04c8

Observation 2121f94a-ed7e-481f-ac14-118ba329eea8 · outbound

This paper cites Prompting segmentation with sound is generalizable audio-visual source localizer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Prompting segmentation with sound is generalizable audio-visual source localizer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.234971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:54.935916Z digest=sha256:90fdacf8f56a04208f14e5ddfed2d9d736a8cf48025afb0b6ecf8db86751d62d

Observation fff28671-19d9-4f53-95d3-f7cedab4b6f5 · outbound

This paper cites Language as queries for referring video object segmentation.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Language as queries for referring video object segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.080900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:55.054466Z digest=sha256:d1a45c99af2aff77b415bf305cf0060859c7885409afcca57093fb086ab738a0

Observation b616ca86-1092-41be-9e2b-4c47089ceeb6 · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Robust referring video object segmentation with cyclic structural consensus

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.912927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:55.159388Z digest=sha256:83b861d794f2966ca28fa77e05a6f62773d660ec5bdd97f021567a68546025be

Observation 28b03452-ecb1-438a-bf2a-f2a1a82b46a0 · outbound

This paper cites TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.274394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.274394Z digest=sha256:5199012f39304e75c0484b35bdc0d05e7e406bf93aa7d7d90e0020d8918cc7c4

Observation 05bc131d-b4c1-4755-91d1-174dceae61f9 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.376208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.376208Z digest=sha256:24639deef923ed8484770ad58dee15ba975d32ea3e8496b0e74c9ab90bba68ad

Observation ed5cd32a-4c1b-4cbd-b21c-9fe3843808f9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVA-OneVision: Easy Visual Task Transfer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.437771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.437771Z digest=sha256:ac8a32e9b8fe86938cfc774857fb4235cfe666848660ac6b4312c43a8bbe4267

Observation a3fcd63b-ceaa-46f9-bd7f-2fcd422481f8 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.535106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.535106Z digest=sha256:d88eefb89f235f61b062df09b8979ce096fa88c904936b4ccf764d1fb3a5baf7

Observation 41704680-dbff-4445-9e9d-9142addf6d12 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.605882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.605882Z digest=sha256:837c69780b83534131454c33ed036f4526e5c9a6710bd99289a791260a56de00

Observation 54bdf31a-b38e-46dd-8339-8ec5817828e8 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.662704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.662704Z digest=sha256:c5ba9752892dcddf9403fef48e95d5d6dd5fcafae788c8a5166e51ac8c2df45f

Observation 30c68b51-6c35-4cf8-b986-ded5388dec21 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.721329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.721329Z digest=sha256:4d2687b990fab7f8821c3cc168cd03b23c3f808efe859464c00baa5de7d54ec2

Observation e8162af0-fae5-4271-94fb-4f446fa0a774 · outbound

This paper cites an unresolved cited work.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:56.762742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:02:55.829573Z digest=sha256:414f7046f1f207c19a3d8e4621e1af3928d999b828a45b726dc62ef109ee04e6

Observation ec44756a-9734-4793-b05a-a27563163f71 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.925858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.925858Z digest=sha256:878b0344413bc453c91d2140b3a2b815159999cbd222a646e6161629bacda2c6

Pith citing papers

Observation 748bcc57-ef43-4c6c-874a-623cafbcc7f9 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.383338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.383338Z digest=sha256:d818f624da6930a1a3d702af6b4cd7dfe5e060391d5d5bed7df3fdfafa019976

Observation 11fa81bf-7cdf-4185-802d-a848fa6ccfb4 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.128382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:5524cb35b66affc86527efb1c302b40930a72784bc058cf61549ddf8efff106b

Observation 9a349702-83c9-405d-ae4c-4abb8a12968e · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.120343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:6316314fd0f27833977cca77cb1ccb25787e29ad7c9fa2da156dde98ca9b0503

Observation 60a8a706-ac62-4312-9138-576e90d3bfca · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:59.982181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:088b7b2ec7b723c2a477e4415bd166f9e6f996d037c09e9073a94003b61f6eaf

Observation a35bda6a-7128-40f8-b307-a518e00bcfc2 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:11:01.973351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:b250a9a970060cce58399f7f5789c8962410597ab695053c93a2e2548fa23cea

Observation 5cd62825-00f7-443b-b17b-2f24b32fc7a8 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.827672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:6107922394fe413a3e894db2cf0f4bb4b54b248ca9529e571c2a801124cb2065

Observation 3eb7c43c-ccbd-4e9a-a055-c0581c6ba1b4 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.298818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:572ac397e703fe34a191344008bf8cbfe1112f3155a705a8c496e965abd3f9a0

Observation 6eb1a9ba-ff0e-40fe-a774-7737a268bd6a · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:f1015ea5c7e9998c8d628acf3fbf30b0486939a74983e3acc8b2c330a40eb01a

Observation 33213efd-ece5-4758-826a-c2693ae4e7c5 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.082628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:56860a3511912b0c40be136e044fdee1981187e0f21f7be19f9dda151c8c4997

Observation 6baa4af0-3229-4b99-b6c1-dfea506983bf · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.211166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:c3337a5f6bcfbfb79fbaf2e810ce13dec7030c20daa67ea4665d42af0b832a47

Observation cf8a1670-c68e-4524-ad51-1587028d1389 · inbound

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition cites this paper.

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:53.325802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:30:44.486435Z digest=sha256:7d480e743f731e831e4cacc767d29ab5617a8aec3e2c50e47a25b49532f38be9

Observation e46a2ae7-4b41-41ec-9c82-fcdb79ddfe19 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.904231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:5f56959c8f91882451ef1ffe749e92016c0bde7736abba1b17618d7001be8c34

Observation 59631231-82f1-48ef-91b3-68a6c3106ccb · inbound

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching cites this paper.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.106597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:f6cddc0b5335db442f75c5ecb6c248a53f468e33ff18e74df368249e3834a9cf

Observation 6c937d94-f391-4895-86bf-b614ce2e5c02 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.773107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:8fe460c69401b3f9161925c2ea3bad4f3e5a96886b03134aa51e32f174db8ad0