Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:01:18.518663Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 35 inbound Pith citation observations for arXiv:2508.04416.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:01:18.518663Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T05:58:12.505711Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
15 of 15 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 9bf4b689-edaa-40a0-b169-434f072761e6 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning As illustrated in Fig
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c06b0847-a345-4f76-b3db-c840703af3f8 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e44997b-56e4-4344-b202-d0a780dbe678 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e95ea3e-1340-477f-b457-ce21aecebbb9 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31ef202-e132-4529-a38c-ec59d6d81794 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 88c03a28-2d84-465a-8aa7-c480a48e061a · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc806f4a-465d-4b52-8987-6886d89fc911 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Do not alter the given ref in any way
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8994c051-27e3-48d7-8da8-aeed4dde6785 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning If the object is unique in the scene, a slightly more detailed description than <s object> is sufficient
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 74af6b7c-0cb2-419a-8ad4-cab413254677 · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning This complex ref must unambiguously and accurately refer to the exact target object uniquely identified by the uid and its associated pixel mask in the video
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 698c8b69-8512-454f-aaaf-0a5c9c398e2d · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Crucially: avoid any simple, direct descrip- tions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0f5df691-9c33-4820-9a98-26ce02fc306c · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b932b4fc-e7c5-45e0-84a1-be76a92f7c6a · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning We observe that the majority of CLIPScore values fall within the range of [0.8, 1], indicating strong se- mantic similarity
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation acaddf52-4655-4fb4-9ffb-019b7e26124e · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning In CVPR, 19108–19118
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b08d946e-b2e5-4dc2-ace7-5b1c264ee31b · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc7d1e3-6c86-42ca-a4ee-6686936fc6ae · outbound
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2ebbb9-6cd8-4aca-b5aa-dcd69f3dfad5 · inbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation decedb93-683f-46c0-ade1-39bc7d6a6c37 · inbound
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e8a7963-12fb-423e-b531-20ddac204a8f · inbound
OneThinker: All-in-one Reasoning Model for Image and Video Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c8449ed7-602c-437b-b133-fa54015da88b · inbound
Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdaf31ee-5295-4b34-94be-02fcc51a19b7 · inbound
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8ef12da1-0c3f-4e51-aeab-d8f38fddd796 · inbound
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2ab52996-c550-47cc-9227-f4c788ad34e3 · inbound
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590a387c-64df-4182-874a-c6d345964da9 · inbound
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b3a6c40-0dfa-4452-a3ea-117cff48f0b4 · inbound
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 70c70560-a7fa-4a35-906f-cfabf5963d6e · inbound
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65bb68ac-ab07-47ad-b95d-ca38c9340817 · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4bd62dca-4a09-499b-bc07-6e22a784805f · inbound
Towards Temporal Compositional Reasoning in Long-Form Sports Videos Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b638aa-bc33-4a04-9004-bec1c65f6b51 · inbound
Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77e937b3-811d-4d06-8936-40a7ddb33a2b · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b94cf810-3ece-406a-9d8e-4ce5a796e368 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation adb14a59-39b4-47f6-b8bf-bd693545d279 · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7bf99d19-731b-4993-bbad-f79538464a1e · inbound
VISD: Enhancing Video Reasoning via Structured Self-Distillation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 76a612d9-b5dd-4dd9-87ac-42cfb4cbbc39 · inbound
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dc24a0bc-2ecd-4ac9-a3cb-69a18fe83560 · inbound
VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db95fa81-b137-4ef0-98d9-dad96a32170f · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb64a801-46c1-4a1a-96ba-865f08b1e2d4 · inbound
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a015ea5d-be13-4fe4-88d9-85fbbb21645c · inbound
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c0ed52b8-4eab-4acb-b59e-023b3e58e831 · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9e4ab436-1ca0-42e6-9997-ffa36ee1fb70 · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c00334ef-161f-491f-be05-6bc80315fd53 · inbound
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4094a22e-e0d7-47df-a43f-674766ef90e5 · inbound
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 368d2bae-e360-4f63-87f2-626cf7a6c461 · inbound
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 26674536-e164-4883-88af-f7ed786e4eae · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6fb61329-55b6-4ac2-96a2-70ba13836a0f · inbound
DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 944ab3b2-5312-479c-b9e6-dd6e3d6c2e17 · inbound
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 683ccd4f-4c63-4a4a-ac9b-5b1c2838605e · inbound
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d72fa15-8089-477a-bab5-02b19d16776e · inbound
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a0e3e4e-fe69-4804-b274-743ef21e3497 · inbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b943b3f8-bd6f-475f-b17a-530aae2b75e3 · inbound
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22419d38-e134-4199-bea9-b3855609dab1 · inbound
Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.