Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.860880Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 9 inbound Pith citation observations for arXiv:2506.01725.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.860880Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:29:11.589395Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T17:40:00.934274Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a15c7f71-a533-4b6f-90c3-a2dca39bc095 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be9052c9-8e93-4861-a699-4f87d8bb6797 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Activitynet: A large-scale video benchmark for human activity understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9c0c257-e40d-4a50-9b57-4793d2ddcc92 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Auroracap: Efficient, performant video detailed captioning and a new benchmark
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87d10cb7-bbce-4570-894c-48ea936d37e7 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge dis- tillation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef9d158-91d2-4043-b920-430cefab504d · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03135c81-6c8b-4354-b373-5950dd5531e4 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb22eed1-3201-432c-9962-4d3ab340b08f · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41784efd-54cd-4147-bb3a-2519b8b8319b · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d9c79d-4ff8-4bf6-a9f4-4936233864b4 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tall: Temporal activity localization via language query
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02ca2843-d1a7-4209-ba52-dea0399d67e4 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Scaling laws for reward model overoptimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7c245ae2-6d72-4855-8bec-7c8f7ad09bc5 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eead944-f033-4430-97d5-61140c8c3ef9 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68037354-34b0-40c7-954e-572ffe521b97 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a3b0e9d-49b5-48eb-889e-a442eaca222c · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903b84b7-4c7b-4e9a-ba12-15ed97a07435 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4o System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b7bb67e-93e3-4e1f-9da3-d259de55f54a · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2b9347-9f47-422d-8df5-d5a7eaa3ad00 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff83fa10-ffe1-48b6-9b70-6a8e139e6a8c · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking A shortest augmenting path algorithm for dense and sparse linear assignment problems
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3210f0e7-20a3-4d2f-9319-37e1bc1b2cad · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-OneVision: Easy Visual Task Transfer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64cd84fb-ff43-4cac-9c69-120eafb54e01 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 00491f7a-f816-44ec-89d2-873dd68e7f75 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54b9f27-0590-4e6d-b7d4-57ac29bbd53b · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Let’s verify step by step
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e962be59-6051-4662-b6ca-b4f3a45d3b9e · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929c53b7-30e9-4924-bd77-7fa1adf09ad0 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463aad97-165e-47ee-95e4-5c492569d533 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c097e315-49c3-4fed-a816-21edd16b6721 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Introducing gemini 2.0: our new ai model for the agentic era, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69971c8d-693e-4318-ba6c-b3316bc9db5f · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Proximal Policy Optimization Algorithms
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77a72e1-7acc-4f7a-a6fb-89a92ca927ce · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e13b280-d600-4202-a8e2-37e22e721598 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecbc943e-6897-4c25-9e4d-093006c3b49f · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 077385f6-de87-4489-81d0-4e9c4156c2e3 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce95458a-a9b4-41d5-8105-9a804e398998 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 835dc834-d0d0-4557-bfba-38afbae77af2 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1befbf9-e216-4ae7-bb19-6a771e49d146 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be61c09-24ad-45a7-95ff-96a0817528a3 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42b6749-3580-4a62-b6eb-d81860a6ce66 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84cb20a6-1314-4606-881c-7b4d6619932a · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2.5 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19a02cd-8847-4a5b-b60b-5160165c8c7f · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Vript: A video is worth thousands of words
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cfe4d595-bfd2-4502-834d-9b58e82f605a · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11382a79-e212-483d-acee-0e4633d8e18a · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b38cb08a-7b3d-4787-9fca-3455a433d275 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Modeling context in referring expressions
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e9c9c5-ee26-4ec7-af1a-7aa1e7509853 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b25c7c-7ad1-44a4-891f-fc297f035177 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b58ff38f-0a06-422c-942e-7381f7305b9d · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26bff47-1afe-46bf-aed8-1efb8589414e · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Improve Vision Language Model Chain-of-thought Reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19ab9372-e3a4-457d-8ffd-1ef64bb7114f · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfc7a347-b255-43b6-97df-b3e9738fef0d · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 568a8494-8f7a-48f4-b86c-3834ef178039 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c016390e-52e4-4925-aa35-151145ec6a97 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48094e61-f3c6-4174-be5d-b770a4d99ebe · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5ca7e3-48c9-404b-97c0-cf111daead75 · outbound
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e5b183dd-b5a6-4e41-9796-1db3ffebe5ec · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 292
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff98ddf-19d9-4d22-a939-46b89e999320 · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36e73833-b17f-49e4-9a20-48e792f14e45 · inbound
Building a Precise Video Language with Human-AI Oversight VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef22b1aa-0882-4abb-b11b-72a052fbcb76 · inbound
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8439cfd-3612-4e2a-a1f4-fcd335241947 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b6a4f68f-becd-49ee-a3bc-6334f5766267 · inbound
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9066af67-d826-4a5c-b0d1-5ef885e470a5 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 194
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b32ac67e-4cd7-421b-a1e9-b83d10d18411 · inbound
ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32461e25-5782-49b6-a010-1cc41b265953 · inbound
PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.