Pith. sign in

Paper Citation Record · LEDGER

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

As of 15 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 10 inbound Pith citation observations for arXiv:2506.01725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01725 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.860880Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:13:44.235558Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.934274Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a15c7f71-a533-4b6f-90c3-a2dca39bc095 · outbound

This paper cites GPT-4 Technical Report.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:54.001986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:54.001986Z digest=sha256:55b417c5e6a6bfdd37f2bdbdcbdcb32c1a30ff5c7d3a83fda3411fdc7dafd2c5

Observation be9052c9-8e93-4861-a699-4f87d8bb6797 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Activitynet: A large-scale video benchmark for human activity understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:08.019749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:08.019749Z digest=sha256:ad8fe5ee1259eb96afec18ac37896cbe87bd1a10ca3ea5ac3703fb84a19bc9ad

Observation b9c0c257-e40d-4a50-9b57-4793d2ddcc92 · outbound

This paper cites Auroracap: Efficient, performant video detailed captioning and a new benchmark.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Auroracap: Efficient, performant video detailed captioning and a new benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.563025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:08.047640Z digest=sha256:bbac2ca3e79c71bfd5572ccdfbcda1441382272c9eaf4b3ed920e98db9e022b0

Observation 87d10cb7-bbce-4570-894c-48ea936d37e7 · outbound

This paper cites M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge dis- tillation.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge dis- tillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.643118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.643118Z digest=sha256:b54a0c599b7fc40cad4f61e506793cd3fadeb4691d9d2a7d437047d48257967f

Observation 2ef9d158-91d2-4043-b920-430cefab504d · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.778162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.778162Z digest=sha256:632697534faebe4dfd67aa55eb98863e36404255a22efc4ecdeb004d9c2c96d7

Observation 03135c81-6c8b-4354-b373-5950dd5531e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.868692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.868692Z digest=sha256:5c454d258a407bfc8fdf9b3fa50d299463013069e9ae102576f4183563f4ebe0

Observation eb22eed1-3201-432c-9962-4d3ab340b08f · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.908851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.908851Z digest=sha256:af90ea80475ec9f46ab0732bc96bd3ced1e721a67b917d26b0be4ea7fca3eb45

Observation 41784efd-54cd-4147-bb3a-2519b8b8319b · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.939333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.939333Z digest=sha256:271ab6a4f9c8880dfc39f0472a5447382310f70db01ffffdc61eed61790a3c18

Observation 83d9c79d-4ff8-4bf6-a9f4-4936233864b4 · outbound

This paper cites Tall: Temporal activity localization via language query.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tall: Temporal activity localization via language query

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.301041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:10.032618Z digest=sha256:bac969553356d1f250fe7e3e90fc8cf3cf18e2b2d65b9fe0e8087096faa7d67e

Observation 02ca2843-d1a7-4209-ba52-dea0399d67e4 · outbound

This paper cites Scaling laws for reward model overoptimization.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Scaling laws for reward model overoptimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.223294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:11.502809Z digest=sha256:bb1701a3ab9811b05c3112d0a838a19ec2dbc9ecaa0f9e0a3ce3a8d17ce093a8

Observation 7c245ae2-6d72-4855-8bec-7c8f7ad09bc5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.701653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.701653Z digest=sha256:31028368f0e4d4eed3982ed3bcef92fcc9ddbf5f9d4fb8a33044a3f7856e88c5

Observation 5eead944-f033-4430-97d5-61140c8c3ef9 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.807797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.807797Z digest=sha256:230494081e952f677557dac1180ed40047e51bedc01d90d783bdd46fb25a838d

Observation 68037354-34b0-40c7-954e-572ffe521b97 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.883387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.883387Z digest=sha256:e789468bb38660691e0e954cb4627047be38cbebfc1908ee3a0017a3b8b8fc1e

Observation 5a3b0e9d-49b5-48eb-889e-a442eaca222c · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.956199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.956199Z digest=sha256:53497fc4034080cef9984912794e6dca4cf232eb051ac4649ed04d35d46414d3

Observation 903b84b7-4c7b-4e9a-ba12-15ed97a07435 · outbound

This paper cites GPT-4o System Card.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.978948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.978948Z digest=sha256:6f0136b4fd0b693681687acb4b2fcadd8218fb68743fdaa7f09a839057f3ceb9

Observation 9b7bb67e-93e3-4e1f-9da3-d259de55f54a · outbound

This paper cites OpenAI o1 System Card.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.995527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.995527Z digest=sha256:c89225d91af6365babf133923732b2ccc7ce67a4f31f5690927834982bdd09bc

Observation 8c2b9347-9f47-422d-8df5-d5a7eaa3ad00 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.012820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.012820Z digest=sha256:6a03058b0023405057de36bd77643a7076ad87bda565b54c23cb8c24674917ce

Observation ff83fa10-ffe1-48b6-9b70-6a8e139e6a8c · outbound

This paper cites A shortest augmenting path algorithm for dense and sparse linear assignment problems.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking A shortest augmenting path algorithm for dense and sparse linear assignment problems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.184267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:12.037643Z digest=sha256:fadf1dd9e2aba845f06c768e5c15686c4bc72e89824520a02bd2b708411bef38

Observation 3210f0e7-20a3-4d2f-9319-37e1bc1b2cad · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.061896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.061896Z digest=sha256:2d80b7e1b116b97fc8db88cb7d940be1081e5da044cb563bb3becb01696e5cf0

Observation 64cd84fb-ff43-4cac-9c69-120eafb54e01 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.084921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:12.097371Z digest=sha256:c6c2cf1ee8a6c70d8404e6045a3e9716d3f7623c790c925c98570c1133c0bd65

Observation 00491f7a-f816-44ec-89d2-873dd68e7f75 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.120821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.120821Z digest=sha256:ea3e655f51e93502892c6d378b0065c5184e89c7ab08dcb4f32695c701ce70c5

Observation e54b9f27-0590-4e6d-b7d4-57ac29bbd53b · outbound

This paper cites Let’s verify step by step.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Let’s verify step by step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.167318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.167318Z digest=sha256:5d9b7ecdb6bfbff5e410aa4172e3c3a489b5e19a580a8acaba1c9793931652ec

Observation e962be59-6051-4662-b6ca-b4f3a45d3b9e · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.227193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.227193Z digest=sha256:304f395ed577643809d1dc6514b0cc137f6b26cb2ce6c3a00705c405272075e0

Observation 929c53b7-30e9-4924-bd77-7fa1adf09ad0 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.317088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.317088Z digest=sha256:f0ec760efc4cc719d7bb1e330a133acda802e860f9cab438d2ef6d5e77da2a3c

Observation 463aad97-165e-47ee-95e4-5c492569d533 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.084340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.084340Z digest=sha256:d3311d90cf2df42f826e8831c28982b127e17ee2bc423b12dcff841d06880bb1

Observation c097e315-49c3-4fed-a816-21edd16b6721 · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.900706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:23.398842Z digest=sha256:a79b3777701d7e469719df5b9d19911302b2f8455590f0dc0582eeaa340c873a

Observation 69971c8d-693e-4318-ba6c-b3316bc9db5f · outbound

This paper cites Proximal Policy Optimization Algorithms.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.659852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.659852Z digest=sha256:1bc56c6c3d2f819625323b3530781f7180503102004492138d43ac6de6b6b2ee

Observation f77a72e1-7acc-4f7a-a6fb-89a92ca927ce · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:24.131641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:24.131641Z digest=sha256:a87041016c834082a8849e9559ec550eba86e336bb96a02f9b02b47a961ebb29

Observation 9e13b280-d600-4202-a8e2-37e22e721598 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:31.437137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:31.437137Z digest=sha256:3f087a7331f3e787fc3cf6e3b54ded7274151b821c56e3c0b2750ff42a701f9c

Observation ecbc943e-6897-4c25-9e4d-093006c3b49f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:32.880469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:32.880469Z digest=sha256:3f0fa3566e1af84e43e48eec835ca579b8597688ce121d0f0e4742d24b5b2caa

Observation 077385f6-de87-4489-81d0-4e9c4156c2e3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:34.931386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:34.931386Z digest=sha256:166f891003fb03ce15bb276da78e2e45037df697badfb06a621ee2522cd8a865

Observation ce95458a-a9b4-41d5-8105-9a804e398998 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.727219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.727219Z digest=sha256:a7403ba88ac293d3092483cc47a0a3ba9eb962f5ccba09e2a80196f497f43901

Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.885400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.885400Z digest=sha256:37a08e1e56f8d81c305106dac81453c7612b9050ff2427871f32357875dea673

Observation 835dc834-d0d0-4557-bfba-38afbae77af2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.947370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.947370Z digest=sha256:5f74c19b79d43a7fb6a5a6f5e831d4df629def567afe30a5513ed03060e7e9be

Observation c1befbf9-e216-4ae7-bb19-6a771e49d146 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.023507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.023507Z digest=sha256:c48c3d14f713a6f22ad335a65adc02b6a70ed6f18c63d09016b0af62b0f1d75c

Observation 7be61c09-24ad-45a7-95ff-96a0817528a3 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.060722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.060722Z digest=sha256:254a40f025c774c4b00ed2ea0be1d32c7d8fc59db1b10fdd9c02d1f072a5ecc9

Observation c42b6749-3580-4a62-b6eb-d81860a6ce66 · outbound

This paper cites CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:45.357734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:45.357734Z digest=sha256:4be37791a84f4c646e7d5378dc268487e0e326b74a235b173ffe6abeab0c0c0d

Observation 84cb20a6-1314-4606-881c-7b4d6619932a · outbound

This paper cites Qwen2.5 Technical Report.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:53.234744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:53.234744Z digest=sha256:a7501dbef85bd7baa7329e451285f5dadf03ac597e8aa4a077d03555fe89a58d

Observation e19a02cd-8847-4a5b-b60b-5160165c8c7f · outbound

This paper cites Vript: A video is worth thousands of words.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Vript: A video is worth thousands of words

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.725983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:54.882956Z digest=sha256:d551e34230a79dbf9f67676f46b8e32f6dc8d6d2fcc0f07bca594e9169520e86

Observation cfe4d595-bfd2-4502-834d-9b58e82f605a · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.543904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.543904Z digest=sha256:4f678a836b754916cffe65917f64e40393008593174df6a84a882ed5c27dff87

Observation 11382a79-e212-483d-acee-0e4633d8e18a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.873734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.873734Z digest=sha256:5ef0b058effcb4bd9387574c175762719ae320fdac4cef7def9adcb567c5ae0a

Observation b38cb08a-7b3d-4787-9fca-3455a433d275 · outbound

This paper cites Modeling context in referring expressions.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Modeling context in referring expressions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.915232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.915232Z digest=sha256:e4111a5ba35497f1d5b97331343c30aca7b60e1a3e75cbde4bc248f8c8980c0f

Observation 71e9c9c5-ee26-4ec7-af1a-7aa1e7509853 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.923555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.923555Z digest=sha256:e5cdf417efe5f3f6f71ad0eec451c6e06acbd219b7f5afbf80d258969fd45c85

Observation 02b25c7c-7ad1-44a4-891f-fc297f035177 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.644904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:55.953394Z digest=sha256:655864885e00ae700bef3d21e78effedc1bfeb703fc9da1bd3659f8b30dd870a

Observation b58ff38f-0a06-422c-942e-7381f7305b9d · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:56.019147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:56.019147Z digest=sha256:b464b12582e4e0ed183cc5110525e3e54e133e475ba2aa14e4a296f653d8c1c8

Observation f26bff47-1afe-46bf-aed8-1efb8589414e · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Improve Vision Language Model Chain-of-thought Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.064569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.064569Z digest=sha256:1b7b7b6c21bb9eab79eb7d5d40993db120e49032dfad859c93719097a3c72a9d

Observation 19ab9372-e3a4-457d-8ffd-1ef64bb7114f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.687621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.687621Z digest=sha256:841d73d339095398ccf72324b88827bd06ef62ed6171ccb666d9ae725c43e204

Observation bfc7a347-b255-43b6-97df-b3e9738fef0d · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.733558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.733558Z digest=sha256:eefdfa3b18f3acf6805fea69d77ba3956fef4d42cdb5f25b681056230a69c7c7

Observation 568a8494-8f7a-48f4-b86c-3834ef178039 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.775507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.775507Z digest=sha256:63a9d5976725c7bca63ce44381a71977c28eae5c5326f9801d1775967c042f6f

Observation c016390e-52e4-4925-aa35-151145ec6a97 · outbound

This paper cites SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.810477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.810477Z digest=sha256:06017c4380b1da1680837fb27de208d203e9f522d2874cf81c5aaa57c6b5a793

Observation 48094e61-f3c6-4174-be5d-b770a4d99ebe · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.860880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.860880Z digest=sha256:dec171e2a202ab41ef96bc03207f1b97d55018591a274c621e9a1ca42d4a2851

Observation eb5ca7e3-48c9-404b-97c0-cf111daead75 · outbound

This paper cites an unresolved cited work.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:40:59.416096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:40:09.314893Z digest=sha256:7148115cbb0d104a4b346437d2ca3b950dd4f4bfcfa521f24fb3f09d6c3548a4

Pith citing papers

Observation e5b183dd-b5a6-4e41-9796-1db3ffebe5ec · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:11.589395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:11.589395Z digest=sha256:eba7e448b45876c40eec520d624f5a3630e7e4254568986065ee77c88bf2cc20

Observation 4ff98ddf-19d9-4d22-a939-46b89e999320 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.561845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:453027fa9fb7fc84995f128ab363f025f84922df5b1017d974ff3ee0aefc7328

Observation 36e73833-b17f-49e4-9a20-48e792f14e45 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.484250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:5a22e6c4467e6621244b0e8055590500945ef4df4e1d7d26231cb6a897940219

Observation ef22b1aa-0882-4abb-b11b-72a052fbcb76 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.493845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:d32fb94e6c6c4c9a2b6a6d7e9baf217e4feb04df115848baf5d19c8a9fc5b5c9

Observation c8439cfd-3612-4e2a-a1f4-fcd335241947 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.755390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:fa09f6c5e80aeb0affb6feebada98d72b69f4d2a1e56cb94b202f856980972d3

Observation b6a4f68f-becd-49ee-a3bc-6334f5766267 · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.935800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:7fad03a4b9d3fd3d9dac7ad359859c3eb50c692d82953fac697211444b82d896

Observation 9066af67-d826-4a5c-b0d1-5ef885e470a5 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.942635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:b89307f12a66c9aaa1a21dd7434a6e0138abba7ca5c4b3d76ee4b4efef3de858

Observation b32ac67e-4cd7-421b-a1e9-b83d10d18411 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:22.196718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:22.196718Z digest=sha256:9dc818f5d2049f6b9bdd3d7c3a8789b4eda6d36c78d6a4027bdba625f3705168

Observation 32461e25-5782-49b6-a010-1cc41b265953 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.207694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.207694Z digest=sha256:a5cf295f0ddbb48b5ebe28bee226e0a21434e61d9fb01202b0346de661ef1c97

Observation 53fcffc3-aa5f-49b5-ac92-20dd6ee28176 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:13:44.235558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:13:44.235558Z digest=sha256:19cb4f8d2659703df2027a0d5df8234aa19a7f7473872aa5d98f85840702f094