Pith. sign in

Paper Citation Record · LEDGER

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 9 inbound Pith citation observations for arXiv:2506.01725.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01725 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:57.860880Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:29:11.589395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:40:00.934274Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a15c7f71-a533-4b6f-90c3-a2dca39bc095 · outbound

This paper cites GPT-4 Technical Report.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:54.001986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:54.001986Z digest=sha256:a91ce976e990a299083f54cd966b2656f0a886690205886f3e5d9064b6095db7

Observation be9052c9-8e93-4861-a699-4f87d8bb6797 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Activitynet: A large-scale video benchmark for human activity understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:08.019749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:08.019749Z digest=sha256:c1c7c4e9d52e9a828a4c627d2d62dba4fe42501fec2ab362152c4711ec4bbbde

Observation b9c0c257-e40d-4a50-9b57-4793d2ddcc92 · outbound

This paper cites Auroracap: Efficient, performant video detailed captioning and a new benchmark.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Auroracap: Efficient, performant video detailed captioning and a new benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.563025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:08.047640Z digest=sha256:5ea0ef235107615e02b4006ab74452ae6b4a67a1a1db9594d55c099243ef46f5

Observation 87d10cb7-bbce-4570-894c-48ea936d37e7 · outbound

This paper cites M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge dis- tillation.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge dis- tillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.643118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.643118Z digest=sha256:9dcd14a0252a03301b7b02f5e31755014628d543c9ff1541aea8b91dbb91d342

Observation 2ef9d158-91d2-4043-b920-430cefab504d · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.778162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.778162Z digest=sha256:6a9fd046bb02fbf82b5d4bddc49da325ec79f213499534d2b478feb8896cebc4

Observation 03135c81-6c8b-4354-b373-5950dd5531e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.868692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.868692Z digest=sha256:4c710e91aa281c7e20de14b4d42e407d1bbc25df29e6f1269fe158ff7d026107

Observation eb22eed1-3201-432c-9962-4d3ab340b08f · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.908851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.908851Z digest=sha256:d5050ad8ab6a0c9b5b914163135ab35b53b7ef54144971345c630e623b861042

Observation 41784efd-54cd-4147-bb3a-2519b8b8319b · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.939333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.939333Z digest=sha256:032b0a182a275a63361b48dd5b17e66cf5101b240558494a890f359f641a3745

Observation 83d9c79d-4ff8-4bf6-a9f4-4936233864b4 · outbound

This paper cites Tall: Temporal activity localization via language query.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tall: Temporal activity localization via language query

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.301041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:10.032618Z digest=sha256:1388a6c8827662ec811c86c9614891d5d923ab0b3cc8153e037fa66ec9bcaa96

Observation 02ca2843-d1a7-4209-ba52-dea0399d67e4 · outbound

This paper cites Scaling laws for reward model overoptimization.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Scaling laws for reward model overoptimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.223294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:11.502809Z digest=sha256:15087351796c8f8fbbe97098fe2740a846ab1633196a40e1e573fb9b3843dbd7

Observation 7c245ae2-6d72-4855-8bec-7c8f7ad09bc5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.701653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.701653Z digest=sha256:0871ebb2d5ea199386b07091ef9a500bdee561b44779d2c9f31dbb466f561399

Observation 5eead944-f033-4430-97d5-61140c8c3ef9 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.807797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.807797Z digest=sha256:cc417bcf5b5e1212f63047457c769fb9cfdc42f7a729e16120ab83f3aea81e32

Observation 68037354-34b0-40c7-954e-572ffe521b97 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.883387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.883387Z digest=sha256:4e0285d180529079b0c6927f61688394963847391d39fb9be0456bef098497a0

Observation 5a3b0e9d-49b5-48eb-889e-a442eaca222c · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.956199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.956199Z digest=sha256:40a650b1774f84601e9468a3bb4df93b4d35dec77c6b6cbe77aaa308494fc7ec

Observation 903b84b7-4c7b-4e9a-ba12-15ed97a07435 · outbound

This paper cites GPT-4o System Card.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.978948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.978948Z digest=sha256:799e5e6ed61e1d39dad45bf180b1ddcfbe0d2da398f86760bff606d7fa466325

Observation 9b7bb67e-93e3-4e1f-9da3-d259de55f54a · outbound

This paper cites OpenAI o1 System Card.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.995527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.995527Z digest=sha256:ff163893a217bf23ee7c6fc18b735a476ee3a13ac35ca915d39f985070cb8b90

Observation 8c2b9347-9f47-422d-8df5-d5a7eaa3ad00 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.012820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.012820Z digest=sha256:9916149ecc46ea75a6484889ce59c1c08d704d93c0a3e4bcf1c55c470332460a

Observation ff83fa10-ffe1-48b6-9b70-6a8e139e6a8c · outbound

This paper cites A shortest augmenting path algorithm for dense and sparse linear assignment problems.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking A shortest augmenting path algorithm for dense and sparse linear assignment problems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.184267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:12.037643Z digest=sha256:9c0d0062abaa9986a54f6151009a901a4344f9fad40794cf277b1a5439fe8c92

Observation 3210f0e7-20a3-4d2f-9319-37e1bc1b2cad · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.061896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.061896Z digest=sha256:3ce3dd0c71be198612cc6fc9b0158ae9776dfb5cddb2d2a285443c39cb3df18f

Observation 64cd84fb-ff43-4cac-9c69-120eafb54e01 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:59.084921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:12.097371Z digest=sha256:c9157bf6190a245589481d8bd4bb9119062282002e675ed14f7487f100f087c9

Observation 00491f7a-f816-44ec-89d2-873dd68e7f75 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.120821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.120821Z digest=sha256:51185ae248ca2a221f45d6c9d56468e6e0e07209c09992f1dbb150200594179d

Observation e54b9f27-0590-4e6d-b7d4-57ac29bbd53b · outbound

This paper cites Let’s verify step by step.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Let’s verify step by step

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.167318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.167318Z digest=sha256:936212c2f050bb2dc800963a9dffe75bbe9cdb4e71df3525fa51deb9a60959f2

Observation e962be59-6051-4662-b6ca-b4f3a45d3b9e · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.227193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.227193Z digest=sha256:e312b26c9c3acba92d6cbb034c90c8e0b5804c902f56a3c1baf3399219d208b5

Observation 929c53b7-30e9-4924-bd77-7fa1adf09ad0 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.317088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.317088Z digest=sha256:bbc94dd6a0f7ec09c5ec7668f1cfd3b1008b2ef4f6f573f586a53b5bcc880cc6

Observation 463aad97-165e-47ee-95e4-5c492569d533 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.084340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.084340Z digest=sha256:025895f0e8e6cf6d641c4f132e375db42112d4433444ec7eb8927b073ec36c68

Observation c097e315-49c3-4fed-a816-21edd16b6721 · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era, 2024.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Introducing gemini 2.0: our new ai model for the agentic era, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.900706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:23.398842Z digest=sha256:77b96048a08a032b5eefd83e6cda1d9c819fcaf9b5c0bcdfa03782f1a23a76aa

Observation 69971c8d-693e-4318-ba6c-b3316bc9db5f · outbound

This paper cites Proximal Policy Optimization Algorithms.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Proximal Policy Optimization Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:23.659852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:23.659852Z digest=sha256:3040ba664f4684888cdc2c660fdec0233059cb653e2ced43c5ca49abfed2eabe

Observation f77a72e1-7acc-4f7a-a6fb-89a92ca927ce · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:24.131641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:24.131641Z digest=sha256:6958b93936a1a25068f49eb331225c6e291f4a84dd134dde801a78ae577aaccb

Observation 9e13b280-d600-4202-a8e2-37e22e721598 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:31.437137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:31.437137Z digest=sha256:2d00d1ac8ba8b6e8ea952ce351ebe308f51c2fd1d6ca8ce1f304dc7299c6f087

Observation ecbc943e-6897-4c25-9e4d-093006c3b49f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:32.880469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:32.880469Z digest=sha256:8c3644f5c802cede2b3771dfaeca341301a14a79996c7b8723b4d1008cf74c74

Observation 077385f6-de87-4489-81d0-4e9c4156c2e3 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:34.931386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:34.931386Z digest=sha256:fca7b08a1b590e94299b7bae9172c827fafd24cf1aaf041e047e36fa31b19a5c

Observation ce95458a-a9b4-41d5-8105-9a804e398998 · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.727219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.727219Z digest=sha256:610b6565ffdeccfd5e6848c26c0a9a18809739a5af8664117f0a49c932cc741c

Observation c314e2d9-ad47-490c-a7b0-549e292d25e3 · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.885400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.885400Z digest=sha256:7acfc25c3863c257b27abeb77a5c09406c61af64c9cebe129b2d6ee5590000ce

Observation 835dc834-d0d0-4557-bfba-38afbae77af2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.947370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:35.947370Z digest=sha256:4a1d17b780d60b9112f8e353b30fc32f8d7833cec493d3ec0dad9387eb8ba39d

Observation c1befbf9-e216-4ae7-bb19-6a771e49d146 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.023507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.023507Z digest=sha256:c43a089a78f93597ddd213d2f556d29a2411f86eedbd3ea5a2f53282d1b7b50d

Observation 7be61c09-24ad-45a7-95ff-96a0817528a3 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:36.060722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:36.060722Z digest=sha256:50089ce40d44dc75ce30a259347ab936e4281ddf20f9e58eb52d9f8374679203

Observation c42b6749-3580-4a62-b6eb-d81860a6ce66 · outbound

This paper cites CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:45.357734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:45.357734Z digest=sha256:9735e0c6cfe11fdda18e7491a61f850ec31609f41fc4371713ddb3d5554365e8

Observation 84cb20a6-1314-4606-881c-7b4d6619932a · outbound

This paper cites Qwen2.5 Technical Report.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:53.234744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:53.234744Z digest=sha256:602736759540a84546d187350d0da9fca56256cfa5ad84262d459cd8d364f252

Observation e19a02cd-8847-4a5b-b60b-5160165c8c7f · outbound

This paper cites Vript: A video is worth thousands of words.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Vript: A video is worth thousands of words

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.725983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:54.882956Z digest=sha256:3dd36bb50bed9a0c648456492ab3f7e548b4cc6102004d9f943f9e5f1ab5f562

Observation cfe4d595-bfd2-4502-834d-9b58e82f605a · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.543904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.543904Z digest=sha256:07add2854a04b7c1bda01cdb984c1c0f58c9620828f8dd725f47e30a0c3cd518

Observation 11382a79-e212-483d-acee-0e4633d8e18a · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.873734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.873734Z digest=sha256:a4aef8a45c8a0759eac551b0b210118024c4811944a1a8629d9da5d62d6e823a

Observation b38cb08a-7b3d-4787-9fca-3455a433d275 · outbound

This paper cites Modeling context in referring expressions.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Modeling context in referring expressions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.915232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.915232Z digest=sha256:bf6aeec72355d6b2147e864ade4fb444279c51c8c03123c21aa1f91496a88223

Observation 71e9c9c5-ee26-4ec7-af1a-7aa1e7509853 · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:55.923555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:55.923555Z digest=sha256:af13f95b578b0286bc569a9cd6ae537add8c26cdd40010a1816f4a178236188a

Observation 02b25c7c-7ad1-44a4-891f-fc297f035177 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.644904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:55.953394Z digest=sha256:9fd073f0cec0697179dc2da5a253e3400d2f72daeb172879aed3493d15a88e3a

Observation b58ff38f-0a06-422c-942e-7381f7305b9d · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:56.019147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:56.019147Z digest=sha256:1b6db512217dad9a9885b1eea1ad8eba77f1a63f7aee0150c42964f36b64ced2

Observation f26bff47-1afe-46bf-aed8-1efb8589414e · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Improve Vision Language Model Chain-of-thought Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.064569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.064569Z digest=sha256:e41efa37f8cf5b17491136d1e2d5ba59c9c9fedac9cd18e82d1f2e86f07d5f28

Observation 19ab9372-e3a4-457d-8ffd-1ef64bb7114f · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.687621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.687621Z digest=sha256:7a44ba16756fae67ece561d70d5e6325ec83ae373d5a19b7e1b186d19c38748d

Observation bfc7a347-b255-43b6-97df-b3e9738fef0d · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.733558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.733558Z digest=sha256:d703b65c10f7532df2ba50118330f7c22496f8a8934ce426da75c74890b9f389

Observation 568a8494-8f7a-48f4-b86c-3834ef178039 · outbound

This paper cites MMVU: Measuring Expert-Level Multi-Discipline Video Understanding.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.775507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.775507Z digest=sha256:d4fe8fa2a69fe89024bfc77abaa723f32ac57f0dc51b299aa3087a65293115d4

Observation c016390e-52e4-4925-aa35-151145ec6a97 · outbound

This paper cites SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.810477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.810477Z digest=sha256:13aa00f5b9ca08c21637e9620b725d08f96927a78bf8edae56f00a50135601ab

Observation 48094e61-f3c6-4174-be5d-b770a4d99ebe · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.860880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:57.860880Z digest=sha256:4f73aa37fd7c1a0032b6230eac3579197872d84d5bbface73a61997d28f041dd

Observation eb5ca7e3-48c9-404b-97c0-cf111daead75 · outbound

This paper cites an unresolved cited work.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:40:59.416096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:40:09.314893Z digest=sha256:78145e79f74cfc2cf7b842e5b7ed4dc5866dc37e0e86a51f7bc021173d23f4d0

Pith citing papers

Observation e5b183dd-b5a6-4e41-9796-1db3ffebe5ec · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 292

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:11.589395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:11.589395Z digest=sha256:c1ad76290ce3b7d4b64e7241803c263cdf390593241446bddb59e6fc77fbdd01

Observation 4ff98ddf-19d9-4d22-a939-46b89e999320 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.561845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:afec25710968c07a0c1582ee60b8a544127da8ee3a11f12714e3f5c706946964

Observation 36e73833-b17f-49e4-9a20-48e792f14e45 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.484250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:2ba9b57ec7da3038d2967a405092888052e2b1174b93b0d25547e83efefaf557

Observation ef22b1aa-0882-4abb-b11b-72a052fbcb76 · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.493845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:593f68da6d443a0c58ae782fd49208b881cbdf37966bf3bcd52fc1e72fe4fa24

Observation c8439cfd-3612-4e2a-a1f4-fcd335241947 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.755390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:4e94d2195163d836c0d94bccfbc576089f6f4a59541e7766dd869dcf9f565d0b

Observation b6a4f68f-becd-49ee-a3bc-6334f5766267 · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.935800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:dd2eb7bc7e7effa7dfe1090b45c26a1fb1f14c49f730d91c1a10000636b67858

Observation 9066af67-d826-4a5c-b0d1-5ef885e470a5 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.942635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d5be0d949cf45c5b24d44f223dd7a57257b330fd111dab13447ca5b7a60eaac3

Observation b32ac67e-4cd7-421b-a1e9-b83d10d18411 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:22.196718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:22.196718Z digest=sha256:1fafcd9d3d1f6f7d80d6ef8377509c87be15f975cbf077547ad6c82c23780f53

Observation 32461e25-5782-49b6-a010-1cc41b265953 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.207694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.207694Z digest=sha256:c2fa4c1cec6a613f9343a74585c43d58a810c3363e418721e27eb8d910492b6f