Pith. sign in

Paper Citation Record · LEDGER

TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2504.09641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.09641 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:47.309529Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:31.402219Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6903ccf-ac39-4c50-b14e-b931bf6a40e2 · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.542292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:0026ca7967eadf29b66251d029aad698cc4f2c49235f44d3d84e82ba04f2d236

Observation 3203da8e-6f23-4b56-81fa-9d2d711bda85 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.563294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:a7651bbadb1e5b036fdd05218c847701424f93bdb1ab6246e5caaa631a0ad371

Observation ff5ef0de-8880-467b-b83f-1b11574f8eaf · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.049230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:2a94550027b9c67e8feb9c433cf72ec37345a88b4985993c460032e04d6052a9

Observation 6b93e37e-a4f1-4806-a581-13cf263521b9 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.309529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.309529Z digest=sha256:93b6f29bdb269552ae26aa55f9447f22d9b7481db5f5c84cc2632da690413bf6

Observation 3aa44a55-2906-47d5-8c55-b4bd0248f676 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.150407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.150407Z digest=sha256:d387caafa0f6640279bd28ebb1edceca27d9a5e04895577bbc2937b7a84f932b

Observation 505ef48c-82e0-40ed-9fc6-f2d6663dd08c · inbound

OctoNav: Towards Generalist Embodied Navigation cites this paper.

OctoNav: Towards Generalist Embodied Navigation TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:43.895513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:43.895513Z digest=sha256:d5666ab19cfde8f781f53a73cd216c165cfb5f1b06ef1d406767042e31ddeacc

Observation e28a9912-3e8c-4a5e-9e30-659ee8039dc9 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.007028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:de7bb90b6b0a1d2df76278636eb05fa8d8d92e98ecd4952643e8c9fee68ab1d1

Observation 7b24de6e-9992-4c1e-8ff4-c7fef8c2dfda · inbound

Compressed Models are NOT Trust-equivalent to Their Large Counterparts cites this paper.

Compressed Models are NOT Trust-equivalent to Their Large Counterparts TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:00:30.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:00:30.163389Z digest=sha256:5bbddefaaa9a711db3d4e177addcf6036fd5b2b00b17518c4b8dc1d65d642236

Observation 4ee90811-fdf6-4cea-836f-44f2efa11dba · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:47.267163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:47.267163Z digest=sha256:c53c470871712a5e73587de1da254a9c75f788998b1079d584d12497177626e3

Observation 4b2b738a-4799-4f6a-8943-d902c017ab8b · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.855828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.855828Z digest=sha256:a416b7d728c9d188b057aed871cd9b3cd31789fc99439d2707cdb2ded6012cd0

Observation f503adc7-1cf5-4832-9b23-ae3118e44777 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.954750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.954750Z digest=sha256:eb7406aec37e2bb890e57bf6b5bc5e7a0cfeed5a2eb80c5738ca5573da8c2aeb

Observation 17f3cfb5-ae42-482d-a31d-a47b4f0a720d · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.544726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:89d13dae066e411566addd3cc3686ecdbf66a1511abbc89cb2001ff38a9256d4

Observation 7e940b54-6378-4f98-b9df-28dc33c57448 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.321509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6838ab2a6b2cdbf998a4f96fc0c1b099339c474860e220a443f2df0b905dd587

Observation d0ab1da0-ad2d-4dc3-bbac-69371a37e1cb · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.581348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:6d2ea7513ac891bb52da9bc9258c40f1dbde461c6b28a8b9eb287ea5f9f58c92

Observation 6705c224-50d7-4c86-bc5c-852e24f6ea36 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.514057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:478d02b3e53a2429669635eb54d7c5adefd54b033453a28999f40b1d1378c22c

Observation b3520bef-61bc-4d47-8a43-dc30f82f04d7 · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:06.511080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:b51bd5bfc27759521da1b4f55069d4dd51d4687596a16bfcd9c0a264075790a1

Observation 542ce35d-b672-4081-af8e-a288041fe7da · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.112134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:4ce48a9ba83e1b52f998b6e59c43ee5c9cca4a38341863288b8282711e710f67

Observation b2b4267c-3ccb-4aab-aad3-7ef8381ec0d5 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.555385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:56e9b7ce63f5206fef4e325e07fcffe26f0f52191bf0b1f75b60583febe95902

Observation b2cf85bd-4380-42a8-9f96-e2c95434cb17 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.741199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:4c8f4833d1127d7f7adce44f476dda3c8a846b9cca779b43358e956bcf72480f

Observation b5854cbb-01ac-42e2-9818-3d7e90e1c590 · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.404380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:5fcc85e3b1bf8c59689f281677295b4c5eafde9de7f142b9561ea4812dadb596