Pith. sign in

Paper Citation Record · LEDGER

TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2504.09641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.09641 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:18.656306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:31.402219Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6903ccf-ac39-4c50-b14e-b931bf6a40e2 · inbound

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding cites this paper.

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.542292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:40:06.454859Z digest=sha256:0376c7dd806c08111c85ddf993057b0e02f02b1dc91952b076ea9f04f1414ed4

Observation ece43dd0-ab59-49e3-80e4-0ba2d1f4d93f · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:18.656306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:18.656306Z digest=sha256:0b485482b84a777c3bab7102eebc7067eb40152619fda341fea26c7749371fdc

Observation 6c710c76-cef8-41fb-ad94-db1d04df5e4e · inbound

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought cites this paper.

Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:15.937040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:15.937040Z digest=sha256:1b6dba9168134e862042e277b7915e25cdcf1a850f126d6bb100cf3c52774b7c

Observation 3203da8e-6f23-4b56-81fa-9d2d711bda85 · inbound

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding cites this paper.

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:17:18.563294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:13:40.485342Z digest=sha256:e1843a34007408ea2e26e680f5c38e257a1c56eb38e4b1b43f2146b4b4dd5ade

Observation ff5ef0de-8880-467b-b83f-1b11574f8eaf · inbound

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? cites this paper.

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:40:56.049230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:40:55.944288Z digest=sha256:3fdb19dc034db4a818002cb916323c480ac76387e6b38ff4174ba516d85b41d5

Observation 4afaad58-44d7-4e64-8e01-b8477d86472e · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:05.300590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:05.300590Z digest=sha256:0392ade69c3d0e751c36af716c7bdc1397e675d8e7f3388a4d1e68f6bf507830

Observation 9024fe1b-a3c3-4883-a50b-a5445afee8c8 · inbound

SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization cites this paper.

SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:28.598898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:28.598898Z digest=sha256:76a68be45bb0941d9507a82cc80aadfb6b08c80dea2b04e24086bddcb2ea5230

Observation 6b93e37e-a4f1-4806-a581-13cf263521b9 · inbound

CyberV: Cybernetics for Test-time Scaling in Video Understanding cites this paper.

CyberV: Cybernetics for Test-time Scaling in Video Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:47.309529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:47.309529Z digest=sha256:adaebad6bbf11c0164879bfaae46d13b25524a80ea8d98879e0613d52bbe6cf1

Observation 3aa44a55-2906-47d5-8c55-b4bd0248f676 · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:52.150407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:52.150407Z digest=sha256:ead5fedb5fc95f3a9629f8a69af38f68cdc0a932909acb0443d1aa936bea57fb

Observation 505ef48c-82e0-40ed-9fc6-f2d6663dd08c · inbound

OctoNav: Towards Generalist Embodied Navigation cites this paper.

OctoNav: Towards Generalist Embodied Navigation TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:43.895513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:43.895513Z digest=sha256:0b6ec4305a1b464a4b35045c6b378391d25c90d43a025e150c1d6ed7dd7128f3

Observation e28a9912-3e8c-4a5e-9e30-659ee8039dc9 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.007028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:89df0afe53cd56a4b8b3b21dc03898a2246e637c9463cc963d59778585f58dce

Observation 7b24de6e-9992-4c1e-8ff4-c7fef8c2dfda · inbound

Compressed Models are NOT Trust-equivalent to Their Large Counterparts cites this paper.

Compressed Models are NOT Trust-equivalent to Their Large Counterparts TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:00:30.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:00:30.163389Z digest=sha256:e5d5ad69af5c30a4eba15f3b1858d6fe5a1623b11dd20d25fd836a98fd9456dc

Observation 4ee90811-fdf6-4cea-836f-44f2efa11dba · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:47.267163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:47.267163Z digest=sha256:b7941ecbff6583c0f344b3ed2911fd2158e99822c53c656c1e59f93f616de8b3

Observation 4b2b738a-4799-4f6a-8943-d902c017ab8b · inbound

Video Reasoning without Training cites this paper.

Video Reasoning without Training TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:09.855828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:12:09.855828Z digest=sha256:eaafc5c522b214b0e7656f7c891892c39f6f1c0d74626e08a1c6bd771baad021

Observation f503adc7-1cf5-4832-9b23-ae3118e44777 · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.954750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.954750Z digest=sha256:245ff3cd811958c1f2545d8daae8f54600919a4b7d807c48d31a3615efe48490

Observation 17f3cfb5-ae42-482d-a31d-a47b4f0a720d · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.544726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:dfe5b465b5dedfdc8a92df9c04bc7d6d6b8dd2d8a7ac93afa9aba44553e22638

Observation 7e940b54-6378-4f98-b9df-28dc33c57448 · inbound

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning cites this paper.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.321509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:19343986bf4e5ac9d3a98b6caabaebf2815007b1a53b5eff124f5ffe211f0085

Observation d0ab1da0-ad2d-4dc3-bbac-69371a37e1cb · inbound

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning cites this paper.

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:52.581348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:40:41.642852Z digest=sha256:fc879885c0a48f7e1a0a8f97a2893f25a17dd8f7e5aeeac133e35ab461d15591

Observation 6705c224-50d7-4c86-bc5c-852e24f6ea36 · inbound

Video-ToC: Video Tree-of-Cue Reasoning cites this paper.

Video-ToC: Video Tree-of-Cue Reasoning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:09.514057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T01:20:29.374012Z digest=sha256:e0c2efe1e5db84046b600af3ba42d86733ce22662c8e8d4f09677cf57be1f979

Observation b3520bef-61bc-4d47-8a43-dc30f82f04d7 · inbound

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs cites this paper.

Beyond Perceptual Shortcuts: Causal-Inspired Debiasing Optimization for Generalizable Video Reasoning in Lightweight MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:06.511080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:30:56.297653Z digest=sha256:a6a5a13ebded315b3f640f77e2566e40d0fc0224d81854ca788e51d0e99d0713

Observation 542ce35d-b672-4081-af8e-a288041fe7da · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.112134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:89c51cab46ec9a33ee283e0ec0bf01017260f8a22b936878efcf20b512ac1753

Observation b2b4267c-3ccb-4aab-aad3-7ef8381ec0d5 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.555385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:5566bb4ed00c04661627a3e2721dbe6dae583c63c94c276d61e4229eb501fa01

Observation b2cf85bd-4380-42a8-9f96-e2c95434cb17 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.741199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:dbfc62b619564376b6441f333d9b920b1261835839f500e3e7f47e7e487e03a9

Observation b5854cbb-01ac-42e2-9818-3d7e90e1c590 · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.404380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:ae97497c7a7d8f631a0702d254694c2dfa2b5e5968c4c760e2f568b965d2eec5