Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2408.15542.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:57.694251Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:29:15.346488Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6afa325e-c4aa-4373-887a-6e3feca6923d · inbound
LVBench: An Extreme Long Video Understanding Benchmark Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · inbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c3f7916-40be-4917-86d3-699740523c97 · inbound
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a75c095b-8f2e-4bbb-815c-15164036c109 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b31e901-5b45-4c98-bc84-9819d5b33790 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3cd6b174-5f79-4bb2-b775-89284f17b9df · inbound
SmolVLM: Redefining small and efficient multimodal models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8335557e-cb75-467b-b018-6d8712a4a5b0 · inbound
ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81ccf352-fc30-40ad-a176-ed4d766704f4 · inbound
Clapper: Compact Learning and Video Representation in VLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06c314bb-4e5b-4358-8610-210596f6fd2f · inbound
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3fcd63b-ceaa-46f9-bd7f-2fcd422481f8 · inbound
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6443ad05-adcd-41d2-b2e6-1dc1fa460b2e · inbound
Reinforcing Video Reasoning with Focused Thinking Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e15b40-6cc5-4825-8b1a-1dea06352fcd · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8b4712-1d7a-4f71-9dcf-3612323dec24 · inbound
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c0dd4b-7efa-417e-8b24-09987e4f0f69 · inbound
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c38d776f-7f12-467b-b91e-4bc3b14c4a9a · inbound
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c7ba70d3-efd9-41c4-bb52-345f72de1a67 · inbound
Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · inbound
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · inbound
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · inbound
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79cdbcc8-6c07-48b9-9e72-e300b01626e0 · inbound
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adce1126-a8a0-4236-b4ca-617e4f826596 · inbound
Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83510f07-e3d2-4591-9f11-3089933d292e · inbound
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9faf2417-4896-46ee-9883-69e0bdcc3e56 · inbound
CAViAR: Critic-Augmented Video Agentic Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c68ae3-7955-4fa6-b102-f464b8aa660a · inbound
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8313bc12-2f58-443d-87b3-565b183979ff · inbound
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1c7aabf-cc5c-4e9e-adc7-4a15863bba87 · inbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82579acb-6fab-499a-978b-dcefacfa5f2e · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e04ced6-a4e9-4ed8-b2f9-66302f455250 · inbound
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92803c71-da0a-4e87-a538-1536a6e706fd · inbound
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8418107-8250-4579-88e0-92a470464e44 · inbound
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1005e51-57e9-4c04-9b1d-862c7d6bfbde · inbound
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88c56c0f-3dab-484a-878e-e22f4a3402a5 · inbound
An Efficient Streaming Video Understanding Framework with Agentic Control Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06199d9c-9c90-4ba0-bca6-6dfab8cdebb0 · inbound
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f5b8170-f9d8-43a5-82cc-fa3598a11b04 · inbound
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 296
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed14d6ca-5e62-4cb0-9eed-d7f32f883594 · inbound
Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a605ff3f-0bea-437f-a1a5-671a2adbaa57 · inbound
Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71114075-e508-49b7-8ba5-8ebae0751e93 · inbound
Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d453a6-8134-402f-ba94-496cc1c508df · inbound
Latent Visual Cache for Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cccac3-b3ac-40f9-adef-3c640cb6616b · inbound
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.