Pith. sign in

Paper Citation Record · LEDGER

Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2408.15542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15542 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:57.694251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:29:15.346488Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6afa325e-c4aa-4373-887a-6e3feca6923d · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.200813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:d08bde6c1d3a07cbca960b35aba3d53aaf24a8eb170b46010c896e551be268c3

Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:e6d9f2ac691d0f62a01ab4aa4cbd021df3974de85ac582877406fdf832963b6d

Observation 4c3f7916-40be-4917-86d3-699740523c97 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.355335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:a2f2ff38902e0f6e544a48b581613b24ab4b4ee7046c97bcec2639c7bea6ae5a

Observation a75c095b-8f2e-4bbb-815c-15164036c109 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:de2672438097b71b0abf1041555d8c7c310a9a45a5294dfbce954ffef2d321fc

Observation 7b31e901-5b45-4c98-bc84-9819d5b33790 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.408504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:08f6ee2bb2286867cf4d5c6e90ee5f0d9a35f5e8823b351d486933662952e0ff

Observation 3cd6b174-5f79-4bb2-b775-89284f17b9df · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:23:51.734057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:d96bd82d45dc91a6e3c2f3eb055922b25bae030223071dc42bb0d830fb83c75e

Observation 8335557e-cb75-467b-b018-6d8712a4a5b0 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.694251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.694251Z digest=sha256:84756d38f7115f1fa94ea7a8ec1e949ba4669388293f4112bb1bd64ffdf02616

Observation 81ccf352-fc30-40ad-a176-ed4d766704f4 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.733144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.733144Z digest=sha256:5ea5dd29b0021c294aa7e7984539ce7ac690be4eef516fc4d3617a550824d87e

Observation 06c314bb-4e5b-4358-8610-210596f6fd2f · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.765688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:01.765688Z digest=sha256:aa85059aa3610e784de9ec660382ba59dd48d629894c946a9429ed298e1dd217

Observation a3fcd63b-ceaa-46f9-bd7f-2fcd422481f8 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.535106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.535106Z digest=sha256:d88eefb89f235f61b062df09b8979ce096fa88c904936b4ccf764d1fb3a5baf7

Observation 6443ad05-adcd-41d2-b2e6-1dc1fa460b2e · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:13.508706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:13.508706Z digest=sha256:0fd8169d13170b9b11659c987f5d1f6768daabf382677cce35de21a99dd45af8

Observation 45e15b40-6cc5-4825-8b1a-1dea06352fcd · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.645844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.645844Z digest=sha256:3a7616e63aa73fa8133d728ad98ff02ef06a865c8890553d8d0fbc5effe2ca10

Observation de8b4712-1d7a-4f71-9dcf-3612323dec24 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:27.753715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:27.753715Z digest=sha256:c60c9d02b39375678705fa3458644802af29bff8c26b481c91623ecc4259ffa6

Observation 82c0dd4b-7efa-417e-8b24-09987e4f0f69 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:03.120334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:03.120334Z digest=sha256:212740a55d53fd7ebb691120a57dc40c69d3eddce481080f56e24d99a629549b

Observation c38d776f-7f12-467b-b91e-4bc3b14c4a9a · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.480720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:2ba94b5515dba8c3ce6420044b2098f628e5434a121e57528e1d2fb9c91725cc

Observation c7ba70d3-efd9-41c4-bb52-345f72de1a67 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.174675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.174675Z digest=sha256:b135e5e709fef7aba66033c15c0f55252106d1abc47e6cd372daccc041d846de

Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.373080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.373080Z digest=sha256:f71c2bd52633e48fc90b3ad8945e650a44259668f64c0d52faec8c67bcf9bdfb

Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.351183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.351183Z digest=sha256:4a437cf645a20083cb68b3c4243890da0f4409b546a1972aad01d0b23f9a5b2a

Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.338028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.338028Z digest=sha256:4a173455a5e016e520853f94516aac398a06e007a8761ac5e6e3ecde5520d37a

Observation 79cdbcc8-6c07-48b9-9e72-e300b01626e0 · inbound

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics cites this paper.

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:32.407540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:05:32.407540Z digest=sha256:9d9ab843d92d3a532320e4b4a3f55def535b716c495bdf30979991894dc0e852

Observation adce1126-a8a0-4236-b4ca-617e4f826596 · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.438516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.438516Z digest=sha256:e42adbe1fc4088c46b99b6ebb21523f8e016666de2f9c202d0f0ec2e6d69f189

Observation 83510f07-e3d2-4591-9f11-3089933d292e · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.769169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.769169Z digest=sha256:23504f24180115a5f164684d26b6bdd2c651cc950b78b8f0a4e5da413ca55b01

Observation 9faf2417-4896-46ee-9883-69e0bdcc3e56 · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.988625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.988625Z digest=sha256:527b069523c021bba0d42b7ca00d8b6ad3e81e444104ea4ff8e73ffc30cd284c

Observation b5c68ae3-7955-4fa6-b102-f464b8aa660a · inbound

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos cites this paper.

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:51:28.132505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T02:49:12.987772Z digest=sha256:d264ff8bba73eae6e25d1ba607d74ed3b1d0b5aeb75f466dc3d3eb76f4602abe

Observation 8313bc12-2f58-443d-87b3-565b183979ff · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.873910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:b79845b510c290be8379be032d746d386c0a696134afbcd89574db8a2e1b0b14

Observation d1c7aabf-cc5c-4e9e-adc7-4a15863bba87 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:15:51.479344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:bf984d626254e8070c299354c9fcc681acf2da10980b61230ed15e33fae9cc14

Observation 82579acb-6fab-499a-978b-dcefacfa5f2e · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.569125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:dc8f9b8d58e73dfd6888af47f3f457278573a686830e950e3260fdcff6fd626e

Observation 3e04ced6-a4e9-4ed8-b2f9-66302f455250 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.759603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:26388d6eb34daeaae29afac18ac1d40423ecf65fcce4c8447e4918f9a28b0dce

Observation 92803c71-da0a-4e87-a538-1536a6e706fd · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.522561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:7dbeb344126a40e945c69dccd19a6b5c51770de22ce499fab7d08f077f7a80a6

Observation c8418107-8250-4579-88e0-92a470464e44 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.479634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:8571d4c9580faa6594064c2ae5de2f8a8f2c53a19d63db433b843020dedc016a

Observation c1005e51-57e9-4c04-9b1d-862c7d6bfbde · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.700487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:51dbac21c647df3a393fc8b31f52007331c1cbec31ff4b71ed330b83273f9e79

Observation 88c56c0f-3dab-484a-878e-e22f4a3402a5 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.452976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:94c471e15641e5ab2ec38acec7ada853ae14950c6947e9c6a26edf459766d7a3

Observation 06199d9c-9c90-4ba0-bca6-6dfab8cdebb0 · inbound

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering cites this paper.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.071338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:9c352e9dbf2bbb83a031ff12945951c28471370c06e9db58ecbe89e82fc97839

Observation 4f5b8170-f9d8-43a5-82cc-fa3598a11b04 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 296

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.192666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:b0fe66a54a8319b6958ed0dd7d7ec4af8394f39a33b0c0931d1b1f7240cc6e22

Observation ed14d6ca-5e62-4cb0-9eed-d7f32f883594 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.726244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:d9e0e392d10d22204fa661b6196418be21f33d9bf67129f163b2e5c1de429bd3

Observation a605ff3f-0bea-437f-a1a5-671a2adbaa57 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:29:15.349201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:14:11.297384Z digest=sha256:60d4f3b5becd9f2bcffe2ca4336b1649321f37f81d0c34b7e2edc68a071bc2cd

Observation 71114075-e508-49b7-8ba5-8ebae0751e93 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:57:19.810078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:57:19.810078Z digest=sha256:caf2fe7b29bfd2fecc6b6efb628bb3b33ae2fee173e988d5f6b8023cbe9594d0

Observation 01d453a6-8134-402f-ba94-496cc1c508df · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:9a43bc18773e4a0e4f6e49728c90f0da15e5ba318a096860b4ae99e1fba51d59

Observation 30cccac3-b3ac-40f9-adef-3c640cb6616b · inbound

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization cites this paper.

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:13.271024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:13.271024Z digest=sha256:b5e1eca2fea7791c53273da47cabb2d537060ef9daf195d00a5faa42d7bc304a