Pith. sign in

Paper Citation Record · LEDGER

Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2408.15542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15542 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:31:16.940927Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:29:15.346488Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6afa325e-c4aa-4373-887a-6e3feca6923d · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.200813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:ba6b66bbd75ad8c8fa433b9def99447eb35af0b97e174f5361900a9a2caf168c

Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · inbound

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling cites this paper.

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:02:43.448546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T04:02:43.261543Z digest=sha256:d11c7e3031600e89e3d9b07f1555e393e919a0a142223f0d60c4f032fd877028

Observation 4c3f7916-40be-4917-86d3-699740523c97 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.355335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:297a5953e53b51ff293a97eee50a11d67554ae92b6ffbe66336b1bedce2c482c

Observation a75c095b-8f2e-4bbb-815c-15164036c109 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.287485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:d8ddf46541d0e3f4a93fc4a0aa82f4ca165ef54e8b26338a6fa20fc4a33f6512

Observation 3ab614dd-274e-44cb-9262-13cb022a3a00 · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.940927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.940927Z digest=sha256:e909057cc80f812d122a63e07478a0800067d4703f3b732e0ac2ae845a05e014

Observation 7b31e901-5b45-4c98-bc84-9819d5b33790 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.408504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:02b0f4f45f36a8ee24b44d5662b9e930c861b5cd5c4e5a223853d54b4917d2b7

Observation 3cd6b174-5f79-4bb2-b775-89284f17b9df · inbound

SmolVLM: Redefining small and efficient multimodal models cites this paper.

SmolVLM: Redefining small and efficient multimodal models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:23:51.734057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:23:50.552549Z digest=sha256:60a06713a56cc52aa39344b518880fe94725b7c8e4dd83d1e61d9c6ffbe6c5a4

Observation 8335557e-cb75-467b-b018-6d8712a4a5b0 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.694251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.694251Z digest=sha256:6e62521ebcc8f2872b30b697d758e50d348350eec660ced53fbfb5b694120e62

Observation 81ccf352-fc30-40ad-a176-ed4d766704f4 · inbound

Clapper: Compact Learning and Video Representation in VLMs cites this paper.

Clapper: Compact Learning and Video Representation in VLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.733144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.733144Z digest=sha256:a73dd1e0f63f733c1905dd874c6fb61f30c8bcfcd6a3ed182ef23e6edcf53ddd

Observation 06c314bb-4e5b-4358-8610-210596f6fd2f · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.765688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:01.765688Z digest=sha256:bde36e2ecc7191a5d21d9a2a0fe6c1a5cf50dfc3f6d38d2b4324363893dbe09c

Observation a3fcd63b-ceaa-46f9-bd7f-2fcd422481f8 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.535106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.535106Z digest=sha256:d88eefb89f235f61b062df09b8979ce096fa88c904936b4ccf764d1fb3a5baf7

Observation 6443ad05-adcd-41d2-b2e6-1dc1fa460b2e · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:13.508706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:13.508706Z digest=sha256:888652dc1893f9d5bdbc3f1c00577eaf14bb475416d651e615bd3dabe6f0eabc

Observation 45e15b40-6cc5-4825-8b1a-1dea06352fcd · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.645844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.645844Z digest=sha256:fb7d9a8c3663fa3e7b0ecfd5fa71f85f75b495a0fd312ed7246692c627ffe314

Observation de8b4712-1d7a-4f71-9dcf-3612323dec24 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:27.753715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:27.753715Z digest=sha256:cb042761a86464d8b7f2b344ab7bb17ae7a8bbeee6d921d0422ef5b2d9570376

Observation 82c0dd4b-7efa-417e-8b24-09987e4f0f69 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:03.120334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:03.120334Z digest=sha256:212740a55d53fd7ebb691120a57dc40c69d3eddce481080f56e24d99a629549b

Observation c38d776f-7f12-467b-b91e-4bc3b14c4a9a · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:22:37.480720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:50a323d578fe626ccbdb9241073b7b3aec1be12762de7c416509620a1197d83f

Observation c7ba70d3-efd9-41c4-bb52-345f72de1a67 · inbound

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency cites this paper.

Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:35:46.174675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:35:46.174675Z digest=sha256:b135e5e709fef7aba66033c15c0f55252106d1abc47e6cd372daccc041d846de

Observation d2d036f9-03e3-4183-a4e4-e3dade6832b8 · inbound

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning cites this paper.

GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.373080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:37.373080Z digest=sha256:f71c2bd52633e48fc90b3ad8945e650a44259668f64c0d52faec8c67bcf9bdfb

Observation 0f33d806-78f9-4ccf-9bd3-34caa259491c · inbound

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs cites this paper.

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:03.351183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:03.351183Z digest=sha256:4a437cf645a20083cb68b3c4243890da0f4409b546a1972aad01d0b23f9a5b2a

Observation 0f4f6773-ee34-41bd-bd60-e24713ee2a90 · inbound

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams cites this paper.

Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:05.338028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:05.338028Z digest=sha256:42a3e1e7bd47aef9c96da20f1becd53af9816aacc51dbc99bb2acbe051703e18

Observation 79cdbcc8-6c07-48b9-9e72-e300b01626e0 · inbound

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics cites this paper.

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:32.407540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:05:32.407540Z digest=sha256:0a9e6ae52143ea3db9b7707b34b43c13e2a0e1afcc40b406ceb6b9fdf6545e85

Observation adce1126-a8a0-4236-b4ca-617e4f826596 · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.438516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.438516Z digest=sha256:e42adbe1fc4088c46b99b6ebb21523f8e016666de2f9c202d0f0ec2e6d69f189

Observation 83510f07-e3d2-4591-9f11-3089933d292e · inbound

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding cites this paper.

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:16.769169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:10:16.769169Z digest=sha256:23504f24180115a5f164684d26b6bdd2c651cc950b78b8f0a4e5da413ca55b01

Observation 9faf2417-4896-46ee-9883-69e0bdcc3e56 · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:34.988625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:34.988625Z digest=sha256:527b069523c021bba0d42b7ca00d8b6ad3e81e444104ea4ff8e73ffc30cd284c

Observation b5c68ae3-7955-4fa6-b102-f464b8aa660a · inbound

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos cites this paper.

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:51:28.132505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:49:12.987772Z digest=sha256:cdba3d4757b7d0c6492402f32c8a1d34d717ce20063d4c0202deaa483108b4a3

Observation 8313bc12-2f58-443d-87b3-565b183979ff · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.873910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:44fe9fe78406f7ee6265bbe1d32085516cdef2be8e0adeaa1751dcbde79c7b43

Observation d1c7aabf-cc5c-4e9e-adc7-4a15863bba87 · inbound

Small Vision-Language Models are Smart Compressors for Long Video Understanding cites this paper.

Small Vision-Language Models are Smart Compressors for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:15:51.479344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:38:49.654073Z digest=sha256:8d23c60f75a305e949589cd2ae5975ed8ddefcd774e81956b49592d6dc706357

Observation 82579acb-6fab-499a-978b-dcefacfa5f2e · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.569125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:391560dfcf632984ee28a94cdcf4da0d86517ec22e05985808c9f223f917998d

Observation 3e04ced6-a4e9-4ed8-b2f9-66302f455250 · inbound

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning cites this paper.

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.759603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:51:52.861981Z digest=sha256:98410621c690789e874f42f1933bd0318c80cadd9d04a25d5146db2b514b2382

Observation 92803c71-da0a-4e87-a538-1536a6e706fd · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.522561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:21429fabd35582eb9b069136c333c3ab2997f7b37003d25fd9338f7fc60d84f0

Observation c8418107-8250-4579-88e0-92a470464e44 · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:20:56.479634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T02:30:55.939351Z digest=sha256:9fe066f63ec8f72c1faf01ad47167c507c67787ee86d06482f1aa48216b9017a

Observation c1005e51-57e9-4c04-9b1d-862c7d6bfbde · inbound

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding cites this paper.

Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:17.700487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:00:34.728880Z digest=sha256:26a9652aef5f0c477c81f6ef567db5312e1b7f0c54089b84f7c2278bd34f98b9

Observation 88c56c0f-3dab-484a-878e-e22f4a3402a5 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.452976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:e560b986f1db19b215732a58e48be2af23f8cd830ab121f4768a2caeca087b87

Observation 06199d9c-9c90-4ba0-bca6-6dfab8cdebb0 · inbound

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering cites this paper.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.071338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:94a2bd108b67b46c20f663bf342679609c43ea107cf1c87c788bd811704680d8

Observation 4f5b8170-f9d8-43a5-82cc-fa3598a11b04 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 296

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.192666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:f99ec0e8c259bae2f37e461719755287fc0ab0171ebf82f02044494599ef9587

Observation ed14d6ca-5e62-4cb0-9eed-d7f32f883594 · inbound

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs cites this paper.

Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.726244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:02:37.632461Z digest=sha256:253e93ef71a24fae0542c8ff5f356ba781fe9334d31fff9a47ccf27cebc3272d

Observation a605ff3f-0bea-437f-a1a5-671a2adbaa57 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:29:15.349201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:14:11.297384Z digest=sha256:b4539b22184f45e12c2734d7622eb1e92c62a7fa38f3988739783b9399ac2064

Observation 71114075-e508-49b7-8ba5-8ebae0751e93 · inbound

Native Active Perception as Reasoning for Omni-Modal Understanding cites this paper.

Native Active Perception as Reasoning for Omni-Modal Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:57:19.810078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:57:19.810078Z digest=sha256:caf2fe7b29bfd2fecc6b6efb628bb3b33ae2fee173e988d5f6b8023cbe9594d0

Observation 01d453a6-8134-402f-ba94-496cc1c508df · inbound

Latent Visual Cache for Video Reasoning cites this paper.

Latent Visual Cache for Video Reasoning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:09:15.248815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:09:15.248815Z digest=sha256:9a43bc18773e4a0e4f6e49728c90f0da15e5ba318a096860b4ae99e1fba51d59

Observation 30cccac3-b3ac-40f9-adef-3c640cb6616b · inbound

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization cites this paper.

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:13.271024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:13.271024Z digest=sha256:b5e1eca2fea7791c53273da47cabb2d537060ef9daf195d00a5faa42d7bc304a