Pith. sign in

Paper Citation Record · LEDGER

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 14 inbound Pith citation observations for arXiv:2506.19225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19225 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:13:11.715462Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:34.999140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.312479Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ac8c17c-ef6b-4bb5-b661-b76f7c694d18 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.005716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.005716Z digest=sha256:2756feb5feae84313c1e96c8ae17e415036b0a48bd76a00f564f6ef60524221c

Observation b639ce59-0d16-4689-a0d5-feccc88d13f4 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.063644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.063644Z digest=sha256:030ceb0cd571b9c0ccb5766a535b4f68f7edaf693d2a66827776c4788d6bb3be

Observation 7e1378af-d07c-4fa4-a737-f20479b34dcf · outbound

This paper cites Claude 3.https://www.anthropic.com/news/claude-3-family, March 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Claude 3.https://www.anthropic.com/news/claude-3-family, March 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.477127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:13:06.126866Z digest=sha256:011f7f1259903d84831729c79054ffb383e13e3fc0bcc23cc4eebb2773e6e115

Observation f4fadd24-7603-4e0a-99b1-1f7d1ceb31ef · outbound

This paper cites Qwen2.5-VL Technical Report.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.296080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.296080Z digest=sha256:afbb1bad090311f4c0d3d3431c3bd81735d74c56fe16a20d243161ea021accf4

Observation fee6a7a4-e6d5-40f7-8eb9-7aafec2469ab · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.349713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.349713Z digest=sha256:9d33e7616596883f5baa5516a0bfef77b4dd83c7ebabe2cc8188433de6393cac

Observation bdf7c58a-339b-43ed-9661-ce81ff93e999 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.438610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.438610Z digest=sha256:9383ff158f6386c7e5ee66fc0aa52bde8a15020484cabdac8c62c1f74b79929f

Observation 700229ed-b13f-4ceb-9151-932718b384b3 · outbound

This paper cites Nvila: Efficient frontier visual language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Nvila: Efficient frontier visual language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.244169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:13:06.555749Z digest=sha256:b11b8235cd08f6d763d2ac0fdce5686b92510a006817409e2b72f8d463f3a885

Observation 97be42b9-d35b-4687-a316-74b8b543f544 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.733535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.733535Z digest=sha256:1626ddf570746cb0ff820e2cb25d2f0aa823a2b14bb454d35e7971ae05b702c0

Observation 15d6f43c-d283-44e8-8993-c24db3336a95 · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.878329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.878329Z digest=sha256:498875e117d3df042fbc2d8d36340cf3fdfd431474b35fa7f267cec4acba5922

Observation db0a7ef3-0626-44b4-8c0a-303153eea28b · outbound

This paper cites Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.962953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.962953Z digest=sha256:87865950cda936fda64b781f1439f908b4e870f72018243f687217c83ad11842

Observation 38fdb9fe-d313-4570-8f2b-565d97b42853 · outbound

This paper cites Long Context Transfer from Language to Vision.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Transfer from Language to Vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.042420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.042420Z digest=sha256:3d848e974c78f9218197d8c7af124e14a8c9e02f947d9968abb227d30ef41e21

Observation f9bd112b-a33b-408d-b03c-4849703a8fb1 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.193817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.193817Z digest=sha256:78e3810cf477d91a2056b25ac903a872fdb0576dd2f136069e9cdd31ec15a7ca

Observation 42939bff-b41f-400e-a552-b7378b9660f8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.311670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.311670Z digest=sha256:8700d40fa4903c5cecfe2650abbf84bf7f29d140ee6de5cf8a254d4aa092e9d1

Observation a6344937-afcc-4870-aefa-9a1db9dde81f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.440858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.440858Z digest=sha256:72a0803cfda2856163b6f83461ede0987e9854dc87ceb58905e041e094216a8d

Observation 82ed5029-2983-493e-ab0a-c3a858ba68dc · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.554817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.554817Z digest=sha256:a03c4e8e3a3c1bb8fb6b496f279da7ce4eeff953afe31e22cb1ae9579864a8df

Observation eb7cc081-46f7-4693-b683-5c2b592d35d0 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.690732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.690732Z digest=sha256:905ccb6c5d9cf5c9d24e8da82d45167f4375d334aeb4742f54d60bdf299bcd9e

Observation 15d543b0-28aa-4420-97a8-ca925f1e9c4a · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.794754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.794754Z digest=sha256:2f73f123d9864bf622c46ed061511b6a9ce64c9b90cc07781ef4086917c0024f

Observation ad403f34-6964-4c60-8b85-f105b25d7c35 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Snapkv: Llm knows what you are looking for before generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.926631Z digest=sha256:ad682a678ff0967d87a7830cd98c0cd72102b0f6520891a3f8e06a4e9bb3139c

Observation a44361d5-4a69-4341-adfa-b6cbd04d5bfb · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Efficient Streaming Language Models with Attention Sinks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.030081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.030081Z digest=sha256:93fee434a28b63c75e624e9a8e2210f2176a7a975c30db8f2f018c64c4b88654

Observation 36b84e53-46fd-45f3-9053-5d382c196328 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.063101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:13:08.152487Z digest=sha256:0dd65d3e6b640f67d88721e764898cdcd29ada9104c01a7aeba59eeddd1c3e53

Observation 69586b48-fd67-484f-b4bf-51fc7821089c · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.220997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.220997Z digest=sha256:8246b7f3e64cae6c835f0287e7f4c646d0535f7cc661a1fae64d10c8e5ed8059

Observation 7182a99b-e30f-4d2d-8389-98e59e56bd28 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.301692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.301692Z digest=sha256:07409dd7aee59638cdfae117c37aaeb8753f8a0d36d17c6808a40826b44a807e

Observation 8140ec18-9ef6-4d59-a03d-33427aef3e36 · outbound

This paper cites Visual Instruction Tuning.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.356194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.356194Z digest=sha256:3009cbdd2f69f2f2b43dc574d6c4b64e2c6f3654d1701aa8f4a39b3835fc0ae6

Observation 57b1f109-c7b6-46c0-a018-99850cf9d76f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.443107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.443107Z digest=sha256:a38706e87200df6cdeaada5bc1f1725cd2f20fe1ed8177a2d9b85eb6a8fdc475

Observation c21a8a70-8e6a-4077-9023-624668c7e886 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.491853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.491853Z digest=sha256:f9899590bc62de124beebc80112cd2d60b3656fe58904cbda74fb64478f8ffe2

Observation 78962752-4e28-4664-8d71-945f60701170 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.548099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.548099Z digest=sha256:fd012043915da5acb6b5c791b1caa71d8f7dc2283763f3d6c9fca51bf787f6f1

Observation 892779bc-8de8-467a-9c2c-783d8d95eb7b · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.631395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.631395Z digest=sha256:c4b7815357b45f1076d3a53baa198df1f1d083d7569f87244fdf88244cb5c9f2

Observation b46f4c0f-ea63-49b5-b8b9-e65a78283748 · outbound

This paper cites Vidtext: Towards comprehensive evaluation for video text understanding.arXiv preprint arXiv:2505.22810, 2025.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vidtext: Towards comprehensive evaluation for video text understanding.arXiv preprint arXiv:2505.22810, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.728844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.728844Z digest=sha256:c11ce0602c968a1c42ddb6c989aa999bfabd2e7f7657e479e7bcc3dc66721c29

Observation cde2b3d5-766a-4eb8-98d8-c27cdfbdd032 · outbound

This paper cites Vid-SME: Membership Inference Attacks against Large Video Understanding Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vid-SME: Membership Inference Attacks against Large Video Understanding Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.790871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.790871Z digest=sha256:d40fe9c7c6b41eacdc659ce5f80d24591db08af46fc3e1689c46fa2aba32f4b8

Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.857202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.857202Z digest=sha256:6d4022acbd9bb72c8f1e2b5b9c4d7f21edbc62ee0e3f5d3ad736edbc7df39b77

Observation e61a192d-bfd2-4f6a-a674-b10c7a491ad3 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Llama-vid: An image is worth 2 tokens in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.928962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.928962Z digest=sha256:4dbe7fbcb902e33a7775242c1cb519c0cc5807d53cd8782c8ec7463a89459392

Observation d632ca7f-a4c0-4de8-9889-4d3d58abaa42 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.989110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.989110Z digest=sha256:57fba7f69d7310b7dcd40e53052dd4666f7c577fa7d00aaa66af608cc34f0cf3

Observation dc11cca0-5f31-4481-8d6d-72be0407d318 · outbound

This paper cites Token Merging: Your ViT But Faster.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Token Merging: Your ViT But Faster

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.033586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.033586Z digest=sha256:911ccb0719922f1d09304ace26632a48eab1c8151012b21465405855584e97e0

Observation 0c1400bf-89e9-49e6-8860-b49ebd31ec38 · outbound

This paper cites Long Context Compression with Activation Beacon.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Compression with Activation Beacon

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.127706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.127706Z digest=sha256:bad31ff9d1a7f6b9d17a4165470f973cb871e608216ff72202d5bfa60376f128

Observation 6ae348ee-dcfc-42ba-8eb1-4c7c705acd17 · outbound

This paper cites Lighter and better: Towards flexible context adaptation for retrieval augmented generation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Lighter and better: Towards flexible context adaptation for retrieval augmented generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.797703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:13:09.293482Z digest=sha256:e5e3212b2734d62579fc745f9daf8dcf2116b6b1513765b642e36a122d37fcfe

Observation 71970c1b-8b03-4058-a55e-e5b57db8986c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.457172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.457172Z digest=sha256:9819fe858886bc7a51aed9da3579fd80f6f8d46eff19c7fa19c1eacd84f060dc

Observation be77acfd-9103-4b5b-8419-9107f89383fb · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.567257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.567257Z digest=sha256:381d43d9f3fca8d049ff884b2a6561a045fe530de6029bdcabaf886c22f16177

Observation d4fe11e0-a9d9-4a45-832a-8ac68d19e073 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.732084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.732084Z digest=sha256:e183f1c9f3e68c8b85cb829e44b256b583f17d7eb6fa6bbeaeee24879d03499a

Observation 7987dfb0-bcf0-4639-9fc3-122cf89bc1c1 · outbound

This paper cites MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.877936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.877936Z digest=sha256:4e6819a6e830a227754a194b1e45275f1f4dbcdc27c41c1a15ccda490bfcb6ca

Observation cb51bb97-22e8-4030-9768-4f0a8f198555 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-xl: Extra-long vision language model for hour-scale video understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.996456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.996456Z digest=sha256:d57adb156d5308df047f74030d1c57c5192c06b23b9aa6bca3545a698363c5a6

Observation afa52dc7-5d5a-4b51-8d3d-97c898209423 · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.091131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.091131Z digest=sha256:5d7e9e943e1ade9034e09f1087968ff1e05a55fac384719813b037df9f808eb9

Observation ff13dda0-1cb8-4aa8-a1d6-dbe5b2d43db2 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.165153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.165153Z digest=sha256:3c37052c2c56e5992946bcff8b83889dac5e2991ab40f8124eec2732ff2eb51d

Observation 96e6cf11-c3aa-4c13-b51b-147e6c544cae · outbound

This paper cites Qwen2.5 technical report.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5 technical report

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.570856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:13:10.225687Z digest=sha256:5a5c8732defaa3aa08b1facca2a01413039ad31e8aca1b0e1a4e48f33781012a

Observation 4ef0ce03-ebad-4ebd-8418-b74ff83892fe · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.317216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.317216Z digest=sha256:279cbe7bf911999e190ffb54b46134a8a9a31f4688147adf9fe7f7bb70384ac0

Observation 36d57352-bccc-4850-955c-2236ca62a090 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.420868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.420868Z digest=sha256:b0fc31c4cd4af1b743b2241c5c3c339f3ff0960104268f3dc1927a7ecd57cb85

Observation 606b8738-b551-45d9-b74b-f46224ce850f · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Internvideo2: Scaling foundation models for multimodal video understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.526472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.526472Z digest=sha256:12c8bc2983f742e08578e741fc8346b78bff1f00df5e81f2f98c1608c65de8eb

Observation 817a48cc-2a7b-4f0d-a292-666c9c209a86 · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.664784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.664784Z digest=sha256:43c6c4b6257df5afc5a3af5fbb6c94948293e597140e86222d4e49674087d960

Observation b3b144f7-02c7-4305-af42-64cee1aca6f3 · outbound

This paper cites Memory-enhanced Retrieval Augmentation for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Memory-enhanced Retrieval Augmentation for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.804795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.804795Z digest=sha256:647d9ec9512ba369327179b8145935d7d4d5e5d52957563ddbacfd71113315d4

Observation a75ceb8b-10b7-4ebc-bc28-60f157a9b6ee · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MLVU: Benchmarking Multi-task Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.893556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.893556Z digest=sha256:86c056228fe822d5a077eaeb946414a01ebc44b61a3ddf9621d977a75f1e5cb3

Observation ed07ec05-ffd7-4a34-8c26-c952a8137a10 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.007680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.007680Z digest=sha256:12830fa0f5adc474997d1a10e017f19cad813512d89aca84f10203f26c35a57c

Observation 21397d08-1c0d-4164-9c4f-9539c132815f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.123012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.123012Z digest=sha256:9849c33212ca2aaa5ae5c8c3e93d0f0cfeb9e920dbb78ca83e83395174a29309

Observation 51b153dd-a8c3-4786-8d0b-f0d8d967a8e5 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LVBench: An Extreme Long Video Understanding Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.200327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.200327Z digest=sha256:4b2aa38201ff1bde6fbb0fba03c1cf764c161729b6b2b24431ebdf9591f08eea

Observation c7fc560d-b42f-45f3-a437-4be1ade2a1d1 · outbound

This paper cites VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.304314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.304314Z digest=sha256:d1d02b72ad653bf12590c4a3a0da4455fc825efd664aac78ce6e03ef39031ea4

Observation 741a5211-f14f-4cc4-adce-e004e9b0feae · outbound

This paper cites Tall: Temporal activity localization via language query.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Tall: Temporal activity localization via language query

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.313796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:13:11.448061Z digest=sha256:a558911b514fbd31a77f84ababd300ec108ae396dd4b1a6a760732904c8052a6

Observation 0c4697dc-4d72-452b-95f8-863bae8de435 · outbound

This paper cites V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.605015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.605015Z digest=sha256:4151b9baea8b7f7c64e8aa68b12e037661793b3230df9d13d81466f572e7f113

Observation 4a909771-f00d-4697-acda-ccf2c2a7a875 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.715462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.715462Z digest=sha256:9d6c86cbcd6c4e28ce47904b0b2fe281062694d4955b1516077141cc04f6abd8

Pith citing papers

Observation af469056-2597-4916-802c-3e9a4e6eb44a · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:398be7f5ec59e0a7989dbe877066b088560eeefdca014b6fc80dd7fc41b5e3d1

Observation bb56a904-e9af-4aa6-bde5-e33822a78dcd · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.477147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.477147Z digest=sha256:05a9e684b8cc2b999ac63047d2e09d090b5421d58c8fe1410153920514cdc884

Observation c6eec0db-231d-4e97-a023-493219000c2a · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.149183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.149183Z digest=sha256:b797bf8b9d0c53ebc18ef9b7f8008a728ce2ca0ebe71aa9d2638f74dbcb89e81

Observation 7e1ac90b-963c-42f6-aea2-6103032d8c14 · inbound

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations cites this paper.

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T09:09:32.932051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:09:32.932051Z digest=sha256:1ed2d7acfe210b0871a80dffc19adf16d98a14394bed991b40a0b4ba7bd87fc8

Observation ddef1e6c-2631-451c-9f2b-f16c6b743166 · inbound

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference cites this paper.

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.769031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:45:26.555395Z digest=sha256:bbf76f3fc30ed8a87a36bedd92657d097c8a784cff21b8225d62b4586275bf8a

Observation b87dd21e-68ef-4313-9a82-2ef1a7024db7 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.566701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:d845cf39a1d92fa22405dea410327dc4f8c416838e3543a80791fc58e338706c

Observation ae7501ad-991c-4e6e-8fac-cec4f55cf572 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.628146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:20a2be235cdb0e5f1fe4b4051eed1e763780ba0399cd593a67a08b23daa25048

Observation 55fdb6c8-43b3-4cec-8dc1-257732f2f102 · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.314625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:0cfa3946321877b18a3f7a31088e6aa25d83b5576027d0b870cc39105241af59

Observation 605428c9-ced5-4e26-b65c-746de9d130bc · inbound

FOLIO: Focused Semantic Memory for Streaming Video Understanding cites this paper.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.344843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.344843Z digest=sha256:232100fbc27a84ebd06600d0c62f3e0d1d5a87ba8b988863ca383413aa3bca3d

Observation 7e95a7ad-809f-46f5-8825-6b314a59a3ec · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.162135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.162135Z digest=sha256:8ad34313cb0324849a9783eb24254ba2c1eb8c71fcaf81030c91ee9dc69cc0c5

Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · inbound

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding cites this paper.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.782850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.782850Z digest=sha256:067cbf6a1b1e9ddef7b766e852d4a3979ae5be578248c2fa2466cab20ced8bad

Observation 64f8a689-5d3e-49c2-b6f5-3ecbe11ffdf0 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:46.830446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:46.830446Z digest=sha256:409cfb542005d1762a2072d1ec6367647a0a07327c1068cb314966f8ae8b99d8

Observation 2bcfcad6-5456-4d17-b706-deb69eae8560 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:15.479594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:15.479594Z digest=sha256:f587613c16a2d273b8c9908aad7933ddb29d3e4bd51d1fba81e4f93c09843bc6

Observation da832fe9-bd47-4155-84ab-af0c83ea644c · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.999140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.999140Z digest=sha256:3f98e0c71b03e0825ed2191a658dd6a3d4505908b962656b2a8e0671bb76fd2b