Pith. sign in

Paper Citation Record · LEDGER

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 14 inbound Pith citation observations for arXiv:2506.19225.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19225 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:13:11.715462Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:34.999140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.312479Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2ac8c17c-ef6b-4bb5-b661-b76f7c694d18 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.005716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.005716Z digest=sha256:a7f6ffa82f9bb31c1ee0d5607ec82a6a5f1eedefa1a565ae16df8ea353e73a2a

Observation b639ce59-0d16-4689-a0d5-feccc88d13f4 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.063644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.063644Z digest=sha256:c8ab2a6a493f215757f354e2f6b480c66632732579c51a64c2924483fc5cbeed

Observation 7e1378af-d07c-4fa4-a737-f20479b34dcf · outbound

This paper cites Claude 3.https://www.anthropic.com/news/claude-3-family, March 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Claude 3.https://www.anthropic.com/news/claude-3-family, March 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.477127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T23:13:06.126866Z digest=sha256:a670f3751e95d5bbef394a8089a59e5221a7f83375432580c054a8e3a39306b3

Observation f4fadd24-7603-4e0a-99b1-1f7d1ceb31ef · outbound

This paper cites Qwen2.5-VL Technical Report.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.296080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.296080Z digest=sha256:3d6c5776cb2045793ecf31b8dd5fddf81bab04412009a0d68780ed9a8952770b

Observation fee6a7a4-e6d5-40f7-8eb9-7aafec2469ab · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.349713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.349713Z digest=sha256:93bea64decb06e83768e00be4bc68eb3036633e21b6b3f61238eb2b484350708

Observation bdf7c58a-339b-43ed-9661-ce81ff93e999 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.438610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.438610Z digest=sha256:25d6645060d162c7cafd7d655fc898b01e78ec546638964c8f22e670a53772f0

Observation 700229ed-b13f-4ceb-9151-932718b384b3 · outbound

This paper cites Nvila: Efficient frontier visual language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Nvila: Efficient frontier visual language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.244169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T23:13:06.555749Z digest=sha256:d0d862fa0b859750cdc39e6773bb41cf71df687f7459c304c521ed350ccc7575

Observation 97be42b9-d35b-4687-a316-74b8b543f544 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.733535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.733535Z digest=sha256:527a0e7042034a69aa0315e134200cc0d38108c1b2ae88ce4f4e3837ea2453f7

Observation 15d6f43c-d283-44e8-8993-c24db3336a95 · outbound

This paper cites Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Eagle 2.5: Boosting long-context post-training for frontier vision-language models.arXiv preprint arXiv:2504.15271, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.878329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.878329Z digest=sha256:6e624eae6172efa90f83eec2bbe959b800cb6a51a0e8a231e336fda1bce48931

Observation db0a7ef3-0626-44b4-8c0a-303153eea28b · outbound

This paper cites Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:06.962953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:06.962953Z digest=sha256:9ebcaf7671edb385b844510c28ace8a3f531e7853e154e5d3dd8ca37cb117992

Observation 38fdb9fe-d313-4570-8f2b-565d97b42853 · outbound

This paper cites Long Context Transfer from Language to Vision.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Transfer from Language to Vision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.042420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.042420Z digest=sha256:27bf0a0b7e9c9cd6e9e1088796a172fa3921dea8c2ab7a55b0c641a29a2555fc

Observation f9bd112b-a33b-408d-b03c-4849703a8fb1 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.193817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.193817Z digest=sha256:f8e6d5ee9c5353be7f42097adb2af4ab44ee8dc336c5e3f01b5f8b62dd3d3416

Observation 42939bff-b41f-400e-a552-b7378b9660f8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.311670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.311670Z digest=sha256:a31aba90366fb15e17847ff3d948884ff9355956cd202757050e4172a1d2fb6b

Observation a6344937-afcc-4870-aefa-9a1db9dde81f · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat: Chat-Centric Video Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.440858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.440858Z digest=sha256:19a67c4311a2cfe9bc775c724b62473d9c029af55df105dc3fa54483561b95f2

Observation 82ed5029-2983-493e-ab0a-c3a858ba68dc · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.554817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.554817Z digest=sha256:ebb7d6d3247c2fd4f0809536e33e1256fa018642613424831c829a10c4d4a518

Observation eb7cc081-46f7-4693-b683-5c2b592d35d0 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.690732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.690732Z digest=sha256:2001e5f155d733c392ab77b5e39c3083ade2c90b6dd850489c7e953567511888

Observation 15d543b0-28aa-4420-97a8-ca925f1e9c4a · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.794754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.794754Z digest=sha256:a641fe1da80fd55a0b2861a7917d3ef334dfd60a94705c22b3533c9ea85d83cf

Observation ad403f34-6964-4c60-8b85-f105b25d7c35 · outbound

This paper cites Snapkv: Llm knows what you are looking for before generation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Snapkv: Llm knows what you are looking for before generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:07.926631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:07.926631Z digest=sha256:eb46c59cdb54491deaea96910c5032d8a52757e5097b50c2b94677bcd77e15ca

Observation a44361d5-4a69-4341-adfa-b6cbd04d5bfb · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Efficient Streaming Language Models with Attention Sinks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.030081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.030081Z digest=sha256:835f86c00f5298e68b86ba1be3b6d3ee7826557b76e14174e8204afab62a6e02

Observation 36b84e53-46fd-45f3-9053-5d382c196328 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:14.063101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T23:13:08.152487Z digest=sha256:17ec8ec9881d8c7cfee78e800486214251e640f6951e1fbb1f1f039190a89ebc

Observation 69586b48-fd67-484f-b4bf-51fc7821089c · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.220997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.220997Z digest=sha256:989abf13a49f7c783856fa60eb54ec34392c4953a2b5ac20e41aa21f00ab7642

Observation 7182a99b-e30f-4d2d-8389-98e59e56bd28 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.301692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.301692Z digest=sha256:781f9f5d75c8ae9e54b6a51a5a686b6bf535f08055bf1043dedecbdc8dc040d9

Observation 8140ec18-9ef6-4d59-a03d-33427aef3e36 · outbound

This paper cites Visual Instruction Tuning.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.356194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.356194Z digest=sha256:2175667f275851fbefb7f170da393d6b2a01ee6dd19c2f310b91ba2ad0134a4a

Observation 57b1f109-c7b6-46c0-a018-99850cf9d76f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.443107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.443107Z digest=sha256:db750ca2a6b8a7aa1a85c8d9f27ea711fa5ff1a9787ac7a13e6458427ad6a0d0

Observation c21a8a70-8e6a-4077-9023-624668c7e886 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.491853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.491853Z digest=sha256:437d7dd176f3aad893a2b262555011d47e843297aab9d06760449b043c27536e

Observation 78962752-4e28-4664-8d71-945f60701170 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.548099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.548099Z digest=sha256:8e643a0c53e275ed196aacde5496287959a8c1e710537b9ff70fdeba145b606b

Observation 892779bc-8de8-467a-9c2c-783d8d95eb7b · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.631395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.631395Z digest=sha256:7163018b655d48601e49c68ab6b29c31ce167ced74d0798ced87e4b58917c363

Observation b46f4c0f-ea63-49b5-b8b9-e65a78283748 · outbound

This paper cites Vidtext: Towards comprehensive evaluation for video text understanding.arXiv preprint arXiv:2505.22810, 2025.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vidtext: Towards comprehensive evaluation for video text understanding.arXiv preprint arXiv:2505.22810, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.728844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.728844Z digest=sha256:323de0b411e9c01ccffa365e7d87fd136dd336ddbbfc5fb6dae2a55cb723b510

Observation cde2b3d5-766a-4eb8-98d8-c27cdfbdd032 · outbound

This paper cites Vid-SME: Membership Inference Attacks against Large Video Understanding Models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Vid-SME: Membership Inference Attacks against Large Video Understanding Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.790871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.790871Z digest=sha256:cad4d45cb7667d051cc8be9c846cf9b8ff47e71452f26504a62c48c9489d8ad1

Observation 8ecbc304-40b5-428f-8234-871d5a1d251c · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.857202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.857202Z digest=sha256:77ce710996b3f978bbdda22e3b30b060e42594416b839360183e029adf7b8854

Observation e61a192d-bfd2-4f6a-a674-b10c7a491ad3 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Llama-vid: An image is worth 2 tokens in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.928962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.928962Z digest=sha256:2235263c0a067494e51bb96480e4bffe7ea965c57880228305e47bc55a6b4950

Observation d632ca7f-a4c0-4de8-9889-4d3d58abaa42 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:08.989110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:08.989110Z digest=sha256:7589bf897f4f83921dc09fa372aea4e8399d9290d38954759daf487ac1d41fdc

Observation dc11cca0-5f31-4481-8d6d-72be0407d318 · outbound

This paper cites Token Merging: Your ViT But Faster.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Token Merging: Your ViT But Faster

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.033586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.033586Z digest=sha256:4651cf0ba63e8826eb8499891425cce6a47c1e70d23ed14c1677b0d580557e8a

Observation 0c1400bf-89e9-49e6-8860-b49ebd31ec38 · outbound

This paper cites Long Context Compression with Activation Beacon.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Long Context Compression with Activation Beacon

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.127706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.127706Z digest=sha256:3676551a1862d66d23f6417ef6087c21409cb4317d3f281fe2434a461b1393fc

Observation 6ae348ee-dcfc-42ba-8eb1-4c7c705acd17 · outbound

This paper cites Lighter and better: Towards flexible context adaptation for retrieval augmented generation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Lighter and better: Towards flexible context adaptation for retrieval augmented generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.797703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T23:13:09.293482Z digest=sha256:00676eeb3dc0cfb39241c6dba8c3452ff6edeb868852835bade4ba5ff765a88b

Observation 71970c1b-8b03-4058-a55e-e5b57db8986c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.457172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.457172Z digest=sha256:c4a90d2c648aeeaaebed3540c0888c9d1c494514b0d19b0815efe48213a8d3b9

Observation be77acfd-9103-4b5b-8419-9107f89383fb · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.567257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.567257Z digest=sha256:80a8b5f871ccfe3f59ccc632c38e2469fb10a730433ea456b8613e57717e1a82

Observation d4fe11e0-a9d9-4a45-832a-8ac68d19e073 · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.732084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.732084Z digest=sha256:4465a857776a4cd23db46d9c698f19b744c338f91f5e26cebede4a2ab15dba59

Observation 7987dfb0-bcf0-4639-9fc3-122cf89bc1c1 · outbound

This paper cites MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.877936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.877936Z digest=sha256:a83d9a2a89f50df3e0611d06bc2488f176a2722edc5c21a3a3797cd426cd4d51

Observation cb51bb97-22e8-4030-9768-4f0a8f198555 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-xl: Extra-long vision language model for hour-scale video understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:09.996456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:09.996456Z digest=sha256:71d6be8c7087588c42d0de4948242ceba90dbb9aebe821934bc8493fabe54a06

Observation afa52dc7-5d5a-4b51-8d3d-97c898209423 · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.091131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.091131Z digest=sha256:dffa4d33779de0aebdb9e9eb8b7eec3c9b96fc28ad7d337c9b62e05d5123ff9a

Observation ff13dda0-1cb8-4aa8-a1d6-dbe5b2d43db2 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.165153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.165153Z digest=sha256:fb4ab9369c92a8790350ab3f0e00ec268cbcdb093f463bda0c20a0db7b2ed525

Observation 96e6cf11-c3aa-4c13-b51b-147e6c544cae · outbound

This paper cites Qwen2.5 technical report.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Qwen2.5 technical report

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.570856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T23:13:10.225687Z digest=sha256:601894539b217c4dc82a4b1f386bfe2bd9c5a77a577cf7d6dc267d545cf3d885

Observation 4ef0ce03-ebad-4ebd-8418-b74ff83892fe · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.317216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.317216Z digest=sha256:f2f4d2df360548d4f944fb1ccda0c68cc72f3e3805a85767d22037f799e5b7bf

Observation 36d57352-bccc-4850-955c-2236ca62a090 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.420868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.420868Z digest=sha256:e08001eaccb281eda3405e69791c1064cd93666fa5303b6387bd0ec8c22a572c

Observation 606b8738-b551-45d9-b74b-f46224ce850f · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Internvideo2: Scaling foundation models for multimodal video understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.526472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.526472Z digest=sha256:11982d1fe43e72b020e95c32d7225ad1056039aeafd9516d7735b65b5b688a7b

Observation 817a48cc-2a7b-4f0d-a292-666c9c209a86 · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.664784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.664784Z digest=sha256:3d5a516486f20f1d08b4d018fce2d2aa08547c6bc95bae76e7873d884d7e8c04

Observation b3b144f7-02c7-4305-af42-64cee1aca6f3 · outbound

This paper cites Memory-enhanced Retrieval Augmentation for Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Memory-enhanced Retrieval Augmentation for Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.804795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.804795Z digest=sha256:5dc879b597d0dcdec819426dd9cbab8a0fdd6f4856f72a205032871683be61b3

Observation a75ceb8b-10b7-4ebc-bc28-60f157a9b6ee · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification MLVU: Benchmarking Multi-task Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:10.893556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:10.893556Z digest=sha256:6df751f77f29b2021b66907fcc89b1d10e87fc81c84832a03fde7c3a68746535

Observation ed07ec05-ffd7-4a34-8c26-c952a8137a10 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.007680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.007680Z digest=sha256:f5667db66fadb36537f5931c034a1a305162a0f69bf5e7149cffc748b4819968

Observation 21397d08-1c0d-4164-9c4f-9539c132815f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.123012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.123012Z digest=sha256:0e6c5c91cfb0178e1e559d683ecf0c8286f90d3a3d87149e5f019c91b17266be

Observation 51b153dd-a8c3-4786-8d0b-f0d8d967a8e5 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification LVBench: An Extreme Long Video Understanding Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.200327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.200327Z digest=sha256:d33f157450d97071698964474917f030ab0815beae34d21aba4d7089a6aa57ee

Observation c7fc560d-b42f-45f3-a437-4be1ade2a1d1 · outbound

This paper cites VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.304314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.304314Z digest=sha256:8c174bc89a52a536d29b4218ccaf72792c2752c79fc809a8aa71ae7f33a8981d

Observation 741a5211-f14f-4cc4-adce-e004e9b0feae · outbound

This paper cites Tall: Temporal activity localization via language query.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification Tall: Temporal activity localization via language query

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:13:13.313796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T23:13:11.448061Z digest=sha256:fea14159e556911e0031c8a42b32d8a568e623b0937ab87ad016453f60d60936

Observation 0c4697dc-4d72-452b-95f8-863bae8de435 · outbound

This paper cites V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.605015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.605015Z digest=sha256:9bc98cbfbd82ca6827182ffd2a1be9c935acb9a305a39553e7d6fb839de1d7be

Observation 4a909771-f00d-4697-acda-ccf2c2a7a875 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:13:11.715462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:13:11.715462Z digest=sha256:478767df814e5bb75444a26e918813196689cadf2669579059ab5d6072d92100

Pith citing papers

Observation af469056-2597-4916-802c-3e9a4e6eb44a · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:df2148e858db39e0f85eb16863e58a322bba72f090742d8c93aef293a665bccf

Observation bb56a904-e9af-4aa6-bde5-e33822a78dcd · inbound

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models cites this paper.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.477147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.477147Z digest=sha256:ef636854a7c5eedd0cc0245f0f65af6630f2d1d337664ccf0d6a42976623cf8f

Observation c6eec0db-231d-4e97-a023-493219000c2a · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.149183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.149183Z digest=sha256:9a2485723f3430e811a9636f0bfc8dcbc92d07e442a187da9ea1578eece70d21

Observation 7e1ac90b-963c-42f6-aea2-6103032d8c14 · inbound

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations cites this paper.

cuRAMSES: Scalable AMR Optimizations for Large-Scale Cosmological Simulations Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T09:09:32.932051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:09:32.932051Z digest=sha256:e841399915832ea432eadefb153d91e496dc8497653e0a5b80941649dae61c02

Observation ddef1e6c-2631-451c-9f2b-f16c6b743166 · inbound

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference cites this paper.

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:52.769031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:45:26.555395Z digest=sha256:01b6fe24cf8634c5495635809b2e4c1860150fa08c3acd09ce34255139da0e0e

Observation b87dd21e-68ef-4313-9a82-2ef1a7024db7 · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.566701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:b316c6ece3858596a00916c7a200bcb46935d0f61cad2e38b0860fe948b13773

Observation ae7501ad-991c-4e6e-8fac-cec4f55cf572 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.628146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:3be91775871852497206e1a42c103db18d2f3ca71407b303c1e76dde8f508587

Observation 55fdb6c8-43b3-4cec-8dc1-257732f2f102 · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.314625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:4ec0789eaa60f0b6d0fe44f11a998d7aa36737a1ffc54c74de4d1cca6bb814d7

Observation 605428c9-ced5-4e26-b65c-746de9d130bc · inbound

FOLIO: Focused Semantic Memory for Streaming Video Understanding cites this paper.

FOLIO: Focused Semantic Memory for Streaming Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:46.344843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:46.344843Z digest=sha256:daf0ea26ade1881fa0b557d8456e52cd117e4ea0dc87c7fdd4bdd8cb118d3270

Observation 7e95a7ad-809f-46f5-8825-6b314a59a3ec · inbound

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding cites this paper.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T00:44:41.162135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:44:41.162135Z digest=sha256:478dedbe718d16c1e595031c59d5c5420a2a02fd795c77f9465babec9fea6555

Observation 3e370153-360e-4d06-9b8a-c84ad2d9ff37 · inbound

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding cites this paper.

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:26:24.782850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:26:24.782850Z digest=sha256:2b209eb81ebbb8753fbc93ec052dd6e3dc48d3e043b530540ceacfc2ea3f33d2

Observation 64f8a689-5d3e-49c2-b6f5-3ecbe11ffdf0 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:12:46.830446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:12:46.830446Z digest=sha256:cd4c8b2da6ba7bd0375de4e489f32f6a4c06c11acad7f1f0979cd8082c2792f4

Observation 2bcfcad6-5456-4d17-b706-deb69eae8560 · inbound

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience cites this paper.

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:15.479594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:15.479594Z digest=sha256:7cfd5433af5ada912a00d4a84d196f003ae09bbd09ea36f6f06f9e0cd039c3b4

Observation da832fe9-bd47-4155-84ab-af0c83ea644c · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.999140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.999140Z digest=sha256:89153c50c2e49caeb62e305347c5342fe908267f9b01b919770956fe7a0dab60