Pith. sign in

Paper Citation Record · LEDGER

Task-Aware KV Compression For Cost-Effective Long Video Understanding

As of 10 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21184 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:39.894521Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bf89560-8927-4e62-a325-3ac8e3dc034c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.772060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:36.457741Z digest=sha256:c9e547e2dcad95037447c00b3802e33a835b30012e5d750f145611138967af35

Observation 89ab8ae2-2bac-47c4-9175-999e557b1902 · outbound

This paper cites Qwen2.5-VL Technical Report.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.521460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.521460Z digest=sha256:fa019aa7ae6c608cf48a158fa9a0f4a93b44680e6e1ebd4011c8d82676fde962

Observation 3c9cd646-4ba3-44ce-8828-ca0f1aeea3ea · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Task-Aware KV Compression For Cost-Effective Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.613827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.613827Z digest=sha256:5c755b7144f7cd8051cdf83def20d03da52a263d8c0915c0f0e04984b33ccabe

Observation 636967fc-0fe2-4008-8d22-6ac784cd54cc · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Task-Aware KV Compression For Cost-Effective Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.671820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.671820Z digest=sha256:00e6907a32d15877cc73ff27eb6aebc1e637b1a8f933427f1a8eb5c5564124e4

Observation df79a092-af4f-40f0-8f24-efedd0d61324 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.711180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.711180Z digest=sha256:af17822731b187412f6124977a84b9cffc5ab2bd1839fe41af33edaa5b42d6d7

Observation 2e9a38c5-4775-4a78-91b0-82d2885203f9 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.783376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.783376Z digest=sha256:0de321faca50fc7a9ec7206f16212beab86ddb7105f40b9fdef3ed4dcc99d31c

Observation 29079ea3-7f5a-45e3-b673-b96d566aed81 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.868717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.868717Z digest=sha256:e946d70d846db99d87c8ccf0a6cc6c14abf0efef9e25557e63145d880e87ecdc

Observation ad67feca-2ef2-4227-929e-8d25852f59fd · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.961899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.961899Z digest=sha256:7e260dfaee7576c0009651788ca861f96d5865d8665f83fa4458ad9a12a091bd

Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.020202Z digest=sha256:e18d0120fdb33a6bde7f916617cc2b3f0605c198817f302255768ac63794a89f

Observation 2140fb08-a1a6-43e6-a721-ec067c175a23 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.156162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.156162Z digest=sha256:a303fe61267f5fa2789b46a0a44c1e349affc86f30f5087ddd4eb46ad30a7ef0

Observation f7a744d3-7a27-4fd3-be55-82d104ffe9b1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.209127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.209127Z digest=sha256:371fcb1f8aba1cd1bb7dc627ff9879a8213a53b9b60f73555951d2a84825b80c

Observation a35b2b65-cf45-416e-b0d5-41cc98cb33d9 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.378601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:37.288226Z digest=sha256:1d9d180eb04ee1d76c3a7899e3e06fa1de50b738e6581532feaeac21a4a47f74

Observation 52d6f597-9a2e-40e1-8491-2f15f6a944e9 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.359101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.359101Z digest=sha256:c0eaa1707b63d86d293bfab9667f1de266b2cdbada8dbe215a620d43ab90d619

Observation e90a2db6-5f6f-4e3e-a0ed-6b1928ac8116 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.422766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.422766Z digest=sha256:1473da1cb8e887c551c8658b8c4ad60fc8fbd7c27fcb760c1068dc4838df459c

Observation 3de1bfdc-6e02-41d3-98f1-9a17c5964c27 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.514961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.514961Z digest=sha256:d5caaa1bc26512b062655700e044e8403084351f41f78458199e464a976b7720

Observation 2fa3dd12-bc54-41ac-8ae1-f4f037053790 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.563665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.563665Z digest=sha256:e68add7b78663b1b51f8a86cb6851fa8d67b47746344c88d9156196f3cfce1ce

Observation f8481452-cfd8-4211-b57b-d2ce89db70bc · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.622776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.622776Z digest=sha256:9115cc806b0af45912a02478fd0974259e976d1d7ab204f24cb2ec0fc897baa7

Observation 98d511ea-937f-486f-8f0b-2f0166c69cb9 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.715027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.715027Z digest=sha256:bfeb72f4805b0da0d10150e3ee4a43cfcd74c7d76f22b0394f3b9d071ea52b63

Observation b1f090d4-785c-47d6-aa1c-6eb6d28552d2 · outbound

This paper cites Visual Instruction Tuning.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.790351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.790351Z digest=sha256:e4481e922dff0dfec3fea6d1d4b053d87a347aaf27d3084a929224e4f98c6656

Observation 699307ac-b544-4401-aa11-366d012c7f29 · outbound

This paper cites Visual instruction tuning, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual instruction tuning, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.899741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.899741Z digest=sha256:2a71c3f57276193b0df3f212e9d4eaacadf98b7ebfd435a1cea4fa9e9a10f8d4

Observation 8e75c140-9583-4936-995e-a5ad9f726e48 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

Task-Aware KV Compression For Cost-Effective Long Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.979162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.979162Z digest=sha256:0b5372cecb6ff5ec160dabd562d36b3750ec4a6eee6336212d13485e657139f2

Observation 7b1cea3d-234c-40a5-8664-57c158420a68 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.084306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.084306Z digest=sha256:b8a1730a0c06db629e4de1234449b8890aada9aa168aeec980fec97d0a8eca40

Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.176899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.176899Z digest=sha256:e59b84d1e6eb2f449335d05e9aa59ae526ebb82a1d4f0a48cec5441457c84be4

Observation 450847d4-afc5-4fe9-9ad1-f79661ed5286 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.221582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.221582Z digest=sha256:caf2c6b527bf48185bc55f4f6eed339b67bd48c301593661fb908f2bf4259bc1

Observation 0589e541-1a20-440f-a926-855b8c3b8573 · outbound

This paper cites Gpt-4 technical report, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4 technical report, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.276888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.276888Z digest=sha256:08b0693df61eb71e9ed3ff054e604f38910c2f104dbb2da58d2d2a30cc5bc906

Observation 6b2c5357-e7ff-4183-bfbd-6f2f9be7bf05 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.041465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:36:38.352178Z digest=sha256:a42aed5f41dc84ae25d8f9fd7a5d00d69cd05ae28b3c48ba7028000efc0ae567

Observation 92e400c0-ffa5-4f75-ac26-451ca4a7d5f8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.387129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.387129Z digest=sha256:0beeb2ea1de92d634b23b4e660bb886a2874c774efdd4d198cd32e59de824043

Observation 69f4681f-c7bd-4a2d-b978-8acdd833a70a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.446461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.446461Z digest=sha256:bc9640decd1b1d93a71e27418711774a435142afaf2c08873770da11f0e2a063

Observation b3b21a56-e1ad-4778-a6ce-c322c3a42042 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.507321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.507321Z digest=sha256:6092814821ae88e361a8c884ebda0f9d9850c98da733a998d0b98d89f3d6a581

Observation b25ca4dc-5cdc-466e-a600-f35ebdf1f7d5 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.567659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.567659Z digest=sha256:579facb6d322956dfa7c4162871239636fdb7818adc80f9feab7320eff3c5ea1

Observation 55d743eb-1d52-4b25-92e5-2c110b521247 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.677891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.677891Z digest=sha256:813ed06146d1b4244c5dacf39cdee67263bfaf9b08738fcce1dd5c441bb57bd0

Observation c0fe3f2f-7ee5-458f-9016-4aeb9d543755 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Task-Aware KV Compression For Cost-Effective Long Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.709295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.709295Z digest=sha256:a523b56aa16a8657cbdd7c143b4749e7d623e4baa2fc4c6587d14952c93b03ec

Observation be7b36a6-61cb-4357-be28-6e974b9a6c70 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.757096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.757096Z digest=sha256:77db36277889f02bc941d017387978db19c5723ea4fbf301f45f9f89462576c9

Observation aad14b21-a422-405e-9593-c7614134390f · outbound

This paper cites Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.823039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.823039Z digest=sha256:d9ee2937dc0baa5f9177ec828a424e30606cb676f520aa0103dcd4b0bc0b5c3e

Observation b69461d4-c78c-4c67-801a-af0204f757d2 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Internvideo2: Scaling foundation models for multimodal video understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.922661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.922661Z digest=sha256:c3670c335ca6973e1b2ce6478dabe7e9a1ed6b5c150313cff989acf424f9aacb

Observation 333a114b-7802-40da-aee7-bd656b079e76 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.968114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.968114Z digest=sha256:6d9331ec38df31529145f0f78df0fe1c6fb1862f1d028362b499a1ed1a975c81

Observation dc1e9d7c-528b-49e6-9fda-5a5a44fc868c · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.027331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.027331Z digest=sha256:ef759af5b6c69ee08e960f7794a503300b2394c745aca4767ea562292f3746b4

Observation 6124ce6e-9411-4742-a532-b56e838e24cf · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.126615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.126615Z digest=sha256:a974659b71191911381df515d5b18e060b9704ead53cb29b5db435fa48d7a4f2

Observation d34835d6-5ad3-46c6-b379-d68751029d8a · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

Task-Aware KV Compression For Cost-Effective Long Video Understanding InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.216006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.216006Z digest=sha256:364caedabb51b421158cb57af621cae97d03b7da02421f2e880ace673d21ffd4

Observation 68ed76ec-112b-4f43-92a6-a1c7b91af129 · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.267392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.267392Z digest=sha256:fae5fa3ff263e1f73e877eabb03a5257ff59a1b0ead9bef9975251ee4649e367

Observation a48fd399-8528-405c-9c19-0202d1e113de · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.329386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.329386Z digest=sha256:25e70c2268d83c4cb8b4bcd75e48d439ce52074a55d10be32a57bc514c6b9194

Observation 122a666a-3a86-4f3c-a925-119f5f0a1358 · outbound

This paper cites Long Context Transfer from Language to Vision.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.426259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.426259Z digest=sha256:b8a050cbe0470a07200d1703a1e5aec0324a90e913c93acc5a7721c90b5c3992

Observation 2ddb7bf0-c48a-4da6-8128-73469e3d4e19 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: A strong zero-shot video understanding model, April 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.503476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.503476Z digest=sha256:a3534f5d5926a393a8ff80fb6c8b1bb82a3def66209df1dadbd93bba6beeddc6

Observation 5360d085-0391-455f-b137-d4d1dc11a146 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.551825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.551825Z digest=sha256:e3c96654b7b2be4315f1ec4281108b2157e555a8e385afa12b84d55230d06ed7

Observation 52688802-2f3f-46f2-8c31-bdd3aad996d9 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.601082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.601082Z digest=sha256:11166910f909ef1294f6924d024761c53ea8659b1668713fcf55e19b6235dc20

Observation 7ef244d8-c636-4f1a-8998-377c3e8507fa · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.692207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.692207Z digest=sha256:29a0ee2ba371c20b6d4f88429f2782b36b79b1f46c96d401d6a426c90bf8e9ae

Observation 03bdf067-a8a2-409b-ace8-51dffb6b97b6 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.772187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.772187Z digest=sha256:a646cba76998d6b88cbb3e6cdb25a7bf444831c38fa27b4b36be0b2f5bb8419c

Observation 0e76717b-f6d5-4173-ba86-00d49d6b5f0d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.829071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.829071Z digest=sha256:1b5ca15627c65cdf5e88b645850f904bf50628e91158afd804845795f1bb166a

Observation 349009e5-9f3b-4f1c-9ff6-80ab18e6d2ae · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.894521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.894521Z digest=sha256:ab570886a580fc7df95ee5b8871459ef69407566dc220713cbf2edd465ba0c73

Pith citing papers

No inbound Pith citation observations are available.