Pith. sign in

Paper Citation Record · LEDGER

Task-Aware KV Compression For Cost-Effective Long Video Understanding

As of 16 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21184.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21184 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:39.894521Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bf89560-8927-4e62-a325-3ac8e3dc034c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.772060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:36:36.457741Z digest=sha256:0fe461dc447173c77c50a95f766bb297b6b9fbb43c163d5fc9ad2b85ac8bd10b

Observation 89ab8ae2-2bac-47c4-9175-999e557b1902 · outbound

This paper cites Qwen2.5-VL Technical Report.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.521460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.521460Z digest=sha256:c9d2ac25fb39c8e63c49899639edb81322b2012542b22c212e1bf8f1348b793a

Observation 3c9cd646-4ba3-44ce-8828-ca0f1aeea3ea · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

Task-Aware KV Compression For Cost-Effective Long Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.613827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.613827Z digest=sha256:c9665bff0a6034d1ac9b3a174e177989dd82f2260678485ee246a9a40758a19b

Observation 636967fc-0fe2-4008-8d22-6ac784cd54cc · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Task-Aware KV Compression For Cost-Effective Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.671820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.671820Z digest=sha256:1c7aecfff6eb9ceaebf2309381aa256071eba0026ccf57fb9d6d66409be56b6f

Observation df79a092-af4f-40f0-8f24-efedd0d61324 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.711180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.711180Z digest=sha256:b9e24e6b5b8d6378814b8f566a39586fba68d7a94a078541ae560e97cc3f1876

Observation 2e9a38c5-4775-4a78-91b0-82d2885203f9 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.783376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.783376Z digest=sha256:50cb1fa09fffef80884bafbf5600461da6443be9b85d30f161297b0c9db92a64

Observation 29079ea3-7f5a-45e3-b673-b96d566aed81 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.868717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.868717Z digest=sha256:11dd46f08a978f9872438cca53b660fefc167d396632a2a1f3662885f9d37330

Observation ad67feca-2ef2-4227-929e-8d25852f59fd · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:36.961899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:36.961899Z digest=sha256:39bc36bd5341cc80420241493273a31f8d6cc4cc03694e26c212973cfa5d11c8

Observation 2800f197-06c6-49eb-bd7d-99a1b40740be · outbound

This paper cites MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.020202Z digest=sha256:66cadda4388cdd8dc14859c6e6e84afe54ac3401f65cdb3d284783bf00a4856a

Observation 2140fb08-a1a6-43e6-a721-ec067c175a23 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.156162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.156162Z digest=sha256:8cf91df63c0866589b6bc89280e42c304aad2ae0f6c8c944bb3c4908d9dca07a

Observation f7a744d3-7a27-4fd3-be55-82d104ffe9b1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.209127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.209127Z digest=sha256:b3d6ccca8f0fa354e6cd889bf4534b5153d8b700eeaa6edca6879f7fa01b49e4

Observation a35b2b65-cf45-416e-b0d5-41cc98cb33d9 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.ICML, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.378601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:36:37.288226Z digest=sha256:ffd0870015e65045c5a1a53f9d0b01b20c983ddf0f838bd8d233675835a2008c

Observation 52d6f597-9a2e-40e1-8491-2f15f6a944e9 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.359101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.359101Z digest=sha256:0c9b9b2e95361c9f87ec9dc01116991cedb70b61a8557a26f9d8d0d1c70c4ee0

Observation e90a2db6-5f6f-4e3e-a0ed-6b1928ac8116 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.422766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.422766Z digest=sha256:abbdd608e7f3a4805af0f6e1ddfd5ed6627f2b90164d64c2d99c0bd152111001

Observation 3de1bfdc-6e02-41d3-98f1-9a17c5964c27 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.514961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.514961Z digest=sha256:08150bb78304202a856f0abba4aebb75df536d7ac63626b55575e971cc092d2b

Observation 2fa3dd12-bc54-41ac-8ae1-f4f037053790 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.563665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.563665Z digest=sha256:1bf4c912949613d5c1dbcbaa4ade01b267706d9c032f85eceedb1b506773c734

Observation f8481452-cfd8-4211-b57b-d2ce89db70bc · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.622776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.622776Z digest=sha256:d6744cb321e7a4e9ec2cb15be330477f6d9431cb5c9701d227d921fb958f2f3a

Observation 98d511ea-937f-486f-8f0b-2f0166c69cb9 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.715027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.715027Z digest=sha256:eb63e7bcd9d1274bde0be1fc2482c624dbbbadb663901ffb9eccb7c4810e215d

Observation b1f090d4-785c-47d6-aa1c-6eb6d28552d2 · outbound

This paper cites Visual Instruction Tuning.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.790351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.790351Z digest=sha256:fda1aa6ce24d3935c235561234f8bc38dcbc7b412e80e24c6e1ef42b6715c691

Observation 699307ac-b544-4401-aa11-366d012c7f29 · outbound

This paper cites Visual instruction tuning, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Visual instruction tuning, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.899741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.899741Z digest=sha256:b40e98d6150bcfbc0f7f4dc1fc46e37bb2f748c01b750596b1a93690eddd106c

Observation 8e75c140-9583-4936-995e-a5ad9f726e48 · outbound

This paper cites ST-LLM: Large Language Models Are Effective Temporal Learners.

Task-Aware KV Compression For Cost-Effective Long Video Understanding ST-LLM: Large Language Models Are Effective Temporal Learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:37.979162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:37.979162Z digest=sha256:2b6d8f62934c51497fe429581b01c9b43a04c8985721ec1fb1ca267447d66af7

Observation 7b1cea3d-234c-40a5-8664-57c158420a68 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.084306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.084306Z digest=sha256:89227901d31cd1381d4a74e5f2fc6e36ddae0e8fca656cfa56b77335c180b9e5

Observation a50540c1-0041-4ace-8dd5-f98bb031f217 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.176899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.176899Z digest=sha256:d8798961aca7cb21c43e813658ee7e7041ac55da8ed44a5391a468c75672911d

Observation 450847d4-afc5-4fe9-9ad1-f79661ed5286 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.221582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.221582Z digest=sha256:f821590760d4a7b011fe08811cacd865964c564957f500867164db4ca878da30

Observation 0589e541-1a20-440f-a926-855b8c3b8573 · outbound

This paper cites Gpt-4 technical report, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4 technical report, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.276888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.276888Z digest=sha256:88d988adbf618c63f73483f8cdb55bebe19f670d74a0fd7e98114320267d60ae

Observation 6b2c5357-e7ff-4183-bfbd-6f2f9be7bf05 · outbound

This paper cites Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gpt-4o.https://openai.com/index/hello-gpt-4o/, May 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:36:41.041465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:36:38.352178Z digest=sha256:387de15cba3a482450a2c4e0e8b372a3d6de81e473e3180e33f7c95112fe1528

Observation 92e400c0-ffa5-4f75-ac26-451ca4a7d5f8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.387129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.387129Z digest=sha256:433ceeb292b075fe32b96a103de0a0e6fe5b19b662a4e98ab278c3b45480750d

Observation 69f4681f-c7bd-4a2d-b978-8acdd833a70a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.446461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.446461Z digest=sha256:ebca29fa26ac72871f1a0beb8ee995038a3c4bcd7b8a57c9f602153e5dbe2715

Observation b3b21a56-e1ad-4778-a6ce-c322c3a42042 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.507321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.507321Z digest=sha256:577061d8015c3134a1eaa4309ca989440653eaa4146195372fd69b419b3c3e97

Observation b25ca4dc-5cdc-466e-a600-f35ebdf1f7d5 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.567659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.567659Z digest=sha256:1ac22d847a24bf3b30f1c7932b4f228284816e405ff9f233ed6807430f02fe97

Observation 55d743eb-1d52-4b25-92e5-2c110b521247 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.677891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.677891Z digest=sha256:281b9d6567106c071c04e59dfc4a06a6b13394d87356a9f770f412034ba3d662

Observation c0fe3f2f-7ee5-458f-9016-4aeb9d543755 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Task-Aware KV Compression For Cost-Effective Long Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.709295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.709295Z digest=sha256:c2ad2bf39d68d892c74a3b8c1013acfa1de896e42c4acfb4fae9193768cd1340

Observation be7b36a6-61cb-4357-be28-6e974b9a6c70 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.757096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.757096Z digest=sha256:a2064b9a3bcce3357fddc1281b1e50cca6a128d6c35fb62954c10bdbc40c8f92

Observation aad14b21-a422-405e-9593-c7614134390f · outbound

This paper cites Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Longllava: Scaling multi- modal llms to 1000 images efficiently via hybrid architecture.arXiv preprint arXiv:2409.02889, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.823039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.823039Z digest=sha256:951763b5454f111b02ae67d20dd1eb754126329f4640a454df6af6df1d2d06f2

Observation b69461d4-c78c-4c67-801a-af0204f757d2 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Internvideo2: Scaling foundation models for multimodal video understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.922661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.922661Z digest=sha256:c11acee173cca2d535cfeda3d6371c67cff823ccde63400851152ede96458f6d

Observation 333a114b-7802-40da-aee7-bd656b079e76 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:38.968114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:38.968114Z digest=sha256:7b4919d3a7eecc897b44e979f4bf9a3cfb30d95096cd666e1913069bc729b650

Observation dc1e9d7c-528b-49e6-9fda-5a5a44fc868c · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.027331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.027331Z digest=sha256:93759ba58e1afc2005ff03d02b4ca6457fb3d64922faf184758db04d20928eb0

Observation 6124ce6e-9411-4742-a532-b56e838e24cf · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.126615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.126615Z digest=sha256:149b715692011c939bc7c4d993e14d4da41c9da9953cd204503be9b278f9f3c1

Observation d34835d6-5ad3-46c6-b379-d68751029d8a · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

Task-Aware KV Compression For Cost-Effective Long Video Understanding InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.216006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.216006Z digest=sha256:a4a0de83f1cafba51703ef40218ecb5204c4a9eb11e55d18554989e7ce52f09c

Observation 68ed76ec-112b-4f43-92a6-a1c7b91af129 · outbound

This paper cites VoCo-LLaMA: Towards Vision Compression with Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding VoCo-LLaMA: Towards Vision Compression with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.267392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.267392Z digest=sha256:0691f8f350e69ddebde3cf196b0322126e060f3a2e8176d569a283a9deab7ad7

Observation a48fd399-8528-405c-9c19-0202d1e113de · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.329386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.329386Z digest=sha256:d780a5c84a70badcf9337b9b19bc0d6d19bfbc7b41bf2d166e7a1056aa604722

Observation 122a666a-3a86-4f3c-a925-119f5f0a1358 · outbound

This paper cites Long Context Transfer from Language to Vision.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.426259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.426259Z digest=sha256:74e94ab68d0cce66f5f9fe4999caab99ae887c4ff1eb18d8a4bda9616cddf80b

Observation 2ddb7bf0-c48a-4da6-8128-73469e3d4e19 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Llava-next: A strong zero-shot video understanding model, April 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.503476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.503476Z digest=sha256:7ccb61d1a901576e4564b08af969a6ba87477878a20aa366af138b0bae4bfcf5

Observation 5360d085-0391-455f-b137-d4d1dc11a146 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023.

Task-Aware KV Compression For Cost-Effective Long Video Understanding H2o: Heavy-hitter oracle for efficient generative inference of large language models.Advances in Neural Information Processing Systems, 36:34661–34710, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.551825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.551825Z digest=sha256:577f99409765dd345b21061592d95fbe7c9de7150e17c31d1aa103feb4a3ad9c

Observation 52688802-2f3f-46f2-8c31-bdd3aad996d9 · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Task-Aware KV Compression For Cost-Effective Long Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.601082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.601082Z digest=sha256:73d870e4b2e659ec415e4bac2f7034f94da5e97dca492e269b1a97ec0fe16375

Observation 7ef244d8-c636-4f1a-8998-377c3e8507fa · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.692207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.692207Z digest=sha256:8a08ab1871361eae134626a747ba7de822b6c8464110ccc6081b46770a0a5bb7

Observation 03bdf067-a8a2-409b-ace8-51dffb6b97b6 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Task-Aware KV Compression For Cost-Effective Long Video Understanding LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.772187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.772187Z digest=sha256:8844f8d90f16f4a0c40ffc7811c2fa3b77d96eafa94e9b46a0de2be523fcdbda

Observation 0e76717b-f6d5-4173-ba86-00d49d6b5f0d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.829071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.829071Z digest=sha256:bf0438b010aa2285fbd727ac254d6497fba7b14f41be29cffe715e47ddea2191

Observation 349009e5-9f3b-4f1c-9ff6-80ab18e6d2ae · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Task-Aware KV Compression For Cost-Effective Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:39.894521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:39.894521Z digest=sha256:4578683ed521a76583cd7eb59269b2bda050836137c342c7e5f6054a546fd84d

Pith citing papers

No inbound Pith citation observations are available.