Pith. sign in

Paper Citation Record · LEDGER

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.15689.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15689 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:38:47.340935Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9d699af-873a-4b22-b63d-7e2243652ff4 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.660878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.660878Z digest=sha256:f29492e77fafe026255a81e1d83594702739764f7e899d63a0e8d52456c69165

Observation d126ce7d-31e6-4b21-8589-6d1ffbc98e4f · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.714721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.714721Z digest=sha256:1c627eb3e3871a114bae1686b2378397661932630a943f314d6994d202feb0ad

Observation dbf9d22f-f5b7-4b34-8c67-4acc3f614be7 · outbound

This paper cites Science China Information Sciences , volume=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Science China Information Sciences , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.772008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.772008Z digest=sha256:03147db8d5508012e5cd1378158d07eac3a1bfe5ae7037c9d4e19d6c50f4d4e7

Observation 982faadf-2023-46fc-aba7-c4cd54575a44 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.828170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.828170Z digest=sha256:e18dca177a4af11f43c60c0ba4b644c05a0a620ef1f8023f079fd9db5fa10b96

Observation 7e6805ad-acac-40b5-9cfb-2befdf91bfa4 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.887011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.887011Z digest=sha256:e11c9dd597568e23cf21e7628c0bb5875ee30356577ed61be1e6da2006a44244

Observation 3f1fa2b7-8292-45f3-96a8-fe26957e6842 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.936578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.936578Z digest=sha256:fc10af1ad2dcadea2b7193bc992b0e982476ea3206d6bc4fde57d51e125b0fb0

Observation 16129752-a127-47ff-a819-6c64e0f6fe8d · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.003198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.003198Z digest=sha256:7b2a0feecec6cff9d2981aa93f161c4d10573cdd337922665df9ac2ce6517c1f

Observation 486bd35f-d07e-44b0-a4db-4a328a3b95d8 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.056803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.056803Z digest=sha256:f38abbd4a50168d57c8e218d2e2c05db99e62c3101c39b3e36351e7787c7058e

Observation 43fe9119-866d-48ab-a661-8e100b2f6cc4 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.121515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.121515Z digest=sha256:e33c207f2c56030cfa8c872dfa2496d6eb69b01647fb4d234f4368cef621124e

Observation 2bbef040-3c0b-47ca-98da-46922382d900 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.181503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.181503Z digest=sha256:d5fb7e114822f4d365669b724b71674b6fde0cbe02e98e4ecac7d1e684ff0ba4

Observation 35115adc-2044-40d7-869a-d71058c4bde1 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.230180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.230180Z digest=sha256:ecf87458cd5788e5a304be036ae3b493554f657f7e835926d7a3a21800d0281b

Observation 9d159b4b-e366-4cc5-8ef7-38491711865f · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.291857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.291857Z digest=sha256:aa2f8b1376788f71c9db1ec878cc4b4e592eb753abddf2f92b2d029bb44d83c0

Observation b75f4b92-3907-478c-bcfe-b74fa1f47dc4 · outbound

This paper cites Kwai Keye-VL Technical Report.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Kwai Keye-VL Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.344173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.344173Z digest=sha256:4b7269f45c2440fc05256c56474514a51f03163e782d6635e9b1e0408048399f

Observation 551218d4-c04f-439e-a53f-3bfd87281b30 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.502073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.502073Z digest=sha256:150dbfc6af9c2f539387d9cefbbcbd5ef5a628815128dfe5e5b8dd3f13625d2f

Observation 023e1e81-9023-4cee-9360-1e5ca636e1a5 · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.688411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.688411Z digest=sha256:9f6d5a1c96a39b306aec223e491a7b33e2d6c64dab29d2772368cfdb85e9c7c2

Observation 542ba471-83a9-47ca-a073-4a16e8c2cb39 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.912377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.912377Z digest=sha256:d6d96c72a22896751125686747cd3369a31adb9087d32445f3cc49df89031cb9

Observation a6bc7367-482b-473f-8b13-5e7a81be1241 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.040367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.040367Z digest=sha256:6b20b243515d19288242c9464a55813a0e83d1649b23c42829b7223c5109f8e1

Observation 9c78edad-fd1b-42a5-8d4e-8cf73009e44f · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.147436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.147436Z digest=sha256:f30745358a6be57446fe18e374f90bc839a6fa2a3a9d7964d65fd3a87eb75768

Observation a7eae351-3594-405c-b82e-254fb07595ec · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.203874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.203874Z digest=sha256:2f2234bb3b6ecb371fc8396c4828fa96fe0bea652bd6f440cfaeebdc36263190

Observation 732a5076-6ed2-4016-b11f-da781c3c7154 · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.297703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.297703Z digest=sha256:3f5d56fead98257b8d81ff9b29598dcef8eb10509cd6058a0bebffd03c966292

Observation 3638ec97-6a7f-47c4-81fd-d04f151aeebf · outbound

This paper cites ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.510039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.510039Z digest=sha256:15cfb378009538b74920c78c6e341075c9e9c65f25b0d81b172fcb6d08cced77

Observation 0b37c1b7-b82d-47ba-88ac-a0f415b44d24 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.617246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.617246Z digest=sha256:2539b2b7c148e83f80631c0807ed502f2a15bd654e508138a5ae4d4ba6c85bc7

Observation 7ab06932-f2dc-42a1-9b7e-ec6315e31dcd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.732788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.732788Z digest=sha256:966d281d98f0b620c3488956e129614d7951020cfba2f9a98d8c140299f27119

Observation 74d5455b-21e8-4602-a0c4-edff0915cd0e · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.902744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.902744Z digest=sha256:c39d29be93adc3284faf11a030a76b7dcbe6dd4071e0821910a7d5539218bfa7

Observation c6a5feda-df2f-4f5c-a1a4-34270efe751c · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.131300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.131300Z digest=sha256:b017e9246f38e1bc65e268c5b66d26602e54087003b4478cf39eefb315e81ce1

Observation ae03ae1c-57e0-4df6-9099-5ea9281c1d45 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.225953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.225953Z digest=sha256:5d029a8ff3082e4dc153b470d4a28c662f357de7d666539a36c93522da814847

Observation 14407740-561b-4ab0-8e05-8d21b0a16b86 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.340870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.340870Z digest=sha256:c06f56936fd60a10db09d57c9f3ac3432ce51f1a35489ba6d8a98a0f8f6480e0

Observation ab762f98-691a-4e91-a1f8-b064bbe9fa96 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.445751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.445751Z digest=sha256:019d7fa9ce78e749cb86866d82548fcaec4048bf3cdf9e7f640f1bb2acf915b6

Observation dabbd4ca-4291-44c8-952c-7d0454feb6e8 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.609118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.609118Z digest=sha256:9975731165cd831558909e4cfdb61300d51c250f1f87c24b5d5c463b4ec38de8

Observation be02a39a-0481-468d-848e-ec29f764cd70 · outbound

This paper cites 2024 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2024 , eprint=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.752176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.752176Z digest=sha256:aa1fc25ac2f14c1f170499f103bbf0cde942da4e95662e4878124ff9e37b1810

Observation c998c41b-a74a-43d2-bfa7-8b7ac441124c · outbound

This paper cites LMMs-Eval: Accelerating the Development of Large Multimodal Models , url =.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LMMs-Eval: Accelerating the Development of Large Multimodal Models , url =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.896998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.896998Z digest=sha256:77465ed4a088709fb1dfc5b355a44207d3cfaf3545a3c33b041f8a2cd48b3097

Observation 2856f8eb-a084-4664-befd-5bb1241ca836 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.017110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.017110Z digest=sha256:de7d6b1c5377cae0e483f18606d517a6b4cd2fec73629e7ecc80832666748aad

Observation 1a5f4efb-2082-46bb-b31e-168becc202de · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.121056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.121056Z digest=sha256:4e739f21df3d0dfc7a4874065ebd5464643694537c8374c098a28358568ac197

Observation 5baadcdd-3acb-44b1-8c0e-f4ff8be78c9a · outbound

This paper cites 2024 , howpublished =.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2024 , howpublished =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.217906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.217906Z digest=sha256:ec4d8e87f2e78dfcb420fc5875c09fdbc9f973bef8f40d481ca6692e983324d8

Observation ea9794b4-b136-4d67-baa3-cbc6469ab891 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.297149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.297149Z digest=sha256:c2b48ca14efa713b25f7767f181907f5fc260fd9801bd43e34e6c6b09f2ed6f8

Observation a2a12512-a6fd-4b8b-a593-cec6c9014f1a · outbound

This paper cites Qwen3 Technical Report.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.400837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.400837Z digest=sha256:9af47833e4a1922e6fc953f1ddf2fbefbdc3b0d74451f58e90a4170b3995fdf5

Observation 39f26be3-53e9-4c38-b469-70f7db0bff8e · outbound

This paper cites 2025 , eprint =.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.508334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.508334Z digest=sha256:1b28126c4701e2dff8c827ac7d148fa81c5f39d19bdb55ba63abdd7becdfe758

Observation cbe4b01c-be6e-410f-b62f-cbb0ef346cb9 · outbound

This paper cites 2024 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2024 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.657730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.657730Z digest=sha256:591c426dbea69823b1e36be4fffa909307d8a6ee0cc828589c9b58a6cf829c9b

Observation 59620e5e-8656-4920-9a29-6bd6d3c46d32 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.828964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.828964Z digest=sha256:54e96e50014abd89e4828395d23e61710a2ba1b5aa4c9b57ed2af23d3000799b

Observation d014b4e7-dd97-4ca6-bfba-9d026d61ca29 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.986561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.986561Z digest=sha256:59d17cfc52052befbca19d2f10b1e59595379ea46868db8f2904a0fd0480c5ea

Observation be6c229a-dfb1-47e4-84a9-823ea7d0fc5e · outbound

This paper cites European conference on computer vision , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors European conference on computer vision , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.096044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.096044Z digest=sha256:3f94e9c8f25ee256cee67184ac2c3849f23beecbb62021918435819fbc735f6a

Observation 08fc3bc3-eab0-45af-a13f-60e86cd4ff4e · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.251209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.251209Z digest=sha256:1f2101a8e9d8098349c6e13c74b659336ef03bf4ab16b2aad4d7f8edc0151c03

Observation 8df1a760-9674-46c4-886d-16f17f13170c · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.380037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.380037Z digest=sha256:0e7644fcfab48a551caa899460924df45859b87adc33ea5d2bb5bed192b67c79

Observation 3f3101e5-e3c1-4138-9464-d00e20841585 · outbound

This paper cites GPT-4o System Card.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors GPT-4o System Card

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.520050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.520050Z digest=sha256:e741f479eabcd85bd225ea570fdddc70fe4159eb0b1dd7972702e2e2ceef7b13

Observation 0f599655-59fa-4af4-b30c-c0633ce361cb · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.621136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.621136Z digest=sha256:4e2d31d4287e9474b56baf176136586b317217c1cb78c7237c24148b89ca5382

Observation 3e637cc5-662e-4c33-b1d2-0dc4bb11fce2 · outbound

This paper cites 2023 , url=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2023 , url=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.718466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.718466Z digest=sha256:72f745e1dc4411254b7daf361b0864f3fdbc338ba55d2cf96e90465f09d9f46e

Observation 50af191e-c0fd-4646-8cfa-999ad45eab79 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.867016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.867016Z digest=sha256:cb7c913100307b47ffec6a8026c06afa813f67422698ea1142ab0bb0361e9538

Observation d5b3bf85-6798-4be6-81e1-ee9a3bbccaae · outbound

This paper cites 2510.04428 , archivePrefix=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2510.04428 , archivePrefix=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.980296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.980296Z digest=sha256:d3cbfd040626d8ef57d0b25ed2800afb0b09402f999fb6cbd756fd6e056ee9e4

Observation 6130b671-47c6-47e8-bfee-7e53335fd434 · outbound

This paper cites 2603.25072 , archivePrefix=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2603.25072 , archivePrefix=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:47.110495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:47.110495Z digest=sha256:2c6b27c48d5c1821d325628ee4b7062caaac9cf6c1fb85b8aac7a1ff6080f8cd

Observation dae80eb8-8c3e-46e5-82af-631ce4a86c8b · outbound

This paper cites 2602.22932 , archivePrefix=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2602.22932 , archivePrefix=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:47.224513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:47.224513Z digest=sha256:e8030cd15a4c93675cfc63bd6b9dc77b571043e5410ace80c6223d17e17ccb6a

Observation 0029752c-f2c1-4cd8-bb54-ae641cb0010e · outbound

This paper cites GPT-4 Technical Report.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors GPT-4 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:47.340935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:47.340935Z digest=sha256:06cdfd0296b890953601037191fc75b7d75e6dc47cb5df7668ad5cdc211102fe

Pith citing papers

No inbound Pith citation observations are available.