Pith. sign in

Paper Citation Record · LEDGER

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors

As of 15 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2607.15689.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15689 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:38:47.340935Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9d699af-873a-4b22-b63d-7e2243652ff4 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.660878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.660878Z digest=sha256:bb6ea65641b5c01ca1c97ce4d02e42e716db4c455172fef47e8959f8bdb548a9

Observation d126ce7d-31e6-4b21-8589-6d1ffbc98e4f · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.714721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.714721Z digest=sha256:a1dadec363e86d842fb55218fcabce8dc84045dc5e3a89505dfb588670ddf5f1

Observation dbf9d22f-f5b7-4b34-8c67-4acc3f614be7 · outbound

This paper cites Science China Information Sciences , volume=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Science China Information Sciences , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.772008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.772008Z digest=sha256:4c3af51d1fe006a67892d86276bede539abf55019751bc5600d9413f5fa99394

Observation 982faadf-2023-46fc-aba7-c4cd54575a44 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.828170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.828170Z digest=sha256:eac1743e13df1b0a234fe13a5c53784708fe404c2c3f4319149fdf82cdf36bd7

Observation 7e6805ad-acac-40b5-9cfb-2befdf91bfa4 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.887011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.887011Z digest=sha256:1d430ab61ae6732bcd303408c65460449f391657b1d660d8e15cacc719ec1b47

Observation 3f1fa2b7-8292-45f3-96a8-fe26957e6842 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:41.936578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:41.936578Z digest=sha256:18608da5023345217b67f8f99d8bb77b87f6ccd5dd0a9a6e7639501eecbef402

Observation 16129752-a127-47ff-a819-6c64e0f6fe8d · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.003198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.003198Z digest=sha256:adffa2d3006394ae8faa1fa3f67919e545fcce139c5b763ef6979ef0523185e6

Observation 486bd35f-d07e-44b0-a4db-4a328a3b95d8 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.056803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.056803Z digest=sha256:eba67a4e457f0a9eb69399284c3ba43eea3e3c28e05651546c65992a509fc0cb

Observation 43fe9119-866d-48ab-a661-8e100b2f6cc4 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.121515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.121515Z digest=sha256:e0932c7225bf9070278b4aee1f546235fc4d74042c67c91c1d65f1bbd1af82a2

Observation 2bbef040-3c0b-47ca-98da-46922382d900 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.181503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.181503Z digest=sha256:843dec8a611fb8021c680c0481140425725fa46d12749f08c830db20e8997c40

Observation 35115adc-2044-40d7-869a-d71058c4bde1 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.230180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.230180Z digest=sha256:33f252736f92012a2deb1a974c46ab828c87dc380fd933f078dbccc9c307da0e

Observation 9d159b4b-e366-4cc5-8ef7-38491711865f · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.291857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.291857Z digest=sha256:184ff995299032b9ee209cbb34d4bdc26caf2c67845432a48c101f48330c287b

Observation b75f4b92-3907-478c-bcfe-b74fa1f47dc4 · outbound

This paper cites Kwai Keye-VL Technical Report.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Kwai Keye-VL Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.344173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.344173Z digest=sha256:409c4d6cd140d4b99d901ed2326d8436b6ba99d53c884095efdf1c05a7915f56

Observation 551218d4-c04f-439e-a53f-3bfd87281b30 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.502073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.502073Z digest=sha256:fac85fd12412e4b303c13b327108a186b8d0099406bd8661aab82423fa173408

Observation 023e1e81-9023-4cee-9360-1e5ca636e1a5 · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.688411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.688411Z digest=sha256:2f295011a3a079334b505e6f7f5fbf4dbba2983d1ec98cef20390f3e7a550464

Observation 542ba471-83a9-47ca-a073-4a16e8c2cb39 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:42.912377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:42.912377Z digest=sha256:1151b973c3a27dbc6b4129c479208595f37f200a7b44f7a5b13be34f2ad9054d

Observation a6bc7367-482b-473f-8b13-5e7a81be1241 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.040367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.040367Z digest=sha256:e26d35e7fed1a309511f5333c9571469e9bdc1908b086143ca52630f48273c5b

Observation 9c78edad-fd1b-42a5-8d4e-8cf73009e44f · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.147436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.147436Z digest=sha256:f2ee2574c032c2518ccc2c7b67f1687ebed52b5b4fafb6e06a13749b405b2cd3

Observation a7eae351-3594-405c-b82e-254fb07595ec · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.203874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.203874Z digest=sha256:b9241056874ac0b6879a8a024e390c99d1006f462f76b751c580c70600112b0f

Observation 732a5076-6ed2-4016-b11f-da781c3c7154 · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.297703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.297703Z digest=sha256:7ebb0516321b859d58772bb097b0d375735b575fea6bfc8cd0d4298f42d3a84b

Observation 3638ec97-6a7f-47c4-81fd-d04f151aeebf · outbound

This paper cites ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.510039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.510039Z digest=sha256:7df0609b737eeb80839db9cfcf21ca28f1b33c61a0e5a0ef695d73bb560b95e4

Observation 0b37c1b7-b82d-47ba-88ac-a0f415b44d24 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.617246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.617246Z digest=sha256:4595025a815cfa4e5679b1528b7cfb2e972bcbec705f992368d26dcef043825f

Observation 7ab06932-f2dc-42a1-9b7e-ec6315e31dcd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.732788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.732788Z digest=sha256:6f94d85cacbe9ec8854b8be9b99d375c0dd89375af8b6fea8a2b6c5e4debf74c

Observation 74d5455b-21e8-4602-a0c4-edff0915cd0e · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2025 , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Findings of the Association for Computational Linguistics: ACL 2025 , pages=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.902744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.902744Z digest=sha256:c148e9f6c68df9a501076c4af604f37ff89d220338581535d80c768853c385cd

Observation c6a5feda-df2f-4f5c-a1a4-34270efe751c · outbound

This paper cites Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.131300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.131300Z digest=sha256:7790c58164b6effb009b76d991344ef295a81d7965fa93d86bc2e744a2486868

Observation ae03ae1c-57e0-4df6-9099-5ea9281c1d45 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.225953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.225953Z digest=sha256:0bbe6231be0d8a369763c4073f15c67bd5f2a7c1595263cfe96eed74c2820767

Observation 14407740-561b-4ab0-8e05-8d21b0a16b86 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Advances in Neural Information Processing Systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.340870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.340870Z digest=sha256:478383148bec07121208d75a1f5161250a5dff3e2fa8c05ff916a3c7538f4a61

Observation ab762f98-691a-4e91-a1f8-b064bbe9fa96 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.445751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.445751Z digest=sha256:7c98c6cd652d851349e2cd04b35081bbe2d9913ca2c21d12fecb4101f9469d9b

Observation dabbd4ca-4291-44c8-952c-7d0454feb6e8 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.609118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.609118Z digest=sha256:ee2aec44b0eddefbec114f79dec0236ce383e52995e549880b75580d757f95d8

Observation be02a39a-0481-468d-848e-ec29f764cd70 · outbound

This paper cites 2024 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2024 , eprint=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.752176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.752176Z digest=sha256:3fa2c88d2ba508a8471869a91cccaac52cc57ecaab3c501d048fd1f9d6c6b526

Observation c998c41b-a74a-43d2-bfa7-8b7ac441124c · outbound

This paper cites LMMs-Eval: Accelerating the Development of Large Multimodal Models , url =.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LMMs-Eval: Accelerating the Development of Large Multimodal Models , url =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:44.896998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:44.896998Z digest=sha256:572705ec8133f05540787a34a4238a96097a402e535bfe05b9682a6aa40beb9b

Observation 2856f8eb-a084-4664-befd-5bb1241ca836 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.017110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.017110Z digest=sha256:006c894a3a2690c07f498e0370287299498ad71315af733ca7db60adb957f05b

Observation 1a5f4efb-2082-46bb-b31e-168becc202de · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.121056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.121056Z digest=sha256:d4589237bc4c9764b98e3f1490b3b4ce80efef77a83ce4ed0d9fbbef5bf810e3

Observation 5baadcdd-3acb-44b1-8c0e-f4ff8be78c9a · outbound

This paper cites 2024 , howpublished =.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2024 , howpublished =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.217906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.217906Z digest=sha256:e9d696963be5b17e7efdc56e08a2fbadf002592dff677a55145f6336ee19ec82

Observation ea9794b4-b136-4d67-baa3-cbc6469ab891 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.297149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.297149Z digest=sha256:f71a30796c4d06d8b13e3181da02e73524460217f594381f2f8e086b220fb6ae

Observation a2a12512-a6fd-4b8b-a593-cec6c9014f1a · outbound

This paper cites Qwen3 Technical Report.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.400837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.400837Z digest=sha256:7aea0b566d7d47f664c260c1f72023e680c0e250e875c239b62ebd0869812f7f

Observation 39f26be3-53e9-4c38-b469-70f7db0bff8e · outbound

This paper cites 2025 , eprint =.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.508334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.508334Z digest=sha256:848473df1a1eb692fa0ff56d2d3a58d67448c17e5426cad4387459c4486f15a8

Observation cbe4b01c-be6e-410f-b62f-cbb0ef346cb9 · outbound

This paper cites 2024 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2024 , eprint=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.657730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.657730Z digest=sha256:8f1eb9a2a368398bbe1289560e1a693f643846607a74f3455dab91eeaef5323a

Observation 59620e5e-8656-4920-9a29-6bd6d3c46d32 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.828964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.828964Z digest=sha256:08a08ee0376e0a00ed94fc7d9ea32e3549847e5f4482c054a425f93503ea6355

Observation d014b4e7-dd97-4ca6-bfba-9d026d61ca29 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:45.986561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:45.986561Z digest=sha256:cb5bc60ff03fd2306a8b1b5c07c5953a779403b5453d2ed92d9497451b82e249

Observation be6c229a-dfb1-47e4-84a9-823ea7d0fc5e · outbound

This paper cites European conference on computer vision , pages=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors European conference on computer vision , pages=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.096044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.096044Z digest=sha256:705cde80a63a0f2a264344f9b7726cea2cd1e8bbefab52f98ce7ace19284d049

Observation 08fc3bc3-eab0-45af-a13f-60e86cd4ff4e · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.251209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.251209Z digest=sha256:453cb10f667b1456f2f922eb02942a9a907c71e570229a96cde1778e86de851d

Observation 8df1a760-9674-46c4-886d-16f17f13170c · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.380037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.380037Z digest=sha256:53efb2e20aede1b9864ffb7934a39d561e2cb3d24b68b325f20e7dd764adf743

Observation 3f3101e5-e3c1-4138-9464-d00e20841585 · outbound

This paper cites GPT-4o System Card.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors GPT-4o System Card

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.520050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.520050Z digest=sha256:8c8825b2516ea927a1386eb8957486a00283807e48d9dae685ad0d6182de5c06

Observation 0f599655-59fa-4af4-b30c-c0633ce361cb · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.621136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.621136Z digest=sha256:a9f0aa2c04378210a2c03a2ce4a421389f320c07714aa5a2f0b4c6a5347ba0fc

Observation 3e637cc5-662e-4c33-b1d2-0dc4bb11fce2 · outbound

This paper cites 2023 , url=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2023 , url=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.718466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.718466Z digest=sha256:181239f2807f67478e3096aac6871e622a4cd61f227836ff93d0dc33e86438e2

Observation 50af191e-c0fd-4646-8cfa-999ad45eab79 · outbound

This paper cites 2025 , eprint=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2025 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.867016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.867016Z digest=sha256:ed27c9582299a3b8b360bd5019026d50a70d4d3eb33f8d00d783c195856b463b

Observation d5b3bf85-6798-4be6-81e1-ee9a3bbccaae · outbound

This paper cites 2510.04428 , archivePrefix=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2510.04428 , archivePrefix=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:46.980296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:46.980296Z digest=sha256:beeaec0ddf9367a029bc83e301da09a01cd1d8d35ca10de0ff41dbd5fd5cf53a

Observation 6130b671-47c6-47e8-bfee-7e53335fd434 · outbound

This paper cites 2603.25072 , archivePrefix=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2603.25072 , archivePrefix=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:47.110495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:47.110495Z digest=sha256:51a124ed37042bcd09477e1b30dcb4b36f287ca48c0cff2794e371b1bc9c596b

Observation dae80eb8-8c3e-46e5-82af-631ce4a86c8b · outbound

This paper cites 2602.22932 , archivePrefix=.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors 2602.22932 , archivePrefix=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:47.224513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:47.224513Z digest=sha256:b1340bf66d7a9e84474f3e0c2cd077c707cdb5598e8da9564d4fe357d693d15c

Observation 0029752c-f2c1-4cd8-bb54-ae641cb0010e · outbound

This paper cites GPT-4 Technical Report.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors GPT-4 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:47.340935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:47.340935Z digest=sha256:2f3a90dcdf83af8bf03ee246edf89bd627f303878b01e9f6d2738de600556cac

Pith citing papers

No inbound Pith citation observations are available.