Pith. sign in

Paper Citation Record · LEDGER

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 5 inbound Pith citation observations for arXiv:2506.00993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00993 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:11.354129Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.141746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T10:11:28.172880Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved42
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2b7e372-c127-4f3a-90d4-efd27c3a854d · outbound

This paper cites Qwen2.5-vl technical report,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5-vl technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.319800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.319800Z digest=sha256:0d2f447368c7d673f7471a21a7ba1896573b686fc9cdb8e6f9930ed2b10cca80

Observation 142e73d1-bebc-49bf-aba7-c98074de1e36 · outbound

This paper cites Fast Differentiable Sorting and Ranking.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Fast Differentiable Sorting and Ranking

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.529141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.529141Z digest=sha256:585ca44b484b87eb719e0fe283c3850e4fcf1b5d9eab0b8965a34cc3c125bd47

Observation 0c9fb958-105e-4065-a867-de5b55fd6aa6 · outbound

This paper cites Principal component analysis.Analytical methods, 6(9): 2812–2831, 2014.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Principal component analysis.Analytical methods, 6(9): 2812–2831, 2014

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:13.275919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:05.644985Z digest=sha256:083c513b5c2106aad1344fb52eca754117ace1429e45a7c272702ff7fdeaf807

Observation ec529799-0e10-4035-b845-b9cc040f1b06 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.780757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.780757Z digest=sha256:1a8e04df9178356fc490b88e485099a6262aead5673b03e198f56da794181cdf

Observation 6fbfa54c-6909-484a-814c-ee2ecabe4087 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.898445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.898445Z digest=sha256:b844f5a1bba356bc51ae62f32c45347eebbfbb702984bb2ba9fd70a876e7458f

Observation 765a8170-44c3-4b5f-8f69-cb30a2eebca8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.001743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.001743Z digest=sha256:e75ab4b5bd4caa56cf8267705535a67c1cc424da31a66aff80fdddc825bb8f8d

Observation d536abaf-3295-4fbb-a432-666bc4a20e34 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.098310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.098310Z digest=sha256:cd47f2aef53834e52ea57b1e083c5a15ca30b16638eed4e1fa07f26dcbc2cf00

Observation 4a6a7819-c052-4f17-a2a0-c9e20b22d499 · outbound

This paper cites ReWind: Understanding Long Videos with Instructed Learnable Memory.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding ReWind: Understanding Long Videos with Instructed Learnable Memory

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.204066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.204066Z digest=sha256:7539abf3797ba910ca0dd427ab0f952d2bfe27248c169065a86677dda09808b6

Observation 0cf1cc51-d3b1-4ccd-95ef-a55c6c0a8e66 · outbound

This paper cites Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.327785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.327785Z digest=sha256:1ca68ddc035b1441216f92acaed59cff19094eade9284b5553ea63024e3a7995

Observation 889779c6-2937-4cab-bb15-295b9ff8aeea · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.421786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.421786Z digest=sha256:e4aa85f75efe74274231067f167a8d09ad7eddb4ef4e3129dab30967dcfa4107

Observation edb3e71d-43c9-4f30-aeab-8340890c5497 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.561620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.561620Z digest=sha256:a594edeedec49221ebaea25aaa8c7966f11858dc21baf7d36a0e1fb20e731c6e

Observation b1499470-2857-43b4-9b4a-d9b606b4758b · outbound

This paper cites Framefusion: Combining similarity and importance for video token reduction on large visual language models, 2024.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Framefusion: Combining similarity and importance for video token reduction on large visual language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:13.133611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:06.662818Z digest=sha256:f23f0ff50db2d5e77ba1d1966c688019152c60e48e62daeaa3032cd76997210b

Observation 7a28d731-d390-4e5c-9c8d-1465ae651bdd · outbound

This paper cites LinVT: Empower Your Image-level Large Language Model to Understand Videos.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LinVT: Empower Your Image-level Large Language Model to Understand Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.759496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.759496Z digest=sha256:eeea32cc9b22810dccc987ec4802cd3f58325f263039c55ff1f5c63d153e7afd

Observation e2fd70d1-368d-4c2e-a2d5-30a39a559de3 · outbound

This paper cites Llava-onevision: Easy visual task transfer,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Llava-onevision: Easy visual task transfer,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:06.898702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:06.898702Z digest=sha256:be1a14bb06f5a10385e9048f5c5fe58f9951827aeaf9b16859ad9682c146593b

Observation 21cc2dfd-dfd4-44a1-b1c5-5e5e52635d85 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.178112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.178112Z digest=sha256:1732f7e89dfab3edc3781f437e69508ef6d40ed587f6414b42c732783d80d14c

Observation cade06a3-70a4-4910-80ba-31e38fac5da2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.303609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.303609Z digest=sha256:8d24e127db0194fcd3c08fcb10a4bf2bcd98f6ba496f8ba58a9ec488147583ae

Observation c0f4db96-2152-4a32-8749-5ef93fdf9da5 · outbound

This paper cites Temporal preference optimization for long-form video understanding, 2025.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Temporal preference optimization for long-form video understanding, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.981096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:07.398143Z digest=sha256:e2c34c80b24e1f0833e4a2d7f31136f1cdd8112d77f4e88ea5e41878a2e873ac

Observation 02ef2d55-2f23-47e2-b7f1-85fa1b1795f0 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.537007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.537007Z digest=sha256:beb8914c17c0bd767fc9d238f9f560fca25c7eb37faa70987f07a0b94beb240a

Observation 45e15b40-6cc5-4825-8b1a-1dea06352fcd · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.645844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.645844Z digest=sha256:fb7d9a8c3663fa3e7b0ecfd5fa71f85f75b495a0fd312ed7246692c627ffe314

Observation 8b3bc549-41c2-46f4-95bc-16ad15f0b52f · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.742693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.742693Z digest=sha256:39b3fb58c2d0be802ad72fe3da7eef31b10466dcc58e994750fb34940da58c68

Observation 53815c49-612f-4fdc-ace2-72e165abdbdc · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.844726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.844726Z digest=sha256:8f21f42dbae22846c4d4e0d8fd34fd70cfd258d14c69173f66ea5c1447dc349a

Observation 76f94cf2-f665-450c-bbba-a82b44d337a1 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.957257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.957257Z digest=sha256:ef36df010bd6a46ef5447717ea15311ac969bb14c15fcc6565ec5c1fa167bb02

Observation b223a9c7-6b65-4bea-8703-b7752342912c · outbound

This paper cites QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:59:11.720174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:08.080043Z digest=sha256:c1ab379596b0fa962882e09944379e49d90b0e9391f538f7b433edb9a1a62b0a

Observation 8c88a420-dec8-4220-ae5d-15760d672cf5 · outbound

This paper cites Hello GPT-4o, 5 2024.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Hello GPT-4o, 5 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.829537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:08.188771Z digest=sha256:eb144999e997899513f11978f6a12144c00ecef6573109bb14078afff1a16591

Observation b5051d15-0865-43d9-a9a2-a195f343f37b · outbound

This paper cites Yarn: Efficient context window extension of large language models, 2023.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Yarn: Efficient context window extension of large language models, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.565984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:08.296154Z digest=sha256:35fa530191f1f5557370fcc8cc032c6417c4e13a4acdb2ab71c3c6dec8db02f8

Observation 44eec4db-7d7a-43fd-86e8-4d3f6f8c8d70 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.402346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.402346Z digest=sha256:7f9373fc397edd592479af5b165703492cadb92f734a20eb7bebee4fb69a907d

Observation 1367ce42-ec36-4505-b7ec-4aae10fc5a3b · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-xl: Extra-long vision language model for hour-scale video understanding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.397566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:08.534974Z digest=sha256:eda9e736a229663d347bfca45667494f1467cb870326b3c8dfe65f092638535c

Observation ac8e8494-8784-4eee-8003-c3e14569b47c · outbound

This paper cites The proof and measurement of association between two things.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding The proof and measurement of association between two things

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.752562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.752562Z digest=sha256:e98d26e5be19e4fb23d0392b6e8c5cd55f72933b0fd63eedc6dc059281a9b5ba

Observation 44971cd9-b215-45e0-bfe9-0d8d6dc5fc3c · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models, 2025.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Dycoke: Dynamic compression of tokens for fast video large language models, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.268873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:08.868437Z digest=sha256:768424419c93f7f3cf187afaf04eb64805b1f8e3f62fbd39639e8c9ce38c4685

Observation 358355fd-9998-444c-aae7-b87c1418de3a · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.637236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.637236Z digest=sha256:73c4e6caab121d9d9fedc05cf61482703aa68f5f4de1600e2eface77f2c2a84a

Observation f1b035eb-1fbc-4349-8fb3-6a2e48470e27 · outbound

This paper cites Qwen2.5 Technical Report.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.068469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.068469Z digest=sha256:c36b732fc26d83b5fb956ef8346d8bf5013d3900d912df23a279bd03c2e21256

Observation 5ee70352-c5af-403b-9cb6-f1beddcc5450 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.173823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.173823Z digest=sha256:2e9be4f6fc36fa40543cf6570668975439bd3d2ba6d00b5e07069aa762f86560

Observation 8e3b39db-a13d-4434-9faa-b797beceb8ab · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:08.992098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:08.992098Z digest=sha256:b49259d226c8bc574ac265fd6bb44baa46b6ef947793dde1ed99dc6f6d565ed8

Observation 59d56e0e-84dd-4318-9897-74785c7466c4 · outbound

This paper cites AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.407693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.407693Z digest=sha256:4fc9b7eae6fa97ea096ca8c9fbb608df8eb01e8092860a14e72bc944c2bf618e

Observation f35f5d8d-1efb-41e6-96bd-6b0121220b77 · outbound

This paper cites Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.565561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.565561Z digest=sha256:450d935c860c6055dd6ca326fa75b49ee22f2b78f0fbe6580bf55563d1fa4183

Observation 85466586-ec1c-4645-8ef9-f0792b8a866d · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.320069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.320069Z digest=sha256:cd3a7b92b3d364809961da1b2f017ac413d8035c1415f874652d876f4cd9f491

Observation 03995c8e-38cc-42d9-8556-b7994c2b04e9 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.820042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.820042Z digest=sha256:fd54d8dcc1b0df11353fce91fda2f8f8234f20e806311f728517e57703021c5a

Observation fdd0506c-935c-43ec-b472-bfb3ddaf2094 · outbound

This paper cites SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.935958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.935958Z digest=sha256:1088e904c96c83682c7389344fb807627e01442a824115d14a71ecc6bd4287ad

Observation 07b2a234-bd0f-4ea4-832d-8429a0bf96d2 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:09.727032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:09.727032Z digest=sha256:de20268ec46a56727cbc5d3f4ee812c00be0be44f6b3c387ef949403ea7ca519

Observation 98b0f7b4-8bb5-4331-ba7a-d661b7b8ed23 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.291998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.291998Z digest=sha256:19b05a7a99e41e180002da2a6b9e2763f1db86fc7200c45bf144fdaa560d172e

Observation 317dbfad-ce0d-4ebf-b807-df11483ea597 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.075062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.075062Z digest=sha256:0d0ffdfd61c8b8b6548f34864939312d01dbdf4aa605d1d0cdb8f8cb33cc0b78

Observation b56cbe40-4a7d-4461-89f2-913ff3f5f7bc · outbound

This paper cites Long context transfer from language to vision,.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long context transfer from language to vision,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:59:12.081596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:59:10.503442Z digest=sha256:f9f35784f60c3e55f8588b139d3efe0a02fbefa82ac405628c5a81238b29a9c1

Observation dfa4e81f-0ea0-49af-aad9-41a331e93e78 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.721952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.721952Z digest=sha256:d3c3433b8fa13ab06e89e19277d6997141bcc4f646773b90c7ba3314d81752ca

Observation 8826dfff-bcbe-4a51-859c-962a6a42fd2b · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.390042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.390042Z digest=sha256:3da2bdca981f8b01d7fc563889bf01c4196535335b1fcf3acb8b3c01849956ba

Observation 947f0de2-438a-4d18-be4e-eb00dfad343f · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.962445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.962445Z digest=sha256:b12c7b2a091f1616a525a1c483d19a64c9e4396049c4ccb04df19a2b10146e00

Observation 31892af6-6997-4b2c-9b72-0e9cc0709858 · outbound

This paper cites Long Context Transfer from Language to Vision.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Long Context Transfer from Language to Vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.606703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.606703Z digest=sha256:fa97cbf1a3708f94ff20260ca82f2a4cd49235477d2d170bf72262c6d525368d

Observation 28156de9-2eef-4bfc-95fd-494707660083 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:11.232538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.232538Z digest=sha256:7c106066ea9b7659cfb9e2b961101b5f28fcee2ce66b33da04d605fc1b35cc4d

Observation cb4cd5d7-3937-4748-ae83-c7fe93378ead · outbound

This paper cites LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:10.869569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:10.869569Z digest=sha256:df2068ff427f28be464743505cc1aa7439503ea08d50ca03efc33452576833c1

Observation 7f6d8b62-bb23-49b7-b3fc-4b9f9fd99e74 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:11.136490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.136490Z digest=sha256:bdd10d53518fb4c240c9aef01a47ae99828d79144f8f3dc9a41c0243362164a3

Observation 0ff19277-4d81-42ea-a6ca-e0b1fdb91644 · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:59:11.354129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:11.354129Z digest=sha256:ffaf46d96c552297efd5ba478ca4fb56f95424a24c600ec937af4a2532454f32

Observation 6e10be65-1ea8-4a7a-b4ef-a848e5b51ac4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.050092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.050092Z digest=sha256:ca1ddcb8875c8cf7e17413ee5cc1b65daa61764ca2b8b3d6c68859e6f733e2f2

Observation 0add14b5-d3c4-4cbf-becc-c9a872c9f851 · outbound

This paper cites Qwen2.5-VL Technical Report.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:05.425676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:05.425676Z digest=sha256:85248ce143a7eac71e8e680fd44582f6066a354b16e8ce9732621d7515cc60f8

Pith citing papers

Observation aab852e4-f310-444f-81f7-d11913bbc228 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:08.270295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:08.270295Z digest=sha256:5896d84747b27c1f9414aa5117ecb6d0feefda0e42a01668a57448eb8269d282

Observation ad69ed62-b0a9-4128-bfef-26435c8257e2 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.045461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.045461Z digest=sha256:11129d58cdfedac432db97b6346967ed9466a833698c278bc542b90be1e2e9b7

Observation 8232cef7-438a-4279-b437-41f59df8bfab · inbound

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG cites this paper.

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.175133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T07:19:44.125479Z digest=sha256:0961872250e541e5d60ddd28d7180bbda3f2b68272476c2283ae63719004d1ab

Observation fd526ce5-1a3b-4b14-91b0-56a7c9883c5e · inbound

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding cites this paper.

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:00.719844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:00.719844Z digest=sha256:1088fda50e0ee41459f11098e8f10e4f977f20c00504c5be82d7dcf8cd6da429

Observation 1b80a305-06b8-4a90-b461-e9b9d104db82 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.141746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.141746Z digest=sha256:43bb0db602521a2eac16c73d3fd01c85ed8e252f51d870b93a234b053870a856