Pith. sign in

Paper Citation Record · LEDGER

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

As of 10 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 0 inbound Pith citation observations for arXiv:2608.05780.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05780 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.360123Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88193e5c-508a-4353-a368-0423a0a7db08 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.967551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.967551Z digest=sha256:f5d7b522e8d53134136fdf7d0a8c05495e1e8adf4456d6087824331c93ad5d56

Observation 322731e2-edaa-4f28-8a97-e0626bd31fe9 · outbound

This paper cites 2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding 2025 IEEE/CVF International Conference on Computer Vision (ICCV) , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.972401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.972401Z digest=sha256:2be0985ed8ab8d81d382a1914f2558068305559d94ae8bd9fbefd4328375c09c

Observation 95c97b24-1dcf-4dab-98bd-8c7ddd47d332 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.976757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.976757Z digest=sha256:82322bdeb2c5ef9693073431a0a43d3053d2434531b86d9942db4a64ed708eda

Observation fe9af7a0-9666-432a-bea8-55eb0c103604 · outbound

This paper cites PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T23:42:38.263471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:34.980982Z digest=sha256:51e48826b3d4fe6dec4addf91ac2c05ea5526eab1793b9f775743ffa05001332

Observation 95e3be72-963b-4e16-8dc1-e63d4aaedbf0 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.986046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.986046Z digest=sha256:d11a69a140b9bf24e15bcdeea530ecbc6dcaf264ce229b0a923d5f75ead0d459

Observation 91da7fbf-cf4c-42e8-859b-57cdbc5c4565 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.990368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.990368Z digest=sha256:04b01e5defc2a0acb7951fd26b4d0e6f9ecdd39c8ba96361a687855963cd5af9

Observation 1468d980-9e42-4d3b-9c1b-9749f003a400 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.994738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.994738Z digest=sha256:006ae3c68a0575bf9bb7a2977fff3891c8beb6cbd773f7002fb7d846e71a65d9

Observation da832fe9-bd47-4155-84ab-af0c83ea644c · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:34.999140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:34.999140Z digest=sha256:f9449b73a28f6af72365b1190bd31bd06ab9f29c256dae15645f6739ca28459b

Observation 791ff586-6a7d-43bc-80a3-1bcf324214ed · outbound

This paper cites arXiv preprint arXiv:2507.07966 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2507.07966 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.003049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.003049Z digest=sha256:bfd9b6f4821def5b7bcc55ee4f01f06e90f06e6f2d9af093c07e9c7d0986b102

Observation 1632a862-58f5-49aa-adfd-28d2d8552ee2 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.006944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.006944Z digest=sha256:77d1b8ecfc51f60205f77da0b5a4f3e01042eb29d854528defe81ac48789d7cd

Observation aece37f1-91a7-455a-af84-5edb8d9f13be · outbound

This paper cites Qwen2.5-VL Technical Report.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.011144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.011144Z digest=sha256:8b21418bf8b9e427570e0591387bb5461ed239e551aa433e8895c9526d382d45

Observation 5e7d020c-7a32-4010-9062-a055c58b0abf · outbound

This paper cites Journal of Visual Communication and Image Representation , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Journal of Visual Communication and Image Representation , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.019850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.019850Z digest=sha256:69ba7b18131f622bb98ca0bd4dc666225eb1c8041875998370f54c244262689c

Observation 93f062a9-c40d-4f58-b11d-11d09aa7a2b2 · outbound

This paper cites Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Pairwise is Not Enough: Hypergraph Neural Networks for Multi-Agent Pathfinding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.024660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.024660Z digest=sha256:f3388d02f2a01482a33c0b3e220fdf33a1a84732c0205d968a22a449ba7c64d0

Observation 076c26d2-d334-4f85-ace0-5774c15b4a2e · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.029043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.029043Z digest=sha256:88ca6d57798c10ced1a48c62986a97ebbd5c4a5aaee86f6436dc1b310d0fb0de

Observation e6977576-efe3-4033-b901-1ae2ef49a408 · outbound

This paper cites CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.033889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.033889Z digest=sha256:a9a62f56db12c2794d13ce383e89870a0a7e2b65617b4fb7cd6ba24d9733f555

Observation 49e25a13-f186-45c8-85a8-c1ed82e9cc6e · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Describe Anything: Detailed Localized Image and Video Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.038915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.038915Z digest=sha256:7b5d6b46e221d71b26b88374754f30a8621fcf1a011f24a6f21893bdd604d8af

Observation 40435e09-dda2-458e-b446-000310e5d923 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.044026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.044026Z digest=sha256:25c4e2205c2ad506a56ac3cfb0f749276ab124ae621c05e7982238323dd2c8f8

Observation b18ccfcd-fb5a-49d3-a109-90f2e479835f · outbound

This paper cites Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF winter conference on applications of computer vision , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.047867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.047867Z digest=sha256:3abea17667af77cf41ed81af73fa84a4930bd801a0954d5b375fc56279a979d8

Observation a50bf22f-14fb-4ea0-bc91-1d05444a9348 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.051735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.051735Z digest=sha256:cee7668eb69fe8df2336374943bec0da928449c91df2cfc3f5249a7391a05304

Observation 67558f6b-def8-4ec3-9a80-11676bf87a67 · outbound

This paper cites arXiv e-prints , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv e-prints , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.055629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.055629Z digest=sha256:9321b331968d5c33efecf48215c7432f777ad87afd9726cbdff9bc52b9113458

Observation 532e0f22-be28-400e-b515-63fca5abadf7 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.060058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.060058Z digest=sha256:b7c558bf3dc8be67af1d584f0a609fc4fcc4a0b6aeda342b6dbcded4e997ab39

Observation bc3329b6-0a19-48fe-9578-f8319e3a2843 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.063900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.063900Z digest=sha256:d01fb97cf183e9c6333a3c3f0a93eed24277add11cb126e1fab5798f0ae05a4f

Observation 9b1d950f-353f-427b-b794-37e0b9e0cecf · outbound

This paper cites arXiv preprint arXiv:2508.04369 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2508.04369 , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.067606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.067606Z digest=sha256:deaaf89f79de19dbe23b5cfcae34f0f7aa8afb97435a2c2a7b56d0bd1f87707b

Observation 577608bd-94de-4b36-a3c6-6450c71c66ed · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.071742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.071742Z digest=sha256:18f687cab1f2b9358b0f63c03582877fbde1bac0e3ae7890c493ae5e9a0f26c7

Observation 9fc0d62d-9555-459d-94bf-0351f54c83b0 · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.075368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.075368Z digest=sha256:c5daf1de8300971a005958f4abcac3c873fd972e21b257eb4ccc424a6a682577

Observation 1df357a6-03e0-4789-a052-6628592dcd04 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.079279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.079279Z digest=sha256:dce646d2791d5abc9ca4182de0dd9e2ceba732434a949195ee7c39c79d033406

Observation a5af1714-1c85-4d3d-8b9e-8d47f4503caf · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.083243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.083243Z digest=sha256:9c7ffdb7300b24c2b34bd747589bb4eef247b8ed939ba2b79ad0a36e46bf52fb

Observation 2faaf74f-9daf-450d-a617-3a4466d330b6 · outbound

This paper cites International conference on machine learning , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding International conference on machine learning , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.086792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.086792Z digest=sha256:fb2fc660a21361e31204f9d5d407abcbd41671a876e8bd0417cfe0da3a054c96

Observation 5a542cce-ab43-4480-b6da-baf3ac9d6403 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.091644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.091644Z digest=sha256:d5384141123a382ee0a78e0ae6159df3176a1c3097323752a383e251635df80e

Observation 308c53c7-5f8d-4242-8dec-85a3af4acfb3 · outbound

This paper cites Long Context Transfer from Language to Vision.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Long Context Transfer from Language to Vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.096042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.096042Z digest=sha256:ae4352caf239c63c17392e55a3e278a4704393606318d245b822f0f98a1524ea

Observation 1a34726a-c37f-4148-b236-4e7398ac6ccc · outbound

This paper cites arXiv preprint arXiv:2502.05177 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2502.05177 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.100870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.100870Z digest=sha256:3bea94ea63016e808d84aa89d583051d9d2b5f545b05fec8b2f346899763b9b8

Observation 62eb75a8-496f-4c5b-9297-5a32e09166bc · outbound

This paper cites European Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding European Conference on Computer Vision , pages=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.104559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.104559Z digest=sha256:ca053fb25c85d2d90383391cbab6fae57452f2d95dd85240fca62c548d064996

Observation 9d0d5924-46e1-4636-bc0c-4ef0ccd7060d · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.108718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.108718Z digest=sha256:49d9ec68910ee993f9ee0e35b626975204c9857b24a18fedc58a06b162a052ac

Observation 23615e5f-63f5-473a-ae1e-29277313f056 · outbound

This paper cites LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.112820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.112820Z digest=sha256:2560ebd7a1d0c82f99cd973848f19acf48946ab0ec07491c4dbcf89be1fa3c27

Observation 3a363538-f653-4df9-86c7-043fdc5295c8 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DINOv2: Learning Robust Visual Features without Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.116786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.116786Z digest=sha256:6b3b3763d7675df57b6e761a2c2d689e796c83d7a207ab6d5e7230440ce02505

Observation 2528b329-713a-44be-acd8-f0ffde781769 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.120837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.120837Z digest=sha256:ab25bbbf5bff753f98924f3375748d3eef51263d926914f62a09e3bc024fd635

Observation 31af810c-b90a-453a-8506-cbaa23064c22 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.438990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.125004Z digest=sha256:2672c65dbd2e240da42622cfb7d669127400ecc6a1fd44184be087e232bee768

Observation fc745383-719e-47c8-9c37-c19f8d4f6247 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.128973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.128973Z digest=sha256:b0614b2935ee2c0956506a622ccd30aeb2412369b080f021978f72e6b177fcba

Observation dd591fcd-b595-4d6b-8978-6099f1586b65 · outbound

This paper cites arXiv preprint arXiv:2510.27280 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.27280 , year=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.133666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.133666Z digest=sha256:e1930b4a64f5e84871ffb93800cb74fc32aad5b036254eb63823d30b14cc5eb0

Observation 83ce1718-bef9-4d79-a24a-13456167b10f · outbound

This paper cites Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.137666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.137666Z digest=sha256:8fd275b40a992685a712b68ded00e121586ce8b30623a6c0651079192bdc5a33

Observation 1b80a305-06b8-4a90-b461-e9b9d104db82 · outbound

This paper cites FlexSelect: Flexible Token Selection for Efficient Long Video Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.141746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.141746Z digest=sha256:65cee65635ee5450969164c834a441dbf33f98ff37f6cead692fbe812f4d0081

Observation d680571f-867a-4a75-881e-03ebc212f6d2 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.426552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.146226Z digest=sha256:0dcb38f6043374c048210cbfcee0c6ee7d0cc0061b203a9310d523db9ce3c820

Observation 104a7dd4-af46-47c8-989c-76bb7ba2552c · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.413824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.149849Z digest=sha256:91f4bfbd5210ae2f15cfad6bc470593440cf0bb3a267c3a22ca9d0386aa1bc32

Observation 9a674896-953e-4cc1-aae0-29c2ccddcae7 · outbound

This paper cites arXiv preprint arXiv:2510.13891 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.13891 , year=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.154446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.154446Z digest=sha256:e8a3787b6dd286e639f5d4560b0ab55139f8b21d01fb65fb812e0c5a9237a4e4

Observation 949b5c8a-2cf7-4ce3-a17a-63fc99fbed09 · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.158387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.158387Z digest=sha256:327864b3038e4df3826ae543c48e9668e9fe4cb22d3277bf197cded326092c8f

Observation b9026e97-5b11-4dc2-88ff-3fe443f527e5 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.162434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.162434Z digest=sha256:9a2a27bb0988819a0954a44a15d596ee551da0bc183aaaaf34dccdb5106af88e

Observation 4ac55cbc-22b6-42bd-9fff-474e2edde272 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.166581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.166581Z digest=sha256:e9845d3f79a8b327e41fcce3e6c78409643d4baf6e6259f4eecacbdb01f326ae

Observation a341499c-793e-48c2-9e41-3b2136d50295 · outbound

This paper cites arXiv preprint arXiv:2512.11534 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.11534 , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.170212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.170212Z digest=sha256:00560183c33e0aa5a419ad418ab56e801be1347d32ebefbeb7c806c410afb443

Observation f24d6f3c-e584-445e-93d1-99d3b7bafd5b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proximal Policy Optimization Algorithms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.173910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.173910Z digest=sha256:1541530d77e6a7e4b6347e4a529a749661a954cae4b150ad7434460c081a149e

Observation 454c92c5-a9de-4550-9945-7fb52ba7f392 · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.177728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.177728Z digest=sha256:330779c63d2df80787939393b2649b0482a13037f24d5818759565723f3378a4

Observation ca11859a-759c-41e5-a5f9-85d0663ace74 · outbound

This paper cites Advances in neural information processing systems , volume=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Advances in neural information processing systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.181274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.181274Z digest=sha256:f34b4321a1e61470f6b25418ce81061076de24bfa69b87c3d54b19f73bcb4ab4

Observation ef3320d5-b85e-4927-9168-7fd73b355b84 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.185360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.185360Z digest=sha256:ea16a2c4833d25c38f8be8e19ed7eb82407eb842a833769385332d8ba25bb270

Observation 9acbfab0-4321-4bc9-b2a6-ab619266c7e4 · outbound

This paper cites Proceedings of the 2024 conference on empirical methods in natural language processing , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the 2024 conference on empirical methods in natural language processing , pages=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.190392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.190392Z digest=sha256:0dff93fb39ee7ab912fe913d1b4bb152c8dcb738e5e79aa89b6aec07d825d07c

Observation 0dfa2b5b-5896-496c-ba1e-69fde5afd8e0 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.195518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.195518Z digest=sha256:c0fa5c46c0fac1b8250c634d242c0f822af9e8a6ab009ce6d21614ecca61c1b4

Observation a3fbdcb7-97d9-4efe-8d32-f232731bec8c · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.376340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.199407Z digest=sha256:0e023eff4ae383909df6353dfdc2c08b10a97a34655c1365a799a975aacb3b63

Observation b44f048c-53c6-4815-8e91-8912f439fc1a · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.203009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.203009Z digest=sha256:f8ef0dc3a66d7a8d1021ef2a3c7315024b056c3ecc8472786ea9a21044458008

Observation dce91058-7b8c-47d8-99cd-fdaad28ec38e · outbound

This paper cites an unresolved cited work.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.207163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.207163Z digest=sha256:3c201300f62d41216ddff4c14d9978a6ac7804690f30ee09983703c1f497b95d

Observation 108275d8-c8cc-4622-9a7c-4a52af0a8537 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.211189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.211189Z digest=sha256:f50eb7850661ad5c4d6f9ac55787ae66d2d2763eee392b45448ff3825b992a87

Observation 2adaf088-b632-4f60-a978-61943e5b8a5f · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.215228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.215228Z digest=sha256:0c0263351fd357a0c5754aa74c9760885a6c64ce80ff311f0f081f37d4f1f6f2

Observation 735d3b90-9c25-467b-b563-83d0b51517eb · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.219360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.219360Z digest=sha256:1d0f38a8c817638ddafba7ed2483fa40d34bcb060b2c4a0cb8473ab32c013abb

Observation 6ff40cdd-5706-4348-972a-deb5b3b8976f · outbound

This paper cites GPT-4V(ision) system card , url=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding GPT-4V(ision) system card , url=

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.353966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.224392Z digest=sha256:8b140e3ebf7f298b9a8fed4eb1e1d0ee16c7927c0e95fb16f0ef013ede59f920

Observation e7ee0b09-26c2-494a-9af5-f554f2efc499 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.228769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.228769Z digest=sha256:b83b6e3b387e144528f22a544ce8f18a3d287dea680a5848a45a860fcbe2fd99

Observation acfb088c-807c-40ce-a2ea-8892e3ff6b96 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.233288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.233288Z digest=sha256:3aa13bc06f276cf830103f4e5031cf8b7a7acd2fa3059fdb8800e1f7dc24ca8f

Observation 29663a2e-d21d-4b52-8965-a6238af709bf · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.343764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.237827Z digest=sha256:d228a69d42e72fb14a0972739a52d117456e212e6f221cf8a23f82d6dc474322

Observation d3ac00bb-b2ca-4bee-bfe6-89ec6ab877b7 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.333721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.241909Z digest=sha256:bc7ddd4d7d7880febb5f4808ad84c424b2113f9b6ca125b7e31a1575442507db

Observation 1b31cd46-dfdf-4aac-94d0-baf39675fb3c · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.246208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.246208Z digest=sha256:2589c949745ed77dd120160b93c923cc46940373fa81b201eb9e30da391dec9e

Observation 4b778ed7-8140-41ab-a1f1-b6e80088b44b · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.250414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.250414Z digest=sha256:8b9610d8cc45aa5493c119a5f5b5e2d3b0e772d8e7e0ae7ff0f8c271f27c94c2

Observation 8be2849b-f821-4698-84c7-b874683bd49a · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.254789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.254789Z digest=sha256:691a1006a9c9b69a1e72e498b1b9d967aa439dd469eec47754570969fb0e8f59

Observation 2adc3bb7-81f2-47ca-942d-a0b06f11b82b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.259354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.259354Z digest=sha256:6b2593efa1e56b2691584d7c4dff28fe4b3dc75cb053fd30cb43b4e8d74a22b7

Observation a3caf4fb-be58-49e2-ac26-315351eaeb75 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.263951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.263951Z digest=sha256:c87c56aef69f78870e09ac66e785bf240ed1df4b1ac1466be20b3b3dc07a97e1

Observation ce3af3d7-e684-4f93-85a3-4ea4f6cb84df · outbound

This paper cites arXiv preprint arXiv:2510.20622 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2510.20622 , year=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.267788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.267788Z digest=sha256:40eaf5395e96d1614b7f745198b25c851ea1558797c0356dc3167af80cab7856

Observation d5d5aada-913e-4dd2-af7c-e63665cd10d5 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.271317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.271317Z digest=sha256:117f67d40de1f7013dfbd59283545472c1a6d3bfd98e5d676a513d1fa6af65ef

Observation 3edf887e-ae98-487f-8ca2-118820cb80a6 · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.275164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.275164Z digest=sha256:2f3524f345a716e34923126b0b9412ba37d40ffa614638a54a1a1a88f5d0c4a2

Observation ea224450-abc1-42c2-b9ae-bca17a4f097c · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.279225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.279225Z digest=sha256:96978bebc0aed055d9ee0e12400b43fdfc5af47d7e1e83ab0fce55c69f4148ec

Observation 64354bbb-8bc1-43d1-8297-6a99e744addb · outbound

This paper cites European Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding European Conference on Computer Vision , pages=

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.282895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.282895Z digest=sha256:a0bb5329e8f32d2e2b131c8473cebb8c672be447d0875be8c267d698e425440a

Observation 70505639-07af-4d2d-ac65-935bf4ad9567 · outbound

This paper cites arXiv preprint arXiv:2512.04000 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.04000 , year=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.286834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.286834Z digest=sha256:a1251c5e73abf52e9147b0d548e9de75fd57508386d9b2bc5d635a04885f1ec2

Observation 7231cc6d-7072-4e4c-bb8f-8dc3ddd4dd39 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.290721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.290721Z digest=sha256:d95c4027e87e241c011d82232dc8ae3f781ea492e1e75bb11a415b631a9582b4

Observation 0f75b74b-7551-4bf3-a952-b6523142cad7 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Findings of the Association for Computational Linguistics: ACL 2024 , pages=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.302472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.294520Z digest=sha256:8ae130e8a0d6015fb9fff8ffa0db1dca239bb3d9a627dd499013d1644a32cb58

Observation 3f747f0b-d0cf-4dcc-9a4d-28bae416a214 · outbound

This paper cites arXiv preprint arXiv:2512.06866 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2512.06866 , year=

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.298403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.298403Z digest=sha256:a185bfc3b74e7656b70987ae3471aa8bb325e827cdda90df51d633f625907f81

Observation e82103f8-687e-4358-be36-fd9b36733719 · outbound

This paper cites International conference on machine learning , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding International conference on machine learning , pages=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.302345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.302345Z digest=sha256:a4b53e45ae9a77894104f587426e6fb47a0bbb6678cc5a6d4c44e92117534d15

Observation 386423a4-cb6a-4ec6-9647-fab143900a02 · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding DeepSeek-OCR: Contexts Optical Compression

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.306068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.306068Z digest=sha256:f790b0d74fae6c41ec0975d547ac18283b2cb0a19cc2f66df47f16880fb27c1f

Observation a8e0168f-e7cf-496b-be61-0b99e380d554 · outbound

This paper cites arXiv preprint arXiv:2502.02770 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2502.02770 , year=

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.310474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.310474Z digest=sha256:52e904cd23cf797d173ef1f292f06d1b9ec7a1f7d92e11e23a6f5a433c08f8e3

Observation 8651f988-a6eb-478d-9397-9933409b53f1 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.314074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.314074Z digest=sha256:c0c1090c4f16a942212cb19d86ea24eeda263a197f347e76134b5ed0f62da84c

Observation 0c39dbf2-088b-4d6a-9299-72b5745b961d · outbound

This paper cites arXiv preprint arXiv:2601.22582 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2601.22582 , year=

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.318166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.318166Z digest=sha256:50aac1d41f10430f9c15cfc387727c5298f9c2b532da299bc50d8956a183744d

Observation 89644a61-823b-477b-9050-d893594e292f · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.322337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.322337Z digest=sha256:1ba4ecb69bb7282b9e2a3b08aa2cbdb7a7e6aa027bb125969dfa477e71dc774a

Observation 0f00d2cd-5d32-4612-8501-582a08c55afb · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.326390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.326390Z digest=sha256:729253bb503195380b3521260abb45f5684923a18a03c988d3717da6bcaff94e

Observation 6e9bdbae-b87b-4365-9025-a400c1fbc98a · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.330473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.330473Z digest=sha256:a0b284ac28e37efe5caede8cd2d343e2494a318249b9c00b937861f1b2c48f99

Observation ef3c325f-9b54-45f3-b907-81c8326c3dbb · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Group-in-Group Policy Optimization for LLM Agent Training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.334661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.334661Z digest=sha256:97af1b6674f5461c05d9029bb51150abe09e3be9e1dea110f486e2eaf26c3092

Observation a46cfd20-1e86-4467-b0c4-b3d875e8158c · outbound

This paper cites On the Emergence of Position Bias in Transformers.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding On the Emergence of Position Bias in Transformers

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.338843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.338843Z digest=sha256:f5b80ed3e870bd5fee13734ccff393c21cd51028f3be819eae09c38519e90b66

Observation df6bbcc7-9e88-4fbf-96b0-9b851f16bda6 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.284992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.343208Z digest=sha256:a2ebe7433291029513dfc924cd8be324e54e9836da423d2da5bc20ddd713ecf9

Observation 8aace232-8f19-4705-b6a3-084255d68645 · outbound

This paper cites Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:42:38.274866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:42:35.347966Z digest=sha256:276357d17ff2151d74ce136871e8b92317a45f0f91ab546ca1d391e2ef844572

Observation 8a677408-6190-4db3-9c70-bd63b83027e4 · outbound

This paper cites arXiv preprint arXiv:2504.18579 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2504.18579 , year=

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.351819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.351819Z digest=sha256:0b81537f1f8f4521b2f31626127058dbfa72250d8b8c6c0e6e3ef57f9c4b1bcd

Observation f91a1cdf-23b2-4ce0-b2ec-332e7591b29c · outbound

This paper cites arXiv preprint arXiv:2603.06199 , year=.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding arXiv preprint arXiv:2603.06199 , year=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.355809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.355809Z digest=sha256:3a03afd065e41752dedbac3f8014dc0d767e2e2a2104f7f320860bd26332dc27

Observation 6083fe37-9c5f-40e6-97e3-d61f110a605a · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.360123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.360123Z digest=sha256:b5bf9047236b44011ce23676f97c28a62c96d9a4506f2991687c3da1d4da55a3

Pith citing papers

No inbound Pith citation observations are available.