Pith. sign in

Paper Citation Record · LEDGER

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding

As of 8 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2607.28463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28463 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:37:33.282545Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b85b768a-13b4-4b7f-93f5-64695f20e3f9 · outbound

This paper cites Video-ChatGPT: Towards detailed video understanding via large vision and language models,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-ChatGPT: Towards detailed video understanding via large vision and language models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.893226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.893226Z digest=sha256:bf5a50d1fa5b327a207027af0e4f9a8c52c93b5c0b9fa1ed7a09d4cc85b71613

Observation cb94bac5-d01f-4503-a2fe-744d3185aec7 · outbound

This paper cites LLaV A-video: Video instruction tuning with synthetic data,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LLaV A-video: Video instruction tuning with synthetic data,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.898930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.898930Z digest=sha256:561e9438bacd13fe65f8fe747e5bdf99f5bb8fcf771f66238f3ed8c1a347b554

Observation ea62da92-f269-41ad-bcee-8fc920916d88 · outbound

This paper cites LLaV A-onevision: Easy visual task transfer,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LLaV A-onevision: Easy visual task transfer,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.903658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.903658Z digest=sha256:845396c6e42c0384323dd9d617f042de468c8200b6e120aa00d513e84b26d9a3

Observation 02c38278-5b66-4c27-a4c0-7dfe9ba73f39 · outbound

This paper cites Qwen2.5-VL Technical Report.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.908495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.908495Z digest=sha256:acc0bc33ae152906ea69e061a32526235ef40165a04e819f0254d8319c6a12ff

Observation 0dea72b5-709d-439d-9d01-f6f95e915957 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.913618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.913618Z digest=sha256:605e8bec60531fd87475d8a3c2ed6c5f5049345fddd66e9822f603bdb54e3259

Observation f1157303-bd66-4240-ab16-66892b8deb5f · outbound

This paper cites Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.918662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.918662Z digest=sha256:29257b94fe3b3a7743c4643123ba55d3937de545535fe39d0590344881864bcb

Observation d9009200-6f30-4bb9-82d8-340c60ad6ba6 · outbound

This paper cites LongVideoBench: A benchmark for long-context interleaved video-language understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LongVideoBench: A benchmark for long-context interleaved video-language understanding,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.923742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.923742Z digest=sha256:3348e1a2dbae4978d3cad6d54f8c680e9ca7d7800be9f5a6b4927b20679625cd

Observation 021b3815-4d18-4ce1-8094-b13d61d8070c · outbound

This paper cites MLVU: Benchmarking multi-task long video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding MLVU: Benchmarking multi-task long video understanding,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.928010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.928010Z digest=sha256:2ab8693c003fa552387678f492365c2066af95befadba6d56484f03f7fc11a0b

Observation b46c9606-395b-442f-8570-7fd1f68bc4fa · outbound

This paper cites SlowFast-LLaV A-1.5: A family of token- efficient video large language models for long-form video understand- ing,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding SlowFast-LLaV A-1.5: A family of token- efficient video large language models for long-form video understand- ing,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.932092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.932092Z digest=sha256:1bf90bdaa17080d022de816e7ff5d3eacf9edd60634bd4d9931118afc41d0255

Observation 5e7a7f3a-19a0-40c9-9c1a-774f61299f3d · outbound

This paper cites Long context transfer from language to vision,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Long context transfer from language to vision,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.936329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.936329Z digest=sha256:76f7a20f3577196e25549ad27d6a1afd1feba934bcc85680e58ad79ac6d10dae

Observation 3d1ad064-51f3-4478-97e7-c2463ed108ca · outbound

This paper cites Adaptive keyframe sampling for long video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Adaptive keyframe sampling for long video understanding,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.940600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.940600Z digest=sha256:cebde48353ca81b4ddea3515cbf89a76ccab5b629f107d739ccbf0b1075793ba

Observation 0060e0c6-ac84-43e5-afbf-a2f61b50c93e · outbound

This paper cites Q-Frame: Query-aware frame selection and multi-resolution adaptation for video-LLMs,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Q-Frame: Query-aware frame selection and multi-resolution adaptation for video-LLMs,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.945112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.945112Z digest=sha256:fa5df47247724e2f60030a88df2b3d54dd2a5a4d01f72d8976e71dcedeb137c0

Observation 53d1db39-be30-4927-a2ba-103836f8a624 · outbound

This paper cites BOLT: Boost large vision- language model without training for long-form video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding BOLT: Boost large vision- language model without training for long-form video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.950792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.950792Z digest=sha256:3ccfa571c41b29436d447dd7127313763b1badc274e0bf4fa541270e4960a44e

Observation ec4378cd-1ac9-40e2-a7a0-18679656ba4e · outbound

This paper cites MDP3: A training-free approach for list- wise frame selection in video-LLMs,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding MDP3: A training-free approach for list- wise frame selection in video-LLMs,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.955257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.955257Z digest=sha256:1f346b16949ff9950b91cd4a1fa39c0086b14dc0c4f2783fc9985bcf83040a1e

Observation e8524394-313c-4732-8c3b-06a91ccf5ab6 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.960181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.960181Z digest=sha256:8950688249e4b832b829e040a1244d50ca870d4338c5f7168b6fa1caecb1e7d1

Observation 4d408267-d69b-4a37-8902-1a9e53a87c11 · outbound

This paper cites MaxInfo: A training-free key-frame selection method using maximum volume for enhanced video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding MaxInfo: A training-free key-frame selection method using maximum volume for enhanced video understanding,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.965907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.965907Z digest=sha256:aae8e720a733e7bb4810986465eb11f2b846b5ceed1dc71b6178f769fc2dbeda

Observation 6659840f-9c13-4b3c-9844-5e3dee41d8f0 · outbound

This paper cites AdaRD-Key: Adaptive relevance- diversity keyframe sampling for long-form video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding AdaRD-Key: Adaptive relevance- diversity keyframe sampling for long-form video understanding,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.970300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.970300Z digest=sha256:5d91a734c9a1a189636b84de282e5429909a9c3ae636d5fc30dccf17f3dc861b

Observation 32732981-9dac-4e5e-b682-39fa7b19ab3a · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.975156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.975156Z digest=sha256:8b9f18b5ff9e812b7c67ad08bc4649471b058ac08e3425942d45b79938ca6fc1

Observation cf471211-9d4a-4e9a-a4c2-231a20393fde · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.980214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.980214Z digest=sha256:f066ee63c6f793e6ac5597167750730d4fb4bd3c043c8ca8ac6afb3c1350d61e

Observation a1204b20-c776-4798-8fe7-6f10d97f197f · outbound

This paper cites LongVU: Spatiotemporal adaptive compression for long video-language understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LongVU: Spatiotemporal adaptive compression for long video-language understanding,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.985690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.985690Z digest=sha256:04fe88e39d4f174dfaafe268e1001dc6c00ac4c1327cce0bb284f3671bd2231f

Observation d8088731-e9df-4580-9e07-dab1a9233c9f · outbound

This paper cites Flexible frame selection for efficient video reasoning,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Flexible frame selection for efficient video reasoning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.990356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.990356Z digest=sha256:21836ca9ac0a3d61d0a18f8ceaed444cbf9a1af67c98b67a393ab705537e8b09

Observation 02b70ab4-dcb9-4bc8-99f1-2cc42e68a8b3 · outbound

This paper cites Frame-V oyager: Learning to query frames for video large language models,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Frame-V oyager: Learning to query frames for video large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:32.995703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:32.995703Z digest=sha256:8e45b21d1c9ed7dbbcb60a87c2e38f76acbca5bff74dc16cb307f36b0f7b41c0

Observation f811a44b-73e0-42e6-b383-f8d559a35a9e · outbound

This paper cites M-LLM based video frame selection for efficient video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding M-LLM based video frame selection for efficient video understanding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.000151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.000151Z digest=sha256:c03c4b6103833cf21d8909961877b3dfa066da07e159983773cc6cffcbe78905

Observation 640fca01-8386-449c-9aed-13b844b1a0ee · outbound

This paper cites Self-chained image-language model for video localization and question answering,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Self-chained image-language model for video localization and question answering,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.014743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.014743Z digest=sha256:be2fc5c5b0a8433bc5de86654879fe84823bf2078dee08a5b1914acf698c77b1

Observation 4f0856bc-5b41-4f4a-bf23-0f2a705c0209 · outbound

This paper cites Generative frame sampler for long video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Generative frame sampler for long video understanding,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.019695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.019695Z digest=sha256:843c810a40e746a12179ed280000c47d99962a607699168047cc7e0543002d00

Observation f78136fb-ea79-4290-acee-5f1f6fefb0fe · outbound

This paper cites K-frames: Scene-driven any- k keyframe selection for long video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding K-frames: Scene-driven any- k keyframe selection for long video understanding,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.024596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.024596Z digest=sha256:390588dc5d44ba59faeeb7111ea47ee82eeee763d6a202fa296f5e49f19a3427

Observation a45a8612-d07a-4653-8684-a7c7a3ea477f · outbound

This paper cites Event-anchored frame selection for effective long-video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Event-anchored frame selection for effective long-video understanding,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.029368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.029368Z digest=sha256:df19d2e047e9fad03465d081e29c4cd0ede501302b4920f5ac763e2c4d763065

Observation bb5544f8-12ce-43b3-ac89-201b1cfa80a2 · outbound

This paper cites Wavelet-based frame selection by detecting semantic boundary for long video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Wavelet-based frame selection by detecting semantic boundary for long video understanding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.034100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.034100Z digest=sha256:85ae730ab7d49f34b08c91198d997b2228cf9b8ef48d43fe1e39fc826b0b702c

Observation 22fbb8da-d30c-463f-b493-f56d00a0e88d · outbound

This paper cites The use of MMR, diversity-based rerank- ing for reordering documents and producing summaries,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding The use of MMR, diversity-based rerank- ing for reordering documents and producing summaries,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.038674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.038674Z digest=sha256:4adef32dac15469d004bdcf5b65543a4df8b491b66253dc9a6c5b4403b67e5f3

Observation 352f5502-243d-4268-9229-1230e2ffa605 · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language mod- els,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.043237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.043237Z digest=sha256:5795c387e93beb6bd5ab37ba8245ab1474ffa208cf85cc10a103f0fdb06f7741

Observation 42ce74f0-32f4-48b5-ad90-54912741af47 · outbound

This paper cites DINOv2: Learning robust visual features without supervision,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding DINOv2: Learning robust visual features without supervision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.055972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.055972Z digest=sha256:322cf5c3e83c0522368d9f3a354822f8f12cb6da33237c03397d09d61d64dfa0

Observation f27c1b9a-4834-48b7-8b62-00864158f02e · outbound

This paper cites Determinantal point processes for machine learning,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Determinantal point processes for machine learning,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.104742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.104742Z digest=sha256:9591a9b49863b968a165383eefbca52a5886d45a6691e6d2dc130257c67121d7

Observation 7d92d894-6d94-46ab-9a1e-9e57c1fdec77 · outbound

This paper cites k-DPPs: Fixed-size determinantal point processes,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding k-DPPs: Fixed-size determinantal point processes,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.144753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.144753Z digest=sha256:d9489afaf5114de92ff08321ffe60d83a60495529d605cbcffd69f51d2d544e7

Observation 454b277f-3406-4a69-b501-b4562d78f71e · outbound

This paper cites Fast greedy MAP inference for determinantal point process to improve recommendation diversity,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Fast greedy MAP inference for determinantal point process to improve recommendation diversity,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.184753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.184753Z digest=sha256:326241a839e43febe240e765b4fe6e26f340ba526c8870dd57dc709753c59725

Observation 2a6d5c48-3ed5-460e-ab82-d798e829484d · outbound

This paper cites LMMs-Eval: Reality check on the evaluation of large multimodal models,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LMMs-Eval: Reality check on the evaluation of large multimodal models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.193017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.193017Z digest=sha256:20295d85e9ea8386445d01d81c144d51a2ecba469864726926076da315336b97

Observation eba41349-aa3c-40b3-97e3-4978d4db6642 · outbound

This paper cites Qwen3-VL Technical Report.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Qwen3-VL Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.197321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.197321Z digest=sha256:df0207cd1a914ca8878ec353be1b407f7cf51c99b839a08783043ad3b21792f9

Observation 561d2acc-260f-4c6c-955c-ba1818d5b686 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.204746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.204746Z digest=sha256:e2ffe3c5a3bdee85712997b8ca0e5f88711fdea24764325419b94a221f605fd8

Observation 56ed70c2-851b-46dd-9b5a-cd8e2a53dc17 · outbound

This paper cites LongVILA: Scaling long-context visual language models for long videos,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding LongVILA: Scaling long-context visual language models for long videos,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.244756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.244756Z digest=sha256:1e26a01a01e198a269601d0735a2c6855bf4b54909f63ba5daaa0fbbaefbb958

Observation b1c71d7b-bc7c-46dc-bb34-a7b0fe49d5a9 · outbound

This paper cites Video-XL: Extra-long vision language model for hour-scale video understanding,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Video-XL: Extra-long vision language model for hour-scale video understanding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.268082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.268082Z digest=sha256:f013c9af6ae63a7473bf3d5c79c0d32181779169356c2e1cc2705204d3ee35d4

Observation bdedef4e-de7a-4baf-8e13-0724e96f652b · outbound

This paper cites Learning transferable visual models from natural language supervision,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Learning transferable visual models from natural language supervision,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.272895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.272895Z digest=sha256:379e8d4e1bc2ffed09f2f8178d75ae761119d85c00b831e0f1b70d2338ed5b65

Observation d5fcc2c9-5ecb-4808-b93f-f6c5de58c6b8 · outbound

This paper cites Sigmoid loss for language image pre-training,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding Sigmoid loss for language image pre-training,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.277099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.277099Z digest=sha256:81fd9e239ec2af5d87e2ae0c2be18fcf784520022470fba0adb79b435f08bd9c

Observation aec06c51-aa56-44c5-803d-55c037d2c96f · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T06:37:33.282545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:37:33.282545Z digest=sha256:e730ee82e89c3901c0edf016770a0600829be560224b7fc9a9cc25177f2cad82

Pith citing papers

No inbound Pith citation observations are available.