Pith. sign in

Paper Citation Record · LEDGER

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

As of 10 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2607.04605.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04605 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T16:20:19.874172Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d69e9b98-ba0a-4fb7-8bb8-3afc01558228 · outbound

This paper cites Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:d35940884429884b906f09e1f32ab8c511219a0be0c1ed63e868dfbb023ed971

Observation c3db9897-eddf-421d-913c-50aeaa3018d4 · outbound

This paper cites Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:e4f5096cf573b677f820f742892b09bfa687c8e00aa1eac02966f5bd526ae9e1

Observation e9abbfd0-264e-45b9-b99f-5ba875813938 · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval ColPali: Efficient Document Retrieval with Vision Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:4bde82c9fddafa0d31e24d2be96b28923c967ceff6a12f02e7c8c19cdbe7260b

Observation 614baf50-3197-4e6c-a184-e83cc3385e28 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:2197860dce47c1807f6e0c9daf2e823c720a94c06497050b1059d3f0dae156bf

Observation e6459299-2443-4e5d-bf89-3fddb94388f4 · outbound

This paper cites International Conference on Learning Representations , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International Conference on Learning Representations , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:9b8e754fafe402a2b994dbff440f47ca8e107f14d6d83f7aa026b8ef487596d4

Observation 72881047-42d5-48b4-ae9d-5f032e5f9f77 · outbound

This paper cites Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:bc79c400360bbadae9ece8c5bf45a787b96fd057731a0306ac22dc5bbee28220

Observation 2b44ce88-7803-4bd0-bb2d-2987fc49004f · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:ce9fe30de313c0fada24b4b59cc7be3b15bb9c2858c82d745a6e623936ad91f5

Observation c5d226f0-ce47-4fe1-8b2f-30c48a48e1e0 · outbound

This paper cites Advances in neural information processing systems , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in neural information processing systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:9e2d05f077e25cbd461bd1dce45f30462028ad0c124f5d1c1dcae32cc89418fa

Observation 7703ee07-54d3-43de-b8f2-4f0b2f26142f · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:69fd6df0a145c001b47d5ad7f91e3507c4f09abf377668ac6074bf164319e09d

Observation dc1234fc-5b86-48a6-be92-0163d25f7af0 · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:ff77a9d4dd57f24d1f7f081a786c49ee1fb5932e06ed35ec897e94260a87d9ae

Observation e80213c8-c79a-44db-a381-bb22539273e3 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:a8e9214a7f95a32bfb9695e936fcd79e75aae08419381aa3b98ce10862e94d50

Observation 4ac8931e-7d9d-4a42-8e35-459216666c32 · outbound

This paper cites Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:84924749c14145bf5f93ad7927c229ab3fca92e1deef2dc32b86841dbb8624d7

Observation 09587334-a962-4f58-b348-dfc3f3705406 · outbound

This paper cites SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:f4dbe8f154dfeb444545cb1a266a886d747e9904af0066bf12b7121e4a1a962d

Observation 48358023-8a27-4865-8203-861a2b2f349a · outbound

This paper cites Proceedings of the 31st ACM International Conference on Information & Knowledge Management , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 31st ACM International Conference on Information & Knowledge Management , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:6a3d84054734b3f7706753928dbcf1ea6d7f9294a05458562cfe9630e53009e1

Observation f1bdc37e-1621-412a-8f23-79f1bf2b187f · outbound

This paper cites Proceedings of the 48th international ACM SIGIR conference on research and development in information retrieval , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 48th international ACM SIGIR conference on research and development in information retrieval , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:ae45e9700e3f07e2185dee5cc203082eed6876fc1d12eaf090673935752bca05

Observation 9266dbf4-b5f7-4987-9eda-3006827c5273 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:6db1611bf16c9af8e23fea9dea0c2838ce6da148e2eeefaeae973e03b0db70e3

Observation 9449c600-a6b1-46f4-970d-158ea7559a7c · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:954973d8e4686165a2cbaea17395537e4cec9e46f1189e3e90818cd92569b268

Observation 39c3bc25-6d97-4a42-8a7d-5398cd0f68d6 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:339a4e6ae6bf425f3c0518d10991593c2051bedadab0576789c2364be1e52bf8

Observation 670d5b3d-9b21-4e27-9358-59f543957974 · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:5c27dfb8fe7be259e76a07de8456eb0d8e34d3dfbc1cd1b6621870c20b34982e

Observation ba733c62-eea7-47ec-ae6c-25dc6e4a0a15 · outbound

This paper cites CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:9953e507e4332c792cc14d19cad9d6a5d6805c3d26d561cd3c0f4bfc38f5df73

Observation 1ac10463-127a-45d7-b11c-e78b436dfdff · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval The Eleventh International Conference on Learning Representations , year=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:95335e4dae6c1d68709b1f64fd546517ad4aaa55d502ca633a10fde41ca4c2d2

Observation 809f288d-ee8b-41b1-851b-d3d07b433fbd · outbound

This paper cites Qwen3-VL Technical Report.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Qwen3-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:0636b72e57ebd719b2a1b8260c102d925aec7d824dab212eefe82b1fabd8cee8

Observation 12098874-ec8f-4ca9-988a-52382304447c · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Transactions of the Association for Computational Linguistics , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:d379ef5c0b28b8f7914e853dbe57f996f3233a620e29198b6f86ba72c31cf0cd

Observation 70a783c4-b83b-49c6-b3a1-bf9c97b06cb4 · outbound

This paper cites Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:78d424bd1d231ce0138bef0f403bdd1c835a22b0dd94efdb56df71ecd7356d8c

Observation 60b43be1-021a-4c63-8c88-cc95ab52d0b8 · outbound

This paper cites European Conference on Computer Vision , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , year =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:55bfa6fa710e4ab25f3c3c785823a26419e4f64bb41e9814bddab16eca48628e

Observation 2ad16f05-d915-4fa0-b319-46bfaae20e41 · outbound

This paper cites Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:422251e601ee47e5b0058aed70cb9bc045826540805ff5f9c8d2f90823249c42

Observation e0eb666d-a6bc-4003-b48a-ee38e51cb817 · outbound

This paper cites Image Retrieval from Contextual Descriptions.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Image Retrieval from Contextual Descriptions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:cd6fdec2209998ee018d75a560c01e1e36cb16dfc37d70cfe9e048e24fef3b34

Observation 1cac47bf-ff03-43c1-868c-ccc2e7c41ce8 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:af123ceb7f7d3dc943db59dc0dd84cb70f0f57aec5021d14c0ded9e041aab934

Observation add979f7-4818-4453-af91-c5d782d59624 · outbound

This paper cites European Conference on Computer Vision , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , year =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:af30ba24c7f2478dd7ffa34e73493ad55faa0ba77cf275156b1ac33889d9af19

Observation 8513aa62-b565-4949-a652-732b1da7900e · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:346183058819c0b8ca1d38bf9dc21fd9a210843bcf7cde0acf5e0fcf155034e1

Observation a1735af3-c2dd-4dae-816b-f95839cfb5de · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:312b714266774396ea24a2c8b374832b1585333e66961a40e9fbc8776b155617

Observation b15c0d7a-21d6-4694-ab37-c72993e0f7f5 · outbound

This paper cites Demystifying.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Demystifying

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:96e2690eeccf6affc1d9f9641665e866c35e0fd633e95ee68b6a8fc6c732a764

Observation fce11417-3c5a-4d9b-8770-bb7346f16fab · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:9c87367218f6bb6f6eca08d9fbfa50ccb89d7254ef67d7c415abcae16f6e2ad1

Observation 3284f6fb-af5d-41e2-abd0-011f00031f52 · outbound

This paper cites Data Filtering Networks.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Data Filtering Networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:97532c9710e9af9ce0c1890a1c6075ac25542a259e131025a5916efffd43bb95

Observation be27854f-dab2-4f0f-9551-9a0f353d78fc · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:ee0144823f831d376ebbcbf828448bf4764c37c2fded49f18fc47359ff073507

Observation e11db6d0-a67e-4419-82ce-a54400f900a8 · outbound

This paper cites 2026 , url=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval 2026 , url=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:14142dac0fa21c1fb916c9670585c7658a6776bf924d305bd5e0e103b26deab5

Observation 2a4f119f-b87f-4274-a54c-eb44ec712845 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:669d4dfea2f316b22caaf681167bb62f805772af31d950f7db6abbe3d3032f3e

Observation 2093b4af-f076-44b3-be8c-cac7b1b3fc25 · outbound

This paper cites an unresolved cited work.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:7b95144d08af55f7396fa912e5874d6ab80cd00ea56b9a649acd4673b929ab76

Observation 89aff243-e945-455a-badc-a74de29dede0 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:206f004adea2b0163cb240bf3b380ddcdba08246f26dc271211037c341176ec7

Observation d6f69e71-bdc3-4e4c-be96-b8a33365b224 · outbound

This paper cites Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:8b0bff0327dff2370dd892033b3ea42e838b72f7815638389d9a4ead595c62b9

Observation 63c43d9d-f918-4abb-bc8d-6d9fe6106649 · outbound

This paper cites International Journal of Computer Vision , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International Journal of Computer Vision , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:6fa68df04c431553ccd75779ea690da3243be27ada6adc18b100e2c9910c6e6e

Observation 6e07c2b9-3d19-496d-9eb2-d3df3499d24a · outbound

This paper cites European Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:18860d880c59ff9f0adfe440377d751822dc2f5eccf5152d9fc999c169a69ade

Observation ca750301-05a1-4f8f-b567-07d07ef89b58 · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:c5bf45756baa40c1597fbd9bc214318f478c1b7766dff7441e58fa15a8266e0f

Observation 33d6b6a4-3f95-482f-8501-340b202a0e67 · outbound

This paper cites International journal of computer vision , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International journal of computer vision , volume=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:d185a5d41263cbc974471ceba8ccbf373de883cd597bbada45736c1b53aae0b4

Observation e1643704-35b0-4bae-a2fa-af4f52a4ac79 · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:727599c701869e1994c736223cd7590b6993d06914ab314c45f11ca3c01374d7

Observation 9ada24e3-bb8a-437b-9214-f508d1bb776a · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in Neural Information Processing Systems , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:26494eb29eba0a954f1db8ffaa40bd0f811e038fdbcd344a509a59914c1118d8

Observation bec6475e-e5dd-41ea-a8d4-a8d073976247 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in Neural Information Processing Systems , volume=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:1d5c01ba7f1427e9c9e21ec2130d1a018ccc3a4864e5e9c9324278966581bb71

Observation d5f3c0f1-6f25-4cf9-b2af-10f80fbbf73b · outbound

This paper cites International Conference on Learning Representations , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International Conference on Learning Representations , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:4ec89974cc54f0d5199e60c409a6ce0123dd54b6c1d4340519667eb6ba062e8d

Observation 855dbae7-fd2f-492f-8af8-23e202e544a5 · outbound

This paper cites European Conference on Computer Vision , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:cc6c137c98f9cf6379d412eb671e6ed2a387eb26223188ba6c71d7a48158be47

Observation c5fccd8b-8d63-44a8-ba83-20b216e24dae · outbound

This paper cites Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:b120f0876b893975fc4203914e7c162ec079060d3defcfb554d7061bde1b41da

Observation 507d476f-c002-4c6e-9e12-cd8785215232 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:3ba978a64f65b501cfedeb930e5e3eb475f6a5d7814a761c665c2bb0339ddc48

Observation b592cfa8-028d-4429-84a4-754274ce5bb4 · outbound

This paper cites European Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:c3d4e9926d2791b01f47b73532324a65b2bda798a3c1c7b92434ac8a20a8d7ac

Observation ffea91b0-0d09-4c19-9b2e-94e75984bced · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:2f07e90075439bb57d05537fc15eb65e459aae807317aa0f150a86c50df09b74

Observation 3f1faae1-c479-4901-b8b4-7231f7803286 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:413cd77116006ff4f57c8b722a13084b4ab285a2b41e36fd21d7a22350c4a4f4

Observation cd18aa89-36ae-415c-bdd9-d87f34618c89 · outbound

This paper cites Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:eebe4cbc7ca2a9328b1f7caebd838e246fc43ee4a31c6a781d24cbe0c70d2db2

Observation 8426ecb0-5c54-48a3-ac15-08cc4b119d94 · outbound

This paper cites Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:7990d03127898cd343aafcd4d5fffdd7351cc9110a600bb11a8040129e2d5f80

Observation 4ce0c88e-ec0b-4238-ad56-e06169975b3b · outbound

This paper cites 2026 , eprint=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval 2026 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:7eeaa5168e9317a4309796a7e763bf1188a096c6661594204a0617199b224984

Observation c8c7517a-ecdb-4aa0-8fe8-e89abb0f064e · outbound

This paper cites arXiv preprint arXiv:2602.21202 , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval arXiv preprint arXiv:2602.21202 , year=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:e996f04e70170b380a1c1422f32eee3a14c75500f49629d7d2c8e69501984f6a

Observation 33f6898f-e8c1-46b6-a794-b7fb5ad3a4ab · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Representation Learning with Contrastive Predictive Coding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:5a8c4cf293b2ff6a0afe7dfa84441c99ec50efd5d3f40935a3782c0ee69e2025

Observation c6cf9eae-3f4a-466a-b2c9-dd54554f0904 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in Neural Information Processing Systems , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:52d0183b42487751c7bd409af27c983880cbbd8d1d991da309f027331d9ed68e

Observation d298ddc2-cea4-40df-ae3d-2ae99672d71e · outbound

This paper cites 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:88f7deb631379f128c11ec7cf0fbc4c65f8e1a23958ef237f3c59c7c24b8ca01

Pith citing papers

No inbound Pith citation observations are available.