Pith. sign in

Paper Citation Record · LEDGER

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2607.04605.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04605 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T16:20:19.874172Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d69e9b98-ba0a-4fb7-8bb8-3afc01558228 · outbound

This paper cites Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:0f32cf15d835d42e11d8cfb3035ea8453793bf7c4f81751c08c1e748621a9973

Observation c3db9897-eddf-421d-913c-50aeaa3018d4 · outbound

This paper cites Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:691e47d2e164a46ffcb843e888e52a903eab1ea10671f2df8dbccceba11469a5

Observation e9abbfd0-264e-45b9-b99f-5ba875813938 · outbound

This paper cites ColPali: Efficient Document Retrieval with Vision Language Models.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval ColPali: Efficient Document Retrieval with Vision Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:e5873e5dba85ca515cceeec8702ec1349e70e1c0aed87e528ea0d6a4d8b8a194

Observation 614baf50-3197-4e6c-a184-e83cc3385e28 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:f982dfc731eb9b0e3e70714d8d028fc92e8bffb3829907b440c65a959f52c805

Observation e6459299-2443-4e5d-bf89-3fddb94388f4 · outbound

This paper cites International Conference on Learning Representations , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International Conference on Learning Representations , year=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:39d85d3a9aebec5f5415a375256e36256d1a5acade37c2e07bdf4c5c1e58f333

Observation 72881047-42d5-48b4-ae9d-5f032e5f9f77 · outbound

This paper cites Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:7e49458682e7e85eb85497c6d885478afa53cee89c86b53de507a335aa38670e

Observation 2b44ce88-7803-4bd0-bb2d-2987fc49004f · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:baecb15280dcb68eccb2b39f631436ba1b8c8524574597b5bc0bd9999f0718ec

Observation c5d226f0-ce47-4fe1-8b2f-30c48a48e1e0 · outbound

This paper cites Advances in neural information processing systems , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in neural information processing systems , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:b80c86316114cdaa1d1af60c4e41b83d868633b4bb5af2008267a5e1d04f1400

Observation 7703ee07-54d3-43de-b8f2-4f0b2f26142f · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:ac28b577b7c13f257a998b43797e90286b971f93498c2f9a2eec9e7d006c75d2

Observation dc1234fc-5b86-48a6-be92-0163d25f7af0 · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:b195354ec4ef16a871252de6efddff5a760478165710feed0193f17aa95f7e53

Observation e80213c8-c79a-44db-a381-bb22539273e3 · outbound

This paper cites Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:ee1f069b78b9c3dae4c0fd40ef69541bac20fef4e05f4cc2d57d1be6910fdd41

Observation 4ac8931e-7d9d-4a42-8e35-459216666c32 · outbound

This paper cites Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:748056c5147bd7aa8342714bc5b89d11d7d27eece7e29120d5b19dbb2fa9f4e0

Observation 09587334-a962-4f58-b348-dfc3f3705406 · outbound

This paper cites SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:4841daf0886e74d634a10c3c865eeda57c772a8dfbf29c2a67662e5d727fbe06

Observation 48358023-8a27-4865-8203-861a2b2f349a · outbound

This paper cites Proceedings of the 31st ACM International Conference on Information & Knowledge Management , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 31st ACM International Conference on Information & Knowledge Management , pages=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:f4ce4ba65c58a364dddf0da91781bf1df7d457ed6b30cc06ac7dd2a614bc361b

Observation f1bdc37e-1621-412a-8f23-79f1bf2b187f · outbound

This paper cites Proceedings of the 48th international ACM SIGIR conference on research and development in information retrieval , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 48th international ACM SIGIR conference on research and development in information retrieval , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:e2a02c40cd6d46e506c676cd2550a00b7b4351ee7f99e686ac7dde4604137be8

Observation 9266dbf4-b5f7-4987-9eda-3006827c5273 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:83f01a8589338abd0dd10a0a4fcb6ae066d9c667230efc2e34c2971dd15f5d3f

Observation 9449c600-a6b1-46f4-970d-158ea7559a7c · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:dc32c398352823691a791f29e35d6bb3424515dc0ae47955eea3639bdc013569

Observation 39c3bc25-6d97-4a42-8a7d-5398cd0f68d6 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:b6095d171009f70f0e3012a49469fdc1d16c48680f83b94e512c7d455893688c

Observation 670d5b3d-9b21-4e27-9358-59f543957974 · outbound

This paper cites ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:36e011d460955aa08dbdfc1c280fca791e16084548e9761088ef0993bd3acc69

Observation ba733c62-eea7-47ec-ae6c-25dc6e4a0a15 · outbound

This paper cites CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:86348d399ec685deffa9980f58761f34123d3a8271d112619eae3d7871cecd1f

Observation 1ac10463-127a-45d7-b11c-e78b436dfdff · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval The Eleventh International Conference on Learning Representations , year=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:791e38e1379e73b717444419593073a2f3c3ab578084b79db5b511b103d3ba60

Observation 809f288d-ee8b-41b1-851b-d3d07b433fbd · outbound

This paper cites Qwen3-VL Technical Report.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Qwen3-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:de40532bdb54a1fd45fff59cb69d9b0d128be1be7a2b54d193db8493dc3bb5e0

Observation 12098874-ec8f-4ca9-988a-52382304447c · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Transactions of the Association for Computational Linguistics , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:39d52c100c49418583b66385762bd1ecf8d902d7eda1d1f61217907ebcd670e3

Observation 70a783c4-b83b-49c6-b3a1-bf9c97b06cb4 · outbound

This paper cites Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:040267afb6bf161681c989eba15e989835b76abba0f55595fa7f7c58921caeb8

Observation 60b43be1-021a-4c63-8c88-cc95ab52d0b8 · outbound

This paper cites European Conference on Computer Vision , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , year =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:8daa9280834e7309cbe3751f2ebcc788d93266a851cfc31e6b774a3698451b9d

Observation 2ad16f05-d915-4fa0-b319-46bfaae20e41 · outbound

This paper cites Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:f3731d0530c46870cb21da17403bbd99c91ee8959064ef1c6c1c4d7fb06dca63

Observation e0eb666d-a6bc-4003-b48a-ee38e51cb817 · outbound

This paper cites Image Retrieval from Contextual Descriptions.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Image Retrieval from Contextual Descriptions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:b6eb02a9573dee0aa3d5e585390a6ce18361de67bd4ed864384401653c961aaf

Observation 1cac47bf-ff03-43c1-868c-ccc2e7c41ce8 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:2ea24e32a41331bd6a76bb04099be7e6c291661175d79fbfdcd479e66ebcc160

Observation add979f7-4818-4453-af91-c5d782d59624 · outbound

This paper cites European Conference on Computer Vision , year =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , year =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:fcdac765e21ae3a2f3eedd233654b7f47c0797b621577d9eb08f216bdbec2bde

Observation 8513aa62-b565-4949-a652-732b1da7900e · outbound

This paper cites International conference on machine learning , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International conference on machine learning , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:74af51248ae3d02f32c16c39864ef79c8370ab3481423bdbc370d1b45c88ed04

Observation a1735af3-c2dd-4dae-816b-f95839cfb5de · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:05c203bac2fffca5e44a9e637b8574f25fa816e54ecc7c9704028a07f33cfcab

Observation b15c0d7a-21d6-4694-ab37-c72993e0f7f5 · outbound

This paper cites Demystifying.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Demystifying

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:f7715e64468d07d314b04bda03b1bcaffeafd4d25fa2738a1519f9d1cb002a24

Observation fce11417-3c5a-4d9b-8770-bb7346f16fab · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:320c3c51e9a36560dfcad6e478faa72567c3e275b9e79dc721fa656881235526

Observation 3284f6fb-af5d-41e2-abd0-011f00031f52 · outbound

This paper cites Data Filtering Networks.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Data Filtering Networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:c834b31280b3aa7a0fad20eaff341216553df5f81819fb9377712feeb0ed5859

Observation be27854f-dab2-4f0f-9551-9a0f353d78fc · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:dc51413499a201c08f51eefe4edce4a4c18a1a77b1fe90b1fdc89e60c9006006

Observation e11db6d0-a67e-4419-82ce-a54400f900a8 · outbound

This paper cites 2026 , url=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval 2026 , url=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:a41dfa14a3751dfec564a84901c6bec5c4c7c5e36d6fe553bb7e4da1833591db

Observation 2a4f119f-b87f-4274-a54c-eb44ec712845 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:e87c27b6b85f956938ed9dd5bca1a8c142c18b95d0cec03615d2cf8e58c8a2f2

Observation 2093b4af-f076-44b3-be8c-cac7b1b3fc25 · outbound

This paper cites an unresolved cited work.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:bb52188681b9bdd85d945890ca1bd62ae9ec423c5af06141f7f15291a9ebbf40

Observation 89aff243-e945-455a-badc-a74de29dede0 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:638443e9f17fad7a536782c8a09f3c27c71aef58a56d5bd59a948c322c35ef59

Observation d6f69e71-bdc3-4e4c-be96-b8a33365b224 · outbound

This paper cites Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:88985918e8e15318b4e66576066f47516e98f4bdbf3bbd45477de18c27552a2f

Observation 63c43d9d-f918-4abb-bc8d-6d9fe6106649 · outbound

This paper cites International Journal of Computer Vision , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International Journal of Computer Vision , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:0121c608ced988d4dbddfb02de50d288d38789ba69f0411675f7189f555c6b6e

Observation 6e07c2b9-3d19-496d-9eb2-d3df3499d24a · outbound

This paper cites European Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:3903f9b994df1fddf8ac12000848b45ade25db2111538a8d4d6bb89e1fdc8386

Observation ca750301-05a1-4f8f-b567-07d07ef89b58 · outbound

This paper cites Proceedings of the IEEE conference on computer vision and pattern recognition , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:8753d923bf70939fc82f629603d07e736e1ec5c9bb936e856fb1192113fb8a3c

Observation 33d6b6a4-3f95-482f-8501-340b202a0e67 · outbound

This paper cites International journal of computer vision , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International journal of computer vision , volume=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:a82eea563f64e6043d886490305ea941c39b1319e706b4037ef37c2dca9252ae

Observation e1643704-35b0-4bae-a2fa-af4f52a4ac79 · outbound

This paper cites Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics , pages=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:fe46cdf37de1f17f4130aa975b22e6f817cf58eb34c3c049568a9f0195a2fc9f

Observation 9ada24e3-bb8a-437b-9214-f508d1bb776a · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in Neural Information Processing Systems , year=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:cbf487c584769312f96821ffccdd9d397cbf4bc48c1b8c58e14d5342b0d5b9de

Observation bec6475e-e5dd-41ea-a8d4-a8d073976247 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in Neural Information Processing Systems , volume=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:b3b645d40a28b321c22bf178e21a8ef5c2d27892c2a68dcf20c706645c667792

Observation d5f3c0f1-6f25-4cf9-b2af-10f80fbbf73b · outbound

This paper cites International Conference on Learning Representations , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval International Conference on Learning Representations , year=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:91d7198f22130685792f1460b2fcdbf3abbacaa1a233357911bf34753dfd0927

Observation 855dbae7-fd2f-492f-8af8-23e202e544a5 · outbound

This paper cites European Conference on Computer Vision , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:d5768279d6f869ea409fb64527e0456f6ebebeb0d558b028b00a6b9495aa6d27

Observation c5fccd8b-8d63-44a8-ba83-20b216e24dae · outbound

This paper cites Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:6bf071a46bb6e1072c636d19dfb6b15c0bfa88e75b0bf9d9ea39ef6eac2eaf82

Observation 507d476f-c002-4c6e-9e12-cd8785215232 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:bc8d075c60aecf4be37ca9a9cdd4ca8596283e2b219ba456493dfafed94ce0c9

Observation b592cfa8-028d-4429-84a4-754274ce5bb4 · outbound

This paper cites European Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval European Conference on Computer Vision , pages=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:51d4fac5d9b96d3348ca3302c6ec55311b6c4b32b55354d398efe3250a8df418

Observation ffea91b0-0d09-4c19-9b2e-94e75984bced · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:cec127a5fe99d30669f30f1e13fe7a71f4ad5e60963580ad363426e395a4bae8

Observation 3f1faae1-c479-4901-b8b4-7231f7803286 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:128694fc4f17755e6fc3fb2c3309ee6b38e9b2f8b71796c5440de85a34672694

Observation cd18aa89-36ae-415c-bdd9-d87f34618c89 · outbound

This paper cites Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:713820f9b5c736ef592138565ad2669ecbafb1e6e7f387761af946dc5dfe2f39

Observation 8426ecb0-5c54-48a3-ac15-08cc4b119d94 · outbound

This paper cites Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:3fd6ee2914fe49d2cd1e59ff7a0e16b9c8f64f44700ab212e8081740bec80473

Observation 4ce0c88e-ec0b-4238-ad56-e06169975b3b · outbound

This paper cites 2026 , eprint=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval 2026 , eprint=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:d6c61c9e65a3a13c9af5151364aaa4f1137b8929a7b09025ebd5a8c9cf4a06fc

Observation c8c7517a-ecdb-4aa0-8fe8-e89abb0f064e · outbound

This paper cites arXiv preprint arXiv:2602.21202 , year=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval arXiv preprint arXiv:2602.21202 , year=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:b4585de6a5f51c6b967a30fdac8e7c2ee26b7537e7554078f700aae9503785c2

Observation 33f6898f-e8c1-46b6-a794-b7fb5ad3a4ab · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Representation Learning with Contrastive Predictive Coding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:1a400fe633dfb8110120e04908136231f76bd9902cf299c004b304ebc9217496

Observation c6cf9eae-3f4a-466a-b2c9-dd54554f0904 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Advances in Neural Information Processing Systems , volume=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:cd78b6ba91dec659949d910f0d47143bc51b075d04f9ebb2feb5ee47839ed55e

Observation d298ddc2-cea4-40df-ae3d-2ae99672d71e · outbound

This paper cites 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages=.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:c0f1342ffe7431acd0de138b5b0d8f1992abf89312ac60d7d9d329c41ef6ea28

Pith citing papers

No inbound Pith citation observations are available.