Pith. sign in

Paper Citation Record · LEDGER

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2607.23052.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23052 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:51:55.185610Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58a3f014-74e5-45b1-8ce7-25ac52ce3384 · outbound

This paper cites H., Kim, Y., and Ghassemi, M.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models H., Kim, Y., and Ghassemi, M

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.040699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.040699Z digest=sha256:6b5487aa9de29112877130734b66f443e5b9f92bee657afcc617ab449700d160

Observation 13533fab-9929-48d4-8652-9622f5cb2882 · outbound

This paper cites Conceptual 12 M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Conceptual 12 M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.047256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.047256Z digest=sha256:5e43765d98f5707d7f88253a438b228f86e258fb7644609df08430ea31c251f4

Observation aa921af7-d115-45ae-8bea-18dc7d70a05b · outbound

This paper cites K., Winn, J., and Zisserman, A.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models K., Winn, J., and Zisserman, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.052134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.052134Z digest=sha256:fe2c587462488e96601219b4dcac36c8c8654e36f085189fb56138ff8c77ed75

Observation 24551c83-1460-4f98-aeba-40aacfa03330 · outbound

This paper cites and Kembhavi, A.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models and Kembhavi, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.056791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.056791Z digest=sha256:2440082d5bc865d14fbe0dd66dee573fccdfa39418e6d50241d4a6802025e4b0

Observation 6b3108a0-c8d8-4af3-8ee9-ba1db47fe5ce · outbound

This paper cites SugarCrepe : Fixing hackable benchmarks for vision-language compositionality.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models SugarCrepe : Fixing hackable benchmarks for vision-language compositionality

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.061416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.061416Z digest=sha256:49dcf94e599d59fbfb9475579aaa951c8a5443f95ab2d52ea4d112cf9f8f7b3e

Observation 7f541cde-4f27-4976-8d43-b2f191a3c8af · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Scaling up visual and vision-language representation learning with noisy text supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.065958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.065958Z digest=sha256:866e9cb2f2dab5912a94db191ecc6718fc652877af00af7a5474502cdd9bee1d

Observation caabac4d-b5dc-4fe0-a7e6-3ef28c0466ab · outbound

This paper cites ComCLIP : Training-free compositional image and text matching.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models ComCLIP : Training-free compositional image and text matching

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.071221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.071221Z digest=sha256:eabbd76bf9510186f329a59ef99d1d7c0a0f1357a5f96f251bf31cfc0246f027

Observation 7c3c3956-8da3-43ee-ace1-f8484f7e78ee · outbound

This paper cites The hard positive truth about vision-language compositionality.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models The hard positive truth about vision-language compositionality

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.075416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.075416Z digest=sha256:f48d1c25c70f07f4e42c821cc25443edaeadf90a3d6af850c07eefefdfdc5c7a

Observation cae2886b-b8dc-4e43-88a5-80d1d012052a · outbound

This paper cites Is CLIP ideal? No.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Is CLIP ideal? No

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.079762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.079762Z digest=sha256:b70c3b3c2f562c700b70f3a6065e88493d409e5ed910d9a5614fed4ecc2fddeb

Observation 2a007bba-9c7a-4924-9416-19d03327575a · outbound

This paper cites an unresolved cited work.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.084523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.084523Z digest=sha256:73ff9bfcf868328d1fca1d0843d2b5dc2f41faebb090ffc51cedc40f9499486e

Observation 74fe2927-dffc-435c-a646-a8a339fd6222 · outbound

This paper cites Does CLIP bind concepts? probing compositionality in large image models.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Does CLIP bind concepts? probing compositionality in large image models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.088705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.088705Z digest=sha256:c160fb1a02ce07407eaa66ec263d761ce7774457b0c1ef972ad63383f03af3a0

Observation 6ae2b34e-a7e2-43b7-83b3-c45d0cc2ff4d · outbound

This paper cites an unresolved cited work.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.093182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.093182Z digest=sha256:00a371d88593e74cb861809c026ee7344d090550a674577c9ca969fcd52e6f04

Observation d12192ed-7958-4fb0-8741-fdf4ed35a1ff · outbound

This paper cites O., Gandhi, M., Gao, I., and Krishna, R.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models O., Gandhi, M., Gao, I., and Krishna, R

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.097450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.097450Z digest=sha256:8191374228eb73bf29b092f6e415f783f448c9933a0c1528c98a23643bc985d0

Observation 1090c92c-04dd-4290-ac91-10cb4884c3b4 · outbound

This paper cites Simple open-vocabulary object detection.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Simple open-vocabulary object detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.101937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.101937Z digest=sha256:904b0da5baafddde25cb7e12b176df444c27ffae9be751121f70812933ae43d8

Observation 4c1da4b2-69ae-424d-b48e-15ffc98e7c42 · outbound

This paper cites GPT-4 Technical Report.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.106219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.106219Z digest=sha256:05f60bddd8b54d7e0cf88bde738194ef9524cca9a1f36b6ee9a7e66effdfd238

Observation 09aec2fb-9d41-4208-84ee-822341bc6ff6 · outbound

This paper cites Know `` No '' better: A data-driven approach for enhancing negation awareness in CLIP.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Know `` No '' better: A data-driven approach for enhancing negation awareness in CLIP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.111104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.111104Z digest=sha256:a7bfa13aa3211154a7af46e831f7368ff9578a27448c256646ad96aff5e8b6b8

Observation 2b9ca59a-fa36-472f-ba1c-06d66830cb88 · outbound

This paper cites D., and Hein, M.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models D., and Hein, M

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.115548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.115548Z digest=sha256:36c725247b64dd6b6371294ad1f25e19591e403fa00bbe1859625398c5122eb5

Observation fd4964f9-92eb-4ed3-9f96-2fc34e05430e · outbound

This paper cites How and where does CLIP process negation? In Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR), 2024.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models How and where does CLIP process negation? In Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR), 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.119841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.119841Z digest=sha256:4f271b8fd48c3064cfb2b148a48d4cfba6a8a5d8602a764bb6f34e9ab0826f7b

Observation 95082a60-a57d-4155-a591-c13bd1f6ff7d · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.124066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.124066Z digest=sha256:c3869004f563965e56ed55956e294d95dde7e54f525e80eeb895819713490f82

Observation 46766c74-500f-48df-bcc8-3207d26a422e · outbound

This paper cites Collecting image annotations using A mazon ' s M echanical T urk.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Collecting image annotations using A mazon ' s M echanical T urk

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.128319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.128319Z digest=sha256:e28a32dc7bf3d05d368bbc0295b4e55f435f8c0c23a5de485d363084d191f9b4

Observation fa702f0c-dd86-4801-843f-45e1e28413ec · outbound

This paper cites T., Argus, M., Fischer, V., and Brox, T.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models T., Argus, M., Fischer, V., and Brox, T

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.132836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.132836Z digest=sha256:ce40a131402e8792e8edaf7e06baa71e40e3d263665c34f66be1fc7eeeea1f4d

Observation 9d4f5c1b-8fab-4d58-817c-c5001745e9c2 · outbound

This paper cites Learning the power of `` No '': Foundation models with negations.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Learning the power of `` No '': Foundation models with negations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.137512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.137512Z digest=sha256:94cb90bbef0bad2050b78730cf437598129dfc93eed67d5383d5b312d92b9333

Observation 86b8ee81-e9e4-48b9-8ace-1dafd5305f5f · outbound

This paper cites ViperGPT : Visual inference via Python execution for reasoning.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models ViperGPT : Visual inference via Python execution for reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.142033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.142033Z digest=sha256:6173f3110c54e1ae9218147333147767c521b36a330daf11da5d60c44552d0c9

Observation 7872f528-de3f-4204-9fbc-128d58e54e64 · outbound

This paper cites Winoground : Probing vision and language models for visio-linguistic compositionality.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Winoground : Probing vision and language models for visio-linguistic compositionality

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.146312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.146312Z digest=sha256:73498d86e26385c55692d147493f5c01622c3cdccc991a9cb76ecb451cca4880

Observation 53d5e963-cea0-4615-b065-5da6c47c9663 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.150635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.150635Z digest=sha256:ec1a5a83e24a26c044904db7843c22d84ea9acfa136e72c0f0f7aa818c54ed7b

Observation a79a2f97-e8b3-4ba6-8924-74995bc2423d · outbound

This paper cites A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.155521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.155521Z digest=sha256:00c0d10aa466e0b5ae54c2130a53113ed6b6717cf7564a5d8bdb994ce0543a19

Observation d70ffbe1-fcd4-4dca-9c33-d94f8b4560e2 · outbound

This paper cites Y., Lee, M.-L., et al.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Y., Lee, M.-L., et al

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.160075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.160075Z digest=sha256:2912c4955cb2e3574590fc83f8f0b72ca5455c05e3eb73029a7d4e493e64b7eb

Observation 36eed525-17d3-4ff9-a9d6-4ae99f7e1fd5 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.164387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.164387Z digest=sha256:d0d1b96eb95810a8ac6e07f1e5da16869c0253310b89b25d76bb638e3340fdf5

Observation d9dbb6c2-5cd6-4cfd-8377-75206d1f2dfb · outbound

This paper cites LiT : Zero-shot transfer with locked-image text tuning.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models LiT : Zero-shot transfer with locked-image text tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.168632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.168632Z digest=sha256:5d4d90776aa4c970a7c8676be24f4f2f67a2e73b23aa94928da39c3dd885fd34

Observation 28425009-383c-4a7c-8af8-eefeed24dcb8 · outbound

This paper cites Sigmoid loss for language image pre-training.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Sigmoid loss for language image pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.172818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.172818Z digest=sha256:be50c7a633d2ca69c1633f9c4bb03a63e94d69576a7e831c445f02dbced876f4

Observation 79305521-3cee-41e5-8875-166d05a4e18a · outbound

This paper cites NegVQA : Can vision language models understand negation? In Findings of the Association for Computational Linguistics: ACL 2025, 2025.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models NegVQA : Can vision language models understand negation? In Findings of the Association for Computational Linguistics: ACL 2025, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.177250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.177250Z digest=sha256:0f31c23502329de7600dab90e7c9ec504ea4d21344b0137ee9ce5db391f6b604

Observation 277d1105-0be4-4aac-862f-7ad1ef78776f · outbound

This paper cites VL-CheckList : Evaluating pre-trained vision-language models with objects, attributes and relations.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models VL-CheckList : Evaluating pre-trained vision-language models with objects, attributes and relations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.181466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.181466Z digest=sha256:5ce75c8ed427c8d4829f9ef771c1ec78944deda0142ee3848954fc0bc2344ad2

Observation 23fbf513-596f-416a-a010-1e59a8d823ce · outbound

This paper cites Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.185610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.185610Z digest=sha256:6870ec33b4fdc98a7a6ca5493c0344de1cc27509edd176c72a7a1b0d7d131348

Pith citing papers

No inbound Pith citation observations are available.