Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models Do Not Understand Negation

As of 11 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 8 inbound Pith citation observations for arXiv:2501.09425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09425 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:08:46.326666Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:58.344234Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T00:18:29.517494Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved11
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 876019cf-4e11-484d-93b7-b82286f09e37 · outbound

This paper cites Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes.

Vision-Language Models Do Not Understand Negation Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.570497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.008226Z digest=sha256:b18b04ce1ba844e12746755c0daf68bb4baaf95a62b3593a69b8012540a3c61a

Observation 76626a12-b8c4-4540-861e-c24ddca9336b · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

Vision-Language Models Do Not Understand Negation Effective conditioned and composed im- age retrieval combining clip-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.548229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.014749Z digest=sha256:e2725d5cde7aac96eea784a7569d58bd22374f56256ee2b6ac0fffa0dc1d1803

Observation 5d08414e-b6b6-492a-ac86-dbac81d7a1fe · outbound

This paper cites FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks.

Vision-Language Models Do Not Understand Negation FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.517244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.020123Z digest=sha256:484bc54d2b2747c6e6211b933b438e5b7222f3363d2134cf1f0bfcf7812adb5e

Observation adc40a9b-4ad3-44b7-96a0-b98ec2b7ff5d · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Vision-Language Models Do Not Understand Negation Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.031905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.031905Z digest=sha256:eed88d3a8a3f462c3215d39c48a5d15dfa751a60d6271700df4190e6888d1664

Observation 19db9dd0-2961-4ed7-ad64-5bae5646be84 · outbound

This paper cites Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach.

Vision-Language Models Do Not Understand Negation Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.453953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.037424Z digest=sha256:fff9086205beec5f48d17e380d451d9f150074b708354de833f686b74e7df60b

Observation 3e0cc5c6-38ab-4ba6-855b-8a80c4030f3d · outbound

This paper cites The Llama 3 Herd of Models.

Vision-Language Models Do Not Understand Negation The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.042836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.042836Z digest=sha256:cfac792575e04e846d8ee47965703e5e5481030466b17b0b525b1eadf05dc4b0

Observation 97d27b26-b7c2-45d5-ab5c-dbe351a3529a · outbound

This paper cites The pascal visual object classes (voc) challenge.

Vision-Language Models Do Not Understand Negation The pascal visual object classes (voc) challenge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.429941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.048534Z digest=sha256:7ac6d982573c3603df551ce1462e900d3901d3bbb2f36c89eed451dd2ec09b9f

Observation 11613278-2c26-4b96-9ae0-c64f4c515b2a · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Vision-Language Models Do Not Understand Negation Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.054919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.054919Z digest=sha256:aa53bee938da9c2e56b6c990c1de16cacc69daaf432efbf2ce3d677bb460d5e2

Observation 43de77e5-ed52-44ba-8ed4-68d1bedef7ab · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Vision-Language Models Do Not Understand Negation Datacomp: In search of the next generation of multimodal datasets

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.402544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.061090Z digest=sha256:def9b68aafd71474511f8417b4d3ce0568ec1d51b8b900d684e7939f254cb9ef

Observation e1117c67-0df2-4e42-a899-fb9d4b5c7bcd · outbound

This paper cites This is not a dataset: A large negation benchmark to challenge large language mod- els.

Vision-Language Models Do Not Understand Negation This is not a dataset: A large negation benchmark to challenge large language mod- els

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.382504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.076557Z digest=sha256:b1368aa7f8690e054956cd3321c9cd07eeb43135ccd75bcf3c7a897a3aa1aedb

Observation 504dc068-fbb3-4026-b1d4-87f3c4e9d131 · outbound

This paper cites Shortcut learning in deep neural networks.

Vision-Language Models Do Not Understand Negation Shortcut learning in deep neural networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.355708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.081422Z digest=sha256:cf0a79dd46899fe982171c1c62d5235f62c48c0b7db0f0699487f4eebfd69754

Observation 6ad3e1d3-2a2b-4509-be2f-5b1e12a3db5e · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

Vision-Language Models Do Not Understand Negation SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.086165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.086165Z digest=sha256:eb9b1809a007fb816150e6dc47e8aa4700b2f65706ac5da6d79af0b671558fca

Observation 863bde2d-1a1e-433d-bfcd-09a3e99ecd6e · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:47.335828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.091030Z digest=sha256:47deaca72c4fb6b1718ebb4fff1b0000acd9af224e7dee614dedabbc8c89c9a9

Observation fd484441-67ad-4796-b2bf-0cc61cb951c7 · outbound

This paper cites Quilt-1m: One million image-text pairs for histopathology.

Vision-Language Models Do Not Understand Negation Quilt-1m: One million image-text pairs for histopathology

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.314683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.096959Z digest=sha256:9fba4aa78fc730fcf358a56928a79a5e23c758a589907eeedaec95102cb781cc

Observation 73531a14-883c-4bee-842c-7c0fc3a33525 · outbound

This paper cites Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison.

Vision-Language Models Do Not Understand Negation Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.296473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.103106Z digest=sha256:9941aed5bb7a914822d34d7eeebbdbebb3297ad67bfa4e43cf15621d39061c6d

Observation a0594ace-e525-4638-b328-70feab20beaa · outbound

This paper cites Generative models as a data source for multiview representa- tion learning.

Vision-Language Models Do Not Understand Negation Generative models as a data source for multiview representa- tion learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.277193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.108782Z digest=sha256:a407ce39df5ff7fec10b7983cc74828f32b2f5d39817d0d251210e4f3fc9f32a

Observation 6825bbb5-161f-4b27-8c15-85c4acb02932 · outbound

This paper cites The power of negation in english: Text, context and relevance.

Vision-Language Models Do Not Understand Negation The power of negation in english: Text, context and relevance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.257383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.113421Z digest=sha256:6eb40653156f0f039d1e961c3699c84eee8172198bdd81717f5adaf9f047bf3c

Observation b40567b7-2ebc-44d1-a989-c1c016b6515d · outbound

This paper cites Negation in syntax–on the na- ture of functional categories and projections.

Vision-Language Models Do Not Understand Negation Negation in syntax–on the na- ture of functional categories and projections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.234939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.118134Z digest=sha256:1ac59853b7bdac6b831576593c08f6318b9a49803a6d7d5e0e44c494108a8139

Observation e4b8857b-2a24-492a-8b01-221b1f34db22 · outbound

This paper cites Naturalbench: Evalu- ating vision-language models on natural adversarial samples.

Vision-Language Models Do Not Understand Negation Naturalbench: Evalu- ating vision-language models on natural adversarial samples

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.214576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.122815Z digest=sha256:33433df8e8882f8a082c62b1281838b2058c5a7ab276c190a921c1d09747dcc1

Observation 51eb013c-e984-4379-8ec6-8ce2bcc29324 · outbound

This paper cites Compre- hending and ordering semantics for image captioning.

Vision-Language Models Do Not Understand Negation Compre- hending and ordering semantics for image captioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.194283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.128486Z digest=sha256:448490c072e77f5bb38832cd722caa8598b740e058bf26e23b78d1ff0ac7c523

Observation e2d73bba-3d18-4099-bf07-0007c3bfce36 · outbound

This paper cites Cross-modal retrieval and semantic re- finement for remote sensing image captioning.

Vision-Language Models Do Not Understand Negation Cross-modal retrieval and semantic re- finement for remote sensing image captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.174919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.133410Z digest=sha256:93053c527878be542f5a7ac669b84bc7594232e9bcc0a4a26874bba05f947ba3

Observation 9c371f82-ffe9-4c86-b92b-0595593f0e24 · outbound

This paper cites Microsoft coco: Common objects in context.

Vision-Language Models Do Not Understand Negation Microsoft coco: Common objects in context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.156422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.138330Z digest=sha256:472b4c2c4db69442a18abc83f3b0581e46bfc2969b1a279d469f7133df4366e5

Observation c376ff10-6159-490e-b8f8-24a60b6a017e · outbound

This paper cites A visual- language foundation model for computational pathology.

Vision-Language Models Do Not Understand Negation A visual- language foundation model for computational pathology

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.134814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.143444Z digest=sha256:284c2207e1bd58df296ba327b6f8ecc3adb46e5b4bf8f6ad269317efe09c3b0e

Observation 92c26dc6-4b83-4553-a6ca-f3662ab335d6 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

Vision-Language Models Do Not Understand Negation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.149164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.149164Z digest=sha256:e961533209c219976924d63f8eb74e2a54c8b6a37e27bcd8c1ba4114812f9607

Observation 46d0c712-e0cf-4ceb-9a8a-be0410d8928e · outbound

This paper cites Fine-tuning llama for multi-stage text retrieval.

Vision-Language Models Do Not Understand Negation Fine-tuning llama for multi-stage text retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.105551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.154599Z digest=sha256:dd8108e7885130e88f957cfe1e6ff549a7b29794f283b823b8550f81ad74c5e5

Observation e298a188-b5e0-4d97-b981-26f8a39358a5 · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023.

Vision-Language Models Do Not Understand Negation Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.077283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.159371Z digest=sha256:cf4b19c82fca64244c798b77f585a1ca4cd42804abfb2694c4b62b4ade14baff

Observation 367d25f1-9eed-41c0-bc92-6b3d86ddcd67 · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

Vision-Language Models Do Not Understand Negation Simple open-vocabulary object detection with vi- sion transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.046531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.164498Z digest=sha256:eccb8b4331c3dcfb7641d12add1d40b835534c9736ae7ec40bb99dcb84e22a3e

Observation 7bb530a6-5b4e-447e-983a-dc3ff370dd56 · outbound

This paper cites Recent advances in processing negation.

Vision-Language Models Do Not Understand Negation Recent advances in processing negation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.021702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.169611Z digest=sha256:c8aa32764270368e4aa6d1a704fa38df70c31212ace863023b0473a114da5460

Observation 98baac9d-107f-4d94-9a8e-fc3102221a5b · outbound

This paper cites Effect of negation in sentences on sentiment analy- sis and polarity detection.

Vision-Language Models Do Not Understand Negation Effect of negation in sentences on sentiment analy- sis and polarity detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.990672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.174089Z digest=sha256:1889e0a9fa5dc9ab19f79bfac78fa6310c20aac0bcc9ac14dc8c9f05ecc65dde

Observation 1e9c7f92-9608-4b9a-bbb0-8f78ca2618b7 · outbound

This paper cites Clip-it! language-guided video summarization.

Vision-Language Models Do Not Understand Negation Clip-it! language-guided video summarization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.967778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.178920Z digest=sha256:4c0bfeb85c36e117d13a8219ad45f7525ae8405b9acd234b5085e7a5e85616d7

Observation 8835517b-641d-4731-9c1e-a04012018817 · outbound

This paper cites Multi-Stage Document Ranking with BERT.

Vision-Language Models Do Not Understand Negation Multi-Stage Document Ranking with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.185183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.185183Z digest=sha256:3bb9ce323f917633856a939d4ba8a0f9b213dd0214b8650f059b339710f7f924

Observation 5869f8b3-04a3-4e44-a339-5be9332767af · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

Vision-Language Models Do Not Understand Negation Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.947397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.190882Z digest=sha256:105eb32c5649cf2039d0468c9bdb2c877dadcb0a436d592450668024a6ba2d99

Observation 3e17222d-e85e-4157-8b69-7f9b531211b2 · outbound

This paper cites On guiding vi- sual attention with language specification.

Vision-Language Models Do Not Understand Negation On guiding vi- sual attention with language specification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.926944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.198307Z digest=sha256:d126fbb267f12ed4bca2f903967e5881ce7c877851e9830897fbf80d9eb41a7a

Observation 877627b6-c6ca-4b13-89d2-df8a6aa91d57 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Vision-Language Models Do Not Understand Negation Learn- ing transferable visual models from natural language super- vision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.909588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.203619Z digest=sha256:d10b487474f4e6a46b6c02ad75cece021ad20e34ef815fca38724599fd73a320

Observation 1a058905-4956-4332-b0bb-e67c043cadab · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Vision-Language Models Do Not Understand Negation Denseclip: Language-guided dense prediction with context- aware prompting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.891949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.208659Z digest=sha256:23c0659ffd2850fa6209a1cd4b89e04e4f22c670a2172bb7025a012653471f48

Observation b2cf1896-10f9-4896-a3a8-551b340ce5b3 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Vision-Language Models Do Not Understand Negation Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.875485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.214930Z digest=sha256:8165778155ae228ef2e9b2a3ffe9a18ea4d45badcc1855737ec973765aca06ae

Observation 986b4a43-0592-49e9-b54d-5fdb299a5c7d · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Vision-Language Models Do Not Understand Negation High-resolution image syn- thesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.855890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.221528Z digest=sha256:710626c2b96c14d4a93be3dc2306c1fb88eac2394fe24cfc5e77848575ec634a

Observation 850b4137-6f56-4bb0-8303-dfb1b0199aa7 · outbound

This paper cites Clip for all things zero-shot sketch-based image retrieval, fine- grained or not.

Vision-Language Models Do Not Understand Negation Clip for all things zero-shot sketch-based image retrieval, fine- grained or not

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.837669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.226461Z digest=sha256:b07c99116a193feec9b74fe207fa84784942110482530c6102128ddced68fd48

Observation 1a841611-2507-45f8-b381-b5259404120b · outbound

This paper cites LAION-5b: An open large-scale dataset for train- ing next generation image-text models.

Vision-Language Models Do Not Understand Negation LAION-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.819673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.231797Z digest=sha256:a47f7c31b420b9eaaddae22f6454f1223114c595ec7804919861dbae7438759a

Observation b1117bc4-936e-443f-b7f3-fcd2367020e7 · outbound

This paper cites How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions.

Vision-Language Models Do Not Understand Negation How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.779307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.238053Z digest=sha256:a49c3f2f20ae44b0f34488a8605e9468f6741d04966c80fe3c63d3151e58682f

Observation 89029596-1604-4d2e-8baa-02b6076886a1 · outbound

This paper cites Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues.

Vision-Language Models Do Not Understand Negation Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.759276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.244630Z digest=sha256:0e5ff3b82a20b098a7c5d6236a8dbc849edf63a45eb894b24409aefa8fd50b4d

Observation 15f7bafa-b741-49ac-964c-95d392e8eba8 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Vision-Language Models Do Not Understand Negation Cliport: What and where pathways for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.739522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.250104Z digest=sha256:c67bf474b02f7f002ad95eaf6b825f8839992269bdbae092b37d0676674ef5b3

Observation 6b9ddda6-4d36-47c2-8059-4939e8fb9485 · outbound

This paper cites Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations.

Vision-Language Models Do Not Understand Negation Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.254987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.254987Z digest=sha256:4bbbeb76fbb908fa181b48ee45d83b48eb4da2e114a0dd4dc23d94d2aa76c2c2

Observation e6d9dcd2-4121-4365-b480-4979e54b9d75 · outbound

This paper cites Stablerep: Synthetic images from text-to- image models make strong visual representation learners.

Vision-Language Models Do Not Understand Negation Stablerep: Synthetic images from text-to- image models make strong visual representation learners

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.707901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.260007Z digest=sha256:c8f6aa784fec69b114ae9d1538384cd63d49cf9d778b07ab4440e651ec27e01f

Observation 3cd798c9-eaeb-4969-820e-39ad73b97f21 · outbound

This paper cites Learning vision from mod- els rivals learning vision from data.

Vision-Language Models Do Not Understand Negation Learning vision from mod- els rivals learning vision from data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.670835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.265791Z digest=sha256:85a23c18e67bb7e0ca1b006e5d8e05dc17ed4dfb6b4f47ac4c39ff199299aa5b

Observation bacd2688-c153-4551-9cd1-9d3ea0705ac0 · outbound

This paper cites Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning.

Vision-Language Models Do Not Understand Negation Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.653013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.271490Z digest=sha256:b2b72bed3bf974be39b038e543324595478aeae69350ff4eef8a3488276c0153

Observation 505de08f-ebd5-4b72-bb76-e06e2100f58f · outbound

This paper cites Language models are not naysayers: an anal- ysis of language models on negation benchmarks.

Vision-Language Models Do Not Understand Negation Language models are not naysayers: an anal- ysis of language models on negation benchmarks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.635529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.277242Z digest=sha256:444dd6c610455dcca601683e0a080f1d61f882f97f1d50f7b4a3e4fce77c4cae

Observation 57d0e0a7-a39f-419a-8995-ab824dc0e9ca · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Vision-Language Models Do Not Understand Negation Msr-vtt: A large video description dataset for bridging video and language

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.283190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.283190Z digest=sha256:ed637302c912e5a8a50ee8b62487729b13ff2a065a6d1e4f4ce934ca8c86a98e

Observation 134bce19-b0fc-4a7e-8e3a-d53033da6fed · outbound

This paper cites Real-fake: Effective training data synthesis through distribution matching.

Vision-Language Models Do Not Understand Negation Real-fake: Effective training data synthesis through distribution matching

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.607105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.289143Z digest=sha256:d4a6f18695a0cec49734c4def88d9b025098aea51e2113ef6c155a50f42da364

Observation 28592488-7af3-4134-a2ea-ce92e7c9575c · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023.

Vision-Language Models Do Not Understand Negation When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.588431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.294239Z digest=sha256:805f0d6dd2c1b5dcaf8769b41a50e51201aa487cb8f11ea107ee3f62acf0c028

Observation 56c7ca48-e364-4fab-afa8-8f56e3500ae9 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Vision-Language Models Do Not Understand Negation Lit: Zero-shot transfer with locked-image text tuning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.569731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.298798Z digest=sha256:75a577503b7b77b72b93c7e54c5aafded6e13bb6be367daec97d8d9a3ed04bfd

Observation 3b85b920-c49b-45cb-b1c9-3d6d2ee36ec1 · outbound

This paper cites Sigmoid loss for language image pre-training.

Vision-Language Models Do Not Understand Negation Sigmoid loss for language image pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.552358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.304481Z digest=sha256:ad27d9de791f215af1649aab150890889e99ef4c6ce27d3af49fdc390bf2614e

Observation 5bdf1839-48ec-41c2-8e4d-f6c92c722db5 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

Vision-Language Models Do Not Understand Negation BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.310828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.310828Z digest=sha256:651f3fd02461957f166778a734202edc6d3bbc3b193178890107077e7f537744

Observation 31792e35-f68e-4868-919a-52efb1ade9b3 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:46.517398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.321808Z digest=sha256:a18645074ab40a4ebad73c33b97352f5ca98805eb3848a159f6ac5c5d4aef2f0

Observation eebb94a9-a692-433c-88eb-a37f3f06fbda · outbound

This paper cites Yes.” over “No.

Vision-Language Models Do Not Understand Negation Yes.” over “No

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.500166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.326666Z digest=sha256:3a4401be69ca202736f16df2e9046765a2fbc30c7ed0e5c12e72eb2766aa2670

Observation 6ff600fa-2015-4ed7-996a-b3c7c3a277ae · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T20:08:47.498368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.025541Z digest=sha256:fc0fe8006d3f46601295078509db133813eca884b128296d0bf0dcbcda576e2a

Observation 1367e721-a6e2-480a-80d9-672a22dd9629 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:08:46.534065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T20:08:46.316277Z digest=sha256:8650aa1f68a0b9c21dd2667638589e199097fe3be27d7ec96a6065688175b276

Pith citing papers

Observation 6a51e034-03ea-4726-a7a1-89714e6234a0 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:58.344234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:58.344234Z digest=sha256:074022c5e8e213c8698dfb376492279ca31ddb1d4eb1e3a5c105537ea6bdbd14

Observation 93ca8fa6-82d7-43c1-853c-34dbd8976c31 · inbound

NegVQA: Can Vision Language Models Understand Negation? cites this paper.

NegVQA: Can Vision Language Models Understand Negation? Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:09.271049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:01:09.271049Z digest=sha256:4547e8586b5b5db90c8b9a43360bfd06a9935dd8109542d95a0a5d75e7c0fe43

Observation af5a841d-b63e-4378-9185-5f06f40dddbf · inbound

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs cites this paper.

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:21.003431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:21.003431Z digest=sha256:134b716b202ea76afb26afc0c5599c1af96ac808da6ceebffb4a2ac41f47ddab

Observation 65823167-c3b5-467e-9557-9740e1845312 · inbound

Disparities In Negation Understanding Across Languages In Vision-Language Models cites this paper.

Disparities In Negation Understanding Across Languages In Vision-Language Models Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:03.990536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T03:31:50.234201Z digest=sha256:aa560391b2be62d42dabe5d7b606663660e795b01364a51d73f652331946431e

Observation 8f28e7c4-ea6d-489f-935b-6c6cc2a6296d · inbound

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface cites this paper.

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:18:29.524884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:18:14.010081Z digest=sha256:71b7e85b2d7d4389195470a5e2ec729a0cbb4b80ab4b84391afe30d06317e8a5

Observation 2471455d-3f34-45b8-8239-b49ecff50ba2 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:09.727332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:aa7f7845fbf6991892623cf7882c65ae4462ce2aab20270e2460580210bd3e2d

Observation 10e8c895-0807-4b6d-a917-130295408816 · inbound

Uneven Evolution of Cognition Across Generations of Generative AI Models cites this paper.

Uneven Evolution of Cognition Across Generations of Generative AI Models Vision-Language Models Do Not Understand Negation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:58.108557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:00:59.086900Z digest=sha256:b24bac8282beed99b94cfe383ceaaea7f738eba7b1692fdb9923263d5cffbef7

Observation 64c36674-3a95-4912-9f7c-3014d8fc0d3a · inbound

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution cites this paper.

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution Vision-Language Models Do Not Understand Negation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:16:49.314099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:16:49.314099Z digest=sha256:a3eb8e6c4dd7df0466eb1cd166799f294a20d5f4426b2bf1a2c858d89791c19e