Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Models Do Not Understand Negation

As of 11 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 8 inbound Pith citation observations for arXiv:2501.09425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09425 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:08:46.326666Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:58.344234Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T00:18:29.517494Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved11
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 876019cf-4e11-484d-93b7-b82286f09e37 · outbound

This paper cites Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes.

Vision-Language Models Do Not Understand Negation Aug- mented reality meets computer vision: Efficient data gen- eration for urban driving scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.570497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.008226Z digest=sha256:19b468b0cfea4b3e4e6f594c0f6cb0a396665d08e56ee36bc9a9d909b69a839d

Observation 76626a12-b8c4-4540-861e-c24ddca9336b · outbound

This paper cites Effective conditioned and composed im- age retrieval combining clip-based features.

Vision-Language Models Do Not Understand Negation Effective conditioned and composed im- age retrieval combining clip-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.548229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.014749Z digest=sha256:9980a097afa4dd9c7f370fcffa2c59bd87efa40030fed787b23a6377715cc804

Observation 5d08414e-b6b6-492a-ac86-dbac81d7a1fe · outbound

This paper cites FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks.

Vision-Language Models Do Not Understand Negation FitCLIP: Refining large- scale pretrained image-text models for zero-shot video un- derstanding tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.517244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.020123Z digest=sha256:43da83ad24cd0e86aa95c436f3eefe23335fec45be210e9622d6ec3a4067c91d

Observation adc40a9b-4ad3-44b7-96a0-b98ec2b7ff5d · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Vision-Language Models Do Not Understand Negation Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.031905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.031905Z digest=sha256:b3235bf29b730bba120fbb817b3ac68e98081ab08dd1e970debad589902dde15

Observation 19db9dd0-2961-4ed7-ad64-5bae5646be84 · outbound

This paper cites Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach.

Vision-Language Models Do Not Understand Negation Learning semantic segmentation from synthetic data: A geo- metrically guided input-output adaptation approach

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.453953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.037424Z digest=sha256:06f5ee1773c1873c42584640562d5ae1d38dcfaf568761de8275bacb0a557f81

Observation 3e0cc5c6-38ab-4ba6-855b-8a80c4030f3d · outbound

This paper cites The Llama 3 Herd of Models.

Vision-Language Models Do Not Understand Negation The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.042836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.042836Z digest=sha256:15a0a0c85327708e90c672349dea7b71f841f6bcda8fdc8142f4175a518b488f

Observation 97d27b26-b7c2-45d5-ab5c-dbe351a3529a · outbound

This paper cites The pascal visual object classes (voc) challenge.

Vision-Language Models Do Not Understand Negation The pascal visual object classes (voc) challenge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.429941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.048534Z digest=sha256:18ef5ca91e88d4ca24e0d91b3c9cfa624af3d7014b1e984edc5fd2ccd1b90c8f

Observation 11613278-2c26-4b96-9ae0-c64f4c515b2a · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Vision-Language Models Do Not Understand Negation Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.054919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.054919Z digest=sha256:7c77bc4a60401558073fb55135eef20fe783dff6b6c7e5fc16486690af0a66b3

Observation 43de77e5-ed52-44ba-8ed4-68d1bedef7ab · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Vision-Language Models Do Not Understand Negation Datacomp: In search of the next generation of multimodal datasets

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.402544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.061090Z digest=sha256:80f5f84d657f53515bae81b0841d6f9ebb91a951578efb32d1a991e6db1eaec5

Observation e1117c67-0df2-4e42-a899-fb9d4b5c7bcd · outbound

This paper cites This is not a dataset: A large negation benchmark to challenge large language mod- els.

Vision-Language Models Do Not Understand Negation This is not a dataset: A large negation benchmark to challenge large language mod- els

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.382504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.076557Z digest=sha256:72e4737892ea802f6e791dee93c4d4d7acbeb6b52c7fe6c6a1f76ac6111bd3f2

Observation 504dc068-fbb3-4026-b1d4-87f3c4e9d131 · outbound

This paper cites Shortcut learning in deep neural networks.

Vision-Language Models Do Not Understand Negation Shortcut learning in deep neural networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.355708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.081422Z digest=sha256:b72d0f082d631440c1d63e85ac313cb8dc4065b7cecf736a9aa2e481d442318f

Observation 6ad3e1d3-2a2b-4509-be2f-5b1e12a3db5e · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

Vision-Language Models Do Not Understand Negation SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.086165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.086165Z digest=sha256:f2fb4155f2e23131e2dfd83e8dff1206147c3415398924a869157fd21ffaa22b

Observation 863bde2d-1a1e-433d-bfcd-09a3e99ecd6e · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:47.335828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.091030Z digest=sha256:98f8a08dd546c7dfbabca1d188f045dd0422dc7022004deed586b1ca13c1d22e

Observation fd484441-67ad-4796-b2bf-0cc61cb951c7 · outbound

This paper cites Quilt-1m: One million image-text pairs for histopathology.

Vision-Language Models Do Not Understand Negation Quilt-1m: One million image-text pairs for histopathology

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.314683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.096959Z digest=sha256:86d1860e610dee823e1d48bebb58bc60999461d4a1c1ae47a57a5c67191d330e

Observation 73531a14-883c-4bee-842c-7c0fc3a33525 · outbound

This paper cites Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison.

Vision-Language Models Do Not Understand Negation Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.296473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.103106Z digest=sha256:2853c8cd34100b59a1841fb620b462bcb09c65e5e4513d197d5ae022359126a3

Observation a0594ace-e525-4638-b328-70feab20beaa · outbound

This paper cites Generative models as a data source for multiview representa- tion learning.

Vision-Language Models Do Not Understand Negation Generative models as a data source for multiview representa- tion learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.277193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.108782Z digest=sha256:f3120e6977d90b52b79dd867ed46bb8333bd7a63cd629f1a187b24c2b48e2434

Observation 6825bbb5-161f-4b27-8c15-85c4acb02932 · outbound

This paper cites The power of negation in english: Text, context and relevance.

Vision-Language Models Do Not Understand Negation The power of negation in english: Text, context and relevance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.257383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.113421Z digest=sha256:7e596287012146738b492ef932af898350b8f1bd40966c128fae3ddf6191d5d9

Observation b40567b7-2ebc-44d1-a989-c1c016b6515d · outbound

This paper cites Negation in syntax–on the na- ture of functional categories and projections.

Vision-Language Models Do Not Understand Negation Negation in syntax–on the na- ture of functional categories and projections

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.234939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.118134Z digest=sha256:e38728484e1c0bd19eeb0895970088a0a2b6f25f1fefd12d9279ac8db449fb3c

Observation e4b8857b-2a24-492a-8b01-221b1f34db22 · outbound

This paper cites Naturalbench: Evalu- ating vision-language models on natural adversarial samples.

Vision-Language Models Do Not Understand Negation Naturalbench: Evalu- ating vision-language models on natural adversarial samples

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.214576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.122815Z digest=sha256:dc3a81ff66dced3dd9e89a31c5c9bdabc29cf97e2044681d279e5974992b6155

Observation 51eb013c-e984-4379-8ec6-8ce2bcc29324 · outbound

This paper cites Compre- hending and ordering semantics for image captioning.

Vision-Language Models Do Not Understand Negation Compre- hending and ordering semantics for image captioning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.194283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.128486Z digest=sha256:34fda7512b2f68bfe309cd815b4914aa550879c883f0d1a8016d5491f3f96730

Observation e2d73bba-3d18-4099-bf07-0007c3bfce36 · outbound

This paper cites Cross-modal retrieval and semantic re- finement for remote sensing image captioning.

Vision-Language Models Do Not Understand Negation Cross-modal retrieval and semantic re- finement for remote sensing image captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.174919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.133410Z digest=sha256:35e1ff232d1dc33a8e84603fb0c7aa2ca861b74453db038adfe5b8231fefdbaa

Observation 9c371f82-ffe9-4c86-b92b-0595593f0e24 · outbound

This paper cites Microsoft coco: Common objects in context.

Vision-Language Models Do Not Understand Negation Microsoft coco: Common objects in context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.156422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.138330Z digest=sha256:befd56ffba895fd8df3e4a8cb3a02e7a27a4a767d5c4980a8d0abb4a5fb5e773

Observation c376ff10-6159-490e-b8f8-24a60b6a017e · outbound

This paper cites A visual- language foundation model for computational pathology.

Vision-Language Models Do Not Understand Negation A visual- language foundation model for computational pathology

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.134814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.143444Z digest=sha256:3f8c2a7221730bb0b133831c51d02b8acc01fe79ae5723c75ea4fe86d2920077

Observation 92c26dc6-4b83-4553-a6ca-f3662ab335d6 · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

Vision-Language Models Do Not Understand Negation CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.149164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.149164Z digest=sha256:a829d67048e5533437c179be1e1239f0bb59db1f4fe4a67f6cc14ce35a0c0b56

Observation 46d0c712-e0cf-4ceb-9a8a-be0410d8928e · outbound

This paper cites Fine-tuning llama for multi-stage text retrieval.

Vision-Language Models Do Not Understand Negation Fine-tuning llama for multi-stage text retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.105551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.154599Z digest=sha256:bea83d2a4f724e0d106af498588135342e11a19bad9984c837235c8833864523

Observation e298a188-b5e0-4d97-b981-26f8a39358a5 · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023.

Vision-Language Models Do Not Understand Negation Crepe: Can vision-language foundation models reason compositionally? In CVPR, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.077283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.159371Z digest=sha256:f80bdd84e61f467095bf54444c8db9c3cb7005c6d79e8ab56af01d437c42043c

Observation 367d25f1-9eed-41c0-bc92-6b3d86ddcd67 · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

Vision-Language Models Do Not Understand Negation Simple open-vocabulary object detection with vi- sion transformers

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.046531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.164498Z digest=sha256:1f18a136de2be62654813dee713f804abcbc4aafdb314e902efccde8152e9fa7

Observation 7bb530a6-5b4e-447e-983a-dc3ff370dd56 · outbound

This paper cites Recent advances in processing negation.

Vision-Language Models Do Not Understand Negation Recent advances in processing negation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:47.021702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.169611Z digest=sha256:6813fd94ff249e0b6798a3e7ee08280f06cb4498ae2f830e473c9c7639e66249

Observation 98baac9d-107f-4d94-9a8e-fc3102221a5b · outbound

This paper cites Effect of negation in sentences on sentiment analy- sis and polarity detection.

Vision-Language Models Do Not Understand Negation Effect of negation in sentences on sentiment analy- sis and polarity detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.990672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.174089Z digest=sha256:d2ec143520d30e0e4d027ac0e5030e0acb5af859a664e2c1155e5e7572712664

Observation 1e9c7f92-9608-4b9a-bbb0-8f78ca2618b7 · outbound

This paper cites Clip-it! language-guided video summarization.

Vision-Language Models Do Not Understand Negation Clip-it! language-guided video summarization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.967778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.178920Z digest=sha256:79936eeb915d275a9914323adeb8fa4708b0eb18963db22e2b6e752b78acbac8

Observation 8835517b-641d-4731-9c1e-a04012018817 · outbound

This paper cites Multi-Stage Document Ranking with BERT.

Vision-Language Models Do Not Understand Negation Multi-Stage Document Ranking with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.185183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.185183Z digest=sha256:62c736ec9c1fe83ad63ceb9d1b1fafa5a080fd6337cadabb78dd33fc93c2a4ce

Observation 5869f8b3-04a3-4e44-a339-5be9332767af · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

Vision-Language Models Do Not Understand Negation Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.947397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.190882Z digest=sha256:d9fbafb29c8dd394566b10854b37d8a15c502d213011eceb3139336dedad0829

Observation 3e17222d-e85e-4157-8b69-7f9b531211b2 · outbound

This paper cites On guiding vi- sual attention with language specification.

Vision-Language Models Do Not Understand Negation On guiding vi- sual attention with language specification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.926944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.198307Z digest=sha256:18df55ec20264d462cd0824e519c06ed4797d03a06e21efd906fc848a30ce48b

Observation 877627b6-c6ca-4b13-89d2-df8a6aa91d57 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Vision-Language Models Do Not Understand Negation Learn- ing transferable visual models from natural language super- vision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.909588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.203619Z digest=sha256:6ff19940b39c6e23a5b85f3fbf1091070a254bd5b946db559b9d5fb00fab55d0

Observation 1a058905-4956-4332-b0bb-e67c043cadab · outbound

This paper cites Denseclip: Language-guided dense prediction with context- aware prompting.

Vision-Language Models Do Not Understand Negation Denseclip: Language-guided dense prediction with context- aware prompting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.891949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.208659Z digest=sha256:4e1c52867bf0db616a3bc0661720f8af7471e2905e84fb249899e027533c0ea8

Observation b2cf1896-10f9-4896-a3a8-551b340ce5b3 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

Vision-Language Models Do Not Understand Negation Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.875485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.214930Z digest=sha256:0278fad365d8300816ae0ea79b97956d27524bd0dd4bd63ef2c656f935b81a5d

Observation 986b4a43-0592-49e9-b54d-5fdb299a5c7d · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Vision-Language Models Do Not Understand Negation High-resolution image syn- thesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.855890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.221528Z digest=sha256:4ad21e421eae977b0eb781ffb6fc538b9bba449cb9a7339a6bf1fb11413619bc

Observation 850b4137-6f56-4bb0-8303-dfb1b0199aa7 · outbound

This paper cites Clip for all things zero-shot sketch-based image retrieval, fine- grained or not.

Vision-Language Models Do Not Understand Negation Clip for all things zero-shot sketch-based image retrieval, fine- grained or not

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.837669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.226461Z digest=sha256:55fd58d9bae2f49781eb10c2e82a0d915be1a14c131e222caddde0ae97f7a558

Observation 1a841611-2507-45f8-b381-b5259404120b · outbound

This paper cites LAION-5b: An open large-scale dataset for train- ing next generation image-text models.

Vision-Language Models Do Not Understand Negation LAION-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.819673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.231797Z digest=sha256:9d2563aa5f0642d61707b2bd94bf0c60266417f4dbb099513e8759ff1602f067

Observation b1117bc4-936e-443f-b7f3-fcd2367020e7 · outbound

This paper cites How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions.

Vision-Language Models Do Not Understand Negation How much can clip benefit vision-and-language tasks? In International Conference on Learning Representa- tions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.779307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.238053Z digest=sha256:22b47d227dd7a69da8bfe5e2bd0dc66b18e5de4548fde361cda9ac09561dc658

Observation 89029596-1604-4d2e-8baa-02b6076886a1 · outbound

This paper cites Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues.

Vision-Language Models Do Not Understand Negation Proposalclip: Unsupervised open-category object pro- posal generation via exploiting clip cues

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.759276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.244630Z digest=sha256:1d7ad7a5fe67a3bddcaad162dec371182d9da62425eba101b1ec182b066baf72

Observation 15f7bafa-b741-49ac-964c-95d392e8eba8 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

Vision-Language Models Do Not Understand Negation Cliport: What and where pathways for robotic manipulation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.739522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.250104Z digest=sha256:c09ce6a6853032a55a166b32250c72e2b80932197dc5de180a611598aeab651e

Observation 6b9ddda6-4d36-47c2-8059-4939e8fb9485 · outbound

This paper cites Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations.

Vision-Language Models Do Not Understand Negation Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.254987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.254987Z digest=sha256:e83aeebfffd57ec2e9fe6e57e6032d8f155d3243ca97df05d52b664ca60d558a

Observation e6d9dcd2-4121-4365-b480-4979e54b9d75 · outbound

This paper cites Stablerep: Synthetic images from text-to- image models make strong visual representation learners.

Vision-Language Models Do Not Understand Negation Stablerep: Synthetic images from text-to- image models make strong visual representation learners

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.707901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.260007Z digest=sha256:37d4253782abc6e78b732af1432df54843c80a557c09dae83ce29386382ed3d3

Observation 3cd798c9-eaeb-4969-820e-39ad73b97f21 · outbound

This paper cites Learning vision from mod- els rivals learning vision from data.

Vision-Language Models Do Not Understand Negation Learning vision from mod- els rivals learning vision from data

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.670835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.265791Z digest=sha256:5ac8d9dbb877a6fb7b82fe30a49106f163fbd3d81f54852982084614f5395c45

Observation bacd2688-c153-4551-9cd1-9d3ea0705ac0 · outbound

This paper cites Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning.

Vision-Language Models Do Not Understand Negation Expert-level detection of pathologies from unannotated chest x-ray images via self- supervised learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.653013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.271490Z digest=sha256:a658cdd0794da24b425af24c0edcf4045799f899a33e6ff4f380879a3f52eaa4

Observation 505de08f-ebd5-4b72-bb76-e06e2100f58f · outbound

This paper cites Language models are not naysayers: an anal- ysis of language models on negation benchmarks.

Vision-Language Models Do Not Understand Negation Language models are not naysayers: an anal- ysis of language models on negation benchmarks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.635529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.277242Z digest=sha256:f38fd640736a46a4ed0a475ccaf286605b0cd80d20f0ebee37d654306fe808c6

Observation 57d0e0a7-a39f-419a-8995-ab824dc0e9ca · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Vision-Language Models Do Not Understand Negation Msr-vtt: A large video description dataset for bridging video and language

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.283190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.283190Z digest=sha256:72db96c41222efa43278d7a903fcf268e0b483565b64ad018a8e5f6c450bf29b

Observation 134bce19-b0fc-4a7e-8e3a-d53033da6fed · outbound

This paper cites Real-fake: Effective training data synthesis through distribution matching.

Vision-Language Models Do Not Understand Negation Real-fake: Effective training data synthesis through distribution matching

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.607105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.289143Z digest=sha256:d346ff8acec62e5ad1240554a98177f2a9381a43fd1dfbef05697071326b9e0b

Observation 28592488-7af3-4134-a2ea-ce92e7c9575c · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023.

Vision-Language Models Do Not Understand Negation When and why vision- language models behave like bags-of-words, and what to do about it? In ICLR, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.588431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.294239Z digest=sha256:13255b08e070368fbfc6fca3da1789b4e4835811cc41d5869b0d9bfb7651202a

Observation 56c7ca48-e364-4fab-afa8-8f56e3500ae9 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Vision-Language Models Do Not Understand Negation Lit: Zero-shot transfer with locked-image text tuning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.569731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.298798Z digest=sha256:55659b9e9e7bf81516895c8b0993145b2f93f1575ff2dd13e99a62db72b027b1

Observation 3b85b920-c49b-45cb-b1c9-3d6d2ee36ec1 · outbound

This paper cites Sigmoid loss for language image pre-training.

Vision-Language Models Do Not Understand Negation Sigmoid loss for language image pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.552358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.304481Z digest=sha256:605f5f8d9568fe28a33370f37249389000767cd32a52004b7f3e9dbc5719bfd1

Observation 5bdf1839-48ec-41c2-8e4d-f6c92c722db5 · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

Vision-Language Models Do Not Understand Negation BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:08:46.310828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:08:46.310828Z digest=sha256:f37631ca1a198e0bf097e50f375d6bc0d76fb759f16fb5511176a60cd2e58444

Observation 31792e35-f68e-4868-919a-52efb1ade9b3 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:08:46.517398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.321808Z digest=sha256:6010a72432dbe7a41e405f3f40fc93fe450cc8ef304df59ba8f9c64727050020

Observation eebb94a9-a692-433c-88eb-a37f3f06fbda · outbound

This paper cites Yes.” over “No.

Vision-Language Models Do Not Understand Negation Yes.” over “No

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:08:46.500166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.326666Z digest=sha256:649f2bc362b3120f449ea55d1b113a67f4029b39ee89cec4f268627313c9af8f

Observation 6ff600fa-2015-4ed7-996a-b3c7c3a277ae · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-10T20:08:47.498368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.025541Z digest=sha256:161decbc47589250ed9695eb0591ed11e028c4d192cde1b8a22c9759e74e2aea

Observation 1367e721-a6e2-480a-80d9-672a22dd9629 · outbound

This paper cites an unresolved cited work.

Vision-Language Models Do Not Understand Negation Unresolved cited work

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T20:08:46.534065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:08:46.316277Z digest=sha256:51511be0c3fa0fd269bab1f893a64c64325429cab6397bdf8191e26e8ca2674d

Pith citing papers

Observation 6a51e034-03ea-4726-a7a1-89714e6234a0 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:58.344234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:58.344234Z digest=sha256:8149bc0a878aa9b7354cc887c18abf6a7a7771849aa91622aecc09ef66e1447f

Observation 93ca8fa6-82d7-43c1-853c-34dbd8976c31 · inbound

NegVQA: Can Vision Language Models Understand Negation? cites this paper.

NegVQA: Can Vision Language Models Understand Negation? Vision-Language Models Do Not Understand Negation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:09.271049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:01:09.271049Z digest=sha256:729db4c0d5ef6ab8c0e9ec1a62e9002bac7177f22844a80ec0f444991915fded

Observation af5a841d-b63e-4378-9185-5f06f40dddbf · inbound

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs cites this paper.

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:21.003431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:21.003431Z digest=sha256:6836d12e21159deb067558925dfdc52918f8a14c4bc7efdb83638cd937ac134b

Observation 65823167-c3b5-467e-9557-9740e1845312 · inbound

Disparities In Negation Understanding Across Languages In Vision-Language Models cites this paper.

Disparities In Negation Understanding Across Languages In Vision-Language Models Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:03.990536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T03:31:50.234201Z digest=sha256:a0969c02c503a31f9f9ec75b4dffc2e558bd57e60837879b95c29927e3a12e49

Observation 8f28e7c4-ea6d-489f-935b-6c6cc2a6296d · inbound

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface cites this paper.

Trace Mutation in Human-LLM Dialogue: The Transcript as Forensic and Mitigation Surface Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:18:29.524884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T00:18:14.010081Z digest=sha256:ddb961e2994b133f602e52277f0d2ce345c2601e7b05a71f7e751af672ce6076

Observation 2471455d-3f34-45b8-8239-b49ecff50ba2 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs Vision-Language Models Do Not Understand Negation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:09.727332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:9543e82c1d70b0c678e85cca50defbd3ff93372f986786a4a395c282d9ca4a45

Observation 10e8c895-0807-4b6d-a917-130295408816 · inbound

Uneven Evolution of Cognition Across Generations of Generative AI Models cites this paper.

Uneven Evolution of Cognition Across Generations of Generative AI Models Vision-Language Models Do Not Understand Negation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:58.108557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-11T01:00:59.086900Z digest=sha256:d3887d08ff7d00741e994d57e6c1b634de9231fedb790636e4ed588b2c417ea9

Observation 64c36674-3a95-4912-9f7c-3014d8fc0d3a · inbound

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution cites this paper.

Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution Vision-Language Models Do Not Understand Negation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:16:49.314099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:16:49.314099Z digest=sha256:1ca284e0f3619f504366af6b25588a293fb6628d0d23d1e7dccb3f45f2e5b77f