Pith. sign in

Paper Citation Record · LEDGER

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2506.08227.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08227 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:35.161833Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:51:55.155521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:07.157025Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1315127a-eb00-4076-a473-6c4635f35f22 · outbound

This paper cites Blindfold Baselines for Embodied QA.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Blindfold Baselines for Embodied QA

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.582772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.631870Z digest=sha256:35f675b033f5317b78b6dd8f3d8f8fd30a9994872f0d1b5177a3d084eddca0e9

Observation 514ba2d7-166e-42c1-aa13-35a90f453105 · outbound

This paper cites VisMin: Visual Minimal-Change Understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VisMin: Visual Minimal-Change Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.697115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.697115Z digest=sha256:26d6b715611feeb42ecc3ce968b4cd8726bcda0491bd6cf34c8b0f3a8e07a4d4

Observation ff43a8e0-0a48-4bd9-9b0b-7175df3f170c · outbound

This paper cites CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.545077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.839756Z digest=sha256:df075eacdc5ddaf8fa2a0b9a395b64a13eb5dd9077ce781b097fbb95341ea2eb

Observation e456ef2e-6b3c-4028-9530-ef8c92b00ed7 · outbound

This paper cites Evil- probe-a composite benchmark for extensive visio-linguistic probing.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Evil- probe-a composite benchmark for extensive visio-linguistic probing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.945624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.925092Z digest=sha256:cf89ae4b256d76e75449b63a56a948f8e8df028eb5e6560ea1c2a33b514eb1e1

Observation 6f3b9892-4235-476c-92a3-98bac5295db2 · outbound

This paper cites ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.962463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.962463Z digest=sha256:671a70dd59723f448b848c6894f9c958265de31ab5bbb4386f7f369a688ed08a

Observation 8c9aacce-3096-449c-9206-cc24b38c8c1f · outbound

This paper cites CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.967677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.967677Z digest=sha256:a88d336088a13d944972fefb4fa717107fc7d9fdf4e17d72126329d711ac1ff1

Observation 4a5409ad-f390-42cf-80ed-30fddc6f4908 · outbound

This paper cites The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.495986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.972843Z digest=sha256:7c2ecf2f5fd27d65454ed6186d95b51ee3259b391cd0be39ab9677a1309bd6db

Observation e06e66ca-e87f-4833-a575-c135a8440cab · outbound

This paper cites Routledge, 2016.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Routledge, 2016

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.932358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.977964Z digest=sha256:5c2c8543417c55d30235f96fbce5098cce32da01c2d017ce48fdb1c44a6e8ffc

Observation 38918694-0f76-466d-950b-0770cb460411 · outbound

This paper cites Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations.Advances in Neural Information Processing Systems, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations.Advances in Neural Information Processing Systems, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.918635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.982501Z digest=sha256:0c150d752da11affe607426de27c32f6a2d767ba8ec974ec6564d14a6f639d6c

Observation 99754130-9e4a-4f79-ab05-62c52a5a0c9e · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Sys- tems, 36:27092–27112, 2023.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Dat- acomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Sys- tems, 36:27092–27112, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.904909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.986765Z digest=sha256:f5640657613819935d2bb44d9d91bf34120e69079d4dbb4438dd4ec8b28ab38d

Observation 7a66031c-fb9e-4c71-a5f1-ff2f5e9d7c3c · outbound

This paper cites Shortcut learning in deep neural networks.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Shortcut learning in deep neural networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.891192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.990808Z digest=sha256:3ab1687c8e68982393712adec76f66ba786dc1236999f69d903bc9058420ea89

Observation d2731f8b-0024-4f47-890d-68ad2b604ddc · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.995102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.995102Z digest=sha256:12177185005f48bca33c3990cde35f8ba2d50ee5a7481cebbe6173b3e50b3a2f

Observation 1136e946-f1a2-43d0-af6f-40eaf60bf8c0 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.869530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:34.999500Z digest=sha256:096f67a9b307665cfed3e95e160901180c39b8798cc380d1259680e2fa31ef84

Observation 177fad96-974e-474e-9b39-9f75b9d9a461 · outbound

This paper cites Probing Image-Language Transformers for Verb Understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Probing Image-Language Transformers for Verb Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.003872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.003872Z digest=sha256:f445276a30fd0d2234cb9dd0198adb8b7a8a355c5f1ea3cd40636f75b7bfd4d5

Observation 4db1d596-d0b0-4a0f-b451-e56cfff7a7c1 · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.NeurIPS,.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.NeurIPS,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.855961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.008007Z digest=sha256:4d040b83e23026626ed2ebc3c48d681516960c19969e3ef8dfdea3b5508aa6b8

Observation 5eb6ce73-4dc4-43e1-978c-e826b9d3dfdd · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Compositional Attention Networks for Machine Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.012585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.012585Z digest=sha256:a9427ea9df9972fe5ce36f0ad339c3b5dcf60598b7b66cbd2a995b1e9e9e40ed

Observation 005709c4-5903-410d-8479-eb5772704c0e · outbound

This paper cites Text encoders bottleneck compositionality in contrastive vision- language models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Text encoders bottleneck compositionality in contrastive vision- language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.842984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.016845Z digest=sha256:1fafb05a052f0e6b30b7fbb25082edf2fbbb1ce23bff149c5d3d560d03eafe11

Observation 14cc2461-dbfe-4320-a637-c8ec647f4f39 · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.021243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.021243Z digest=sha256:7ab959394b98f2ae9725f60edfe63bfe565d50d61a0dfd87d9ef238eea22d115

Observation ebc2b2ef-b1fb-4778-aa38-06d7686a4dce · outbound

This paper cites The hard positive truth about vision-language compositionality.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks The hard positive truth about vision-language compositionality

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.830314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.026678Z digest=sha256:a35caa60a0b1a458c28670f59635a397f18cb3587c58003da1ceb8712a8a61da

Observation 9c5d8276-f12b-47e7-8b59-3f0301333ac8 · outbound

This paper cites Clip behaves like a bag-of-words model cross-modally but not uni-modally.arXiv preprint arXiv:2502.03566, 2025.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Clip behaves like a bag-of-words model cross-modally but not uni-modally.arXiv preprint arXiv:2502.03566, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.030959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.030959Z digest=sha256:a8bf116876dd7949dbcf10e49a6de255a14bf559a752cd56720eb9f8c71b3bd3

Observation 89abbafb-e363-47cb-9007-e6bbc3e3731a · outbound

This paper cites Building machines that learn and think like people.Behavioral and brain sciences, 40:e253,.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Building machines that learn and think like people.Behavioral and brain sciences, 40:e253,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.035604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.035604Z digest=sha256:c20a262ef1fcb516dd19e8aeb6b0eee5b1dd98d03d2ac2e46ed3ea0ad7e30953

Observation 458b3f24-8cbf-4a93-8b7d-60019b2fd7a3 · outbound

This paper cites Coco- counterfactuals: Automatically constructed counterfactual examples for image-text pairs.Advances in Neural Infor- mation Processing Systems, 2023.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Coco- counterfactuals: Automatically constructed counterfactual examples for image-text pairs.Advances in Neural Infor- mation Processing Systems, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.808751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.039797Z digest=sha256:760cdf04a439b2db50e8fa2dba0761474bb2d0aca027c2a7d80d47ec45adaa21

Observation 93a25d74-a6cf-4136-8f00-08f039e07271 · outbound

This paper cites Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.043702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.043702Z digest=sha256:90a56b1f0fdecfede07c7efc564f3c7d32423134d16502fc46572e1069c919cf

Observation 72da09fc-fc4c-44b3-87e4-9ccfdc15f9fd · outbound

This paper cites Remov- ing distributional discrepancies in captions improves image- text alignment.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Remov- ing distributional discrepancies in captions improves image- text alignment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.795629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.048382Z digest=sha256:8901814a9a0aa153e0f027a882c113634c748e0cc4e3391ef7ed1be95d72ffa6

Observation dee2741e-8289-4c75-98ec-5d5620aa12a9 · outbound

This paper cites Microsoft coco: Common objects in context.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Microsoft coco: Common objects in context

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.781736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.053390Z digest=sha256:10134c8170d2feef264317a2658754e1e6122e3186a0574c663b0897fdc36e2b

Observation ae714ec9-4090-46e2-923a-b3a46c576065 · outbound

This paper cites Vera: A general- purpose plausibility estimation model for commonsense statements.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Vera: A general- purpose plausibility estimation model for commonsense statements

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.769409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.057677Z digest=sha256:7f2b598028d4a83671fea0a5ea38ff6af15900d770046f73f594ad04e5a88091

Observation aa6a9c88-529b-4294-a2d2-21b848ec309f · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally? InCVPR, 2023.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Crepe: Can vision-language foundation models reason compositionally? InCVPR, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.755690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.061946Z digest=sha256:476634a52613b1935d84d68953732d1f965134c5939254cc512132e13196077a

Observation 81902300-1f71-49e1-b56c-c86878d31958 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Compositional chain-of-thought prompting for large multimodal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.742994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.066367Z digest=sha256:53663c17e8d3ce8a3bf9fc28dabba5e4c9d7bc89f2cd70d70c7483a028d3a06f

Observation 6ad48054-2851-4434-8254-b444e1fc6a33 · outbound

This paper cites TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.071014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.071014Z digest=sha256:bc6367e6e7a5f11e078e3985dcdb232d4a3a3964021fa9a19eb6576d4d0023f6

Observation 75edd83d-e0b1-49ea-a639-88d5cbf2042a · outbound

This paper cites Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.075784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.075784Z digest=sha256:183f8431da442c9b5c221422af2151f1aa6f72fb09940e51461c2f8f26d920db

Observation c2d9873b-fb69-4996-a5ea-eeeff6c3f773 · outbound

This paper cites VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.079871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.079871Z digest=sha256:3b5b7eb6222aa6be5f17efcd8c09621062fda86c448d9211344d8fede19ec73e

Observation f78dec15-7bb9-42b2-a5f5-a716e191633a · outbound

This paper cites Triplet- clip: Improving compositional reasoning of clip via synthetic vision-language negatives.Advances in Neural Information Processing Systems, 37:32731–32760, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Triplet- clip: Improving compositional reasoning of clip via synthetic vision-language negatives.Advances in Neural Information Processing Systems, 37:32731–32760, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.730065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.084098Z digest=sha256:70a158e52ebf943f2a4750f8a294a2381064d0d63f4ba059e219298dc8bea4c1

Observation e94f47ec-3ca1-44e7-bfb7-f685dbd980f8 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Learn- ing transferable visual models from natural language super- vision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.088354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.088354Z digest=sha256:1c3fca9abae553fbb03a7a7530966c5d00f88b99fcd5f58d46eda0998b0941fa

Observation 50f17bcd-ba9a-48db-97ff-4fa4316e24c9 · outbound

This paper cites cola: A bench- mark for compositional text-to-image retrieval.NeurIPS,.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks cola: A bench- mark for compositional text-to-image retrieval.NeurIPS,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.707422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.092289Z digest=sha256:9cf55fb5427a7444aa8dd0065f671076dcde980b1ba21f9425416347f74f328a

Observation 6622fd16-4bca-4cb5-adce-d1f7c9e7cceb · outbound

This paper cites ColorFoil: Investigating Color Blindness in Large Vision and Language Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks ColorFoil: Investigating Color Blindness in Large Vision and Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.096863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.096863Z digest=sha256:70d7a857727fa441b6e9186c82f5f246d5c46ca6b27baeb897359fa7d057f066

Observation 01b8eed0-e6a3-46b4-9a55-64940aa8889f · outbound

This paper cites Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.100979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.100979Z digest=sha256:302a38d230333b730348c1054df3612b4744ac0d7394c3655f8a81cbe690e6c8

Observation 48eca0ff-40c0-49d5-a724-6e4abdaad982 · outbound

This paper cites Teaching composition- ality to cnns.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Teaching composition- ality to cnns

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.694423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.105148Z digest=sha256:facfc9f89ecca400ebba51ba5a1ec0203d0c0ebd7934fb0ce9bfcb958c43e8b1

Observation 679d7e41-8db7-4673-8be4-4eb0ea8d00b2 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.681245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.109185Z digest=sha256:1bcf9e2b69ea30fed38c0b79e08eea7553232ff59bba8813df83a5c0e8311b47

Observation a3698515-14c1-45aa-a30f-59ac2fb2f97a · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.113327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.113327Z digest=sha256:2d948a0a2c799ab8470755c462abb2283905daeb7c476c3a438d2371c88bf6ae

Observation b6d1e98b-3f4c-44db-b8a4-a016b56c09a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.117743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.117743Z digest=sha256:e34f6db8cde18256c21728afeef4d98cbf4a00426a8d65bfb3c1f75c8eab6e6d

Observation 2075b482-a3c3-4074-860f-e7c0d92f2ff9 · outbound

This paper cites Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.668355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.121739Z digest=sha256:9ca9f91c01cfd8629c29e4fa621e345188925aad2aed63c63675eca5e25db74b

Observation dbb94afc-f6c8-4b71-bcd1-584bfa3c92e0 · outbound

This paper cites Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.654603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.126072Z digest=sha256:35b538891132a149bc609210455f6f0bb530dfa0cd0b97ad294131c290537c78

Observation 0ca2d9da-5c56-47dc-badc-5060aaec599a · outbound

This paper cites Equivariant similarity for vision-language foundation models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Equivariant similarity for vision-language foundation models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.640730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.130067Z digest=sha256:abefcf4d4112eb45a16f5edead2a350a2a20c127c718aae71cb50786055051d1

Observation 4cce305e-31aa-4962-8c5e-eaacd241025a · outbound

This paper cites Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.249852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.134083Z digest=sha256:d38d038253043687452c66d9dd9c8ab717632734e6ff9d9f37e9a6ba1efba781

Observation c88e7dc5-7c2d-4a58-8584-5ff64e0d218f · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2022.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.627745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.139226Z digest=sha256:5702a754f5c9b5ae637568192145c642959406e87789092400d74666319165ac

Observation bfcb0d5d-3e2b-4b3d-aa1b-a8702e7af27f · outbound

This paper cites Investigating compositional chal- lenges in vision-language models for visual grounding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Investigating compositional chal- lenges in vision-language models for visual grounding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.613423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.143356Z digest=sha256:5e619f79a7dc10f5a876679df6c74d1dd8f6845865cf0d822248fb6384b5fc13

Observation e2f2da20-3329-4311-a1f6-e899b82cfd34 · outbound

This paper cites CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.148475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.148475Z digest=sha256:71fe0c3341582c1b362966aedba7322b43c28e2bd5f89394f4cd43400cba0ab7

Observation 8c230ce9-de7e-47c6-9477-98231b1afb40 · outbound

This paper cites Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.153214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.153214Z digest=sha256:4be691439e3ca1348f383d1cf34ee9886b6d6dee36a353d0fd37b82aeb34a607

Observation 42821a0d-6d10-4218-a568-1379fe734f45 · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.157557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.157557Z digest=sha256:47e06dc00912313a54c6b69973358a93ce8504bc95fb10204c4aa833bffdd00f

Observation 30cc44a3-8be7-42e3-864c-83382a985605 · outbound

This paper cites Iterated learning improves composition- ality in large vision-language models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Iterated learning improves composition- ality in large vision-language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.598478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:21:35.161833Z digest=sha256:9f4bfd1342bba5e115eb83f7cf60ffd3e985095e05e9b2afb7f99d548bd1934d

Pith citing papers

Observation 773f9160-465e-4d0e-bd18-743ac80e09ce · inbound

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference cites this paper.

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.787520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:38:00.522094Z digest=sha256:e11d092556f1089448c5fb9764a5c948a708e8b1c0ec9ade347f1c7228ce7d73

Observation 6547264c-3fe5-413d-aa13-6fc142f74696 · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.158536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:de01208631bc08f06a8475897e19669e38f0baaceda68ad4f670ea1bae34c481

Observation a79a2f97-e8b3-4ba6-8924-74995bc2423d · inbound

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models cites this paper.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.155521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.155521Z digest=sha256:4da9a951e013ff7b8ef19a724a6bb0b9ed3fe0e0f692d56036bf1267c5cf38a5