Pith. sign in

Paper Citation Record · LEDGER

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

As of 10 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2506.08227.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08227 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:35.161833Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:51:55.155521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:07.157025Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1315127a-eb00-4076-a473-6c4635f35f22 · outbound

This paper cites Blindfold Baselines for Embodied QA.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Blindfold Baselines for Embodied QA

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.582772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.631870Z digest=sha256:34709eb15c071ff7c9652b17b9b71a32540270ac2a2165a8d78585f589aea33d

Observation 514ba2d7-166e-42c1-aa13-35a90f453105 · outbound

This paper cites VisMin: Visual Minimal-Change Understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VisMin: Visual Minimal-Change Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.697115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.697115Z digest=sha256:a0a81ed734e9e9abf7bcd01aa7bc9e4740b5cf81257f0de02225cfce2527f268

Observation ff43a8e0-0a48-4bd9-9b0b-7175df3f170c · outbound

This paper cites CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.545077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.839756Z digest=sha256:f16caaa0de870712d5634472b0b8ed37f632adc2466709e65a6dd449e35f70b2

Observation e456ef2e-6b3c-4028-9530-ef8c92b00ed7 · outbound

This paper cites Evil- probe-a composite benchmark for extensive visio-linguistic probing.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Evil- probe-a composite benchmark for extensive visio-linguistic probing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.945624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.925092Z digest=sha256:0470ec03d9af941378cfb2dd13dd55c75dd7eb9e4e17f6407af4c4b5bc06060c

Observation 6f3b9892-4235-476c-92a3-98bac5295db2 · outbound

This paper cites ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.962463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.962463Z digest=sha256:acdef0bda8502433349e74617dc22d127def030be7d203e34db8d85ef073dc7c

Observation 8c9aacce-3096-449c-9206-cc24b38c8c1f · outbound

This paper cites CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.967677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.967677Z digest=sha256:b28029322d872ad3f68cbd57d7615fe3c4dd2774c031f392daa2108c8c2163da

Observation 4a5409ad-f390-42cf-80ed-30fddc6f4908 · outbound

This paper cites The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.495986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.972843Z digest=sha256:9250f284c84ad097145806b4949d2a84e0e2cfc7a23aa7b54e8b2fde40b8c163

Observation e06e66ca-e87f-4833-a575-c135a8440cab · outbound

This paper cites Routledge, 2016.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Routledge, 2016

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.932358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.977964Z digest=sha256:b9d12b9a667038fd7493b54303661df0bb14a8b28387ae7f1fedceadc0222013

Observation 38918694-0f76-466d-950b-0770cb460411 · outbound

This paper cites Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations.Advances in Neural Information Processing Systems, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Sugarcrepe++ dataset: Vision-language model sensitivity to semantic and lexical alterations.Advances in Neural Information Processing Systems, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.918635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.982501Z digest=sha256:f4dd50a2c1efc9e7b2803cfd95f2a535e76c79d1320fb362e14cbf8be26b602c

Observation 99754130-9e4a-4f79-ab05-62c52a5a0c9e · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Sys- tems, 36:27092–27112, 2023.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Dat- acomp: In search of the next generation of multimodal datasets.Advances in Neural Information Processing Sys- tems, 36:27092–27112, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.904909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.986765Z digest=sha256:6db7e38c816a3261cf6e4dfdc9969c49e24c3e3219faf0c6b69e37d0b644a3f6

Observation 7a66031c-fb9e-4c71-a5f1-ff2f5e9d7c3c · outbound

This paper cites Shortcut learning in deep neural networks.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Shortcut learning in deep neural networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.891192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.990808Z digest=sha256:45cfa0e305fd9ba09b4863468ea7422efac04b684cb0478edc781fc377756718

Observation d2731f8b-0024-4f47-890d-68ad2b604ddc · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:34.995102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:34.995102Z digest=sha256:0e7ed42d819c5c1ecbeaff0eada045ac6ff7229e23145e2555af2a59d50e0a2d

Observation 1136e946-f1a2-43d0-af6f-40eaf60bf8c0 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.869530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:34.999500Z digest=sha256:9fe66d465f7aa96614fe96cec5e217cbd0dbbf0e7c0496941f9ffb860a8b77af

Observation 177fad96-974e-474e-9b39-9f75b9d9a461 · outbound

This paper cites Probing Image-Language Transformers for Verb Understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Probing Image-Language Transformers for Verb Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.003872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.003872Z digest=sha256:3ee7006b7f5fc871c2518e347c8bf1c22f4b0c4c82a541a79173a424b0f869be

Observation 4db1d596-d0b0-4a0f-b451-e56cfff7a7c1 · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.NeurIPS,.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.NeurIPS,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.855961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.008007Z digest=sha256:e5358ac74544c668d6bc670b81c4bceab29716e61c041674b8b84b181d1c0c68

Observation 5eb6ce73-4dc4-43e1-978c-e826b9d3dfdd · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Compositional Attention Networks for Machine Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.012585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.012585Z digest=sha256:7d9cb950df0d33018ae202e04b1ac891b3ac645aee4c69de4b08434ea72c8d0d

Observation 005709c4-5903-410d-8479-eb5772704c0e · outbound

This paper cites Text encoders bottleneck compositionality in contrastive vision- language models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Text encoders bottleneck compositionality in contrastive vision- language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.842984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.016845Z digest=sha256:303d9f30a025f00af42d083fd7ce6f61ad0e89b60cc1132dac2ac1d3310effb0

Observation 14cc2461-dbfe-4320-a637-c8ec647f4f39 · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.021243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.021243Z digest=sha256:caf492714e5f7d9e15bbe79ccc254e57000a4427258fb990e70fc3737b40d142

Observation ebc2b2ef-b1fb-4778-aa38-06d7686a4dce · outbound

This paper cites The hard positive truth about vision-language compositionality.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks The hard positive truth about vision-language compositionality

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.830314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.026678Z digest=sha256:772e1717d3fc24dd857e87581b5d719e266ab87eb15b7f6334126533e4b5fe13

Observation 9c5d8276-f12b-47e7-8b59-3f0301333ac8 · outbound

This paper cites Clip behaves like a bag-of-words model cross-modally but not uni-modally.arXiv preprint arXiv:2502.03566, 2025.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Clip behaves like a bag-of-words model cross-modally but not uni-modally.arXiv preprint arXiv:2502.03566, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.030959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.030959Z digest=sha256:e7e9ea43d1d87029a4cbbb2dfb96eb4b45445cd54afb7ad8f1c84a6a397e1805

Observation 89abbafb-e363-47cb-9007-e6bbc3e3731a · outbound

This paper cites Building machines that learn and think like people.Behavioral and brain sciences, 40:e253,.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Building machines that learn and think like people.Behavioral and brain sciences, 40:e253,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.035604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.035604Z digest=sha256:92c6e3392f76e52db7c294337fb6f500a4bc4f966baa51da9e524d0adf3c9408

Observation 458b3f24-8cbf-4a93-8b7d-60019b2fd7a3 · outbound

This paper cites Coco- counterfactuals: Automatically constructed counterfactual examples for image-text pairs.Advances in Neural Infor- mation Processing Systems, 2023.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Coco- counterfactuals: Automatically constructed counterfactual examples for image-text pairs.Advances in Neural Infor- mation Processing Systems, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.808751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.039797Z digest=sha256:c1c21927ce9a2d7e28fed6956f76edc0e3e3ac35856aa0b38173c971387a464f

Observation 93a25d74-a6cf-4136-8f00-08f039e07271 · outbound

This paper cites Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.043702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.043702Z digest=sha256:1b0c83f7b37dcb9b9f97fc4ced6936f183093c42aa449a2df15f0a019049bf4f

Observation 72da09fc-fc4c-44b3-87e4-9ccfdc15f9fd · outbound

This paper cites Remov- ing distributional discrepancies in captions improves image- text alignment.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Remov- ing distributional discrepancies in captions improves image- text alignment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.795629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.048382Z digest=sha256:ebd8564aba42946bc390225eb3d724942ba36f0ab1ebf0783b41332642ae86f0

Observation dee2741e-8289-4c75-98ec-5d5620aa12a9 · outbound

This paper cites Microsoft coco: Common objects in context.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Microsoft coco: Common objects in context

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.781736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.053390Z digest=sha256:bb1921cbe8ef4353ead9f31bdf2e019205930d1046f817e69580a8e31b8e50b6

Observation ae714ec9-4090-46e2-923a-b3a46c576065 · outbound

This paper cites Vera: A general- purpose plausibility estimation model for commonsense statements.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Vera: A general- purpose plausibility estimation model for commonsense statements

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.769409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.057677Z digest=sha256:1bf412838945a0bf6546e427d80df798113381307fb6071a2804937b65276893

Observation aa6a9c88-529b-4294-a2d2-21b848ec309f · outbound

This paper cites Crepe: Can vision-language foundation models reason compositionally? InCVPR, 2023.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Crepe: Can vision-language foundation models reason compositionally? InCVPR, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.755690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.061946Z digest=sha256:d387c9abab7adb2917674fda1b2fe1646e975848e005fff9fa322f0af20ef188

Observation 81902300-1f71-49e1-b56c-c86878d31958 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Compositional chain-of-thought prompting for large multimodal models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.742994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.066367Z digest=sha256:53cdbbd475d51e9c40f6de864f8288bf541bbe7423919b1de700065dd8114fb7

Observation 6ad48054-2851-4434-8254-b444e1fc6a33 · outbound

This paper cites TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.071014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.071014Z digest=sha256:ccfb426d650b9c411316bf5e40d8acbbbc957d4cbde64f20b72b4fa1d0dbf1b4

Observation 75edd83d-e0b1-49ea-a639-88d5cbf2042a · outbound

This paper cites Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.075784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.075784Z digest=sha256:042614fb786d7ba973dc540af79ac1c4fd705b5c0a9c38a5b4fcddab1828ad4d

Observation c2d9873b-fb69-4996-a5ea-eeeff6c3f773 · outbound

This paper cites VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.079871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.079871Z digest=sha256:862a34cdc612e0617a87f99becd008e1c0211d936293a936e29b32120e304618

Observation f78dec15-7bb9-42b2-a5f5-a716e191633a · outbound

This paper cites Triplet- clip: Improving compositional reasoning of clip via synthetic vision-language negatives.Advances in Neural Information Processing Systems, 37:32731–32760, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Triplet- clip: Improving compositional reasoning of clip via synthetic vision-language negatives.Advances in Neural Information Processing Systems, 37:32731–32760, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.730065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.084098Z digest=sha256:75d342af66ae099e9f3dfe311ebad938951937e9a718e3986389bbd82cfe7750

Observation e94f47ec-3ca1-44e7-bfb7-f685dbd980f8 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Learn- ing transferable visual models from natural language super- vision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.088354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.088354Z digest=sha256:1a73442e6e72caa46614ce5087a3ffda3330b1e277b949a2d95e19b14f5974f2

Observation 50f17bcd-ba9a-48db-97ff-4fa4316e24c9 · outbound

This paper cites cola: A bench- mark for compositional text-to-image retrieval.NeurIPS,.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks cola: A bench- mark for compositional text-to-image retrieval.NeurIPS,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.707422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.092289Z digest=sha256:dc907e218b54369ad606c9a2cfd9f40591a58560067bfd6c8b565f8cc96cfc80

Observation 6622fd16-4bca-4cb5-adce-d1f7c9e7cceb · outbound

This paper cites ColorFoil: Investigating Color Blindness in Large Vision and Language Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks ColorFoil: Investigating Color Blindness in Large Vision and Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.096863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.096863Z digest=sha256:37c5214fba2736352470aca3ef3b117942566a49dfa0c65ec3a6cf7fa1b76d98

Observation 01b8eed0-e6a3-46b4-9a55-64940aa8889f · outbound

This paper cites Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.100979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.100979Z digest=sha256:2e3e2a445daf8eb3d9ef9196f0a51bb1be137dd1211ecef1f17bf74ed2ea39fd

Observation 48eca0ff-40c0-49d5-a724-6e4abdaad982 · outbound

This paper cites Teaching composition- ality to cnns.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Teaching composition- ality to cnns

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.694423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.105148Z digest=sha256:068867e4b21d04b2ccb945d4536835940e9e07f9b6df62fc43a597d9edb4f13c

Observation 679d7e41-8db7-4673-8be4-4eb0ea8d00b2 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.681245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.109185Z digest=sha256:8bc52404dbd09ae6bf6a30e208a08aa17fe61961d5cea32871d83dc9d92b0889

Observation a3698515-14c1-45aa-a30f-59ac2fb2f97a · outbound

This paper cites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.113327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.113327Z digest=sha256:c3b4e4d949a4c2ecbd6540996494973e72d9ab58f71bfaeafec74963c8fc5735

Observation b6d1e98b-3f4c-44db-b8a4-a016b56c09a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.117743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.117743Z digest=sha256:ae3d830e37be809d69ecc427affa6644578459f5b6ccf32d8b328396e0c5523b

Observation 2075b482-a3c3-4074-860f-e7c0d92f2ff9 · outbound

This paper cites Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.668355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.121739Z digest=sha256:7c3191dad42c29111cdb091cb02bd74aeb8b84a40f1d6b007c2e60fdad51f22f

Observation dbb94afc-f6c8-4b71-bcd1-584bfa3c92e0 · outbound

This paper cites Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Image captioners are scalable vision learners too.Advances in Neural Infor- mation Processing Systems, 36, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.654603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.126072Z digest=sha256:4dbacc5d805715181b6cc25365c061a557e7d6ad6b976ab920e7f9b866bbd1bf

Observation 0ca2d9da-5c56-47dc-badc-5060aaec599a · outbound

This paper cites Equivariant similarity for vision-language foundation models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Equivariant similarity for vision-language foundation models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.640730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.130067Z digest=sha256:f40922b4e5d7ed28b5def92064e4e87eefff91d7b027eeb6b41e03ba9b1b1908

Observation 4cce305e-31aa-4962-8c5e-eaacd241025a · outbound

This paper cites Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:35.249852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.134083Z digest=sha256:b196cd145a88685a6c8c0b2781efdf9371d75a2f041e29184bac4c2930840c9a

Observation c88e7dc5-7c2d-4a58-8584-5ff64e0d218f · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2022.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.627745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.139226Z digest=sha256:9a8fa386c8e0052328a7ea7bed4767ab22d43facb9a01937c728118093208615

Observation bfcb0d5d-3e2b-4b3d-aa1b-a8702e7af27f · outbound

This paper cites Investigating compositional chal- lenges in vision-language models for visual grounding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Investigating compositional chal- lenges in vision-language models for visual grounding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.613423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.143356Z digest=sha256:0bc6323a7364d68414f45fe78cf9d8023f25fa6b8bdaf54baf6e898986587413

Observation e2f2da20-3329-4311-a1f6-e899b82cfd34 · outbound

This paper cites CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.148475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.148475Z digest=sha256:5ea160bad5bc5e8b74edc168dbb5e3c1cbeae3d57bcd531795029f614580f8e6

Observation 8c230ce9-de7e-47c6-9477-98231b1afb40 · outbound

This paper cites Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.153214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.153214Z digest=sha256:902cc2ed8f648b5df1b9a5d577640ebfa60ef6cd12021ecc5e296f6631b9c789

Observation 42821a0d-6d10-4218-a568-1379fe734f45 · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.157557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.157557Z digest=sha256:cee50ea3f415023618d2aebbdc7574363c4ee477ea37d31cd3b4899bb8726580

Observation 30cc44a3-8be7-42e3-864c-83382a985605 · outbound

This paper cites Iterated learning improves composition- ality in large vision-language models.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Iterated learning improves composition- ality in large vision-language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:35.598478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:21:35.161833Z digest=sha256:497826f0900201569b6fb22c172bdddf7e44903621aae6f5de343d6b96d3084c

Pith citing papers

Observation 773f9160-465e-4d0e-bd18-743ac80e09ce · inbound

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference cites this paper.

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.787520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:38:00.522094Z digest=sha256:1506b8923cf5fc339be1b3b6b926cd8c9ed50ecdf0a3d6585d271f52f7a536b6

Observation 6547264c-3fe5-413d-aa13-6fc142f74696 · inbound

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation cites this paper.

Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.158536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:05:36.836361Z digest=sha256:695f994c51637a47575cf7a9d2a96cde910b2aad2ff7e75b017091b4bc20f42e

Observation a79a2f97-e8b3-4ba6-8924-74995bc2423d · inbound

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models cites this paper.

Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T03:51:55.155521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:51:55.155521Z digest=sha256:00c0d10aa466e0b5ae54c2130a53113ed6b6717cf7564a5d8bdb994ce0543a19