Pith. sign in

Paper Citation Record · LEDGER

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

As of 17 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2507.04699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04699 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:45:53.015838Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e541e223-0112-4f2b-bbc9-a9d01ba6a773 · outbound

This paper cites GPT-4 Technical Report.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:51.395705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:51.395705Z digest=sha256:6bac07ff908539b6df550b25e07e3dfaaa5aa7c3cfa988fdef219268362c8db5

Observation 227d747f-b29e-408f-a3b9-4f62ce9cb1cb · outbound

This paper cites Vismin: Visual minimal-change understanding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vismin: Visual minimal-change understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.566352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:51.925349Z digest=sha256:2240b8ffded4c0bfd1b542c771b41576d641998e14b127b7557ef6f8c3b16970

Observation 3880f207-f9cf-4cb5-b486-27be3270d16f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.329853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.329853Z digest=sha256:5a8046e1dd5ece081186c83c02efd6f6b4eb1033c37ac01a14842834a3aa2a44

Observation 1c97579e-5679-49a4-bffb-f94465bededd · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.383239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.383239Z digest=sha256:210143a82c7765e1ae4f37c6c7b79a04b34d5082d36f40cfa1800c174f80c09d

Observation 9a65d044-fa8b-44fe-9df2-d56f09da7c57 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.555147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.413537Z digest=sha256:856b783c6100ecdfb1da57ec6637cecb54b067b7e21274afd4e98462029d0c50

Observation 92458099-43b3-46b0-9043-596746f5c15e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A simple framework for contrastive learning of visual representations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.543996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.487593Z digest=sha256:0ebe6775fbf685c13bc17eee9f4f2c901e84fbdd9e071f1d81a0548487bb9fab

Observation 3f519064-4f4e-4d59-9604-802108f68bf0 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.563298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.563298Z digest=sha256:58cfdf36b3f0ea9c1d6c358ea4962f4ee52e4aa09f0ea70b0e1bef73a820caac

Observation d1ed6828-d923-4682-b1cd-7e32a88ecf1c · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.624039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.624039Z digest=sha256:5ee72185fdab48ae6bb4bc2ab9edd571b77aab85dc231fa539f590687d78a75f

Observation 255f7f21-4cc0-4c52-b8bb-4f8f972f554a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.695270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.695270Z digest=sha256:0ddeab25276b9ec4a697192873d2bde58bd4d70fbc60e0df796e00529d8821df

Observation 907d0124-c006-4743-8192-02908b83132e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.531705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.724114Z digest=sha256:941baf1405f8977f2585aecc8f1efda0a2cf6a05e150163a290f422dacf790d3

Observation 8fe1138b-f0b8-424d-8daf-e48f4b130349 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.761831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.761831Z digest=sha256:ce03e56b3de5c4a75e51a2a49774402aae9273fdad1d6939d90b385b93406d35

Observation 2870a00f-fa24-45d3-8a1f-15b453fb94fe · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36:76137–76150, 2023.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36:76137–76150, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.520358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.779677Z digest=sha256:a0558c8ba8840c30d841f7dc0676c8aa4afe28930b1ff323ef43068a582c6765

Observation 53581e1d-5568-463a-a2ed-2a178875b130 · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Teaching structured vision & language concepts to vision & language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.509035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.873754Z digest=sha256:ab00cd9097136f2323c0a82c5cc6e6df04de77d24d2a476919822b73aef830cc

Observation b3cbef5f-286e-4ca8-abc8-4dc10f2b40d7 · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.497514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.904104Z digest=sha256:fece7cd1c4be1222351bc5bb0e5beecff55325739503e28a1ea5c322e62345ee

Observation ff1e9ff5-461c-410f-99de-bdf4a9bf2b57 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.484864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.907558Z digest=sha256:efa4401a0a88d47edb1f150d306a0a2c08e43844b8b3cb4a099ab9d86e482a9f

Observation bc526a0e-7a79-435c-b2c1-9d0b1222b699 · outbound

This paper cites Advances in deep concealed scene understanding.Visual Intelligence, 1(1):16,.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Advances in deep concealed scene understanding.Visual Intelligence, 1(1):16,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.472951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.910679Z digest=sha256:05064584112754fccbb06f0254abe0996417c65fb6175ced5223f26824370ffe

Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.913989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.913989Z digest=sha256:9525388fdeac1ed83c61bb415a5ab6a4d50efae3a365431699ba72afd1d468dc

Observation ff800481-a222-45c8-80db-7d4383db1b2e · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual program- ming: Compositional visual reasoning without training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.917752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.917752Z digest=sha256:729020fac40c4f67d86fac264557820a564a2f851af0cf628b4c511159fc9d72

Observation 034cbaee-71a7-48be-bae9-c85b71a30a3a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.453433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.921818Z digest=sha256:29af435b7da5d44e9f4213b334f9654a2ddcc182dac8ea6435f5a89a463c898f

Observation 37d2f436-fce2-4aec-9234-3c5a9aa599ab · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.925907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.925907Z digest=sha256:d54e7b0b77d8a6904b8a549d4a6653ff0c25643057f98cd44fa2a929e3e7e419

Observation 784193e0-7fa7-4c41-970e-390885a53cf6 · outbound

This paper cites Semantic to Structure: Learning Structural Representations for Infringement Detection.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Semantic to Structure: Learning Structural Representations for Infringement Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.929510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.929510Z digest=sha256:543a3e068ba99f1ba38b5c72d06c5b11f6288496d2f7f1324fcbb724f185093f

Observation 56f713ab-86a7-4915-b5d1-57a18790bbda · outbound

This paper cites Secret lies in color: Enhancing ai-generated images detection with color distribution anal- ysis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Secret lies in color: Enhancing ai-generated images detection with color distribution anal- ysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.441738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.933305Z digest=sha256:72f2862556194650653b19c6993f0cc8814721436f9d85a4b5f0183df3b44d30

Observation d19b3da0-5082-4a78-aa13-25f335ec430d · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.Advances in Neural Information Processing Systems, 35:9287–9301,.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Elevater: A benchmark and toolkit for evaluating language-augmented visual models.Advances in Neural Information Processing Systems, 35:9287–9301,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.936493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.936493Z digest=sha256:61a90a10199572938ce9a4f9e92ab7038992a7cf256a9ecb2fcaabd5d1de6a37

Observation 3398b2d4-3be4-426a-a529-90823f509958 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.939887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.939887Z digest=sha256:340a676f7389e70d347355c72ab2c3db9690555fa588b4cad6ffece8cc3bb618

Observation d0c50298-4162-40d0-a531-ad5605347184 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.415920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.943579Z digest=sha256:d93cdc415578d5f9d42ed63efe8118af01bed57db442c52f6d409cf00b4fb071

Observation 8a3fd584-54a8-4495-8687-75df5a833d64 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vila: On pre-training for vi- sual language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.403872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.947716Z digest=sha256:6440a11f71306e8e723b5326d09f2604d208c263f8672136c9a7205325378a32

Observation e46b6ba9-fbad-431f-a29a-0b7d99b39aa6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.392740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.951149Z digest=sha256:e626c32fffac8ac9099dd71118f63cdad7cbf442d0190d096e45eab5bf245f27

Observation 2280f4f3-b2b0-4be1-838b-7c667339728c · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.956170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.956170Z digest=sha256:a264d82f3dd8150da020fd41007c4f10e0ee3a1059ac6a8b906255d5cf75ba38

Observation 0c1b60e1-3386-462d-a2d7-9aa857a4ddf5 · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.381845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.960323Z digest=sha256:3b74b0f663e2e37a950a38a7cedbf2ac5d70557b922820b82e45a686b92ec799

Observation d93327d3-3d0c-4963-85bc-6477f84f3fa1 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.963677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.963677Z digest=sha256:f6c7209ca052e21e90737bb53d34b90e938ab4604112588e826d4eeb8611f36c

Observation 66df1fac-28a3-490d-ab5e-9321e882a381 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.370750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.967774Z digest=sha256:386c74b75f04afc1d7471844f945c2993858146eacd9f826bb6a40518b746f76

Observation 5f425da8-9735-4cc6-9291-eed2103a7e35 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.971449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.971449Z digest=sha256:38a5b085d7b50afd072d5c8662dcaebf3120633eac4a153507bb5951aae5132c

Observation e4bcd0cc-39f7-4c3a-9b0d-7c18cb48a168 · outbound

This paper cites Flava: A foundational language and vision alignment model.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Flava: A foundational language and vision alignment model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.359516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.975175Z digest=sha256:346f16e00411be9b23601caada10a3d5d3af7f125d2eddf52a84a5fa91becadf

Observation 76d185f6-1d9b-4154-9f2f-b3e45f382573 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Yfcc100m: The new data in multimedia research

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.348027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.978562Z digest=sha256:4815cdc2754fb4a8da07a91d09d6fb6b2994b8a210616710ca5d0ebaba0ea8f1

Observation 82a8c44b-5adc-4a9c-babc-da6136a43e44 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.336292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.982108Z digest=sha256:843538bbb773aae28206db6464075f150a754e31fefb70569e99a032f9afb358

Observation d6a37b8d-b988-4cf3-8ecc-a2c11124e0e8 · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.322934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.985873Z digest=sha256:c1d4bb506b1011c9569d76248b1a3b1910b503e32d45135fbb85c5f4ead52a62

Observation eefbe169-f499-4469-9781-0b24b3d440ee · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.989662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.989662Z digest=sha256:a30f32c4e46707f60a7dce231ba3b4b9a729571d26db05e5d09710432f96375f

Observation 5d2c2618-0cf4-4fb8-a16d-838f722ac5b8 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.Advances in Neural Informa- tion Processing Systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets What you see is what you read? improving text- image alignment evaluation.Advances in Neural Informa- tion Processing Systems, 36, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.311324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.993852Z digest=sha256:13311df992d41f1116e018d4292fd8728ba9fb1389e0175956c9af4f33c59725

Observation 3b220d1b-48cf-46f1-8116-12b54b6a27ea · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2023.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.299125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:52.997084Z digest=sha256:6de638e5176d53d86c47abc3aea804937e2ebc4358c359e8e1fa29918800ee06

Observation ed6d5e29-6cf9-4f0b-bf71-73cdb6b0e7d3 · outbound

This paper cites Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.000363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.000363Z digest=sha256:61c1cfb7dbce92846c54c9879b0faad2f41d76593434fe23c4e2407961138a3c

Observation 9b14da76-f3fe-4ef7-9011-0ab3907d6449 · outbound

This paper cites Investigating compositional chal- lenges in vision-language models for visual grounding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Investigating compositional chal- lenges in vision-language models for visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.287164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:53.003972Z digest=sha256:0da01ee025867b9049693e57255830350ecb64c41d6ebbd3fed8b39f847e7123

Observation 6bec6939-e9bd-4efe-a87d-d4f9fac82942 · outbound

This paper cites Sigmoid loss for language image pre-training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sigmoid loss for language image pre-training

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.274342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:45:53.007487Z digest=sha256:dc1bb3c74ae13def8d1b01d0d32b3a98a2ec44fe134cab80ed0a0900887ced8f

Observation b54f18df-cf1d-496b-8186-365b28cf9c0a · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.011960Z digest=sha256:dcb487ce1cfffb5ebdc5383083336068a3028351e4e515c2716fca116764cc20

Observation 6faa031f-bd95-4ea0-9866-de705e61db24 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.015838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.015838Z digest=sha256:69ed74a6522b4fc7a78b119a6552969a7c52d2a25604964c8c479bf354afb39d

Pith citing papers

No inbound Pith citation observations are available.