Pith. sign in

Paper Citation Record · LEDGER

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2507.04699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04699 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:45:53.015838Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e541e223-0112-4f2b-bbc9-a9d01ba6a773 · outbound

This paper cites GPT-4 Technical Report.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:51.395705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:51.395705Z digest=sha256:630a251689673462c187ac4dfa003dc6338402c83677ebfc9ba21badfd69d946

Observation 227d747f-b29e-408f-a3b9-4f62ce9cb1cb · outbound

This paper cites Vismin: Visual minimal-change understanding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vismin: Visual minimal-change understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.566352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:51.925349Z digest=sha256:f72e2f6d67f8e88a7ca86c73627671d9c73fb9eed5ed5ee3c9af2c4152bfa48a

Observation 3880f207-f9cf-4cb5-b486-27be3270d16f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.329853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.329853Z digest=sha256:32c97d6490a0c7a398ac7392df70119b755305c45ea2861a3b732519550e5631

Observation 1c97579e-5679-49a4-bffb-f94465bededd · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.383239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.383239Z digest=sha256:702e0cb1340a83bd323c7605d66b30e7c0a7d3b14dacfafdfb4d5febda00c139

Observation 9a65d044-fa8b-44fe-9df2-d56f09da7c57 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.555147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.413537Z digest=sha256:bb74b5d471b35356d6952eb8be08d5855132ef1aec4f03fc33ca027e41313fc4

Observation 92458099-43b3-46b0-9043-596746f5c15e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A simple framework for contrastive learning of visual representations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.543996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.487593Z digest=sha256:04392a93a608e21a21bdf4a7b7edfb9574ca1cb6eddc0d1a61be0612d19eaab4

Observation 3f519064-4f4e-4d59-9604-802108f68bf0 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.563298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.563298Z digest=sha256:ba1485f5fda614e80de2aaf49cc404f7e842d128379da1b8944d3fbfaa3e1120

Observation d1ed6828-d923-4682-b1cd-7e32a88ecf1c · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.624039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.624039Z digest=sha256:2f01d53ecd14a78a7e742f68160b75801ce939d2e3da36121be05e32f2be6878

Observation 255f7f21-4cc0-4c52-b8bb-4f8f972f554a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.695270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.695270Z digest=sha256:0acceb3b211af2601f97b44136aa10a1f10e900a4319f2f651eb138d284f161a

Observation 907d0124-c006-4743-8192-02908b83132e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.531705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.724114Z digest=sha256:1b793ae792249bc56ec24c7eba5ec8bb35e3265a96dc5657ff6f029923faa9d3

Observation 8fe1138b-f0b8-424d-8daf-e48f4b130349 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.761831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.761831Z digest=sha256:90c0293023217d327ce916b41d0c4384203d0c80bfc943738359563b221c0962

Observation 2870a00f-fa24-45d3-8a1f-15b453fb94fe · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36:76137–76150, 2023.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36:76137–76150, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.520358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.779677Z digest=sha256:ea26f4c6b086dd5c12c048e44e96e7b65810b27725f4c1959597ab66a517d3e1

Observation 53581e1d-5568-463a-a2ed-2a178875b130 · outbound

This paper cites Teaching structured vision & language concepts to vision & language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Teaching structured vision & language concepts to vision & language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.509035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.873754Z digest=sha256:c574ffe07e2d59015fb5fa943d0bd8d58c6c9b6dbcc611eac1b51d7ce9f0e83d

Observation b3cbef5f-286e-4ca8-abc8-4dc10f2b40d7 · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Dense and aligned captions (dac) promote compositional reasoning in vl models.Advances in Neural Information Processing Systems, 36, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.497514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.904104Z digest=sha256:c7e3bc783c74b95398f405fd71519745e955790ecc6df65f2d8875c15c9e20b1

Observation ff1e9ff5-461c-410f-99de-bdf4a9bf2b57 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.484864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.907558Z digest=sha256:eeb76553a4ddb8ef68ec2f52dd916147699b5481fc6187aa74fd2f0b85f8e79d

Observation bc526a0e-7a79-435c-b2c1-9d0b1222b699 · outbound

This paper cites Advances in deep concealed scene understanding.Visual Intelligence, 1(1):16,.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Advances in deep concealed scene understanding.Visual Intelligence, 1(1):16,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.472951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.910679Z digest=sha256:fe0f014c4c51f4abe5854ecbdeedd46f5d88674d610770b2468568c7260b02b1

Observation 7883666e-4ec6-4b91-a616-a090dfa88957 · outbound

This paper cites Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.913989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.913989Z digest=sha256:e1d1343fccebd4bdb358ed79094c1c8db71d46df1ab874e675636f6ff227fbcf

Observation ff800481-a222-45c8-80db-7d4383db1b2e · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual program- ming: Compositional visual reasoning without training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.917752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.917752Z digest=sha256:820f4edc972b8f3cb2a8327694d617428efb98b4c31196e5a8ecbbb638092f22

Observation 034cbaee-71a7-48be-bae9-c85b71a30a3a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.453433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.921818Z digest=sha256:9a6f54b843ce978d6f8fcf8a259009f623f7da4c06cc10f25dcee470b5d3d6f9

Observation 37d2f436-fce2-4aec-9234-3c5a9aa599ab · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LoRA: Low-Rank Adaptation of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.925907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.925907Z digest=sha256:c90579043d2b9d1efb5240eff1ad431bfd6ef93ea2199d10147ed4863e322939

Observation 784193e0-7fa7-4c41-970e-390885a53cf6 · outbound

This paper cites Semantic to Structure: Learning Structural Representations for Infringement Detection.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Semantic to Structure: Learning Structural Representations for Infringement Detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.929510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.929510Z digest=sha256:93ad62076e86c11b2f03bfe1501d617957701a6274ee2ef11fd13e3e67f3826e

Observation 56f713ab-86a7-4915-b5d1-57a18790bbda · outbound

This paper cites Secret lies in color: Enhancing ai-generated images detection with color distribution anal- ysis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Secret lies in color: Enhancing ai-generated images detection with color distribution anal- ysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.441738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.933305Z digest=sha256:3fb55fe8565111b301cb65793dafd654c2186447cdd5b6cf395b89afa26fc505

Observation d19b3da0-5082-4a78-aa13-25f335ec430d · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.Advances in Neural Information Processing Systems, 35:9287–9301,.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Elevater: A benchmark and toolkit for evaluating language-augmented visual models.Advances in Neural Information Processing Systems, 35:9287–9301,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.936493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.936493Z digest=sha256:bad69e31f8ffa75618584c7a847f4a559325b1d5d594d25dc3f77e4dbe7a4d8f

Observation 3398b2d4-3be4-426a-a529-90823f509958 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.939887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.939887Z digest=sha256:c38b01b60e6214e95a772e16fdb31ea2a03489991527aa0fa4e99a5d3fb58bb2

Observation d0c50298-4162-40d0-a531-ad5605347184 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.415920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.943579Z digest=sha256:1588141ad3d9150f31601d2f5866a26b15dae99b17e6a5174cfd5d400a5b545e

Observation 8a3fd584-54a8-4495-8687-75df5a833d64 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Vila: On pre-training for vi- sual language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.403872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.947716Z digest=sha256:649ab37a391299b525d42063436ceeb8153531998ca4005d7bfe3e5fd74ff94f

Observation e46b6ba9-fbad-431f-a29a-0b7d99b39aa6 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.392740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.951149Z digest=sha256:0cdcb96edd8f35ac85da9898074290c0663e6d726a90d4b287db0902a2f37dd3

Observation 2280f4f3-b2b0-4be1-838b-7c667339728c · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.956170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.956170Z digest=sha256:25324c4632440fd715ee1ac1e1ade2763409f90910893f9d4f98ada85bce15dd

Observation 0c1b60e1-3386-462d-a2d7-9aa857a4ddf5 · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.381845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.960323Z digest=sha256:57eb02ea8bcc663ed2bab360f427862dd5377a517ee68e629786f7fedf98ae1f

Observation d93327d3-3d0c-4963-85bc-6477f84f3fa1 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.963677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.963677Z digest=sha256:8a245b11d3cc42b9641a4997a38019a0bfc6a2765d7ee087ae1986c321491228

Observation 66df1fac-28a3-490d-ab5e-9321e882a381 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.370750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.967774Z digest=sha256:3cce8de351bdf4b4c05093a8d42f4c1d6dda402efbd97fdb2b33b6f05644865b

Observation 5f425da8-9735-4cc6-9291-eed2103a7e35 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.971449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.971449Z digest=sha256:5ed3da6451802cc4e60a17a056e93234dd0e902561db8ff49a8e0603e25b5f2d

Observation e4bcd0cc-39f7-4c3a-9b0d-7c18cb48a168 · outbound

This paper cites Flava: A foundational language and vision alignment model.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Flava: A foundational language and vision alignment model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.359516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.975175Z digest=sha256:78ce7f15013f8077d29d0d664f8e71eb8428df2ef19c03ae52575081a8ae0380

Observation 76d185f6-1d9b-4154-9f2f-b3e45f382573 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Yfcc100m: The new data in multimedia research

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.348027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.978562Z digest=sha256:93ef75662f522a540f7c5625a21651ce2c1043f2cddc8d13e3770b260d6b1ca8

Observation 82a8c44b-5adc-4a9c-babc-da6136a43e44 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.336292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.982108Z digest=sha256:27be72e1e20daf71fb47b2e27c57446f764437b68799ead054dd81ff9f906dce

Observation d6a37b8d-b988-4cf3-8ecc-a2c11124e0e8 · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.322934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.985873Z digest=sha256:6ecdac5f71483e87a5590b3f6508958d58d267e2d96da11b2beb281655949df0

Observation eefbe169-f499-4469-9781-0b24b3d440ee · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets xgen-mm (blip-3): A family of open large multimodal models.arXiv preprint arXiv:2408.08872, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:52.989662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:52.989662Z digest=sha256:529fbf1b1abcabb73cee9bc72b17d92b9dccdfae00ad2ee74c9fc094a514b604

Observation 5d2c2618-0cf4-4fb8-a16d-838f722ac5b8 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.Advances in Neural Informa- tion Processing Systems, 36, 2024.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets What you see is what you read? improving text- image alignment evaluation.Advances in Neural Informa- tion Processing Systems, 36, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.311324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.993852Z digest=sha256:3bd7d61d29604ddaa64ad2873245407d242578bc48b77a0cde7758b9befb57f3

Observation 3b220d1b-48cf-46f1-8116-12b54b6a27ea · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2023.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets When and why vision- language models behave like bags-of-words, and what to do about it? InThe Eleventh International Conference on Learning Representations, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.299125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:52.997084Z digest=sha256:37b02221420d4eed4d6f3eab4563b1a3ee4d4d50d5b723aa81933538a6c00a9d

Observation ed6d5e29-6cf9-4f0b-bf71-73cdb6b0e7d3 · outbound

This paper cites Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.000363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.000363Z digest=sha256:53df115e41857ec9194d028a6e2561fb819966346b918694c450f0dedab2b77e

Observation 9b14da76-f3fe-4ef7-9011-0ab3907d6449 · outbound

This paper cites Investigating compositional chal- lenges in vision-language models for visual grounding.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Investigating compositional chal- lenges in vision-language models for visual grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.287164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:53.003972Z digest=sha256:40840804c98baaffba74c135c715d3adf9f6ee85eedf64331d497db9f7d1abcf

Observation 6bec6939-e9bd-4efe-a87d-d4f9fac82942 · outbound

This paper cites Sigmoid loss for language image pre-training.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets Sigmoid loss for language image pre-training

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:45:53.274342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:45:53.007487Z digest=sha256:a23715bab88898a0d30b04b844e65cd4aac56b65ee7d03d8c47e1152a1342c40

Observation b54f18df-cf1d-496b-8186-365b28cf9c0a · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.011960Z digest=sha256:5948dfdea17e6dc3ce94c9a5500f89a5b718ced182a9052526216d379cdcfdb9

Observation 6faa031f-bd95-4ea0-9866-de705e61db24 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:45:53.015838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:45:53.015838Z digest=sha256:bc35d055cd7586d9de287d2a6e85065f3b74e33c7e086ca3011b32c7a193041e

Pith citing papers

No inbound Pith citation observations are available.