Pith. sign in

Paper Citation Record · LEDGER

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.08000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08000 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:48.742052Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b196439a-1861-4339-b27d-57c23a0700be · outbound

This paper cites Learning transferable visual models from natural language supervision.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.403863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:44.404148Z digest=sha256:3039de6e0e2283be0588f124011f1a9fbd8623ebf8356d8f53f9fa6e4b35fd3c

Observation c88cb4d2-17f0-4076-a45d-90484707a1c3 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Reproducible scaling laws for contrastive language-image learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.526529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.526529Z digest=sha256:625fba887902eebff111660e4a44ea814adaf65626959964bdf620499e162b2f

Observation c011eae3-7836-4a3b-b2ea-870a6c0caea6 · outbound

This paper cites GPT-4 Technical Report.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.629556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.629556Z digest=sha256:313793448cd65c5d2fb10fc6920ec5b441126c63196bb670055808ad53ee4ed1

Observation 0a4e9da3-a4ec-46ae-bc55-b4275ba7210d · outbound

This paper cites Visualinstructiontuning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Visualinstructiontuning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.252715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:44.743931Z digest=sha256:788373bfc9fc9895f762cba4dc42bc840436f33edff84a37de7ddb6b9e692b50

Observation c41c3b77-5c8b-437a-9899-718f0553e9b6 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.845309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.845309Z digest=sha256:003b717473ff31e642c611110a781dc537c316960c65a427595569345081ae76

Observation 40584af3-8ba4-42ac-aa75-446cec0fdab3 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.933692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.933692Z digest=sha256:23e7b2a0891d288b7f4b39cd6d098d9cfb00c5c3d89d3a3c9bddd716600e090b

Observation 83d88bc5-375e-4231-84ed-706503d94dec · outbound

This paper cites Measuring Robustness to Natural Distribution Shifts in Image Classification.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Measuring Robustness to Natural Distribution Shifts in Image Classification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.071477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.071477Z digest=sha256:c263eb007874369af139a7a1bdaa119337fc1ed4e4c4601499c83d5446ab1d60

Observation cdeac91e-f2f1-46a2-b25d-d012f3d9d14f · outbound

This paper cites The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.197907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.197907Z digest=sha256:bd825b9d8b2063fac21746c93e79701a19bac564f8def62f2f9ccf1d3f3662b7

Observation 9b5d6ad1-921b-4dcc-8427-8a7f921c63ca · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.174976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:45.324151Z digest=sha256:d8b1b75bd37b1f9bf6bbd90408476000c0e9de0bd927e5f392ad57ebd3760e27

Observation 1e89ab05-e4e0-417c-a567-7af021e4aec6 · outbound

This paper cites Data determines distributional robustness in contrastive language image pre-training (clip).

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Data determines distributional robustness in contrastive language image pre-training (clip)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.066083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:45.443919Z digest=sha256:5ad6000affe30bf83b5f9252b4120ad4e73de6e36f22891a83eb390bc66e98f7

Observation 6e24f533-29e8-4405-80ef-4b3ca13b1423 · outbound

This paper cites The neglected tails in vision-language models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The neglected tails in vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:45.583592Z digest=sha256:d998d04ed92f0ae2e1e28b29fee3eb9e2b95a9576e4fa5b3287a4581c5e80375

Observation 98585498-f97c-4134-b086-41f08b14aaf2 · outbound

This paper cites zero-shot.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models zero-shot

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.840210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:45.659673Z digest=sha256:a50065d618ee5106172790245d169ebaa241a797350784b2c76e01f6ea58be6e

Observation 4ee03997-7c78-4265-b26f-1a83e7a71c79 · outbound

This paper cites Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.745787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.745787Z digest=sha256:8f746509212ddbd416570983312700d8f49bed0ce3bbc0fbd8401eae50017ed8

Observation fd56ea5e-25a5-4d91-9f4c-6d6187645e1b · outbound

This paper cites Deciphering the role of representation disentanglement: Investigating compositional generalization in clip models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Deciphering the role of representation disentanglement: Investigating compositional generalization in clip models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.676249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:45.842927Z digest=sha256:d1c3b9800786354f0bf9a165125e22a898f278b226fd4bedcec7689842943659

Observation 25838902-9322-43f2-ba82-ce2c655a5bce · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.923104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.923104Z digest=sha256:e70b8c7d488b89b8d67456886540826eb573acd49b080ea130c3a45e89c97c47

Observation 6f1add8e-df0b-4439-932e-52027ab2d4e3 · outbound

This paper cites @ crepe: Can vision-language foundation models reason compositionally?2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10910–10921, 2022.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models @ crepe: Can vision-language foundation models reason compositionally?2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10910–10921, 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.096827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.096827Z digest=sha256:a691df848a45a4ee662817b001e9251997386b4a0d70c625c10e73a1639d0444

Observation 207f01f4-20ed-485a-81b0-a50c81a42ec9 · outbound

This paper cites Sugar- crepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Sugar- crepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.545135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:46.233303Z digest=sha256:a80b2528c1df6912cfc20b24ce41eea3812e48a6b0d59712893f4f2a25748a85

Observation 60d29984-4d5e-4197-be41-e06da54d3ebe · outbound

This paper cites A Sober Look at the Robustness of CLIPs to Spurious Features.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models A Sober Look at the Robustness of CLIPs to Spurious Features

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.319612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.319612Z digest=sha256:4244a169032d80beca0db27ed2760815c66077ecf963f22f753eeae262bd8f1a

Observation ba669502-0b95-4bcf-994b-d60523b9e356 · outbound

This paper cites Word association norms, mutual information, and lexicography.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Word association norms, mutual information, and lexicography

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.407432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:46.419843Z digest=sha256:a9b4d8b349af77cd8a9ca75c0f8fd906f92085dae937dff7cf7e4b7c156430fb

Observation 28885a78-4344-42a2-8dc0-c171b6de1729 · outbound

This paper cites Improved baselines with visual instruction tuning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Improved baselines with visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.490093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.490093Z digest=sha256:270cdaeea9ddfd54bc6566664034b4c9123e86b43db57c99164e6a275b338ea1

Observation f16bf04a-f7e2-4076-9a3b-7985e3121b76 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.563950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.563950Z digest=sha256:52c4c9e0bd35f91a47801950af93061b583095f6268b3d5c5cbb58120732f3b2

Observation b72d4b65-5d59-420b-a6f5-e07168d0e385 · outbound

This paper cites Two multivariate generalizations of pointwise mutual information.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Two multivariate generalizations of pointwise mutual information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.247942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:46.701812Z digest=sha256:8c1648fbf30d260edbdc79576f1b4ac3178fca88c34fb3538a8469091fa8131d

Observation 1dd04bef-0152-489f-bd1f-a6b49fee3bc1 · outbound

This paper cites Lawrence Zitnick, Devi Parikh, and Dhruv Batra.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Lawrence Zitnick, Devi Parikh, and Dhruv Batra

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.128937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:46.817205Z digest=sha256:ab51cb9534551986f877aa5a966b53124d44c63217b285dbc7ee8d43a0e0017e

Observation 2121567b-7e69-49a4-bde3-1243b2b32d97 · outbound

This paper cites The Llama 3 Herd of Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.919345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.919345Z digest=sha256:7170cbd322cbe2ad53244b5436bfad1813ecbe15eddf39e683efe6cfb2f36e0b

Observation 2ca3ebd8-02d4-4d88-aed2-02bb93b9094d · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Flux.https://github.com/black-forest-labs/flux, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.040675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.040675Z digest=sha256:0c56050a5529bec4ed8e975a625d9cd3dc56876b6e0410642bf5b9ddfcd6ff21

Observation 29bd6e86-cc0b-483e-b38c-3a47cc355b84 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.160787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.160787Z digest=sha256:c6011a2adb10d8fd3699a36f5de04c5c22a4fbba82bc8d9765911d20e9989327

Observation 250aebb7-906b-4c67-9cf9-ee4314c75f4f · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.262926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.262926Z digest=sha256:536a9843eaf0e4544054efc1de015fe64b2ed68159e9e959572cb71e261ef59d

Observation 339cc28c-f88e-461b-a376-9e1d25dc31cc · outbound

This paper cites Towards vqa models that can read.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Towards vqa models that can read

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.999667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:47.314389Z digest=sha256:0ca9fc86963f1d505d92548fdd2464658343e30a187bb94b1b081daf1787e935

Observation 09724e1e-11fb-4555-88d7-7bb5ee6c3fcb · outbound

This paper cites Recognition in terra incognita.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Recognition in terra incognita

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.800409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:47.418010Z digest=sha256:29d0a9e869c4d526c90b141684af50dc5ce2596911cfe988ff8861b47311be1a

Observation 266d094f-4272-4fc4-ad49-9c6785ed3322 · outbound

This paper cites Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study.PLoS medicine, 15(11):e1002683, 2018.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study.PLoS medicine, 15(11):e1002683, 2018

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.443100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:47.520665Z digest=sha256:b9d5637ab38a6e1524c06668429b12da1d69b3eaa144e38352fd7d4066ff93b3

Observation 47a9c4c8-3021-4671-9691-255ae3e0b563 · outbound

This paper cites Hashimoto, and Percy Liang.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Hashimoto, and Percy Liang

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.603989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.603989Z digest=sha256:410d4531b4dc3a33cdca3ba2a95eda0a7073f7864c1b77e486d90486f20b613e

Observation 3eb9b301-b9c7-4258-9f7d-2ec1efb4b97e · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2020.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Shortcut learning in deep neural networks.Nature Machine Intelligence, 2020

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.185552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:47.674980Z digest=sha256:689a2ca71e78ef19adab9bce6719c2b7a31a47cbf962be58416e36731d667208

Observation cd90cc6b-1161-4104-862c-38303358836a · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.863554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:47.791137Z digest=sha256:4a050373e15338205e59c2fd8b140a5fb2aeaf9205310f2717b2b283013aac3d

Observation a353761a-8675-4775-ba49-8e40afb2a124 · outbound

This paper cites COVR: A test-bed for Visually Grounded Compositional Generalization with real images.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models COVR: A test-bed for Visually Grounded Compositional Generalization with real images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.888275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.888275Z digest=sha256:11c61f53e7210a95b3d35be20653d361fe8dd0a6d3647c03646233e854fe122b

Observation cddf270b-ab97-4fbf-939b-36699f877d03 · outbound

This paper cites Does CLIP Bind Concepts? Probing Compositionality in Large Image Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.029297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.029297Z digest=sha256:1d0e21eed9fe7f7732b23864d9e36fe4d041166014f6a86b923e6e31c50170c5

Observation e72a3b76-2f3c-414e-b73d-298ca1be321a · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.147656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.147656Z digest=sha256:ddded60da81eab83b624ee883660350ada4e7a3752a5274356c715cb1f53f216

Observation 9cc3661c-57d7-4361-ab1d-fe160a633f48 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.223641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.223641Z digest=sha256:a5b0290614eaa0a74e6839227d65c8a7abdb91af315127b8946d64d8328f08d6

Observation c98760e6-266c-46ff-82b8-da61dc20f370 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.297322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.297322Z digest=sha256:fe09f51c8678c4ac1b41ced64c11e1d7d6ae17c08178dc8bddd1bb19034927c5

Observation 8b821a8e-93e2-45bc-a01f-1e1abb7fc48b · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.326309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.326309Z digest=sha256:84c3a63c3f3763cc7f064f4a72f7334b8f9f31addbeb0e34482c7edcb68fbf8a

Observation 08c3f604-a780-46fe-b6b7-cb5b067c7cd4 · outbound

This paper cites MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.390311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.390311Z digest=sha256:f74465fa2814048be1fa1d3559d86415005c78dd57f15b35aaf51bfc1d7c84dc

Observation 353242e8-6191-4d6f-b25c-f2d98f403f79 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.461452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.461452Z digest=sha256:ee6351855164bbc8c884a541c81efadfaf1483e0d676d65cb0a2821e69b7305c

Observation 63fcb153-04db-4f4e-a61f-25e58fd4085f · outbound

This paper cites Segment anything.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Segment anything

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.683465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:48.536131Z digest=sha256:50762bbbc6ed25faccefa771b4ac3bedfe57999481eb03764a2eb423d836fe2a

Observation 0293cf60-53c7-4324-bdb7-91f2a2f60ea7 · outbound

This paper cites Robust fine-tuning of zero-shot models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Robust fine-tuning of zero-shot models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.612236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.612236Z digest=sha256:15596b5803ef0f902e38cae74d2d8bdcf22c8793e4a1638f443db0ce7893fec6

Observation 9adf4580-b315-4ca6-9b5c-3bc902cd207a · outbound

This paper cites Openclip, July 2021.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Openclip, July 2021

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:49.533161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:48.687684Z digest=sha256:7e836b22574668f217aea18125cf873b2dd1b7260cc84c612c8d300981ed9513

Observation 2026a014-5c8c-4b95-bae7-1d3c8799e045 · outbound

This paper cites visualizable.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models visualizable

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.387084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:32:48.742052Z digest=sha256:f55c83960acca245efe61e26ae4fbbb90194c8b4324aab19b8e2ce8b2f8322aa

Pith citing papers

No inbound Pith citation observations are available.