Pith. sign in

Paper Citation Record · LEDGER

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models

As of 16 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2507.08000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08000 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:32:48.742052Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b196439a-1861-4339-b27d-57c23a0700be · outbound

This paper cites Learning transferable visual models from natural language supervision.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Learning transferable visual models from natural language supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.403863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:44.404148Z digest=sha256:7444b25f0f517cd8fdd3dbe2a74c1da4a39753ea3b3fc4800f6c695cbd599426

Observation c88cb4d2-17f0-4076-a45d-90484707a1c3 · outbound

This paper cites Reproducible scaling laws for contrastive language-image learning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Reproducible scaling laws for contrastive language-image learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.526529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.526529Z digest=sha256:efe75a6e847944b806bae6f0e7308058afd23e5b199b2273e57cc4924d056c1c

Observation c011eae3-7836-4a3b-b2ea-870a6c0caea6 · outbound

This paper cites GPT-4 Technical Report.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.629556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.629556Z digest=sha256:7e9d5d7c1656f0d4b85c011a958d2a5830260c3efe94973a3dec60a34c7417d6

Observation 0a4e9da3-a4ec-46ae-bc55-b4275ba7210d · outbound

This paper cites Visualinstructiontuning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Visualinstructiontuning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.252715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:44.743931Z digest=sha256:c7f77b8a9a5e34a62ec307a5bdda5d34808f1397fcc1b475aac34fa8317db49b

Observation c41c3b77-5c8b-437a-9899-718f0553e9b6 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.845309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.845309Z digest=sha256:1d468300896aba1a18de6a26d517b7863bb844c27e907fb5ca36e5d08347a584

Observation 40584af3-8ba4-42ac-aa75-446cec0fdab3 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.933692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.933692Z digest=sha256:3b515ea10072a44155c6fb06b73306ae71c25888c09440e23ba6571e52375dec

Observation 83d88bc5-375e-4231-84ed-706503d94dec · outbound

This paper cites Measuring Robustness to Natural Distribution Shifts in Image Classification.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Measuring Robustness to Natural Distribution Shifts in Image Classification

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.071477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.071477Z digest=sha256:bc02cf4503186c4dbc67fcee69ebc10dad88681947fedc6061104bd787f2227c

Observation cdeac91e-f2f1-46a2-b25d-d012f3d9d14f · outbound

This paper cites The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.197907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.197907Z digest=sha256:b0d8e1b472cffc79427d8902d0a2d9bb599689a0bf005a099a9016b60199909a

Observation 9b5d6ad1-921b-4dcc-8427-8a7f921c63ca · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.174976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:45.324151Z digest=sha256:4d7e93c4c7fb337a4a6c53bf51468c97fbadeb1be1c58772b7f08baddfa8d022

Observation 1e89ab05-e4e0-417c-a567-7af021e4aec6 · outbound

This paper cites Data determines distributional robustness in contrastive language image pre-training (clip).

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Data determines distributional robustness in contrastive language image pre-training (clip)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:52.066083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:45.443919Z digest=sha256:ac6ddae2847ecc66f477be20dffd19c77286c9198ae6376ce1035243ec3df79d

Observation 6e24f533-29e8-4405-80ef-4b3ca13b1423 · outbound

This paper cites The neglected tails in vision-language models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The neglected tails in vision-language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.945896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:45.583592Z digest=sha256:119756b0e28546a7ea6c6e0bda43f1b03229ac28a5d40a0d2d11e0789a7b49b5

Observation 98585498-f97c-4134-b086-41f08b14aaf2 · outbound

This paper cites zero-shot.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models zero-shot

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.840210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:45.659673Z digest=sha256:55d928530fa97d9664116b34f253a745e1201c7f19234e2ac0b27dcd0c478fbc

Observation 4ee03997-7c78-4265-b26f-1a83e7a71c79 · outbound

This paper cites Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.745787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.745787Z digest=sha256:b708599ca615d38ce1d53663be94fd505a88e86d67617777b88fa66b6ec008e6

Observation fd56ea5e-25a5-4d91-9f4c-6d6187645e1b · outbound

This paper cites Deciphering the role of representation disentanglement: Investigating compositional generalization in clip models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Deciphering the role of representation disentanglement: Investigating compositional generalization in clip models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.676249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:45.842927Z digest=sha256:bc6db4ee78a9b1d20973793b905ecd48b77b0239e90ee7908e1e7cdedecf0ab2

Observation 25838902-9322-43f2-ba82-ce2c655a5bce · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:45.923104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:45.923104Z digest=sha256:8c07fa6dabfaee5178c5b8ece6ba27c888bc1d6166148f9fb9646aa95307165a

Observation 6f1add8e-df0b-4439-932e-52027ab2d4e3 · outbound

This paper cites @ crepe: Can vision-language foundation models reason compositionally?2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10910–10921, 2022.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models @ crepe: Can vision-language foundation models reason compositionally?2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10910–10921, 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.096827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.096827Z digest=sha256:307f4aef3d2e236fc9b3759b13a0343595dca9ecfcc0eed4624425c3454fe7aa

Observation 207f01f4-20ed-485a-81b0-a50c81a42ec9 · outbound

This paper cites Sugar- crepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Sugar- crepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.545135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:46.233303Z digest=sha256:9a1e179868ed9be0499b719f49b57899b1efbe8078a5f37bfe4c83b218fc8991

Observation 60d29984-4d5e-4197-be41-e06da54d3ebe · outbound

This paper cites A Sober Look at the Robustness of CLIPs to Spurious Features.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models A Sober Look at the Robustness of CLIPs to Spurious Features

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.319612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.319612Z digest=sha256:3c4bd25a20f2934c8bb5e5708540c465f6168642060bb190ac4a780800be1399

Observation ba669502-0b95-4bcf-994b-d60523b9e356 · outbound

This paper cites Word association norms, mutual information, and lexicography.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Word association norms, mutual information, and lexicography

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.407432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:46.419843Z digest=sha256:617266fa05aa6cf2b4ec90af41c55b175b633b0f1d5a252ff2b5024d12b0f85a

Observation 28885a78-4344-42a2-8dc0-c171b6de1729 · outbound

This paper cites Improved baselines with visual instruction tuning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Improved baselines with visual instruction tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.490093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.490093Z digest=sha256:263162aa21443c224e73333ca41da1372fd59e641389a6fc7a085f7018a9429f

Observation f16bf04a-f7e2-4076-9a3b-7985e3121b76 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.563950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.563950Z digest=sha256:811a46c7454eac968f72a2be9e9626450f4ae6882c4366c01321f25a8cfc9f87

Observation b72d4b65-5d59-420b-a6f5-e07168d0e385 · outbound

This paper cites Two multivariate generalizations of pointwise mutual information.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Two multivariate generalizations of pointwise mutual information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.247942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:46.701812Z digest=sha256:98845eeb3074214e7fd5214a7d6d4dddd1726f42f545f186328c9c764396b8e5

Observation 1dd04bef-0152-489f-bd1f-a6b49fee3bc1 · outbound

This paper cites Lawrence Zitnick, Devi Parikh, and Dhruv Batra.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Lawrence Zitnick, Devi Parikh, and Dhruv Batra

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:51.128937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:46.817205Z digest=sha256:72597edacbcdf6ef9b91fd9d12970d2dab3df60cd7bb17791cebf834976fa1a3

Observation 2121567b-7e69-49a4-bde3-1243b2b32d97 · outbound

This paper cites The Llama 3 Herd of Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:46.919345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:46.919345Z digest=sha256:25b2a5846cbf7177f7bd83e78b71a740e9b4166292dc28afae6874f529f004a6

Observation 2ca3ebd8-02d4-4d88-aed2-02bb93b9094d · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Flux.https://github.com/black-forest-labs/flux, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.040675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.040675Z digest=sha256:616644353b8945deb2d6c36fe274d5a9dd22909456badb6efe5412e3c1a72549

Observation 29bd6e86-cc0b-483e-b38c-3a47cc355b84 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.160787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.160787Z digest=sha256:23571467807795f7db7d4c44a711918453059c3ad8a9d95ce94ec8b5f1ece9c4

Observation 250aebb7-906b-4c67-9cf9-ee4314c75f4f · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.262926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.262926Z digest=sha256:d037baf3252ff60d98e7461ab47ba156e6fde9e20c01064d1624297709232a21

Observation 339cc28c-f88e-461b-a376-9e1d25dc31cc · outbound

This paper cites Towards vqa models that can read.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Towards vqa models that can read

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.999667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:47.314389Z digest=sha256:26dbded75fbf44bbe206de355c09e4a58c4bd0267da4d579cb0ddb92110167b6

Observation 09724e1e-11fb-4555-88d7-7bb5ee6c3fcb · outbound

This paper cites Recognition in terra incognita.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Recognition in terra incognita

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.800409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:47.418010Z digest=sha256:3053da0951310ca035953e033481fdac2cee485e93025cd97a8d20ff6e1af946

Observation 266d094f-4272-4fc4-ad49-9c6785ed3322 · outbound

This paper cites Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study.PLoS medicine, 15(11):e1002683, 2018.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study.PLoS medicine, 15(11):e1002683, 2018

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.443100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:47.520665Z digest=sha256:b069ffca99c4bba41410a4527c4cbd8777a08421e22dbd0f4350b8a1f8aede9d

Observation 47a9c4c8-3021-4671-9691-255ae3e0b563 · outbound

This paper cites Hashimoto, and Percy Liang.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Hashimoto, and Percy Liang

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.603989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.603989Z digest=sha256:c59bbf7a02204d48018f756d4d9060362ffa347eaca350764339c21c63dceb7f

Observation 3eb9b301-b9c7-4258-9f7d-2ec1efb4b97e · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2020.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Shortcut learning in deep neural networks.Nature Machine Intelligence, 2020

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:50.185552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:47.674980Z digest=sha256:120f35374cf187cc2a008efe3c149d5a787e9f6a46bf9453df7c4d08ced50f76

Observation cd90cc6b-1161-4104-862c-38303358836a · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.863554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:47.791137Z digest=sha256:135824e8b3bb38ff909e0c1e8723f75ff5aa0684dd2b0ce8cf61fbb3f9a06f26

Observation a353761a-8675-4775-ba49-8e40afb2a124 · outbound

This paper cites COVR: A test-bed for Visually Grounded Compositional Generalization with real images.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models COVR: A test-bed for Visually Grounded Compositional Generalization with real images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:47.888275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:47.888275Z digest=sha256:196203b92ffc927b0172320fff6393d5a2c85d5ca0c28bb3aa98713d709a8eab

Observation cddf270b-ab97-4fbf-939b-36699f877d03 · outbound

This paper cites Does CLIP Bind Concepts? Probing Compositionality in Large Image Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.029297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.029297Z digest=sha256:1d1ca8ac52ed89e95e5dec6f18b157ce723aea8d148c80c150d00bad458b9dd9

Observation e72a3b76-2f3c-414e-b73d-298ca1be321a · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.147656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.147656Z digest=sha256:78a25ca55dc1460d09159ed77b782bac3ab73081771ad382ec87d92d991e0642

Observation 9cc3661c-57d7-4361-ab1d-fe160a633f48 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.223641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.223641Z digest=sha256:76f1cde39bbc459ff76a131967f1c7d627ca624d4ff4a9ceb466c78c7e9d6f0c

Observation c98760e6-266c-46ff-82b8-da61dc20f370 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.297322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.297322Z digest=sha256:62ea4135c6680e22fa4c7efe462e3a0b69bfa3f3baa51e8b4ace88dffd67faee

Observation 8b821a8e-93e2-45bc-a01f-1e1abb7fc48b · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.326309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.326309Z digest=sha256:3e3e54540b22bf07d3049cca03f216c166042e38f0293f1d9c35a3d599d4aaad

Observation 08c3f604-a780-46fe-b6b7-cb5b067c7cd4 · outbound

This paper cites MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.390311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.390311Z digest=sha256:122bc1b8b7c241b5e21c56aa9261b2badd5bbb0651be7359c71cdb02811e68c1

Observation 353242e8-6191-4d6f-b25c-f2d98f403f79 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.461452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.461452Z digest=sha256:fe024827bff8518487a958484f8672170e495c27aac891ec46159e3c826bc6a5

Observation 63fcb153-04db-4f4e-a61f-25e58fd4085f · outbound

This paper cites Segment anything.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Segment anything

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.683465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:48.536131Z digest=sha256:bb0ded3269e609b253f407fd8d8d1015c26abac54a2008d260a5c5d9a42cbbcf

Observation 0293cf60-53c7-4324-bdb7-91f2a2f60ea7 · outbound

This paper cites Robust fine-tuning of zero-shot models.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Robust fine-tuning of zero-shot models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.612236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.612236Z digest=sha256:5ad3f8da92def13fff2f20bc784065de718491a77c16a152ccf930cb89c67e88

Observation 9adf4580-b315-4ca6-9b5c-3bc902cd207a · outbound

This paper cites Openclip, July 2021.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models Openclip, July 2021

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:32:49.533161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:48.687684Z digest=sha256:898f4b2e74ba82d8a5bd9c527c8ad88b8e06f482b1c84d283e9f70d4164cd2c6

Observation 2026a014-5c8c-4b95-bae7-1d3c8799e045 · outbound

This paper cites visualizable.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models visualizable

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:32:49.387084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:32:48.742052Z digest=sha256:84e6bc7cd2cf7e013a1d6b36c8f6a82397c58a3d92a57da98ab7ab53f89152d9

Pith citing papers

No inbound Pith citation observations are available.