Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:59:01.698073Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2412.01814.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:59:01.698073Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T22:42:28.802465Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T16:27:09.235019Z
94 of 94 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2acbd428-3830-45fe-93d7-7c9c84b351b5 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Flamingo: a visual language model for few-shot learning.NeurIPS, 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c5260d-c378-44c8-8aab-632a635aa5f4 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Single-stage semantic segmentation from image labels
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7af55be0-2496-45be-b0f5-91b5cfa3523f · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Beit: Bert pre-training of image transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb619a82-6ccc-4ec8-b9fb-ae8d943b713f · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Food-101–mining discriminative components with random forests
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f0282ec-5e69-4cd3-bb3d-83ea39aee8fd · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Coco- stuff: Thing and stuff classes in context
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb87e830-9a4b-41c8-9b49-5b76c83f38cc · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Unsupervised learn- ing of visual features by contrasting cluster assignments
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1691d2f1-92b1-4bee-86bb-b7168eef9727 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Emerg- ing properties in self-supervised vision transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d626e36a-75bb-4367-adba-49148d239221 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146a3257-d629-41cb-946e-a8dab6caad62 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af539555-6e7f-4cf7-81ce-21dd28a48688 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sharegpt4v: Improving large multi-modal models with better captions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05374346-b9bb-4f52-9ec7-07d24727a305 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training A simple framework for contrastive learning of visual representations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51cd77d6-93a2-448a-8d1d-57492678cfa7 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Intriguing properties of contrastive losses
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c66352b4-f5a5-494d-af9c-d757c89906f0 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Reproducible scal- ing laws for contrastive language-image learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac515e9e-1cc8-417a-a635-90ae73bfe593 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Describing textures in the wild
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f047899-dd01-40f4-bd39-2fe402c8ff26 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training The cityscapes dataset for semantic urban scene understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d4c787a-b8b4-44b8-98af-ca5352b00144 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Democratizing contrastive language-image pre- training: A clip benchmark of data, model, and supervision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d92ab704-8f91-40fc-b28e-11168bc64594 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training InstructBLIP: Towards general- purpose vision-language models with instruction tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 65aa734b-2a1e-4806-87c8-a67628a91d5a · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Imagenet: A large-scale hierarchical image database
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9755562e-6caf-4b31-9f68-4c6310447869 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Redcaps: Web-curated image-text data created by the people, for the people
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9fdc2215-99ad-47db-b60b-e55efbff62bb · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Bert: Pre-training of deep bidirectional trans- formers for language understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 54b74ef3-05b9-41ce-9a5f-16e5f7372504 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Maskclip: Masked self- distillation advances contrastive language-image pretraining
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1790a824-4b22-4503-b7a3-b9e8bdc441a1 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training The pascal visual object classes challenge: A retrospective
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b5a8220-4685-42bb-8275-7fb415448414 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Improving clip training with language rewrites
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e5c1d5e0-84cc-4f6e-b252-745503d192ec · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b77ac433-95a3-4132-8d29-870e71ba727c · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Dat- acomp: In search of the next generation of multimodal datasets
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d4242c3e-b862-44ba-a27d-147635fc8e42 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretraining
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 47fc3fd9-cb86-4b10-ac76-f65eef6c157a · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Softclip: Softer cross-modal alignment makes clip stronger
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b3efd4d9-4696-45e9-b65c-f24e80455272 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training HiCLIP: Contrastive language-image pre- training with hierarchy-aware attention
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d451149-5157-482a-ace4-e7c9c3d023d9 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Bootstrap your own latent-a new approach to self-supervised learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d814104-8839-4172-9320-a1fed898f0ba · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training A survey on self-supervised learning: Algorithms, applications, and future trends
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b0170a05-e30f-49b9-b94e-122ecd1a81d9 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Masked autoencoders are scalable vision learners
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2335dbc5-f7fe-40a0-a4dd-cca0ff73c0aa · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Probing image- language transformers for verb understanding
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5d231367-4c5b-417e-8351-394b21ff393b · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ca07a299-f29a-421c-9e43-0a0971ec2ec0 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fffa3571-1830-41bf-aa93-51728aa71ca9 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Scaling up visual and vision-language representation learning with noisy text supervision
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1aca144b-b25e-4dee-9f8a-40c48c91ee2b · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 28205008-6967-4e81-aef1-7cb092c5ae7b · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Expediting contrastive language-image pretraining via self-distilled encoders
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 40f8976a-0b8f-46b7-9d89-b3f3bced11f2 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training 3d object representations for fine-grained categorization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b4732ce-b8e6-4329-becc-b55ac72e6be0 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learning multiple layers of features from tiny images
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85dcf65-9397-4d82-9549-8f34ebb734f1 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Veclip: Improving clip training via visual-enriched captions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7e59a81b-f756-43c3-8aec-400d0b2d4ea2 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Modeling caption diversity in contrastive vision- language pretraining
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7fd87592-cd6e-4d5e-9886-16ee48c13452 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Uni- clip: Unified framework for contrastive language-image pre- training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db64e3b-95be-4353-9b3d-d4b5e785db79 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43df981f-28fb-4fa4-8172-f298f464b8fa · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Addressing feature suppression in unsupervised visual representations
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cf4fba8c-a748-4ffe-9b38-93ed858ea971 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ee8fa206-f536-4ad2-aaa9-7756346109bf · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Evaluating object hallucination in large vision-language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e977972d-a9a7-491c-aa37-5f47d14a930e · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Scaling language-image pre-training via masking
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 270ad24b-2677-43c9-ab8d-e79c5d77fab9 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Microsoft coco: Common objects in context
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 914a2adc-52a7-4af9-9e17-3b34d0f2a532 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Improved baselines with visual instruction tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ae13cc-3bc3-4405-9b80-5b2596a65bac · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Visual instruction tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b2c064d-bf1e-444d-9ba4-aa33aaf72ab8 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training MLLMs-Augmented Visual-Language Representation Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5896950-44b3-45f2-afbe-4760779e1c28 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Decoupled Weight Decay Regularization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef9242d-acb6-4c19-bba0-a6fea2e89433 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373a3b4f-5f65-4bc8-adfc-055248b60599 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Fine-Grained Visual Classification of Aircraft
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3a5c50e-a21e-4dae-a5f6-609b1c42243a · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training The role of context for object detection and se- mantic segmentation in the wild
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0f800bc9-b72f-427b-a897-2cc42c6e7b24 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Slip: Self-supervision meets language-image pre- training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2a966b05-2b8d-4cad-9d34-84187d5fb175 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Silc: Improving vision language pretraining with self-distillation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d879a00-adcf-4924-86b2-9d24fce288ba · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Automated flower classification over a large number of classes
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0d779462-8cdc-4406-b167-3846f737a30e · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Docci: De- scriptions of connected and contrasting images
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6fed06a7-ab62-4720-9f9a-80f41a01e56b · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Representation Learning with Contrastive Predictive Coding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc963c2f-7d71-4e18-b421-44f0cac54772 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Cats and dogs
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 56516dae-85b5-4fc2-96f8-d28eebef2ac6 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Language models are unsu- pervised multitask learners
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f7c7423-52ef-4fd7-b347-5671b729535b · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learn- ing transferable visual models from natural language super- vision
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e41a84c-d228-4333-9866-2c8738021ce2 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Can contrastive learning avoid shortcut solutions? NeurIPS, 2021
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 23470b4c-4ffa-4658-a4ac-1c38628389b6 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Building vision-language models on solid foundations with masked distillation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5150e07e-858b-4834-82d4-a874478e2811 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb4a333-6b53-431b-9cea-627257c7ccfd · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d10c965-c3a8-47ef-8fb9-e9bad0d5d146 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9386d0-38c8-4069-96b9-d70be07cb39c · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Towards vqa models that can read
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c88223e4-099f-478e-afbf-79b9f07c9bf9 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training From Pixels to Prose: A Large Dataset of Dense Image Captions
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2089c6b-2dbf-432e-b1e6-afb7b05f77a8 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Feature dropout: Revisiting the role of augmentations in contrastive learning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c7b6a6dc-e200-46fd-a1ff-931d19e6bebc · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Yfcc100m: The new data in multimedia research
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4880a0f0-34e0-4f3d-85d4-a49d81b36f41 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Winoground: Probing vision and language models for visio- linguistic compositionality
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3259ba9a-d2b6-4802-b975-1155a34081fd · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0036db66-678d-4c50-9f3e-36526ca122af · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training A picture is worth more than 77 text tokens: Evaluating clip- style models on dense captions
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2900c910-458d-49d5-b188-0413a190181e · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Mobile- clip: Fast image-text models through multi-modal reinforced training
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6849a673-9f1b-4338-b2d0-f7d815f05342 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sclip: Rethinking self-attention for dense vision-language inference
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a1d627ca-69a2-4ce2-a4f2-b832652a5576 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Lotlip: Improving language-image pre- training for long text understanding
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ef549f9d-6e36-4b5f-848d-041213f458dd · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sun database: Large-scale scene recognition from abbey to zoo
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 720276a8-722b-40da-b1ec-19d16d439e14 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Demystifying clip data
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 84b90c2c-6d4c-426c-bd13-23a655ec38fa · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Which features are learnt by con- trastive learning? on the role of simplicity bias in class col- lapse and feature suppression
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3757f023-7e51-4f95-9c2e-f3aaf559ceff · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Alip: Adaptive language-image pre-training with synthetic cap- tion
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2730cf4a-de97-418a-ab08-573ba4d2a5d3 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Filip: Fine-grained interactive language-image pre-training
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 01343171-1b51-4250-be3f-a5fbf6dddce0 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5beb5583-4a5c-4ff2-beb3-70e86ae85bde · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Coca: Contrastive captioners are image-text foundation models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 97cffcac-a975-49b9-9869-0fb1cd7e82b1 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6f3fa02b-44ab-492e-a293-65fd3b14c078 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training When and why vision- language models behave like bags-of-words, and what to do about it? ICLR, 2023
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c23bf18c-567b-448c-a311-acb698f5eac9 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sigmoid loss for language image pre-training
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2c698dfb-1afa-46f2-88ff-d0bfde6255f5 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Long-clip: Unlocking the long-text capability of clip
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ec95625c-92c6-4024-9419-61aac0f964d3 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learning the unlearned: Mitigating feature suppression in contrastive learning
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1ffc3780-82ce-402e-84b0-b1c8857f3bc4 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Dreamlip: Language- image pre-training with long captions
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f2f462a3-d922-4a66-bd84-0fea60865499 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Semantic under- standing of scenes through the ade20k dataset
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0bf5983b-acfa-47f0-a1fe-d5ee5c46ed42 · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Extract free dense labels from clip
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a3912940-b9a1-4d7d-88a2-3b51915297de · outbound
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b4cfeb0-ea76-4d11-a3a5-b6f6099bc7ba · inbound
Multi-FRuGaL: Multimodal Flexible Redundancy-aware Decomposed Gated Learning for Cancer Diagnosis and Prognosis COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.