Pith. sign in

Paper Citation Record · LEDGER

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

As of 15 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 2 inbound Pith citation observations for arXiv:2412.01814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01814 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:59:01.698073Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:05:39.135041Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:27:09.235019Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact0
  • verified fuzzy59
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2acbd428-3830-45fe-93d7-7c9c84b351b5 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.432127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.432127Z digest=sha256:79fb92dbd19a113c0762984e2ef211cd777bed36f3dca87228330958d15b0544

Observation f1c5260d-c378-44c8-8aab-632a635aa5f4 · outbound

This paper cites Single-stage semantic segmentation from image labels.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Single-stage semantic segmentation from image labels

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.435687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.435687Z digest=sha256:54d5efd1ad0bfa1c9a8faf544eb9f81046fdc63c49cdc3c1a825c6e081007ad0

Observation 7af55be0-2496-45be-b0f5-91b5cfa3523f · outbound

This paper cites Beit: Bert pre-training of image transformers.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Beit: Bert pre-training of image transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.438584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.438584Z digest=sha256:3a69b1c3f051e5b1d7595a8bad9ae700f5097c42c27c67bbcf39c84522dbca2e

Observation cb619a82-6ccc-4ec8-b9fb-ae8d943b713f · outbound

This paper cites Food-101–mining discriminative components with random forests.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Food-101–mining discriminative components with random forests

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.441778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.441778Z digest=sha256:fb66702995d650ec2d1b0d76a4cb201361fb152541424856911cb17dedb99f5c

Observation 5f0282ec-5e69-4cd3-bb3d-83ea39aee8fd · outbound

This paper cites Coco- stuff: Thing and stuff classes in context.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Coco- stuff: Thing and stuff classes in context

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.444870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.444870Z digest=sha256:108f099330f0296828d58fd412633ff4814fd32d807246dfe2138399ae016f4f

Observation fb87e830-9a4b-41c8-9b49-5b76c83f38cc · outbound

This paper cites Unsupervised learn- ing of visual features by contrasting cluster assignments.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Unsupervised learn- ing of visual features by contrasting cluster assignments

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.448023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.448023Z digest=sha256:01decd918498dedf9aec174b5a36a17af5e9fd2dbcc37d074c09432254a7e4d3

Observation 1691d2f1-92b1-4bee-86bb-b7168eef9727 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.451411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.451411Z digest=sha256:411a826c156c68ae35cae9c51ee99021c757313dd625a6dc9039d3f5ca3eac37

Observation d626e36a-75bb-4367-adba-49148d239221 · outbound

This paper cites Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learn- ing to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.454495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.454495Z digest=sha256:e31d558d18ae4e616c9212e49531a17ee368ea1a1db6257e520aba50609d3253

Observation 146a3257-d629-41cb-946e-a8dab6caad62 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.457364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.457364Z digest=sha256:f06867c708fea89a62921b24f705a8820d811bf3fefe7e32bb929a61608985d7

Observation af539555-6e7f-4cf7-81ce-21dd28a48688 · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sharegpt4v: Improving large multi-modal models with better captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.460432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.460432Z digest=sha256:694435933c24a677379931f4af5b3ebf8dc0a8af8d02ad8ca15f61a6387cb1e0

Observation 05374346-b9bb-4f52-9ec7-07d24727a305 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training A simple framework for contrastive learning of visual representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.463365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.463365Z digest=sha256:549679b6177f694125fdb2cb2ba73e1dd4e23abf70e31db82bcaa290f43dea7d

Observation 51cd77d6-93a2-448a-8d1d-57492678cfa7 · outbound

This paper cites Intriguing properties of contrastive losses.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Intriguing properties of contrastive losses

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.466833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.466833Z digest=sha256:a06c7454f10dd5c7fa1914b12fa5c85707efef7bb3f857594970e34c2d4858ed

Observation c66352b4-f5a5-494d-af9c-d757c89906f0 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Reproducible scal- ing laws for contrastive language-image learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.469700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.469700Z digest=sha256:70471ae49af6509dad1ee721a3d6dc23eb0c858ce6b8be5c50c5f26b6d921006

Observation ac515e9e-1cc8-417a-a635-90ae73bfe593 · outbound

This paper cites Describing textures in the wild.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Describing textures in the wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.472620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.472620Z digest=sha256:ad967c253979631df9b1719ed70c5772cf16c277dfdf0bdf1f1a1a123a790c76

Observation 5f047899-dd01-40f4-bd39-2fe402c8ff26 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training The cityscapes dataset for semantic urban scene understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.475572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.475572Z digest=sha256:fd7d034dc4a473fef04338cfe8fc2d94b7e23e050b5b4c65fa2f984b5681f132

Observation 1d4c787a-b8b4-44b8-98af-ca5352b00144 · outbound

This paper cites Democratizing contrastive language-image pre- training: A clip benchmark of data, model, and supervision.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Democratizing contrastive language-image pre- training: A clip benchmark of data, model, and supervision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.353984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.478518Z digest=sha256:4fb1785b692eb7c0368ef5dbaf15fb8df551478c8bea71bbf1cecc6918f3adbd

Observation d92ab704-8f91-40fc-b28e-11168bc64594 · outbound

This paper cites InstructBLIP: Towards general- purpose vision-language models with instruction tuning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training InstructBLIP: Towards general- purpose vision-language models with instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.346004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.481577Z digest=sha256:cffe1ae6e67b2a5054e27581c6fb17c4909181fb6e079f3028f61b11e2dd1810

Observation 65aa734b-2a1e-4806-87c8-a67628a91d5a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Imagenet: A large-scale hierarchical image database

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.337905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.484471Z digest=sha256:79515e05a444dbc9806170005195d8fad9d4418e0a545905133d8558b9e1d841

Observation 9755562e-6caf-4b31-9f68-4c6310447869 · outbound

This paper cites Redcaps: Web-curated image-text data created by the people, for the people.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Redcaps: Web-curated image-text data created by the people, for the people

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.329703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.487324Z digest=sha256:50bb34608ad3ac3b10bd0dd8eac8445807ee3d84be9bb0a77172f3abc35f04d1

Observation 9fdc2215-99ad-47db-b60b-e55efbff62bb · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.321406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.490080Z digest=sha256:2dfba0befe8b190a3aa96559b790ab61a5f07d716ba1ef816d781b827a127401

Observation 54b74ef3-05b9-41ce-9a5f-16e5f7372504 · outbound

This paper cites Maskclip: Masked self- distillation advances contrastive language-image pretraining.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Maskclip: Masked self- distillation advances contrastive language-image pretraining

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.313247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.492548Z digest=sha256:d07c403fdcab765480dfff05775b4c778b6b110ff9f9b28a3e1c0441e56409c1

Observation 1790a824-4b22-4503-b7a3-b9e8bdc441a1 · outbound

This paper cites The pascal visual object classes challenge: A retrospective.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training The pascal visual object classes challenge: A retrospective

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.305043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.495242Z digest=sha256:447fabbd9af8c97396340eaebc10e1c80f338860ec6ba8fee3a91a823590ae4e

Observation 8b5a8220-4685-42bb-8275-7fb415448414 · outbound

This paper cites Improving clip training with language rewrites.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Improving clip training with language rewrites

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.296907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.497970Z digest=sha256:7eac13bbe12ed37e94d8f35c933845736d09f9b62d3c05bc6ae3d95d92d2f8a9

Observation e5c1d5e0-84cc-4f6e-b252-745503d192ec · outbound

This paper cites Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learning gener- ative visual models from few training examples: An incre- mental bayesian approach tested on 101 object categories

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.288559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.500653Z digest=sha256:1e072cf528434faa02e4c87ef26aeb77d7afe58a1537ec4775c21d5468e644c2

Observation b77ac433-95a3-4132-8d29-870e71ba727c · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Dat- acomp: In search of the next generation of multimodal datasets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.279224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.503384Z digest=sha256:9a2b6e6a8695900954f811ec89c28d4d3c77666530371094eeba7d82cd140aa6

Observation d4242c3e-b862-44ba-a27d-147635fc8e42 · outbound

This paper cites Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretraining.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Pyramidclip: Hierarchi- cal feature alignment for vision-language model pretraining

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.270340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.506039Z digest=sha256:2bab424a88c454bbc75586bdf554a7375b921898ffb82978a75ce43018b769df

Observation 47fc3fd9-cb86-4b10-ac76-f65eef6c157a · outbound

This paper cites Softclip: Softer cross-modal alignment makes clip stronger.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Softclip: Softer cross-modal alignment makes clip stronger

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.261682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.508676Z digest=sha256:e0816ab79c7528763e61ecfa2dff24d0bc79a0e58b67bf25b86a0933319a3303

Observation b3efd4d9-4696-45e9-b65c-f24e80455272 · outbound

This paper cites HiCLIP: Contrastive language-image pre- training with hierarchy-aware attention.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training HiCLIP: Contrastive language-image pre- training with hierarchy-aware attention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.253398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.511435Z digest=sha256:1b6277566bc241c5f7e5c6e235c37b174e1b4ce8a270e537f50949ecc01ceeba

Observation 6d451149-5157-482a-ace4-e7c9c3d023d9 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Bootstrap your own latent-a new approach to self-supervised learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.245173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.514078Z digest=sha256:a727c0e0c3efb1c4d89d247c89340860f5b17367b435a6c7e3d460416eb7b64c

Observation 0d814104-8839-4172-9320-a1fed898f0ba · outbound

This paper cites A survey on self-supervised learning: Algorithms, applications, and future trends.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training A survey on self-supervised learning: Algorithms, applications, and future trends

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.236823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.516730Z digest=sha256:1f1791f8fe7ab3deb9fd1cb2358c46c232b872f7c10e644784a9bd3b8fd708e4

Observation b0170a05-e30f-49b9-b94e-122ecd1a81d9 · outbound

This paper cites Masked autoencoders are scalable vision learners.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Masked autoencoders are scalable vision learners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.519556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.519556Z digest=sha256:0d6e71a6bcaae002d73da06e96d72902946af6026c53dd65a4c4e4807b021299

Observation 2335dbc5-f7fe-40a0-a4dd-cca0ff73c0aa · outbound

This paper cites Probing image- language transformers for verb understanding.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Probing image- language transformers for verb understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.223327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.522570Z digest=sha256:d3ca0e789b415894ab5de78c9c4c8625bc362e76500df20c2dd27350e03c4367

Observation 5d231367-4c5b-417e-8351-394b21ff393b · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.213999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.525264Z digest=sha256:fac0715f15adbdf3934c15c41e9f25752a119f2118f689c314a70a5896427a39

Observation ca07a299-f29a-421c-9e43-0a0971ec2ec0 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.205124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.528211Z digest=sha256:42409b46305a4df1459e107171ffebb5e60f12a7130040a700865bf638ba45a3

Observation fffa3571-1830-41bf-aa93-51728aa71ca9 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Scaling up visual and vision-language representation learning with noisy text supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.196645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.530839Z digest=sha256:d6ccb538b3dc9bbb005db244f96b2d86c8df4ed1e31d6ed1401e027bb1a37536

Observation 1aca144b-b25e-4dee-9f8a-40c48c91ee2b · outbound

This paper cites JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.188170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.533475Z digest=sha256:79e3c9fbbf682d463b2b46c70281d7927522f4d98bf26d621116dc4917739cfd

Observation 28205008-6967-4e81-aef1-7cb092c5ae7b · outbound

This paper cites Expediting contrastive language-image pretraining via self-distilled encoders.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Expediting contrastive language-image pretraining via self-distilled encoders

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.179877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.536104Z digest=sha256:d2305fe8d73ec83eaef13b17169a3c37384a8a22c03a69745dc08037f5076b72

Observation 40f8976a-0b8f-46b7-9d89-b3f3bced11f2 · outbound

This paper cites 3d object representations for fine-grained categorization.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training 3d object representations for fine-grained categorization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.538549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.538549Z digest=sha256:c7418d836f8f321519601e10179b5a1452951f6b133730cec287a75c62630b65

Observation 7b4732ce-b8e6-4329-becc-b55ac72e6be0 · outbound

This paper cites Learning multiple layers of features from tiny images.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learning multiple layers of features from tiny images

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.541200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.541200Z digest=sha256:57909771ff054fc2ae37d4adeff44c5ad046218f6ddecb2e324797d1ebdb31d6

Observation b85dcf65-9397-4d82-9549-8f34ebb734f1 · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Veclip: Improving clip training via visual-enriched captions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.162101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.543893Z digest=sha256:150c6f09888763cd6ba72b5c226a3ae4693d2bb3abf9420d35870d265b30ddc6

Observation 7e59a81b-f756-43c3-8aec-400d0b2d4ea2 · outbound

This paper cites Modeling caption diversity in contrastive vision- language pretraining.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Modeling caption diversity in contrastive vision- language pretraining

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.153526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.546561Z digest=sha256:7c966a06979aa54f2bfd93fb2e6431a528bf1e99174feb97b3d65f13f5df2d54

Observation 7fd87592-cd6e-4d5e-9886-16ee48c13452 · outbound

This paper cites Uni- clip: Unified framework for contrastive language-image pre- training.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Uni- clip: Unified framework for contrastive language-image pre- training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.549111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.549111Z digest=sha256:62288af105aa6f6209329a53277e6c29064c2cb4febf2570789296b46d444a54

Observation 6db64e3b-95be-4353-9b3d-d4b5e785db79 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.551755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.551755Z digest=sha256:9cfd0dc11aaeeb87b637f7f1ce5fa82fa3c6a5a487a0525bacce722215066bac

Observation 43df981f-28fb-4fa4-8172-f298f464b8fa · outbound

This paper cites Addressing feature suppression in unsupervised visual representations.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Addressing feature suppression in unsupervised visual representations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.134536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.554813Z digest=sha256:ee343dbe9503e1fda57480355c18c20fb5464a925131cbf977a85d2b7d4146fd

Observation cf4fba8c-a748-4ffe-9b38-93ed858ea971 · outbound

This paper cites Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Su- pervision exists everywhere: A data efficient contrastive language-image pre-training paradigm

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.126074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.557540Z digest=sha256:d4aa25dddc195d922ff8b847fa4e11b459f59674f2e466e301dd135908a1384c

Observation ee8fa206-f536-4ad2-aaa9-7756346109bf · outbound

This paper cites Evaluating object hallucination in large vision-language models.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Evaluating object hallucination in large vision-language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.117155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.560282Z digest=sha256:3b46bc656f9fd85ec5edd43e0b0c14d2e4e541aa27b25618adaa98385b00c917

Observation e977972d-a9a7-491c-aa37-5f47d14a930e · outbound

This paper cites Scaling language-image pre-training via masking.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Scaling language-image pre-training via masking

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.563041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.563041Z digest=sha256:886a95a76c0d8db0d12d915169324d74b2241e2466145cb56d2cf4099e9f145b

Observation 270ad24b-2677-43c9-ab8d-e79c5d77fab9 · outbound

This paper cites Microsoft coco: Common objects in context.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Microsoft coco: Common objects in context

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.104111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.565660Z digest=sha256:fa3b9d9dc5c250817022b617b64f7befd5f82a17501d71c931509ef351e036bc

Observation 914a2adc-52a7-4af9-9e17-3b34d0f2a532 · outbound

This paper cites Improved baselines with visual instruction tuning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Improved baselines with visual instruction tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.568397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.568397Z digest=sha256:1eb8f1b69ea9cdad1acf7ccb6821550874a856f63c9919a2f5438b2a177c9771

Observation a7ae13cc-3bc3-4405-9b80-5b2596a65bac · outbound

This paper cites Visual instruction tuning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Visual instruction tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.090060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.570985Z digest=sha256:90bd7848cc16032f66eec5bc3788cd5486e3a8160173d7dd3c5f46858113f3cf

Observation 9b2c064d-bf1e-444d-9ba4-aa33aaf72ab8 · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training MLLMs-Augmented Visual-Language Representation Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.573702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.573702Z digest=sha256:95e4ca79a871b3a8f11125ca116be1f25c5e4661634e59059095fdd912d832bd

Observation d5896950-44b3-45f2-afbe-4760779e1c28 · outbound

This paper cites Decoupled Weight Decay Regularization.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Decoupled Weight Decay Regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.576976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.576976Z digest=sha256:6fe03eabdf2ba5caf8675d04dc5fd7d24e3e9410de5d9cae862bf9d6b3b95da3

Observation 0ef9242d-acb6-4c19-bba0-a6fea2e89433 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.580026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.580026Z digest=sha256:7ede6c60e6857a3f0cf2280ed973f1996aba679b438d5764057c35816ee47259

Observation 373a3b4f-5f65-4bc8-adfc-055248b60599 · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Fine-Grained Visual Classification of Aircraft

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.582789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.582789Z digest=sha256:d536c6869ea29c82c7853171cd28a1358f367f0e413d16b675f3253a6daf8cb5

Observation c3a5c50e-a21e-4dae-a5f6-609b1c42243a · outbound

This paper cites The role of context for object detection and se- mantic segmentation in the wild.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training The role of context for object detection and se- mantic segmentation in the wild

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.075691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.585944Z digest=sha256:11940ff08088aa685a0d4b0bf1ee3c2c21dcd84807f00438500a246941e7f7c8

Observation 0f800bc9-b72f-427b-a897-2cc42c6e7b24 · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Slip: Self-supervision meets language-image pre- training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.066074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.588548Z digest=sha256:bfebe115f5283d08bc9598e74691862a75b5bc1de7daced5931507ef4c4eaeb9

Observation 2a966b05-2b8d-4cad-9d34-84187d5fb175 · outbound

This paper cites Silc: Improving vision language pretraining with self-distillation.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Silc: Improving vision language pretraining with self-distillation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.056308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.591328Z digest=sha256:87e25225c313224fefe0ff75e6f63a29fb2711369eb4f573ee894a08b9afedd6

Observation 6d879a00-adcf-4924-86b2-9d24fce288ba · outbound

This paper cites Automated flower classification over a large number of classes.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Automated flower classification over a large number of classes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.047364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.593960Z digest=sha256:da8747b5ffdedfef93dece02a733b05b8318fd8de8eaab4a59971a70b44fc047

Observation 0d779462-8cdc-4406-b167-3846f737a30e · outbound

This paper cites Docci: De- scriptions of connected and contrasting images.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Docci: De- scriptions of connected and contrasting images

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.038430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.596705Z digest=sha256:b4eb945d6246d4e7392b220ca8de51cadf532658d080622259a54e24a413726a

Observation 6fed06a7-ab62-4720-9f9a-80f41a01e56b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Representation Learning with Contrastive Predictive Coding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.599753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.599753Z digest=sha256:2be3750b23a33651279979f6bf4ffecea77b2fd18a134dcb172be3a82f6b9150

Observation dc963c2f-7d71-4e18-b421-44f0cac54772 · outbound

This paper cites Cats and dogs.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Cats and dogs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.029089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.602628Z digest=sha256:ac25db4210796d7314daa0f0b2f40390af3eda573b53d21b1f360971b8218856

Observation 56516dae-85b5-4fc2-96f8-d28eebef2ac6 · outbound

This paper cites Language models are unsu- pervised multitask learners.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Language models are unsu- pervised multitask learners

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.605473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.605473Z digest=sha256:9cddc775fdeb86849d325c6fe263befe524d3ab708aea1471d4207f2b81a87b1

Observation 4f7c7423-52ef-4fd7-b347-5671b729535b · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learn- ing transferable visual models from natural language super- vision

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.014123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.608384Z digest=sha256:38b254bd360a83e87b9328767b85e07ca21ed58db568d70d5c6e0b1c28fdb265

Observation 3e41a84c-d228-4333-9866-2c8738021ce2 · outbound

This paper cites Can contrastive learning avoid shortcut solutions? NeurIPS, 2021.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Can contrastive learning avoid shortcut solutions? NeurIPS, 2021

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:02.003410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.611260Z digest=sha256:d6fcd29c85901ffec5b37b6b8f3dae588bbe0ee794fdf154878884ec22408354

Observation 23470b4c-4ffa-4658-a4ac-1c38628389b6 · outbound

This paper cites Building vision-language models on solid foundations with masked distillation.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Building vision-language models on solid foundations with masked distillation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.994418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.613979Z digest=sha256:400f2feb07a811c04edbaee1e826e00814ed6b699fbb69fb543ed55407855c8d

Observation 5150e07e-858b-4834-82d4-a874478e2811 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.616766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.616766Z digest=sha256:ce4d83b605d1baec1f06b1800cd064c26e9d9b3af2ce8cbff7567a4f06ada7b0

Observation 9cb4a333-6b53-431b-9cea-627257c7ccfd · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.619938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.619938Z digest=sha256:3a239410b73a4dfa701e52a1e3f252b8bf78f748e58aa18ef0db2c8817b14791

Observation 0d10c965-c3a8-47ef-8fb9-e9bad0d5d146 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.622729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.622729Z digest=sha256:7e47fb5102d9d3d2e6f60f9af275103d5ce9d91fa34389dd677f2e28e86d541b

Observation fe9386d0-38c8-4069-96b9-d70be07cb39c · outbound

This paper cites Towards vqa models that can read.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Towards vqa models that can read

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.625540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.625540Z digest=sha256:fb526616dff00809849e812fbf6ee5b16ece0313c284fa6c0644896a1fa9f989

Observation c88223e4-099f-478e-afbf-79b9f07c9bf9 · outbound

This paper cites From Pixels to Prose: A Large Dataset of Dense Image Captions.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training From Pixels to Prose: A Large Dataset of Dense Image Captions

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.629788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.629788Z digest=sha256:a2395fcc99788f5b6f19bd767830f64ac2b52145dcafb22363a1c011e9cde30f

Observation a2089c6b-2dbf-432e-b1e6-afb7b05f77a8 · outbound

This paper cites Feature dropout: Revisiting the role of augmentations in contrastive learning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Feature dropout: Revisiting the role of augmentations in contrastive learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.969736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.632981Z digest=sha256:51d4ff37317fc70806ff2d27beb4af3f961723e6554a1aa597605a130f01acd7

Observation c7b6a6dc-e200-46fd-a1ff-931d19e6bebc · outbound

This paper cites Yfcc100m: The new data in multimedia research.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Yfcc100m: The new data in multimedia research

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.961726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.635784Z digest=sha256:ff82a8ad49300aeb14a92921b65c52dbc9da5e8d49e4ced23c7c3aeef04ca429

Observation 4880a0f0-34e0-4f3d-85d4-a49d81b36f41 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.953374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.638585Z digest=sha256:95cf98a0ecd7b7b993be414a3e6aef4234accccd15fbab5d60781b94d27787ed

Observation 3259ba9a-d2b6-4802-b975-1155a34081fd · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.945166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.641225Z digest=sha256:3d69238aae4bfe7e5a9e7f8dd1c3c9e98657fd850dcd7f7b700743d871548e90

Observation 0036db66-678d-4c50-9f3e-36526ca122af · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip- style models on dense captions.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training A picture is worth more than 77 text tokens: Evaluating clip- style models on dense captions

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.936681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.644041Z digest=sha256:b554aa82454995045d04a3e3411cdef5c355e151d6a20cf51bc1a6c9f0e36448

Observation 2900c910-458d-49d5-b188-0413a190181e · outbound

This paper cites Mobile- clip: Fast image-text models through multi-modal reinforced training.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Mobile- clip: Fast image-text models through multi-modal reinforced training

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.927259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.647216Z digest=sha256:4e59f81642caedbf18265b7d8c6e2af25952f534dba5e08dffc9e4234a664476

Observation 6849a673-9f1b-4338-b2d0-f7d815f05342 · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sclip: Rethinking self-attention for dense vision-language inference

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.918378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.650242Z digest=sha256:21621233cf735b6edde6d8c7b44ad67e3ae36ebbd0abfd58defcce3511c02f4c

Observation a1d627ca-69a2-4ce2-a4f2-b832652a5576 · outbound

This paper cites Lotlip: Improving language-image pre- training for long text understanding.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Lotlip: Improving language-image pre- training for long text understanding

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.909611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.653129Z digest=sha256:867ccb79ea324e5e5f18400b26bd5e3a1297e535fb7359a28a0fa53168458813

Observation ef549f9d-6e36-4b5f-848d-041213f458dd · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sun database: Large-scale scene recognition from abbey to zoo

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.900051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.656036Z digest=sha256:dc9dd3752f2b507263e4652d8f6ad9905fbf746e8347af282ca7c7f89905966d

Observation 720276a8-722b-40da-b1ec-19d16d439e14 · outbound

This paper cites Demystifying clip data.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Demystifying clip data

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.891558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.658633Z digest=sha256:d0580ad0a2260b8a32a02a685832a816bb761bf5d7d398cdd6026132b3f44d1d

Observation 84b90c2c-6d4c-426c-bd13-23a655ec38fa · outbound

This paper cites Which features are learnt by con- trastive learning? on the role of simplicity bias in class col- lapse and feature suppression.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Which features are learnt by con- trastive learning? on the role of simplicity bias in class col- lapse and feature suppression

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.882225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.662014Z digest=sha256:15bdeab72c6d7dbaba4ae262eb6dbaaf28326ca2763bdf6b72eb9c6b93edb3b0

Observation 3757f023-7e51-4f95-9c2e-f3aaf559ceff · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic cap- tion.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Alip: Adaptive language-image pre-training with synthetic cap- tion

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T00:59:01.664819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:59:01.664819Z digest=sha256:ad644ffd70a914c67b0cb3e1b048669376d626f88344df03b1e4d388b6f2dd4b

Observation 2730cf4a-de97-418a-ab08-573ba4d2a5d3 · outbound

This paper cites Filip: Fine-grained interactive language-image pre-training.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Filip: Fine-grained interactive language-image pre-training

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.868632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.667403Z digest=sha256:4f71c242bcc9805e33e65caecad64264b36383772316b0c0c565cbe70b98db5d

Observation 01343171-1b51-4250-be3f-a5fbf6dddce0 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.860321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.670158Z digest=sha256:061104dcc79cf4cc87b235abdc6015c8614a7af4c29fe27ab2e281786fed1fc0

Observation 5beb5583-4a5c-4ff2-beb3-70e86ae85bde · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Coca: Contrastive captioners are image-text foundation models

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.851879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.672871Z digest=sha256:f400bef3a88f3be617c59df5ce68b54d3ab9f0813a12e53b972c8b18a161e0f8

Observation 97cffcac-a975-49b9-9869-0fb1cd7e82b1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.842833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.675547Z digest=sha256:767c4fb581f6fbe1e6536f5f16f583114969162f4eac033eea08a5fddb59b4a0

Observation 6f3fa02b-44ab-492e-a293-65fd3b14c078 · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? ICLR, 2023.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training When and why vision- language models behave like bags-of-words, and what to do about it? ICLR, 2023

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.833392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.678286Z digest=sha256:73c751f6e6a894aea9ac6820fc7888a3efcbed685c4a42555a22f85ceb7322ef

Observation c23bf18c-567b-448c-a311-acb698f5eac9 · outbound

This paper cites Sigmoid loss for language image pre-training.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Sigmoid loss for language image pre-training

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.824877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.680908Z digest=sha256:bdea0436e3363dd84067e325018621419355350858fc9ab627e5ef00bdc99572

Observation 2c698dfb-1afa-46f2-88ff-d0bfde6255f5 · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Long-clip: Unlocking the long-text capability of clip

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.815993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.683643Z digest=sha256:78fb0de6145abc9defd6a4cc73e84bd47bbca577ee74b261ee0e23e3c5ee6446

Observation ec95625c-92c6-4024-9419-61aac0f964d3 · outbound

This paper cites Learning the unlearned: Mitigating feature suppression in contrastive learning.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Learning the unlearned: Mitigating feature suppression in contrastive learning

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.806542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.686390Z digest=sha256:552ba23b55399e3f29f5fbefd9e8eb8e6a0f8f192fb10748e880549cbc8f951d

Observation 1ffc3780-82ce-402e-84b0-b1c8857f3bc4 · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Dreamlip: Language- image pre-training with long captions

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.797925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.689027Z digest=sha256:e1d5095ea52796ec281f89319588c14561245c2e6d59127220e40284b16093e2

Observation f2f462a3-d922-4a66-bd84-0fea60865499 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Semantic under- standing of scenes through the ade20k dataset

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.789021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.692388Z digest=sha256:8bb1f78a5ff41286a3bad449cd4ba8045a6b96bbc4657ee59a71cfcf3c0434b5

Observation 0bf5983b-acfa-47f0-a1fe-d5ee5c46ed42 · outbound

This paper cites Extract free dense labels from clip.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Extract free dense labels from clip

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:59:01.779397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.695349Z digest=sha256:5073d96c987c953d09516ce36d15bf93821d39eaf1dd6b0a2be0804f372b359d

Observation a3912940-b9a1-4d7d-88a2-3b51915297de · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 94

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:59:01.770126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:59:01.698073Z digest=sha256:1e820d3fdc94524f27a945fde274e3514b700dc29db9083c348863f132a66433

Pith citing papers

Observation 9b4cfeb0-ea76-4d11-a3a5-b6f6099bc7ba · inbound

Multi-FRuGaL: Multimodal Flexible Redundancy-aware Decomposed Gated Learning for Cancer Diagnosis and Prognosis cites this paper.

Multi-FRuGaL: Multimodal Flexible Redundancy-aware Decomposed Gated Learning for Cancer Diagnosis and Prognosis COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.236693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T22:42:28.802465Z digest=sha256:c5b28e01743ab727299d112733fc6066679e15c5dcbad382da339f0cfa9869ca

Observation 246c281d-8ea9-41bc-ad28-d4d32daf4b45 · inbound

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression cites this paper.

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:05:39.135041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:05:39.135041Z digest=sha256:8ef36ed6a644a3e9f15ec5751942ebb053726599ba2cba9b42e77bdfc86b5c6e