Pith. sign in

Paper Citation Record · LEDGER

Context-Aware Multimodal Pretraining

As of 15 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 1 inbound Pith citation observation for arXiv:2411.15099.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15099 v1

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:37:57.214117Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:43.443062Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T19:24:44.625903Z

Reference resolution

100 of 113 outbound references displayed

  • verified exact1
  • verified fuzzy44
  • unresolved55
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70c765bd-bb09-4136-8981-ce3cf2be048d · outbound

This paper cites Towards in-context scene understanding.

Context-Aware Multimodal Pretraining Towards in-context scene understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.805098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.805098Z digest=sha256:6722e5e7ef85e9cb08b4de79135925213a86060766a6620c44bb4ff4a076c8a6

Observation c0e9a56c-d66c-4990-8c39-9eac67e6e79a · outbound

This paper cites Food-101 – mining discriminative components with ran- dom forests.

Context-Aware Multimodal Pretraining Food-101 – mining discriminative components with ran- dom forests

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.810558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.810558Z digest=sha256:aeedaf6e5e3df54821e0736318d0ab0ef66d65a374cc819949dd917f59de69da

Observation 77018c0c-9f9a-4993-8ebe-d7d729fea3ab · outbound

This paper cites JAX: composable transformations of Python+NumPy programs, 2018.

Context-Aware Multimodal Pretraining JAX: composable transformations of Python+NumPy programs, 2018

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.814909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.814909Z digest=sha256:852969ffbd3648d823715b8923c7f61e3cbe6b1b24be03559dcd2b23c0aab9f9

Observation 45091425-3262-43a2-9787-d4ab8e48c3f3 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Context-Aware Multimodal Pretraining Emerging properties in self-supervised vision transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.818931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.818931Z digest=sha256:1da5efa75a1b286abd8a9e406313e091aefd116e95e63e2f745b2ec1253006ba

Observation 203c6aed-dd69-44da-872c-2cc3c2938ed6 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Context-Aware Multimodal Pretraining A simple framework for contrastive learning of visual representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.823242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.823242Z digest=sha256:2af1ca3c6c50818e5c2aadc65cda399b0a059bdfaa44695d3d9553ed1d669fb5

Observation 002d7e9c-d335-436e-8bfe-d192d5a83437 · outbound

This paper cites PaLI: A jointly- scaled multilingual language-image model.

Context-Aware Multimodal Pretraining PaLI: A jointly- scaled multilingual language-image model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.827291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.827291Z digest=sha256:a87e23fb7fd674da7814409524ad9d19b93045fac85b59d1822c4d76b51b7c42

Observation 1983bc03-c6d4-4a5c-868e-40201bee3ca5 · outbound

This paper cites Meta-baseline: Exploring simple meta- learning for few-shot learning.

Context-Aware Multimodal Pretraining Meta-baseline: Exploring simple meta- learning for few-shot learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.831414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.831414Z digest=sha256:81aa8b139169c0141b6675c3e6412e63a17481cfbc4c5eea0cfabe0e80916326

Observation efd77a6e-def5-4984-8e05-27718a21572b · outbound

This paper cites Remote sens- ing image scene classification: Benchmark and state of the art.

Context-Aware Multimodal Pretraining Remote sens- ing image scene classification: Benchmark and state of the art

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.835669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.835669Z digest=sha256:48ff5e08cf2044287a4f5c862fa74483edc950f8f89af30c8d64a03bd7f1d8de

Observation 7cf14394-847a-4cfa-b7e4-d329de04a286 · outbound

This paper cites Cimpoi, S.

Context-Aware Multimodal Pretraining Cimpoi, S

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.839717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.839717Z digest=sha256:2a270f7340f300068f32d1519cd7d87213e6e8995ec86f288e7a09b9151cf3d9

Observation caa77555-f1f3-4b56-a4f5-c5308f884a12 · outbound

This paper cites Embedding arithmetic of multi- modal queries for image retrieval.

Context-Aware Multimodal Pretraining Embedding arithmetic of multi- modal queries for image retrieval

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.843837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.843837Z digest=sha256:bd00cae6a1c501cf03b9ea6d25fc55324b279b9a7cec20b0055d92dcd2bfa321

Observation 0aeefcd9-4f7b-48a5-8438-b8a75a987473 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.847839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.847839Z digest=sha256:8c93e98836a7ce9e3d9ee7ffb5fbebaf1c48e55bb9c7fce31b02515743c3bde7

Observation be71853a-f6b3-41e6-ac97-e522472987fa · outbound

This paper cites Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation.

Context-Aware Multimodal Pretraining Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.851770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.851770Z digest=sha256:773b104730d636c041fbeaf93dbe21270e73916672f4b3886db57d4eaa1e2007

Observation 48e47d30-883a-4748-a93a-f4daee470184 · outbound

This paper cites An im- age is worth 16x16 words: Transformers for image recog- nition at scale.

Context-Aware Multimodal Pretraining An im- age is worth 16x16 words: Transformers for image recog- nition at scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.856394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.856394Z digest=sha256:35c50393609a6bafc24763f42bdbb5cba2a6dd8b0e624cfd64719204c0ba4487

Observation 791e29dc-2e3a-4abd-88d3-8197b9bbf257 · outbound

This paper cites With a little help from my friends: Nearest-neighbor contrastive learning of vi- sual representations.

Context-Aware Multimodal Pretraining With a little help from my friends: Nearest-neighbor contrastive learning of vi- sual representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.860414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.860414Z digest=sha256:2a0371f9eda70f1c1619ddaf84caacf513c274613217fc0fd016ebe655d97aa8

Observation b1ab9afa-8d8e-4ce1-a16f-17109cceba86 · outbound

This paper cites Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual Understanding.

Context-Aware Multimodal Pretraining Bad Students Make Great Teachers: Active Learning Accelerates Large-Scale Visual Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.864270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.864270Z digest=sha256:c4ec6bfe7ade81088673b84f7c92d63fe580a0d2f07184b63636c8d1ce8353d9

Observation 76b9f8e3-6223-46ff-b754-0bbe8ed0982f · outbound

This paper cites Data curation via joint example selection further accelerates multimodal learning.

Context-Aware Multimodal Pretraining Data curation via joint example selection further accelerates multimodal learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.868240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.868240Z digest=sha256:b7b312358a13f5566020e60e1794be312f998a6f8b8b5eb6e00ea8b713976c2f

Observation 1eebf32f-e751-4382-9dac-0216922ce997 · outbound

This paper cites Data determines distributional robustness in contrastive lan- guage image pre-training (clip).

Context-Aware Multimodal Pretraining Data determines distributional robustness in contrastive lan- guage image pre-training (clip)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.872532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.872532Z digest=sha256:d3396c9a051fa5c22e97c310e1ddb1ed563e3ab836c161f95b5435bcd8602a77

Observation 3f82e4a1-ddc8-48d1-a892-5ea3aecd3922 · outbound

This paper cites Caption supervision enables robust learners.

Context-Aware Multimodal Pretraining Caption supervision enables robust learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.876708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.876708Z digest=sha256:ad6edfc33a614f49b8db052fbbd9616fc2f58d4862f17ad185bdfb4a05d2dd1b

Observation 3344ca7d-cb0b-4469-958f-9d0f9eb12bf9 · outbound

This paper cites Context-aware meta-learning.

Context-Aware Multimodal Pretraining Context-aware meta-learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.881328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.881328Z digest=sha256:7082ed438a77c50982ef16c88a6fc90836024fab4bdfe9b768b2277126634753

Observation 3ecec771-1e71-44f9-9788-8068b4d7989e · outbound

This paper cites Model- agnostic meta-learning for fast adaptation of deep networks.

Context-Aware Multimodal Pretraining Model- agnostic meta-learning for fast adaptation of deep networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.885032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.885032Z digest=sha256:f60770d0aad093274644763b58e68e5943e2f32fe979585d5f04dc8d4a7bbea0

Observation 4ce5867a-c163-46ac-90e3-90ae46b03a87 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Context-Aware Multimodal Pretraining Clip-adapter: Better vision-language models with feature adapters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.889441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.889441Z digest=sha256:df9425832ceb26a6658cc859bdf0285162259dce270acad3196db92f40ab85e6

Observation ea69ba92-37d5-4ba2-ae99-f7e68baa6f06 · outbound

This paper cites Towards flexible perception with visual memory.

Context-Aware Multimodal Pretraining Towards flexible perception with visual memory

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.893039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.893039Z digest=sha256:b6e064af63c3b69346c791ce4ffb58da3a4b210e46b39758dd176f09fdb1a59b

Observation c1af1f92-3eac-4a0b-91d9-ba0b27715a8d · outbound

This paper cites Cyclip: Cyclic con- trastive language-image pretraining.

Context-Aware Multimodal Pretraining Cyclip: Cyclic con- trastive language-image pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.896970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.896970Z digest=sha256:a651f8b5f0f2931cbdeef89c184dd6ae775fdd2d861346528062ea6e37239a23

Observation 9b34ef7e-5e95-4083-9bcb-6d08c7968367 · outbound

This paper cites kNN-CLIP: Retrieval enables training-free segmenta- tion on continually expanding large vocabularies.

Context-Aware Multimodal Pretraining kNN-CLIP: Retrieval enables training-free segmenta- tion on continually expanding large vocabularies

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.900622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.900622Z digest=sha256:858e81279977e4cad78f4d4889033f17b11a1af16470e43929127272ad3bf5d9

Observation 72718a6f-1246-4ebd-941a-9664a307b1c3 · outbound

This paper cites Calip: zero-shot enhancement of clip with parameter-free attention.

Context-Aware Multimodal Pretraining Calip: zero-shot enhancement of clip with parameter-free attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.904118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.904118Z digest=sha256:54791cc116ecd9554039b0ae3661e7fbc07e0cb57186c1befe3b7b39aedca932

Observation c3d15297-de6e-461b-895b-ae9edcc6fb1e · outbound

This paper cites Anchor-based robust finetuning of vision-language models.

Context-Aware Multimodal Pretraining Anchor-based robust finetuning of vision-language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.907731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.907731Z digest=sha256:0a0a2f33ff95301bbc38cd0198580252fee2643006e3182f4acd79d495370e54

Observation 73352938-bd53-4430-858a-da6a8b60f17d · outbound

This paper cites Dota: Dis- tributional test-time adaptation of vision-language models.

Context-Aware Multimodal Pretraining Dota: Dis- tributional test-time adaptation of vision-language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.911882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.911882Z digest=sha256:43f1aedc47865b92ad63a204fd1e78690a81130dfa2d3b225f52e2c6702e7da1

Observation 31521a41-7721-487b-b3dc-687f24777925 · outbound

This paper cites Momentum Contrast for Unsupervised Visual Representation Learning.

Context-Aware Multimodal Pretraining Momentum Contrast for Unsupervised Visual Representation Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.915601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.915601Z digest=sha256:8711b2ced4e1aae514c597c2de2a30b3062cbb75ca4cb97c8e2b457a199957a7

Observation 04fa7485-abd6-4db2-843f-8b06ca16ce79 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.

Context-Aware Multimodal Pretraining Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.919978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.919978Z digest=sha256:ce7c0a8a59a71ec7304b8f71d4c68bd13a085de1df87732fb26c57f9664fbc3a

Observation 0f11c2db-bca9-4a8d-a902-59f97392e053 · outbound

This paper cites Ross, and Alireza Fathi.

Context-Aware Multimodal Pretraining Ross, and Alireza Fathi

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.923750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.923750Z digest=sha256:c84b083c4c47090ddc68579100f32b2ddea98082145a65baa32c6a5bd1347e71

Observation 712c3c31-e308-433f-b43e-98f09a666d95 · outbound

This paper cites An open access repository of images on plant health to enable the development of mobile disease diagnostics.

Context-Aware Multimodal Pretraining An open access repository of images on plant health to enable the development of mobile disease diagnostics

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.927566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.927566Z digest=sha256:7e2ee14d799194526ffad7510a62efced6c4e30d6d10b7654c41b54674eb1cb6

Observation 65be5d2f-e5e3-4871-8b7c-c4c697f0b0cc · outbound

This paper cites Retrieval-enhanced contrastive vision-text mod- els.

Context-Aware Multimodal Pretraining Retrieval-enhanced contrastive vision-text mod- els

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.932485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.932485Z digest=sha256:2cf34003ae07451de352de30dfcd3d6ad7b4c187eb3614e252868ca52fb12aa7

Observation 630f52a3-f4d3-4bba-bc2c-aabb5916cdb2 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Context-Aware Multimodal Pretraining Scaling up visual and vision-language representation learning with noisy text supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.936218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.936218Z digest=sha256:cd16c52b1eba3095014dfc867e5551d4d14af32d35f40c0a669448f40985f917

Observation 301c0320-f0e9-4ad3-927e-e8c3da1ea784 · outbound

This paper cites Billion- scale similarity search with gpus.

Context-Aware Multimodal Pretraining Billion- scale similarity search with gpus

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.939932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.939932Z digest=sha256:6f183a8922d6370c8311ad4393fa35d59faa32b2abb05a6dd640427ea8babeef

Observation 4e7093eb-6566-4511-bab5-8903fb954295 · outbound

This paper cites Multi-class texture analysis in colorectal cancer histology.

Context-Aware Multimodal Pretraining Multi-class texture analysis in colorectal cancer histology

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.943995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.943995Z digest=sha256:b78bda3e99886cb9568c76365e97a2484e3a283b332a27dccd4158ca6cb30c5a

Observation a12c991e-ddef-4b57-b744-3b4c493c9bdd · outbound

This paper cites Maple: Multi-modal prompt learning.

Context-Aware Multimodal Pretraining Maple: Multi-modal prompt learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.948744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.948744Z digest=sha256:0ba763ddfc21642e124656f9a791bd4ceaf230666d7ee9803090f90b5d731615

Observation 3f0178d9-5fb9-4878-9322-7773ebf4b246 · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting.

Context-Aware Multimodal Pretraining Self-regulating prompts: Foundational model adaptation without forgetting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.953882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.953882Z digest=sha256:add370a32c088dbb70f987fc426f38b5df8aec6d53cd948f6e9981f3ec8a4f45

Observation 40bbfb34-0c1b-4bb5-9f79-d237c8071b8b · outbound

This paper cites Datadream: Few-shot guided dataset generation.

Context-Aware Multimodal Pretraining Datadream: Few-shot guided dataset generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.958180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.958180Z digest=sha256:2a97b02b1b3334a2a052e0fb4f594fad0889a216db534f1586d9ca615663333c

Observation 8083119b-cd98-48a2-8670-f63045aa784d · outbound

This paper cites Kirchhof, K.

Context-Aware Multimodal Pretraining Kirchhof, K

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.961933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.961933Z digest=sha256:8939283394af10213482c0686891d2a98f073f4c8f4b194b1f37e9e4c32b76cd

Observation 21f16633-77be-402c-b7ee-a9f2837d006b · outbound

This paper cites Wilds: A benchmark of in-the-wild distri- bution shifts.

Context-Aware Multimodal Pretraining Wilds: A benchmark of in-the-wild distri- bution shifts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.965448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.965448Z digest=sha256:851cb573f634e5b6f552466b5c2d427f0eec0c354a97d2d2d696d180c89108cb

Observation 52bea425-8108-4dfb-bd98-0a052fa4d840 · outbound

This paper cites 3d object representations for fine-grained categorization.

Context-Aware Multimodal Pretraining 3d object representations for fine-grained categorization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:56.969771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:56.969771Z digest=sha256:7bf8c10d93e08e31dc337e9aec7653c2a1d71665827057fa4b6a53414dc052a2

Observation a56237f0-2a9a-4d2b-b657-7fa37b0660a4 · outbound

This paper cites Learning multiple layers of features from tiny images.

Context-Aware Multimodal Pretraining Learning multiple layers of features from tiny images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.346155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:56.973726Z digest=sha256:da49490cb99b5dabe63dd11db77a21f09b09bc518e7f4ad9fd4e62cb18b03326

Observation b070a9d6-f5af-4f61-91b8-64f339a3df6e · outbound

This paper cites SentencePiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing.

Context-Aware Multimodal Pretraining SentencePiece: A sim- ple and language independent subword tokenizer and deto- kenizer for neural text processing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.335232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:56.977722Z digest=sha256:4ae0c2482009ef672ead55aeab8ce7934cd069ba8f8de37e1e7647714de8d545

Observation 84fac99f-09e7-4b1c-99ad-abc348ab5315 · outbound

This paper cites Meta-learning with differentiable con- vex optimization.

Context-Aware Multimodal Pretraining Meta-learning with differentiable con- vex optimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.323606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:56.981678Z digest=sha256:d702a2ae8edd04bb888b738e1abb0814bf7b46719cf8cd7e44903bc3a3e36afc

Observation 05d17675-4868-4313-be29-a586dc535122 · outbound

This paper cites Universal representation learning from multiple domains for few- shot classification.

Context-Aware Multimodal Pretraining Universal representation learning from multiple domains for few- shot classification

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.310950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:56.986973Z digest=sha256:ae818e9230771802c5e65eb8419e7f46c8afc28eb8c53ecd8c57559a62e11934

Observation cc5f9c01-79bb-4ce3-b473-f5a0519ef951 · outbound

This paper cites The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning.

Context-Aware Multimodal Pretraining The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:37:57.462781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:56.991378Z digest=sha256:8733d3d01767fbdc58daa123e0d4e50706825073998ff87b62c64601618412e8

Observation 100e03d1-66a3-4846-938d-ae726ac4617d · outbound

This paper cites Decoupled weight de- cay regularization.

Context-Aware Multimodal Pretraining Decoupled weight de- cay regularization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.299402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:56.995575Z digest=sha256:606ff00d47dd4a9cf13c410b319cf693ba6f74c940c779c01b5704a7230aff1f

Observation 64fc52eb-12c5-477a-8321-317b216e4e17 · outbound

This paper cites A closer look at few-shot classification again.

Context-Aware Multimodal Pretraining A closer look at few-shot classification again

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.287880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:56.999335Z digest=sha256:a0b678827d80fec619f77180f4403a74af2cc5b8f75674aad5d1947e8461cd7a

Observation 66be19f4-ae75-41fa-a267-c0469a025420 · outbound

This paper cites Efficient and ro- bust approximate nearest neighbor search using hierarchi- cal navigable small world graphs.

Context-Aware Multimodal Pretraining Efficient and ro- bust approximate nearest neighbor search using hierarchi- cal navigable small world graphs

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.274989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.004285Z digest=sha256:19bd1b1fc51d168d3e306a9c08f0d86156091ce900c7914e0760e3f1cae63ef7

Observation 4c316190-2606-4644-8e32-d02c12dc136c · outbound

This paper cites Visual classification via description from large language models.

Context-Aware Multimodal Pretraining Visual classification via description from large language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.262648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.008287Z digest=sha256:406046e8584ed896eafd7e3c8314cfd1196ceb9aa12841c5ed0e55935c24366a

Observation 1c6ae813-cfd4-497c-915c-53081821f268 · outbound

This paper cites Understanding retrieval- augmented task adaptation for vision-language models.

Context-Aware Multimodal Pretraining Understanding retrieval- augmented task adaptation for vision-language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.249862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.012415Z digest=sha256:a52a097aa8c0d7ca23c4367a4a1729602846e759e479663bb32a1308ae2c7bbc

Observation 64f891d0-f16e-4700-9fda-89deececb47a · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

Context-Aware Multimodal Pretraining Slip: Self-supervision meets language-image pre- training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.238857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.016068Z digest=sha256:289e2b319b5b7d5448162fc78d68ad38a3baa9c84b3a3d54272715eddcf15977

Observation 34267b4b-0cd9-4537-88cb-0d25cf3faaee · outbound

This paper cites iCassava 2019 Fine-Grained Visual Categorization Challenge.

Context-Aware Multimodal Pretraining iCassava 2019 Fine-Grained Visual Categorization Challenge

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.019995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.019995Z digest=sha256:1146a362fb0186c2c1384b5eff94a5e242161a729cd92c3f59a04ae67c319444

Observation 6a6f778a-d042-4d47-b99e-0728aea139c9 · outbound

This paper cites Revisiting knn- based image classification system with high-capacity stor- age.

Context-Aware Multimodal Pretraining Revisiting knn- based image classification system with high-capacity stor- age

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.227426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.023854Z digest=sha256:28a52fa03ce8f4ff754839de6e62d55be6ec31da35a64d815ff2bfa5f02ceae5

Observation a4ba33e8-925d-4752-8028-e3c9221f3ba8 · outbound

This paper cites On First-Order Meta-Learning Algorithms.

Context-Aware Multimodal Pretraining On First-Order Meta-Learning Algorithms

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.031237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.031237Z digest=sha256:103162ebebab5843154c5f03b5771ec59d634f43a80ffcace0b2190a2f7baafc

Observation a516c8b1-56ff-49db-ac39-ed9e9667bbbc · outbound

This paper cites CHiLS: Zero-shot image classifica- tion with hierarchical label sets.

Context-Aware Multimodal Pretraining CHiLS: Zero-shot image classifica- tion with hierarchical label sets

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.202661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.035558Z digest=sha256:95d686d2d607e80665089addf3d39904542b6898134ad59714f22c4841857ecc

Observation ffc3d409-feb7-4a51-9ff2-9f8763f5a71b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Context-Aware Multimodal Pretraining Representation Learning with Contrastive Predictive Coding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.039478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.039478Z digest=sha256:f4979a54c8580a90e8be1941ba38657f8d66fbd93b7fbcd3f1c1de0672cc6c0a

Observation 442b0a8e-c71a-4a6a-91e7-640c56d7afa1 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.191074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.043705Z digest=sha256:c2b3b7a0f302633fb8c5f45ca3d5e24a8ddb58cc887c927b2e9e49150f92ec96

Observation b27edba5-7226-4014-bcc7-df36401432a1 · outbound

This paper cites Svl-adapter: Self-supervised adapter for vision-language pretrained models.

Context-Aware Multimodal Pretraining Svl-adapter: Self-supervised adapter for vision-language pretrained models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.179609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.047926Z digest=sha256:db9f2d2e068051deee095191fe9d35fd610bdad669bb134ba65ba67dc175759c

Observation aad77fb2-7711-4292-a136-f79e59bc68a7 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.167209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.051654Z digest=sha256:f19d8967250d49443db72fc60835af46c504533c250fadeb65715b4a517dd80f

Observation fb1fd241-37be-4acc-b86f-2b62a2cce57a · outbound

This paper cites Moment matching for multi-source domain adaptation.

Context-Aware Multimodal Pretraining Moment matching for multi-source domain adaptation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.154247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.056471Z digest=sha256:a50ccac68692bd12c815c36a75a1b49853be4b79b15fd039b3fb4318ed36bf7e

Observation 711b7eb8-5a14-40ff-825d-1fee0f1c1dbc · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.142966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.060344Z digest=sha256:aedc460134a74cea421d855dab3c3d784cc6a2592352faaf50386a355f5bee7d

Observation a5a555e5-f8dc-4cf2-9f8a-6f8744cc603f · outbound

This paper cites Online Continual Learning Without the Storage Constraint.

Context-Aware Multimodal Pretraining Online Continual Learning Without the Storage Constraint

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.064283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.064283Z digest=sha256:265503dcd7676c050d9c52f6087a69d9d7f785897911ade73e0aae3386d29cb2

Observation e006fed8-7986-4efa-a2df-edc5dbae5f72 · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classification.

Context-Aware Multimodal Pretraining What does a platypus look like? generating customized prompts for zero-shot image classification

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.131420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.068671Z digest=sha256:a7af2457d1ea374cdf3a99f538283e91178ad4639f19871e9bdfe1bd986ab2ef

Observation 9a24563c-aeb9-4cdb-9f12-218f0817d7ea · outbound

This paper cites Learning transferable visual models from natural language supervision.

Context-Aware Multimodal Pretraining Learning transferable visual models from natural language supervision

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.119743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.072389Z digest=sha256:28f52490800132f2a6ecbcf5088d7c2e1d650e1ecff2bcd88c959dcf293b9755

Observation 17366306-0fec-4a1c-bb20-4dd743293929 · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:58.108238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.076389Z digest=sha256:d8c9ec77e1981901a93df4b769aa4a393795831461a728becdeb8b9d59feded3

Observation 11a6dafd-f568-480b-8ea3-4b1e87e4bfd6 · outbound

This paper cites Meta-learning with implicit gradients.

Context-Aware Multimodal Pretraining Meta-learning with implicit gradients

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.097235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.080772Z digest=sha256:2bd04b1e0090f0262dfab0efd2ff6ec03e067bd3bd5f502ed03aefa776577870

Observation 5a3eec50-6345-4d87-9c7a-e59d352db629 · outbound

This paper cites Towards to- tal recall in industrial anomaly detection.

Context-Aware Multimodal Pretraining Towards to- tal recall in industrial anomaly detection

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.086217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.084489Z digest=sha256:74a51bf74978f073bc6f849f859792058e78b3e012c70b7075200c8fda16b79e

Observation f5547801-b058-45c5-b9aa-41767054ec72 · outbound

This paper cites Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata.

Context-Aware Multimodal Pretraining Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.075211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.088022Z digest=sha256:e1372534435bc312e45557bcd35eaffb9b08170ed7c68e19f9c9f3f3b3afeffd

Observation e2ce42b8-6507-4191-bd56-8497b31970a0 · outbound

This paper cites A Practitioner's Guide to Continual Multimodal Pretraining.

Context-Aware Multimodal Pretraining A Practitioner's Guide to Continual Multimodal Pretraining

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.092746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.092746Z digest=sha256:797580a25f5b91bffcdd6800e9bd3fe812c03c214acdef37da8190f41b05f64a

Observation 03000239-5940-4e84-a6ab-09e07e705872 · outbound

This paper cites Berg, and Li Fei-Fei.

Context-Aware Multimodal Pretraining Berg, and Li Fei-Fei

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.063806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.097341Z digest=sha256:cdcc72be6e967f67f33be028548d3c6aba3bfdf7f8769117a1834a296dde48b2

Observation d6ecb8b8-8445-4d45-9567-e9ac8637df2e · outbound

This paper cites Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Had- sell.

Context-Aware Multimodal Pretraining Rusu, Dushyant Rao, Jakub Sygnowski, Oriol Vinyals, Razvan Pascanu, Simon Osindero, and Raia Had- sell

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.052211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.100964Z digest=sha256:0ca2859214a25ecabf65830312aed01338eb1976b7b495691f8f01e6ce8285e0

Observation f2d45028-895b-4b9e-93e3-9e2990279037 · outbound

This paper cites Is a caption worth a thou- sand images? a study on representation learning.

Context-Aware Multimodal Pretraining Is a caption worth a thou- sand images? a study on representation learning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.041055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.104716Z digest=sha256:ef12da8688e64fae9aad763074d63fda3c0de81e17004b02ad44ed63b276e6b0

Observation d93cf45e-2090-402e-95ef-4a92bd597ee6 · outbound

This paper cites Scott, Andrew C.

Context-Aware Multimodal Pretraining Scott, Andrew C

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.030097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.109119Z digest=sha256:43ded2e78b2d5ee801d24fb103bfdc36994c3a503006ccb0695a33068df383ae

Observation c9b90e25-f56c-4b06-9818-850fc62b1c62 · outbound

This paper cites statsmodels: Econo- metric and statistical modeling with python.

Context-Aware Multimodal Pretraining statsmodels: Econo- metric and statistical modeling with python

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:58.012305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.113383Z digest=sha256:18e6c0a8689fcf44d8fc59f59502adcc91afa572ba1aa48cc24913f0e013fbc0

Observation 81001eac-2be7-4ef4-a53c-7430a15c25a1 · outbound

This paper cites Prototyp- ical networks for few-shot learning.

Context-Aware Multimodal Pretraining Prototyp- ical networks for few-shot learning

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.999678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.117264Z digest=sha256:cb560e853c0182c8853ed4d604eaad3901c4ec06ca867f9cd7629e52b0dc327e

Observation b477796c-5448-4111-9aaf-d90044714688 · outbound

This paper cites CLIP models are few-shot learners: Empirical stud- ies on VQA and visual entailment.

Context-Aware Multimodal Pretraining CLIP models are few-shot learners: Empirical stud- ies on VQA and visual entailment

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.976059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.124616Z digest=sha256:60cd2ea16306053aa4c2fd1b53e195a25c62a9658883cd0a7333d5c959f52623

Observation 736c5b70-6251-4dd3-aa01-ad1ea67a5694 · outbound

This paper cites Momentum-based Weight Interpolation of Strong Zero-Shot Models for Continual Learning.

Context-Aware Multimodal Pretraining Momentum-based Weight Interpolation of Strong Zero-Shot Models for Continual Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.128358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.128358Z digest=sha256:4df8d51599e7f012a3500116573013506edaf2de53413ed2f73af27dd0a49d79

Observation 007a8fe0-30bc-4ac9-9eae-401bb59f09b5 · outbound

This paper cites Torr, and Timothy M.

Context-Aware Multimodal Pretraining Torr, and Timothy M

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.964324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.132369Z digest=sha256:07c40928525dcf30fde886a79e35614d4e1e3d72b40566028aa1ba11d6436117

Observation 95c8b417-a95c-409d-8def-04ac52074166 · outbound

This paper cites A Fistful of Words: Learning Transferable Visual Models from Bag-of-Words Supervision.

Context-Aware Multimodal Pretraining A Fistful of Words: Learning Transferable Visual Models from Bag-of-Words Supervision

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.136457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.136457Z digest=sha256:ab27d4c825badf08bb32669b3de8a0dd4c366ff051880a1ac505009364447484

Observation dc99c10a-e7c6-494d-817a-f680068153a7 · outbound

This paper cites Reflecting on the state of rehearsal-free continual learning with pretrained models.

Context-Aware Multimodal Pretraining Reflecting on the state of rehearsal-free continual learning with pretrained models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.140969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.140969Z digest=sha256:12301160abbf2dce656259b7960994058870c7cdf2d5e294ff2ce7139a42590a

Observation 2f283015-664d-4491-be18-b0dd2ff33416 · outbound

This paper cites Tenenbaum, and Phillip Isola.

Context-Aware Multimodal Pretraining Tenenbaum, and Phillip Isola

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.952387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.145029Z digest=sha256:010607ac8d03b5e60bb52836d81bd632fea6587d0f9543e168760b5c678271fc

Observation 723165a3-c8b8-48bf-90e5-3e7bd85928a8 · outbound

This paper cites Learning a universal template for few-shot dataset generalization.

Context-Aware Multimodal Pretraining Learning a universal template for few-shot dataset generalization

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.939953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.149162Z digest=sha256:958e21f5066388f24cabab15ededbbd9d1ad1b9e507965e5b9f26707aa4b7240

Observation 3ab390aa-eb50-4dae-a438-3083e747ee7e · outbound

This paper cites Sus-x: Training-free name-only transfer of vision-language models.

Context-Aware Multimodal Pretraining Sus-x: Training-free name-only transfer of vision-language models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.926951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.152895Z digest=sha256:c2e579c37287fdcdf281ca17234ff9dcbb887db7845b98e8f1d8d73c99050aa0

Observation dea92e67-8c29-41a8-877e-579c86f2bb22 · outbound

This paper cites No ”zero-shot” without exponential data: Pretraining concept frequency determines multimodal model performance.

Context-Aware Multimodal Pretraining No ”zero-shot” without exponential data: Pretraining concept frequency determines multimodal model performance

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.914761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.156538Z digest=sha256:d90e1a92baf73a6a3a2dfb07cb59e36fd53dc50437c893c92b1f2768aa14d215

Observation 82c60407-9136-4ee8-bcb4-085c2e53757a · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Context-Aware Multimodal Pretraining Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.903118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.160216Z digest=sha256:5938a3aa800ae7b347228ab32de3e5d9e49e5151b29c2f6e240beb917cbdbc0d

Observation 9d0d702b-5736-4283-8759-998f518144f6 · outbound

This paper cites Matching networks for one shot learning.

Context-Aware Multimodal Pretraining Matching networks for one shot learning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.891598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.164036Z digest=sha256:d33e7f49417d04104f130a87d30b58f54d9e2eb4bb9e60c0258404b4c4a440a3

Observation 79ea8769-f9a9-4ad6-932c-235425f2753b · outbound

This paper cites Learning robust global representations by penalizing local predictive power.

Context-Aware Multimodal Pretraining Learning robust global representations by penalizing local predictive power

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.879394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.167806Z digest=sha256:8c9410d5837225def46ffb91b87ff5492fb41e84062d5fb249e9e8cd27eed85a

Observation 58d52a3f-8c90-4eb9-ba33-21098263e4b3 · outbound

This paper cites SimpleShot: Revisiting Nearest-Neighbor Classification for Few-Shot Learning.

Context-Aware Multimodal Pretraining SimpleShot: Revisiting Nearest-Neighbor Classification for Few-Shot Learning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:57.172045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:57.172045Z digest=sha256:c14fbb1960457dec103911aad3a4a403fc1257e526c654e3400ee0eda88ac514

Observation 51f8e675-7350-412a-94a5-0b679013241e · outbound

This paper cites A hard-to-beat baseline for training- free CLIP-based adaptation.

Context-Aware Multimodal Pretraining A hard-to-beat baseline for training- free CLIP-based adaptation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.868102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.175826Z digest=sha256:7176371cd1c8c0cc2f733d840cb8482229d1751a395bd2080e1ca9aa3baf7354

Observation 181e1fb5-0c3b-47f4-b543-a17c2dd03b85 · outbound

This paper cites Welinder, S.

Context-Aware Multimodal Pretraining Welinder, S

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.857105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.179462Z digest=sha256:d3ac55b8d0a798b38438b216b8204a9af2cfa2287a331e340bdcd5023b884c47

Observation c99e5433-c297-4b5d-ba9b-e6f829b23660 · outbound

This paper cites Cascade prompt learning for vision-language model adaptation.

Context-Aware Multimodal Pretraining Cascade prompt learning for vision-language model adaptation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.846173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.183149Z digest=sha256:a43855b86016a79dfe74850f678dbf565de8207e403cf6f6207525805f11d5c2

Observation 2ef29238-f14e-4282-954b-6baa244b04ed · outbound

This paper cites Yu, and Dahua Lin.

Context-Aware Multimodal Pretraining Yu, and Dahua Lin

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.835015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.186746Z digest=sha256:cb438ab4e79befcdfe8e256eb72f392f7a8c17dc2b6670b21c739a4d8e41cf4d

Observation deb49dfa-4408-4827-811a-446eb5d7debb · outbound

This paper cites an unresolved cited work.

Context-Aware Multimodal Pretraining Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:37:57.824058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.190332Z digest=sha256:3d77874050196d9ffcf898da80f0e5406720baa753f1c01349186a2465d7133c

Observation 575cea0f-38ef-425c-a187-745270fb36a0 · outbound

This paper cites Ra-clip: Retrieval augmented contrastive language-image pre-training.

Context-Aware Multimodal Pretraining Ra-clip: Retrieval augmented contrastive language-image pre-training

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.812976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.193803Z digest=sha256:882cb621ddd4f4692328402e0d4a5c20ea31ef04dd4346ee1e37f24203b1c9f2

Observation e67708da-7c0d-4613-a188-06e4cdf8e229 · outbound

This paper cites MetaFun: Meta-learning with itera- tive functional updates.

Context-Aware Multimodal Pretraining MetaFun: Meta-learning with itera- tive functional updates

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.801712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.197717Z digest=sha256:bb6d2d001aaf1ce1726859bef41d677eca1aee273cf45fdf4e376cb70b6dfcfb

Observation 6127758c-8ad4-44f4-9215-885eb7169f91 · outbound

This paper cites Bag-of-visual-words and spatial extensions for land-use classification.

Context-Aware Multimodal Pretraining Bag-of-visual-words and spatial extensions for land-use classification

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.790291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.201604Z digest=sha256:c49b66e8d8f4b891ecbe6ec560e5d59639895095ec699e9a55c877f7f3adfe77

Observation 1b7a84d9-5cdd-41d1-9862-cf12abb957c2 · outbound

This paper cites TapNet: Neural network augmented with task-adaptive projection for few-shot learning.

Context-Aware Multimodal Pretraining TapNet: Neural network augmented with task-adaptive projection for few-shot learning

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.777850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.205556Z digest=sha256:e7570566686a631940a470675c12299a03bb5caa8ef5ba6889f469ebba712db9

Observation df4d8c5c-6d8b-41b0-a767-f9c6521dc099 · outbound

This paper cites Sigmoid loss for language image pre-training.

Context-Aware Multimodal Pretraining Sigmoid loss for language image pre-training

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.765024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.210229Z digest=sha256:27eb0068c7f3b0c661eb3c3d919014aacaafdfa940c9e8701a498af52a85c66b

Observation e6204512-50f0-4d9c-b946-a3f74601fde8 · outbound

This paper cites Deepemd: Few-shot image classification with differen- tiable earth mover’s distance and structured classifiers.

Context-Aware Multimodal Pretraining Deepemd: Few-shot image classification with differen- tiable earth mover’s distance and structured classifiers

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:37:57.753232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:37:57.214117Z digest=sha256:4cdeebb385c176695a23317a105493767f6d2b8fea6be384413743beb88b40da

Pith citing papers

Observation b9b8c456-4792-410d-b998-f07c992c7b0e · inbound

How to Merge Your Multimodal Models Over Time? cites this paper.

How to Merge Your Multimodal Models Over Time? Context-Aware Multimodal Pretraining

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:24:44.634180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T19:24:43.443062Z digest=sha256:9139baef29bc4bd73d6ae6b3f5811c57ca5f515c1b0878d1a07cc28e7f00f34f