Pith. sign in

Paper Citation Record · LEDGER

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks

As of 21 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2412.20682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20682 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:19:02.251395Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:32:20.408599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:32:21.103085Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6080f1f-968c-4258-abc4-92ad33c29ca7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Learning transferable visual models from natural language supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.853203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.062335Z digest=sha256:f20f696e95fdf69e9a7d015e7accc7a229e261b661760ea1c903e44a63a7a95f

Observation 24a9c0f8-f2fe-4376-a37e-59c0a3822ab3 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.844441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.067085Z digest=sha256:8cc25d983ce752ea224ed683c4e6efebab8c2e57d42496fc897d4253d8ebc63f

Observation 1496bf16-a508-4a7d-9e0c-b9a0d8b127c3 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Sigmoid loss for language image pre-training,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.836193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.070476Z digest=sha256:00e1213137e636761674edcaa5b88ea84b20b0ccdc21cd85f27192970b0accc4

Observation 1d7dc89c-8759-4ea7-bd9f-7f6e10e24a54 · outbound

This paper cites Sgva-clip: Semantic- guided visual adapting of vision-language models for few-shot image classification,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Sgva-clip: Semantic- guided visual adapting of vision-language models for few-shot image classification,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.827397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.074512Z digest=sha256:7e09b19585ff77192a4a60944bb5e7fb59c5e362beac295e0ae25078eec3dcce

Observation f125e2cf-d345-4ff8-8bf6-1d65bf841a07 · outbound

This paper cites Clip-vg: Self-paced curriculum adapting of clip for visual grounding,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Clip-vg: Self-paced curriculum adapting of clip for visual grounding,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.818810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.078024Z digest=sha256:d12dcf905fdf705e9dfe0c9ed9d3e1a49de836d8e5772f49aabef12939e785f1

Observation 4f0ff8ac-de93-4d10-bae8-4f677414e50c · outbound

This paper cites Effective end-to-end vision language pre- training with semantic visual loss,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Effective end-to-end vision language pre- training with semantic visual loss,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.809768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.081342Z digest=sha256:7876b4b64d86d3f8c40901c3a2995b9acc74eb47fab636a3ccc210eba5bbeaf3

Observation 39e2d1f9-77ae-45d4-8827-e40e13faa139 · outbound

This paper cites Neural logic vision language explainer,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Neural logic vision language explainer,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.799422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.085792Z digest=sha256:08086fca1158e6545995f337f65af7b87167933be5c36b5f8ef1f8be8538631d

Observation 45f4ac16-7638-4079-96fe-a911d77bd1f4 · outbound

This paper cites Lovm: Language- only vision model selection,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Lovm: Language- only vision model selection,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.788305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.088657Z digest=sha256:dd9896da231a27d8224fdb0dc90080b1e44e930a9e5817dcb80f111129cb73e2

Observation ed98bd14-cf8f-4e55-a8c6-7b2f33da4b48 · outbound

This paper cites Bridge the Modality and Capability Gaps in Vision-Language Model Selection.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Bridge the Modality and Capability Gaps in Vision-Language Model Selection

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.092268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.092268Z digest=sha256:63c9ffb83b23a6accc0e17c7135a74727952c42332bedef6133e70f989d217e7

Observation a9663b70-939f-42e3-b691-b47a749e5d47 · outbound

This paper cites Imagenet large scale visual recognition challenge,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Imagenet large scale visual recognition challenge,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.096576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.096576Z digest=sha256:7854b0530c8966489deb48836bb4812bc2ba9a962876d081655dee1a207d8469

Observation 59c80df0-e92e-4a65-bf70-fb67a9d8e0d8 · outbound

This paper cites GPT-4 Technical Report.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks GPT-4 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.099108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.099108Z digest=sha256:742a816e98146e2aecd34949bd111d370f4deb03d008ab530ee742a9a54bdaf1

Observation a1291799-f858-46f5-b114-950b4b26af82 · outbound

This paper cites Leveraging unlabeled data to predict out-of-distribution performance,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Leveraging unlabeled data to predict out-of-distribution performance,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.776644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.102861Z digest=sha256:cd231099861510739cc9c618062c4e9bf0b8eb687591d3a5b66ca63976109691

Observation a13c2413-be53-44db-b299-acae43f47066 · outbound

This paper cites Are labels always necessary for classifier accuracy evaluation?.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Are labels always necessary for classifier accuracy evaluation?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.769117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.106942Z digest=sha256:f01efcb0991d9360200092036ad76e28003c440d0b76d444bbefaa486b8e6314

Observation 327ab1b6-d597-49f8-b18e-7b59a2ddeee1 · outbound

This paper cites Predicting out-of- distribution error with the projection norm,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Predicting out-of- distribution error with the projection norm,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.760868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.109667Z digest=sha256:09139dcddec4aa875d6e29631523676af9f13caa29e74cea89304533a4d4f600

Observation c3384c91-811e-458a-8cee-ce8ae80128a6 · outbound

This paper cites Data determines distributional robustness in contrastive language image pre-training (clip),.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Data determines distributional robustness in contrastive language image pre-training (clip),

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.752385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.112361Z digest=sha256:fb7e92ecbebf6179e20c0958f3c72e2598b84b4108476d5b26ce309ccaae760b

Observation 18584f6a-cff5-4044-b263-60498db4077c · outbound

This paper cites Does clip’s generalization performance mainly stem from high train- test similarity?.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Does clip’s generalization performance mainly stem from high train- test similarity?

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.744442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.115218Z digest=sha256:b2ad4a23d4c6241b1ecc36e41970093ca7f720555c91bfb8a1f81acf7b384d78

Observation 3bde1082-2188-4a4f-a509-69890d4088ec · outbound

This paper cites A Survey on Evaluation of Out-of-Distribution Generalization.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks A Survey on Evaluation of Out-of-Distribution Generalization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.117741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.117741Z digest=sha256:d03c01db0c73d2b299810a8281ed1f39202e7dd2620216a695efd8f23f3dab8e

Observation 23247191-a828-4b72-9d61-913dc3761483 · outbound

This paper cites Which Model to Transfer? A Survey on Transferability Estimation.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Which Model to Transfer? A Survey on Transferability Estimation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.120616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.120616Z digest=sha256:dc2c4c2f9a3237c5b88fb91ed4074b0e69660d0079813f9f301842be697284d3

Observation 67e21d47-42b9-47c7-841a-68ede3983428 · outbound

This paper cites Rankme: Assessing the downstream performance of pretrained self-supervised representations by their rank,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Rankme: Assessing the downstream performance of pretrained self-supervised representations by their rank,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.736740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.123697Z digest=sha256:7b8d4148f2f189df71bbf316ab2ed2a6e6bb1bc347b221b0af16b8b9c252a9fc

Observation 1a5b9b43-afd2-4f53-83e0-2523ef975313 · outbound

This paper cites Identifying useful learnwares for heterogeneous label spaces,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Identifying useful learnwares for heterogeneous label spaces,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.727669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.126635Z digest=sha256:cc147988ca9109daa90812e93d06cfa8bbf2b0698a85d39ffc6cdc739598b46f

Observation bac8c173-40be-4e03-930c-158b51cd1e34 · outbound

This paper cites Etran: Energy-based transferability estimation,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Etran: Energy-based transferability estimation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.720403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.129072Z digest=sha256:704633f3a43ff65582a305b67e301cdc85affba14fbb868de0aa5559be66d2f1

Observation 27d233fe-7d78-4b5d-a545-499a6c8ed392 · outbound

This paper cites Predicting out-of-distribution error with confidence optimal transport,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Predicting out-of-distribution error with confidence optimal transport,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.713537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.132874Z digest=sha256:bffa1626323d1dde6c0522f59f9539dccd60e414770a1c041957acdcccf6202f

Observation e3096a9c-fbdd-48cf-9bb4-7fa431d6cd33 · outbound

This paper cites Data analysis and regression. a second course in statistics,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Data analysis and regression. a second course in statistics,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.706480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.136115Z digest=sha256:126badc564e851e24a87526c9b18c4a7c093cba9e797ba888aa236b2009d9422

Observation 07e1f4b7-5bf3-4985-9975-d8648163fea9 · outbound

This paper cites Tune it the right way: Unsupervised validation of domain adaptation via soft neighborhood density,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Tune it the right way: Unsupervised validation of domain adaptation via soft neighborhood density,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.699122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.138546Z digest=sha256:9e7e3c5f71aa61751070b32bea0457344273ae5dd0dd1235c30d03dc5f5d30f8

Observation 50eff583-92e6-4574-bb0a-40ebbff8ca86 · outbound

This paper cites Covariate shift adap- tation by importance weighted cross validation.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Covariate shift adap- tation by importance weighted cross validation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.689740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.142516Z digest=sha256:0572eccf3ce0fb28713dd70f590379c64d3498abf25eae26fbb243d50bd436fc

Observation 591eb3d8-46f2-42c8-9f10-df2da74637e7 · outbound

This paper cites Towards accurate model selection in deep unsupervised domain adaptation,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Towards accurate model selection in deep unsupervised domain adaptation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.680320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.146091Z digest=sha256:a5d400d5ea7d2e6509e381d534d7f42a129f7b86e929e78740ea598ba290eee8

Observation 7b35aab9-69dc-4412-a0f5-3156d5e7e5c2 · outbound

This paper cites Stochastic gradient methods for dis- tributionally robust optimization with f-divergences,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Stochastic gradient methods for dis- tributionally robust optimization with f-divergences,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.672181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.148593Z digest=sha256:27be5bc12e6eb4e3c5407220fba1596e4062151e6dc6ee715bea997b168264bc

Observation add8cb75-5990-468f-ae31-0288f3206bb9 · outbound

This paper cites Invariant Risk Minimization.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Invariant Risk Minimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.151663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.151663Z digest=sha256:1ce7eee097ac92a2b315c9e52e94e441333e4b3c5907ee7334fef4a55c16e61e

Observation 4c547453-dc72-408a-92b8-62b718f6e057 · outbound

This paper cites Stable learning via sample reweighting,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Stable learning via sample reweighting,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.664007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.155701Z digest=sha256:52359bee7c5463db5d95b6d6d3bef806433b78548e779f15f4472ba669b79d20

Observation a61de187-c254-4af9-8795-bcda9ebcb386 · outbound

This paper cites A baseline for detecting misclassified and out-of-distribution examples in neural networks,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks A baseline for detecting misclassified and out-of-distribution examples in neural networks,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.655597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.159008Z digest=sha256:d872d06357ae8410b1fb8e912906bd8aab4e648898d96fdc04245e8d40e8a4ce

Observation b578e508-0b05-421a-8fdb-39e58819caf8 · outbound

This paper cites What does rotation prediction tell us about classifier accuracy under varying testing environments?.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks What does rotation prediction tell us about classifier accuracy under varying testing environments?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.647830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.161587Z digest=sha256:112378bda98721e4e951db75e29c91eb79760a88d28992e52ad175e7e54cc366

Observation 42141123-3547-4e66-940d-9e732916013f · outbound

This paper cites Agreement-on-the- line: Predicting the performance of neural networks under distribution shift,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Agreement-on-the- line: Predicting the performance of neural networks under distribution shift,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.637822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.164953Z digest=sha256:d31c234b256c4ad861b39602f30ed13263a3a5fa0bd17f8c00ad141586ed2cd6

Observation 69d6d889-6245-47ff-84ed-d09684bf3972 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.630112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.168153Z digest=sha256:530bc8178f1c39c9d4420f71e68c1458e8fcfab0fd9b90d430c480e914972345

Observation d0258e10-b13b-47bd-aa46-36d5a80fa574 · outbound

This paper cites I. mathematical contributions to the theory of evolu- tion.—vii. on the correlation of characters not quantitatively measur- able,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks I. mathematical contributions to the theory of evolu- tion.—vii. on the correlation of characters not quantitatively measur- able,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.621254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.171155Z digest=sha256:949b8bc031d060e55ff4828ae9c6829324535c39d53f53ce9fb90ff552d8380a

Observation 3a956975-6cb9-4cde-a6d6-06c8787a739b · outbound

This paper cites Learning multiple layers of features from tiny images,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Learning multiple layers of features from tiny images,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.173527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.173527Z digest=sha256:f7782ab5d1575426b2b77f080365150e6c527cc10707cb16dc97512b4819e6e5

Observation 7101e1b9-9972-4dd6-9fd8-8c9d9db1b391 · outbound

This paper cites Cats and dogs,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Cats and dogs,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.610138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.176231Z digest=sha256:e02e74c866b717f9ab70029991924ffe95455c6e3ec20b463c38548faf69b77e

Observation 2c700529-10b4-4c1e-8269-d91c1adf3467 · outbound

This paper cites Automated flower classification over a large number of classes,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Automated flower classification over a large number of classes,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.603200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.178912Z digest=sha256:21c0cbb3d6aff3a10065d518d5d81b44d7982acd5cea0f4568ebef28a0b2196b

Observation ba33f10b-0d1f-43da-aa42-52ceab64c3e8 · outbound

This paper cites Reading digits in natural images with unsupervised feature learning,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Reading digits in natural images with unsupervised feature learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.595212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.181346Z digest=sha256:4916936ece0019ef637c27e3a8e3156f0430cbe32d0da6edfc392d466d1d9f79

Observation fc7e523e-c7f9-45ea-ac45-d9a27606758f · outbound

This paper cites Detection of traffic signs in real-world images: The german traffic sign detection benchmark,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Detection of traffic signs in real-world images: The german traffic sign detection benchmark,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.585413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.185130Z digest=sha256:b5d61c09324f8ef5a5e29de4a4ff790dbfb12e642b3a2d503b1dd52286ba27a2

Observation 7ca0654a-f47a-4828-81f5-6b5bfc88c71a · outbound

This paper cites Describing textures in the wild,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Describing textures in the wild,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.576201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.189062Z digest=sha256:88d23256a16f383004105ef01c438e0581b31ebd18b0365c1b26528ca0689a39

Observation e08b47f2-ef10-484a-b07f-6fc1400b7c45 · outbound

This paper cites Yfcc100m: The new data in multimedia research,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Yfcc100m: The new data in multimedia research,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.566206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.192584Z digest=sha256:e022f2628454cbfed87a107e53164b1681a519f7361c4fdd1a19f8c1c70235b1

Observation 0c4880aa-d471-4885-a240-66142a8a7251 · outbound

This paper cites Sun database: Large-scale scene recognition from abbey to zoo,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Sun database: Large-scale scene recognition from abbey to zoo,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.554864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.198741Z digest=sha256:c273bdc997be2471719d1a0f17d8f4a3235098ce333da875de29127e56568e68

Observation 2ad95db0-05d0-42b8-b66c-0ea03d373dca · outbound

This paper cites Gradient-based learning applied to document recognition,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Gradient-based learning applied to document recognition,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.544234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.201847Z digest=sha256:a78484db8f8cf6794f1cec0d6c4b15ef14aefa0fb8a42746e822c38f5abba886

Observation a387ba64-2dae-4eb9-a1ef-9d21e78f5e04 · outbound

This paper cites Challenges in representation learning: Facial expression recognition challenge,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Challenges in representation learning: Facial expression recognition challenge,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.535101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.205503Z digest=sha256:b2d11a6bfa224812113429a7c764785b60ae101e2943d5f9206baf88213f8be1

Observation 45258e42-4956-4e3a-aa1a-87627e74055d · outbound

This paper cites On the importance of feature separability in predicting out-of-distribution error,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks On the importance of feature separability in predicting out-of-distribution error,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.527167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.209061Z digest=sha256:69cdf68dad476b46f667f15967e1ccc91b82f5a818f1145106836a87f1e6e06d

Observation 315a85bc-873d-463f-97c0-b37bf5325664 · outbound

This paper cites Unsupervised representation learning by predicting image rotations,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Unsupervised representation learning by predicting image rotations,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.517108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.212042Z digest=sha256:e3722b9d45d315f48e891095bac271cc7fa54ed68d88787a93d31e5caa639264

Observation a88f3d46-786e-4e23-acdb-27219d5b6f06 · outbound

This paper cites The use of multiple measurements in taxonomic prob- lems,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks The use of multiple measurements in taxonomic prob- lems,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.508560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.214634Z digest=sha256:e842a47e506adb3ed68d68f7e05ddd6edff1d61c63e3f59974b2e646769bb135

Observation 218e96b0-d58b-4d22-9b7a-d5f41067cfa0 · outbound

This paper cites Silhouettes: a graphical aid to the interpretation and validation of cluster analysis,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Silhouettes: a graphical aid to the interpretation and validation of cluster analysis,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.217025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.217025Z digest=sha256:3ea670d2840debdc249019482286af305e07236d58d711fbc329ed194c008ff9

Observation c39edd1e-2beb-494c-b54b-d9f797f5ae56 · outbound

This paper cites Deep residual learning for image recognition,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Deep residual learning for image recognition,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.219537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.219537Z digest=sha256:4bd20930178de1f0e3e1ced8a14af3da7ba7a9c7b7942035ad55723424b08298

Observation 7995dc37-9bf7-441f-b9fc-c302ac82e98e · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.422639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.222327Z digest=sha256:ea3f9959f10336de33c836cac9422214845195c7316b8e82f5fb57426e22144b

Observation 119a2684-2c0c-439d-9f05-922bad6cc5cc · outbound

This paper cites A convnet for the 2020s,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks A convnet for the 2020s,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.382172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.225312Z digest=sha256:e6f85fb220c0706f061b2d4ebee770ae9329db642524b94b2ab4d740bf9c3b2e

Observation 675996f6-57fd-4a99-ae23-b26f0b48b90a · outbound

This paper cites Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Laion-400m: Open dataset of clip-filtered 400 million image-text pairs,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.372675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.227883Z digest=sha256:a370091536c3ef1a4d7b657638b3a7c22e38e22bfa43e65776dd9540a33be2fe

Observation 82cc4b91-9e7e-4ad3-8f3c-c26e6c5d8102 · outbound

This paper cites AltCLIP: Altering the language encoder in CLIP for extended language capabilities,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks AltCLIP: Altering the language encoder in CLIP for extended language capabilities,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.362356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.230309Z digest=sha256:2a73fc01c1685399c9181ab74644818186447409997a8e37956b4f57f3d2dfaa

Observation d7237b80-b802-4761-859c-448fcacc02af · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Groupvit: Semantic segmentation emerges from text supervision,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.350740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.232910Z digest=sha256:7434ecb755c1317ef0c5e43d91bcf7dd0b19920232057ae02b705e8005ce984b

Observation a3b48329-45f1-4ab2-856a-bec9f04e4f0e · outbound

This paper cites Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.235754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.235754Z digest=sha256:b85a4515bbb0e3459895352d8d1cc135008ad804dfcf695e51742359520ade94

Observation b63e8fcc-ea48-4330-8d0f-23b2ae3a5fef · outbound

This paper cites Demystifying CLIP Data.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Demystifying CLIP Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.239160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.239160Z digest=sha256:8be417d7b50b4d9d1f370e34d8b10d489659ed8e304d3dc3ab00b5e17720ed25

Observation 8829e10d-c0a3-4c87-a026-2ede33fedede · outbound

This paper cites BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.242350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.242350Z digest=sha256:b9e96390dce32ba2b399708e7f7afa61ed967bb76ec6c8c674681c85215920cf

Observation 60cb4d52-1c1d-4b93-8581-ff0c846425f1 · outbound

This paper cites Quilt-1M: One Million Image-Text Pairs for Histopathology.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Quilt-1M: One Million Image-Text Pairs for Histopathology

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:02.245892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:02.245892Z digest=sha256:a309434b0eeaeab319df02b03a195cf005b5c866c7b87c4fb2ddbd5b90468a5f

Observation 8e2650a0-8233-468f-8cc3-09fc29526389 · outbound

This paper cites BioCLIP: A vision foundation model for the tree of life,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks BioCLIP: A vision foundation model for the tree of life,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.341422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.248744Z digest=sha256:8a6968cf6f8e4842c67cc000ab57f351b62c8d8316167cb37774b876f819df0d

Observation cb36a3e1-a99f-4dcc-8527-2511724e5102 · outbound

This paper cites Gpt-4: Generative pre-trained transformer,.

Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks Gpt-4: Generative pre-trained transformer,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:02.331857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T23:19:02.251395Z digest=sha256:ca5f8e76826b602c317817d19c344cf58066a5600a2e702925bab7ef215feb2e

Pith citing papers

Observation 12f4859a-834e-4817-a2cf-bfb3385c5b4e · inbound

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection cites this paper.

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:32:21.110763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T00:32:20.408599Z digest=sha256:b1016c546eb86400220fb11c6fdec4e9b4660af69c058e67c51526dd90a48d47