Pith. sign in

Paper Citation Record · LEDGER

Image Recognition with Vision and Language Embeddings of VLMs

As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2509.09311.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09311 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:24:21.549089Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9883c8b3-ce2f-4d7d-ba90-b22152ec868c · outbound

This paper cites Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design.

Image Recognition with Vision and Language Embeddings of VLMs Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.431803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.431803Z digest=sha256:cb8a0a4f40f5bd44b677b94f73ea539a279fa2dc8f4d798bab60771f2622737d

Observation 79b454bc-263c-4864-bfb1-dc73f3bc623d · outbound

This paper cites Are we done with ImageNet?.

Image Recognition with Vision and Language Embeddings of VLMs Are we done with ImageNet?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.435965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.435965Z digest=sha256:ac451da6a3c9482d7496bbe123b2b645855825a718d33c5656ed8fc9da7fc951

Observation a130b777-9e34-4895-93cc-42df65db3df9 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Image Recognition with Vision and Language Embeddings of VLMs Imagenet: A large-scale hierarchical image database

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.439635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.439635Z digest=sha256:dc44ec9235d255f6af0c6eda6e1bf4b59077f6c4d37167545e284967d4a345c8

Observation a158f1a2-b443-42da-a6ab-6d59617aef43 · outbound

This paper cites Eva-02: A visual representation for neon genesis.

Image Recognition with Vision and Language Embeddings of VLMs Eva-02: A visual representation for neon genesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.443012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.443012Z digest=sha256:bb5170ee5ee9a0203c90f1e0e2dc61694f7a29d4486d0bc5e79f9eb248cc7e62

Observation f82c2987-ceb4-48f8-b189-b0354759ffde · outbound

This paper cites Renovating names in open-vocabulary segmentation benchmarks.

Image Recognition with Vision and Language Embeddings of VLMs Renovating names in open-vocabulary segmentation benchmarks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.446516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.446516Z digest=sha256:ed9b557216226bf0edad7c4c790a651ce0a7a861c4f51aa41b84dfdb06f844fe

Observation fd80e3a1-1bf5-4e67-acb5-6911634527ee · outbound

This paper cites RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models.

Image Recognition with Vision and Language Embeddings of VLMs RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.450187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.450187Z digest=sha256:3b778a6205382707627ca3417aedf9a4ad67cd2ee27bffe59a2d90eb2a2ccf60

Observation 61a2475a-a115-4991-afc0-1744b5160b9a · outbound

This paper cites Openclip, July 2021.

Image Recognition with Vision and Language Embeddings of VLMs Openclip, July 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.454305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.454305Z digest=sha256:fe1dd78dbbbb3e03bd117dbdfa05441a03a06297a537d26ad61144213dccb348

Observation 03b350dc-7bd6-48b1-b7df-0cdd057b1163 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

Image Recognition with Vision and Language Embeddings of VLMs Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.457434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.457434Z digest=sha256:634d4e2eda0a096667b38795834d2258cb10943b703d7b94e47289848a68a607

Observation f61acd8f-110f-4bc2-83f9-bf00489db786 · outbound

This paper cites Flaws of imagenet, computer vision’s favorite dataset.

Image Recognition with Vision and Language Embeddings of VLMs Flaws of imagenet, computer vision’s favorite dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.462404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.462404Z digest=sha256:b338dce0fc628997ba319d3e177bbb5f55c201e15d56823d50cb40a447a70933

Observation 8ddf585f-62c6-434f-822c-ccb3333993be · outbound

This paper cites Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models.

Image Recognition with Vision and Language Embeddings of VLMs Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.466174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.466174Z digest=sha256:06f6552dabd2008f99dae13b75094f681fafaafb6b393b22a0d6f481151b4541

Observation 0070f8cf-d175-4caa-934d-1c5f9598f731 · outbound

This paper cites Visual instruction tuning,.

Image Recognition with Vision and Language Embeddings of VLMs Visual instruction tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.469313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.469313Z digest=sha256:4e71b593bb8a2db76ff00461e6893f02abb34da3cbecf16b10394b58c6e5dadd

Observation a1c85b76-e1e6-4e7f-b84f-3a570739fec6 · outbound

This paper cites Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks.

Image Recognition with Vision and Language Embeddings of VLMs Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.475663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.475663Z digest=sha256:49c8dc68e7e6c08c935f0fb85006ef428a8b29ed2ef3353fb19a7a8e3e177747

Observation a15fc33b-2c4b-441a-abc2-ff9225f78a20 · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, 2022.

Image Recognition with Vision and Language Embeddings of VLMs Chatgpt: Optimizing language models for dialogue, 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.479479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.479479Z digest=sha256:f838ddd6e831300cae22fdf3b993eeba50c3bd4189a4179eca21727a4aee43d6

Observation f0fcf758-83ae-4563-850a-aaa7e5619bee · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Image Recognition with Vision and Language Embeddings of VLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.482265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.482265Z digest=sha256:92f2c6aa5c23ead1bb89d138ddfc032608356db615ca2ef992ee45d373471dac

Observation 540d8450-2e48-4db0-9e36-8ad336aebf26 · outbound

This paper cites Learning to name classes for vision and language models, 2023.

Image Recognition with Vision and Language Embeddings of VLMs Learning to name classes for vision and language models, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.485402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.485402Z digest=sha256:1971c0bd7879dabe45d9317c029303a490b57c54f0c36f2f8881dad81e1f565c

Observation 878ad9ef-f2ab-48e1-8663-6300a301619d · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Image Recognition with Vision and Language Embeddings of VLMs Learning Transferable Visual Models From Natural Language Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.488592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.488592Z digest=sha256:99f125168f74603c92db970596a9005a1c6e6c26eba781c2086ed5229af8c613

Observation 7d8e7bdd-e8a1-4ef9-a9a0-4f0ddc0315b9 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge.

Image Recognition with Vision and Language Embeddings of VLMs ImageNet Large Scale Visual Recognition Challenge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.491601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.491601Z digest=sha256:33676ba4e8467c3a0ed5c222ab54e36cd662683592189bf5c6c5f94abfd01e0d

Observation cadf1b4e-088b-4dcc-a538-116ebe94679b · outbound

This paper cites Evaluating machine accuracy on ImageNet.

Image Recognition with Vision and Language Embeddings of VLMs Evaluating machine accuracy on ImageNet

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.494917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.494917Z digest=sha256:76d5659a982ba7d60d18311ea4e4b8263cc82507c1560faf2503e9155bb0ddd6

Observation f4258173-5baa-4ad9-9d49-88d1b5c7885f · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

Image Recognition with Vision and Language Embeddings of VLMs PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.501065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.501065Z digest=sha256:c66c5255a7e0ec8ba4909d1744b9ee486379dafc563b8d8f503da7766e6e0958

Observation 0e3e8dc7-5b40-4a0c-8a6a-fb6ce18526d0 · outbound

This paper cites an unresolved cited work.

Image Recognition with Vision and Language Embeddings of VLMs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.504476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.504476Z digest=sha256:730e105104ab4513421e5733fe3c2ffb105cd6163b28a89f3793f6080c451be1

Observation cb45ebac-444f-4468-9565-8b2156c2d286 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Image Recognition with Vision and Language Embeddings of VLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.510846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.510846Z digest=sha256:472a43a345d71a4d3ac6e5859b67707cf22b03e9835d07791ea16a3bba96f459

Observation 4966a22b-7840-4dbc-ab81-3d99222383a4 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Image Recognition with Vision and Language Embeddings of VLMs SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.514128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.514128Z digest=sha256:14c9917acda9fc579dbc91220377e28afaa22b5d1750aad905e2a3ad188f3c9d

Observation 6ccb2a77-e206-49d0-898d-6cd2b0d04462 · outbound

This paper cites From ImageNet to Image Classification: Contextualizing Progress on Benchmarks.

Image Recognition with Vision and Language Embeddings of VLMs From ImageNet to Image Classification: Contextualizing Progress on Benchmarks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.517237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.517237Z digest=sha256:631b5736bcf40ea492bfaf8d7f45178169727038e6ad863586b802283a1f4888

Observation 98fc8d27-3340-4fa4-b60e-9b4b440f9092 · outbound

This paper cites When does dough become a bagel? Analyzing the remaining mistakes on ImageNet.

Image Recognition with Vision and Language Embeddings of VLMs When does dough become a bagel? Analyzing the remaining mistakes on ImageNet

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.520369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.520369Z digest=sha256:93c415f33e4a28f5b5c193402466286e24b1eaea832d174c0f446c61ff5559e1

Observation 5c6f9b95-61e4-42ee-ba3e-65f622c2b323 · outbound

This paper cites ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet Accuracy.

Image Recognition with Vision and Language Embeddings of VLMs ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet Accuracy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.523599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.523599Z digest=sha256:3cda8f095e6c1c08b74c7f3eb4c302e49f55ab2d46ff4f3d5e44e6f0d4d20b7a

Observation e414068e-e0e8-4937-abf9-1501886f0e84 · outbound

This paper cites Pytorch image models.

Image Recognition with Vision and Language Embeddings of VLMs Pytorch image models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.526595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.526595Z digest=sha256:21fd619ef56ea71ecfb089673b692ff8dfd412b3c11b377f7f5a48c6d616e7b8

Observation 11799de6-8e27-453a-92f4-d13f1bc23230 · outbound

This paper cites Self-training with Noisy Student improves ImageNet classification.

Image Recognition with Vision and Language Embeddings of VLMs Self-training with Noisy Student improves ImageNet classification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.529702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.529702Z digest=sha256:d469193b6c332affb2580adc04deea3e199e2c78b9d720f1620b989a611d276a

Observation 9e96df1d-d9e6-4e7f-9663-d8f563d65d6d · outbound

This paper cites Leveraging cross-modal neigh- bor representation for improved clip classification.

Image Recognition with Vision and Language Embeddings of VLMs Leveraging cross-modal neigh- bor representation for improved clip classification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.533090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.533090Z digest=sha256:40d4977f14bd2e0fe522436d60adf65c6c9085dc0c135bc0ec5d8d66c4920cd0

Observation 11bd0161-0444-46c4-a228-89f9118114e2 · outbound

This paper cites Sigmoid loss for language image pre-training, 2023.

Image Recognition with Vision and Language Embeddings of VLMs Sigmoid loss for language image pre-training, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.535869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.535869Z digest=sha256:ed113475659455e0d96cafc6516fc216602a89e951507b9bece50b97f01a6df2

Observation b6fc6248-da4c-4349-a485-ec227cf8cc84 · outbound

This paper cites Vision-language models for vision tasks: A survey.

Image Recognition with Vision and Language Embeddings of VLMs Vision-language models for vision tasks: A survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.538976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.538976Z digest=sha256:576b3f11504640e5bd326c2a26b238b1d3f0af06284f91bf5f356b145a1f3c1e

Observation a6e85d05-2b34-4bf3-95a5-50a12d6426be · outbound

This paper cites Tip-adapter: Training-free adaption of clip for few-shot classification.

Image Recognition with Vision and Language Embeddings of VLMs Tip-adapter: Training-free adaption of clip for few-shot classification

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.542732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.542732Z digest=sha256:3b13b825424ccaaae597c29dccd09d1c9cb8729d6d6a14826391603e55fc11a8

Observation 0b041f9d-4f58-4ac6-bf92-e96c1a93d2fe · outbound

This paper cites Conditional Prompt Learning for Vision-Language Models.

Image Recognition with Vision and Language Embeddings of VLMs Conditional Prompt Learning for Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.545774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.545774Z digest=sha256:da8ab4f6c59870a69ead3ae75aa3cd377c4f3db37feb95e672646aa4c617d641

Observation 9172bc4a-4b87-486d-8d06-370a2c1e901a · outbound

This paper cites {class name}.

Image Recognition with Vision and Language Embeddings of VLMs {class name}

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.549089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.549089Z digest=sha256:a30743b07cf2feaf6e1c33f3226a70a3b7ce5df17e4272001f1ef79ade0b23f1

Observation 1e4496f4-70f1-4cd2-ab1c-2ca998b2b319 · outbound

This paper cites EfficientNetV2: Smaller Models and Faster Training.

Image Recognition with Vision and Language Embeddings of VLMs EfficientNetV2: Smaller Models and Faster Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.507572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.507572Z digest=sha256:8d7246c212b365c5119d15a5963cb033bcdd2998dba99cdd890c5a84f76cf918

Observation 6c372c09-a47f-49c3-aefa-25cedcfdfc32 · outbound

This paper cites Visual Instruction Tuning.

Image Recognition with Vision and Language Embeddings of VLMs Visual Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.472351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.472351Z digest=sha256:9aa1bd0d7ffce3fc4bc61e4b27b19b2e357ff17a8dc7f8d29ba6db44d37343d5

Observation c0db88a8-b69d-4bea-9f53-723bc2abef3f · outbound

This paper cites URL https://proceedings.mlr.press/ v119/shankar20c.html.

Image Recognition with Vision and Language Embeddings of VLMs URL https://proceedings.mlr.press/ v119/shankar20c.html

Reference 8644

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:21.498040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:21.498040Z digest=sha256:1430cdbd10e3c454c6c0ef6389c88dc33737369ca62900cb20616ca42beddcea

Pith citing papers

No inbound Pith citation observations are available.