Pith. sign in

Paper Citation Record · LEDGER

Robust image classification with multi-modal large language models

As of 17 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2412.10353.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10353 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:00:13.103101Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c8f0814-7449-4a9f-af37-ed9a0d5a9160 · outbound

This paper cites Deep neural rejection against adver- sarial examples,.

Robust image classification with multi-modal large language models Deep neural rejection against adver- sarial examples,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.594571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:13.013540Z digest=sha256:e8da3138e256de92ebe8cccbbf678bc3fc5ef2cf83e46843c72f236b54ff6f49

Observation 84bb4307-1660-44c8-8125-eb2f712e35f7 · outbound

This paper cites Image-text Retrieval: A Survey on Recent Research and Development.

Robust image classification with multi-modal large language models Image-text Retrieval: A Survey on Recent Research and Development

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.040800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.040800Z digest=sha256:bcf51cdae443ad51e33d24c0b73f4a1a2d20c106a5b4cb362c07c75e3b914ffe

Observation 60c69d0f-88ce-4201-a803-ab13abd4585b · outbound

This paper cites VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks.

Robust image classification with multi-modal large language models VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.052533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.052533Z digest=sha256:2cfe2aecbe8b3be778c4fe2195522801f70fd2fb0880906878a80eb90a6d5b11

Observation f0282e34-d82e-4341-a3d1-f10c78907782 · outbound

This paper cites On Evaluating Adversarial Robustness.

Robust image classification with multi-modal large language models On Evaluating Adversarial Robustness

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.059287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.059287Z digest=sha256:eec6fd369408b1a6d66706f2f271a6a056fc6e3963195563dc619cf90fe08ab7

Observation e3767c66-8306-4766-8642-51167b2a8c16 · outbound

This paper cites An analysis of single-layer networks in unsupervised feature learning,.

Robust image classification with multi-modal large language models An analysis of single-layer networks in unsupervised feature learning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.515652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:13.070621Z digest=sha256:6cb272004a82aef251df9e2b652afb7e1f3c3bb34ee1c48a583c2e8d1ffb6633

Observation 86033ae1-2dec-4428-966e-2f14dc0b4370 · outbound

This paper cites Improving robustness using generated data,.

Robust image classification with multi-modal large language models Improving robustness using generated data,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.487170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:13.075645Z digest=sha256:c8fc18ceed5f09f26b8d815e2d1b656e39463b80342f0ba9e0b6374e2cfa7390

Observation a80416b1-e043-4b94-a4d2-fdf6bd256249 · outbound

This paper cites Adversarial robustness: From self-supervised pre-training to fine-tuning,.

Robust image classification with multi-modal large language models Adversarial robustness: From self-supervised pre-training to fine-tuning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.465157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:13.082583Z digest=sha256:79ddd2c9cd553c2dca3b97be70042647249b0cd98f9449a5567f59582fda80ec

Observation dc5aa641-55d7-47b6-9234-ac34f1b736b7 · outbound

This paper cites Fusionbench: A comprehensive benchmark of deep model fusion,.

Robust image classification with multi-modal large language models Fusionbench: A comprehensive benchmark of deep model fusion,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.088859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.088859Z digest=sha256:6a27a729468f56589e20d2711ef1035b39994237bbd1ccfb7450edd2f0db4a18

Observation 2a39b60f-6d87-4c68-baa6-66f024898cab · outbound

This paper cites CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy.

Robust image classification with multi-modal large language models CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \$10,000 Budget; An Extra \$4,000 Unlocks 81.8% Accuracy

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.094737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.094737Z digest=sha256:8f9268d1d6ad9a0dff9eb544cd0d9b2dfc7a4adbf6831ec40e916a045dfcb7af

Observation dac85a6e-423f-41a1-9f33-3bb62ac9547f · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Robust image classification with multi-modal large language models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.103101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.103101Z digest=sha256:349ca7da3d3d15c4b30f93fd8894ce82af5d9dcf6525cbe027a7a3ddb54b9385

Observation 28b0a0ed-11a8-435c-964a-785d74169182 · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,.

Robust image classification with multi-modal large language models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.544798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:13.065711Z digest=sha256:9e44fb8da4315f8b543e9ab6d15a36c81930ab1235ef6d6645abddd6ed5de000

Observation 460ee490-2aec-4398-95b3-c508bebe017f · outbound

This paper cites Towards evaluating the robustness of neural networks,.

Robust image classification with multi-modal large language models Towards evaluating the robustness of neural networks,

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.675015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:12.984402Z digest=sha256:6d1866c84091ed10e0a76b9772c9d04869b96a20540aab2cc8ab75cf6bc79ea8

Observation 775bfcc7-9d62-4945-ab92-753b48169323 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Robust image classification with multi-modal large language models Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.002316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.002316Z digest=sha256:fce51e41d4345790059656976cfc0af9129fec129cbb684eeb518733adb2f14d

Observation aacb442f-a626-47a5-8a97-f93dd6079843 · outbound

This paper cites Safetynet: Detecting and rejecting adversarial examples robustly,.

Robust image classification with multi-modal large language models Safetynet: Detecting and rejecting adversarial examples robustly,

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.613574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:13.008087Z digest=sha256:2fc519e3436b8b3f28ce9cc93615556b3e25e8c36b3b15806078e303b3384f1b

Observation 970b03d7-ed8a-40c4-9025-18827242f426 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Robust image classification with multi-modal large language models A Survey on Multimodal Large Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.029837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.029837Z digest=sha256:edca4545c1c27dd199613d792b218236e401bfbc789c4f78a7484851e81471fb

Observation 8f44f540-b9c8-4e4d-968c-17fa362054f7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Robust image classification with multi-modal large language models Learning transferable visual models from natural language supervision,

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.566416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:13.019822Z digest=sha256:75c90df5b1bc3aaaf308049836ea305213f2b9984c2a20200fea1e8f46f43c1d

Observation 9ae533e3-f133-4521-a039-162509643c0a · outbound

This paper cites Evasion attacks against 8 machine learning at test time,.

Robust image classification with multi-modal large language models Evasion attacks against 8 machine learning at test time,

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.845651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:12.977010Z digest=sha256:cc486f70a1310c97ad43a5d3fd4f711d4a21c5ae61155bcec2751b64480f3097

Observation 6e033151-2dd2-4d03-ae19-fb02cfd3de9e · outbound

This paper cites Wild patterns: Ten years after the rise of adversarial machine learning,.

Robust image classification with multi-modal large language models Wild patterns: Ten years after the rise of adversarial machine learning,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.633314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:12.997383Z digest=sha256:fa46ee4714c2785c44384c31516a3d829554baf4bc9f004c5f6c8d37456f8586

Observation 4f7405e4-7fd0-4d69-acee-cdab0767910a · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Robust image classification with multi-modal large language models VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.046276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.046276Z digest=sha256:c8d243b33bd0f9a7580e693c69a9cc971ad91e0779f9929f0522dac8b6b3defb

Observation 1edc973a-b1b2-42cc-89e4-663c6096cc65 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

Robust image classification with multi-modal large language models A Survey of Vision-Language Pre-Trained Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.035351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.035351Z digest=sha256:837cd1a4e74dd9402ccd0ad048df46d91fe2b452621ea6b74d3e0314f096d1c5

Observation 93b7f950-d061-4b8f-a711-164aab124dc5 · outbound

This paper cites Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning.

Robust image classification with multi-modal large language models Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.024471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.024471Z digest=sha256:4ee9540c016ce0399a1710f116a12bd60c640dfe16805529a09da8542f2126a1

Observation 10eed4f2-559b-4a97-a16a-7e7067022bdc · outbound

This paper cites Attackbench: Evaluating gradient-based attacks for adversarial examples,.

Robust image classification with multi-modal large language models Attackbench: Evaluating gradient-based attacks for adversarial examples,

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:00:13.654844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T16:00:12.992098Z digest=sha256:a75d163633be48909f1fe3be50048a457041462495c415cd1e2a15084caf6628

Pith citing papers

No inbound Pith citation observations are available.