Pith. sign in

Paper Citation Record · LEDGER

Why are Visually-Grounded Language Models Bad at Image Classification?

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.18415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.18415 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:14.376214Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T18:26:55.252818Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d7c1eef8-82cb-4789-8adb-af3f9e23165c · inbound

De-biased Multimodal Electrocardiogram Analysis cites this paper.

De-biased Multimodal Electrocardiogram Analysis Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:58:10.084164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:58:10.084164Z digest=sha256:7e8985d882b6942da9f64264752205765534314ad0655a3ddaa3d597063725c5

Observation 8fb2b9bb-2758-4d91-bdc8-85307ab7a811 · inbound

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models cites this paper.

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T14:52:57.345045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:52:57.345045Z digest=sha256:a6920f592f2d344b8e906a7e257e082be3b3a37e4fc545062a2aca6ec22b51c1

Observation eafa2640-0500-469a-9c8d-5b60b542520c · inbound

DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users cites this paper.

DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:50:51.523727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:50:51.523727Z digest=sha256:8a33a57e1d64e680d1705e059dd166e646dd9db2518939cc94e0ca6a36e33436

Observation 8a1675ab-2c43-475f-9b3b-d6bc147de72c · inbound

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features cites this paper.

Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-12T10:24:06.940093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:24:06.940093Z digest=sha256:c923e692747c1fd89294fafe1f5c8188d12105d859471dbbe54ebf4052bcbc80

Observation 879d9d47-ff82-46aa-9669-be9f00d0af6c · inbound

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging cites this paper.

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:45.287415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:45.287415Z digest=sha256:5b31ba5d821ea751db0457be7bc9620882658997f5fc4031f73a01661f65a4c7

Observation 7048e534-36ba-4009-97b0-65f362a3a0f4 · inbound

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization cites this paper.

Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:14.376214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:14.376214Z digest=sha256:a32ea5c1284fafabc33a617550cbfe0e737ff8ee50045c875eefa1bef7c1be39

Observation a00e4c4e-a295-454d-ac90-bf1040b28d75 · inbound

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation cites this paper.

Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:26:55.254753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T18:26:12.597756Z digest=sha256:c98ac81b41c2fa9e3d46ef550981e88d78d7cec993faf40a62609e074c01165e

Observation c38dcdf1-237c-439e-94e0-91c601db2e2e · inbound

Single Domain Generalization for Few-Shot Counting via Universal Representation Matching cites this paper.

Single Domain Generalization for Few-Shot Counting via Universal Representation Matching Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:02.227715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:02.227715Z digest=sha256:115c805a91c5107fbcc1bfc79053388ea194fa786009727595e55e06c76e68aa

Observation 22352f41-f8c3-4b2c-851f-73ded0b38e90 · inbound

GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models cites this paper.

GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:30.556929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:30.556929Z digest=sha256:5853370e47025ba31a857bd58b50e2649f09004e2e1f9dbe34c1b817bf520035

Observation 4abf9151-01b7-4781-ac18-c573087a4c8d · inbound

AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models cites this paper.

AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-06T20:32:59.911874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:32:59.911874Z digest=sha256:9ae633c3731192d1b5e4c511afc09825048c85885b95572ee4f0714896487cfa

Observation 2f4deb52-f0f0-4b76-8427-1491df391ee8 · inbound

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models cites this paper.

Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:27:46.828449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:27:46.828449Z digest=sha256:a6cb90dbc120e9381b7ef6563c16bd6e24631ac953c52a90bf879c33c6f02d5f

Observation 5b47dfc0-6f42-40af-9456-f23975f858c7 · inbound

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation cites this paper.

Filter-And-Refine: A MLLM Based Cascade System for Industrial-Scale Video Content Moderation Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:44.118082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:58:44.118082Z digest=sha256:f0a71da2a4fd3c833a49f7036bd6f60a72b6258bb3a93f3b1f9c115153113e57

Observation 9df4ce5f-c840-4476-becc-a370ef38184e · inbound

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs cites this paper.

Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:22.267079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:22.267079Z digest=sha256:ffc2dd98ad21ca8c4ce4ee5ffceba654e0c5e7719d2bd3efeaef597de9be24e1

Observation 62803457-e2fe-4e8e-8a31-9fc12368c114 · inbound

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes cites this paper.

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T15:47:21.277306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:47:21.277306Z digest=sha256:3c6023ba147f0f641b977b1afb0e6339f181f4ad469a646d2bce6b10539998e9

Observation 118e2245-8dfb-4c81-b2d5-eeebd7a5a3b7 · inbound

Unpacking Hateful Memes: Presupposed Context and False Claims cites this paper.

Unpacking Hateful Memes: Presupposed Context and False Claims Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T10:26:23.813098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:26:23.813098Z digest=sha256:aaa0b7a01d2012798c23e50a720fd17dc2a47b683f3c223198fd16fb8488cad2

Observation f71b06a0-b4fc-496a-a091-da1324cc9c3b · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Why are Visually-Grounded Language Models Bad at Image Classification?

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:49.155307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:49.155307Z digest=sha256:e5ee87994740d09922e37691de5b90a8c26ef3ec8f3fbc534809016f5a460dc7