Pith. sign in

Paper Citation Record · LEDGER

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.03542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03542 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:13:47.973601Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 514e91e1-0e31-4f17-ae39-97a34219086b · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning,.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Flamingo: a Visual Language Model for Few-Shot Learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.909807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:44.886696Z digest=sha256:d5d638916867b566301d9df8bf48976f0c1f004d26ba7e299d479f7a851a270e

Observation b2230a53-3e9d-4ba8-86b6-b639a7ca850b · outbound

This paper cites Vision-language models do not understand negation, 2025.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Vision-language models do not understand negation, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.715376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:45.066252Z digest=sha256:d0f96e633b5e83ec4b2f138078350fe2d3185596fe9c2b67918da27d27f296f5

Observation d092a9c8-e782-4bf8-9063-2ac28c24e2ce · outbound

This paper cites Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE).

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.126162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.126162Z digest=sha256:0951dae5e626c006b60935e211b5c536a5d0c2ff208fe546537b05675a750988

Observation 320ef00d-9715-47d0-be98-97bb7e4d2b81 · outbound

This paper cites Evolving interpretable visual classifiers with large language models.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Evolving interpretable visual classifiers with large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.657578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:45.211704Z digest=sha256:8608bebb6e0f023efb00809d4619aa74862f1fda8bb8af3be1bc549cbdaa4880

Observation 8c94f42e-3714-46a5-b39f-d5ff3fe20e6c · outbound

This paper cites The platonic representation hypothesis, 2024.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor The platonic representation hypothesis, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.594457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:45.304699Z digest=sha256:560650e7d7372b50d6b10aac7fcad7af9dd15ce52d216fa94ce042089143d3c1

Observation 4d74d4bf-483a-4ca9-9d77-c693da85ec94 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.396579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.396579Z digest=sha256:a8cf3a30ecdc36fc94396aa8b0cb6bfe1c94ad8ad41f76bf9bb149b3f42a532b

Observation d82a0fa9-6153-4417-9d18-8c851062505e · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.458422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.458422Z digest=sha256:6c763d1ac475f6956de56cbde6bc882fa394bef8ff1171670f294e55e66b7474

Observation 12f6097e-142c-47d1-b06c-549592b1a386 · outbound

This paper cites Large language models struggle to learn long-tail knowledge, 2023.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Large language models struggle to learn long-tail knowledge, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.538319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:45.556686Z digest=sha256:8ccef376860ca14a9ce47750a2aab81af0eae972aad40072c824b5452096b2aa

Observation e4a5d569-e6ae-401d-93b7-87eec66b9952 · outbound

This paper cites Cifar- 100 (canadian institute for advanced research).

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Cifar- 100 (canadian institute for advanced research)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.485661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:45.616748Z digest=sha256:283f0bac4f0fdceb5a4fbbd5a93dff76754f7ae86933d3e119de283a573a615f

Observation ae75167e-f446-46e0-a689-3ca8eba802ac · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.629794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.629794Z digest=sha256:4d42363325a0c5c27a0068666c509ed8c20b8d840db70984f2a26e137cb42784

Observation 8910940d-96a0-4bda-be08-75a09dfa55d2 · outbound

This paper cites Improved baselines with visual instruction tuning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Improved baselines with visual instruction tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.654032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.654032Z digest=sha256:42c1e31503800f624d55fc17164fa7c12a6e8987af9cf49d212f43f45376d1f8

Observation cffcfc55-cc6d-4671-93b8-a982536abc0a · outbound

This paper cites Visual instruction tuning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Visual instruction tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.841570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.841570Z digest=sha256:80ec88adcc7f1f1d17af4483e532c1dd3ae6eabe038496dd3db23bf0299ce962

Observation e51f2db3-70d4-4c48-a252-12925c062df8 · outbound

This paper cites Visual classification via description from large language models, 2022.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Visual classification via description from large language models, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.435605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:45.945125Z digest=sha256:4505951082a45ebed98fbdec1d180982fc089d7f4571e5d0ec2d402988fc5dfb

Observation c50dfedd-ed59-455a-b089-498faed52d00 · outbound

This paper cites Visual Classification via Description from Large Language Models.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Visual Classification via Description from Large Language Models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.412046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:46.106812Z digest=sha256:679bdb01e22344e4bf2b3aac601d532a84b5b6a96df15b5e51fa184b4e81973c

Observation 56026ccf-c83c-4f76-8e27-db45f747e6e7 · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Slip: Self-supervision meets language-image pre- training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:46.234593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:46.234593Z digest=sha256:262db4e1c19df4f98a920150e772f405dc2c6a9a6f7db4d4f9436aeec27e4bfb

Observation d5a8d458-3ca0-4235-a84f-6e78a3ee3794 · outbound

This paper cites Dinov2: Learning robust visual features with- out supervision, 2024.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Dinov2: Learning robust visual features with- out supervision, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:46.420003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:46.420003Z digest=sha256:844e05cca49432be09f7c7b233437fc9480e4c22cda701d6b0a22fb47b25fb3b

Observation 58fbfcf4-8fab-4580-9678-2f8346bafff2 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning Transferable Visual Models From Natural Language Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:46.524666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:46.524666Z digest=sha256:0e373330cd4ec263a5e170fbdf34ca0aba76b6a6fa9d5586a86ef871b5a43e93

Observation 9c9fdb4a-06e6-4ba2-8ebd-64d9158315fe · outbound

This paper cites Learning transferable visual models from natural language supervision.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning transferable visual models from natural language supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.306332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:46.571481Z digest=sha256:708ed9c0caeb5705a0ea8d416d402c7585ee18ffe431a7b578c82f39e33fb579

Observation f61589dd-d120-4037-aa20-442ae9ede13b · outbound

This paper cites Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.056043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:46.636919Z digest=sha256:3d343648e006d633944b3ab39b8e786c46996d971befb03bf22e99bd3fe04c4f

Observation a567c6e1-e12f-4ac8-880e-9a6f36b24095 · outbound

This paper cites Sun, and Swarat Chaudhuri.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Sun, and Swarat Chaudhuri

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.849513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:46.835514Z digest=sha256:575d90191c54eb769c1c94d87b4a2eeaa7642631e7da4a30e9c6d978b936dcbf

Observation faf6d0b9-3eef-4c57-b727-a935747683a1 · outbound

This paper cites Love, Christo- pher J.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Love, Christo- pher J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.657671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:46.965074Z digest=sha256:2c1a03e36c7231ef940949be4bad736b4731e6aeac9f59d44be1cf5063a27ee8

Observation 335682cf-102e-4107-8b7d-25d1d37a5bdc · outbound

This paper cites Understanding the emergence of multimodal representation alignment, 2025.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Understanding the emergence of multimodal representation alignment, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.400989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.109632Z digest=sha256:ea7aaaab8fdd1c580a831553a5fbddc4dfe376c5eddeb03de17333747567897d

Observation 98ee8c2b-cd79-4be2-852d-65a645b5068b · outbound

This paper cites Eyes Wide Shut? Exploring the Vi- sual Shortcomings of Multimodal LLMs.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Eyes Wide Shut? Exploring the Vi- sual Shortcomings of Multimodal LLMs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.239394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.212467Z digest=sha256:6047a9ec14b2d957b87334708070a27e93b66529632163c380ff1272d36b2a7a

Observation 6835cae2-5dcb-4d53-9c6f-1e658f69a0ca · outbound

This paper cites Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:49.923528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.298571Z digest=sha256:b3a134273d74dd206ee86b9cdf78e2ab4427097a4beb52f52b0cf48b87a624c7

Observation f318744b-75a4-477c-9b73-e1fdda7efc10 · outbound

This paper cites an unresolved cited work.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:13:49.601688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.470192Z digest=sha256:2a6d336160edd5a9d0fe656d51c694069949fe1b2e035e0c43e4beca5f244ff9

Observation 044a7e41-72bb-41ee-9e8f-f053afa823bf · outbound

This paper cites Learning Concise and Descriptive Attributes for Visual Recognition.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning Concise and Descriptive Attributes for Visual Recognition

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:13:48.328678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.524543Z digest=sha256:78fd06b8ce61856aa4a35105ce81c4a17ef0f0c8c24fe5d7401d9b8d7fd35950

Observation 40433328-06c2-4aa4-8d6f-d9e4f40a3787 · outbound

This paper cites Learning concise and descriptive attributes for visual recognition.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning concise and descriptive attributes for visual recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:47.625348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:47.625348Z digest=sha256:688a56b0a1041af76725b128da275accd582a162b7ab54268cb07bf5b914ec0e

Observation 2511dd34-242d-489f-a07b-a6243ef53d17 · outbound

This paper cites Language in a bottle: Language model guided concept bottlenecks for in- terpretable image classification, 2023.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Language in a bottle: Language model guided concept bottlenecks for in- terpretable image classification, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:49.284244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.722998Z digest=sha256:784bd3591799783facde1114b2e4a24f8041cee260c59f95d2279a02b1ce6ea5

Observation d193fc5f-09a5-488c-9559-38c6f818b9a8 · outbound

This paper cites Filip: Fine-grained interactive language-image pre-training.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Filip: Fine-grained interactive language-image pre-training

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:48.982821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.871811Z digest=sha256:5e769d7bcc65aeb8e9ac47a4c3b5f2913d35d30763cfcbbbfa654d245f6315b3

Observation 19274119-75c2-4d67-9c60-8d015a7e368e · outbound

This paper cites laysan albatross, which is a.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor laysan albatross, which is a

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:48.658859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:13:47.973601Z digest=sha256:9a489d4e32b37d20de05126c77690e9c5c136da219518d3f68290263b186de74

Observation 0bdd19ee-a710-47ff-b531-5ce9f118ab5d · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:44.966493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:44.966493Z digest=sha256:ac1e8cf243a3787766d0222abedcd09a0fe1e872d5da2953621c5d7ba8171539

Pith citing papers

No inbound Pith citation observations are available.