Pith. sign in

Paper Citation Record · LEDGER

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor

As of 20 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.03542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03542 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:13:47.973601Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 514e91e1-0e31-4f17-ae39-97a34219086b · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning,.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Flamingo: a Visual Language Model for Few-Shot Learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.909807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:44.886696Z digest=sha256:6547991b0dc85fba3e60689a4730a8d968d9b0aff154aa3cd7152c570d2758c1

Observation b2230a53-3e9d-4ba8-86b6-b639a7ca850b · outbound

This paper cites Vision-language models do not understand negation, 2025.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Vision-language models do not understand negation, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.715376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:45.066252Z digest=sha256:9b0278af5f70f131fd73ddce0f04c7fdc378441ef3f341feae16bcc6863f3cc8

Observation d092a9c8-e782-4bf8-9063-2ac28c24e2ce · outbound

This paper cites Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE).

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.126162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.126162Z digest=sha256:0951dae5e626c006b60935e211b5c536a5d0c2ff208fe546537b05675a750988

Observation 320ef00d-9715-47d0-be98-97bb7e4d2b81 · outbound

This paper cites Evolving interpretable visual classifiers with large language models.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Evolving interpretable visual classifiers with large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.657578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:45.211704Z digest=sha256:60bfac6b9c127e578b995738fdfbfd4fccd0796846fa3adc90b3ded11dde2ff4

Observation 8c94f42e-3714-46a5-b39f-d5ff3fe20e6c · outbound

This paper cites The platonic representation hypothesis, 2024.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor The platonic representation hypothesis, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.594457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:45.304699Z digest=sha256:31ea91905ee91b968670296d66c7b5b2509dc816ff8ccc9e2dea0fed843d98b6

Observation 4d74d4bf-483a-4ca9-9d77-c693da85ec94 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.396579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.396579Z digest=sha256:a8cf3a30ecdc36fc94396aa8b0cb6bfe1c94ad8ad41f76bf9bb149b3f42a532b

Observation d82a0fa9-6153-4417-9d18-8c851062505e · outbound

This paper cites What's "up" with vision-language models? Investigating their struggle with spatial reasoning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor What's "up" with vision-language models? Investigating their struggle with spatial reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.458422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.458422Z digest=sha256:6c763d1ac475f6956de56cbde6bc882fa394bef8ff1171670f294e55e66b7474

Observation 12f6097e-142c-47d1-b06c-549592b1a386 · outbound

This paper cites Large language models struggle to learn long-tail knowledge, 2023.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Large language models struggle to learn long-tail knowledge, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.538319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:45.556686Z digest=sha256:1007d6b73fa2924c8deac98c5e847632546dff66348945bb3d21b2b95632807c

Observation e4a5d569-e6ae-401d-93b7-87eec66b9952 · outbound

This paper cites Cifar- 100 (canadian institute for advanced research).

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Cifar- 100 (canadian institute for advanced research)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.485661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:45.616748Z digest=sha256:a38ea5a00a0e817a095e7becc898613ed37f9475df1889593226678e8dd1f2f3

Observation ae75167e-f446-46e0-a689-3ca8eba802ac · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Align before fuse: Vision and language representation learn- ing with momentum distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.629794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.629794Z digest=sha256:4d42363325a0c5c27a0068666c509ed8c20b8d840db70984f2a26e137cb42784

Observation 8910940d-96a0-4bda-be08-75a09dfa55d2 · outbound

This paper cites Improved baselines with visual instruction tuning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Improved baselines with visual instruction tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.654032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.654032Z digest=sha256:42c1e31503800f624d55fc17164fa7c12a6e8987af9cf49d212f43f45376d1f8

Observation cffcfc55-cc6d-4671-93b8-a982536abc0a · outbound

This paper cites Visual instruction tuning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Visual instruction tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:45.841570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:45.841570Z digest=sha256:80ec88adcc7f1f1d17af4483e532c1dd3ae6eabe038496dd3db23bf0299ce962

Observation e51f2db3-70d4-4c48-a252-12925c062df8 · outbound

This paper cites Visual classification via description from large language models, 2022.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Visual classification via description from large language models, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.435605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:45.945125Z digest=sha256:e7af95428a81584ef35f2b1e25ffe3f0054f7ce9b64831c30f6be29fb96821c0

Observation c50dfedd-ed59-455a-b089-498faed52d00 · outbound

This paper cites Visual Classification via Description from Large Language Models.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Visual Classification via Description from Large Language Models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.412046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:46.106812Z digest=sha256:c0296e3170a1539bb831e71e005eb1878875639d0b6ac8753849b1964e21571a

Observation 56026ccf-c83c-4f76-8e27-db45f747e6e7 · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Slip: Self-supervision meets language-image pre- training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:46.234593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:46.234593Z digest=sha256:262db4e1c19df4f98a920150e772f405dc2c6a9a6f7db4d4f9436aeec27e4bfb

Observation d5a8d458-3ca0-4235-a84f-6e78a3ee3794 · outbound

This paper cites Dinov2: Learning robust visual features with- out supervision, 2024.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Dinov2: Learning robust visual features with- out supervision, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:46.420003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:46.420003Z digest=sha256:844e05cca49432be09f7c7b233437fc9480e4c22cda701d6b0a22fb47b25fb3b

Observation 58fbfcf4-8fab-4580-9678-2f8346bafff2 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning Transferable Visual Models From Natural Language Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:46.524666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:46.524666Z digest=sha256:0e373330cd4ec263a5e170fbdf34ca0aba76b6a6fa9d5586a86ef871b5a43e93

Observation 9c9fdb4a-06e6-4ba2-8ebd-64d9158315fe · outbound

This paper cites Learning transferable visual models from natural language supervision.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning transferable visual models from natural language supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.306332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:46.571481Z digest=sha256:bba2deb6850f7af0fc2fc9afad6bc2f3b8609e1b8dbb6121c0c73d8de63e9ed6

Observation f61589dd-d120-4037-aa20-442ae9ede13b · outbound

This paper cites Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:51.056043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:46.636919Z digest=sha256:14162d15d436704c4cbc6a4ff878dbd286e3d65efddc044c48d3e5c9f2c3363b

Observation a567c6e1-e12f-4ac8-880e-9a6f36b24095 · outbound

This paper cites Sun, and Swarat Chaudhuri.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Sun, and Swarat Chaudhuri

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.849513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:46.835514Z digest=sha256:cb5b34f17180eb4e01f4468f69f50f515c87a183aabb4481cb0de610d925bc82

Observation faf6d0b9-3eef-4c57-b727-a935747683a1 · outbound

This paper cites Love, Christo- pher J.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Love, Christo- pher J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.657671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:46.965074Z digest=sha256:5b7ba7585caca37aa34e73774725f97662d4eb400438cc7ea1f7d96548688053

Observation 335682cf-102e-4107-8b7d-25d1d37a5bdc · outbound

This paper cites Understanding the emergence of multimodal representation alignment, 2025.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Understanding the emergence of multimodal representation alignment, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.400989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.109632Z digest=sha256:0638ccf1e951e3a63e09acda9dd1b598cfbd3bb028da83e4e0f1abc655d5572c

Observation 98ee8c2b-cd79-4be2-852d-65a645b5068b · outbound

This paper cites Eyes Wide Shut? Exploring the Vi- sual Shortcomings of Multimodal LLMs.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Eyes Wide Shut? Exploring the Vi- sual Shortcomings of Multimodal LLMs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:50.239394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.212467Z digest=sha256:c529a3a566635be421ae01563038f8dc7d18453b1a7fe868892543118fc2d889

Observation 6835cae2-5dcb-4d53-9c6f-1e658f69a0ca · outbound

This paper cites Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:49.923528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.298571Z digest=sha256:eb69940bb88d2cd0fdea20c1e6e7a3fa8de9e1d92b9ab4d2cfd830e6f844817c

Observation f318744b-75a4-477c-9b73-e1fdda7efc10 · outbound

This paper cites an unresolved cited work.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:13:49.601688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.470192Z digest=sha256:be06dc40743f84c1affa11579bdc84f88b12928a552b7c90725555323c1435bc

Observation 044a7e41-72bb-41ee-9e8f-f053afa823bf · outbound

This paper cites Learning Concise and Descriptive Attributes for Visual Recognition.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning Concise and Descriptive Attributes for Visual Recognition

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:13:48.328678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.524543Z digest=sha256:79e19201359fa1b6d86a17d4c07c494e10f6579489e5694d1623eee23aef77ef

Observation 40433328-06c2-4aa4-8d6f-d9e4f40a3787 · outbound

This paper cites Learning concise and descriptive attributes for visual recognition.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Learning concise and descriptive attributes for visual recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:47.625348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:47.625348Z digest=sha256:688a56b0a1041af76725b128da275accd582a162b7ab54268cb07bf5b914ec0e

Observation 2511dd34-242d-489f-a07b-a6243ef53d17 · outbound

This paper cites Language in a bottle: Language model guided concept bottlenecks for in- terpretable image classification, 2023.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Language in a bottle: Language model guided concept bottlenecks for in- terpretable image classification, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:49.284244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.722998Z digest=sha256:fe5e7d578449a52b117b38f9953be9816006bf1b0da9d8e0e274509f49d42440

Observation d193fc5f-09a5-488c-9559-38c6f818b9a8 · outbound

This paper cites Filip: Fine-grained interactive language-image pre-training.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Filip: Fine-grained interactive language-image pre-training

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:48.982821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.871811Z digest=sha256:c3c0e121be939d37de17c97f434bd05eb61c273b7abb95d770645657e23c2a73

Observation 19274119-75c2-4d67-9c60-8d015a7e368e · outbound

This paper cites laysan albatross, which is a.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor laysan albatross, which is a

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:13:48.658859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T20:13:47.973601Z digest=sha256:e49eea0d0bea0784418cad0cc16675694fcc6a9ca57950ef6f724d863dcc0dc2

Observation 0bdd19ee-a710-47ff-b531-5ce9f118ab5d · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:13:44.966493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:13:44.966493Z digest=sha256:ac1e8cf243a3787766d0222abedcd09a0fe1e872d5da2953621c5d7ba8171539

Pith citing papers

No inbound Pith citation observations are available.