Pith. sign in

Paper Citation Record · LEDGER

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition

As of 21 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2412.13947.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13947 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:42:01.307694Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57f7fd7a-a03f-4465-8301-1340777590ff · outbound

This paper cites Food-101 – mining discriminative components with random forests.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Food-101 – mining discriminative components with random forests

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:02.011963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.107324Z digest=sha256:bd8f344ca256e211039320244090a64a0e09e328a7ee10ed82900dd265570021

Observation 5785fb7b-5673-467b-91cf-9c21af468b69 · outbound

This paper cites Crossvit: Cross-attention multi-scale vision transformer for image classification, 2021.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Crossvit: Cross-attention multi-scale vision transformer for image classification, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.993424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.114009Z digest=sha256:6fb6f440271395cfaa858cd166e81b2e9db9f437a4766922b6c58dac191e1ed7

Observation 71b3f697-bf71-48ab-9d68-97a77c8f1807 · outbound

This paper cites Ovarnet: Towards open- vocabulary object attribute recognition.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Ovarnet: Towards open- vocabulary object attribute recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.975533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.120273Z digest=sha256:ad793fe5dac476bee273a5c5e89d5d34db6b8947a8c089f81f76d04bbaa1c734

Observation 1ec9753c-8f5d-43ff-bfe4-c423ccd0c27d · outbound

This paper cites Multi- modal classifiers for open-vocabulary object detection.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Multi- modal classifiers for open-vocabulary object detection

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.957269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.127491Z digest=sha256:13843d1e87df2dbc22e5544bb26427242f1d694ad825da4d10868f1a15af4561

Observation 5f26d611-452d-4b0a-a489-45fc7211787c · outbound

This paper cites Novel dataset for fine-grained image categorization: Stanford dogs.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Novel dataset for fine-grained image categorization: Stanford dogs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.936374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.135394Z digest=sha256:5f51f6a90146a4c07b9bf4e01179c7324d14ce624d9a87de7be0c70c293bafd4

Observation a32dff49-1dc5-4eac-8f82-d7d0d9f189b8 · outbound

This paper cites 3d object representations for fine-grained categorization.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition 3d object representations for fine-grained categorization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.913556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.145234Z digest=sha256:3078177807a00e4227edb53d0824b72f12a1827ee25f4846ff2d2a4e448c3d03

Observation 6b7d32d4-d92f-4e24-bcf4-8a35e5c7d59c · outbound

This paper cites Descriptor and Word Soups: Overcoming the Parameter Efficiency Accuracy Tradeoff for Out-of-Distribution Few-shot Learning.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Descriptor and Word Soups: Overcoming the Parameter Efficiency Accuracy Tradeoff for Out-of-Distribution Few-shot Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:42:01.459296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.152748Z digest=sha256:1cfce2407b2b537bcabbb4efae9dfa4d7b6a068c91912cfafd5e00806d816552

Observation 49d266cc-7619-4fd0-b5b7-67c81d3d18cb · outbound

This paper cites Feature pyramid networks for object detection, 2017.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Feature pyramid networks for object detection, 2017

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.888766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.160023Z digest=sha256:0f85943936254acbaf13b83bb0c39997f01cad518a2bb8928a4f484f71c1cb49

Observation 64612c1f-0577-43dd-be12-ab3b7620beba · outbound

This paper cites Visual Classification via Description from Large Language Models.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Visual Classification via Description from Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:42:01.165695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:42:01.165695Z digest=sha256:3b713a23f9da74adb7fba671cea830b44e31bd9d9ed343184fc63d63dffbd740

Observation 3e1d83d8-e508-4672-ba79-8f9f2c40b04a · outbound

This paper cites Automated flower classification over a large number of classes.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Automated flower classification over a large number of classes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.863219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.171566Z digest=sha256:69ae820da94f3ae7dbd36906b7864d1f08878533b7f5c6f4436cac1afdad731c

Observation 90582daa-7a69-415f-815c-9bb25566ef98 · outbound

This paper cites Parkhi, Andrea Vedaldi, Andrew Zisserman, and C.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Parkhi, Andrea Vedaldi, Andrew Zisserman, and C

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.840799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.179209Z digest=sha256:803e83f41e7ba8e13e787e9404a37d9fbe5ab4ba0350b09dd7da7f211fd7d50f

Observation 2822b140-c3bd-4417-ae51-b7e881f74c26 · outbound

This paper cites What does a platypus look like? generating customized prompts for zero-shot image classification.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition What does a platypus look like? generating customized prompts for zero-shot image classification

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.817385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.185865Z digest=sha256:b9615eab6558dff3b83bc4a4d17d09410af0b61f5113c3227b72965bec683559

Observation ca6f0a4b-d1be-42e8-88f5-e1070f37daf3 · outbound

This paper cites What is the limitation of multimodal llms? a deeper look into multimodal llms through prompt prob- ing.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition What is the limitation of multimodal llms? a deeper look into multimodal llms through prompt prob- ing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.793639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.192499Z digest=sha256:14dd35832eb7e649038c9d5fc7cc182d6c47b802e5260f7294c4e2ce759958ca

Observation 044e5b52-e0b7-455d-9909-db804f121917 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Learning transferable visual models from natural language supervi- sion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:42:01.201618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:42:01.201618Z digest=sha256:4c6611fc5504c045ec42adddcc3f4a788f1f8f1eb3c08b551f25319cea80648c

Observation 43d5c935-1826-40df-a087-c8ff82670836 · outbound

This paper cites Paco: Parts and attributes of common objects, 2023.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Paco: Parts and attributes of common objects, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.755939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.208226Z digest=sha256:076ecf8d5e1d92917e659b0e04178fee8f9518356512f2ae87018a4d73f89868

Observation 55b7f611-7b0d-4623-afe3-5f3d02067bb5 · outbound

This paper cites Imagenet-21k pretraining for the masses.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Imagenet-21k pretraining for the masses

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.730711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.214448Z digest=sha256:d483909946e070eb36852bd9029348f66fd02b8ffc703903682c7abf4e746be4

Observation 11c97817-e1b7-414f-8fd6-81dc73e2ebff · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition U-net: Convolutional networks for biomedical image segmentation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:42:01.219846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:42:01.219846Z digest=sha256:4d84397e221c61dff8bae553a8b494260dc4a2ca8d7f40435308e4ee39b742b7

Observation a140945d-003c-4216-99e9-edbaa9c998b0 · outbound

This paper cites Waffling around for performance: Visual classification with random words and broad concepts.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Waffling around for performance: Visual classification with random words and broad concepts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.696296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.225717Z digest=sha256:03a0f82973c693c55e4c7c82378cb6ac7e2609443baf1a377df0a3beb80715fc

Observation cce4af3c-f0bc-4f1f-9cba-31b08cb50ec8 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.674498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.234945Z digest=sha256:4a74488136bc6edcf2dc172973366eb78134f9525894dca6f9a67ca0a110c0d5

Observation 8c46673c-c832-4e37-93dd-c0fa975f4fdc · outbound

This paper cites When do we not need larger vision models?, 2024.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition When do we not need larger vision models?, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.654920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.244146Z digest=sha256:2d8dae5d8d67a0e054fd67cbef0cba1aab22b96543ebb054b7e8c912f8be7a44

Observation 2d682432-b4a2-4a7c-ba9b-9acc6e281ca4 · outbound

This paper cites Alpha-CLIP: A CLIP Model Focusing on Wherever You Want.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:42:01.250402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:42:01.250402Z digest=sha256:08700092bec114b6016e9e12f0f3e8794b1eed0afea06eb94548a54154438d1f

Observation 2d9f8e97-8c4d-443b-8e2a-4e7701721f1f · outbound

This paper cites ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:42:01.377095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.256552Z digest=sha256:c243ef94a7a7e852074b548fe19df7c048c609b9bfe0acae37a8a1200e5f0438

Observation 97b0b32b-9eaa-4840-96f6-0807da876d59 · outbound

This paper cites Efficient object localization using convolutional networks, 2015.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Efficient object localization using convolutional networks, 2015

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.627699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.262792Z digest=sha256:b7664f5a75fd132a40f320e8960772ebc3e5360c4ab38bc2c3c7dd08fc84b642

Observation a569b525-0438-41e8-9a15-979427efe460 · outbound

This paper cites an unresolved cited work.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:42:01.606992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.270465Z digest=sha256:5d6e039deb3fcc482911adb23b59b0739899c93f68f11e6090ee4b9fd4f8e160

Observation 0b8e864f-4e75-4d24-b9b1-3ca99974aef7 · outbound

This paper cites Learning concise and descriptive attributes for visual recognition.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Learning concise and descriptive attributes for visual recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:42:01.276390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:42:01.276390Z digest=sha256:74c8cc2eacbf15395e413ae9a4a99fddb067e96e090532833ca06f564c68e97b

Observation c5059604-3cba-49d4-84af-d05be2df3d68 · outbound

This paper cites Focal self-attention for local-global interactions in vision transformers, 2021.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Focal self-attention for local-global interactions in vision transformers, 2021

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.569481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.282297Z digest=sha256:b69eb4c46c45de87e81c5b74745d7d8d2903d6cadaa38864c1f7b787dba4fe10

Observation 60c42623-3596-4e84-b6e0-e0f2cd0384e3 · outbound

This paper cites Language in a bottle: Language model guided concept bottlenecks for interpretable image classification.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Language in a bottle: Language model guided concept bottlenecks for interpretable image classification

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.547832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.288241Z digest=sha256:ca732786687acfcc3581703687cc3dff9b4a6856a7e41dd90bbe6320fefff236

Observation 70ceeb75-2be3-491a-8560-29ff5d6a44c1 · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2022.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition When and why vision- language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.527141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.294066Z digest=sha256:b2fd3329f79a71523d650411aa00f8e099b178d9586ab189d579be851af0ed6a

Observation 47fb1570-e149-4915-960e-25642ac14be4 · outbound

This paper cites Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.502464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.300124Z digest=sha256:32ca34f670ce6f716e9d88af9e8f50cc0679bcc6a92369313c1019edffef942b

Observation cf47ab0c-65a8-45c7-9986-e6f037f5e101 · outbound

This paper cites Vl- checklist: Evaluating pre-trained vision-language models with objects, attributes and relations, 2023.

Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition Vl- checklist: Evaluating pre-trained vision-language models with objects, attributes and relations, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:42:01.481399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T12:42:01.307694Z digest=sha256:a68705a3ee8e4110668c4f82724f2e19a68cce85f99fccb02ac96612da9cf122

Pith citing papers

No inbound Pith citation observations are available.