Pith. sign in

Paper Citation Record · LEDGER

On the rankability of visual embeddings

As of 16 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2507.03683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03683 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:11:51.717772Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact11
  • verified fuzzy29
  • unresolved29
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72fcb7d1-1878-4a5b-987a-7a8956bbd1b7 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

On the rankability of visual embeddings Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:48.623426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:48.623426Z digest=sha256:2571c941ab3374b9770b84a46f8669aa38e6ee2e07121592a889880b3d1a7750

Observation 72f8616e-862e-4547-92c6-4f8ae4d82f13 · outbound

This paper cites Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features.

On the rankability of visual embeddings Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.679728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:48.662843Z digest=sha256:158c08a61e8e54132f3cf833be1be5151d1045019e712134cba167218809e051

Observation 671d080d-4f91-48ef-b91b-bc3ff30dd0f1 · outbound

This paper cites Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models.

On the rankability of visual embeddings Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.140579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:48.742002Z digest=sha256:f9a1fb6bff7b117adeb1c5fbe6a33d78218d4c658bb8969a05d1c583da9a681f

Observation 82b92998-93e2-45a5-9a65-d480c670884b · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

On the rankability of visual embeddings A Simple Framework for Contrastive Learning of Visual Representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.663979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:48.803202Z digest=sha256:c30ad585eacae490e063d53d642d20bab92069f394fbe6df595d5d565a8ea3e9

Observation 511cf3cc-d07f-4c12-a0ae-b1d11993093b · outbound

This paper cites Deep Learning for Instance Retrieval: A Survey.

On the rankability of visual embeddings Deep Learning for Instance Retrieval: A Survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.647220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:48.894792Z digest=sha256:7af1c4ee2bc100b42eb0829f89d2186c4e98e431f6df688b54d74165c7cc19ed

Observation 0208f5c9-189e-4893-a53b-1e5093e7b4a4 · outbound

This paper cites Deep learning for instance retrieval: A survey.

On the rankability of visual embeddings Deep learning for instance retrieval: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.630656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:48.981497Z digest=sha256:ddcdaf4c7616d12c195e6be14203005e5938c4ec06e0cb773e49491eacc11eab

Observation af97d169-f7c6-4cc0-8ed5-efa72c98a39c · outbound

This paper cites Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds.

On the rankability of visual embeddings Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:49.072090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.072090Z digest=sha256:cbca4ba98c7d1d68c21306ef59b95e177403550c7febd6af275d720858869ed2

Observation 94a871ac-e6e3-41c2-b3b3-06ddd5a78bde · outbound

This paper cites Hyperbolic Image-Text Representations.

On the rankability of visual embeddings Hyperbolic Image-Text Representations

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.112878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:49.128084Z digest=sha256:b9bb684ccc7ec8ea795efdf0efabf6e9efeefd175fd8e37c9c730112c1e77d45

Observation 2d825a8c-c3ea-4a01-a78d-877669bf5ce3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

On the rankability of visual embeddings An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.229769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.229769Z digest=sha256:03b8c73e43110e5e313fa3b63d5135efa1e0d81cb47965b724080d4cf8e4c69d

Observation b6cc826c-934f-489b-94e1-f44b9422ddc1 · outbound

This paper cites Teach CLIP to Develop a Number Sense for Ordinal Regression.

On the rankability of visual embeddings Teach CLIP to Develop a Number Sense for Ordinal Regression

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.059438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:49.329152Z digest=sha256:f7e89453c99728769a9164a5ff3e279072548e9bb8a36df8a1281d5b6038749c

Observation 8883acf4-b68b-4343-903e-8706086d6021 · outbound

This paper cites Age and Gender Estimation of Unfiltered Faces.

On the rankability of visual embeddings Age and Gender Estimation of Unfiltered Faces

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:53.088420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:49.450137Z digest=sha256:b126880450d5a8e6c22eb673da72e48246dfe9b2d41646282175469654673076

Observation 9cfaf071-0d3c-4582-83b4-b5dd8d280749 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

On the rankability of visual embeddings It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.531103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.531103Z digest=sha256:1ff97165108b44b005f0260b74881a4bb8342a5d8a1d311e96cfdb8caf3e3336

Observation 9dd7610a-3d37-4134-a268-e1343c54c0f3 · outbound

This paper cites Heterogeneous face attribute estimation: A deep multi-task learning approach.

On the rankability of visual embeddings Heterogeneous face attribute estimation: A deep multi-task learning approach

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.613502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:49.721697Z digest=sha256:429e37e2c2c9ac538903d010aff70dd8e972429d6d0e67c9197c19ac3f41821c

Observation 477da1c2-50ff-4f16-8b1d-05c3201567d5 · outbound

This paper cites Deep Residual Learning for Image Recognition.

On the rankability of visual embeddings Deep Residual Learning for Image Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.805256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.805256Z digest=sha256:95a67a1f001354fbc4d9ab786e4cc4e28d85b199be8b8a4e66be01a8ae8da4fe

Observation 307ea798-7ea6-4281-aa6e-2432bac0103f · outbound

This paper cites Multilayer feedforward networks are universal approximators.

On the rankability of visual embeddings Multilayer feedforward networks are universal approximators

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.597532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:49.928316Z digest=sha256:9c79a538eadf2be6339ed0fccecb1b92280958b18061e47b952956fe93c72885

Observation bb2b5fda-2abe-47f1-b5e9-2c4c8f7635f0 · outbound

This paper cites KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment.

On the rankability of visual embeddings KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.582027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:50.055806Z digest=sha256:8945b46b076e68fae53af591853c94c218433eff0297c2de4301e00acb80fc4a

Observation 97fb22fc-96d4-4760-9b18-e002309f4483 · outbound

This paper cites Lp++: A surprisingly strong linear probe for few-shot clip.

On the rankability of visual embeddings Lp++: A surprisingly strong linear probe for few-shot clip

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.566871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:50.309455Z digest=sha256:ca87ad917bd0286f0485a5f735f1828ca491f6aefdde1478a79d9922641744e9

Observation 9d9125c5-7393-4fdd-90fd-9213f230accf · outbound

This paper cites The Platonic Representation Hypothesis.

On the rankability of visual embeddings The Platonic Representation Hypothesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.427501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.427501Z digest=sha256:e942eb59ee612a7313a0a55f01e182426877a37391b5e23502e035acc819ea43

Observation 1e02910c-73b7-4adf-a889-8aa29d70c8f0 · outbound

This paper cites CLIP-Count: Towards Text-Guided Zero- Shot Object Counting.

On the rankability of visual embeddings CLIP-Count: Towards Text-Guided Zero- Shot Object Counting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.544893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.544893Z digest=sha256:6d5ab4a9becbef76e70da1c79a4cf2eefff18ab6010f7d4994953b1ff543e636

Observation 22adfc6d-f764-4110-8eea-9146e82b38b7 · outbound

This paper cites Hyperbolic Image Embeddings.

On the rankability of visual embeddings Hyperbolic Image Embeddings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.549000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:50.625607Z digest=sha256:92bf251dc018e0a55082b2a3019f6bfe9a38a19919c1c850cd9661134d011da8

Observation 86d92a71-4aea-4f83-ad09-41ba0cd792a0 · outbound

This paper cites MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs.

On the rankability of visual embeddings MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.693133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.693133Z digest=sha256:ac8ec65e84097191be3c17933729e2d274037f4f2fff62c5119674787ddc8711

Observation 4c0d978c-48c2-461d-9110-e11b77b6153b · outbound

This paper cites Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav).

On the rankability of visual embeddings Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.529295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:50.813651Z digest=sha256:6e3824c7d64218d8fca6f3bc585663880b71ca103e0a5bdb8f188cf7e9fa2eea

Observation d74f666a-fe1f-4221-be11-5cb58b0a6fd4 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.960886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.960886Z digest=sha256:3730222c44a9c8b2b9931039cc2ce26e1fb5844f451041d0190ac8f6ea299ffc

Observation de4d0571-cc07-4331-a0c2-034a6b53f7b1 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.002032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.002032Z digest=sha256:7373c87ab5e1ba196de6aa8e53efab14ee503d5df0d2c6014540c906b04e1b3c

Observation d160e3ef-edc5-4886-b0e0-5c3704ab4ed2 · outbound

This paper cites Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning.

On the rankability of visual embeddings Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.512741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.033881Z digest=sha256:99747cc5fbc7db0c340cfc194c85c45f698a7cb0649fca41faaf98ac72df6e70

Observation c5ada6c4-84e0-41ad-b510-1f10d8b09f98 · outbound

This paper cites MiVOLO: Multi-input Transformer for Age and Gender Estimation.

On the rankability of visual embeddings MiVOLO: Multi-input Transformer for Age and Gender Estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.039704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.039704Z digest=sha256:501a848ddbf26a3c2750dcc45f93a0c68f651637d7054827d87211ba3d209e71

Observation 9e393176-8412-4a5a-816e-ea283721eaac · outbound

This paper cites The Double-Ellipsoid Geometry of CLIP.

On the rankability of visual embeddings The Double-Ellipsoid Geometry of CLIP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.045027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.045027Z digest=sha256:98e8690ff182c1a9b0ef5bdd6253a7aeb625be89e355028a89fbe8f0357602d5

Observation 4d502446-6440-47de-b3a7-37f7735ac5f9 · outbound

This paper cites Does CLIP Bind Concepts? Probing Compositionality in Large Image Models.

On the rankability of visual embeddings Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.071811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.071811Z digest=sha256:b287dcd466b1c2bc1c5f0aa3a63cd3f61782e15eeeca0abb2b0356e29d5a28a3

Observation 2624fd97-d6cf-4533-b4f8-02e1e764d3dd · outbound

This paper cites Align before fuse: Vision and language representation learning with mo- mentum distillation.

On the rankability of visual embeddings Align before fuse: Vision and language representation learning with mo- mentum distillation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.494930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.171974Z digest=sha256:1618c675e9f0211ac786bd87ade3c72e3b43bd08d07bc79f46e6d46f26f8a2e2

Observation bf413c30-217d-4123-95e4-6fe4efb58ffe · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation.

On the rankability of visual embeddings BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.478419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.232772Z digest=sha256:cc2fc86dbcc2229c18ff1732aa033c076bc6436d72b737a509cf1a23072e36bf

Observation f366b8f4-65fd-4e5b-8ae5-52c09621ca3b · outbound

This paper cites OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression.

On the rankability of visual embeddings OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.956760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.302801Z digest=sha256:f476bda4aeea6749adf8f79b9832f9428c2ee10f6f689bba35d5544369b3478f

Observation 00efac1b-a70d-4467-b8f4-0ba5c81b0c6d · outbound

This paper cites CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model.

On the rankability of visual embeddings CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.933505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.386439Z digest=sha256:74277269a200cd28b15279714e0da271de9bc9189d5ca210ea7f4898c0d093e2

Observation ffb1b424-5803-447c-983e-274b6ec42009 · outbound

This paper cites Beyond comparing image pairs: Setwise active learning for relative attributes.

On the rankability of visual embeddings Beyond comparing image pairs: Setwise active learning for relative attributes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.462867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.467629Z digest=sha256:73248e04d1906c579f37ca96443d7af3d8f48c50a557c7ce69cf72e04b80a8b6

Observation 9cdd71ea-a5ce-4033-93fd-6d02518fd554 · outbound

This paper cites CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification.

On the rankability of visual embeddings CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.500131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.500131Z digest=sha256:47f69d36f43d8db6aa9e1734753ebe77bce79e9cfb653ff9d684687070582eec

Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.511161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.511161Z digest=sha256:df4d0de779b1108e70106a433a01412c29f346349897addc8e2dbd40b826f20f

Observation ead6e43e-4d76-4484-8b92-a25d7b2ded35 · outbound

This paper cites A V A: A Large-Scale Database for Aesthetic Visual Analysis.

On the rankability of visual embeddings A V A: A Large-Scale Database for Aesthetic Visual Analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.516129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.516129Z digest=sha256:4fdf144b232f70feb0590fa204037caa804e405655d72a3ed3944de3fb9c8170

Observation 55c49fff-8449-4320-8028-83eba01abd3f · outbound

This paper cites A metric learning reality check.

On the rankability of visual embeddings A metric learning reality check

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.441046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.520788Z digest=sha256:e5405cc49ae351ea036a310697420d610fd027cd7ef27b28f1dca673c4c189f3

Observation f8ee04c5-2aa8-46e1-be7d-9572b7fa1707 · outbound

This paper cites Parts of Speech-Grounded Subspaces in Vision-Language Models.

On the rankability of visual embeddings Parts of Speech-Grounded Subspaces in Vision-Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.533367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.526319Z digest=sha256:be1cb942df59676bd02ba82d0efad1b9042983e0482f0331a0b842a475222346

Observation 18e7cfd6-4087-4ef9-beec-b45a1a211e6a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

On the rankability of visual embeddings DINOv2: Learning Robust Visual Features without Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.531605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.531605Z digest=sha256:000f5e0d449598092d504c7d7cd25c3d9dce0b7163ff3165d75e88694fb271c7

Observation bd3a1569-da58-45a3-826f-5ab8b6110eaf · outbound

This paper cites Teaching CLIP to Count to Ten.

On the rankability of visual embeddings Teaching CLIP to Count to Ten

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.537121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.537121Z digest=sha256:a65d31f4ed299a9f134101be4e0b466a6dc20df941ddbed4dcd0e37a43cba738

Observation 75fa21f4-3934-4458-8b3d-b3fd6aaebe20 · outbound

This paper cites Dating Historical Color Images.

On the rankability of visual embeddings Dating Historical Color Images

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.542600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.542600Z digest=sha256:25d55439eec321389a29148f5adc7c96892952306ca8bb4a163864419110d1a8

Observation 5e8fa665-8c00-4a0b-b22f-f10365a35b98 · outbound

This paper cites Relative Attributes.

On the rankability of visual embeddings Relative Attributes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.548487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.548487Z digest=sha256:7ac019494eb1d2d0711bf7507cef18eafd450e2412206b182bd9921f4bdd8dc8

Observation 0babb719-b950-468c-a839-5b56f2279032 · outbound

This paper cites HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports.

On the rankability of visual embeddings HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.422955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.553683Z digest=sha256:2278d0040e57375fe8a2c15dcc63dad97e8cd2b319ee43fa43c64e68a9f89663

Observation a0e88aef-b4aa-4aff-90ec-4a7203a7dd5e · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.559058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.559058Z digest=sha256:631c42aa7ce1b9284756c056e9b8a8a92acd487958de02d29546a9d2fb04a5f6

Observation 025ec07b-d7e3-459d-99a9-9abf9a1a44e9 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Super- vision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.404122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.564088Z digest=sha256:7392dda1f7c9576fc743636da8b41459207cececfde2820ed4a681fe9f789411

Observation e2d4c7c3-cb2a-4eed-bf32-218603569642 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

On the rankability of visual embeddings Steering Llama 2 via Contrastive Activation Addition

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.569573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.569573Z digest=sha256:7bff404226d989935895f2a526eae9b288ce86739f2e37447b605c8c8ea14730

Observation fe3b8e31-c769-4e73-ae32-dd0e949665a4 · outbound

This paper cites Finetuning CLIP to Reason about Pairwise Differences.

On the rankability of visual embeddings Finetuning CLIP to Reason about Pairwise Differences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.575686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.575686Z digest=sha256:c2dffa68915a135c88c7fbf4ed6169d9c3304bdae69ac6c10911ddbdd101d2a7

Observation 3fd03dbe-f488-4d9f-be8d-8c43e02a1b47 · outbound

This paper cites Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification.

On the rankability of visual embeddings Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification

Reference 49

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T20:11:52.391746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.583085Z digest=sha256:cda99998c5d53d4299e4b40cea5e73213832e391bc8e4d87f12e7d330a1cb6a1

Observation bae4703a-499c-46ee-8035-7a7ab5c85cd7 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.386329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.588973Z digest=sha256:85c7240ffff2d6de6ef93d40f7c622c8010fd7cf37979ee4bd246a28f2a70867

Observation 3238f178-81da-4011-99c2-d6fe4be79ac2 · outbound

This paper cites Linear Spaces of Meanings: Compositional Structures in Vision- Language Models.

On the rankability of visual embeddings Linear Spaces of Meanings: Compositional Structures in Vision- Language Models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.368075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.601038Z digest=sha256:739f28e8ca1c319fb3e84001ac3d9693170760e7b363b2978e37dc4cce22c4b5

Observation 1b0cdd12-01f9-4d10-9361-19f08ebb1ee4 · outbound

This paper cites Deep image prior.

On the rankability of visual embeddings Deep image prior

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.348792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.607467Z digest=sha256:e23775330357425894d7d04cb2b64c0a0cd82ebc3f23f224133b43677c6bbfde

Observation e4fe7ec6-2185-47a2-af26-0d7a5a3c787e · outbound

This paper cites BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION.

On the rankability of visual embeddings BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.330263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.612865Z digest=sha256:f135e0ea27aba7d37443123de40a3aa6cab366e36efbd698d58512a4df2a8d25

Observation 4230b46f-7018-4a08-8e07-12086fe5af61 · outbound

This paper cites Intermediate Layer Classifiers for OOD Generalization.

On the rankability of visual embeddings Intermediate Layer Classifiers for OOD Generalization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.308628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.617981Z digest=sha256:0f2b371f256eb28307e73e02b02dcd11832803390be08ddb69f690aee07e7c95

Observation 9247635d-3ab4-46b7-847f-bf53ff49548e · outbound

This paper cites Order-Embeddings of Images and Language.

On the rankability of visual embeddings Order-Embeddings of Images and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.623271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.623271Z digest=sha256:5e5b99159454850d38d5b8eb6abd3a60e257fad570c31597c35a84aae6e6fee3

Observation 6b5e8000-421a-459c-a408-8b76f589aed6 · outbound

This paper cites Exploring CLIP for Assessing the Look and Feel of Images.

On the rankability of visual embeddings Exploring CLIP for Assessing the Look and Feel of Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.628732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.628732Z digest=sha256:6c1c9704579ee6961dd3a496da0e9a4938d3e00690be16339860e646d801313c

Observation 936ac420-362e-4fbf-8e94-e91a8469f375 · outbound

This paper cites Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification.

On the rankability of visual embeddings Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.290182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.634171Z digest=sha256:130d225425546682fde58725c86e9e17bb09dbde5bccca0977d58ca11fbdc38d

Observation e4732314-04f8-4044-ba8f-e92acfb33c58 · outbound

This paper cites Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification.

On the rankability of visual embeddings Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.271360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.640208Z digest=sha256:c5e2e8ad9ea663a4968c41a26a1996f5876f81e1d096fa20c22c913259653d00

Observation be34882b-7800-4c43-bc56-e814d8e5511c · outbound

This paper cites Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere.

On the rankability of visual embeddings Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.646034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.646034Z digest=sha256:a316c2740384b009e079ae3284acc6a7d2bc11e0de1c04d0e801e1f63f041b77

Observation 6f1bb15f-0122-4804-9671-9633a12c7e06 · outbound

This paper cites Disentangled representation learning.

On the rankability of visual embeddings Disentangled representation learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.252755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.651508Z digest=sha256:3982c1cb542ed20cb49af4a2cdde03740f9abef4cec5fb808878c8007979477a

Observation f0ff0cf9-3c50-4858-9037-a99cb6c6d5d9 · outbound

This paper cites ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders.

On the rankability of visual embeddings ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.656002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.656002Z digest=sha256:270734e8e3be0df14272b23c930841c0d97c36746e4fb7fe70b7d361b99790b7

Observation c36c12cd-79c4-42e7-a276-151ea491f4ee · outbound

This paper cites CLIP Brings Better Features to Visual Aesthetics Learners.

On the rankability of visual embeddings CLIP Brings Better Features to Visual Aesthetics Learners

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.250347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.661258Z digest=sha256:abbbac8fa2fdfc697e193fe7622f5182db73091bc41f3f91b00af30296a8056f

Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.667082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.667082Z digest=sha256:20963230d3e1390a8aaec92fc47bee0796f0ad89da0131f2aedbd7719a4bcf4b

Observation 01c94fe3-87ce-4cef-8ba2-51249241ef8a · outbound

This paper cites Just noticeable differences in visual attributes.

On the rankability of visual embeddings Just noticeable differences in visual attributes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.233523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.673658Z digest=sha256:bfaaa829108babc44995f4be7288d93018213a386b805db202959b0ad67d2ed7

Observation 82ad9c1e-42a9-416c-b056-06aac0007fb0 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

On the rankability of visual embeddings CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.216137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.679214Z digest=sha256:29c0cc0717da6302f8575c2ccb15ccd7c58f128f66943acdcbcee9134cc4451a

Observation db9bfc04-6741-4d31-9d31-7c52aab8b687 · outbound

This paper cites RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP.

On the rankability of visual embeddings RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.200041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.684176Z digest=sha256:d7d01b6fae923949166179db6054c91fb4a789ad0e935b0dbee9796bbdd57067

Observation 3f14a96a-9b7b-416c-97c0-65bfe82f4afc · outbound

This paper cites Ranking-aware adapter for text-driven image ordering with CLIP.

On the rankability of visual embeddings Ranking-aware adapter for text-driven image ordering with CLIP

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.801956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.689502Z digest=sha256:38afabccba70189461ec59e2355536da2a330bbf0883f3b3b55d1de50b519c3a

Observation 3ed8fddf-ba6d-48cb-8c12-f3952ea698e1 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

On the rankability of visual embeddings When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.695389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.695389Z digest=sha256:c906e2efc10ef347d5a475131e1f7a999ec42da38ee2e25d15b44387a1862199

Observation 68d75df2-e218-4e43-b214-cd4e2f6db13d · outbound

This paper cites Single-Image Crowd Counting via Multi-Column Convolutional Neural Network.

On the rankability of visual embeddings Single-Image Crowd Counting via Multi-Column Convolutional Neural Network

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.701459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.701459Z digest=sha256:7448ae51e7e7603ab06430f8c8ea5defb10567479f6b8dbcb8932c8e7ee135b4

Observation ef6fcadf-a217-4378-ae55-e1e091615786 · outbound

This paper cites an unresolved cited work.

On the rankability of visual embeddings Unresolved cited work

Reference 70

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:52.187722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.706634Z digest=sha256:ffcab500920245d513d4aecc7bfef4860fcad1f6d78a13280d9c775e63262b48

Observation e330120c-f2ee-478c-809d-c67b5537b483 · outbound

This paper cites Learning Ordinal Relationships for Mid-Level Vision.

On the rankability of visual embeddings Learning Ordinal Relationships for Mid-Level Vision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.183574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.717772Z digest=sha256:b75f1d959b2eb2301685da22975c570339cfac9f4c9c85775c42f94fe5b2c3ec

Observation 73c93712-e3f5-4205-97ab-3869d8d60f4d · outbound

This paper cites Age Progression/Regression by Conditional Adversarial Autoencoder.

On the rankability of visual embeddings Age Progression/Regression by Conditional Adversarial Autoencoder

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.767240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T20:11:51.712243Z digest=sha256:32230e7086f1d1370ef0b65e95a5434abc6c328a393dcef905082f2526d211d5

Observation ec1d0391-a4a0-4c85-84d5-c2f4ae582f49 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.594734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.594734Z digest=sha256:87595273a008e6983bc4ff9142a83d82c4da984c3b26fc4cf700a5824bc68a69

Observation 7d03c2f3-6c39-41dc-89c0-7120d1ade950 · outbound

This paper cites DOI: 10.1109/TIP.2020.2967829.

On the rankability of visual embeddings DOI: 10.1109/TIP.2020.2967829

Reference 4056

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.176194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.176194Z digest=sha256:1cb5eb1d77e2b9b2dc5569f6d4ac2059e636b9ea5d13ce676dda9355b924939c

Pith citing papers

No inbound Pith citation observations are available.