Pith. sign in

Paper Citation Record · LEDGER

On the rankability of visual embeddings

As of 10 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2507.03683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03683 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:11:51.717772Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact11
  • verified fuzzy29
  • unresolved29
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72fcb7d1-1878-4a5b-987a-7a8956bbd1b7 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

On the rankability of visual embeddings Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:48.623426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:48.623426Z digest=sha256:e586edf5d6bc5f68d7d543d8ce1dbdc3ba71445b08a9d274399b1248c51ba5f3

Observation 72f8616e-862e-4547-92c6-4f8ae4d82f13 · outbound

This paper cites Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features.

On the rankability of visual embeddings Conditioned and composed image retrieval combining and partially fine-tuning CLIP-based features

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.679728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:48.662843Z digest=sha256:d18101b0698c85e9f4ee222f624ac1219ebc40826b3f5cd943b727e5f7c5c7d2

Observation 671d080d-4f91-48ef-b91b-bc3ff30dd0f1 · outbound

This paper cites Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models.

On the rankability of visual embeddings Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.140579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:48.742002Z digest=sha256:6e98d83b0ff5b691f8b428c71f8c6dfaebdb3537b617032e23fc5a657c939921

Observation 82b92998-93e2-45a5-9a65-d480c670884b · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

On the rankability of visual embeddings A Simple Framework for Contrastive Learning of Visual Representations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.663979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:48.803202Z digest=sha256:6e1e5286f6dacb3aa55564a490a21587ea47060a68352dcd9902249a8a698762

Observation 511cf3cc-d07f-4c12-a0ae-b1d11993093b · outbound

This paper cites Deep Learning for Instance Retrieval: A Survey.

On the rankability of visual embeddings Deep Learning for Instance Retrieval: A Survey

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.647220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:48.894792Z digest=sha256:9f038d1fb9a6e735516d220324c038f8a203a86cce4932cc2ff062e21a5d90e5

Observation 0208f5c9-189e-4893-a53b-1e5093e7b4a4 · outbound

This paper cites Deep learning for instance retrieval: A survey.

On the rankability of visual embeddings Deep learning for instance retrieval: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.630656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:48.981497Z digest=sha256:673f90d4190b01b014f1f6c51d1b499529ea41ac956eadd4835a846b4c5eea9c

Observation af97d169-f7c6-4cc0-8ed5-efa72c98a39c · outbound

This paper cites Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds.

On the rankability of visual embeddings Composition Loss for Counting, Density Map Estimation and Localization in Dense Crowds

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:49.072090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.072090Z digest=sha256:995c0483baaddf345b633e76ba0ecfa830e0405fb0548aabc671870f2351573b

Observation 94a871ac-e6e3-41c2-b3b3-06ddd5a78bde · outbound

This paper cites Hyperbolic Image-Text Representations.

On the rankability of visual embeddings Hyperbolic Image-Text Representations

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:53.112878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:49.128084Z digest=sha256:9cd9a19547310f3d4c647a056281ed0b43768ed05a8a084b7642061760b7cd68

Observation 2d825a8c-c3ea-4a01-a78d-877669bf5ce3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

On the rankability of visual embeddings An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.229769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.229769Z digest=sha256:8e614b1e801ef71ad05982adf6629ef0c2f0c2bdb709cabd321bb90260ba099f

Observation b6cc826c-934f-489b-94e1-f44b9422ddc1 · outbound

This paper cites Teach CLIP to Develop a Number Sense for Ordinal Regression.

On the rankability of visual embeddings Teach CLIP to Develop a Number Sense for Ordinal Regression

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.059438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:49.329152Z digest=sha256:a54171bed0d108537b09b14aa3c35b1a85a0b20594abe03099b79fce8409c5ee

Observation 8883acf4-b68b-4343-903e-8706086d6021 · outbound

This paper cites Age and Gender Estimation of Unfiltered Faces.

On the rankability of visual embeddings Age and Gender Estimation of Unfiltered Faces

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:53.088420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:49.450137Z digest=sha256:fd09875609163206dfd60a5a85398206859a669751daa5281ff715fe58469513

Observation 9cfaf071-0d3c-4582-83b4-b5dd8d280749 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

On the rankability of visual embeddings It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.531103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.531103Z digest=sha256:b8de18fbac05a089bf0edd277c62e262d1568caa0284d212c5c63bb5b6fbc758

Observation 9dd7610a-3d37-4134-a268-e1343c54c0f3 · outbound

This paper cites Heterogeneous face attribute estimation: A deep multi-task learning approach.

On the rankability of visual embeddings Heterogeneous face attribute estimation: A deep multi-task learning approach

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.613502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:49.721697Z digest=sha256:95c74632e7d3ffc1ca3e4691e6e6bd00cad8e7ba244154a6f3aef6c47e1fe761

Observation 477da1c2-50ff-4f16-8b1d-05c3201567d5 · outbound

This paper cites Deep Residual Learning for Image Recognition.

On the rankability of visual embeddings Deep Residual Learning for Image Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:49.805256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:49.805256Z digest=sha256:8a089100446a4d5262cdd04a38ce933eb316b851e8544032fe79f6c8830c68d2

Observation 307ea798-7ea6-4281-aa6e-2432bac0103f · outbound

This paper cites Multilayer feedforward networks are universal approximators.

On the rankability of visual embeddings Multilayer feedforward networks are universal approximators

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.597532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:49.928316Z digest=sha256:efda763b53f8ea31fa56d313ae5fba4aef779d5325a32ba3f9982cc22375461f

Observation bb2b5fda-2abe-47f1-b5e9-2c4c8f7635f0 · outbound

This paper cites KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment.

On the rankability of visual embeddings KonIQ-10k: An Ecologically Valid Database for Deep Learning of Blind Image Quality Assessment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.582027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:50.055806Z digest=sha256:1f2729b0cfc9eaf680e7371cbeb80c0cb3f12b3e22d405f53cf4cf9298581fea

Observation 97fb22fc-96d4-4760-9b18-e002309f4483 · outbound

This paper cites Lp++: A surprisingly strong linear probe for few-shot clip.

On the rankability of visual embeddings Lp++: A surprisingly strong linear probe for few-shot clip

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.566871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:50.309455Z digest=sha256:98042344eb3320f2c51bcc7a2d87f639895e263764881f73d725f50c3d4eb844

Observation 9d9125c5-7393-4fdd-90fd-9213f230accf · outbound

This paper cites The Platonic Representation Hypothesis.

On the rankability of visual embeddings The Platonic Representation Hypothesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.427501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.427501Z digest=sha256:837a8965343f1ff86a9bdf76bf2af011c34ee94e713f5db14795efe409699464

Observation 1e02910c-73b7-4adf-a889-8aa29d70c8f0 · outbound

This paper cites CLIP-Count: Towards Text-Guided Zero- Shot Object Counting.

On the rankability of visual embeddings CLIP-Count: Towards Text-Guided Zero- Shot Object Counting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.544893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.544893Z digest=sha256:c93511d0ceb4029fcb82954eaa417bb62f44269750414966a54e99dc0f048e5b

Observation 22adfc6d-f764-4110-8eea-9146e82b38b7 · outbound

This paper cites Hyperbolic Image Embeddings.

On the rankability of visual embeddings Hyperbolic Image Embeddings

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.549000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:50.625607Z digest=sha256:abda77660712f561b37254ef7df2a96a65211fc9d3286c5b1ef6ad9de2885df3

Observation 86d92a71-4aea-4f83-ad09-41ba0cd792a0 · outbound

This paper cites MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs.

On the rankability of visual embeddings MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.693133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.693133Z digest=sha256:8f6317c9d8815eb924370331d50fd9f4cfcf07a1867f71b30000e31242218d2d

Observation 4c0d978c-48c2-461d-9110-e11b77b6153b · outbound

This paper cites Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav).

On the rankability of visual embeddings Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.529295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:50.813651Z digest=sha256:73c7e294f74559f6d6503faed2b979aeac55bca2004a824c9be6106d44aaae18

Observation d74f666a-fe1f-4221-be11-5cb58b0a6fd4 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but Not Uni-modally

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.960886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.960886Z digest=sha256:8b728b658a927026a0e44e00fd116c2e91e0c5506c5373eb4ad4b7d31241630a

Observation de4d0571-cc07-4331-a0c2-034a6b53f7b1 · outbound

This paper cites CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally.

On the rankability of visual embeddings CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.002032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.002032Z digest=sha256:d4ff88b06092d6bc6885d3588b6b2d2227107375f1f99603781f3f39e4ad9e1e

Observation d160e3ef-edc5-4886-b0e0-5c3704ab4ed2 · outbound

This paper cites Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning.

On the rankability of visual embeddings Beyond a Pre-Trained Object Detector: Cross-Modal Textual and Visual Context for Image Captioning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.512741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.033881Z digest=sha256:742ec4c92b7b36dfafe20e577dc2bfd243545455e9f272afd409f00686fa0a4f

Observation c5ada6c4-84e0-41ad-b510-1f10d8b09f98 · outbound

This paper cites MiVOLO: Multi-input Transformer for Age and Gender Estimation.

On the rankability of visual embeddings MiVOLO: Multi-input Transformer for Age and Gender Estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.039704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.039704Z digest=sha256:f9b1b434c2c50b8dbddb4f1d25cdb1337d026649b004df6426daece0b7891557

Observation 9e393176-8412-4a5a-816e-ea283721eaac · outbound

This paper cites The Double-Ellipsoid Geometry of CLIP.

On the rankability of visual embeddings The Double-Ellipsoid Geometry of CLIP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.045027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.045027Z digest=sha256:852609e9df97ee67c3c2666514342f5a24872aa01b39add77fddc7f4de0452ba

Observation 4d502446-6440-47de-b3a7-37f7735ac5f9 · outbound

This paper cites Does CLIP Bind Concepts? Probing Compositionality in Large Image Models.

On the rankability of visual embeddings Does CLIP Bind Concepts? Probing Compositionality in Large Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.071811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.071811Z digest=sha256:f4797b535f16321319873a9cfdd36fed80ad7c0a32b46c30df9d229243fd821d

Observation 2624fd97-d6cf-4533-b4f8-02e1e764d3dd · outbound

This paper cites Align before fuse: Vision and language representation learning with mo- mentum distillation.

On the rankability of visual embeddings Align before fuse: Vision and language representation learning with mo- mentum distillation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.494930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.171974Z digest=sha256:ad3048392f0c954f62ef7a9d407c458200bad2ceb905fd67ec7a4df35ca4dff5

Observation bf413c30-217d-4123-95e4-6fe4efb58ffe · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation.

On the rankability of visual embeddings BLIP: Bootstrapping Language-Image Pre-training for Unified Vision- Language Understanding and Generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.478419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.232772Z digest=sha256:f6b93ea5ae04a391f2201099160119bea0a06e9c23fc740112212cca30671e43

Observation f366b8f4-65fd-4e5b-8ae5-52c09621ca3b · outbound

This paper cites OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression.

On the rankability of visual embeddings OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.956760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.302801Z digest=sha256:432b88522ed68d0f3a2375e8bd7a0f4997fc1bcd33b94d07acb8c41aab9de477

Observation 00efac1b-a70d-4467-b8f4-0ba5c81b0c6d · outbound

This paper cites CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model.

On the rankability of visual embeddings CrowdCLIP: Unsupervised Crowd Counting via Vision-Language Model

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.933505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.386439Z digest=sha256:a1323dc7b1a8e7b79b73fcbb7d8d9e70e24075901af8724585c819698f57d7ed

Observation ffb1b424-5803-447c-983e-274b6ec42009 · outbound

This paper cites Beyond comparing image pairs: Setwise active learning for relative attributes.

On the rankability of visual embeddings Beyond comparing image pairs: Setwise active learning for relative attributes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.462867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.467629Z digest=sha256:f9b81c7196e027890b5cc543f7d787ea77af72049540ecae56b4dcd0076ef700

Observation 9cdd71ea-a5ce-4033-93fd-6d02518fd554 · outbound

This paper cites CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification.

On the rankability of visual embeddings CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.500131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.500131Z digest=sha256:129bfaad6c8a61ceb0176e6559798b9446e02365c688889564786ff909978843

Observation ca8ccce3-3a66-418b-96d1-4e4cd7f48a02 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

On the rankability of visual embeddings ClipCap: CLIP Prefix for Image Captioning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.511161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.511161Z digest=sha256:89b05099011a786ca7635d9c3532428f0efc3a68dbbeec842c7ddcfaebce83c5

Observation ead6e43e-4d76-4484-8b92-a25d7b2ded35 · outbound

This paper cites A V A: A Large-Scale Database for Aesthetic Visual Analysis.

On the rankability of visual embeddings A V A: A Large-Scale Database for Aesthetic Visual Analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.516129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.516129Z digest=sha256:3964dccadd1645197c72a243e7b566689f88cf1d161960a840fdb00e70b87a71

Observation 55c49fff-8449-4320-8028-83eba01abd3f · outbound

This paper cites A metric learning reality check.

On the rankability of visual embeddings A metric learning reality check

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.441046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.520788Z digest=sha256:e4c44a6c53c1eed9c073c11e0e7a8bcf78d85aadbc64c474eaa772cc8fd93f39

Observation f8ee04c5-2aa8-46e1-be7d-9572b7fa1707 · outbound

This paper cites Parts of Speech-Grounded Subspaces in Vision-Language Models.

On the rankability of visual embeddings Parts of Speech-Grounded Subspaces in Vision-Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.533367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.526319Z digest=sha256:58f2e9ac58bbd2089c822503324b80a082cb386adc313323e2f8672a5793750c

Observation 18e7cfd6-4087-4ef9-beec-b45a1a211e6a · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

On the rankability of visual embeddings DINOv2: Learning Robust Visual Features without Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.531605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.531605Z digest=sha256:135366c2166a332ad34f8265fca062b1510e6827065e2e10b6938d077c3dc77c

Observation bd3a1569-da58-45a3-826f-5ab8b6110eaf · outbound

This paper cites Teaching CLIP to Count to Ten.

On the rankability of visual embeddings Teaching CLIP to Count to Ten

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.537121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.537121Z digest=sha256:c989fa2406c19355928b1376499577cd8b9b18301320b2516983326db47ff1c7

Observation 75fa21f4-3934-4458-8b3d-b3fd6aaebe20 · outbound

This paper cites Dating Historical Color Images.

On the rankability of visual embeddings Dating Historical Color Images

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.542600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.542600Z digest=sha256:a0538da1c891eaa434962accbadd03d09697e1499d16c873381fbe6efbb4f3b6

Observation 5e8fa665-8c00-4a0b-b22f-f10365a35b98 · outbound

This paper cites Relative Attributes.

On the rankability of visual embeddings Relative Attributes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.548487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.548487Z digest=sha256:9f6a2230dea28313d02343ac33211b429795b131c353d8b87f99e76018b699e4

Observation 0babb719-b950-468c-a839-5b56f2279032 · outbound

This paper cites HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports.

On the rankability of visual embeddings HYDEN: Hyperbolic Density Representations for Medical Images and Re- ports

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.422955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.553683Z digest=sha256:bd21ad6dfa48a39a6bbf0af28ac58cc65f42239dbc55e7730d0ffc507f2f27b2

Observation a0e88aef-b4aa-4aff-90ec-4a7203a7dd5e · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.559058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.559058Z digest=sha256:62b13796e188f4f017a20bdbd68e4c03b1b989e28fabbac3fa8849ed3bdfae6d

Observation 025ec07b-d7e3-459d-99a9-9abf9a1a44e9 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Super- vision.

On the rankability of visual embeddings Learning Transferable Visual Models From Natural Language Super- vision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.404122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.564088Z digest=sha256:db52a5685ce90f6619de941f35bc87e96aa715203b391896a2efc8797bb5fe63

Observation e2d4c7c3-cb2a-4eed-bf32-218603569642 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

On the rankability of visual embeddings Steering Llama 2 via Contrastive Activation Addition

Reference 47

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:11:51.569573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.569573Z digest=sha256:91f2b2c9e37108e487d0177df18a0a96d7ddfd55f34494eb9e0041f1ffc8f324

Observation fe3b8e31-c769-4e73-ae32-dd0e949665a4 · outbound

This paper cites Finetuning CLIP to Reason about Pairwise Differences.

On the rankability of visual embeddings Finetuning CLIP to Reason about Pairwise Differences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.575686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.575686Z digest=sha256:5cadbdc4fd59bca1c933714ffad178148e6dbdb75bd6be1df9eb23e4987fc2c6

Observation 3fd03dbe-f488-4d9f-be8d-8c43e02a1b47 · outbound

This paper cites Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification.

On the rankability of visual embeddings Improving Image Encoders for General-Purpose Nearest Neighbor Search and Classification

Reference 49

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T20:11:52.391746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.583085Z digest=sha256:23c1f6ef28b6a17a92bb8438c73952ba76ecc2166a24180e68472542e143d460

Observation bae4703a-499c-46ee-8035-7a7ab5c85cd7 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.386329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.588973Z digest=sha256:ac2d95b90a60c1f8f6744c609b5cc3588b991f5aa2c3d8eb4775e07bccef0817

Observation 3238f178-81da-4011-99c2-d6fe4be79ac2 · outbound

This paper cites Linear Spaces of Meanings: Compositional Structures in Vision- Language Models.

On the rankability of visual embeddings Linear Spaces of Meanings: Compositional Structures in Vision- Language Models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.368075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.601038Z digest=sha256:53de7554278a7ac2791d2ea50555cf73a548397b4c961341fd3ec133596f4e4e

Observation 1b0cdd12-01f9-4d10-9361-19f08ebb1ee4 · outbound

This paper cites Deep image prior.

On the rankability of visual embeddings Deep image prior

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.348792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.607467Z digest=sha256:1c70074a61460d573b2b741187dcf1b633c52649ad82cb74c3a4a51af6e628a8

Observation e4fe7ec6-2185-47a2-af26-0d7a5a3c787e · outbound

This paper cites BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION.

On the rankability of visual embeddings BEYOND DECODABILITY: LIN- EAR FEATURE SPACES ENABLE VISUAL COMPOSITIONAL GENERALIZATION

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.330263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.612865Z digest=sha256:7f25abff6b45f27e635d17b01f0032db41ee6b4a40bc81545d45ec00c06d8770

Observation 4230b46f-7018-4a08-8e07-12086fe5af61 · outbound

This paper cites Intermediate Layer Classifiers for OOD Generalization.

On the rankability of visual embeddings Intermediate Layer Classifiers for OOD Generalization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.308628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.617981Z digest=sha256:a636a171c9cbf6f51c1cf7c8fb4722cfb4654b3ff53588774b4088fface8f698

Observation 9247635d-3ab4-46b7-847f-bf53ff49548e · outbound

This paper cites Order-Embeddings of Images and Language.

On the rankability of visual embeddings Order-Embeddings of Images and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.623271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.623271Z digest=sha256:a45bb38257dd79f15870aac252c0035023d74b0f23ead4631a8ebe7fa6471cab

Observation 6b5e8000-421a-459c-a408-8b76f589aed6 · outbound

This paper cites Exploring CLIP for Assessing the Look and Feel of Images.

On the rankability of visual embeddings Exploring CLIP for Assessing the Look and Feel of Images

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.628732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.628732Z digest=sha256:d8408ee9d0d3c6e6d7de149971bc1498cbf898b673118ebc48cea3ee036e1e86

Observation 936ac420-362e-4fbf-8e94-e91a8469f375 · outbound

This paper cites Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification.

On the rankability of visual embeddings Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.290182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.634171Z digest=sha256:d402c497a1039dd3a4373f49f56714d4746af293709c9cd7688eb038996f17a2

Observation e4732314-04f8-4044-ba8f-e92acfb33c58 · outbound

This paper cites Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification.

On the rankability of visual embeddings Learning-to-rank meets language: Boosting language-driven ordering align- ment for ordinal classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.271360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.640208Z digest=sha256:2e8a5ac417624bf1e4f2102490bcc5381d3495117a44bf79be85c5c92c6175f6

Observation be34882b-7800-4c43-bc56-e814d8e5511c · outbound

This paper cites Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere.

On the rankability of visual embeddings Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.646034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.646034Z digest=sha256:0c34005421b3afb122f6fff0e61325c7b7dd23cec230dccbc9a20a7dcb3af552

Observation 6f1bb15f-0122-4804-9671-9633a12c7e06 · outbound

This paper cites Disentangled representation learning.

On the rankability of visual embeddings Disentangled representation learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.252755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.651508Z digest=sha256:4af8fccbdfd708f72b895d43d6b779ce16890980a6f33b50b6a845b2fcfb280e

Observation f0ff0cf9-3c50-4858-9037-a99cb6c6d5d9 · outbound

This paper cites ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders.

On the rankability of visual embeddings ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.656002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.656002Z digest=sha256:2928f24fc833fccc01b9f764e41b9a70dda5c75749fc27cc8d49aa4ab065dc79

Observation c36c12cd-79c4-42e7-a276-151ea491f4ee · outbound

This paper cites CLIP Brings Better Features to Visual Aesthetics Learners.

On the rankability of visual embeddings CLIP Brings Better Features to Visual Aesthetics Learners

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:52.250347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.661258Z digest=sha256:aee0cc929f5e1b42b2e57fdd0a8303ea29dd405789c721cff82bfc7793b52015

Observation 2dd22aa8-820d-4c84-b04a-02be321730dd · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

On the rankability of visual embeddings FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.667082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.667082Z digest=sha256:f7c9494713bfa09a5478de6ebf9adcbd70e6cccb20abf4aee26b5b38167cf94e

Observation 01c94fe3-87ce-4cef-8ba2-51249241ef8a · outbound

This paper cites Just noticeable differences in visual attributes.

On the rankability of visual embeddings Just noticeable differences in visual attributes

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.233523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.673658Z digest=sha256:12e1571912f47ec75b6377cadc8c7bb7f38ce47969ba5853a61fde638815e739

Observation 82ad9c1e-42a9-416c-b056-06aac0007fb0 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

On the rankability of visual embeddings CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.216137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.679214Z digest=sha256:fde074609e3674b050c701fb3b02021865bb53dac13861b157c3a7beb3cf6952

Observation db9bfc04-6741-4d31-9d31-7c52aab8b687 · outbound

This paper cites RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP.

On the rankability of visual embeddings RANKING-AWARE ADAPTER FOR TEXT-DRIVEN IMAGE OR- DERING WITH CLIP

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.200041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.684176Z digest=sha256:3a377e30f8477b838112c05f3933e374da501642c0999e70642410c7dd9fa7c9

Observation 3f14a96a-9b7b-416c-97c0-65bfe82f4afc · outbound

This paper cites Ranking-aware adapter for text-driven image ordering with CLIP.

On the rankability of visual embeddings Ranking-aware adapter for text-driven image ordering with CLIP

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.801956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.689502Z digest=sha256:60a8c64b922b22c367f88cef45bc971f2736b9832f90d23ea9aa7f79288c9cfb

Observation 3ed8fddf-ba6d-48cb-8c12-f3952ea698e1 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

On the rankability of visual embeddings When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.695389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.695389Z digest=sha256:615b60f6c02f2e34014c7367e45cd9bc8b4a287f8b05372c25b0837232bb6826

Observation 68d75df2-e218-4e43-b214-cd4e2f6db13d · outbound

This paper cites Single-Image Crowd Counting via Multi-Column Convolutional Neural Network.

On the rankability of visual embeddings Single-Image Crowd Counting via Multi-Column Convolutional Neural Network

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.701459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.701459Z digest=sha256:f3b683ff22db690ed501e4293b08fc45b8dcc9ccb219b828e5cca4afd8c46596

Observation ef6fcadf-a217-4378-ae55-e1e091615786 · outbound

This paper cites an unresolved cited work.

On the rankability of visual embeddings Unresolved cited work

Reference 70

Resolution
verified exact
raw_fallback, observed 2026-08-06T20:11:52.187722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.706634Z digest=sha256:65f1b4582fcf396bec6d02e28d32463fd4563290df75cd47d9d4b04fedc33374

Observation e330120c-f2ee-478c-809d-c67b5537b483 · outbound

This paper cites Learning Ordinal Relationships for Mid-Level Vision.

On the rankability of visual embeddings Learning Ordinal Relationships for Mid-Level Vision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:11:53.183574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.717772Z digest=sha256:6dc046528e16fec70dfacfa03e412c1e35880912737eb71b8941b99a48351ebd

Observation 73c93712-e3f5-4205-97ab-3869d8d60f4d · outbound

This paper cites Age Progression/Regression by Conditional Adversarial Autoencoder.

On the rankability of visual embeddings Age Progression/Regression by Conditional Adversarial Autoencoder

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:11:51.767240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:11:51.712243Z digest=sha256:f5cd7136dd484fd81ad524b0917983210dbd1bcbb75da7350824b4bff8de28c9

Observation ec1d0391-a4a0-4c85-84d5-c2f4ae582f49 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

On the rankability of visual embeddings Linear Representations of Sentiment in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.594734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.594734Z digest=sha256:149344af6c9d182ca27fca3b7aebb6d81da1634f9cc8f9a34482aeebc1ce2d21

Observation 7d03c2f3-6c39-41dc-89c0-7120d1ade950 · outbound

This paper cites DOI: 10.1109/TIP.2020.2967829.

On the rankability of visual embeddings DOI: 10.1109/TIP.2020.2967829

Reference 4056

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:50.176194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:50.176194Z digest=sha256:956bf973fa77fba4cfed62179748c725b0dfed985e6005a9eff3ca03726b1083

Pith citing papers

No inbound Pith citation observations are available.