Pith. sign in

Paper Citation Record · LEDGER

The Double-Ellipsoid Geometry of CLIP

As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 7 inbound Pith citation observations for arXiv:2411.14517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14517 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:27:08.913531Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:11:51.045027Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2e69f3b2-d031-4ee4-bdb6-532e7acfb594 · outbound

This paper cites write newline.

The Double-Ellipsoid Geometry of CLIP write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.692056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.692056Z digest=sha256:30239c3249ea14f96cfcd1451bc47ada246064ec3258036bb2c6c687c4438f24

Observation 3a3ce743-e330-4b24-9ea3-a0c5475141b9 · outbound

This paper cites Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding.

The Double-Ellipsoid Geometry of CLIP Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.552537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.696999Z digest=sha256:2af8cf8dbc759598ac6ced7cbe04bd155c1f167598f0fabc49e3c405e0d292c8

Observation 8450b623-64af-41eb-9e49-c0db6daa6e6e · outbound

This paper cites A Theoretical Analysis of Contrastive Unsupervised Representation Learning.

The Double-Ellipsoid Geometry of CLIP A Theoretical Analysis of Contrastive Unsupervised Representation Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.700684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.700684Z digest=sha256:c920c237a52f9bb8b859613826e141b2127611801529517922fd47d0f521cc28

Observation 441a6acf-f127-45c5-b805-63727d1bcbcf · outbound

This paper cites Grit-vlp: Grouped mini-batch sampling for efficient vision and language pre-training.

The Double-Ellipsoid Geometry of CLIP Grit-vlp: Grouped mini-batch sampling for efficient vision and language pre-training

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.542365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.705196Z digest=sha256:fca79dab3c99b1487eb6eee98baeacbc20c620ac63644dbc33d1af6c10d740b8

Observation 8da093b7-fcd1-4f8a-94b7-4d34386714f9 · outbound

This paper cites Mafa: Managing false negatives for vision-language pre-training.

The Double-Ellipsoid Geometry of CLIP Mafa: Managing false negatives for vision-language pre-training

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.531238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.708525Z digest=sha256:399e867ec5b48391a2ec31ddbbae58d16481fd42e6223456fa680471f560dafc

Observation f194b376-7684-4c4f-b993-172018eca4e5 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

The Double-Ellipsoid Geometry of CLIP Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.520727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.712059Z digest=sha256:94a91dbe89de5bd0abaeac668d19932a399b817e4d702cd932f6aaa333d95072

Observation 15d7b1fa-e38b-479e-9d0b-7bdcc41fde77 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

The Double-Ellipsoid Geometry of CLIP A simple framework for contrastive learning of visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.715475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.715475Z digest=sha256:c028130da4e6e914bcbb1314df13678271393928265d2ef00ea92eb8db26a9d0

Observation 82ae280a-d73c-41e1-9c54-c9f95123f773 · outbound

This paper cites Fine-grained Image Captioning with CLIP Reward.

The Double-Ellipsoid Geometry of CLIP Fine-grained Image Captioning with CLIP Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.719025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.719025Z digest=sha256:fd645a30414ea9961e35522884b8c50954f0a4587ecbea80448ec2be67b2733c

Observation d2415993-83b1-4ded-8681-44e71a5fd964 · outbound

This paper cites Debiased contrastive learning.

The Double-Ellipsoid Geometry of CLIP Debiased contrastive learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.502088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.722558Z digest=sha256:685523990230565c2930db8351f7c5ba0b977c5b9809baef9f677bb6a8fbb6af

Observation 15a06a48-4505-42e8-960e-787fc7ac50b1 · outbound

This paper cites Eccv caption: Correcting false negatives by collecting machine-and-human-verified image-caption associations for ms-coco.

The Double-Ellipsoid Geometry of CLIP Eccv caption: Correcting false negatives by collecting machine-and-human-verified image-caption associations for ms-coco

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.490912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.726139Z digest=sha256:6c4189675d6e0f9e307e9e60dfba79531902bb5f0367c3287b5905bda8a5fc34

Observation fa8d2de3-14c7-4db0-b694-8a37b0cbce0f · outbound

This paper cites Eldar and Alan V.

The Double-Ellipsoid Geometry of CLIP Eldar and Alan V

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.480059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.729900Z digest=sha256:9689dc665e3fbf8c188bf1461d8e0706c492a66c50aea270d674aa221cc5d0c8

Observation 985352e5-5476-4be7-8f86-7144a0f14e28 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

The Double-Ellipsoid Geometry of CLIP It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.733351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.733351Z digest=sha256:a3eee5e657b5d20f393bd13310549da1c4d05017d33154710ba486bbe989baa2

Observation a3affa10-9b1b-4eab-bb4c-8a37f79ab46c · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

The Double-Ellipsoid Geometry of CLIP Datacomp: In search of the next generation of multimodal datasets

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.469451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.737235Z digest=sha256:b060fac95b0d8b318ccfa026fb26a86e40fadea8cab3fd941bca9ff7207e0879

Observation d7c749ab-375c-48b7-a5dd-8ae3ee7bbb22 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

The Double-Ellipsoid Geometry of CLIP An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.740577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.740577Z digest=sha256:7260018495c4652840458d66b6ba9b611b4600d51db625000b41d98b1b645b99

Observation 1f1010a0-12ba-4911-8912-0e17bfa4c948 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

The Double-Ellipsoid Geometry of CLIP SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.744074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.744074Z digest=sha256:f058a0454757223fc73963c4a862455b1cc7239a4c9f33ead9aae45491b4b036

Observation c9cae812-ed5c-449a-a7c1-ef754b47dea6 · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

The Double-Ellipsoid Geometry of CLIP Audioclip: Extending clip to image, text and audio

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.458805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.747702Z digest=sha256:0fc56ed91086b3f671871ec7fb8542916dab960ce2f5f3d083df72e8a2609d93

Observation 026a5c38-ac85-4969-811f-20f7a01314f7 · outbound

This paper cites Proxedit: Improving tuning-free real image editing with proximal guidance.

The Double-Ellipsoid Geometry of CLIP Proxedit: Improving tuning-free real image editing with proximal guidance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.447304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.751124Z digest=sha256:e54fde54364f79a22c3acc32020e464d2048935b15ba58568040dd6ca9a8f0cb

Observation 359a2066-106d-4655-9038-55ac7475dc5e · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

The Double-Ellipsoid Geometry of CLIP Momentum contrast for unsupervised visual representation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.754774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.754774Z digest=sha256:bb9a5714936513a646934877f7a92bc64e0c6b7d903faafd096e93699d0a3b80

Observation 9bb68dc2-63d7-4ba5-9134-f0f50addab42 · outbound

This paper cites Open-vocabulary multi-label classification via multi-modal knowledge transfer.

The Double-Ellipsoid Geometry of CLIP Open-vocabulary multi-label classification via multi-modal knowledge transfer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.429653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.758531Z digest=sha256:c5f6ee0548b7de04cfc3c884920cacb799aa0b22e11957269591452898ad3c8b

Observation 9185dd05-1fcd-4b25-bac3-669db242fd80 · outbound

This paper cites Clip goes 3d: Leveraging prompt tuning for language grounded 3d recognition.

The Double-Ellipsoid Geometry of CLIP Clip goes 3d: Leveraging prompt tuning for language grounded 3d recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.418553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.762049Z digest=sha256:d1187f457365c1f7df9378874a98ae2e8dfbe55c130703b972c712b8a285d9e8

Observation 0bd0dae5-f0c7-4029-9d7a-a6384d233b9b · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

The Double-Ellipsoid Geometry of CLIP The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.765412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.765412Z digest=sha256:aede82f2e9e9ac3b5a8f026c24fc75202f2e8dafa54d042d366d10fe41f081ec

Observation df71112c-7046-449c-bda2-81eeab463582 · outbound

This paper cites Natural adversarial examples.

The Double-Ellipsoid Geometry of CLIP Natural adversarial examples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.768831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.768831Z digest=sha256:1091d8bfaed880131aa933e1cdcae161a75ad8b11fe35c820164ffc69a8f27c6

Observation b2ea9f17-f111-4014-8ca2-9b5e0529f250 · outbound

This paper cites Boosting contrastive self-supervised learning with false negative cancellation.

The Double-Ellipsoid Geometry of CLIP Boosting contrastive self-supervised learning with false negative cancellation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.394316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.772696Z digest=sha256:516086bcc0e5e0c1726712d24a94d72cf58a803d98f48b4574171849a8995bfd

Observation 59434c5c-02c4-4a45-842d-a43261a26660 · outbound

This paper cites A Slightly Improved Bound for the KLS Constant.

The Double-Ellipsoid Geometry of CLIP A Slightly Improved Bound for the KLS Constant

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.776197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.776197Z digest=sha256:e67d73038f16ec1a26419ce72b2242ce959fbefd49a38d85ee314c0cad284b59

Observation 9f37cd12-7d20-441f-b51c-7bfb46e4f61c · outbound

This paper cites The power of contrast for feature learning: A theoretical analysis.

The Double-Ellipsoid Geometry of CLIP The power of contrast for feature learning: A theoretical analysis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.383804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.780045Z digest=sha256:8b1ba3cb86c4113daa9639789e3f9d441b963dd6786346a897a5dbbe7a86cd2d

Observation a10cefdd-fdc2-4ded-8f89-b9778d31862d · outbound

This paper cites Hard negative mixing for contrastive learning.

The Double-Ellipsoid Geometry of CLIP Hard negative mixing for contrastive learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.372683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.783478Z digest=sha256:2562c5a27454affc8cb9c510ff83ecb1517494ac536d045f92d4092accfebbbb

Observation f4cfce5a-a4b6-4409-b94e-618a8f75b103 · outbound

This paper cites Isoperimetric problems for convex bodies and a localization lemma.

The Double-Ellipsoid Geometry of CLIP Isoperimetric problems for convex bodies and a localization lemma

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.788113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.788113Z digest=sha256:ecd6830b144cdcdc45b4ca85ed15fd1abaf813cedd45b792fbf68752aeb61b32

Observation 85ee44e1-97c5-4fd2-8325-423f2e540f2c · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

The Double-Ellipsoid Geometry of CLIP Imagic: Text-based real image editing with diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.355146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.791658Z digest=sha256:efdc333dff73fe3aa13f81e282c8a52e085b62d73a3a0d238c80b46bc6bafbe4

Observation ac23b385-0acd-4187-84d8-29824a7c9145 · outbound

This paper cites Optimal whitening and decorrelation.

The Double-Ellipsoid Geometry of CLIP Optimal whitening and decorrelation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.795199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.795199Z digest=sha256:273d4b0850419aec8d6a063f6e485ea02fea47bca56745cc5d86590bf1cb4a0a

Observation 914c90c9-b426-4262-bbc9-7497d2043241 · outbound

This paper cites Diffusionclip: Text-guided diffusion models for robust image manipulation.

The Double-Ellipsoid Geometry of CLIP Diffusionclip: Text-guided diffusion models for robust image manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.337163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.798626Z digest=sha256:0a029da7653bd9fc1ada31f60d1386de7a80cfcaeac9475d2ba9e725bf20450a

Observation 9bf2ca50-2804-42a3-942d-39c2e064c152 · outbound

This paper cites Self-Guided Contrastive Learning for BERT Sentence Representations.

The Double-Ellipsoid Geometry of CLIP Self-Guided Contrastive Learning for BERT Sentence Representations

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:27:09.045781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.803202Z digest=sha256:dd5311ff8bdd16c5f28c5ab5f142a6bcd36c81535801513623a5604dd8f437eb

Observation ac978ad9-78b0-45f4-ba51-d7387df81320 · outbound

This paper cites Logarithmic bounds for isoperimetry and slices of convex sets.

The Double-Ellipsoid Geometry of CLIP Logarithmic bounds for isoperimetry and slices of convex sets

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.326378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.807373Z digest=sha256:814cfabed5e593dbaa2c2d2c6e61efb032923395fbdb36c937fdb906da95380e

Observation 5675096a-f93d-473a-b7c8-846bbbcf5260 · outbound

This paper cites Bourgain’s slicing problem and kls isoperimetry up to polylog.

The Double-Ellipsoid Geometry of CLIP Bourgain’s slicing problem and kls isoperimetry up to polylog

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.315268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.811016Z digest=sha256:d1c065d4c42172305e6a5cc90cb0f83bbc5415391f54a32e32fc27a4fb824f6c

Observation 0810a8d6-d28d-40f1-b972-355fd1516f9b · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

The Double-Ellipsoid Geometry of CLIP Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.814599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.814599Z digest=sha256:17ef53bd1dd09c51deed75b88f76bf18b2c8754d1ddefdc9310386f3583e520c

Observation 082272bb-9e5b-4e0d-bd12-0a44b48d582b · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

The Double-Ellipsoid Geometry of CLIP Open-vocabulary semantic segmentation with mask-adapted clip

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.296988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.818367Z digest=sha256:d30f488e30ec9a93d6c7eca055b624c8a962ea8d37e3a0e22a54ec74f87dd3d4

Observation b9e9b5b3-673a-4d8e-97ce-fe9a3c4defeb · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

The Double-Ellipsoid Geometry of CLIP Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.285480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.821762Z digest=sha256:1e5468a0795b95f04c361ae42ee9d299d6c23debca7788468eeb644dc5e42999

Observation 1573240e-e676-4e06-9164-c71bfc98bb2b · outbound

This paper cites Microsoft coco: Common objects in context.

The Double-Ellipsoid Geometry of CLIP Microsoft coco: Common objects in context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.825500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.825500Z digest=sha256:19d133b2f4ef51c154a7b46eb8d758bb62087b124f751d9c677273ce36c25f5b

Observation 7eaba0d9-0525-47ad-829d-b1456cb58ee5 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

The Double-Ellipsoid Geometry of CLIP Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.829241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.829241Z digest=sha256:259e7db8638e565f61e15d9837063dc55a40384efe62bae46cb34f542882f704

Observation 23bc8428-552b-496b-adcd-0b06efbf0260 · outbound

This paper cites T-MARS: Improving Visual Representations by Circumventing Text Feature Learning.

The Double-Ellipsoid Geometry of CLIP T-MARS: Improving Visual Representations by Circumventing Text Feature Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.832852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.832852Z digest=sha256:a178869ecdeca94d18ea0db5e1fd3ae1c1c20f3f0a492f9f8eabd15e4909d164

Observation 1173e5b2-1281-4833-b1b2-fd21cff77965 · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

The Double-Ellipsoid Geometry of CLIP Null-text inversion for editing real images using guided diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.261120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.837390Z digest=sha256:ae6362a7a8800bd647c548b770018f17b81ba69701ff88f86a5d2d5095aecad5

Observation 8b205510-a7a7-4003-926f-b1e29ce5f987 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

The Double-Ellipsoid Geometry of CLIP ClipCap: CLIP Prefix for Image Captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.841244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.841244Z digest=sha256:b1414c03afb8d046200a3471f88fa815858a6266e17f710fc72df59445339e52

Observation 87a77cfb-2324-4d8b-aa5b-2b4d489a5b3f · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

The Double-Ellipsoid Geometry of CLIP GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.845177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.845177Z digest=sha256:fcd305f21b3dc20425ef6ea9de6f7e49f0928242673f470ceafa8db87629c6db

Observation f5b25e2e-85d3-42b3-82b1-4f4a6176caac · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

The Double-Ellipsoid Geometry of CLIP Representation Learning with Contrastive Predictive Coding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.849329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.849329Z digest=sha256:b0aca4f7bf7ce37c2c2fe4d9f2773e58db17a8cc63dfed7c328b455f19fdfc68

Observation 46acc012-612c-4832-bf2b-f5765c1ca293 · outbound

This paper cites Concentration of mass on convex bodies.

The Double-Ellipsoid Geometry of CLIP Concentration of mass on convex bodies

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.249574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.853808Z digest=sha256:b7faaef4d9d35712d96972bdb72e18c5708f007e746a6c254143eb44e72cd2ed

Observation f7cf45cd-38b8-4a73-b3f0-4f0f19a1a8d6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

The Double-Ellipsoid Geometry of CLIP Learning transferable visual models from natural language supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.857514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.857514Z digest=sha256:1a5851545ec3e2a0099640fae75c3363c00ae0eedcd7fecdb67b005e0ab85ff5

Observation 7bbcc534-37a2-488c-a04d-0cebc1ed2d0e · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

The Double-Ellipsoid Geometry of CLIP Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.861394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.861394Z digest=sha256:472f9b95ec9040a77abb6c20c88beff2df3a8a5d705f275c3a348e8015798333

Observation be457e78-b581-46f3-a84a-0688e61ce576 · outbound

This paper cites Contrastive Learning with Hard Negative Samples.

The Double-Ellipsoid Geometry of CLIP Contrastive Learning with Hard Negative Samples

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.865285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.865285Z digest=sha256:58f6008ba481525636dc744967fc22def5949fee5013688f97865e8ee3595d58

Observation ecae586e-f37f-46c1-93b3-84815eab7a7a · outbound

This paper cites Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models.

The Double-Ellipsoid Geometry of CLIP Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.869500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.869500Z digest=sha256:1fa47fc9cd8596f959e10fb28ad8ddf5d95e94ce8e1fd80f56f89f12568eca28

Observation 756ae1ac-b357-4503-b1a8-e97284e17410 · outbound

This paper cites Towards understanding the modality gap in clip.

The Double-Ellipsoid Geometry of CLIP Towards understanding the modality gap in clip

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.230577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.873464Z digest=sha256:0409baef09a01f520b54fc688150fd6f4c8b8ac79c7aefeff115494029f5287a

Observation 25320019-feb7-4ccd-9f99-02c2f237008f · outbound

This paper cites Clip4caption: Clip for video caption.

The Double-Ellipsoid Geometry of CLIP Clip4caption: Clip for video caption

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.218980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.878435Z digest=sha256:4b9aecb5f71321c8d4f63e5c1744cbdc899b972f85d0b382cbeb7163dbe5c4a9

Observation baba9d6b-8b96-4fc6-9a68-04b543150ee4 · outbound

This paper cites Too large; data reduction for vision-language pre-training.

The Double-Ellipsoid Geometry of CLIP Too large; data reduction for vision-language pre-training

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.208723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.883741Z digest=sha256:89393d4b6b96a7c726fcd4ba98bf612b6557d1d77aac5dc63143fbc30c1e8108

Observation 3e3708c2-1e12-4209-855d-170ae8df0a23 · outbound

This paper cites Understanding contrastive representation learning through alignment and uniformity on the hypersphere.

The Double-Ellipsoid Geometry of CLIP Understanding contrastive representation learning through alignment and uniformity on the hypersphere

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.197495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.887310Z digest=sha256:963255fe7bef3bd94a8616b7aa1d93b716d6544965d68280e19aa4dd19b5b6af

Observation 8c992d2e-eab8-47dd-9c91-8350b99fd5dd · outbound

This paper cites Chaos is a Ladder: A New Theoretical Understanding of Contrastive Learning via Augmentation Overlap.

The Double-Ellipsoid Geometry of CLIP Chaos is a Ladder: A New Theoretical Understanding of Contrastive Learning via Augmentation Overlap

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.890954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.890954Z digest=sha256:cd3887a9053e3b6d2ca57fbb4f0e4469019bca8dac058496e12ac30501363af3

Observation 87f4022b-f9f6-4cee-8f98-949bc6b43588 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip.

The Double-Ellipsoid Geometry of CLIP Wav2clip: Learning robust audio representations from clip

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.186237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.894992Z digest=sha256:49815dad46e4fb36bb0140c0b379f2045c013c309024dd0fca9e6c1f9fe3a620

Observation d34d1666-21b4-4268-ab99-45d680099ddd · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

The Double-Ellipsoid Geometry of CLIP Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.174912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.898772Z digest=sha256:8f8ea8ce02013d9e086dc926f411f33c25d4a272c9c9982ce862d5192e611d0f

Observation 740e3585-3cf8-451c-867e-7b9219c6ddeb · outbound

This paper cites Pointcontrast: Unsupervised pre-training for 3d point cloud understanding.

The Double-Ellipsoid Geometry of CLIP Pointcontrast: Unsupervised pre-training for 3d point cloud understanding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.163432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.902233Z digest=sha256:2e5fb633187ed19e08923d063a487163e4fe7da494129ab50b6cb1c4f00b9e98

Observation ee0d79db-2944-4174-9518-2ad3fd31c606 · outbound

This paper cites Vision-language pre-training with triple contrastive learning.

The Double-Ellipsoid Geometry of CLIP Vision-language pre-training with triple contrastive learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.151554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.905851Z digest=sha256:cb56b66b652ddb6186dfa28be710567021604ceedc54d729d0d21869ed6e3812

Observation fdeb8284-5b63-4f74-bf67-40715d3e6dd1 · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.

The Double-Ellipsoid Geometry of CLIP Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.136920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.909555Z digest=sha256:2db35df884d7f3a9d47a072031197a329776a57992f9cf6403d00b7b8ff7eac6

Observation d4736b66-74a3-4761-b6da-12d79b23e592 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

The Double-Ellipsoid Geometry of CLIP Pointclip: Point cloud understanding by clip

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.122362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.913531Z digest=sha256:7d0cef1e58ce0c0ceff7db6c354de38eda55cce6165ec9fbe76b02674ea71c91

Pith citing papers

Observation 9e393176-8412-4a5a-816e-ea283721eaac · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings The Double-Ellipsoid Geometry of CLIP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.045027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.045027Z digest=sha256:3a6860df003ee9e0168294ec85d15d73205e05eb4de84ede72a5db1678e8636f

Observation 249358b8-1e94-48ba-b496-86084158aa02 · inbound

Consistency Regularised Gradient Flows for Inverse Problems cites this paper.

Consistency Regularised Gradient Flows for Inverse Problems The Double-Ellipsoid Geometry of CLIP

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:25:54.899664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T03:21:35.082352Z digest=sha256:d68b174bf2a95c103bafa64b7f1f25883b5eef631db4cc07ad39e6cf70d070bb

Observation e08923c8-03ac-4056-b529-ae2e7e7fbe62 · inbound

HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models cites this paper.

HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models The Double-Ellipsoid Geometry of CLIP

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:54.923804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T03:12:10.701349Z digest=sha256:23060e0c434b435b13401ec4f308e902d8f43bdd2cc409538a18c5cbc4faec69

Observation a51214b1-5ced-44eb-8ff2-f786b450dc8d · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning The Double-Ellipsoid Geometry of CLIP

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T14:10:57.780723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:1f794e4fa99ad3846fee13c3f291e37397a15a264276a4024cfe31608912e21d

Observation be83d961-77ab-4be8-8e6e-5b083317ec00 · inbound

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms cites this paper.

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms The Double-Ellipsoid Geometry of CLIP

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T03:34:13.360064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T03:27:31.177489Z digest=sha256:07608008a5a4135d6a0ec6dd9180b474a496dccbc9e61fb1d05140856b73ce93

Observation fd5d053f-626a-41a5-b30a-92475cfc6c00 · inbound

On the modality gap and the contrastive loss in multi-modal representation learning cites this paper.

On the modality gap and the contrastive loss in multi-modal representation learning The Double-Ellipsoid Geometry of CLIP

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:340207eb08b9de41028768d25f58dc832082cdaab70e77eb24e35e16467adee9

Observation 742c16b9-db8c-4f45-9197-72f6a802c9c9 · inbound

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model cites this paper.

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model The Double-Ellipsoid Geometry of CLIP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T04:33:53.106844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:33:53.106844Z digest=sha256:eded88d7cc7cccb9a374688262f5b7b1644d6319499b1ed0535e994f2eb3d2d9