Pith. sign in

Paper Citation Record · LEDGER

The Double-Ellipsoid Geometry of CLIP

As of 14 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 7 inbound Pith citation observations for arXiv:2411.14517.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14517 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:27:08.913531Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:11:51.045027Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2e69f3b2-d031-4ee4-bdb6-532e7acfb594 · outbound

This paper cites write newline.

The Double-Ellipsoid Geometry of CLIP write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.692056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.692056Z digest=sha256:fdcd889ff1937072a8e0cea65e479fc77ffb48a0b472a6acff8b40b5d85383e1

Observation 3a3ce743-e330-4b24-9ea3-a0c5475141b9 · outbound

This paper cites Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding.

The Double-Ellipsoid Geometry of CLIP Crosspoint: Self-supervised cross-modal contrastive learning for 3d point cloud understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.552537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.696999Z digest=sha256:d6a2c8105249b07c191a2517f18700fcf09dc2facf7d4acee7a212f0e9e334bb

Observation 8450b623-64af-41eb-9e49-c0db6daa6e6e · outbound

This paper cites A Theoretical Analysis of Contrastive Unsupervised Representation Learning.

The Double-Ellipsoid Geometry of CLIP A Theoretical Analysis of Contrastive Unsupervised Representation Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.700684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.700684Z digest=sha256:85b42dac3c3f657b03540cb8206a8b2c3e3573e798dba55a41581f5baaa13da2

Observation 441a6acf-f127-45c5-b805-63727d1bcbcf · outbound

This paper cites Grit-vlp: Grouped mini-batch sampling for efficient vision and language pre-training.

The Double-Ellipsoid Geometry of CLIP Grit-vlp: Grouped mini-batch sampling for efficient vision and language pre-training

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.542365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.705196Z digest=sha256:22122e0eb6feefba22d17dd67bfd5aae13de8dc048dcbd063f8b5c4bb37c1ea1

Observation 8da093b7-fcd1-4f8a-94b7-4d34386714f9 · outbound

This paper cites Mafa: Managing false negatives for vision-language pre-training.

The Double-Ellipsoid Geometry of CLIP Mafa: Managing false negatives for vision-language pre-training

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.531238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.708525Z digest=sha256:f6acbd43ff2910fd6261f614d07c1419053c4eba617bdcf0b335248e76a455d8

Observation f194b376-7684-4c4f-b993-172018eca4e5 · outbound

This paper cites Clip2scene: Towards label-efficient 3d scene understanding by clip.

The Double-Ellipsoid Geometry of CLIP Clip2scene: Towards label-efficient 3d scene understanding by clip

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.520727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.712059Z digest=sha256:e8ee0a884635c430db22fa1157193c2144ff495b3b31b44d8773f6f63408b665

Observation 15d7b1fa-e38b-479e-9d0b-7bdcc41fde77 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

The Double-Ellipsoid Geometry of CLIP A simple framework for contrastive learning of visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.715475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.715475Z digest=sha256:54e082688a9ba5336eeb6695784ad87f84fda94a24f4c5cb52af84f94ab419b8

Observation 82ae280a-d73c-41e1-9c54-c9f95123f773 · outbound

This paper cites Fine-grained Image Captioning with CLIP Reward.

The Double-Ellipsoid Geometry of CLIP Fine-grained Image Captioning with CLIP Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.719025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.719025Z digest=sha256:3516a48990e1624f772d5c1af5ec022d0c041740db983360a4bb1805519220f0

Observation d2415993-83b1-4ded-8681-44e71a5fd964 · outbound

This paper cites Debiased contrastive learning.

The Double-Ellipsoid Geometry of CLIP Debiased contrastive learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.502088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.722558Z digest=sha256:bee1531b5358ee0620dfa5a9250e6c2b745dbddffc3a916ae14f5e5e093f42e3

Observation 15a06a48-4505-42e8-960e-787fc7ac50b1 · outbound

This paper cites Eccv caption: Correcting false negatives by collecting machine-and-human-verified image-caption associations for ms-coco.

The Double-Ellipsoid Geometry of CLIP Eccv caption: Correcting false negatives by collecting machine-and-human-verified image-caption associations for ms-coco

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.490912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.726139Z digest=sha256:a795d4d7ff646b404d7d3a87bf85fc55d9b7d50e271562f264fc1470f41d981a

Observation fa8d2de3-14c7-4db0-b694-8a37b0cbce0f · outbound

This paper cites Eldar and Alan V.

The Double-Ellipsoid Geometry of CLIP Eldar and Alan V

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.480059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.729900Z digest=sha256:5e63663ca1e05d7672f7ca266400014d9a7e9308e2f8f409bef5c52b41e50d83

Observation 985352e5-5476-4be7-8f86-7144a0f14e28 · outbound

This paper cites It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap.

The Double-Ellipsoid Geometry of CLIP It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.733351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.733351Z digest=sha256:dc1ce470176de61687575265870762a81f475f09713c883656d1de495a893194

Observation a3affa10-9b1b-4eab-bb4c-8a37f79ab46c · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

The Double-Ellipsoid Geometry of CLIP Datacomp: In search of the next generation of multimodal datasets

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.469451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.737235Z digest=sha256:6f22ffb77e55d7f701f1af362ddc144bfbf9547ae14a8510f25a937f4d62281f

Observation d7c749ab-375c-48b7-a5dd-8ae3ee7bbb22 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

The Double-Ellipsoid Geometry of CLIP An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.740577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.740577Z digest=sha256:b98d7d1a500034cea1870861f6fb2d61409981d1ed89ad18f8000cf3369df33e

Observation 1f1010a0-12ba-4911-8912-0e17bfa4c948 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

The Double-Ellipsoid Geometry of CLIP SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.744074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.744074Z digest=sha256:fd1cde60d587215495aaef19e773705d702a533e0f8fa7f726107e04177d5f4d

Observation c9cae812-ed5c-449a-a7c1-ef754b47dea6 · outbound

This paper cites Audioclip: Extending clip to image, text and audio.

The Double-Ellipsoid Geometry of CLIP Audioclip: Extending clip to image, text and audio

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.458805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.747702Z digest=sha256:3279ee383e662d9d814d2f72498537fe07cdc8d2cda0ac9402d658834106c435

Observation 026a5c38-ac85-4969-811f-20f7a01314f7 · outbound

This paper cites Proxedit: Improving tuning-free real image editing with proximal guidance.

The Double-Ellipsoid Geometry of CLIP Proxedit: Improving tuning-free real image editing with proximal guidance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.447304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.751124Z digest=sha256:953065ade716a137c11db73f434d6a74b7a94b6e721801ff8d43a88dee8635eb

Observation 359a2066-106d-4655-9038-55ac7475dc5e · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

The Double-Ellipsoid Geometry of CLIP Momentum contrast for unsupervised visual representation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.754774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.754774Z digest=sha256:fc177f8c8ba88637548d4f708b4cf505d82cb53a6c19e1404668ed60fa1669d0

Observation 9bb68dc2-63d7-4ba5-9134-f0f50addab42 · outbound

This paper cites Open-vocabulary multi-label classification via multi-modal knowledge transfer.

The Double-Ellipsoid Geometry of CLIP Open-vocabulary multi-label classification via multi-modal knowledge transfer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.429653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.758531Z digest=sha256:83b6fe565c000a4471973816ba5c6f6d5117533fad16a92f77726a41eb2eb785

Observation 9185dd05-1fcd-4b25-bac3-669db242fd80 · outbound

This paper cites Clip goes 3d: Leveraging prompt tuning for language grounded 3d recognition.

The Double-Ellipsoid Geometry of CLIP Clip goes 3d: Leveraging prompt tuning for language grounded 3d recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.418553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.762049Z digest=sha256:60d6652d45cf95f33f97c4d033b82c517fb259b153c326d4d12d01858983f73e

Observation 0bd0dae5-f0c7-4029-9d7a-a6384d233b9b · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

The Double-Ellipsoid Geometry of CLIP The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.765412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.765412Z digest=sha256:c2a8886d1bf0e210662cf1d27769013e797fa68ee73eb582a14bf6edf1a6d113

Observation df71112c-7046-449c-bda2-81eeab463582 · outbound

This paper cites Natural adversarial examples.

The Double-Ellipsoid Geometry of CLIP Natural adversarial examples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.768831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.768831Z digest=sha256:85723ad02fab3c15e2aacfd2caeeb603582b95463611174820551df7d2f1e3ad

Observation b2ea9f17-f111-4014-8ca2-9b5e0529f250 · outbound

This paper cites Boosting contrastive self-supervised learning with false negative cancellation.

The Double-Ellipsoid Geometry of CLIP Boosting contrastive self-supervised learning with false negative cancellation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.394316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.772696Z digest=sha256:aeb85aa00a549f2a6b5c1bd0d4c9078669cf9ce542bd219535062c28ed1ea305

Observation 59434c5c-02c4-4a45-842d-a43261a26660 · outbound

This paper cites A Slightly Improved Bound for the KLS Constant.

The Double-Ellipsoid Geometry of CLIP A Slightly Improved Bound for the KLS Constant

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.776197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.776197Z digest=sha256:87d1d902b0c3c3ecc0e2b842e5fd9cbf29efef80d343f2941510a40b4e6299f9

Observation 9f37cd12-7d20-441f-b51c-7bfb46e4f61c · outbound

This paper cites The power of contrast for feature learning: A theoretical analysis.

The Double-Ellipsoid Geometry of CLIP The power of contrast for feature learning: A theoretical analysis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.383804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.780045Z digest=sha256:84c27a1b757a3f5999ee1126f803d5c4aa80e6013b00115ac61ea6c9458ec321

Observation a10cefdd-fdc2-4ded-8f89-b9778d31862d · outbound

This paper cites Hard negative mixing for contrastive learning.

The Double-Ellipsoid Geometry of CLIP Hard negative mixing for contrastive learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.372683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.783478Z digest=sha256:21c9b015623db36c4c97b718233fa347fc40ce085259fb70fec56527f4b14144

Observation f4cfce5a-a4b6-4409-b94e-618a8f75b103 · outbound

This paper cites Isoperimetric problems for convex bodies and a localization lemma.

The Double-Ellipsoid Geometry of CLIP Isoperimetric problems for convex bodies and a localization lemma

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.788113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.788113Z digest=sha256:35a18e08ff41a49153836b28f191ef8160ec216d2f32904686d358953d931661

Observation 85ee44e1-97c5-4fd2-8325-423f2e540f2c · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

The Double-Ellipsoid Geometry of CLIP Imagic: Text-based real image editing with diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.355146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.791658Z digest=sha256:ac11c9d2befc3669d95ae37eb5326bc7aa0b0453c30a572ab1166dc0f0af409d

Observation ac23b385-0acd-4187-84d8-29824a7c9145 · outbound

This paper cites Optimal whitening and decorrelation.

The Double-Ellipsoid Geometry of CLIP Optimal whitening and decorrelation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.795199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.795199Z digest=sha256:69b005439726899a4292099809cbc4436dab6467a837c88992bf37a92a36525a

Observation 914c90c9-b426-4262-bbc9-7497d2043241 · outbound

This paper cites Diffusionclip: Text-guided diffusion models for robust image manipulation.

The Double-Ellipsoid Geometry of CLIP Diffusionclip: Text-guided diffusion models for robust image manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.337163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.798626Z digest=sha256:845bdd21c52d65aae250ccb5626b86fba7bb16d2aaaeece82b83dc76fb64176d

Observation 9bf2ca50-2804-42a3-942d-39c2e064c152 · outbound

This paper cites Self-Guided Contrastive Learning for BERT Sentence Representations.

The Double-Ellipsoid Geometry of CLIP Self-Guided Contrastive Learning for BERT Sentence Representations

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:27:09.045781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.803202Z digest=sha256:0fa4e484397bcc1d09df0d80d0bea3a98b1cbdabff28ec0e12c2479aec4bcf45

Observation ac978ad9-78b0-45f4-ba51-d7387df81320 · outbound

This paper cites Logarithmic bounds for isoperimetry and slices of convex sets.

The Double-Ellipsoid Geometry of CLIP Logarithmic bounds for isoperimetry and slices of convex sets

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.326378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.807373Z digest=sha256:b5946187ee28c46209a03cd4337dc58007afab7a213ddc89f41a6aa23ff51ceb

Observation 5675096a-f93d-473a-b7c8-846bbbcf5260 · outbound

This paper cites Bourgain’s slicing problem and kls isoperimetry up to polylog.

The Double-Ellipsoid Geometry of CLIP Bourgain’s slicing problem and kls isoperimetry up to polylog

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.315268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.811016Z digest=sha256:208fbdeaa4f16a57a8857fa669752cf17c04dca00aedb65c57a80732fc32ac01

Observation 0810a8d6-d28d-40f1-b972-355fd1516f9b · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

The Double-Ellipsoid Geometry of CLIP Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.814599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.814599Z digest=sha256:d96aa00cffdff5f048349b280c7f04e348d42cdde6aec579c52e05a30c54cca2

Observation 082272bb-9e5b-4e0d-bd12-0a44b48d582b · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

The Double-Ellipsoid Geometry of CLIP Open-vocabulary semantic segmentation with mask-adapted clip

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.296988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.818367Z digest=sha256:16d3b673e18d1a5da25289d89200fee92e50c791a13115f580abd161150dc459

Observation b9e9b5b3-673a-4d8e-97ce-fe9a3c4defeb · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

The Double-Ellipsoid Geometry of CLIP Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.285480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.821762Z digest=sha256:191e819ab500a02c5a1c1027d10c0909dec6ff64f0d1a728ea39b109a470c5fb

Observation 1573240e-e676-4e06-9164-c71bfc98bb2b · outbound

This paper cites Microsoft coco: Common objects in context.

The Double-Ellipsoid Geometry of CLIP Microsoft coco: Common objects in context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.825500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.825500Z digest=sha256:c2ad5f127024b4d619f1265511af07f1cd4044e270af865c102ab72f89f7713b

Observation 7eaba0d9-0525-47ad-829d-b1456cb58ee5 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

The Double-Ellipsoid Geometry of CLIP Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.829241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.829241Z digest=sha256:46a8103e7322d8ba81ed9dfdabb19acf1062df9b6aea3275887995306b01c050

Observation 23bc8428-552b-496b-adcd-0b06efbf0260 · outbound

This paper cites T-MARS: Improving Visual Representations by Circumventing Text Feature Learning.

The Double-Ellipsoid Geometry of CLIP T-MARS: Improving Visual Representations by Circumventing Text Feature Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.832852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.832852Z digest=sha256:bfeae3ebb66e28375034a2819f5eea0de8b48a9931a58b69bfeea70ab891eda5

Observation 1173e5b2-1281-4833-b1b2-fd21cff77965 · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

The Double-Ellipsoid Geometry of CLIP Null-text inversion for editing real images using guided diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.261120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.837390Z digest=sha256:82360ebf92f425267e3f29d3b2873911b6c1548f1a755678c3413a0b934f630c

Observation 8b205510-a7a7-4003-926f-b1e29ce5f987 · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

The Double-Ellipsoid Geometry of CLIP ClipCap: CLIP Prefix for Image Captioning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.841244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.841244Z digest=sha256:7a7b9a99dc7c5afac56daeadf78edc8309dd80a17a2221bd9cd64dcc546f3c39

Observation 87a77cfb-2324-4d8b-aa5b-2b4d489a5b3f · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

The Double-Ellipsoid Geometry of CLIP GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.845177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.845177Z digest=sha256:f2d1d8fe4f7c4a28835f9c70acf0db68633d8ab1e9c4a897219acaadf0bb2f58

Observation f5b25e2e-85d3-42b3-82b1-4f4a6176caac · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

The Double-Ellipsoid Geometry of CLIP Representation Learning with Contrastive Predictive Coding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.849329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.849329Z digest=sha256:5668638d3d36420bd300328ed402e9ca7a90c2a33b2f99fcac0970b84f4fb070

Observation 46acc012-612c-4832-bf2b-f5765c1ca293 · outbound

This paper cites Concentration of mass on convex bodies.

The Double-Ellipsoid Geometry of CLIP Concentration of mass on convex bodies

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.249574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.853808Z digest=sha256:1bef3c6e2621fe29c0da4bf83e584b812f79ccf092fe0c3c149b05f564a8fdf9

Observation f7cf45cd-38b8-4a73-b3f0-4f0f19a1a8d6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

The Double-Ellipsoid Geometry of CLIP Learning transferable visual models from natural language supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.857514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.857514Z digest=sha256:3b0dad7882feaff357204feeed19b062179878e84d5cdbe2442327bfaf4be205

Observation 7bbcc534-37a2-488c-a04d-0cebc1ed2d0e · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

The Double-Ellipsoid Geometry of CLIP Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.861394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.861394Z digest=sha256:91f15d019943a71180f6c257aaab881080e0c678dbb68f0f098cad559be04ea1

Observation be457e78-b581-46f3-a84a-0688e61ce576 · outbound

This paper cites Contrastive Learning with Hard Negative Samples.

The Double-Ellipsoid Geometry of CLIP Contrastive Learning with Hard Negative Samples

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.865285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.865285Z digest=sha256:cd68730772b9d6718d92fcbc2dda48e6e343ce915df5e85cc3fc3d6d2f915afe

Observation ecae586e-f37f-46c1-93b3-84815eab7a7a · outbound

This paper cites Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models.

The Double-Ellipsoid Geometry of CLIP Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.869500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.869500Z digest=sha256:19de47683a87d7a42e08fc806daa17fd6587e098ff6f0855b7bf54d2c48a1db1

Observation 756ae1ac-b357-4503-b1a8-e97284e17410 · outbound

This paper cites Towards understanding the modality gap in clip.

The Double-Ellipsoid Geometry of CLIP Towards understanding the modality gap in clip

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.230577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.873464Z digest=sha256:7035d6423fb9ea8a893063f3c0cb9d2f82c9ae7a20a4ff419058c5c08efb0758

Observation 25320019-feb7-4ccd-9f99-02c2f237008f · outbound

This paper cites Clip4caption: Clip for video caption.

The Double-Ellipsoid Geometry of CLIP Clip4caption: Clip for video caption

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.218980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.878435Z digest=sha256:edc933843713a6d94d368ea16a029afdf01bfea2ad86737541076fa8d1ee8376

Observation baba9d6b-8b96-4fc6-9a68-04b543150ee4 · outbound

This paper cites Too large; data reduction for vision-language pre-training.

The Double-Ellipsoid Geometry of CLIP Too large; data reduction for vision-language pre-training

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.208723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.883741Z digest=sha256:64d0b999226e1a9147d16ded4c9a351ec41537fe1a490ba0d348122d3981c135

Observation 3e3708c2-1e12-4209-855d-170ae8df0a23 · outbound

This paper cites Understanding contrastive representation learning through alignment and uniformity on the hypersphere.

The Double-Ellipsoid Geometry of CLIP Understanding contrastive representation learning through alignment and uniformity on the hypersphere

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.197495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.887310Z digest=sha256:ddb0c117624dc14c8049dc349f6ef20aeb03e3685e0f94ef77e2dd1f274b9300

Observation 8c992d2e-eab8-47dd-9c91-8350b99fd5dd · outbound

This paper cites Chaos is a Ladder: A New Theoretical Understanding of Contrastive Learning via Augmentation Overlap.

The Double-Ellipsoid Geometry of CLIP Chaos is a Ladder: A New Theoretical Understanding of Contrastive Learning via Augmentation Overlap

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.890954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.890954Z digest=sha256:c6a3fe58e855b6861ccdc1358eea9e3e612cf1ca215cc0a3aeaffc99ea49510a

Observation 87f4022b-f9f6-4cee-8f98-949bc6b43588 · outbound

This paper cites Wav2clip: Learning robust audio representations from clip.

The Double-Ellipsoid Geometry of CLIP Wav2clip: Learning robust audio representations from clip

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.186237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.894992Z digest=sha256:e9b17451d5e5368e9248c16fee7026b7fefbef371b9ef43d68865618f19add3c

Observation d34d1666-21b4-4268-ab99-45d680099ddd · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

The Double-Ellipsoid Geometry of CLIP Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.174912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.898772Z digest=sha256:df5b43e5e4549ad6ce01c48625499dc91bd2c36dafebbdbd4ad999927cae96ab

Observation 740e3585-3cf8-451c-867e-7b9219c6ddeb · outbound

This paper cites Pointcontrast: Unsupervised pre-training for 3d point cloud understanding.

The Double-Ellipsoid Geometry of CLIP Pointcontrast: Unsupervised pre-training for 3d point cloud understanding

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.163432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.902233Z digest=sha256:5492300352d827e7bd5709ac0656bbb658fddf485d052f0c0919abbc39569977

Observation ee0d79db-2944-4174-9518-2ad3fd31c606 · outbound

This paper cites Vision-language pre-training with triple contrastive learning.

The Double-Ellipsoid Geometry of CLIP Vision-language pre-training with triple contrastive learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.151554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.905851Z digest=sha256:7bea4bff70781d35ef73ce724e301db3e6b98c89ae07cd271e968bb44b267e27

Observation fdeb8284-5b63-4f74-bf67-40715d3e6dd1 · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.

The Double-Ellipsoid Geometry of CLIP Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.136920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.909555Z digest=sha256:7fde8de5a7fa0c742ccfc4d2ac8d7c664b30877b8db915dd9f545e75c3118e7d

Observation d4736b66-74a3-4761-b6da-12d79b23e592 · outbound

This paper cites Pointclip: Point cloud understanding by clip.

The Double-Ellipsoid Geometry of CLIP Pointclip: Point cloud understanding by clip

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:27:09.122362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T15:27:08.913531Z digest=sha256:5ccb6417f235e5ab72d5278c6d1c0d620cc25cddb02913a2e1c3de850a244ab2

Pith citing papers

Observation 9e393176-8412-4a5a-816e-ea283721eaac · inbound

On the rankability of visual embeddings cites this paper.

On the rankability of visual embeddings The Double-Ellipsoid Geometry of CLIP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:51.045027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:11:51.045027Z digest=sha256:0aad4cfe282fa5c5e67626f59572ea46efa45ef98d000dcf4528a9dc0e4ff8ce

Observation 249358b8-1e94-48ba-b496-86084158aa02 · inbound

Consistency Regularised Gradient Flows for Inverse Problems cites this paper.

Consistency Regularised Gradient Flows for Inverse Problems The Double-Ellipsoid Geometry of CLIP

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:25:54.899664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T03:21:35.082352Z digest=sha256:cb5e45bb14cfe594f26deda228f871b5aeaf1fa9db953c2a6834423e0620deb6

Observation e08923c8-03ac-4056-b529-ae2e7e7fbe62 · inbound

HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models cites this paper.

HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models The Double-Ellipsoid Geometry of CLIP

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:54.923804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T03:12:10.701349Z digest=sha256:0d6056a21b70329cca96c35f7ce0d558aab6db0eb425e7e5a4817c5f60065136

Observation a51214b1-5ced-44eb-8ff2-f786b450dc8d · inbound

When to Align, When to Predict: A Phase Diagram for Multimodal Learning cites this paper.

When to Align, When to Predict: A Phase Diagram for Multimodal Learning The Double-Ellipsoid Geometry of CLIP

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T14:10:57.780723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T14:08:33.542700Z digest=sha256:b877b47fa9558a9aa858397c910e5d3770d3ff1f98fcf8bcdc1fd2fa1a72c2f4

Observation be83d961-77ab-4be8-8e6e-5b083317ec00 · inbound

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms cites this paper.

Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms The Double-Ellipsoid Geometry of CLIP

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T03:34:13.360064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T03:27:31.177489Z digest=sha256:0d6608e78e9d414cbf6add73857042a1f0bd1cf84344331c048bce0cf74e3d2c

Observation fd5d053f-626a-41a5-b30a-92475cfc6c00 · inbound

On the modality gap and the contrastive loss in multi-modal representation learning cites this paper.

On the modality gap and the contrastive loss in multi-modal representation learning The Double-Ellipsoid Geometry of CLIP

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T09:55:43.412471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T09:55:43.412471Z digest=sha256:53e6196fa6f7112470fd4852105e75098a4aa2d95c8872eca1557f9351c6dacf

Observation 742c16b9-db8c-4f45-9197-72f6a802c9c9 · inbound

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model cites this paper.

The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model The Double-Ellipsoid Geometry of CLIP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T04:33:53.106844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:33:53.106844Z digest=sha256:8bf92defc68bab40f3778e75f06e5e02e3a0c61d3e8d798aa93a3c08b588850b