Pith. sign in

Paper Citation Record · LEDGER

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 5 inbound Pith citation observations for arXiv:2502.04263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04263 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:05:28.366500Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:24.087885Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:13:29.951494Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e67da87-1f65-4a36-9c84-2ed7d83ab3bc · outbound

This paper cites write newline.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.158867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.158867Z digest=sha256:5746d78dabe246d8a69670a9bd5998f1e5155863390b35271e1ecd8a926ef7da

Observation df186c1b-5f45-42da-ae80-8a73e10a113e · outbound

This paper cites iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.164106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.164106Z digest=sha256:97dbceb0c0d6c7b01975796967801842d51112cccc87c29c2e6c8ea49f4e33ab

Observation 6a124118-1d9a-4700-8267-fe7b4076ded6 · outbound

This paper cites nocaps: novel object captioning at scale.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion nocaps: novel object captioning at scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.802665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.168444Z digest=sha256:0d420736d00339aca49e416abbf478a1fb8ed1f38a0dbbb4f18b817956f85bfe

Observation 2b8ea1ac-88fe-458b-a1f0-5357f4f48c1e · outbound

This paper cites Zero-Shot Composed Image Retrieval with Textual Inversion.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Zero-Shot Composed Image Retrieval with Textual Inversion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.792661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.172838Z digest=sha256:e19c2993b61070211662d3563dad2f84e0c8b01f3a1fd5d12feb2583c67c4446

Observation e24f9d0c-9719-4aea-bc9e-27794b8a251a · outbound

This paper cites Food-101--mining discriminative components with random forests.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Food-101--mining discriminative components with random forests

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.782507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.176227Z digest=sha256:7d4934d9322a4ed6f7b1a64bc04f70d49eed3a845d4635ed7cdd6fd0b549900a

Observation ff880db6-7a66-46f4-bf95-40dda6bfbad5 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion A simple framework for contrastive learning of visual representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.179775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.179775Z digest=sha256:46c62a2f8cfaa32458e82df728529f0b73b5beb066d7fdc6e7302d7ba376f458

Observation 320cc0d9-1416-4192-bc80-9746e1222e7d · outbound

This paper cites Describing textures in the wild.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Describing textures in the wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.183453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.183453Z digest=sha256:7d5c9446d0e71ebe297bb459dbe8f4d35450bb2b4b0fb96d6c8a76e6e2f1e2ed

Observation b34db35c-6e3d-4e2c-887f-31bb92027d55 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Imagenet: A large-scale hierarchical image database

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.187018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.187018Z digest=sha256:e51a6c68f0bdd0a2a8fa847dbe30673065c7ff4c5e2f607c73728b9c834c1df3

Observation f4ebc51a-6dc0-4995-9654-972d35bd0975 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion An image is worth 16x16 words: Transformers for image recognition at scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.190767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.190767Z digest=sha256:9bf783a9d66e02e2fb91547b7cba70ef1e9e441684b87f39644c029c7d211b17

Observation 4d93ccfe-64ce-41ec-ab1a-36897b528558 · outbound

This paper cites The Llama 3 Herd of Models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.194653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.194653Z digest=sha256:b75718435766de56d78f5f7b2ac1a0a8d5e384c780d0d64e78dd9a0aa775371a

Observation b092ac40-efb0-402d-93e4-5b1ee8646544 · outbound

This paper cites Asirra: a captcha that exploits interest-aligned manual image categorization.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Asirra: a captcha that exploits interest-aligned manual image categorization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.748820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.198269Z digest=sha256:1c069290d426524d35b759b098d8c6e22b101f285e9451272b3d9b3801184076

Observation f16a91e5-fbf0-4781-a47f-eeae517f014e · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Structure and content-guided video synthesis with diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.201903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.201903Z digest=sha256:8f894009bf6b8a2e8309dd77820690d1433d24082046ece0cf52e30c8630aa5c

Observation d639af3c-d7a2-4186-a0a7-ab65a62c1bf1 · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.205084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.205084Z digest=sha256:e2329ff5f02eac6831d95f697b8d525d294f6cb40cedc44a18672a4a39f6dfde

Observation 8fe38a14-23d5-424a-ba2d-e2223f2e2a5d · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Datacomp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.208342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.208342Z digest=sha256:775bce478b8ac3b24aa0cf15044c749464e74cc89b67ce642f3bad66964ec447

Observation ad8985fd-34e1-4311-8b00-0f681239b6b8 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.211690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.211690Z digest=sha256:f987d996d3e65d362dbcd395e39b57359080238d25517598927850e9737523cf

Observation 63a55951-2600-47fd-b55b-351820567691 · outbound

This paper cites Towards flexible perception with visual memory.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Towards flexible perception with visual memory

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-08T23:05:31.393813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.215151Z digest=sha256:f64744ed93c3b0e025d29cbe6f46874a3d9dce793e61352bb396a6159026b175

Observation 38e593f4-eeec-400b-a8f8-6acc321ef0dc · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.218897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.218897Z digest=sha256:8f22b53788b4f07169f8c740b48d588b89e75a5b90de6755433dbd453496fb4d

Observation cedcd4ea-f5da-4c1e-86ab-a417dec8c4b9 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Scaling up visual and vision-language representation learning with noisy text supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.222020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.222020Z digest=sha256:b55de9e58678a52b7d78f4135b3cdea9df1cd91c2dfa508b6ebcccec1fceef83

Observation 6e08ae28-74fa-4ae2-ad32-1a8c14e0cbd5 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Deep visual-semantic alignments for generating image descriptions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.225249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.225249Z digest=sha256:77615aa77e93f64952883ff0dacf548ccf2bd492fcccc3109e3fcded340cb905

Observation d2b1de36-b920-436e-b296-197dd7e96096 · outbound

This paper cites 3d object representations for fine-grained categorization.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion 3d object representations for fine-grained categorization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.228575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.228575Z digest=sha256:d279fdba4a419a9ad2eadb760311119d001b6d4a4e38932f03d4d6256fa24995

Observation 25920617-61a7-448d-bf15-b68c9858cfac · outbound

This paper cites Newsweeder: Learning to filter netnews.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Newsweeder: Learning to filter netnews

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.231717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.231717Z digest=sha256:11584e86c634b7ea7982e3e1ed420d5a9f8230ecb1d5718b08688beb2a86ed6c

Observation 6a1cca45-c804-4898-b714-2d95db06f8fb · outbound

This paper cites DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.234921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.234921Z digest=sha256:edaa5a92e5532675bd6bc17c0102adcbd0cf9f7e923ce0bfa30a0c4a0f10c329

Observation 7eaf461c-ec4e-4ac5-bf04-695ec715f0a5 · outbound

This paper cites Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.238464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.238464Z digest=sha256:d1de3e942034873af43b6e7d9c9316a44ccb74bf94c32543f7ca4226610ab846

Observation dcffd5eb-efe3-4146-b392-479e4d08bcbe · outbound

This paper cites Mind the Gap: Understanding the Modality Gap in Multi-Modal Contrastive Representation Learning.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Mind the Gap: Understanding the Modality Gap in Multi-Modal Contrastive Representation Learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.690268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.242157Z digest=sha256:cd8daa5db8620498d2b130ffb0c05704e400e4d3730cbfa04b19065f3d3ac7c6

Observation 666c1b54-5e94-4533-8dc4-f61fa7837b40 · outbound

This paper cites Microsoft coco: Common objects in context.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Microsoft coco: Common objects in context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.245374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.245374Z digest=sha256:c02d960fed15472d1e91bcf466630b4b41a882a6cf80f4f145b0ad087fb80207

Observation 6f157b15-ef77-420f-8058-d6bb58d629d3 · outbound

This paper cites Visual instruction tuning.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.248562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.248562Z digest=sha256:1cd1a677494fee3d7ef1525c8d5fda509b330c1755af618f4d8f890af1d8adbb

Observation c4a5a5ad-23f6-4684-9168-b3a6e7bac4a8 · outbound

This paper cites Decoupled weight decay regularization.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Decoupled weight decay regularization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.251727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.251727Z digest=sha256:b122ab3665eae82f3422c4b9a9298abbf8b5338e842d612e5aa8cefd75f62f69

Observation 93e0a649-9035-4687-808c-41c3eed86e4d · outbound

This paper cites Image segmentation using text and image prompts.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Image segmentation using text and image prompts

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.662166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.255108Z digest=sha256:12bf7f2c4a89d3b325f872d3c1f7b932b482449bac7a76405625263d1769ef64

Observation 5d68d289-afc7-4060-af7d-2faaa7000cb5 · outbound

This paper cites Learning word vectors for sentiment analysis.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning word vectors for sentiment analysis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.258345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.258345Z digest=sha256:90577f4c8f6fb6ff4ebee201a8eefa6b6cab6d541414d19f7002b12fe9bfddc2

Observation 52dd46dc-1bfa-4692-aa54-33ba31143ca7 · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Fine-Grained Visual Classification of Aircraft

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.261742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.261742Z digest=sha256:8d0e3b8f2a466b8617ea8135be565511e65cb77084cff6c411399c401343526b

Observation ad756d82-f0fa-4950-8361-5aa369639e19 · outbound

This paper cites Slip: Self-supervision meets language-image pre-training.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Slip: Self-supervision meets language-image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.645695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.265463Z digest=sha256:8681430211b792b3ac25680fba64df3f3e61811d3dbc050928f9fdcdec8c7560

Observation 30629513-27b0-4a96-a759-3b43540c0eb0 · outbound

This paper cites Automated flower classification over a large number of classes.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Automated flower classification over a large number of classes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.268889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.268889Z digest=sha256:07ede950361a51707e3298278b2357a1342acd1c04f4e44aee6b1b9a1d80456a

Observation b433afe8-3ce4-4f3a-a751-116285151618 · outbound

This paper cites Deep metric learning via lifted structured feature embedding.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Deep metric learning via lifted structured feature embedding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.628790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.272247Z digest=sha256:cea9ab443a7d05985ea2e6353d3a373fb5b8b75f187ab4ebfeb65db66a02b447

Observation 9cd5f98e-8640-4594-a7a9-b3b55ac0fb10 · outbound

This paper cites Clip-guided vision-language pre-training for question answering in 3d scenes.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Clip-guided vision-language pre-training for question answering in 3d scenes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.618213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.275652Z digest=sha256:952a7ff4c01d5260f09fe7721e8ece5d0d37fb77596ab8a49688eef066a8a15a

Observation 43f0499c-47ad-4c36-973b-d9908769a5c9 · outbound

This paper cites Cats and dogs.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Cats and dogs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.278965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.278965Z digest=sha256:82edf274ee261760de74fdc58b9aebb5b5f56d582b5760c5929290a0295e5c94

Observation b8336ba0-9c42-49cc-92d2-75f4c91b39d7 · outbound

This paper cites Eclipse: A resource-efficient text-to-image prior for image generations.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Eclipse: A resource-efficient text-to-image prior for image generations

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.601859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.282437Z digest=sha256:e6e62864cb905781ffdeb45d79e73f076dba195b4708b520fa67ea64c8cb1eaa

Observation 1d5606bf-8a1d-4a78-8014-a40faae32aa0 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.285695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.285695Z digest=sha256:7ab2dc26155269d130704d14b4530247b9767ac0bc9cc38319e50b1335179af5

Observation e71decce-b0da-4846-bd81-5920527175a3 · outbound

This paper cites Revisiting oxford and paris: Large-scale image retrieval benchmarking.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Revisiting oxford and paris: Large-scale image retrieval benchmarking

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.585707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.288917Z digest=sha256:eb70cab593c841d66ac24102a2bfee00616493b7b9a1e4b4c10685a2a5bc366b

Observation 2d8504b3-ee9d-4792-8b83-9c299a793a3f · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning Transferable Visual Models From Natural Language Supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.574565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.292099Z digest=sha256:45168f8f618c551f3ce422a1dd8193b790052566a9c2cf8d3731cd907240e08d

Observation aea494b7-c64a-473e-a9f6-36c9950c2d25 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.295365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.295365Z digest=sha256:601c2ac500aeec0bf3a5ffe094675dc9e3c8da02cae043b38c9695043de2ac83

Observation a2a37b1b-bd28-48f8-b7ac-0a1f0cd53173 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.298693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.298693Z digest=sha256:70052037adabc2bd232218186e447dc17271d50b7c9d4a771774de8380c7ca96

Observation 1ef2bc42-8645-4506-9f70-e9690e01dfe7 · outbound

This paper cites Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.302016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.302016Z digest=sha256:ed0d323969f56d798cb0756eb00fc1116fdb81c8bfd0ff0bf57d5eaf5cd7123a

Observation 12435e46-267e-49ac-9880-830c2f685cb7 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.305486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.305486Z digest=sha256:a299aac34c951ea42db9f16153ea25edc69eb0a39d4883b0519d879c39aed663

Observation 14bfbac4-28ab-479e-bb57-685723b527d1 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.309008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.309008Z digest=sha256:786b7b6554fe5f01806ce5153a63233241eade9e132d17b15894f253cb16194e

Observation b6c45df0-393b-405a-8615-fafd2d6bacd6 · outbound

This paper cites Towards Understanding the Modality Gap in CLIP.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Towards Understanding the Modality Gap in CLIP

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.545840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.312316Z digest=sha256:a5b2c95fb08f0b8acde16f011f9ff9683bbec3fc9113398e2e56dfa55cdacbc5

Observation 07ff24e2-ce6f-499d-ba4c-4df030201343 · outbound

This paper cites Improved deep metric learning with multi-class n-pair loss objective.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Improved deep metric learning with multi-class n-pair loss objective

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.534046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.315355Z digest=sha256:c468d2f11c700751f9352c17142c5eb51a27de6d9773dd81d4e5508d7973d625

Observation adad59bb-9f9f-406c-9c95-8386147827c9 · outbound

This paper cites CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.318675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.318675Z digest=sha256:32abee9f6e33951279b571e84062635e8b8ff5129a28f690d1542f2d8e6dfb8f

Observation edbcfd55-15c3-4edb-b6ed-461eb7da6929 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.322291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.322291Z digest=sha256:4533675ed2c6926647dbd7f3498f3347fa4758745a4a0ccaa07acf106718a851

Observation fac08640-37ec-48d0-b3ff-5667bca5d0d8 · outbound

This paper cites SuS-X: Training-Free Name-Only Transfer of Vision-Language Models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion SuS-X: Training-Free Name-Only Transfer of Vision-Language Models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.524274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.325835Z digest=sha256:754aed79fc2cd1222aa8ddf9706b1d0ab988739c07b099268f8102cc06408f82

Observation 5efc861e-25af-419d-b972-9ab6a20fc525 · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion The caltech-ucsd birds-200-2011 dataset

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.329093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.329093Z digest=sha256:7c63e0a9023b971355fc2e9aaa476954b168dac16e7883d553ff8facd760e04d

Observation 53627af5-6126-4204-a195-5dbaf1e5e96b · outbound

This paper cites Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.508001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.332228Z digest=sha256:5193b61b00b252e72e935417c3323ac54357e744eb630c1fa0bb226bb6587c97

Observation 11fbbb18-74b0-4e16-a09f-43f52b2f9afd · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Coca: Contrastive captioners are image-text foundation models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.497851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.335778Z digest=sha256:992264dda27b955b97e8107cdca7f7b887f16c7c533aac70effcb02ee1841e36

Observation 692e86c5-9045-40d1-9ca3-544cc62727a8 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Lit: Zero-shot transfer with locked-image text tuning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.487903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.339318Z digest=sha256:35c27fa5e3c6f44fb185ebcfb549d94e39188cfbb8b55e92bd58ed060a80cb9b

Observation dc84c6d1-b7e2-415e-8467-e0ebbee1d78a · outbound

This paper cites Sigmoid loss for language image pre-training.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Sigmoid loss for language image pre-training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.342500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.342500Z digest=sha256:d058a22baef5045ca4567445c17f374e341183136d8dd9d6f7de5068a0cab16a

Observation 64c93fef-c480-4d13-886b-c78aefefe651 · outbound

This paper cites Diagnosing and Rectifying Vision Models using Language.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Diagnosing and Rectifying Vision Models using Language

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.471591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.345739Z digest=sha256:344be9f16bbd5322f915100d09a4bc3d9cb8f0708aacceede16cf86f38c7a151

Observation 18179f04-a4c0-4af9-925e-1a51f971e059 · outbound

This paper cites Avid: Any-length video inpainting with diffusion model.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Avid: Any-length video inpainting with diffusion model

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.461236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.348990Z digest=sha256:d487c93fba586d19a313a1961928164a8ad552023aae0b16f41988d0467312d5

Observation 5538da71-045c-4d56-bdb5-3b5ca66cb48f · outbound

This paper cites Extract free dense labels from clip.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Extract free dense labels from clip

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T23:05:31.451094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T23:05:28.352254Z digest=sha256:2ce48682cb51d683229a2edc0dc8b8c2018584cd3606dc38f72e02cc66ca5103

Observation 7847be78-e3e5-45a1-b39d-08130c794dbb · outbound

This paper cites Learning to prompt for vision-language models.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning to prompt for vision-language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.355537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.355537Z digest=sha256:8613ec7dabffa847a1f5ff8dc0e9e46f04040fb9c84306d9d6dffd568247b3b3

Observation cbbb8424-0553-484c-b80a-48770af66919 · outbound

This paper cites @esa (Ref.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion @esa (Ref

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.358826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.358826Z digest=sha256:01643e44cf21230e0cc82c689b02067d29207c4c6bfaba3f8d29b80b9cc46957

Observation 2965186f-a05c-4125-a819-6c8c989c558c · outbound

This paper cites an unresolved cited work.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.362890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.362890Z digest=sha256:802e88720e138aedcd1c584739eca0a9280f10a17305210051ff25480b40ef1b

Observation 647fabc1-5e1a-4907-aeef-00cde2700bf4 · outbound

This paper cites a photo of.

Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion a photo of

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T23:05:28.366500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:05:28.366500Z digest=sha256:e8eddaed6c10071735691b6fa9e51f8bd1e72aca836b978c3e06a725da59b325

Pith citing papers

Observation 421bfe05-7e31-41b5-ad92-4b487724c3f8 · inbound

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models cites this paper.

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:24.087885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:24.087885Z digest=sha256:1ee46520a91abbb0970b8978d4aefdc082378eaf8ab2dbefc31b6127a43f92bd

Observation a3bd064c-3c80-479b-a405-5813f35c1596 · inbound

Global and Local Entailment Learning for Natural World Imagery cites this paper.

Global and Local Entailment Learning for Natural World Imagery Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:33:33.700571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:33:33.700571Z digest=sha256:51e0999666387dc555d1a3b54a893561827d4886b70f659ff5f0715740b90662

Observation f64889f6-b20e-45ba-98ec-b6bcaf0766dd · inbound

Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning cites this paper.

Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:17.391053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:17.391053Z digest=sha256:c2767dc341058e63188731cc741e7775ca3e8d427f1742d9432ce9a65ace92d0

Observation 9528a14c-23ee-49b2-b3de-e52fd5152860 · inbound

Best Segmentation Buddies for Image-Shape Correspondence cites this paper.

Best Segmentation Buddies for Image-Shape Correspondence Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:13.515460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:10:29.360224Z digest=sha256:6c0484f930a4362283344d0583c9535d93a6d3f638300b11424d12ea6e905c98

Observation 963adb28-36cf-4c0d-894e-b1dc113a3e37 · inbound

LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment cites this paper.

LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:13:29.953305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T14:12:23.837961Z digest=sha256:6db9a90b7879deb670d26761a23df8f1f064ede90347e1afec7fe5ec8bcc2459