Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T23:05:28.366500Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 5 inbound Pith citation observations for arXiv:2502.04263.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T23:05:28.366500Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:24.087885Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T14:13:29.951494Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8e67da87-1f65-4a36-9c84-2ed7d83ab3bc · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df186c1b-5f45-42da-ae80-8a73e10a113e · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a124118-1d9a-4700-8267-fe7b4076ded6 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion nocaps: novel object captioning at scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2b8ea1ac-88fe-458b-a1f0-5357f4f48c1e · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Zero-Shot Composed Image Retrieval with Textual Inversion
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e24f9d0c-9719-4aea-bc9e-27794b8a251a · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Food-101--mining discriminative components with random forests
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff880db6-7a66-46f4-bf95-40dda6bfbad5 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion A simple framework for contrastive learning of visual representations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320cc0d9-1416-4192-bc80-9746e1222e7d · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Describing textures in the wild
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b34db35c-6e3d-4e2c-887f-31bb92027d55 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Imagenet: A large-scale hierarchical image database
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ebc51a-6dc0-4995-9654-972d35bd0975 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion An image is worth 16x16 words: Transformers for image recognition at scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d93ccfe-64ce-41ec-ab1a-36897b528558 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b092ac40-efb0-402d-93e4-5b1ee8646544 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Asirra: a captcha that exploits interest-aligned manual image categorization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f16a91e5-fbf0-4781-a47f-eeae517f014e · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Structure and content-guided video synthesis with diffusion models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d639af3c-d7a2-4186-a0a7-ab65a62c1bf1 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe38a14-23d5-424a-ba2d-e2223f2e2a5d · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Datacomp: In search of the next generation of multimodal datasets
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad8985fd-34e1-4311-8b00-0f681239b6b8 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a55951-2600-47fd-b55b-351820567691 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Towards flexible perception with visual memory
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 38e593f4-eeec-400b-a8f8-6acc321ef0dc · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cedcd4ea-f5da-4c1e-86ab-a417dec8c4b9 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Scaling up visual and vision-language representation learning with noisy text supervision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e08ae28-74fa-4ae2-ad32-1a8c14e0cbd5 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Deep visual-semantic alignments for generating image descriptions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b1de36-b920-436e-b296-197dd7e96096 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion 3d object representations for fine-grained categorization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25920617-61a7-448d-bf15-b68c9858cfac · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Newsweeder: Learning to filter netnews
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1cca45-c804-4898-b714-2d95db06f8fb · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eaf461c-ec4e-4ac5-bf04-695ec715f0a5 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcffd5eb-efe3-4146-b392-479e4d08bcbe · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Mind the Gap: Understanding the Modality Gap in Multi-Modal Contrastive Representation Learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 666c1b54-5e94-4533-8dc4-f61fa7837b40 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Microsoft coco: Common objects in context
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f157b15-ef77-420f-8058-d6bb58d629d3 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Visual instruction tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4a5a5ad-23f6-4684-9168-b3a6e7bac4a8 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Decoupled weight decay regularization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93e0a649-9035-4687-808c-41c3eed86e4d · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Image segmentation using text and image prompts
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d68d289-afc7-4060-af7d-2faaa7000cb5 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning word vectors for sentiment analysis
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52dd46dc-1bfa-4692-aa54-33ba31143ca7 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Fine-Grained Visual Classification of Aircraft
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad756d82-f0fa-4950-8361-5aa369639e19 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Slip: Self-supervision meets language-image pre-training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 30629513-27b0-4a96-a759-3b43540c0eb0 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Automated flower classification over a large number of classes
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b433afe8-3ce4-4f3a-a751-116285151618 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Deep metric learning via lifted structured feature embedding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9cd5f98e-8640-4594-a7a9-b3b55ac0fb10 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Clip-guided vision-language pre-training for question answering in 3d scenes
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43f0499c-47ad-4c36-973b-d9908769a5c9 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Cats and dogs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8336ba0-9c42-49cc-92d2-75f4c91b39d7 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Eclipse: A resource-efficient text-to-image prior for image generations
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1d5606bf-8a1d-4a78-8014-a40faae32aa0 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e71decce-b0da-4846-bd81-5920527175a3 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Revisiting oxford and paris: Large-scale image retrieval benchmarking
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d8504b3-ee9d-4792-8b83-9c299a793a3f · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning Transferable Visual Models From Natural Language Supervision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aea494b7-c64a-473e-a9f6-36c9950c2d25 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a37b1b-bd28-48f8-b7ac-0a1f0cd53173 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef2bc42-8645-4506-9f70-e9690e01dfe7 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12435e46-267e-49ac-9880-830c2f685cb7 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14bfbac4-28ab-479e-bb57-685723b527d1 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6c45df0-393b-405a-8615-fafd2d6bacd6 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Towards Understanding the Modality Gap in CLIP
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 07ff24e2-ce6f-499d-ba4c-4df030201343 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Improved deep metric learning with multi-class n-pair loss objective
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation adad59bb-9f9f-406c-9c95-8386147827c9 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion CLIP Models are Few-shot Learners: Empirical Studies on VQA and Visual Entailment
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edbcfd55-15c3-4edb-b6ed-461eb7da6929 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac08640-37ec-48d0-b3ff-5667bca5d0d8 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion SuS-X: Training-Free Name-Only Transfer of Vision-Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5efc861e-25af-419d-b972-9ab6a20fc525 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion The caltech-ucsd birds-200-2011 dataset
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53627af5-6126-4204-a195-5dbaf1e5e96b · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11fbbb18-74b0-4e16-a09f-43f52b2f9afd · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Coca: Contrastive captioners are image-text foundation models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 692e86c5-9045-40d1-9ca3-544cc62727a8 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Lit: Zero-shot transfer with locked-image text tuning
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dc84c6d1-b7e2-415e-8467-e0ebbee1d78a · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Sigmoid loss for language image pre-training
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c93fef-c480-4d13-886b-c78aefefe651 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Diagnosing and Rectifying Vision Models using Language
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 18179f04-a4c0-4af9-925e-1a51f971e059 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Avid: Any-length video inpainting with diffusion model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5538da71-045c-4d56-bdb5-3b5ca66cb48f · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Extract free dense labels from clip
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7847be78-e3e5-45a1-b39d-08130c794dbb · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Learning to prompt for vision-language models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbbb8424-0553-484c-b80a-48770af66919 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion @esa (Ref
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2965186f-a05c-4125-a819-6c8c989c558c · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647fabc1-5e1a-4907-aeef-00cde2700bf4 · outbound
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion a photo of
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 421bfe05-7e31-41b5-ad92-4b487724c3f8 · inbound
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3bd064c-3c80-479b-a405-5813f35c1596 · inbound
Global and Local Entailment Learning for Natural World Imagery Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f64889f6-b20e-45ba-98ec-b6bcaf0766dd · inbound
Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9528a14c-23ee-49b2-b3de-e52fd5152860 · inbound
Best Segmentation Buddies for Image-Shape Correspondence Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 963adb28-36cf-4c0d-894e-b1dc113a3e37 · inbound
LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.