Pith. sign in

Paper Citation Record · LEDGER

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

As of 10 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 4 inbound Pith citation observations for arXiv:2506.02557.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02557 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:14.795852Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T22:07:35.021986Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T22:09:07.132811Z

Reference resolution

98 of 98 outbound references displayed

  • verified exact4
  • verified fuzzy35
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d16e6ee-d894-4bda-8dcc-7acfeb66ae2e · outbound

This paper cites Tallyqa: Answering complex counting questions.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Tallyqa: Answering complex counting questions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.076810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.076810Z digest=sha256:85f1759bbd269e555eec8df0be1c4ca54d5b5ec536860a72fa3b512ac6d8818a

Observation 9ac6405e-c393-4fe7-be5e-44e9302a72b6 · outbound

This paper cites Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.168645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.168645Z digest=sha256:9a70c269ebeafdd05c00f38bb3171a4fcd4b5550a2df748f599a355ed5e6438f

Observation 5b8e08a7-39c5-472a-bf0b-f4305a43ef72 · outbound

This paper cites Multi-label cluster discrimination for visual representation learning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Multi-label cluster discrimination for visual representation learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.268183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.268183Z digest=sha256:3b313df6c018ed51648aceeee82e596ff5224c018fded55a33e901f46c2a4a96

Observation dd2aea2c-c9b9-4bd1-894b-5ac1d6191e03 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.323907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.323907Z digest=sha256:bbe8658902fcdbd85eae7313d8013078d8009e04d73e939067397609ea9d6e72

Observation ae69c505-1c92-4049-8542-abc0a16de6f6 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.361069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.361069Z digest=sha256:e9c87f7c4009360d5288b9493d1b558b414578824e156f43a74f363682fd602b

Observation 93b3c931-1656-477b-91ee-e9678554b1ea · outbound

This paper cites BE it: BERT pre-training of image transformers.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models BE it: BERT pre-training of image transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.366454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.366454Z digest=sha256:46a7c505fb1a270b2c56233f40f72268fbb243ce5ceba040f301302b94489715

Observation 15631b84-da0c-4158-890f-8af4e5d2f0cf · outbound

This paper cites Demystifying MMD GANs.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Demystifying MMD GANs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.372034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.372034Z digest=sha256:709db19a3a99bf0beddc124f89e82308df127585e31a1c95025542762ba0245e

Observation 687fedb5-d726-43a1-a2cd-2cb12ad73329 · outbound

This paper cites Domain prompt learning with quaternion networks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Domain prompt learning with quaternion networks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.379895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.379895Z digest=sha256:d2f3559bbfa9d731c5d0b7a4dba404149c1d0c7c4d9f8034af2d16bc63c51847

Observation 18af799d-c07f-4a26-8766-b41a6ea96b33 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Emerging properties in self-supervised vision transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.388470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.388470Z digest=sha256:a16d9cdf2912f767c093d513f094ad08c4b441f201c9f8dd029cf742b8bfaec9

Observation 4ae0c78d-9be3-4ee3-84a6-54fc288609ed · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.395742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.395742Z digest=sha256:7c62ccc570605fd16f7a45e122aa449d1f686b7c798f845cdbe3843972642f10

Observation 5bc761d1-b06c-4caa-a25b-ce9284b1b416 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.404320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.404320Z digest=sha256:3cc44cc31391bed00d152b63c58f9b8e28c7a744185ab3607a698cf80d6ee93e

Observation 9aa90695-3be6-400e-943d-09083641276e · outbound

This paper cites Remote sensing image scene classification: Benchmark and state of the art.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Remote sensing image scene classification: Benchmark and state of the art

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.413877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.413877Z digest=sha256:ae561730f1fb0f7a8c0aa85fe40635cbca7a3bb760a1a79c6bffca64b1fc1650

Observation 2bc4a2da-efac-488c-a0ec-b5892fa02b4a · outbound

This paper cites Describing textures in the wild.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Describing textures in the wild

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.424318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.424318Z digest=sha256:9c13c05532ef2a13a4f295147df0eb9acc1a576e8ee5392d5962bdc658352cb2

Observation 01f684b7-2458-4e0c-844a-5c245287180d · outbound

This paper cites Locality alignment improves vision-language models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Locality alignment improves vision-language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.433254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.433254Z digest=sha256:f410b4659db8013ca7c6097c4e412bfc6b8264aa80695ec3e6f92d361ae41aa1

Observation 9ea5f798-9654-40d3-b615-2e300fe98c8e · outbound

This paper cites Vision transformers need registers.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Vision transformers need registers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.442273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.442273Z digest=sha256:291c49109915b776dfe66211a106bc978ab10875357fe3dbc18177fdd80b2079

Observation 76f1de1f-3955-4271-956b-0e71f1f8478a · outbound

This paper cites FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.449072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.449072Z digest=sha256:9ff27ad69a7c9cef1e869ad1e02a21036fdfe278833310175f8fa4095cc3c2c8

Observation 29001f20-7d10-4a44-aa2e-3a212bc2928b · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Imagenet: A large-scale hierarchical image database

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.458149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.458149Z digest=sha256:2e07d69bdecc2e55bb0edc545b3c6ccc57e31a6dda2e5c2675730e94bf0336ac

Observation 0f19e8af-fdf4-4222-b513-e6808e566af5 · outbound

This paper cites Data Filtering Networks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Data Filtering Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.466593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.466593Z digest=sha256:2e1c4cd5a7efd46d21be8df4dcda8be22b328348a32ef4585ccb7e7c0b27019e

Observation e1768fd8-cc1c-4a4b-8fb5-cac422a01f1c · outbound

This paper cites Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.474213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.474213Z digest=sha256:0682443a485bbd33204934522ba88e0a9c19d092d029615d5305e8c277964af7

Observation f27885c4-874f-4f66-80cd-ca21c6c67bfc · outbound

This paper cites The Vendi Score: A Diversity Evaluation Metric for Machine Learning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models The Vendi Score: A Diversity Evaluation Metric for Machine Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.482039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.482039Z digest=sha256:58bf4ef4280f03d03aaa1c69f9bbe9ebc0f29e2af9173669d2da54e48861bb51

Observation fe8321d5-a764-48b8-9d49-79270960917b · outbound

This paper cites Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., et al

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.489672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.489672Z digest=sha256:4092792d4c90c193d642ccbaab943d2937428b0c3f9bd097370492fd31a097d5

Observation 5e9ed5a2-454f-4ca2-afc6-0660046e0e4a · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip-adapter: Better vision-language models with feature adapters

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.498808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.498808Z digest=sha256:b513a0a914fdf7d465c0b73aee254be8a9b498287a25ba269e6047fc332e13c1

Observation 01bb1b83-88d7-447e-a6f6-e6d5055328d6 · outbound

This paper cites 3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models 3dsam-adapter: Holistic adaptation of sam from 2d to 3d for promptable tumor segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.507737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.507737Z digest=sha256:77d1dd97a2b8ab7c1e437f26486a70bd1c00fbae469a47c42d592cabe47e96a0

Observation 1b00e8df-c653-435c-8b8a-207190ba5e94 · outbound

This paper cites Boosting the visual interpretability of clip via adversarial fine-tuning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Boosting the visual interpretability of clip via adversarial fine-tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.514659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.514659Z digest=sha256:f290054755e1249d34bd45c5a9fd086028d9c2937604a6b5720f2cc7269b489e

Observation d9c80418-3b83-47c1-9d71-9fe959274c87 · outbound

This paper cites J., Erhan, D., Carrier, P.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., Erhan, D., Carrier, P

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.523302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.523302Z digest=sha256:cd8b42ec244b5b94f68d20311fc3e8f3a6b048d97eb8aa9320f873c4648eb233

Observation 6930017c-288a-4408-b09a-498b9b8db2e4 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.531936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.531936Z digest=sha256:8f02f0538d09930e1df8efea86e94196ad245d2a05987fe5309800e8efb58840

Observation f2ecbac2-a02a-4822-91ae-6e21c4010444 · outbound

This paper cites Recovering low-rank matrices from few coefficients in any basis.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Recovering low-rank matrices from few coefficients in any basis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.543285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.543285Z digest=sha256:dac157ce5a18bda66fe3f1d9ea2ce57d7815d1423d66c9041e5d8f246834829b

Observation 366ffc2e-2196-4e70-8d5f-1cc98c69c510 · outbound

This paper cites Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.322547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.551090Z digest=sha256:d8a391a3c8d71577bb57702cb91e451c62dc4661759c7d330daf3beb5beeb28e

Observation 3a7fc68b-0e04-4eca-a9b0-df2e2203c2b7 · outbound

This paper cites J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.559218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.559218Z digest=sha256:9390902cd89f05d7095993ffe9a7d46bb6c73e02e50f35b47754101c68517eba

Observation b69abfb6-bd6c-44df-9b2b-e0d23c980b20 · outbound

This paper cites and Ozay, M.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models and Ozay, M

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.293617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.569083Z digest=sha256:246e5943552f39601ed06cc636fd18ab0db1afa276541b984f8ac90ddb0a8e72

Observation 121c676e-fa5f-48ad-9c50-690fa20805bb · outbound

This paper cites Masked autoencoders are scalable vision learners.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Masked autoencoders are scalable vision learners

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.577818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.577818Z digest=sha256:c4f5a0e07fc4ee3857c9ca2c1021977ea35d9411202336020a9404dfda52cc2d

Observation 2a254a0e-9dcf-4e16-949a-9b1c1abc6c1c · outbound

This paper cites Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Introducing eurosat: A novel dataset and deep learning benchmark for land use and land cover classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.263624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.585481Z digest=sha256:839d164b01432ab12fe70c0eb530e87de165794bc106369926c845799bb98297

Observation 4d7b9426-0970-4aae-b9d7-3d9b4b0597f1 · outbound

This paper cites Natural adversarial examples.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Natural adversarial examples

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.245014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.594682Z digest=sha256:8374caaecc37d88efa6c4e38a024f09e4a5e4a8413b6621163b916d33453cf26

Observation cf580be8-ad12-443a-9c32-81712a0531de · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Probability inequalities for sums of bounded random variables

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.227681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.603893Z digest=sha256:003dc77ed67b5b83651f458e86508077618d8f2f52b2f512cfe74a55328951fe

Observation 4416e393-5b99-45c6-9ea9-dba6335057b2 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:16.203710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.611022Z digest=sha256:a13f95936e376e4a7dec615ddd31ecc7f09b590be0ab7adcb41efc35998e4a9c

Observation dc363e2b-3006-4734-a536-73a19962b57a · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.622913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.622913Z digest=sha256:897be210f757ab4cb00cfd968f8dcd6e83a0e93f0716d4b22b80551607b13cca

Observation ec4452a5-ec39-44de-b667-91adde54424a · outbound

This paper cites T., and Farnia, F.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models T., and Farnia, F

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.171988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.629823Z digest=sha256:05fff6288f7d3f99feab6e6f332953ba88a46bb077384319eb5985cc015d7786

Observation eae863a3-ed91-4854-9112-b5ceaa723332 · outbound

This paper cites T., and Farnia, F.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models T., and Farnia, F

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.156234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.639656Z digest=sha256:cb7fd52130b93db986cf3f68d2d531f6a8076ea92edcc8027502a8b727129925

Observation dbba5313-2aad-4768-8d8a-e44aa9ffa657 · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.676257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.676257Z digest=sha256:75ecd5098506a7d8f35b75b965bf84a5488e6cc4fb77111cde390cda537054bc

Observation 206e2dba-54a7-45c6-8e8b-d3e48972387e · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.137088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.710170Z digest=sha256:914cb587089331706c7f7cdea3cad1c651016eb5a4b3d57ac3b9b016b7e296d8

Observation ad4df0c3-f888-4b8c-81b7-7cf62eae08cc · outbound

This paper cites What's ''up'' with vision-language models? investigating their struggle with spatial reasoning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models What's ''up'' with vision-language models? investigating their struggle with spatial reasoning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.118412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.745597Z digest=sha256:58934c3a189a6cc1af27ab60a8a3c2f2c511bfd02c11ce709fe78ce00eeeb4a8

Observation bc348745-03d8-4a07-95c6-203ae6f6619c · outbound

This paper cites Studiogan: A taxonomy and benchmark of gans for image synthesis.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Studiogan: A taxonomy and benchmark of gans for image synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.769979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.769979Z digest=sha256:5c5bbfb7693c5692be82336ebe5251e2f01256e5bee696427cb0ae9a0b8f1796

Observation e2a1764d-f340-4029-ab33-673f15dd52c5 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Referitgame: Referring to objects in photographs of natural scenes

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.098644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.788031Z digest=sha256:7c119725b78261c7cb4301131edafa4b668e5a6e847909bee4d99bdfa8f92a3b

Observation fd87e030-16f8-4a32-8da4-d7c713a2ebcb · outbound

This paper cites A diagram is worth a dozen images.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models A diagram is worth a dozen images

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.076987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.814240Z digest=sha256:fbfb9c20a2d8d7b84d34c20d16bde2664e948796436c1d0357d34f1ee94a9e4d

Observation 1d4571e4-a51e-4c97-a34b-51d71ddfc5ed · outbound

This paper cites The hateful memes challenge: Detecting hate speech in multimodal memes.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models The hateful memes challenge: Detecting hate speech in multimodal memes

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.058144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.852337Z digest=sha256:a41cb35918aa28ac82251dedd59a4bd844d8a57b2023f28b1dee954b6e8742cb

Observation 4f904f8d-1662-4281-94b0-c76d3ac765c4 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.878916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.878916Z digest=sha256:d77c4499210179ed56ca38cb79e4966fb1c8911f831d2a9a067954f23cfc6fd9

Observation da99cb35-99f2-40e9-b0d1-58ece4309242 · outbound

This paper cites Learning multiple layers of features from tiny images.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Learning multiple layers of features from tiny images

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.897490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.897490Z digest=sha256:f1325e8f50ef61070cf96e9b86cb24b8c4065f6c9efe0307cd498f48e7747e73

Observation fdc3bc36-8274-4109-a257-e0b499b0d560 · outbound

This paper cites Clip benchmark: Clip-like model evaluation, 2022.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip benchmark: Clip-like model evaluation, 2022

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:16.019546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.931847Z digest=sha256:a68df76e9d7e4465c058acfe9369b6cd97a19f10807ddbb6b784d39f1457b82c

Observation 986ea017-a1a9-423e-a162-80aa191b2015 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.958430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.958430Z digest=sha256:36ad2965f17c3d160178f3f73f9cc5bd4abd96f1dcd75344049193619fb7fedf

Observation ab00bb5c-ea22-4fe4-b469-420a53fd3470 · outbound

This paper cites Transfer learning in computer vision tasks: Remember where you come from.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Transfer learning in computer vision tasks: Remember where you come from

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.987195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:13.985748Z digest=sha256:8b1dc89f1e9d2ecad400e8b9fba1bf2309fc45a20882e6eaa729134571ae27dc

Observation 0e1a228a-284d-4834-acf0-c75ca3829e9a · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Evaluating object hallucination in large vision-language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.971091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.022976Z digest=sha256:b43a8e2d7f92b5aac317ae04f92a14b6e8f3c57dda420b127a61eb17dcb373f0

Observation b55b8aa9-9a3c-492d-85ad-173108fd5013 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.058425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.058425Z digest=sha256:c2ee42d3b21f232735e096362fafbd397d3468e26997a5b3459f72a6820859c3

Observation 941bbed5-eb97-4269-8ad7-47c1f58e61e1 · outbound

This paper cites Visual spatial reasoning.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Visual spatial reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.943536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.093939Z digest=sha256:902741eab33453e6aea00656d39eede9a3cbfb24015ec62360dd8ce03b8ad5e9

Observation 95581358-a2dc-4a8f-8e01-5f39a1289963 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.130078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.130078Z digest=sha256:b1b7c32057a35b56c6ae575ae650715fc1f7f9a9e68da8b7ba18ea3bf8f801d6

Observation ad658fa4-a1d9-4e11-88d2-2c6fa4d826b1 · outbound

This paper cites Decoupled Weight Decay Regularization.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Decoupled Weight Decay Regularization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.157884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.157884Z digest=sha256:de43053b18d90cae30bec8e51fae79bd4c41763aed0e1df02e143004abd1572c

Observation 2b9b5496-ad13-4cc1-b52b-2c82d1c81616 · outbound

This paper cites Understanding Zero-Shot Adversarial Robustness for Large-Scale Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Understanding Zero-Shot Adversarial Robustness for Large-Scale Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.191544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.191544Z digest=sha256:cbf1910c45d7059581e714d90dc4b6d4fa63713e11ce0a7df3f139a37632ca7d

Observation 97d40f26-43b0-4ba1-b6c8-bbf546f4a1d4 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.214004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.214004Z digest=sha256:57fea8feb1555c40382c194397664cca2139b404b472a11a7defa5fcf6b12840

Observation 53d26859-b591-4ff0-8722-4807ae2a4770 · outbound

This paper cites Y., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Y., et al

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.234636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.234636Z digest=sha256:7765cdcec414d6f1e1d631dd7f54740da8cf871f4a5769fad4d6dd591aa894ba

Observation 76527488-a35d-443a-aab9-e2a81759f282 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Dinov2: Learning robust visual features without supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.894122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.250170Z digest=sha256:ae086a8dbb3a59b66bf06ad28351a4a2a8bcd8e1101e56674ca87f9570511e06

Observation dcf1e64b-a01b-4377-bd5c-3489b89a64c3 · outbound

This paper cites Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.272780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.272780Z digest=sha256:1c6e28f8ce26dac065a6bbcb5bfd835a05a75943a990e1c55ed53e9bb82907c4

Observation 1a5fae69-ee38-41bc-b5dc-14c10c800486 · outbound

This paper cites Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Scendi Score: Prompt-Aware Diversity Evaluation via Schur Complement of CLIP Embeddings

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:15.081496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.289306Z digest=sha256:adc8782e18d385c26ca1a088099dc0f814fb5ad1d2b03f72f5e8f4660740f200

Observation 74d19f35-21aa-471d-910a-247e7ee840c8 · outbound

This paper cites Towards a scalable reference-free evaluation of generative models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards a scalable reference-free evaluation of generative models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.875218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.310303Z digest=sha256:d069ffee3e8b130a83c0a9f8d092e43c5f8a3b84bab833fe1f06df70a9b0a1a9

Observation 43bac356-ca4d-42a1-81aa-2d9b908b5fc7 · outbound

This paper cites M., Vedaldi, A., Zisserman, A., and Jawahar, C.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models M., Vedaldi, A., Zisserman, A., and Jawahar, C

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.857821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.315818Z digest=sha256:2c7aba2ace3d580138f456dcec9c13b93536ae62eb4024e384b31c9e0499ec5b

Observation 2618be62-08dc-4987-b436-f6c8b8b33a5d · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.327491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.327491Z digest=sha256:393d4911d20adca08fc35b4ee4467cc9f28a07d770cd2211575768e63016e4a6

Observation 36b7853d-b122-4b04-98f4-4a8a043db632 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.828165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.337840Z digest=sha256:75f64870b1b93cac29fd67ad9528bb5b9e57970aa97773ba9f7f86498561f9ff

Observation 918a3279-df71-4ef6-937d-790a1d9df14b · outbound

This paper cites Be More Diverse than the Most Diverse: Optimal Mixtures of Generative Models via Mixture-UCB Bandit Algorithms.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Be More Diverse than the Most Diverse: Optimal Mixtures of Generative Models via Mixture-UCB Bandit Algorithms

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:15.058806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.349429Z digest=sha256:6d125ccbe43703d916bcbf3f1f3eb0eddd206667885a562c62310983f61c9fb4

Observation 6ac9159c-f5d3-405d-81b3-fcf795583ee8 · outbound

This paper cites Improved zero-shot classification by adapting vlms with text descriptions.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Improved zero-shot classification by adapting vlms with text descriptions

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.809295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.362400Z digest=sha256:05317fbe0d0ed17a272b2d575a1ca3cded25dd2ae5153994f755be938404b77e

Observation 4e5c131a-6d20-4eff-b873-ca932678061f · outbound

This paper cites CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:15.035097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.375483Z digest=sha256:79f56d4b11a995a6c41847e00244f226eabe1f5e74f50ca0835b3ec605123cd3

Observation ebf584df-d99f-4edd-a28a-dc9ee716d36c · outbound

This paper cites Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.409174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.409174Z digest=sha256:3de674964bcd760bfd562aa8d57e6cbcb3bee2e0aaa429309035e539f696d436

Observation b6b6a7b7-2294-4b70-ba5a-04f1a26dbd1c · outbound

This paper cites MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.425638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.425638Z digest=sha256:c9e02d1c9952d7f9cba4ee757bc82d16740ce4f350e81b5f20fc5ccc69cb0658

Observation 73f02e0a-268b-47b9-8f9d-0ec95bc9ce56 · outbound

This paper cites Finetuning Text-to-Image Diffusion Models for Fairness.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Finetuning Text-to-Image Diffusion Models for Fairness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.450325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.450325Z digest=sha256:0135c14040f1a8cd98ab58314980c51b37a0288553f88e48ad1558eb1e7c2132

Observation 5efff50b-5682-40c5-b5d9-0ed382beb549 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.486555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.486555Z digest=sha256:8b83940caf9add086be9e3d5351a955929d403e981529908fa7832d880f41dc4

Observation 86c03fab-1dc6-414c-9dd9-3cf325655264 · outbound

This paper cites Towards vqa models that can read.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards vqa models that can read

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.514075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.514075Z digest=sha256:fb31de5afdcd57057a2d6efcff78d0c40b8cb90ed37c7d33b8826be67d4a665d

Observation d5eeb4e4-5a06-4026-845f-44b8f6fe2d19 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:15.779942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.532330Z digest=sha256:5eeb3b0c2846e291086bd18bffe868396b01120082344b4db2f91b5b1704b0c3

Observation a43fc90d-4703-49a7-8e49-220310d8585c · outbound

This paper cites L., Taylor, E., and Loaiza-Ganem, G.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models L., Taylor, E., and Loaiza-Ganem, G

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.758280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.567592Z digest=sha256:60a1c4c841ae5c8f317bd11e051942656b21533b5f451457105f2bfd9100a919

Observation 92360aea-3301-4b3a-8ae3-14ba88afdfd5 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.595926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.595926Z digest=sha256:92b965b796669107899e34e890628b52f0287a9f256f385703c726b9bc71f14d

Observation 75fce3a4-6058-4509-a68a-2ce192512d2d · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Winoground: Probing vision and language models for visio-linguistic compositionality

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.740074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.633074Z digest=sha256:0794db7ba78c461ec8c601ed61af117040733eb7116abff92adaec0073f2ba9a

Observation 8a11379d-27de-4437-81c4-d11d76d335b6 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.668972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.668972Z digest=sha256:8ef81500e08bc8832cc4f910c8e71f5fdbb1bea33d8cfade6cf99e8844335620

Observation 8958d4da-88b8-4d83-a640-75d27891eb51 · outbound

This paper cites S., Linmans, J., Winkens, J., Cohen, T., and Welling, M.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models S., Linmans, J., Winkens, J., Cohen, T., and Welling, M

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.709139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.688699Z digest=sha256:2c101d81f3693cd5265598787189c5c81e302beed7707618d01e67f051e014f2

Observation bd473008-39b0-4582-86b0-9cf57d3e5756 · outbound

This paper cites Clip the gap: A single domain generalization approach for object detection.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Clip the gap: A single domain generalization approach for object detection

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.691181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.693990Z digest=sha256:161445e0d906f8b9d0d96740a814f2d02a73823e821d5e7b756cb80d39bf67b5

Observation e5740819-6d52-4e95-8a60-38da9a6d39f4 · outbound

This paper cites S., Steiner, A.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models S., Steiner, A

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.673241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.699343Z digest=sha256:efaac4430ce86c6f3b4ea9797a59433e24a0cfaf0476f752f79ffdeac2212833

Observation b7a586b6-5dc2-43c8-bc52-d27d87f85af7 · outbound

This paper cites an unresolved cited work.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:15.655991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.708069Z digest=sha256:960a843ebed2cb83ccdec8c861cad20aa7e00b3accca5ff647f8476822fab8fc

Observation 65186142-1aa3-4919-9f62-ca235b46f4c6 · outbound

This paper cites Diffusion feedback helps clip see better.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Diffusion feedback helps clip see better

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.640372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.714678Z digest=sha256:08df512cd41141a1c268bd4c1bbb4ae96516644dadbf1323be424466c4482e21

Observation 62d63c46-d754-4f26-828f-ce1d23533941 · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.720443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.720443Z digest=sha256:c39f347cd7470c90d3982cce5ea269bbc2cc0f73389c15435d1a23531d9e50b0

Observation a43c192a-2ac7-447c-a69d-9cc493a6cfcf · outbound

This paper cites Demystifying CLIP data.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Demystifying CLIP data

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.623105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.725923Z digest=sha256:ad47b8219b66cb59a4d248a0c7a13790b768ebfd2dbea25250d9ba8fc918d220

Observation c367537a-4415-49b2-9169-6420db8b0cea · outbound

This paper cites Explicit inductive bias for transfer learning with convolutional networks.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Explicit inductive bias for transfer learning with convolutional networks

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.606495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.731229Z digest=sha256:2b4295089c84fbb1d911b0b26a501c77974592962cc94d1c2c475b3185ea1180

Observation 984b5b48-16a6-4047-ac45-5dbea0e582a2 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.587517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.736867Z digest=sha256:f7d18bc74097bdee19dad5c7ed8a92b5ad0108b65f00378f4c133e545abebc32

Observation a6135983-a5c7-4bad-a1fd-8200c9309782 · outbound

This paper cites C., and Berg, T.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models C., and Berg, T

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.570210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.741696Z digest=sha256:e6da480c9a51813921c803c8afc33d98c72a506d59cc219ee99e934a57be64dc

Observation 74155806-e21b-4baf-881f-117b867164ee · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.554372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.746368Z digest=sha256:0f195af70faf23e4dec98f81e99f535b7cfba6fc700c10c3bf475e2ef6fa8934

Observation 215a7a43-1c7f-4c53-bc74-17924b037985 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.538506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.752572Z digest=sha256:d23fbb2f307c5f28afa580d52f30cfd22d9495a5c22acc4250315783faa8f0a6

Observation 2cb6e895-9f87-4c8d-8c4d-69b22d6ba8ec · outbound

This paper cites Sigmoid loss for language image pre-training.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Sigmoid loss for language image pre-training

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.757778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.757778Z digest=sha256:d64835ec59ac63cea0ee6c6c913b0930ddc98ddda55ffa5dd09395db4e8e32ef

Observation 4bfa7c29-6693-46f8-b2ca-bd214c111d9c · outbound

This paper cites Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Unveiling Differences in Generative Models: A Scalable Differential Clustering Approach

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.763166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.763166Z digest=sha256:ac0525a5bb3259db1ae992b0b8e846c15325128720f85eb98444e7af624e8f39

Observation f147c74c-a66a-443c-a7e1-333c75204a53 · outbound

This paper cites An Interpretable Evaluation of Entropy-based Novelty of Generative Models.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models An Interpretable Evaluation of Entropy-based Novelty of Generative Models

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:26:14.871675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.769062Z digest=sha256:8c01f8f86c97dfb790d018167f1877cf7d2926a33943da05fc10aa1fa106c641

Observation dc3159f4-8721-495b-b9dd-4cb65b7c68ef · outbound

This paper cites Tip-adapter: Training-free adaption of clip for few-shot classification.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Tip-adapter: Training-free adaption of clip for few-shot classification

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.510047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.774232Z digest=sha256:9918c331cce6030227a10a991a848bdeaa10e98205cf471566b492fe103487f6

Observation bd3c76e1-ac43-4256-a6c3-cb6b3b7ec446 · outbound

This paper cites H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models H., Zhou, L., Dai, X., Yuan, L., Li, Y., et al

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:15.489436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T11:26:14.780451Z digest=sha256:2d8217ba2373675d722cdb2d4e6384a6cc14a753dbe757b77003e6e1855d1856

Observation aabcead0-936a-4ad9-9929-e5d6cdbdf76a · outbound

This paper cites C., and Liu, Z.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models C., and Liu, Z

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.786001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.786001Z digest=sha256:cea8ff1b460fa7e9de4fd0ee901709e18ee95afa245edc766b33d8808654fb63

Observation ef46bbe8-0ec8-4a86-b913-da9bd885d844 · outbound

This paper cites Rethinking Centered Kernel Alignment in Knowledge Distillation.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Rethinking Centered Kernel Alignment in Knowledge Distillation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.790735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.790735Z digest=sha256:7e846293a0ec5a882472bf9533e8248972316baebed2e8bc3cf308b66c4b746c

Observation 32880931-6e19-4a41-8cc9-ebf2088f0d25 · outbound

This paper cites write newline.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models write newline

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.795852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.795852Z digest=sha256:68c0bbc8556f60d8a1da3875beb0a8a9aab0003878ca200dd556dc51869c71b0

Pith citing papers

Observation d62736ac-a819-41bb-a7c7-eb0abf546778 · inbound

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels cites this paper.

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:52.003486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:32:36.729122Z digest=sha256:a23687c57f5a1befc44cbd31d88a2c9b4b7ec91bb7b9bc79761d85ce6a5a5f91

Observation a3fff1ec-2d20-4ddc-ad54-cc96cf462af6 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.700458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:d1b6c0af2a54186b9e6a1fc830c49ac16810363ccd94a30c883d1d8bd706063c

Observation 5b84b43d-5e64-4629-a595-4b924cbc06b4 · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:51.075371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:14:03.809826Z digest=sha256:8f198f665fb27b025bd5095e1aebea085f460db8f06ca09e1f817ba9cc15d498

Observation 885178fa-8e7d-4394-a495-a60a3a927002 · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.135657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:07:35.021986Z digest=sha256:03557aeb4a06dcf193ffd994c76caadf09a1f85fe81ebec9f7f335d1b4c60450