Pith. sign in

Paper Citation Record · LEDGER

Illuminating Visual Identity in Universal Multimodal Embeddings

As of 13 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.01794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01794 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:43:17.507439Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved55
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2363b7d3-81bd-4afc-a3da-8fe99693d3db · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Illuminating Visual Identity in Universal Multimodal Embeddings Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.346553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.346553Z digest=sha256:02656bc7b946b3828263edbdc4b5e9d21fdc4c8f83a0ebccd3e3ce06cd2382b8

Observation bd4f6e0e-b313-4b2c-89cf-b70b0de3e3dc · outbound

This paper cites Unicom: Universal and Compact Representation Learning for Image Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Unicom: Universal and Compact Representation Learning for Image Retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.403233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.403233Z digest=sha256:5d2d1f4513fb27abe9be320b1b59c039a8f97a67fd6212a5cf27c5f6cd42d481

Observation 88636538-b85b-48d0-8d11-55c14ce56b5b · outbound

This paper cites Qwen2.5-VL Technical Report.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.493030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.493030Z digest=sha256:be1ce2d619ac91e87c4aa05b06ca24259a0b54541500e4be27f168335b2b320c

Observation 6dd964af-ecc0-4d72-b6a0-0b501e81ebdb · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Illuminating Visual Identity in Universal Multimodal Embeddings LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.611224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.611224Z digest=sha256:22056bd10306ba12f5ba10b8995977993086f8870e3ad9e66bc62b508fcd3d3e

Observation a64de592-e3dc-479b-8d0b-49625f0cf298 · outbound

This paper cites Flame: Frozen large language models enable data-efficient language-image pre- training.

Illuminating Visual Identity in Universal Multimodal Embeddings Flame: Frozen large language models enable data-efficient language-image pre- training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.727986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.727986Z digest=sha256:6ce72fb8ccfa44654b21781b7b94f7e224098e7ee57e55218e2f02ec89d89e97

Observation 9e7d481e-e174-4d49-a57f-18746709650d · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

Illuminating Visual Identity in Universal Multimodal Embeddings mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.812234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.812234Z digest=sha256:eb8709d021a6beb781b1e3f1ddba29ebb6cacfe5ecb3f12dc0eb7b09bc9ffcbd

Observation ce543fff-5dde-46e1-bc81-2031f0eb2874 · outbound

This paper cites Murag: Multimodal retrieval-augmented generator for open question answering over images and text.

Illuminating Visual Identity in Universal Multimodal Embeddings Murag: Multimodal retrieval-augmented generator for open question answering over images and text

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.932581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.932581Z digest=sha256:ff626aee0247b50317d83e3ec9bc9c501bd34f5ca35873576f94a0bb2a3fc8bf

Observation c0ab3143-6da5-4689-99ea-34a6d71ff4a8 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Illuminating Visual Identity in Universal Multimodal Embeddings Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.058096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.058096Z digest=sha256:988096a6a713f31624d43c259f7694adef5c498165b60f7e845583b3eef6c20e

Observation e2d25ea6-cd73-4d60-9e7a-3ef01090b387 · outbound

This paper cites Opengpt-4o-image: A com- prehensive dataset for advanced image generation and edit- ing.arXiv preprint arXiv:2509.24900, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Opengpt-4o-image: A com- prehensive dataset for advanced image generation and edit- ing.arXiv preprint arXiv:2509.24900, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.225875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.225875Z digest=sha256:cc712f0b07c1b1453a6a114352432977b7d7e2132d5d6ee628ba6065879d24ea

Observation ced016ae-bb8a-4efa-8861-ef6d4e4a0e19 · outbound

This paper cites Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014,.

Illuminating Visual Identity in Universal Multimodal Embeddings Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.394360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.394360Z digest=sha256:99b829d72fc2e82d4bd451b9238564f85849f6826311c6f48238108067b1a5c6

Observation 1503431a-f199-4920-89fd-2aa1bdbb0418 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Arcface: Additive angular margin loss for deep face recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.494829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.494829Z digest=sha256:6a192184d74f7d63acf13e548b137a7489d3c5ee24291cf56f857af2175085c3

Observation 84c5eb15-19db-4409-9a10-8ec3f53b88d7 · outbound

This paper cites Efficient and discriminative image feature extrac- tion for universal image retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Efficient and discriminative image feature extrac- tion for universal image retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.665447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.665447Z digest=sha256:8924f3541dca87f73c47868e690059f891e4f58575fe7157066e9f067a359e9a

Observation 6675d07d-4f80-441a-9de8-552a76f22541 · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.

Illuminating Visual Identity in Universal Multimodal Embeddings DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.825424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.825424Z digest=sha256:7dd7b43ec663bf151377d201dae0b397ee2bf432399a5e2d2d1574deab06a471

Observation 3e1ac40a-362b-4648-b680-886b4dcee7b4 · outbound

This paper cites Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup.

Illuminating Visual Identity in Universal Multimodal Embeddings Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.991222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.991222Z digest=sha256:34beb0b695177a557c4f4915449e3cfe3b102f55278d88e5199318da51e3b310

Observation 036f4965-2c3d-44eb-95c6-f35eb5cbe0a1 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Illuminating Visual Identity in Universal Multimodal Embeddings SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.155373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.155373Z digest=sha256:ba278baf21274af1408e481bcdb20701a22c63e4bc3250d212181752519e34fd

Observation a2d55248-5423-4de2-bd52-c4b22dd1ce2b · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.239595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.239595Z digest=sha256:28fddac6b36ac6bff707e3c808193d1353fccb0f228e3422bd61ce992187eb06

Observation 809043bb-8bd1-4a11-8630-d35cce3060a4 · outbound

This paper cites Unime-v2: Mllm-as-a-judge for uni- versal multimodal embedding learning.arXiv preprint arXiv:2510.13515, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Unime-v2: Mllm-as-a-judge for uni- versal multimodal embedding learning.arXiv preprint arXiv:2510.13515, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.281060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.281060Z digest=sha256:aabb7cec3c08f0db174dc38e56cb7ab7d8883901b365ddcd3e872a496df88cb4

Observation c6a7724b-16fd-42e0-8e92-c11485f0cb06 · outbound

This paper cites jina-embeddings-v4: Universal embeddings for multimodal multilingual retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-embeddings-v4: Universal embeddings for multimodal multilingual retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.356047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.356047Z digest=sha256:125cf18a28df7801df5f71371386e870564fe0077ec13a7540b662fc3587cded

Observation 44a9300e-5520-4cee-ae65-51ce664fb9c2 · outbound

This paper cites Ms-celeb-1m: A dataset and benchmark for large-scale face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Ms-celeb-1m: A dataset and benchmark for large-scale face recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.495874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.495874Z digest=sha256:089a9ed801643d61baaa744ad90a0c123efdf1c0b411c659b73b26e0bc45d4d8

Observation f71a18a8-5f8a-4d10-ae47-c58f4fb9ad44 · outbound

This paper cites In Defense of the Triplet Loss for Person Re-Identification.

Illuminating Visual Identity in Universal Multimodal Embeddings In Defense of the Triplet Loss for Person Re-Identification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.565316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.565316Z digest=sha256:1905e89cc803c39b16fd898dc43e3f636f8f26502f60c2fafd680843731faed8

Observation 07e17851-bee6-46b0-aa54-affa8b247483 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

Illuminating Visual Identity in Universal Multimodal Embeddings Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.678112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.678112Z digest=sha256:8e5a17bd1511a8d86b630096f43079d42cc3a6c7de020b8e1624ac39bd8af491

Observation 31614ed6-7a1a-41c3-bc9d-bc694a947a15 · outbound

This paper cites Llm2clip: Powerful language model unlocks richer visual representation.arXiv preprint arXiv:2411.04997, 2024.

Illuminating Visual Identity in Universal Multimodal Embeddings Llm2clip: Powerful language model unlocks richer visual representation.arXiv preprint arXiv:2411.04997, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.751502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.751502Z digest=sha256:571ac676588471e28cbc33d458936d1293c0bb620c27c4187da157ea032c07bf

Observation 7e5169f5-85b8-4aad-a134-ca2008b0f153 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Illuminating Visual Identity in Universal Multimodal Embeddings Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.861413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.861413Z digest=sha256:f143fbb1ded97a0ad1abdf56e377fce40ed166ea3377a29b9210c65bf2be55b7

Observation aec1c32d-63af-40d8-a722-e6794830ce73 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Illuminating Visual Identity in Universal Multimodal Embeddings E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.906861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.906861Z digest=sha256:1ef826abd29c81b6ecc7e80c7f9842c4ff6666b0d2f3281b30cb445291f0bef9

Observation 8c618cf0-6d98-493e-af7c-dc803f18a35b · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

Illuminating Visual Identity in Universal Multimodal Embeddings Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.560698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:16.006633Z digest=sha256:43fd4d1e317bf88d7f8150f85738750347da8cec5d70629aec37e4cd6c800c68

Observation 68e4c1a0-4d0d-436a-b85c-8bffd421c7ac · outbound

This paper cites Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.106644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.106644Z digest=sha256:f2010e101ac8a674fcb07ebb6b7776f879f3cd028c795809ee91ab1415ffdbfc

Observation 69f665ef-4a45-4d80-a72b-0487e08fa2e7 · outbound

This paper cites Ilias: Instance-level image retrieval at scale.

Illuminating Visual Identity in Universal Multimodal Embeddings Ilias: Instance-level image retrieval at scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.326487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:16.217922Z digest=sha256:bf5a36fdc4fd4e0b68b97de8b604efe6ef001505c2613276c4aa186dd97091ec

Observation 73ae2735-f4c0-4927-ba64-3ea619c019e7 · outbound

This paper cites jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.254863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.254863Z digest=sha256:0570c810a2d12a18959875ad593dda5e421c44f8f467fe0b57b680360a513c1a

Observation a65a327a-6dda-4007-b95c-ae815c88f5d0 · outbound

This paper cites 3d object representations for fine-grained categorization.

Illuminating Visual Identity in Universal Multimodal Embeddings 3d object representations for fine-grained categorization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.112459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:16.403310Z digest=sha256:ae145d64778437cf7adb6f9a8bb17478733db0b68ffbc363a17aa0fa1078e60f

Observation 4d378f39-698e-463e-b460-d32df3b19ecd · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.514852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.514852Z digest=sha256:f6dfd3bc88a648eee0c324f4122bccd46dd2d07c5f8686ae1b08812bf265e2dd

Observation 8fdaf9ed-6939-417d-b030-3aef457a9870 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Illuminating Visual Identity in Universal Multimodal Embeddings NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.588928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.588928Z digest=sha256:038965e42f7d146c3b4fb567a75428122ea18b6508326fd909334ac823d6ee70

Observation c5aa0e80-d869-447d-9cf1-3cc72948598d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Illuminating Visual Identity in Universal Multimodal Embeddings LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.658249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.658249Z digest=sha256:be4b5a9849d9a985b3b0c5e0fe4e5af7dffaed08b44d55700f2da906c4bfb1e4

Observation b1c7a46d-5e3a-4422-ba4e-2b86744e564a · outbound

This paper cites MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization.

Illuminating Visual Identity in Universal Multimodal Embeddings MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-04T20:43:17.744970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:16.771390Z digest=sha256:82c91b94857081dd2f0db1c684c41be069e1b515332d8b2de8979255c4c289e7

Observation 6dc4ce12-4e7a-495b-8c25-13ae83d0a9db · outbound

This paper cites Personalvideo: High id-fidelity video customization without dynamic and semantic degradation.

Illuminating Visual Identity in Universal Multimodal Embeddings Personalvideo: High id-fidelity video customization without dynamic and semantic degradation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.903587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:16.813054Z digest=sha256:a0317228873aec76e49d62a46754e03d3491eba0693381bebec8ea7b0bb91dbf

Observation 77dbe72f-dc1e-4ac4-ae64-923d92c34748 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Illuminating Visual Identity in Universal Multimodal Embeddings Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.848115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.848115Z digest=sha256:96b99db4c2b63829e5cf1b450c21b73539d96ebc920036d58c1b745c0a8842de

Observation 20c6ace6-e135-4bb5-a263-204f179e002a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Illuminating Visual Identity in Universal Multimodal Embeddings Open-Sora Plan: Open-Source Large Video Generation Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.957129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.957129Z digest=sha256:61e865733ba11212c175c07faedc3ca3090273f204dffa21a4acc36295c17c15

Observation f6134962-d0e8-4c11-8465-dabdb4092ea0 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Illuminating Visual Identity in Universal Multimodal Embeddings UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.032612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.032612Z digest=sha256:08c5a597d234156d99528d1320b6dfbc9077b88abaeb9947db544f59731efe01

Observation 55759e70-7d51-48d1-ae24-01e839854cf1 · outbound

This paper cites Mm-embed: Universal multimodal retrieval with multimodal llms.

Illuminating Visual Identity in Universal Multimodal Embeddings Mm-embed: Universal multimodal retrieval with multimodal llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.674734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.102764Z digest=sha256:898935ca821ef0584ff84c9ee45daed41f8703146cc4e706d21079e47c3484df

Observation 97524d2e-81b4-4eaf-b53f-630d677d6522 · outbound

This paper cites IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.179548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.179548Z digest=sha256:30e0ed0c6d173232e1df4842fb65319321d9513a69c06f3f4b4a49001d0744ae

Observation 2c0a19dd-6584-43e3-8e78-f2306bde46d5 · outbound

This paper cites Automatic synthetic data and fine- grained adaptive feature alignment for composed person re- trieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Automatic synthetic data and fine- grained adaptive feature alignment for composed person re- trieval

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.464572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.257562Z digest=sha256:88eda5226e511484a6b836c2abc7ae728bd8596db24e4ac34b68f43df6a5aa4f

Observation 9f181751-f6ec-46ee-8eeb-b9440fc26489 · outbound

This paper cites Sphereface: Deep hypersphere embedding for face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Sphereface: Deep hypersphere embedding for face recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.260553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.260553Z digest=sha256:327673f810730ff950e5d393dbbd381e86e78b32611e2d03afc9e8b55aa0de78

Observation 775382e5-b8cc-4a14-a21b-127769462100 · outbound

This paper cites Large- scale vehicle re-identification in urban surveillance videos.

Illuminating Visual Identity in Universal Multimodal Embeddings Large- scale vehicle re-identification in urban surveillance videos

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.260878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.264145Z digest=sha256:8625bb88fa535e258158a7755c3c243441265dbfb2fca4343b8449d4b8c43aaa

Observation bcfdc93d-0712-4382-8ac6-8d4129837f68 · outbound

This paper cites Lamra: Large multimodal model as your advanced retrieval assistant.

Illuminating Visual Identity in Universal Multimodal Embeddings Lamra: Large multimodal model as your advanced retrieval assistant

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.013085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.267112Z digest=sha256:11d47a702c1e8f8eb5505c2fcc1f2b688d9dab30161a021401b266e762e50ed8

Observation 45cf1584-619d-4bbf-bf45-0b149e00568f · outbound

This paper cites Deepfashion: Powering robust clothes recognition and retrieval with rich annotations.

Illuminating Visual Identity in Universal Multimodal Embeddings Deepfashion: Powering robust clothes recognition and retrieval with rich annotations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.814995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.270047Z digest=sha256:89cf886452a088ceafd6d459f4ca03fba387f7258b635e8e793f94c6798546a6

Observation 8091a9df-dbd6-42d8-81e4-7712ab156d92 · outbound

This paper cites VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents.

Illuminating Visual Identity in Universal Multimodal Embeddings VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.273022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.273022Z digest=sha256:135bf32f38103964e487f0aaa60cf691eb6adc3a67b44362d56eccbf0c79e723

Observation d6eef85d-e283-41e8-8d54-c2db472faef7 · outbound

This paper cites A culturally-aware benchmark for person re-identification in modest attire.Engineering Ap- plications of Artificial Intelligence, 158:111494, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings A culturally-aware benchmark for person re-identification in modest attire.Engineering Ap- plications of Artificial Intelligence, 158:111494, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.638218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.276038Z digest=sha256:9c0f21804f5d6811d325c573b151b63013a8fa5344357daeddd07299e7f2771c

Observation 421feeaf-c31b-42ed-a906-f54525ffb7e8 · outbound

This paper cites Deep metric learning via lifted structured fea- ture embedding.

Illuminating Visual Identity in Universal Multimodal Embeddings Deep metric learning via lifted structured fea- ture embedding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.392243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.281806Z digest=sha256:323f735cd159e80ec9eaa7beceb09f53522d483716bf75237f9f6cab3703851f

Observation 5898f116-c7ce-4c52-b079-08ad080bd6a7 · outbound

This paper cites RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification.

Illuminating Visual Identity in Universal Multimodal Embeddings RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.284647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.284647Z digest=sha256:d3ac1ad09e9658d794f82524dd8c80eee58bdaca3a06293cad18f174ada9247c

Observation 4eaf0447-583f-46d7-8672-60cb21b66218 · outbound

This paper cites Revisiting oxford and paris: Large-scale image retrieval benchmarking.

Illuminating Visual Identity in Universal Multimodal Embeddings Revisiting oxford and paris: Large-scale image retrieval benchmarking

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.266199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.288092Z digest=sha256:c53ae18ca3ac18f396c13df3d8bf54c746cb3c6214e79a4ff2ebe74b76e441c9

Observation 624f29d7-4559-42a1-9e5c-1285ade0939e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Illuminating Visual Identity in Universal Multimodal Embeddings Learning transferable visual models from natural language supervi- sion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.290997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.290997Z digest=sha256:34f0f73207f53b1fa1ad83125eaf5c805b04215f5d4f4681b7b919d07295712d

Observation e24db9c0-0477-404c-b96c-4fa04bdd4ca2 · outbound

This paper cites Performance measures and a data set for multi-target, multi-camera tracking.

Illuminating Visual Identity in Universal Multimodal Embeddings Performance measures and a data set for multi-target, multi-camera tracking

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.078583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.293921Z digest=sha256:818794ae5a2b9d4f9be5b311aa8f7ad171c393cf12f914aee96f42205acf90bb

Observation 3aa65bd2-642e-44e9-9f56-f9c3284cf5c8 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Illuminating Visual Identity in Universal Multimodal Embeddings EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.297183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.297183Z digest=sha256:3d3920646e1adc9401dd0971aeda2c3b9659cdfb011e6ad8d6124673d6595c90

Observation 51c43246-bcb5-42d8-85e2-138c2169789b · outbound

This paper cites Visual named entity linking: A new dataset and a baseline.

Illuminating Visual Identity in Universal Multimodal Embeddings Visual named entity linking: A new dataset and a baseline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.950019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.300330Z digest=sha256:044c88f2f23515a8421dd8e9fe5a9242c8a2f68a366cc98e01858e0e5820e093

Observation 9297161b-4cfe-43c4-8db0-90ab72749cd6 · outbound

This paper cites Breaking the batch barrier (b3) of contrastive learning via smart batch mining.arXiv preprint arXiv:2505.11293, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Breaking the batch barrier (b3) of contrastive learning via smart batch mining.arXiv preprint arXiv:2505.11293, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.303217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.303217Z digest=sha256:55cdb99f67d2690950f69d3dfb2c3ef8d3ca51783c87de28088e7e47252df33e

Observation 7b175f18-ce65-49f6-b80d-f97439b0ad29 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Illuminating Visual Identity in Universal Multimodal Embeddings SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.306059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.306059Z digest=sha256:89e6048f6e90709c45cae8cb97fb3f0e726ef7a0bf1d2077619a6e95685be131

Observation 610fae48-b01c-4259-8dc5-e769b16bfb37 · outbound

This paper cites The inaturalist species classification and de- tection dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings The inaturalist species classification and de- tection dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.309133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.309133Z digest=sha256:4dd2934184a26df151639b15433c3266bb8057433be4d111453ce777dc38f69b

Observation f4055274-555e-4706-a121-2c528e67943e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.312239Z digest=sha256:eeb78d5ae134c40c437ea05940bb7da217bf4633d50aba539e530648ccdc633f

Observation 60351181-dd15-4ad5-9198-04f7cc56be39 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Illuminating Visual Identity in Universal Multimodal Embeddings InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.315397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.315397Z digest=sha256:92ac38d351b5bc89d30b5b2e5916ee0f7bcd914e0543dd1f0481bf1d785cd695

Observation e875a0a6-b8b5-442d-8947-11f2f9821327 · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.318655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.318655Z digest=sha256:55373ecff8542701c76c16490ff086fae3c72a30549fe06be66c2df524806bc0

Observation 932e528f-e291-479a-afd3-bda7ba484c26 · outbound

This paper cites Uniir: Train- ing and benchmarking universal multimodal information re- trievers.

Illuminating Visual Identity in Universal Multimodal Embeddings Uniir: Train- ing and benchmarking universal multimodal information re- trievers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.787105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.322048Z digest=sha256:c0bc724b9c30e7a9795208b165c987c4a507990bbcb3fa2633aeb50c8a77e89b

Observation 6755346c-bcb1-4235-8fe7-e6c08722c841 · outbound

This paper cites Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.613019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.324853Z digest=sha256:d3b738bdb089f8bb711d2aa81ffec14e2463ebc592fd7d3bafa4791c61107e13

Observation 99e89c04-a22a-4938-bea5-2df5c91abf9e · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

Illuminating Visual Identity in Universal Multimodal Embeddings Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.467311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.327787Z digest=sha256:187fe793796b3fd8c5cd604713ec15927504d483dac260c5083422548867a87c

Observation d03ce222-b0c5-458c-aa79-c6143ef507e5 · outbound

This paper cites Forb: a flat object retrieval benchmark for universal im- age embedding.Advances in Neural Information Processing Systems, 36:25448–25460, 2023.

Illuminating Visual Identity in Universal Multimodal Embeddings Forb: a flat object retrieval benchmark for universal im- age embedding.Advances in Neural Information Processing Systems, 36:25448–25460, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.366074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.330652Z digest=sha256:4c1ad0c92088fa4dbccc2840772a800d8ed0de189806a911001ba6504874e676

Observation 46449807-f27c-469a-8b8c-f762153cc6e0 · outbound

This paper cites Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.333489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.333489Z digest=sha256:5eb327fa2834373e46e3af99f6128ec04796873e8c90d55713e9d66bacdf3ede

Observation d2e52caf-a117-4d20-864f-ee1847e6c1db · outbound

This paper cites Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying.

Illuminating Visual Identity in Universal Multimodal Embeddings Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.336608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.336608Z digest=sha256:8a48724ab16bdaf8501518acdbb10aec2fa1973ebc654b463243a546667989d4

Observation 87e9efb2-e9f0-4622-8961-558d1fdb1567 · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Illuminating Visual Identity in Universal Multimodal Embeddings Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.339889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.339889Z digest=sha256:77ea3abb62ecdffeba77ab9ef32190cb9d2627e94b557ba73f9d132a4291bb62

Observation 7af4341d-dbac-4eda-b7b4-20a07289fc47 · outbound

This paper cites A large-scale car dataset for fine-grained categorization and verification.

Illuminating Visual Identity in Universal Multimodal Embeddings A large-scale car dataset for fine-grained categorization and verification

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.229304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.342861Z digest=sha256:f16ed9bafe0744184fb4243cbb00862523ccde50720520315ae711e1800a7584

Observation e90943c0-354b-48fb-82e6-fcba2ef6a499 · outbound

This paper cites Deep learning for person re- identification: A survey and outlook.IEEE transactions on pattern analysis and machine intelligence, 44(6):2872–2893,.

Illuminating Visual Identity in Universal Multimodal Embeddings Deep learning for person re- identification: A survey and outlook.IEEE transactions on pattern analysis and machine intelligence, 44(6):2872–2893,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.085925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.345722Z digest=sha256:80b5d71eeb695616d0520a1c14689cd6a7b23353b2799df4bd2bcb38178954b7

Observation 7cd99d99-95d9-4e89-a6ef-3f036daaf510 · outbound

This paper cites Learning Face Representation from Scratch.

Illuminating Visual Identity in Universal Multimodal Embeddings Learning Face Representation from Scratch

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.348652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.348652Z digest=sha256:911a174b30cab632b1368014a3ee0e7c7817d9e1ae413f829a4eda92e358d0fb

Observation ac953921-ea05-4cee-a1ac-eed9274e9d03 · outbound

This paper cites The met dataset: Instance-level recognition for artworks.

Illuminating Visual Identity in Universal Multimodal Embeddings The met dataset: Instance-level recognition for artworks

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.941829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.351877Z digest=sha256:7967f3e316c8f813f0efeaa94d49422565ec8dbdcb75d3cbcaf5ff48362000f2

Observation 2bd6fb3a-74be-4cc9-a278-c3e40dd7f15b · outbound

This paper cites Towards universal image embed- dings: A large-scale dataset and challenge for generic image representations.

Illuminating Visual Identity in Universal Multimodal Embeddings Towards universal image embed- dings: A large-scale dataset and challenge for generic image representations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.824469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.354680Z digest=sha256:ea5d614dbb5cf20c73807d6e837bbdddfcdd6e426cdeafe5010991f4bbd9d8ce

Observation 9cbd0a3e-e878-45fd-857a-ffbc63cac070 · outbound

This paper cites Udon: Universal dynamic online distilla- tion for generic image representations.Advances in Neural Information Processing Systems, 37:86836–86859, 2024.

Illuminating Visual Identity in Universal Multimodal Embeddings Udon: Universal dynamic online distilla- tion for generic image representations.Advances in Neural Information Processing Systems, 37:86836–86859, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.694822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.357587Z digest=sha256:cfe376be47c96531025008a872ee848635941344331d737fee7242bdafaf5468

Observation 931fdf07-6b8d-4b55-98f8-1f24ea3029e9 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Illuminating Visual Identity in Universal Multimodal Embeddings CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.360998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.360998Z digest=sha256:355511fce2063d3317c9fa2ecbc92bc1243bed1a9f866cb50cec4e718423732a

Observation 8178b8ee-26cc-4913-9162-c5f2d8ad71e5 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

Illuminating Visual Identity in Universal Multimodal Embeddings VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.364610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.364610Z digest=sha256:90fccc2c1c11009bbb4f2b9bebe66dd93d679db82f14271f824b3286716aaef1

Observation 6c629151-d8cf-4adc-8784-379eba4fed65 · outbound

This paper cites Identity- preserving text-to-video generation by frequency decompo- sition.

Illuminating Visual Identity in Universal Multimodal Embeddings Identity- preserving text-to-video generation by frequency decompo- sition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.558554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.367779Z digest=sha256:9b145cdd57e9573b597354845e17da315f4f4826c2f0a978ba4593b206293a4e

Observation 40690f5a-428d-4b23-8919-477e13ee1eb5 · outbound

This paper cites Sigmoid loss for language image pre-training.

Illuminating Visual Identity in Universal Multimodal Embeddings Sigmoid loss for language image pre-training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.370533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.370533Z digest=sha256:5d3056074211a8a4f32b154cc8f8982930d44e3293444dbdfff91781f196f7c3

Observation b18f833e-7fe4-4b36-9661-ab9b29429545 · outbound

This paper cites Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining.

Illuminating Visual Identity in Universal Multimodal Embeddings Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.413590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.373366Z digest=sha256:40c3d6b08243e5031ea8db027b3e93f5758ad9bdad99089f6bc112ca1f025748

Observation 5fd02a67-b5ed-461a-99b7-0bc1bf3d7fdd · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

Illuminating Visual Identity in Universal Multimodal Embeddings MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.376305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.376305Z digest=sha256:675f37667fa7265a96e227928a5b201d510ee853f94d1993022cfb0cce4e9795

Observation 9fd7f93b-194a-406e-b802-404beba054a2 · outbound

This paper cites Assess- ing and learning alignment of unimodal vision and language models.

Illuminating Visual Identity in Universal Multimodal Embeddings Assess- ing and learning alignment of unimodal vision and language models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.379454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.379454Z digest=sha256:c1cc50d608bd38ef6ed1b46a6fe88f5f307c8f12c2734fdf438ec348438fa284

Observation 1dde37f5-d3b3-4440-9a43-fc552cb4ac59 · outbound

This paper cites Beyond frontal faces: Improving per- son recognition using multiple cues.

Illuminating Visual Identity in Universal Multimodal Embeddings Beyond frontal faces: Improving per- son recognition using multiple cues

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.275526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.382385Z digest=sha256:32ff30a4c0296df1563548ae6324b02204e5880d2b4a2b9ec810a45f7ad23162

Observation a68bac3d-d1c0-410e-b9bb-f932ca113872 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

Illuminating Visual Identity in Universal Multimodal Embeddings GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.485401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.485401Z digest=sha256:9ec9cd281e54894143ff4332339590e983da8b0d48049314d093fbc4da78f606

Observation 1ea36a60-e6aa-4864-a5f7-160629ce6583 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.489048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.489048Z digest=sha256:12ae472450371da318c4db50fbbd70ec4dafae9a4f3e656371fbfe966b26dc01

Observation e18f69cb-8fdc-4673-917f-1354f340451b · outbound

This paper cites Magicmirror: Id-preserved video generation in video diffusion transformers.

Illuminating Visual Identity in Universal Multimodal Embeddings Magicmirror: Id-preserved video generation in video diffusion transformers

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.119414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.492298Z digest=sha256:7425b9131137f9feaca9ba8d7f0196b844d4a2216b7a9adeeb19785450e9afef

Observation 0d85fe8b-7bc2-406a-87ae-95816f61eebe · outbound

This paper cites Scalable person re-identification: A benchmark.

Illuminating Visual Identity in Universal Multimodal Embeddings Scalable person re-identification: A benchmark

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.035923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.495323Z digest=sha256:9409bcd2e386843225563c4f82a8e09a7618ed20a93f14d285f4faac625f2698

Observation da79ef23-b212-4b4f-999e-de3b37100501 · outbound

This paper cites Scalable person re-identification: A benchmark.

Illuminating Visual Identity in Universal Multimodal Embeddings Scalable person re-identification: A benchmark

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.010772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.498322Z digest=sha256:f5da2eb796bc8d577a8f31dc4697db9ae2db6545613f20d90381d840b6d3d9cf

Observation 10ef6a1b-c816-419f-b476-c14682660d44 · outbound

This paper cites Cartoon face recognition: A bench- mark dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings Cartoon face recognition: A bench- mark dataset

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:17.983167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.501286Z digest=sha256:5cc56ae182d3f7d4fd107c0cfa4daf76dfe0f2e204e5e4073a1d4727307aa2ee

Observation 47e166c9-8827-4454-9a0a-97499ad8221c · outbound

This paper cites Megapairs: Massive data synthesis for universal multi- modal retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Megapairs: Massive data synthesis for universal multi- modal retrieval

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:17.963805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T20:43:17.504442Z digest=sha256:8d797065790309cc1f57e9e718a0e30fac86e6b0aafd3a2cbfb55487f20642d4

Observation 7f78a490-0d0c-41d6-bd8c-2321a9113e37 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Illuminating Visual Identity in Universal Multimodal Embeddings InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 89

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:43:17.507439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.507439Z digest=sha256:449a7f92a08b35945a7488a794b1f40634dd16954832a15d2c5a83cc39417730

Pith citing papers

No inbound Pith citation observations are available.