Pith. sign in

Paper Citation Record · LEDGER

Illuminating Visual Identity in Universal Multimodal Embeddings

As of 9 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.01794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01794 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:43:17.507439Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved55
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2363b7d3-81bd-4afc-a3da-8fe99693d3db · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Illuminating Visual Identity in Universal Multimodal Embeddings Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.346553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.346553Z digest=sha256:da6b5c1610d00dc8f22611402d9e83cc0b6a375be08eccb9742b8c72d2b389b1

Observation bd4f6e0e-b313-4b2c-89cf-b70b0de3e3dc · outbound

This paper cites Unicom: Universal and Compact Representation Learning for Image Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Unicom: Universal and Compact Representation Learning for Image Retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.403233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.403233Z digest=sha256:093a99a0cf3f79ba2ee93de8622e6ea86a418aff57acbc5de75bdb3423f7d832

Observation 88636538-b85b-48d0-8d11-55c14ce56b5b · outbound

This paper cites Qwen2.5-VL Technical Report.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.493030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.493030Z digest=sha256:98b216c12b65c9401e3af19cafd0f211e1ab27eb1a4929c1d6a48ed506610bd0

Observation 6dd964af-ecc0-4d72-b6a0-0b501e81ebdb · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Illuminating Visual Identity in Universal Multimodal Embeddings LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.611224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.611224Z digest=sha256:728c97e742c56a4023d5cede63824eeae96c7f49d89a815549e28f9eb77ab635

Observation a64de592-e3dc-479b-8d0b-49625f0cf298 · outbound

This paper cites Flame: Frozen large language models enable data-efficient language-image pre- training.

Illuminating Visual Identity in Universal Multimodal Embeddings Flame: Frozen large language models enable data-efficient language-image pre- training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.727986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.727986Z digest=sha256:6ce72fb8ccfa44654b21781b7b94f7e224098e7ee57e55218e2f02ec89d89e97

Observation 9e7d481e-e174-4d49-a57f-18746709650d · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

Illuminating Visual Identity in Universal Multimodal Embeddings mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.812234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.812234Z digest=sha256:60dd3565ac5b8aae0f1b55bb64ef1260f719a677d0c4ce30e53384be779769de

Observation ce543fff-5dde-46e1-bc81-2031f0eb2874 · outbound

This paper cites Murag: Multimodal retrieval-augmented generator for open question answering over images and text.

Illuminating Visual Identity in Universal Multimodal Embeddings Murag: Multimodal retrieval-augmented generator for open question answering over images and text

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:13.932581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:13.932581Z digest=sha256:ff626aee0247b50317d83e3ec9bc9c501bd34f5ca35873576f94a0bb2a3fc8bf

Observation c0ab3143-6da5-4689-99ea-34a6d71ff4a8 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Illuminating Visual Identity in Universal Multimodal Embeddings Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.058096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.058096Z digest=sha256:988096a6a713f31624d43c259f7694adef5c498165b60f7e845583b3eef6c20e

Observation e2d25ea6-cd73-4d60-9e7a-3ef01090b387 · outbound

This paper cites Opengpt-4o-image: A com- prehensive dataset for advanced image generation and edit- ing.arXiv preprint arXiv:2509.24900, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Opengpt-4o-image: A com- prehensive dataset for advanced image generation and edit- ing.arXiv preprint arXiv:2509.24900, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.225875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.225875Z digest=sha256:cc712f0b07c1b1453a6a114352432977b7d7e2132d5d6ee628ba6065879d24ea

Observation ced016ae-bb8a-4efa-8861-ef6d4e4a0e19 · outbound

This paper cites Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014,.

Illuminating Visual Identity in Universal Multimodal Embeddings Think then embed: Generative context improves multimodal embedding.arXiv preprint arXiv:2510.05014,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.394360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.394360Z digest=sha256:99b829d72fc2e82d4bd451b9238564f85849f6826311c6f48238108067b1a5c6

Observation 1503431a-f199-4920-89fd-2aa1bdbb0418 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Arcface: Additive angular margin loss for deep face recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.494829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.494829Z digest=sha256:6a192184d74f7d63acf13e548b137a7489d3c5ee24291cf56f857af2175085c3

Observation 84c5eb15-19db-4409-9a10-8ec3f53b88d7 · outbound

This paper cites Efficient and discriminative image feature extrac- tion for universal image retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Efficient and discriminative image feature extrac- tion for universal image retrieval

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.665447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.665447Z digest=sha256:8924f3541dca87f73c47868e690059f891e4f58575fe7157066e9f067a359e9a

Observation 6675d07d-4f80-441a-9de8-552a76f22541 · outbound

This paper cites DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.

Illuminating Visual Identity in Universal Multimodal Embeddings DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.825424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.825424Z digest=sha256:060131b73e0358904d5bcd9181f19b6b9a4ee5032a424222e5fe8250e11634cb

Observation 3e1ac40a-362b-4648-b680-886b4dcee7b4 · outbound

This paper cites Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup.

Illuminating Visual Identity in Universal Multimodal Embeddings Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:14.991222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:14.991222Z digest=sha256:344c0411d17bad7f448479ab12e951a2b39d297240723fe0641fba6b71dc854b

Observation 036f4965-2c3d-44eb-95c6-f35eb5cbe0a1 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Illuminating Visual Identity in Universal Multimodal Embeddings SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.155373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.155373Z digest=sha256:ba9d644cb1dce26e497ea3f806485725cdfc290e8ce9b03e092a78c6226c9c1f

Observation a2d55248-5423-4de2-bd52-c4b22dd1ce2b · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Breaking the modality barrier: Universal embedding learning with multimodal llms.arXiv preprint arXiv:2504.17432, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.239595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.239595Z digest=sha256:28fddac6b36ac6bff707e3c808193d1353fccb0f228e3422bd61ce992187eb06

Observation 809043bb-8bd1-4a11-8630-d35cce3060a4 · outbound

This paper cites Unime-v2: Mllm-as-a-judge for uni- versal multimodal embedding learning.arXiv preprint arXiv:2510.13515, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Unime-v2: Mllm-as-a-judge for uni- versal multimodal embedding learning.arXiv preprint arXiv:2510.13515, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.281060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.281060Z digest=sha256:aabb7cec3c08f0db174dc38e56cb7ab7d8883901b365ddcd3e872a496df88cb4

Observation c6a7724b-16fd-42e0-8e92-c11485f0cb06 · outbound

This paper cites jina-embeddings-v4: Universal embeddings for multimodal multilingual retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-embeddings-v4: Universal embeddings for multimodal multilingual retrieval

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.356047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.356047Z digest=sha256:125cf18a28df7801df5f71371386e870564fe0077ec13a7540b662fc3587cded

Observation 44a9300e-5520-4cee-ae65-51ce664fb9c2 · outbound

This paper cites Ms-celeb-1m: A dataset and benchmark for large-scale face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Ms-celeb-1m: A dataset and benchmark for large-scale face recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.495874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.495874Z digest=sha256:089a9ed801643d61baaa744ad90a0c123efdf1c0b411c659b73b26e0bc45d4d8

Observation f71a18a8-5f8a-4d10-ae47-c58f4fb9ad44 · outbound

This paper cites In Defense of the Triplet Loss for Person Re-Identification.

Illuminating Visual Identity in Universal Multimodal Embeddings In Defense of the Triplet Loss for Person Re-Identification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.565316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.565316Z digest=sha256:1905e89cc803c39b16fd898dc43e3f636f8f26502f60c2fafd680843731faed8

Observation 07e17851-bee6-46b0-aa54-affa8b247483 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

Illuminating Visual Identity in Universal Multimodal Embeddings Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.678112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.678112Z digest=sha256:8e5a17bd1511a8d86b630096f43079d42cc3a6c7de020b8e1624ac39bd8af491

Observation 31614ed6-7a1a-41c3-bc9d-bc694a947a15 · outbound

This paper cites Llm2clip: Powerful language model unlocks richer visual representation.arXiv preprint arXiv:2411.04997, 2024.

Illuminating Visual Identity in Universal Multimodal Embeddings Llm2clip: Powerful language model unlocks richer visual representation.arXiv preprint arXiv:2411.04997, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.751502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.751502Z digest=sha256:571ac676588471e28cbc33d458936d1293c0bb620c27c4187da157ea032c07bf

Observation 7e5169f5-85b8-4aad-a134-ca2008b0f153 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Illuminating Visual Identity in Universal Multimodal Embeddings Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.861413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.861413Z digest=sha256:f143fbb1ded97a0ad1abdf56e377fce40ed166ea3377a29b9210c65bf2be55b7

Observation aec1c32d-63af-40d8-a722-e6794830ce73 · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Illuminating Visual Identity in Universal Multimodal Embeddings E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:15.906861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:15.906861Z digest=sha256:1ef826abd29c81b6ecc7e80c7f9842c4ff6666b0d2f3281b30cb445291f0bef9

Observation 8c618cf0-6d98-493e-af7c-dc803f18a35b · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

Illuminating Visual Identity in Universal Multimodal Embeddings Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.560698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:16.006633Z digest=sha256:7f1382acae7e87aa17640041e39fa4aa00ded5944836ba7021f56ea48f09512e

Observation 68e4c1a0-4d0d-436a-b85c-8bffd421c7ac · outbound

This paper cites Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.106644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.106644Z digest=sha256:6db115af74fbe718b7f4c2aac7dfbb6cbd03f61d18b6b3c10bca4a59da55761c

Observation 69f665ef-4a45-4d80-a72b-0487e08fa2e7 · outbound

This paper cites Ilias: Instance-level image retrieval at scale.

Illuminating Visual Identity in Universal Multimodal Embeddings Ilias: Instance-level image retrieval at scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.326487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:16.217922Z digest=sha256:4d8c1c62db6e627e409d362ec2b00bf09831155f8fea4f281a9ae1731d885f72

Observation 73ae2735-f4c0-4927-ba64-3ea619c019e7 · outbound

This paper cites jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images.

Illuminating Visual Identity in Universal Multimodal Embeddings jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.254863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.254863Z digest=sha256:327b9325701a2c25374537ee8d0b0bd90eebbeb351a45ca31dcf22c95c86c407

Observation a65a327a-6dda-4007-b95c-ae815c88f5d0 · outbound

This paper cites 3d object representations for fine-grained categorization.

Illuminating Visual Identity in Universal Multimodal Embeddings 3d object representations for fine-grained categorization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:22.112459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:16.403310Z digest=sha256:63f78647a5cd4da2de6fdbbe90cd8a671ac73b9ee15713385cf1c60d5953aa8b

Observation 4d378f39-698e-463e-b460-d32df3b19ecd · outbound

This paper cites Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Llave: Large language and vision embedding models with hardness-weighted contrastive learning.arXiv preprint arXiv:2503.04812, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.514852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.514852Z digest=sha256:f6dfd3bc88a648eee0c324f4122bccd46dd2d07c5f8686ae1b08812bf265e2dd

Observation 8fdaf9ed-6939-417d-b030-3aef457a9870 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

Illuminating Visual Identity in Universal Multimodal Embeddings NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.588928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.588928Z digest=sha256:038965e42f7d146c3b4fb567a75428122ea18b6508326fd909334ac823d6ee70

Observation c5aa0e80-d869-447d-9cf1-3cc72948598d · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Illuminating Visual Identity in Universal Multimodal Embeddings LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.658249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.658249Z digest=sha256:be4b5a9849d9a985b3b0c5e0fe4e5af7dffaed08b44d55700f2da906c4bfb1e4

Observation b1c7a46d-5e3a-4422-ba4e-2b86744e564a · outbound

This paper cites MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization.

Illuminating Visual Identity in Universal Multimodal Embeddings MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-04T20:43:17.744970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:16.771390Z digest=sha256:cd957ef5b48e88c69879342c28995ac8dc6113d27f0fd3d081a598a1992b6755

Observation 6dc4ce12-4e7a-495b-8c25-13ae83d0a9db · outbound

This paper cites Personalvideo: High id-fidelity video customization without dynamic and semantic degradation.

Illuminating Visual Identity in Universal Multimodal Embeddings Personalvideo: High id-fidelity video customization without dynamic and semantic degradation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.903587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:16.813054Z digest=sha256:a783840d0b5e765b3b87692cc18d472f3372584951fb96b2a38999a84ef7835e

Observation 77dbe72f-dc1e-4ac4-ae64-923d92c34748 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Illuminating Visual Identity in Universal Multimodal Embeddings Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.848115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.848115Z digest=sha256:96b99db4c2b63829e5cf1b450c21b73539d96ebc920036d58c1b745c0a8842de

Observation 20c6ace6-e135-4bb5-a263-204f179e002a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Illuminating Visual Identity in Universal Multimodal Embeddings Open-Sora Plan: Open-Source Large Video Generation Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:16.957129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:16.957129Z digest=sha256:e388147ed1cc4cd6640a29b85ad1a0914aaecb523e90226c5c93aca2f21cdad9

Observation f6134962-d0e8-4c11-8465-dabdb4092ea0 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Illuminating Visual Identity in Universal Multimodal Embeddings UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.032612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.032612Z digest=sha256:f85806c4528314c4a6e3e2e1e3dc6ae0987468f092e69151dadab2384a972653

Observation 55759e70-7d51-48d1-ae24-01e839854cf1 · outbound

This paper cites Mm-embed: Universal multimodal retrieval with multimodal llms.

Illuminating Visual Identity in Universal Multimodal Embeddings Mm-embed: Universal multimodal retrieval with multimodal llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.674734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.102764Z digest=sha256:2c4db9c317cd2858c4b54849e95ba794d4b1d10d351d6775f333a10117a6e3c6

Observation 97524d2e-81b4-4eaf-b53f-630d677d6522 · outbound

This paper cites IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.179548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.179548Z digest=sha256:30e0ed0c6d173232e1df4842fb65319321d9513a69c06f3f4b4a49001d0744ae

Observation 2c0a19dd-6584-43e3-8e78-f2306bde46d5 · outbound

This paper cites Automatic synthetic data and fine- grained adaptive feature alignment for composed person re- trieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Automatic synthetic data and fine- grained adaptive feature alignment for composed person re- trieval

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.464572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.257562Z digest=sha256:ea1ef305f371ae92663bf8bc10c7dd0f56898ce35ba6c5e87dcb2e7c39a992f3

Observation 9f181751-f6ec-46ee-8eeb-b9440fc26489 · outbound

This paper cites Sphereface: Deep hypersphere embedding for face recognition.

Illuminating Visual Identity in Universal Multimodal Embeddings Sphereface: Deep hypersphere embedding for face recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.260553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.260553Z digest=sha256:327673f810730ff950e5d393dbbd381e86e78b32611e2d03afc9e8b55aa0de78

Observation 775382e5-b8cc-4a14-a21b-127769462100 · outbound

This paper cites Large- scale vehicle re-identification in urban surveillance videos.

Illuminating Visual Identity in Universal Multimodal Embeddings Large- scale vehicle re-identification in urban surveillance videos

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.260878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.264145Z digest=sha256:bceaa188c65b57ae3331f7e011469a996a02216e1f74b7ed6acb8f454dc75532

Observation bcfdc93d-0712-4382-8ac6-8d4129837f68 · outbound

This paper cites Lamra: Large multimodal model as your advanced retrieval assistant.

Illuminating Visual Identity in Universal Multimodal Embeddings Lamra: Large multimodal model as your advanced retrieval assistant

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:21.013085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.267112Z digest=sha256:45c786fbfe2fd463a7bbbed1650ccf6a99b6053b66853117a0d850f00785ed82

Observation 45cf1584-619d-4bbf-bf45-0b149e00568f · outbound

This paper cites Deepfashion: Powering robust clothes recognition and retrieval with rich annotations.

Illuminating Visual Identity in Universal Multimodal Embeddings Deepfashion: Powering robust clothes recognition and retrieval with rich annotations

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.814995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.270047Z digest=sha256:7bbf2792e10261f83978714b2a590c074a317b08eb9f99dc82b9eaaa41a3ab2c

Observation 8091a9df-dbd6-42d8-81e4-7712ab156d92 · outbound

This paper cites VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents.

Illuminating Visual Identity in Universal Multimodal Embeddings VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.273022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.273022Z digest=sha256:135bf32f38103964e487f0aaa60cf691eb6adc3a67b44362d56eccbf0c79e723

Observation d6eef85d-e283-41e8-8d54-c2db472faef7 · outbound

This paper cites A culturally-aware benchmark for person re-identification in modest attire.Engineering Ap- plications of Artificial Intelligence, 158:111494, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings A culturally-aware benchmark for person re-identification in modest attire.Engineering Ap- plications of Artificial Intelligence, 158:111494, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.638218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.276038Z digest=sha256:27f0869f65a1b23460727b8a5e2f0f0f2a6eb81de5b33d3821b6aace5fcb27fc

Observation 421feeaf-c31b-42ed-a906-f54525ffb7e8 · outbound

This paper cites Deep metric learning via lifted structured fea- ture embedding.

Illuminating Visual Identity in Universal Multimodal Embeddings Deep metric learning via lifted structured fea- ture embedding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.392243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.281806Z digest=sha256:4cc72224e1237ddbc43b386f0c77e9fddc3b5d3c7f2a4ef46e897d9b9330c5be

Observation 5898f116-c7ce-4c52-b079-08ad080bd6a7 · outbound

This paper cites RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification.

Illuminating Visual Identity in Universal Multimodal Embeddings RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.284647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.284647Z digest=sha256:9d3dd3b21b4d7b6c3c325d9ef8e6345fb4c8558f22f1c7cbce5399dc07994b9e

Observation 4eaf0447-583f-46d7-8672-60cb21b66218 · outbound

This paper cites Revisiting oxford and paris: Large-scale image retrieval benchmarking.

Illuminating Visual Identity in Universal Multimodal Embeddings Revisiting oxford and paris: Large-scale image retrieval benchmarking

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.266199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.288092Z digest=sha256:265dc370f624aa1d71bebafd8b71ab232e0db69cc733eb40b4aaab4c8d7e0e34

Observation 624f29d7-4559-42a1-9e5c-1285ade0939e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Illuminating Visual Identity in Universal Multimodal Embeddings Learning transferable visual models from natural language supervi- sion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.290997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.290997Z digest=sha256:34f0f73207f53b1fa1ad83125eaf5c805b04215f5d4f4681b7b919d07295712d

Observation e24db9c0-0477-404c-b96c-4fa04bdd4ca2 · outbound

This paper cites Performance measures and a data set for multi-target, multi-camera tracking.

Illuminating Visual Identity in Universal Multimodal Embeddings Performance measures and a data set for multi-target, multi-camera tracking

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:20.078583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.293921Z digest=sha256:01b39f99d7683321f86eebfa8fca05b50e109ccfab550979f04626db882e00dc

Observation 3aa65bd2-642e-44e9-9f56-f9c3284cf5c8 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Illuminating Visual Identity in Universal Multimodal Embeddings EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.297183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.297183Z digest=sha256:3d3920646e1adc9401dd0971aeda2c3b9659cdfb011e6ad8d6124673d6595c90

Observation 51c43246-bcb5-42d8-85e2-138c2169789b · outbound

This paper cites Visual named entity linking: A new dataset and a baseline.

Illuminating Visual Identity in Universal Multimodal Embeddings Visual named entity linking: A new dataset and a baseline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.950019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.300330Z digest=sha256:d7126ed1112b71a9d260318de7b47b5ac0669be6b79e542d51986b131ebfd77c

Observation 9297161b-4cfe-43c4-8db0-90ab72749cd6 · outbound

This paper cites Breaking the batch barrier (b3) of contrastive learning via smart batch mining.arXiv preprint arXiv:2505.11293, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Breaking the batch barrier (b3) of contrastive learning via smart batch mining.arXiv preprint arXiv:2505.11293, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.303217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.303217Z digest=sha256:55cdb99f67d2690950f69d3dfb2c3ef8d3ca51783c87de28088e7e47252df33e

Observation 7b175f18-ce65-49f6-b80d-f97439b0ad29 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Illuminating Visual Identity in Universal Multimodal Embeddings SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.306059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.306059Z digest=sha256:89e6048f6e90709c45cae8cb97fb3f0e726ef7a0bf1d2077619a6e95685be131

Observation 610fae48-b01c-4259-8dc5-e769b16bfb37 · outbound

This paper cites The inaturalist species classification and de- tection dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings The inaturalist species classification and de- tection dataset

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.309133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.309133Z digest=sha256:4dd2934184a26df151639b15433c3266bb8057433be4d111453ce777dc38f69b

Observation f4055274-555e-4706-a121-2c528e67943e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.312239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.312239Z digest=sha256:eeb78d5ae134c40c437ea05940bb7da217bf4633d50aba539e530648ccdc633f

Observation 60351181-dd15-4ad5-9198-04f7cc56be39 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Illuminating Visual Identity in Universal Multimodal Embeddings InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.315397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.315397Z digest=sha256:92ac38d351b5bc89d30b5b2e5916ee0f7bcd914e0543dd1f0481bf1d785cd695

Observation e875a0a6-b8b5-442d-8947-11f2f9821327 · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.318655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.318655Z digest=sha256:437567545efd34233e1ecab7563fb65526f6fd41f0f32a8ddbd2bcda7be30eef

Observation 932e528f-e291-479a-afd3-bda7ba484c26 · outbound

This paper cites Uniir: Train- ing and benchmarking universal multimodal information re- trievers.

Illuminating Visual Identity in Universal Multimodal Embeddings Uniir: Train- ing and benchmarking universal multimodal information re- trievers

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.787105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.322048Z digest=sha256:59529d4e13639b62ae6e0f1d7b9b43be398edf6761244e426cc0725f803fe640

Observation 6755346c-bcb1-4235-8fe7-e6c08722c841 · outbound

This paper cites Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.613019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.324853Z digest=sha256:33f3476c4991ca33cba8f573dc3c3c23236abc9b3407ef7f802311f2d8f3b1b0

Observation 99e89c04-a22a-4938-bea5-2df5c91abf9e · outbound

This paper cites Fashion iq: A new dataset towards retrieving images by natural language feedback.

Illuminating Visual Identity in Universal Multimodal Embeddings Fashion iq: A new dataset towards retrieving images by natural language feedback

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.467311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.327787Z digest=sha256:ed14ddae2fe0bbe8b2d6d2764f377813e42755a86cc021d91594d8eabdd5cc02

Observation d03ce222-b0c5-458c-aa79-c6143ef507e5 · outbound

This paper cites Forb: a flat object retrieval benchmark for universal im- age embedding.Advances in Neural Information Processing Systems, 36:25448–25460, 2023.

Illuminating Visual Identity in Universal Multimodal Embeddings Forb: a flat object retrieval benchmark for universal im- age embedding.Advances in Neural Information Processing Systems, 36:25448–25460, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.366074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.330652Z digest=sha256:d46c468603b23ade503317ac3858fc35e49b2f86004de1ac3381d45a4acc0089

Observation 46449807-f27c-469a-8b8c-f762153cc6e0 · outbound

This paper cites Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025.

Illuminating Visual Identity in Universal Multimodal Embeddings Withanyone: Towards control- lable and id consistent image generation.arXiv preprint arXiv:2510.14975, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.333489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.333489Z digest=sha256:5eb327fa2834373e46e3af99f6128ec04796873e8c90d55713e9d66bacdf3ede

Observation d2e52caf-a117-4d20-864f-ee1847e6c1db · outbound

This paper cites Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying.

Illuminating Visual Identity in Universal Multimodal Embeddings Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.336608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.336608Z digest=sha256:13356f1effa91f0bd6087daf8f986dce67aae6970943b23d88c986137f17b70d

Observation 87e9efb2-e9f0-4622-8961-558d1fdb1567 · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

Illuminating Visual Identity in Universal Multimodal Embeddings Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.339889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.339889Z digest=sha256:8e7e17ba33a0989293367fdcfc1c45c527efe1108938b3d5609289d941b1f667

Observation 7af4341d-dbac-4eda-b7b4-20a07289fc47 · outbound

This paper cites A large-scale car dataset for fine-grained categorization and verification.

Illuminating Visual Identity in Universal Multimodal Embeddings A large-scale car dataset for fine-grained categorization and verification

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.229304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.342861Z digest=sha256:e15d90c4f62f98e6d9c17bc1b652fe93fd3bab7ede0b8d80255912cea13d7c26

Observation e90943c0-354b-48fb-82e6-fcba2ef6a499 · outbound

This paper cites Deep learning for person re- identification: A survey and outlook.IEEE transactions on pattern analysis and machine intelligence, 44(6):2872–2893,.

Illuminating Visual Identity in Universal Multimodal Embeddings Deep learning for person re- identification: A survey and outlook.IEEE transactions on pattern analysis and machine intelligence, 44(6):2872–2893,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:19.085925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.345722Z digest=sha256:89f7e140e863463f9aef989efedcd41c03985f873b943910b3e6fd898686026f

Observation 7cd99d99-95d9-4e89-a6ef-3f036daaf510 · outbound

This paper cites Learning Face Representation from Scratch.

Illuminating Visual Identity in Universal Multimodal Embeddings Learning Face Representation from Scratch

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.348652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.348652Z digest=sha256:911a174b30cab632b1368014a3ee0e7c7817d9e1ae413f829a4eda92e358d0fb

Observation ac953921-ea05-4cee-a1ac-eed9274e9d03 · outbound

This paper cites The met dataset: Instance-level recognition for artworks.

Illuminating Visual Identity in Universal Multimodal Embeddings The met dataset: Instance-level recognition for artworks

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.941829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.351877Z digest=sha256:9a652bcaccfbd6e89680d3c0639e62a2a7c3ed0cff15f50a4c6dc51b2f555628

Observation 2bd6fb3a-74be-4cc9-a278-c3e40dd7f15b · outbound

This paper cites Towards universal image embed- dings: A large-scale dataset and challenge for generic image representations.

Illuminating Visual Identity in Universal Multimodal Embeddings Towards universal image embed- dings: A large-scale dataset and challenge for generic image representations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.824469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.354680Z digest=sha256:f972a783b6fabe91dfa98119ebbec9f465896b834bd5d20554fb2d834f63cb9a

Observation 9cbd0a3e-e878-45fd-857a-ffbc63cac070 · outbound

This paper cites Udon: Universal dynamic online distilla- tion for generic image representations.Advances in Neural Information Processing Systems, 37:86836–86859, 2024.

Illuminating Visual Identity in Universal Multimodal Embeddings Udon: Universal dynamic online distilla- tion for generic image representations.Advances in Neural Information Processing Systems, 37:86836–86859, 2024

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.694822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.357587Z digest=sha256:a788a43013c83515f7ece15d09c671ac469fb329ae3215f1ea48782e6395b5e0

Observation 931fdf07-6b8d-4b55-98f8-1f24ea3029e9 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Illuminating Visual Identity in Universal Multimodal Embeddings CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.360998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.360998Z digest=sha256:2c7536075827aa1bdb3cf7560641213f417e052c3ec9ea9171dd13242e9c9fc9

Observation 8178b8ee-26cc-4913-9162-c5f2d8ad71e5 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

Illuminating Visual Identity in Universal Multimodal Embeddings VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.364610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.364610Z digest=sha256:80cea5e9b9b92240e6fb11d92ba06f6bfca28a2da3d2fa65044d0473d9017c0b

Observation 6c629151-d8cf-4adc-8784-379eba4fed65 · outbound

This paper cites Identity- preserving text-to-video generation by frequency decompo- sition.

Illuminating Visual Identity in Universal Multimodal Embeddings Identity- preserving text-to-video generation by frequency decompo- sition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.558554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.367779Z digest=sha256:02b20117781f01d4e35e9ce92b62196e7ecf4f5189537750b23ccfc85dfaa35a

Observation 40690f5a-428d-4b23-8919-477e13ee1eb5 · outbound

This paper cites Sigmoid loss for language image pre-training.

Illuminating Visual Identity in Universal Multimodal Embeddings Sigmoid loss for language image pre-training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.370533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.370533Z digest=sha256:5d3056074211a8a4f32b154cc8f8982930d44e3293444dbdfff91781f196f7c3

Observation b18f833e-7fe4-4b36-9661-ab9b29429545 · outbound

This paper cites Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining.

Illuminating Visual Identity in Universal Multimodal Embeddings Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.413590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.373366Z digest=sha256:07401cbb65ad60eb590e3cc49377d8af91c04cbba1012a7319e9e8b76e2bcc42

Observation 5fd02a67-b5ed-461a-99b7-0bc1bf3d7fdd · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

Illuminating Visual Identity in Universal Multimodal Embeddings MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.376305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.376305Z digest=sha256:44b55d5be5f0dde31e90e81c3b4564f17a5ce40b49b875a1133db0243cd223ee

Observation 9fd7f93b-194a-406e-b802-404beba054a2 · outbound

This paper cites Assess- ing and learning alignment of unimodal vision and language models.

Illuminating Visual Identity in Universal Multimodal Embeddings Assess- ing and learning alignment of unimodal vision and language models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.379454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.379454Z digest=sha256:c1cc50d608bd38ef6ed1b46a6fe88f5f307c8f12c2734fdf438ec348438fa284

Observation 1dde37f5-d3b3-4440-9a43-fc552cb4ac59 · outbound

This paper cites Beyond frontal faces: Improving per- son recognition using multiple cues.

Illuminating Visual Identity in Universal Multimodal Embeddings Beyond frontal faces: Improving per- son recognition using multiple cues

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.275526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.382385Z digest=sha256:b8f3bbd66e5c960b1e282507ff5818dbd51b27849f344dab9205ebb11120dea9

Observation a68bac3d-d1c0-410e-b9bb-f932ca113872 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

Illuminating Visual Identity in Universal Multimodal Embeddings GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.485401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.485401Z digest=sha256:9ec9cd281e54894143ff4332339590e983da8b0d48049314d093fbc4da78f606

Observation 1ea36a60-e6aa-4864-a5f7-160629ce6583 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

Illuminating Visual Identity in Universal Multimodal Embeddings Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.489048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.489048Z digest=sha256:ec2c66ca37c751b28b215dc3daf9dec7861031525eb48d763f9646c91ab84a3a

Observation e18f69cb-8fdc-4673-917f-1354f340451b · outbound

This paper cites Magicmirror: Id-preserved video generation in video diffusion transformers.

Illuminating Visual Identity in Universal Multimodal Embeddings Magicmirror: Id-preserved video generation in video diffusion transformers

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.119414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.492298Z digest=sha256:e33ff0615907d913d1572191086f2dfb5f8e746b317a6492d394ae73938ca7c1

Observation 0d85fe8b-7bc2-406a-87ae-95816f61eebe · outbound

This paper cites Scalable person re-identification: A benchmark.

Illuminating Visual Identity in Universal Multimodal Embeddings Scalable person re-identification: A benchmark

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.035923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.495323Z digest=sha256:2022295c20503bc312834fe746b214b816e2f24a0671ac4cfb35b51f35f1bfc5

Observation da79ef23-b212-4b4f-999e-de3b37100501 · outbound

This paper cites Scalable person re-identification: A benchmark.

Illuminating Visual Identity in Universal Multimodal Embeddings Scalable person re-identification: A benchmark

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:18.010772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.498322Z digest=sha256:c5d6d9e9ca20c5f1d891ff1a30d36a4b7357277394fc8054e43dbf4c5fbd9c5a

Observation 10ef6a1b-c816-419f-b476-c14682660d44 · outbound

This paper cites Cartoon face recognition: A bench- mark dataset.

Illuminating Visual Identity in Universal Multimodal Embeddings Cartoon face recognition: A bench- mark dataset

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:17.983167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.501286Z digest=sha256:3dc30678374e2825112fa31622ef8580c1e1fd50598808dadd328f197d76ec73

Observation 47e166c9-8827-4454-9a0a-97499ad8221c · outbound

This paper cites Megapairs: Massive data synthesis for universal multi- modal retrieval.

Illuminating Visual Identity in Universal Multimodal Embeddings Megapairs: Massive data synthesis for universal multi- modal retrieval

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:43:17.963805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T20:43:17.504442Z digest=sha256:e913060fea79159038e828a9761113f803282c3ceaac5cd0f0da2beacff17309

Observation 7f78a490-0d0c-41d6-bd8c-2321a9113e37 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Illuminating Visual Identity in Universal Multimodal Embeddings InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 89

Resolution
malformed identifier
no resolver link, observed 2026-08-04T20:43:17.507439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.507439Z digest=sha256:c6cc67b66dd333f6366211e44f0a85376eef2169e88186dbee4d37af824e015b

Pith citing papers

No inbound Pith citation observations are available.