Pith. sign in

Paper Citation Record · LEDGER

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.04292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04292 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:19:53.886239Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.252122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e78c84d3-b0fa-4346-abd8-c07bf685350a · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:21.010148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:754cfa18e7de28911a19d72973f5a11d66485e8a782f8b4c26afaa1e29d13c3e

Observation 527ce582-fe2c-4f10-88d6-e09feb5f7953 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.859394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:499d021292a53362925ff5d415d4d40739e23893a7ce3eb61c705b12e6e87a56

Observation 62fa00f9-d82b-4ced-a41e-43fc4c60dcbe · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:44.622251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:e648ed09f0798eeeab5d57525c7872efdb1c7a1999099ce5294c8ff6292e48aa

Observation 79da0965-9f65-40f9-9e30-40f02bcb645f · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.218371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.218371Z digest=sha256:d18fc7ab7f9687a37d97a115fecb7fde490448d17fcd24723e5ec5a1a5524d0d

Observation 06be4929-c3df-439c-8cd8-4a42f4069f70 · inbound

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG cites this paper.

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.170740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T07:19:44.125479Z digest=sha256:101a8679ea9a95f64237a82105cbee6e7275255eb1549e6e710c5e007dbb892a

Observation 0ce5236c-ea7f-4cb5-88f4-1e587a8e86e5 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.423530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:e011148afe01e7e3e5afb3b742b52a5b3be1bace1d3ea723299395ff2e4f3ad4

Observation c764cc1d-0431-4631-ac7a-56970e5dd1b3 · inbound

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception cites this paper.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.255293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:8f6839e264d27acae3bc95547b9027e76a2e19edde1913354a2a70cebfd9c5b2

Observation 39f0b923-903b-4568-9ae5-62839bfda0a8 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.822604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.822604Z digest=sha256:06ad35297e0053ed95acd162773bb3453e00b584aa3b58c8b39826d156a0d0c8

Observation cac9af94-b7cd-47cc-baff-5f9517ab36f1 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:53.886239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:53.886239Z digest=sha256:62a5e87d51ccb10b64513522153bcc95aa3b74139562075542fa5b1a54c61bf9