Pith. sign in

Paper Citation Record · LEDGER

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2406.04292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04292 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:25:12.449077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.252122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e78c84d3-b0fa-4346-abd8-c07bf685350a · inbound

E5-V: Universal Embeddings with Multimodal Large Language Models cites this paper.

E5-V: Universal Embeddings with Multimodal Large Language Models VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:52:21.010148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:52:20.935555Z digest=sha256:a5b5967ff457182525c124781f6e0224aacafc291a3232b88f3db707d5a2b871

Observation 5224de59-74ab-49fe-bb42-b9ca6e9f4159 · inbound

O1 Embedder: Let Retrievers Think Before Action cites this paper.

O1 Embedder: Let Retrievers Think Before Action VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:12.449077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:12.449077Z digest=sha256:0f7b942ec0a4407146b66218ea56fdac829b4142b325f942b5ae9011fafa0fd4

Observation 527ce582-fe2c-4f10-88d6-e09feb5f7953 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.859394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:1e97a735d7234660101434d82369415f41e188b47c6cfff6bf914a3ca096eb86

Observation 62fa00f9-d82b-4ced-a41e-43fc4c60dcbe · inbound

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding cites this paper.

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:17:44.622251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:14:15.589472Z digest=sha256:9dac3460f20615de3bafdae11b831d1133bfd3d9c445e52f15a0dbcfd9b253ee

Observation 79da0965-9f65-40f9-9e30-40f02bcb645f · inbound

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning cites this paper.

Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T19:41:35.218371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:41:35.218371Z digest=sha256:fc6170c85c98e38844845232d85c2a52e617de27a1aad3ecc50a46ca3fa32630

Observation 06be4929-c3df-439c-8cd8-4a42f4069f70 · inbound

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG cites this paper.

Purifying Multimodal Retrieval: Fragment-Level Evidence Selection for RAG VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.170740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T07:19:44.125479Z digest=sha256:979ad91585e69d77876b7e18b7e003f1c4207e600873f4025e7c5a99ea8181d5

Observation 0ce5236c-ea7f-4cb5-88f4-1e587a8e86e5 · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.423530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:94d945d8ae1889ed3d662087e7e2e0f77b6bf3841c4a9af5f9ed8c44d7a06d89

Observation c764cc1d-0431-4631-ac7a-56970e5dd1b3 · inbound

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception cites this paper.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.255293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:6da26c1731c579985076cc2e08ade70b951e025a04c8e76c043193025fdf9e31

Observation 39f0b923-903b-4568-9ae5-62839bfda0a8 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.822604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.822604Z digest=sha256:63042567904b8f602b56713591ed397a5640ad741137ebf6e2d1404139146fee

Observation cac9af94-b7cd-47cc-baff-5f9517ab36f1 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:53.886239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:53.886239Z digest=sha256:1f9181a3cd59d088aca90935ea0d6245997c42bc89f659df213a709e31554660