Pith. sign in

Paper Citation Record · LEDGER

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

As of 13 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2509.07538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07538 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:14.726012Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:12.301504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T22:06:14.962988Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 606b6d8b-d461-4960-af5b-173171b3b925 · outbound

This paper cites TextlessRAG: End-to-End Visual Document RAG by Speech Without Text.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-04T22:06:15.028832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:12.301504Z digest=sha256:df8ce640c7bdf771211e85a9fe070461f5dd4e0a896b8354088c5aaaf3a45714

Observation 600a6e06-3e4b-4834-b9ad-dbfe3a638a6e · outbound

This paper cites an unresolved cited work.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:06:17.673515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:12.385386Z digest=sha256:22e6f93162b63a4ae498a772f50f2cae08ad8767be37e9521c3df88c336584a0

Observation 47eab786-962c-4950-91ee-8c1849fab26a · outbound

This paper cites QA” and “Pool.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text QA” and “Pool

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.545134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:12.476150Z digest=sha256:2eadbe4cd94dca3268ae957ea3f1c9d386dc5d59a0eb9472ac15060fbea2c7b4

Observation 1bacaeaf-5a0f-4afa-abb0-3abaca15851a · outbound

This paper cites T”, “I”, and “A.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text T”, “I”, and “A

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-04T22:06:17.451643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:12.554832Z digest=sha256:83b2bafa336ef731016d5fc5fa1ab65b1e07ff219514b500600c158f33009c89

Observation 3440a9c0-d2b6-40a1-a534-d80f8976fdb2 · outbound

This paper cites We also introduce the first bilingual bench- mark for this task and release the first open-source Chinese visual document RAG dataset.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text We also introduce the first bilingual bench- mark for this task and release the first open-source Chinese visual document RAG dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.317046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:12.640657Z digest=sha256:6469f964f38e580de256abf3c103b5bfc65253821278c1f9bf1c720c7b7fbe6e

Observation 1423b882-83c0-40a3-8715-4419875f8612 · outbound

This paper cites Qwen2.5-VL Technical Report.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:12.726138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:12.726138Z digest=sha256:d9776491ec47650f93a693e29259628db4605673fb17e7ace03e1fc724d594a7

Observation 0f0f9cdd-5cb6-45df-a3b4-b80af33a98cd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:12.859876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:12.859876Z digest=sha256:f83d43e47b4d1a970721902ff01d920f048961cfbf955afaef18811b72912ad8

Observation 871b664e-c294-4bcb-b8d9-0f3f6e942b40 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.214049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:12.949476Z digest=sha256:2e6d3ccd12f89dff1859f740a299199cfdaf08686fe54981e4d3507c4f7d865a

Observation d34ea977-7fb2-4015-9f58-671224254057 · outbound

This paper cites Qwen2.5-Omni Technical Report.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-Omni Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:13.082058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:13.082058Z digest=sha256:3c2ad6e10dbf9fa371ac94d6e356740fcb7fbc1d268b1d191a9479b2a539934a

Observation 32547a26-8c4f-457a-99e8-5ac6ce85f793 · outbound

This paper cites Multimodal large language models for text-rich image understanding: A comprehensive review,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Multimodal large language models for text-rich image understanding: A comprehensive review,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.120494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:13.228865Z digest=sha256:10c03a6095d1dbb5d983313a133a2dd49e28dc67f2bef4bf991ac623e5533e79

Observation 95a827ec-44a4-4fe2-bbe2-090653ed6d23 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context doc- ument understanding with visualizations,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Mmlongbench-doc: Benchmarking long-context doc- ument understanding with visualizations,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.990451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:13.373610Z digest=sha256:50d08ed40bfb1340da0e0482aaf8493f895d24ab44e5af3af663861a0c66edee

Observation 562fc4ec-99ff-4236-b3c1-ac4a53790dd8 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple im- ages,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Slidevqa: A dataset for document visual question answering on multiple im- ages,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.851276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:13.458860Z digest=sha256:91f401d284c3c4fc160899bdcd681020d995dc21ffc02b9bb3643a66422a0bba

Observation 480d65f9-44d7-4389-96f5-54c4f5ff02ee · outbound

This paper cites Visrag: Vision-based retrieval-augmented gener- ation on multi-modality documents,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Visrag: Vision-based retrieval-augmented gener- ation on multi-modality documents,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.748304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:13.685022Z digest=sha256:ef576f3d8f14d29be78c59ebabfdadb204309c0b334a580f64414f62b9f414be

Observation cf69189f-c193-4101-ac81-a1f6822cfb10 · outbound

This paper cites Vdocrag: Retrieval-augmented generation over visually-rich documents,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Vdocrag: Retrieval-augmented generation over visually-rich documents,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.643795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:13.763678Z digest=sha256:8153b372aca8253f5d0d0a1bf82fb2c1f3e45a667a72e5e6ebfe75aea3dd2e81

Observation 7cbf5f85-c8b1-4fce-a48e-2d61d87c3c8c · outbound

This paper cites ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:13.857070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:13.857070Z digest=sha256:d0e04992ea6d8bcbccf47dd51e19cda88e4ed4d2a4f8a642e62caa137da19f6f

Observation c36e5c04-4cf5-4a1e-b3f3-c55cc9e501dd · outbound

This paper cites Towards multilingual spoken visual question answering system using cross-attention,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Towards multilingual spoken visual question answering system using cross-attention,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.545105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:13.956121Z digest=sha256:0944578b6c3104e619acd76dc1847ca7b242b44cc5b058d584c6b42b41aafbd8

Observation 3629ed68-c9d4-4e57-899b-ccbee08d4cd5 · outbound

This paper cites Spoken question answering for visual queries,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Spoken question answering for visual queries,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.411574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.052491Z digest=sha256:d8f4b810298ca26928846a9c4db3138fb47ebcc3a60be9324ae84d6c8a58ae95

Observation 7cd47da3-cd8c-4f36-9db2-675f04a48a10 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.316678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.163239Z digest=sha256:6679d4d5f0824ec650e075fcf7a940b261972673da737bfd9771a9dd4ed886ab

Observation 643ea55e-f9f3-40c9-a69c-72a08ff89757 · outbound

This paper cites Infographicvqa,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Infographicvqa,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.186822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.279547Z digest=sha256:7a7ac719b9bd44f6b02588b655a4e71b1b7d4c06db5239adc30d9c8d5904324a

Observation b53390ec-f5f9-4d6a-9ad6-698038c1d832 · outbound

This paper cites ICDAR 2023 competition on document understanding of everything (dude),.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ICDAR 2023 competition on document understanding of everything (dude),

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.083954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.365990Z digest=sha256:40a6c16cd29b840990a40471ff7d23c67a366f372c0d37f688fd28a65c35b33d

Observation 451b199f-c2f9-4bf3-b9a5-4bffa7b36efa · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text The probabilistic relevance framework: Bm25 and beyond,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.934745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.495250Z digest=sha256:6b08da88542437d0a3dfd4cc165d815d7c2dd1696094248e6d951138c35e2790

Observation 4298f707-c53d-4f71-9f43-94128b5750ac · outbound

This paper cites Text embeddings by weakly-supervised contrastive pre-training,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Text embeddings by weakly-supervised contrastive pre-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.783246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.561399Z digest=sha256:2ed4bc0a498f22b5d000993c029a507195c2cb4d14c48995ce9c8726cfa78518

Observation 61a5874e-f27b-4497-a356-f2e1dfdda83f · outbound

This paper cites Nv-embed: Improved tech- niques for training llms as generalist embedding mod- els,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Nv-embed: Improved tech- niques for training llms as generalist embedding mod- els,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.560850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.635933Z digest=sha256:1cadc589ea305c878e0370b8c9898ffb235d4b5d87e6fc559c95872e266d804a

Observation a17f529f-8511-49a9-9759-ce5e9170135b · outbound

This paper cites Learning transferable vi- sual models from natural language supervision,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Learning transferable vi- sual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.372478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.705712Z digest=sha256:2f5f0e9a0622de6baa875f96a45cb6a3de8dcdd4115770ef000fbc99197982e5

Observation 79f12e5e-06ae-4220-b8e0-5c89fe9d0eca · outbound

This paper cites Uni- fying multimodal retrieval via document screenshot em- bedding,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Uni- fying multimodal retrieval via document screenshot em- bedding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.231025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:14.726012Z digest=sha256:12a1d663ca6ca235ed9a93f6e7cf821a821dc86d74028df6489b0ef448ee0365

Pith citing papers

Observation 606b6d8b-d461-4960-af5b-173171b3b925 · inbound

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text cites this paper.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-04T22:06:15.028832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-04T22:06:12.301504Z digest=sha256:df8ce640c7bdf771211e85a9fe070461f5dd4e0a896b8354088c5aaaf3a45714