Pith. sign in

Paper Citation Record · LEDGER

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

As of 9 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2509.07538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07538 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:14.726012Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:06:12.301504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T22:06:14.962988Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 606b6d8b-d461-4960-af5b-173171b3b925 · outbound

This paper cites TextlessRAG: End-to-End Visual Document RAG by Speech Without Text.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-04T22:06:15.028832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:12.301504Z digest=sha256:716863ba8589e73e97142f4c6f5edb6109e076940d540328dbef034a626a97b1

Observation 600a6e06-3e4b-4834-b9ad-dbfe3a638a6e · outbound

This paper cites an unresolved cited work.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-04T22:06:17.673515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:12.385386Z digest=sha256:dcb7bccd6dc2aa9b3a15f4d1e73c9002368a345584978eae64c57782629b4c3b

Observation 47eab786-962c-4950-91ee-8c1849fab26a · outbound

This paper cites QA” and “Pool.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text QA” and “Pool

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.545134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:12.476150Z digest=sha256:de05af4d7c58b235757342122815bbd398619dd73b38f5349f9b0306e07aa82a

Observation 1bacaeaf-5a0f-4afa-abb0-3abaca15851a · outbound

This paper cites T”, “I”, and “A.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text T”, “I”, and “A

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-04T22:06:17.451643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:12.554832Z digest=sha256:eec19b1bd9db117b9aea2fc50b11ff2da73152d12c923f20743d1754a04c2840

Observation 3440a9c0-d2b6-40a1-a534-d80f8976fdb2 · outbound

This paper cites We also introduce the first bilingual bench- mark for this task and release the first open-source Chinese visual document RAG dataset.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text We also introduce the first bilingual bench- mark for this task and release the first open-source Chinese visual document RAG dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.317046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:12.640657Z digest=sha256:2c3566cebfc7b1b528c9ccdfa5d3180b130c0de8a78f3dc77d3639884b2c7774

Observation 1423b882-83c0-40a3-8715-4419875f8612 · outbound

This paper cites Qwen2.5-VL Technical Report.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:12.726138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:12.726138Z digest=sha256:cee0ad18629486d80deb33f741fc453d103b954ceace405b9099cca62864d2bb

Observation 0f0f9cdd-5cb6-45df-a3b4-b80af33a98cd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:12.859876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:12.859876Z digest=sha256:ae14d52ecd9b68472756fae9664828b7abe7db0f062e8ee111c2cfe163d6e169

Observation 871b664e-c294-4bcb-b8d9-0f3f6e942b40 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.214049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:12.949476Z digest=sha256:89257da8d7207317da939d4e1046cfd60a3ca32824d5bbbc9d6f6e9e1df81bac

Observation d34ea977-7fb2-4015-9f58-671224254057 · outbound

This paper cites Qwen2.5-Omni Technical Report.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Qwen2.5-Omni Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:13.082058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:13.082058Z digest=sha256:ac00fe9d33f893dd9a6819c28dfd7b2206eb268b8d91edf1ea6c3c884e0cbf54

Observation 32547a26-8c4f-457a-99e8-5ac6ce85f793 · outbound

This paper cites Multimodal large language models for text-rich image understanding: A comprehensive review,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Multimodal large language models for text-rich image understanding: A comprehensive review,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:17.120494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:13.228865Z digest=sha256:c86729269503dd37dea99c4cfcb80a0cdaff87de6340ffcaec7969cbe9fd953d

Observation 95a827ec-44a4-4fe2-bbe2-090653ed6d23 · outbound

This paper cites Mmlongbench-doc: Benchmarking long-context doc- ument understanding with visualizations,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Mmlongbench-doc: Benchmarking long-context doc- ument understanding with visualizations,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.990451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:13.373610Z digest=sha256:a603496e3f5f7aeaeba9fd18f3441b3d761d54c3c8d658700a1d006cd13deda1

Observation 562fc4ec-99ff-4236-b3c1-ac4a53790dd8 · outbound

This paper cites Slidevqa: A dataset for document visual question answering on multiple im- ages,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Slidevqa: A dataset for document visual question answering on multiple im- ages,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.851276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:13.458860Z digest=sha256:3ce4147df4b96a9c0eeea67dfbef099575fec0d6ecf7c7795c43da80d1505f64

Observation 480d65f9-44d7-4389-96f5-54c4f5ff02ee · outbound

This paper cites Visrag: Vision-based retrieval-augmented gener- ation on multi-modality documents,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Visrag: Vision-based retrieval-augmented gener- ation on multi-modality documents,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.748304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:13.685022Z digest=sha256:011382e7afeacdab936c7d07f2602a09d6dca1063027d4091c0ea006a199a521

Observation cf69189f-c193-4101-ac81-a1f6822cfb10 · outbound

This paper cites Vdocrag: Retrieval-augmented generation over visually-rich documents,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Vdocrag: Retrieval-augmented generation over visually-rich documents,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.643795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:13.763678Z digest=sha256:104cd812cf3a484709bb315e4c906253e0ac91f2a7b1e046d2d40655034cb804

Observation 7cbf5f85-c8b1-4fce-a48e-2d61d87c3c8c · outbound

This paper cites ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:06:13.857070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:06:13.857070Z digest=sha256:263d525fbba8dc99c98063a16c58599ca91684c6b315230b06019953fa5a6a62

Observation c36e5c04-4cf5-4a1e-b3f3-c55cc9e501dd · outbound

This paper cites Towards multilingual spoken visual question answering system using cross-attention,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Towards multilingual spoken visual question answering system using cross-attention,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.545105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:13.956121Z digest=sha256:9ce6c06479f08368772f8c023d15aa48e10a1ca2980e131059f3da19f8c1b089

Observation 3629ed68-c9d4-4e57-899b-ccbee08d4cd5 · outbound

This paper cites Spoken question answering for visual queries,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Spoken question answering for visual queries,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.411574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.052491Z digest=sha256:520eb25babad62e5ccb2a9e029244de48a1e3c6aad2d7f14fdcfe9b443f11d63

Observation 7cd47da3-cd8c-4f36-9db2-675f04a48a10 · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.316678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.163239Z digest=sha256:e0bca5c68dc11301cb34c087bfe4ccb69df57d0ad929e820adb81d0e7ef70b91

Observation 643ea55e-f9f3-40c9-a69c-72a08ff89757 · outbound

This paper cites Infographicvqa,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Infographicvqa,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.186822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.279547Z digest=sha256:429b7205cdc6956fe10fd9a9f9fdc32204f2a96033cba675fd121abc3902f35e

Observation b53390ec-f5f9-4d6a-9ad6-698038c1d832 · outbound

This paper cites ICDAR 2023 competition on document understanding of everything (dude),.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text ICDAR 2023 competition on document understanding of everything (dude),

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:16.083954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.365990Z digest=sha256:88c3e8d0ab7d9427c036944536e8680acd8d97927f90efb2c231b02a5279dcbf

Observation 451b199f-c2f9-4bf3-b9a5-4bffa7b36efa · outbound

This paper cites The probabilistic relevance framework: Bm25 and beyond,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text The probabilistic relevance framework: Bm25 and beyond,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.934745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.495250Z digest=sha256:8af5229be171d218237e5705ab00d874279bf71ea71675d2014c454dfc3694b0

Observation 4298f707-c53d-4f71-9f43-94128b5750ac · outbound

This paper cites Text embeddings by weakly-supervised contrastive pre-training,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Text embeddings by weakly-supervised contrastive pre-training,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.783246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.561399Z digest=sha256:5faffedbd34ce91a8e7c202a0863cb8f032808a709232069c4e37c931928fb00

Observation 61a5874e-f27b-4497-a356-f2e1dfdda83f · outbound

This paper cites Nv-embed: Improved tech- niques for training llms as generalist embedding mod- els,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Nv-embed: Improved tech- niques for training llms as generalist embedding mod- els,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.560850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.635933Z digest=sha256:b69337737979843243e5e99bf280acf035b82f28ab09c540e25566d45d16e2ee

Observation a17f529f-8511-49a9-9759-ce5e9170135b · outbound

This paper cites Learning transferable vi- sual models from natural language supervision,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Learning transferable vi- sual models from natural language supervision,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.372478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.705712Z digest=sha256:ed3dcb41ffa995020ed974a08df64071df7c22a1161ef90a219fda6993f0a64b

Observation 79f12e5e-06ae-4220-b8e0-5c89fe9d0eca · outbound

This paper cites Uni- fying multimodal retrieval via document screenshot em- bedding,.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text Uni- fying multimodal retrieval via document screenshot em- bedding,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:06:15.231025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:14.726012Z digest=sha256:41cdbf87e115f64325f7c479ada91ad15c6829e154cb7e31a39834c85e778dae

Pith citing papers

Observation 606b6d8b-d461-4960-af5b-173171b3b925 · inbound

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text cites this paper.

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-04T22:06:15.028832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T22:06:12.301504Z digest=sha256:716863ba8589e73e97142f4c6f5edb6109e076940d540328dbef034a626a97b1