Pith. sign in

Paper Citation Record · LEDGER

Speech Retrieval-Augmented Generation without Automatic Speech Recognition

As of 20 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2412.16500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16500 v3

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:34:01.593732Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:01.375758Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:55:02.618954Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 80eb3137-2701-400b-962c-0f608cdad2d5 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval- augmented generation for knowledge-intensive nlp tasks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.180136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.422535Z digest=sha256:4c2e86c957bf754afb6b0f8ab211491eb1d62fc34df5e937d537e28d8356bacc

Observation 84b65983-1dc2-4833-9c0b-4da3151f8c4c · outbound

This paper cites MuRAG: Multimodal retrieval-augmented generator for open question answering over images and text,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition MuRAG: Multimodal retrieval-augmented generator for open question answering over images and text,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.162110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.428686Z digest=sha256:fa298d1f2c0ba8abe534c9ec7776d7bf3de8c89f7eebf1d134bf801798b8f134

Observation 39ca2426-c6be-42bc-922d-23e41597bfe7 · outbound

This paper cites Robust multi model rag pipeline for documents containing text, table & images,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Robust multi model rag pipeline for documents containing text, table & images,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.144348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.434116Z digest=sha256:11de11f4e8a628125e4722f06afe133071881f9eff1ead2cc3c25bc5a78dcb48

Observation 60b72f7b-0e0d-43e5-b74e-15ce10ff1212 · outbound

This paper cites Spo- ken content retrieval—beyond cascading speech recognition with text retrieval,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Spo- ken content retrieval—beyond cascading speech recognition with text retrieval,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.126941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.439667Z digest=sha256:ca52e8e4e454a10689ed6d50498d6076df8af0759018942b6ac669c8b1fef4f7

Observation c4c1b470-422d-4b5a-80eb-72c500128016 · outbound

This paper cites Retrieval and browsing of spoken content,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval and browsing of spoken content,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.110068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.445300Z digest=sha256:009106909c600ddb670b7621dfbb71c6b8d24c2abcd4a76a6179634736441ab9

Observation ba7ae7a4-2f8e-4c14-9420-1e49c7dd19aa · outbound

This paper cites Robust speech recognition via large- scale weak supervision,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Robust speech recognition via large- scale weak supervision,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T10:34:01.451168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:34:01.451168Z digest=sha256:0651a1552773fbfff3adb04e7ae98c8ae221238be1958fdef50318821e483295

Observation d03f8691-a22a-42f2-be8b-0e29458e1921 · outbound

This paper cites MTEB: Massive text embedding benchmark,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition MTEB: Massive text embedding benchmark,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.082895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.457101Z digest=sha256:0ae58ed6879b1651a362048a2ccd20d96ae2d11ac12b7ade85c286bf0c84bc8b

Observation 99dec45b-bbb6-4086-adbf-3055a5464571 · outbound

This paper cites OLISIA: a cascade system for spoken dialogue state tracking,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition OLISIA: a cascade system for spoken dialogue state tracking,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.065497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.462412Z digest=sha256:b355ec0ec286cb2110e93008440ca68d0c11b041873139af6486a66e7c8f74db

Observation edf4fadd-4547-48ed-af7d-26f78f191f11 · outbound

This paper cites Towards end-to-end spoken language understanding,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Towards end-to-end spoken language understanding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.048320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.467782Z digest=sha256:7f84b1f19aaa8aa5771a5fefee82485c3d1d7146b24be27336e72170dc77f21b

Observation a9112a81-faf9-42f6-9787-23fc027e9026 · outbound

This paper cites Why aren’t we NER yet? artifacts of ASR errors in named entity recognition in spontaneous speech transcripts,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Why aren’t we NER yet? artifacts of ASR errors in named entity recognition in spontaneous speech transcripts,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.031321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.472945Z digest=sha256:aa07dbd52273714860fb0c18e4faf33ee1dc297e11de53360b3229e20b8ecd83

Observation 8ae5eaac-430e-4c54-8f36-a34d044740a0 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Learning transferable visual models from natural language supervision,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:34:01.478392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:34:01.478392Z digest=sha256:bf2ddcd366cf6dca290e406a5c71dd0a0bcaaf0c2926074338c6c9b222157ed8

Observation 2514954a-6bc6-46f5-87bb-86f39ee438c8 · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:02.003175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.483448Z digest=sha256:16dd61ee9c61ee446e8924f9b795b7d938523b4664e724918858d085ef7f0ef9

Observation ed0286f4-3792-4a2c-aaa7-1bc7387d4cfd · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Clap learning audio concepts from natural language supervision,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.987577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.488158Z digest=sha256:0418f49035937c13a39693901302b64db311a235ec753f08f28af3f271d60050

Observation 26c5c16a-708f-4cb6-8f4d-84e0183889e4 · outbound

This paper cites Contrastive learning with hard negative samples,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Contrastive learning with hard negative samples,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.971950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.492989Z digest=sha256:f03ea51bd9f5ed6eedb1cf375b989452def231c1b1afb94276ddf90cb7c5974b

Observation c7b95544-11e1-4f3c-a95d-1d8e3de65d64 · outbound

This paper cites Why do we need large batchsizes in contrastive learning? a gradient-bias perspective,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Why do we need large batchsizes in contrastive learning? a gradient-bias perspective,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.956476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.497949Z digest=sha256:8c180def74530540ebd0dfa9c41a5e55db6b8f793727098d54b8fd73043f8a90

Observation 9a0790f5-60b6-42ed-a15b-b1d6987ce91d · outbound

This paper cites SONAR: sentence-level multimodal and language-agnostic represen- tations,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition SONAR: sentence-level multimodal and language-agnostic represen- tations,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.940828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.502602Z digest=sha256:4b28882eea3a5183e0282525d8d9b2611bf2e2d07d4aef4e4bc6a48055d8261e

Observation 67e43f76-6cad-4edd-8452-8aac1dac2249 · outbound

This paper cites Audio retrieval with natural language queries,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Audio retrieval with natural language queries,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.924405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.507919Z digest=sha256:385a14fcd10c2b5d2cb7816666319d146ee53546036c10936f7837830082b4cd

Observation 51a67ffe-e91a-40c1-aa01-4775cd0ba3ec · outbound

This paper cites SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:34:01.513076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:34:01.513076Z digest=sha256:68bb2ea76d180242079877f83a08b402b9f24d5e12543874ded2f9943fad06d7

Observation a1f07396-8ef3-4204-81a1-bbfb223a65b0 · outbound

This paper cites Recap: Retrieval-augmented audio captioning,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Recap: Retrieval-augmented audio captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.907394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.518207Z digest=sha256:16f0ee57d2e7366e01499cfe85a53bd2968fb806458fbfad4f3f6c137f1af856

Observation 46ee1afd-8d3f-44e2-979d-f4bd80ab0a9f · outbound

This paper cites Retrieval Augmented End-to-End Spoken Dialog Models.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval Augmented End-to-End Spoken Dialog Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:34:01.522941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:34:01.522941Z digest=sha256:b8e26c8df755c3f0c99c9de6877856ed4806b18da571d731c56087dc0bbc4a58

Observation f30801ac-8bc5-4c37-8c01-246e62884db1 · outbound

This paper cites Speechdpr: End-to-end spoken passage retrieval for open-domain spoken question answering,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Speechdpr: End-to-end spoken passage retrieval for open-domain spoken question answering,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.891128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.528930Z digest=sha256:318fdf8fe551d2e52bcbe9afd7c0512498a31974991e6ddfe1f36eb6bb2206d2

Observation cc8d95bd-aaa4-4521-b796-b4ca13d099fb · outbound

This paper cites Retrieval augmented end-to-end spoken dialog models,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval augmented end-to-end spoken dialog models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.873780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.533776Z digest=sha256:45228c3e49fd1017d5b60b2a2ac4206b6283d67f2fce75269e99b2599521e8b5

Observation 80b3516b-b725-4286-bbdd-2c265c5709ff · outbound

This paper cites Speechverse: A large-scale generalizable audio language model,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Speechverse: A large-scale generalizable audio language model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.856407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.538960Z digest=sha256:5383484928276c819f55e3856d7b2e4bbc87dd826adc5303384922334be8b734

Observation aec0e83a-8cee-4213-8c15-47daf8c489c3 · outbound

This paper cites Hubert: Self- supervised speech representation learning by masked prediction of hidden units,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Hubert: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.838272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.543232Z digest=sha256:db527c2cf45c2a52a755f61b1b98c14758cf7ddb36f38db67a32651c9b0f40d3

Observation 52cd75fa-1e3a-4fb5-bb42-2cebc847cdde · outbound

This paper cites Prompting large language models with audio for general-purpose speech summarization,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Prompting large language models with audio for general-purpose speech summarization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.820592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.547778Z digest=sha256:ace029fd803497729e2768d29d4c782bef6822e1283d9f068954eef0f4e1ae32

Observation 7ae6f0a0-52b8-47fd-a428-b97948288b39 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Sentence-bert: Sentence embeddings using siamese bert-networks,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.803413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.552100Z digest=sha256:96695d1640fcdad21f5b9b51cccde12c67d1b95fce59967de99210d164762748

Observation 005b85e5-cd82-4d68-98c1-fa2b911e6bdb · outbound

This paper cites Improving text embeddings with large language models,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Improving text embeddings with large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.784231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.557277Z digest=sha256:15e59e25280ec7bfa84aa907f98850db1fa9b5dc28f2cf79d9e72a9f09644b94

Observation 1fed0f5e-61b2-4c79-93c5-9b65e5f0c3f6 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T10:34:01.561642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:34:01.561642Z digest=sha256:0b7d501eedd84127a8f68350d43cb8bf41bd92ff2a93051c34099ddf1bec2454

Observation 52ffbcbc-bb4c-47e2-9821-5f39c24c8c18 · outbound

This paper cites Spoken squad: A study of mitigating the impact of speech recognition errors on listening comprehension,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Spoken squad: A study of mitigating the impact of speech recognition errors on listening comprehension,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.764343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.567131Z digest=sha256:fa223a75ca171f18795167563f7f616ff1b571f85b86635304b4371b25ef5a18

Observation d7e6ad55-28be-4e19-80e7-489c99aa0ab2 · outbound

This paper cites V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.746702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.572675Z digest=sha256:0775cc16962f10019553a2613eca6b382f1b669bb77ed691129c27068fc1650e

Observation 3b0cdaea-edae-486a-b9a4-8340b3ec3e03 · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition SQuAD: 100,000+ questions for machine comprehension of text,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.729565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.578224Z digest=sha256:08e1c723451c3b942e32bf0c6732eda35ecb22fe210813673371378e6f1ee23f

Observation 86d84a70-c3a6-4037-a220-6e57cb779aa8 · outbound

This paper cites Introduction to the CoNLL-2003 shared task: Language-independent named entity recogni- tion,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Introduction to the CoNLL-2003 shared task: Language-independent named entity recogni- tion,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.712449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.583421Z digest=sha256:bba5de7a635df0b97ed728cb0e5b1908e6b532e777838aa0db9a76a2e2d5df90

Observation 3a5c6793-995a-4c02-9eee-4141dfd3d35a · outbound

This paper cites Evaluation of rag metrics for question answering in the telecom domain,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Evaluation of rag metrics for question answering in the telecom domain,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:34:01.694861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:34:01.588611Z digest=sha256:531079acf81ac87189553fee3c702141725b32b24b821e809868bde6e66c2979

Observation b959645b-2a80-4083-bac5-148f25e17a30 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books,.

Speech Retrieval-Augmented Generation without Automatic Speech Recognition Librispeech: An asr corpus based on public domain audio books,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:34:01.593732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:34:01.593732Z digest=sha256:ebfdd7d5e3f4c39219036217924f1fb9f102322c98cfb4db7aa8deb13f2fd24a

Pith citing papers

Observation b97890f7-17c2-4981-acd4-fe1cc4b5d397 · inbound

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering cites this paper.

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering Speech Retrieval-Augmented Generation without Automatic Speech Recognition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:55:02.723063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T14:55:01.375758Z digest=sha256:2b3f096b28c05e1201f2f70048dabf56de92e4f32138431f62b39bab5e41e88b