Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:34:01.593732Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2412.16500.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:34:01.593732Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:01.375758Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T14:55:02.618954Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 80eb3137-2701-400b-962c-0f608cdad2d5 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval- augmented generation for knowledge-intensive nlp tasks,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 84b65983-1dc2-4833-9c0b-4da3151f8c4c · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition MuRAG: Multimodal retrieval-augmented generator for open question answering over images and text,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 39ca2426-c6be-42bc-922d-23e41597bfe7 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Robust multi model rag pipeline for documents containing text, table & images,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 60b72f7b-0e0d-43e5-b74e-15ce10ff1212 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Spo- ken content retrieval—beyond cascading speech recognition with text retrieval,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c4c1b470-422d-4b5a-80eb-72c500128016 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval and browsing of spoken content,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ba7ae7a4-2f8e-4c14-9420-1e49c7dd19aa · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Robust speech recognition via large- scale weak supervision,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d03f8691-a22a-42f2-be8b-0e29458e1921 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition MTEB: Massive text embedding benchmark,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 99dec45b-bbb6-4086-adbf-3055a5464571 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition OLISIA: a cascade system for spoken dialogue state tracking,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation edf4fadd-4547-48ed-af7d-26f78f191f11 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Towards end-to-end spoken language understanding,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a9112a81-faf9-42f6-9787-23fc027e9026 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Why aren’t we NER yet? artifacts of ASR errors in named entity recognition in spontaneous speech transcripts,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8ae5eaac-430e-4c54-8f36-a34d044740a0 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Learning transferable visual models from natural language supervision,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2514954a-6bc6-46f5-87bb-86f39ee438c8 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ed0286f4-3792-4a2c-aaa7-1bc7387d4cfd · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Clap learning audio concepts from natural language supervision,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 26c5c16a-708f-4cb6-8f4d-84e0183889e4 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Contrastive learning with hard negative samples,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c7b95544-11e1-4f3c-a95d-1d8e3de65d64 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Why do we need large batchsizes in contrastive learning? a gradient-bias perspective,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9a0790f5-60b6-42ed-a15b-b1d6987ce91d · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition SONAR: sentence-level multimodal and language-agnostic represen- tations,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 67e43f76-6cad-4edd-8452-8aac1dac2249 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Audio retrieval with natural language queries,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 51a67ffe-e91a-40c1-aa01-4775cd0ba3ec · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f07396-8ef3-4204-81a1-bbfb223a65b0 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Recap: Retrieval-augmented audio captioning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 46ee1afd-8d3f-44e2-979d-f4bd80ab0a9f · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval Augmented End-to-End Spoken Dialog Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30801ac-8bc5-4c37-8c01-246e62884db1 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Speechdpr: End-to-end spoken passage retrieval for open-domain spoken question answering,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cc8d95bd-aaa4-4521-b796-b4ca13d099fb · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Retrieval augmented end-to-end spoken dialog models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 80b3516b-b725-4286-bbdd-2c265c5709ff · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Speechverse: A large-scale generalizable audio language model,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation aec0e83a-8cee-4213-8c15-47daf8c489c3 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Hubert: Self- supervised speech representation learning by masked prediction of hidden units,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 52cd75fa-1e3a-4fb5-bb42-2cebc847cdde · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Prompting large language models with audio for general-purpose speech summarization,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7ae6f0a0-52b8-47fd-a428-b97948288b39 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Sentence-bert: Sentence embeddings using siamese bert-networks,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 005b85e5-cd82-4d68-98c1-fa2b911e6bdb · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Improving text embeddings with large language models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1fed0f5e-61b2-4c79-93c5-9b65e5f0c3f6 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52ffbcbc-bb4c-47e2-9821-5f39c24c8c18 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Spoken squad: A study of mitigating the impact of speech recognition errors on listening comprehension,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d7e6ad55-28be-4e19-80e7-489c99aa0ab2 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3b0cdaea-edae-486a-b9a4-8340b3ec3e03 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition SQuAD: 100,000+ questions for machine comprehension of text,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 86d84a70-c3a6-4037-a220-6e57cb779aa8 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Introduction to the CoNLL-2003 shared task: Language-independent named entity recogni- tion,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3a5c6793-995a-4c02-9eee-4141dfd3d35a · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Evaluation of rag metrics for question answering in the telecom domain,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b959645b-2a80-4083-bac5-148f25e17a30 · outbound
Speech Retrieval-Augmented Generation without Automatic Speech Recognition Librispeech: An asr corpus based on public domain audio books,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b97890f7-17c2-4981-acd4-fe1cc4b5d397 · inbound
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering Speech Retrieval-Augmented Generation without Automatic Speech Recognition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.