REVIEW 3 cited by
HeySQuAD: A Spoken Question Answering Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
HeySQuAD: A Spoken Question Answering Dataset
read the original abstract
Spoken question answering (SQA) systems are critical for digital assistants and other real-world use cases, but evaluating their performance is a challenge due to the importance of human-spoken questions. This study presents a new large-scale community-shared SQA dataset called HeySQuAD, which includes 76k human-spoken questions, 97k machine-generated questions, and their corresponding textual answers from the SQuAD QA dataset. Our goal is to measure the ability of machines to accurately understand noisy spoken questions and provide reliable answers. Through extensive testing, we demonstrate that training with transcribed human-spoken and original SQuAD questions leads to a significant improvement (12.51%) in answering human-spoken questions compared to training with only the original SQuAD textual questions. Moreover, evaluating with a higher-quality transcription can lead to a further improvement of 2.03%. This research has significant implications for the development of SQA systems and their ability to meet the needs of users in real-world scenarios.
Forward citations
Cited by 3 Pith papers
-
ATIR: Towards Audio-Text Interleaved Contextual Retrieval
Defines ATIR task and benchmark for mixed audio-text queries; MLLM model with token compression shows substantial gains over strong baselines.
-
LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering
Multi-source and voice-design TTS training data yield the strongest Luxembourgish spoken QA on real speakers, while no-reference MOS scores fail to rank systems by QA utility.
-
AuRA: Internalizing Audio Understanding into LLMs as LoRA
AuRA uses LoRA and layer-wise distillation from an ASR teacher to internalize audio encoding into LLMs for improved speech-language performance.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.