Pith. sign in

REVIEW 3 cited by

LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.10390 v4 pith:JZ6RVAZS submitted 2023-08-20 cs.CL

classification cs.CL
keywords llmslibrisqadatasetframeworkmultimodalansweringexistinghandling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While Large Language Models (LLMs) have demonstrated commendable performance across a myriad of domains and tasks, existing LLMs still exhibit a palpable deficit in handling multimodal functionalities, especially for the Spoken Question Answering (SQA) task which necessitates precise alignment and deep interaction between speech and text features. To address the SQA challenge on LLMs, we initially curated the free-form and open-ended LibriSQA dataset from Librispeech, comprising Part I with natural conversational formats and Part II encompassing multiple-choice questions followed by answers and analytical segments. Both parts collectively include 107k SQA pairs that cover various topics. Given the evident paucity of existing speech-text LLMs, we propose a lightweight, end-to-end framework to execute the SQA task on the LibriSQA, witnessing significant results. By reforming ASR into the SQA format, we further substantiate our framework's capability in handling ASR tasks. Our empirical findings bolster the LLMs' aptitude for aligning and comprehending multimodal information, paving the way for the development of universal multimodal LLMs. The dataset and demo can be found at https://github.com/ZihanZhaoSJTU/LibriSQA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

    cs.MM 2026-08 conditional novelty 5.0 of 10

    An audio agent trained with trajectory-based SFT and multi-turn GRPO improves tool-use and reasoning on a new AI-generated audio agent benchmark, including tasks with unseen tools and workflows.

  2. LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs

    cs.AI 2025-05 conditional novelty 5.0 of 10

    LiSTEN shows that dynamically selecting a few learnable prompt tokens from a shared pool can replace LoRA fine-tuning for audio-language models, matching or beating it with less training data.

  3. Speechless: Speech Instruction Training Without Speech for Low Resource Languages

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Fine-tuning an LLM on text instructions converted to Whisper semantic tokens enables it to understand spoken instructions at inference, bypassing TTS and speech instruction data.

Pith tools