REVIEW 5 cited by
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system with a speech encoder, which is trained to produce token embeddings such that the LLM's response to an expressive speech prompt is aligned with its response to a semantically matching text prompt that has also been conditioned on the user's speaking style. This framework enables the encoder to generate tokens that capture both linguistic and paralinguistic information and effectively convey them to the LLM, even when the LLM's weights remain completely frozen. To the best of our knowledge, our work is the first to explore how to induce a frozen LLM to understand more than just linguistic content from speech inputs in a general interaction setting. Experiments demonstrate that our system is able to produce higher quality and more empathetic responses to expressive speech prompts compared to several baselines.
Forward citations
Cited by 5 Pith papers
-
Dual Information Speech Language Models for Emotional Conversations
A dual-adapter design with equivalence replacement regularization lets frozen LLMs perceive both paralinguistic and linguistic information from speech for emotional conversation.
-
Towards Reliable Large Audio Language Model
Training a large audio language model to say 'I don't know' on one audio type (speech, music, or sound) makes it more likely to refuse uncertain questions on the other types.
-
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
Training a speech-LLM on question-answer pairs generated with both discrete and continuous emotion labels improves its contextual emotion reasoning as scored by an LLM judge.
-
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs
A frozen LLM with a small EMG adaptor converts unvoiced EMG to text at 0.49 average word error rate on a 67-word closed vocabulary without any voiced audio.
-
MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages
A LoRA-tuned speech encoder with an ECAPA-TDNN head outperforms prior gender-recognition systems on most English and Southeast Asian test sets.
Discussion (0). Continue with ORCID to comment.