Pith. sign in

REVIEW 5 cited by

Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.01162 v2 pith:66KU3SRK submitted 2024-10-02 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords speechfrozenparalinguisticaspectsencoderexpressivelanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights. We utilize an end-to-end system with a speech encoder, which is trained to produce token embeddings such that the LLM's response to an expressive speech prompt is aligned with its response to a semantically matching text prompt that has also been conditioned on the user's speaking style. This framework enables the encoder to generate tokens that capture both linguistic and paralinguistic information and effectively convey them to the LLM, even when the LLM's weights remain completely frozen. To the best of our knowledge, our work is the first to explore how to induce a frozen LLM to understand more than just linguistic content from speech inputs in a general interaction setting. Experiments demonstrate that our system is able to produce higher quality and more empathetic responses to expressive speech prompts compared to several baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual Information Speech Language Models for Emotional Conversations

    cs.CL 2025-08 conditional novelty 7.0 of 10

    A dual-adapter design with equivalence replacement regularization lets frozen LLMs perceive both paralinguistic and linguistic information from speech for emotional conversation.

  2. Towards Reliable Large Audio Language Model

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Training a large audio language model to say 'I don't know' on one audio type (speech, music, or sound) makes it more likely to refuse uncertain questions on the other types.

  3. Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models

    cs.CL 2025-08 conditional novelty 5.0 of 10

    Training a speech-LLM on question-answer pairs generated with both discrete and continuous emotion labels improves its contextual emotion reasoning as scored by an LLM judge.

  4. Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A frozen LLM with a small EMG adaptor converts unvoiced EMG to text at 0.49 average word error rate on a 67-word closed vocabulary without any voiced audio.

  5. MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

    cs.CL 2026-08 conditional novelty 4.0 of 10

    A LoRA-tuned speech encoder with an ECAPA-TDNN head outperforms prior gender-recognition systems on most English and Southeast Asian test sets.

Pith tools