REVIEW 4 cited by
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable. Specifically, CLAP-style models are able to `answer' a diverse set of language queries, extending the capabilities of audio models beyond a closed set of labels. However, CLAP relies on a large set of (audio, query) pairs for pretraining. While such sets are available for general audio tasks, like captioning or sound event detection, there are no datasets with matched audio and text queries for computational paralinguistic (CP) tasks. As a result, the community relies on generic CLAP models trained for general audio with limited success. In the present study, we explore training considerations for ParaCLAP, a CLAP-style model suited to CP, including a novel process for creating audio-language queries. We demonstrate its effectiveness on a set of computational paralinguistic tasks, where it is shown to surpass the performance of open-source state-of-the-art models.
Forward citations
Cited by 4 Pith papers
-
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
Dual-encoder speech-text models trained on rich intrinsic and situational style captions outperform prior CLAP-style baselines on retrieval, classification, and inference-time TTS style guidance.
-
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
Authors propose ESS-CLAP and RA-CLAP, contrastive speech-text models for emotional speaking style retrieval, evaluated on PromptSpeech, TextrolSpeech, and SpeechCraft.
-
AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
AutoSIFT disentangles text-describable style categories from residual speech styles and selectively infills only the categories the user specifies.
-
CLEP-DG: Contrastive Learning for Speech Emotion Domain Generalization via Soft Prompt Tuning
CLEP-DG fine-tunes CLAP on emotional speech and augments it with acoustic-context prompt tuning, reporting improved speech emotion recognition and domain generalization.
Discussion (0). Sign in to comment.