Pith. sign in

REVIEW 2 cited by

CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.03587 v1 pith:MQYNRGFB submitted 2022-02-08 eess.AS cs.SDeess.SP

classification eess.AScs.SDeess.SP
keywords representationsaudiomultimodalcontrastivecalmembeddinglexicalmultirate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deriving multimodal representations of audio and lexical inputs is a central problem in Natural Language Understanding (NLU). In this paper, we present Contrastive Aligned Audio-Language Multirate and Multimodal Representations (CALM), an approach for learning multimodal representations using contrastive and multirate information inherent in audio and lexical inputs. The proposed model aligns acoustic and lexical information in the input embedding space of a pretrained language-only contextual embedding model. By aligning audio representations to pretrained language representations and utilizing contrastive information between acoustic inputs, CALM is able to bootstrap audio embedding competitive with existing audio representation models in only a few hours of training time. Operationally, audio spectrograms are processed using linearized patches through a Spectral Transformer (SpecTran) which is trained using a Contrastive Audio-Language Pretraining objective to align audio and language from similar queries. Subsequently, the derived acoustic and lexical tokens representations are input into a multimodal transformer to incorporate utterance level context and derive the proposed CALM representations. We show that these pretrained embeddings can subsequently be used in multimodal supervised tasks and demonstrate the benefits of the proposed pretraining steps in terms of the alignment of the two embedding spaces and the multirate nature of the pretraining. Our system shows 10-25\% improvement over existing emotion recognition systems including state-of-the-art three-modality systems under various evaluation objectives.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Contrastive Learning for Task-Independent SpeechLLM-Pretraining

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Contrastive pre-training that aligns speech and text across all model layers beats ASR-based pre-training and, with 10% of task data, matches or exceeds specialized models on translation and question answering.

  2. Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.

Pith tools