REVIEW 4 cited by
Pengi: An Audio Language Model for Audio Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques. These approaches have led to the development of versatile models capable of tackling a wide array of tasks, while delivering state-of-the-art performance. However, current models inherently lack the capacity to produce the requisite language for open-ended tasks, such as Audio Captioning or Audio Question & Answering. We introduce Pengi, a novel Audio Language Model that leverages Transfer Learning by framing all audio tasks as text-generation tasks. It takes as input, an audio recording, and text, and generates free-form text as output. The input audio is represented as a sequence of continuous embeddings by an audio encoder. A text encoder does the same for the corresponding text input. Both sequences are combined as a prefix to prompt a pre-trained frozen language model. The unified architecture of Pengi enables open-ended tasks and close-ended tasks without any additional fine-tuning or task-specific extensions. When evaluated on 22 downstream tasks, our approach yields state-of-the-art performance in several of them. Our results show that connecting language models with audio models is a major step towards general-purpose audio understanding
Forward citations
Cited by 4 Pith papers
-
Adaptive Perturbation Selection for Contrastive Audio Decoding
A learned per-example router over a 105-perturbation audio library improves contrastive decoding for audio-LLM hallucination, with task-dependent best distortions (e.g., reverse audio for temporal order).
-
RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity
RA-QA converts 11 public respiratory-audio datasets into 9M template-generated QA pairs and shows current audio-language models score near zero on clinical task accuracy.
-
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
LiSTEN shows that dynamically selecting a few learnable prompt tokens from a shared pool can replace LoRA fine-tuning for audio-language models, matching or beating it with less training data.
-
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
The paper organizes negation phenomena into a four-category taxonomy for multimodal foundation models and lists open research questions, but provides no new experimental results.
Discussion (0). Continue with ORCID to comment.