REVIEW 3 cited by
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse audio and visual data, missing important semantic cues from each modality. To address this, we introduce LAVCap, a large language model (LLM)-based audio-visual captioning framework that effectively integrates visual information with audio to improve audio captioning performance. LAVCap employs an optimal transport-based alignment loss to bridge the modality gap between audio and visual features, enabling more effective semantic extraction. Additionally, we propose an optimal transport attention module that enhances audio-visual fusion using an optimal transport assignment map. Combined with the optimal training strategy, experimental results demonstrate that each component of our framework is effective. LAVCap outperforms existing state-of-the-art methods on the AudioCaps dataset, without relying on large datasets or post-processing. Code is available at https://github.com/NAVER-INTEL-Co-Lab/gaudi-lavcap.
Forward citations
Cited by 3 Pith papers
-
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
OT couplings that map Whisper and AV-HuBERT features onto LLaMA token embeddings, used as soft contrastive labels, yield SOTA LRS3-TED AVSR under clean and noisy SNRs.
-
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Entropy-aware gating and shuffled audio-video training pairs improve robustness to audiovisual mismatch in video-guided audio captioning.
-
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
WhisQ uses Whisper and Qwen with co-attention and optimal transport to predict music quality and text-alignment scores, but its reported improvements do not match its own data.
Discussion (0). Sign in to comment.