Pith. sign in

REVIEW 3 cited by

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09291 v2 pith:UHXQFTPG submitted 2025-01-16 cs.MM cs.AIcs.SDeess.AS

classification cs.MMcs.AIcs.SDeess.AS
keywords audiocaptioningoptimallavcapvisualaudio-visualtransporteffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse audio and visual data, missing important semantic cues from each modality. To address this, we introduce LAVCap, a large language model (LLM)-based audio-visual captioning framework that effectively integrates visual information with audio to improve audio captioning performance. LAVCap employs an optimal transport-based alignment loss to bridge the modality gap between audio and visual features, enabling more effective semantic extraction. Additionally, we propose an optimal transport attention module that enhances audio-visual fusion using an optimal transport assignment map. Combined with the optimal training strategy, experimental results demonstrate that each component of our framework is effective. LAVCap outperforms existing state-of-the-art methods on the AudioCaps dataset, without relying on large datasets or post-processing. Code is available at https://github.com/NAVER-INTEL-Co-Lab/gaudi-lavcap.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

    cs.SD 2026-07 conditional novelty 5.0 of 10

    OT couplings that map Whisper and AV-HuBERT features onto LLaMA token embeddings, used as soft contrastive labels, yield SOTA LRS3-TED AVSR under clean and noisy SNRs.

  2. Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning

    cs.MM 2025-05 conditional novelty 5.0 of 10

    Entropy-aware gating and shuffled audio-video training pairs improve robustness to audiovisual mismatch in video-guided audio captioning.

  3. WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

    cs.SD 2025-06 conditional novelty 4.0 of 10

    WhisQ uses Whisper and Qwen with co-attention and optimal transport to predict music quality and text-alignment scores, but its reported improvements do not match its own data.

Pith tools