Pith. sign in

REVIEW 6 cited by

Enhancing Multimodal Unified Representations for Cross Modal Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.05168 v3 pith:BZESA6EE submitted 2024-03-08 cs.CV cs.AI

classification cs.CVcs.AI
keywords representationsunifieddiscreteinformationmultimodalachievingcharacteristicsdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To enhance the interpretability of multimodal unified representations, many studies have focused on discrete unified representations. These efforts typically start with contrastive learning and gradually extend to the disentanglement of modal information, achieving solid multimodal discrete unified representations. However, existing research often overlooks two critical issues: 1) The use of Euclidean distance for quantization in discrete representations often overlooks the important distinctions among different dimensions of features, resulting in redundant representations after quantization; 2) Different modalities have unique characteristics, and a uniform alignment approach does not fully exploit these traits. To address these issues, we propose Training-free Optimization of Codebook (TOC) and Fine and Coarse cross-modal Information Disentangling (FCID). These methods refine the unified discrete representations from pretraining and perform fine- and coarse-grained information disentanglement tailored to the specific characteristics of each modality, achieving significant performance improvements over previous state-of-the-art models. The code is available at https://github.com/haihuangcode/CMG.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage prompt-tuning method with low-rank and contrastive prompt enhancement claims all-in-one adverse weather removal at 2.75M parameters.

  2. Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    By aligning category-level information across video, audio, and flow into one unified representation while keeping modality-specific details separate, this paper shows that standard domain generalization methods impro...

  3. IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A Gaussian-path transition equation lets a pretrained Stable Diffusion model serve as the denoiser inside image restoration bridges, cutting per-task training to a lightweight ControlNet.

  4. Offline Map Matching Based on Localization Error Distribution Modeling

    cs.SI 2025-05 conditional novelty 6.0 of 10

    LNSP models city-wide GPS error distributions from fixed bus routes and uses them, plus detour detection, to match sparse trajectories more accurately than three existing map matching methods.

  5. Open-set Cross Modal Generalization via Multimodal Unified Representation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The authors propose OSCMG, an open-set version of Cross Modal Generalization, and show their MICU method with masked contrastive learning and unified jigsaw puzzles outperforms prior methods.

  6. MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding

    cs.CV 2025-06 reject novelty 5.0 of 10

    A multimodal retrieval pipeline that projects video and audio into text and claims near-optimal context selection, with reported gains of up to 22.6% on Video-MME that rest on circular theory and unreleased data.

Pith tools