REVIEW 3 cited by
ReconBoost: Boosting Can Achieve Modality Reconcilement
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper explores a novel multi-modal alternating learning paradigm pursuing a reconciliation between the exploitation of uni-modal features and the exploration of cross-modal interactions. This is motivated by the fact that current paradigms of multi-modal learning tend to explore multi-modal features simultaneously. The resulting gradient prohibits further exploitation of the features in the weak modality, leading to modality competition, where the dominant modality overpowers the learning process. To address this issue, we study the modality-alternating learning paradigm to achieve reconcilement. Specifically, we propose a new method called ReconBoost to update a fixed modality each time. Herein, the learning objective is dynamically adjusted with a reconcilement regularization against competition with the historical models. By choosing a KL-based reconcilement, we show that the proposed method resembles Friedman's Gradient-Boosting (GB) algorithm, where the updated learner can correct errors made by others and help enhance the overall performance. The major difference with the classic GB is that we only preserve the newest model for each modality to avoid overfitting caused by ensembling strong learners. Furthermore, we propose a memory consolidation scheme and a global rectification scheme to make this strategy more effective. Experiments over six multi-modal benchmarks speak to the efficacy of the method. We release the code at https://github.com/huacong/ReconBoost.
Forward citations
Cited by 3 Pith papers
-
Boosting Multimodal Learning via Disentangled Gradient Learning
Disentangled gradient learning replaces the multimodal gradient to each encoder with a unimodal gradient computed via modality dropout, improving both unimodal and multimodal accuracy across several tasks.
-
Improving Multimodal Learning via Imbalanced Learning
Asymmetric Representation Learning reweights each modality's gradient by the inverse of its prediction variance, improving multimodal accuracy on CREMA-D, Kinetics-Sounds, AVE, MOSI, and UCF101.
-
RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer
RollingQ rotates the classification query in a multimodal Transformer toward a rebalanced direction so attention stops over-favoring a single modality, restoring dynamic fusion and improving accuracy.
Discussion (0). Continue with ORCID to comment.