Pith. sign in

REVIEW 1 cited by

MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.10589 v1 pith:LIRSIX3I submitted 2024-10-14 cs.CV

classification cs.CV
keywords temporalgeneralizationknowledgemotevideozero-shotclose-setexperts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transferring visual-language knowledge from large-scale foundation models for video recognition has proved to be effective. To bridge the domain gap, additional parametric modules are added to capture the temporal information. However, zero-shot generalization diminishes with the increase in the number of specialized parameters, making existing works a trade-off between zero-shot and close-set performance. In this paper, we present MoTE, a novel framework that enables generalization and specialization to be balanced in one unified model. Our approach tunes a mixture of temporal experts to learn multiple task views with various degrees of data fitting. To maximally preserve the knowledge of each expert, we propose \emph{Weight Merging Regularization}, which regularizes the merging process of experts in weight space. Additionally with temporal feature modulation to regularize the contribution of temporal feature during test. We achieve a sound balance between zero-shot and close-set video recognition tasks and obtain state-of-the-art or competitive results on various datasets, including Kinetics-400 \& 600, UCF, and HMDB. Code is available at \url{https://github.com/ZMHH-H/MoTE}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    CleanPose combines front-door causal adjustment with ULIP-2 knowledge distillation to improve category-level object pose estimation, reaching 61.7% on REAL275 5°2cm.

Pith tools