REVIEW 2 cited by
Open-Vocabulary Temporal Action Localization using Multimodal Guidance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Open-Vocabulary Temporal Action Localization (OVTAL) enables a model to recognize any desired action category in videos without the need to explicitly curate training data for all categories. However, this flexibility poses significant challenges, as the model must recognize not only the action categories seen during training but also novel categories specified at inference. Unlike standard temporal action localization, where training and test categories are predetermined, OVTAL requires understanding contextual cues that reveal the semantics of novel categories. To address these challenges, we introduce OVFormer, a novel open-vocabulary framework extending ActionFormer with three key contributions. First, we employ task-specific prompts as input to a large language model to obtain rich class-specific descriptions for action categories. Second, we introduce a cross-attention mechanism to learn the alignment between class representations and frame-level video features, facilitating the multimodal guided features. Third, we propose a two-stage training strategy which includes training with a larger vocabulary dataset and finetuning to downstream data to generalize to novel categories. OVFormer extends existing TAL methods to open-vocabulary settings. Comprehensive evaluations on the THUMOS14 and ActivityNet-1.3 benchmarks demonstrate the effectiveness of our method. Code and pretrained models will be publicly released.
Forward citations
Cited by 2 Pith papers
-
MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization
A multi-grained network that recognizes seen actions with a supervised classifier and unseen actions via video-level coarse filtering plus proposal-level matching achieves state-of-the-art open-vocabulary temporal act...
-
GASP: A Gradient-Aware Shortest Path Algorithm for Boundary-Confined Visualization of 2-Manifold Reeb Graphs
GASP is a new algorithm that draws Reeb graphs of 2-manifold scalar fields so they hug the shape's boundary, stay compact, and align with the function's gradient, beating TTK's barycenter layout in evaluation.
Discussion (0). Continue with ORCID to comment.