Pith. sign in

REVIEW 3 cited by

Temporal Action Localization with Enhanced Instant Discriminability

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05590 v1 pith:77353ALW submitted 2023-09-11 cs.CV cs.AIcs.MM

classification cs.CVcs.AIcs.MM
keywords actionboundariesdiscriminabilityinstantproposeboundarycontextdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Temporal action detection (TAD) aims to detect all action boundaries and their corresponding categories in an untrimmed video. The unclear boundaries of actions in videos often result in imprecise predictions of action boundaries by existing methods. To resolve this issue, we propose a one-stage framework named TriDet. First, we propose a Trident-head to model the action boundary via an estimated relative probability distribution around the boundary. Then, we analyze the rank-loss problem (i.e. instant discriminability deterioration) in transformer-based methods and propose an efficient scalable-granularity perception (SGP) layer to mitigate this issue. To further push the limit of instant discriminability in the video backbone, we leverage the strong representation capability of pretrained large models and investigate their performance on TAD. Last, considering the adequate spatial-temporal context for classification, we design a decoupled feature pyramid network with separate feature pyramids to incorporate rich spatial context from the large model for localization. Experimental results demonstrate the robustness of TriDet and its state-of-the-art performance on multiple TAD datasets, including hierarchical (multilabel) TAD datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DisTime: Distribution-based Time Representation for Video Large Language Models

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A single learnable time token, decoded into a probability distribution over time bins, improves temporal grounding in Video-LLMs and is trained partly on a new 1.25M-event pseudo-labeled dataset.

  2. ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization

    cs.HC 2025-05 conditional novelty 6.0 of 10

    A drag-and-link interface lets users define rules for actions from body-part and object relations, generating frame labels to train temporal action localization models.

  3. DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DEL is a new audio-visual transformer framework that reports state-of-the-art temporal action localization on UnAV-100, THUMOS14, ActivityNet 1.3, and EPIC-Kitchens-100.

Pith tools