Pith. sign in

REVIEW 3 cited by

Micro-gesture Online Recognition using Learnable Query Points

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04490 v1 pith:WVH5WDMD submitted 2024-07-05 cs.CV

classification cs.CV
keywords micro-gestureonlinerecognitiontaskmicro-gesturessolutionstarttimes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we briefly introduce the solution developed by our team, HFUT-VUT, for the Micro-gesture Online Recognition track in the MiGA challenge at IJCAI 2024. The Micro-gesture Online Recognition task involves identifying the category and locating the start and end times of micro-gestures in video clips. Compared to the typical Temporal Action Detection task, the Micro-gesture Online Recognition task focuses more on distinguishing between micro-gestures and pinpointing the start and end times of actions. Our solution ranks 2nd in the Micro-gesture Online Recognition track.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

    cs.CV 2026-07 conditional novelty 3.0 of 10

    MAC 2026 reports a three-track micro-action challenge, adding a fine-grained MLLM-based understanding track evaluated on MA-Bench, with top-3 leaderboard results for each track.

  2. Online Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention

    cs.CV 2025-07 reject novelty 3.0 of 10

    The paper claims a first-place micro-gesture detection result from data augmentation and spatial-temporal attention, but its own table shows the winning F1 comes from the unmodified AdaTAD baseline, while the proposed...

  3. Exploiting Ensemble Learning for Cross-View Isolated Sign Language Recognition

    cs.CV 2025-02 conditional novelty 2.0 of 10

    An ensemble of three Video Swin Transformer sizes with RGB and depth fusion achieves 20.29% RGB and 24.53% RGB-D top-1 accuracy, ranking third in the CV-ISLR challenge.

Pith tools