Pith. sign in

REVIEW 4 cited by

Language Knowledge-Assisted Representation Learning for Skeleton-Based Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.12398 v1 pith:BGF5LUMI submitted 2023-05-21 cs.CV

classification cs.CV
keywords knowledgeactionbrainhumansinformationnodeprioritopology
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

How humans understand and recognize the actions of others is a complex neuroscientific problem that involves a combination of cognitive mechanisms and neural networks. Research has shown that humans have brain areas that recognize actions that process top-down attentional information, such as the temporoparietal association area. Also, humans have brain regions dedicated to understanding the minds of others and analyzing their intentions, such as the medial prefrontal cortex of the temporal lobe. Skeleton-based action recognition creates mappings for the complex connections between the human skeleton movement patterns and behaviors. Although existing studies encoded meaningful node relationships and synthesized action representations for classification with good results, few of them considered incorporating a priori knowledge to aid potential representation learning for better performance. LA-GCN proposes a graph convolution network using large-scale language models (LLM) knowledge assistance. First, the LLM knowledge is mapped into a priori global relationship (GPR) topology and a priori category relationship (CPR) topology between nodes. The GPR guides the generation of new "bone" representations, aiming to emphasize essential node information from the data level. The CPR mapping simulates category prior knowledge in human brain regions, encoded by the PC-AC module and used to add additional supervision-forcing the model to learn class-distinguishable features. In addition, to improve information transfer efficiency in topology modeling, we propose multi-hop attention graph convolution. It aggregates each node's k-order neighbor simultaneously to speed up model convergence. LA-GCN reaches state-of-the-art on NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Topological Symmetry Enhanced Graph Convolution for Skeleton-Based Action Recognition

    cs.CV 2024-11 conditional novelty 5.0 of 10

    TSE-GCN uses symmetry-aware graph reactivation and per-frame deformable temporal convolution to reach 90.0 and 91.1 percent on NTU RGB+D 120 with 4.4 million parameters.

  2. Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Spatio-temporal foundation models are organized into a pipeline of data harmonization, model design, training, and adaptation, with a data property taxonomy for model selection.

  3. SMART-Vision: Survey of Modern Action Recognition Techniques in Vision

    cs.CV 2025-01 conditional novelty 4.0 of 10

    The SMART-Vision survey organizes vision-based human action recognition into a hybrid Venn-diagram taxonomy and reviews the emerging open-set/open-world HAR literature.

  4. 3D Skeleton-Based Action Recognition: A Review

    cs.CV 2025-06 reject novelty 3.0 of 10

    A task-oriented review of skeleton-based action recognition that reorganizes known methods along a data processing pipeline and contains no new experimental result.

Pith tools