Pith. sign in

REVIEW 2 cited by

All About Knowledge Graphs for Actions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.12432 v1 pith:Z56SY7CY submitted 2020-08-28 cs.CV

classification cs.CV
keywords knowledgeactiondifferentembeddingsfew-shotgraphsrecognitionzero-shot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current action recognition systems require large amounts of training data for recognizing an action. Recent works have explored the paradigm of zero-shot and few-shot learning to learn classifiers for unseen categories or categories with few labels. Following similar paradigms in object recognition, these approaches utilize external sources of knowledge (eg. knowledge graphs from language domains). However, unlike objects, it is unclear what is the best knowledge representation for actions. In this paper, we intend to gain a better understanding of knowledge graphs (KGs) that can be utilized for zero-shot and few-shot action recognition. In particular, we study three different construction mechanisms for KGs: action embeddings, action-object embeddings, visual embeddings. We present extensive analysis of the impact of different KGs in different experimental setups. Finally, to enable a systematic study of zero-shot and few-shot approaches, we propose an improved evaluation paradigm based on UCF101, HMDB51, and Charades datasets for knowledge transfer from models trained on Kinetics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Kronecker Mask and Interpretive Prompts are Language-Action Video Learners

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CLAVER adds a cross-frame temporal attention mask (Kronecker mask) and LLM-generated interpretive action prompts to CLIP, improving video action recognition.

  2. Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task Perspectives

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Hier-EgoPack extends EgoPack's task-prototype transfer to multiple temporal granularities with a hierarchical GNN, improving Moment Queries and Long-Term Anticipation on Ego4D.

Pith tools