Pith. sign in

REVIEW 18 cited by

A Short Note on the Kinetics-700 Human Action Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.06987 v2 pith:GLLDPC5K submitted 2019-07-15 cs.CV

classification cs.CV
keywords datasetactionclasseshumanarchitecturebaselinechangesclass
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We describe an extension of the DeepMind Kinetics human action dataset from 600 classes to 700 classes, where for each class there are at least 600 video clips from different YouTube videos. This paper details the changes introduced for this new release of the dataset, and includes a comprehensive set of statistics as well as baseline results using the I3D neural network architecture.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The TIME Machine: On The Power of Motion for Efficient Perception

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    TIME is a motion-based embedding from point tracks, trained only on synthetic data via masked autoencoding, that matches state-of-the-art video model performance with up to 10,000x less training data.

  2. Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Identity-specific free-throw motion signatures exist and are learnable by video models, but models prefer static appearance shortcuts unless silhouettes or skeletons suppress them.

  3. Self-Supervised Learning of Structured Dynamics from Videos

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A two-token 'primary/residual' future-feature predictor separates camera from object motion, outpacing frozen-feature baselines and matching larger supervised models on several probes.

  4. Cambrian-P: Pose-Grounded Video Understanding

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Cambrian-P adds per-frame camera pose tokens and a regression head to video MLLMs, delivering 4.5-6.5% gains on spatial benchmarks, generalization to other video QA tasks, and SOTA streaming pose estimation on ScanNet.

  5. Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A learned human-action manifold, combining 3D pose, 2D keypoints, appearance, and motion derivatives, scores generated videos by distance to real-action centroids and embedding smoothness, beating prior metrics on hum...

  6. A Comprehensive Survey on World Models for Embodied AI

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.

  7. No More Sibling Rivalry: Debiasing Human-Object Interaction Detection

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A detection transformer for human-object interactions gains 9.18 mAP on HICO-DET by adding contrastive-then-calibration and merge-then-split training objectives against a diagnosed 'toxic siblings' interference bias.

  8. Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Category-level retrieval improves when text queries are converted into multiple generated images, aggregated with a learned attention module, and fused with CLIP text similarity.

  9. VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A person-anchored tree plus multi-agent LLM pipeline lets a system answer cross-video queries about the same person, and it beats single-video models on the authors' new CrossVideoQA benchmark.

  10. SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SPACT18 is claimed to be the first action recognition dataset captured with a spike camera, paired with synchronized RGB and thermal video.

  11. Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LG-CAV-MAE trains a tri-modal audio-visual-text model on CLAP-filtered auto-generated captions and beats existing audio-visual masked autoencoders on retrieval and classification.

  12. Simplifying Traffic Anomaly Detection with Video Foundation Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An encoder-only Video ViT with self-supervised masked video pretraining matches or beats specialized traffic anomaly detectors and is more efficient.

  13. Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using an inverse dynamics model to label and verify frame transitions lets vision-language models learn forward dynamics and outperform specialized image editors on Aurora-Bench.

  14. Large-scale Self-supervised Video Foundation Model for Intelligent Surgery

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SurgVISTA is a masked-reconstruction surgical video foundation model whose joint spatiotemporal pretraining plus expert distillation outperforms image-level and natural-video pretrained models on 13 surgical benchmarks.

  15. CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks

    cs.CV 2025-07 reject novelty 4.0 of 10

    Adding ground-truth object localization, tracking, and relation annotations as context improves VideoLLM answers, but the models never generate these steps themselves in the experiments.

  16. Grounding Intelligence in Movement

    cs.AI 2025-07 conditional novelty 4.0 of 10

    Movement should be treated as a first-class AI modeling modality, and a unified, biomechanically grounded movement foundation model built from aggregated data across species and sensors is the proposed path forward.

  17. DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition

    cs.CV 2025-07 conditional novelty 3.0 of 10

    DVFL-Net, a 22M-parameter distilled video model, reaches near-teacher accuracy on five action recognition benchmarks with 27 GFLOPs, but its claimed state-of-the-art status is not fully supported by the reported numbers.

  18. OpenMAP-BrainAge: Generalizable and Interpretable Brain Age Predictor

    cs.CV 2025-06

Pith tools