REVIEW 18 cited by
A Short Note on the Kinetics-700 Human Action Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We describe an extension of the DeepMind Kinetics human action dataset from 600 classes to 700 classes, where for each class there are at least 600 video clips from different YouTube videos. This paper details the changes introduced for this new release of the dataset, and includes a comprehensive set of statistics as well as baseline results using the I3D neural network architecture.
Forward citations
Cited by 18 Pith papers
-
The TIME Machine: On The Power of Motion for Efficient Perception
TIME is a motion-based embedding from point tracks, trained only on synthetic data via masked autoencoding, that matches state-of-the-art video model performance with up to 10,000x less training data.
-
Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study
Identity-specific free-throw motion signatures exist and are learnable by video models, but models prefer static appearance shortcuts unless silhouettes or skeletons suppress them.
-
Self-Supervised Learning of Structured Dynamics from Videos
A two-token 'primary/residual' future-feature predictor separates camera from object motion, outpacing frozen-feature baselines and matching larger supervised models on several probes.
-
Cambrian-P: Pose-Grounded Video Understanding
Cambrian-P adds per-frame camera pose tokens and a regression head to video MLLMs, delivering 4.5-6.5% gains on spatial benchmarks, generalization to other video QA tasks, and SOTA streaming pose estimation on ScanNet.
-
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
A learned human-action manifold, combining 3D pose, 2D keypoints, appearance, and motion derivatives, scores generated videos by distance to real-action centroids and embedding smoothness, beating prior metrics on hum...
-
A Comprehensive Survey on World Models for Embodied AI
A unified three-axis taxonomy — functionality, temporal modeling, spatial representation — organizes the world-model literature for embodied AI.
-
No More Sibling Rivalry: Debiasing Human-Object Interaction Detection
A detection transformer for human-object interactions gains 9.18 mAP on HICO-DET by adding contrastive-then-calibration and merge-then-split training objectives against a diagnosed 'toxic siblings' interference bias.
-
Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
Category-level retrieval improves when text queries are converted into multiple generated images, aggregated with a learned attention module, and fused with CLIP text similarity.
-
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
A person-anchored tree plus multi-agent LLM pipeline lets a system answer cross-video queries about the same person, and it beats single-video models on the authors' new CrossVideoQA benchmark.
-
SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities
SPACT18 is claimed to be the first action recognition dataset captured with a spike camera, paired with synchronized RGB and thermal video.
-
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
LG-CAV-MAE trains a tri-modal audio-visual-text model on CLAP-filtered auto-generated captions and beats existing audio-visual masked autoencoders on retrieval and classification.
-
Simplifying Traffic Anomaly Detection with Video Foundation Models
An encoder-only Video ViT with self-supervised masked video pretraining matches or beats specialized traffic anomaly detectors and is more efficient.
-
Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics
Using an inverse dynamics model to label and verify frame transitions lets vision-language models learn forward dynamics and outperform specialized image editors on Aurora-Bench.
-
Large-scale Self-supervised Video Foundation Model for Intelligent Surgery
SurgVISTA is a masked-reconstruction surgical video foundation model whose joint spatiotemporal pretraining plus expert distillation outperforms image-level and natural-video pretrained models on 13 surgical benchmarks.
-
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
Adding ground-truth object localization, tracking, and relation annotations as context improves VideoLLM answers, but the models never generate these steps themselves in the experiments.
-
Grounding Intelligence in Movement
Movement should be treated as a first-class AI modeling modality, and a unified, biomechanically grounded movement foundation model built from aggregated data across species and sensors is the proposed path forward.
-
DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition
DVFL-Net, a 22M-parameter distilled video model, reaches near-teacher accuracy on five action recognition benchmarks with 27 GFLOPs, but its claimed state-of-the-art status is not fully supported by the reported numbers.
- OpenMAP-BrainAge: Generalizable and Interpretable Brain Age Predictor
Discussion (0). Continue with ORCID to comment.