CoordTok encodes a 128-frame video into three 2D triplane latents (1280 tokens total) and reconstructs randomly sampled patch coordinates, enabling efficient long-video tokenization and 128-frame generation.
Quo vadis, action recognition? a new model and the kinetics dataset
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
CoordTok encodes a 128-frame video into three 2D triplane latents (1280 tokens total) and reconstructs randomly sampled patch coordinates, enabling efficient long-video tokenization and 128-frame generation.