Pith. sign in

REVIEW 6 cited by

Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.03638 v4 pith:TYLU7LGT submitted 2022-04-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords videoslongvideoframesgenerategeneratinginformationprogress
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Videos are created to express emotion, exchange information, and share experiences. Video synthesis has intrigued researchers for a long time. Despite the rapid progress driven by advances in visual synthesis, most existing studies focus on improving the frames' quality and the transitions between them, while little progress has been made in generating longer videos. In this paper, we present a method that builds on 3D-VQGAN and transformers to generate videos with thousands of frames. Our evaluation shows that our model trained on 16-frame video clips from standard benchmarks such as UCF-101, Sky Time-lapse, and Taichi-HD datasets can generate diverse, coherent, and high-quality long videos. We also showcase conditional extensions of our approach for generating meaningful long videos by incorporating temporal information with text and audio. Videos and code can be found at https://songweige.github.io/projects/tats/index.html.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Frozen video foundation features can be compressed into reconstruction-capable, generation-friendly latents that improve video generation quality and convergence speed.

  2. ELT: Elastic Looped Transformers for Visual Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    Weight-shared looped transformers trained with intra-loop self-distillation match MaskGIT-class FID/FVD at roughly 4x fewer parameters and support any-time inference across loop counts.

  3. Masked Generative Nested Transformers with Decode Time Scaling

    cs.CV 2025-02 conditional novelty 6.0 of 10

    MaGNeTS schedules progressively larger nested transformer sub-models over decode iterations and caches key-value pairs of unmasked tokens, achieving 2.5-3.7x compute reduction with competitive FID/FVD.

  4. Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A temporal attention purification and skip-connection rerouting method that separates motion learning from appearance learning in text-to-video diffusion model customization.

  5. Learnings from Scaling Visual Tokenizers for Reconstruction and Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A ViT-based visual tokenizer study shows latent code size drives reconstruction quality, encoder scaling gives little benefit for generation, and ViTok reaches competitive or state-of-the-art results with fewer FLOPs.

  6. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools