Pith. sign in

REVIEW 12 cited by

Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02757 v2 pith:QA55MJPS submitted 2024-10-03 cs.CV

classification cs.CV
keywords videosautoregressivelongvideogeneratinggeneratelanguagellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

It is desirable but challenging to generate content-rich long videos in the scale of minutes. Autoregressive large language models (LLMs) have achieved great success in generating coherent and long sequences of tokens in the domain of natural language processing, while the exploration of autoregressive LLMs for video generation is limited to generating short videos of several seconds. In this work, we conduct a deep analysis of the challenges that prevent autoregressive LLM-based video generators from generating long videos. Based on the observations and analysis, we propose Loong, a new autoregressive LLM-based video generator that can generate minute-long videos. Specifically, we model the text tokens and video tokens as a unified sequence for autoregressive LLMs and train the model from scratch. We propose progressive short-to-long training with a loss re-weighting scheme to mitigate the loss imbalance problem for long video training. We further investigate inference strategies, including video token re-encoding and sampling strategies, to diminish error accumulation during inference. Our proposed Loong can be trained on 10-second videos and be extended to generate minute-level long videos conditioned on text prompts, as demonstrated by the results. More samples are available at: https://yuqingwang1029.github.io/Loong-video.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A single video diffusion model, trained with flexible temporal–timestep chunking and noise-aligned keys, supports bidirectional, autoregressive, and hybrid inference with better quality–efficiency than Self-Forcing.

  2. TokensGen: Harnessing Condensed Tokens for Long Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    TokensGen generates consistent long videos by representing each clip as condensed semantic tokens, generating all tokens jointly from text, and stitching clips with adaptive FIFO denoising.

  3. VideoMAR: Autoregressive Video Generatio with Continuous Tokens

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A decoder-only autoregressive video model with continuous tokens, frame-wise causal attention, and a next-frame diffusion loss reports a higher VBench-I2V score than Cosmos I2V with a much smaller model and dataset.

  4. UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

    cs.CV 2025-06 conditional novelty 6.0 of 10

    UltraVideo provides the first UHD 4K/8K text-to-video dataset with 10 structured captions per clip, and LoRA fine-tuning on it produces native 1K/4K video generation with improved visual quality.

  5. SpectralAR: Spectral Autoregressive Visual Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive image generator that tokenizes images in the DCT frequency domain into nested 1D spectral sequences and generates them coarse-to-fine, reaching 3.02 gFID with 64 tokens on ImageNet-1K.

  6. Video World Models with Long-term Spatial Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive video world model with a persistent static point-cloud spatial memory and sparse episodic keyframes improves revisit consistency over point-cloud-conditioned baselines.

  7. Long-Context State-Space Video World Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A hybrid state-space and local-attention architecture gives autoregressive video diffusion models long-term spatial memory with constant per-frame inference cost, demonstrated on Maze and Minecraft.

  8. BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    BlockVid generates minute-long videos with a semantic sparse KV cache, Block Forcing training, and chunk-level noise scheduling, reporting large gains on its own LV-Bench and on VBench.

  9. HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    HumanGenesis is a claimed state-of-the-art framework that couples 3D Gaussian reconstruction, LLM-based critique, pose guidance, and diffusion-based harmonization for synthetic human video generation; unverified in th...

  10. Enhancing Scene Transition Awareness in Video Generation via Post-Training

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A post-training dataset of transition-centered clips increases the number of scenes an open-source video generator produces for multi-scene prompts, with mixed effects on quality.

  11. Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Hi-MAR adds a low-resolution token prediction phase and a diffusion transformer head to masked autoregressive image generation, improving FID and cutting autoregressive steps.

  12. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.

Pith tools