Pith. sign in

REVIEW 31 cited by

Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.04233 v1 pith:QDBOGXFK submitted 2024-05-07 cs.CV cs.LG

classification cs.CVcs.LG
keywords generationvidugeneratortext-to-videovideoscapablediffusionvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Vidu, a high-performance text-to-video generator that is capable of producing 1080p videos up to 16 seconds in a single generation. Vidu is a diffusion model with U-ViT as its backbone, which unlocks the scalability and the capability for handling long videos. Vidu exhibits strong coherence and dynamism, and is capable of generating both realistic and imaginative videos, as well as understanding some professional photography techniques, on par with Sora -- the most powerful reported text-to-video generator. Finally, we perform initial experiments on other controllable video generation, including canny-to-video generation, video prediction and subject-driven generation, which demonstrate promising results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A film-academy cinematic taxonomy and reverse-engineered multi-shot prompts expose large gaps in leading video generators that web-style benchmarks miss.

  2. GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    GroundShot introduces entity-grounded shot scheduling with online visual memory to improve consistency in multi-shot video generation and presents GroundBench for entity-level evaluation.

  3. BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    BiWM is the first full-stack open-source bidirectional autoregressive framework for interactive video world models, reducing training stages to two while adding camera control and efficiency features across several backbones.

  4. Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    Causal Forcing uses an autoregressive teacher for ODE initialization in diffusion distillation to close the causal attention gap and deliver better real-time video generation than Self Forcing.

  5. WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Across 1,474 cases and 20 models, WorldExam shows that video world models split along paradigm lines — camera-, action-, and language-driven models each dominate one capability, and none combines strong reactivity wit...

  6. MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    MultiRef-Compass is a 350-sample benchmark and 14-metric protocol for multi-reference-to-audio-video generation; current models still fail at reference binding and audio-visual consistency.

  7. ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    ARGen uses AU-guided prompts and a reinforcement-learned diffusion strategy to synthesize scarce-class facial expression videos that improve dynamic emotion recognition.

  8. Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SideInfo-RoPE encodes reference-identity agreement as an extra rotary axis, disambiguating similar characters in multi-reference multi-shot video generation while keeping full semantic attention.

  9. RefAlign: Representation Alignment for Reference-to-Video Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Explicit training-time alignment of DiT reference features to a VFM (with pull/push loss) raises OpenS2V-Eval TotalScore over prior R2V methods with no inference cost.

  10. FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FreeLong++ extends short-video diffusion models to 4x to 8x longer clips, without retraining, by fusing multiple windowed attention branches through frequency-domain filters and a spectral noise initialization.

  11. FastInit: Fast Noise Initialization for Temporally Consistent Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A single-pass learned noise predictor, trained to imitate FreeInit's outputs, gives temporally more consistent text-to-video generation at near-zero added inference cost.

  12. Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Context-as-Memory conditions video generation on selected historical frames chosen by camera FOV overlap, improving scene consistency in long generated videos.

  13. VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VCapsBench is a video caption quality benchmark with 109,796 QA pairs across 21 fine-grained dimensions on 5,677 videos, evaluating caption accuracy, inconsistency, and coverage.

  14. Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer

    cs.LG 2025-05 conditional novelty 6.0 of 10

    One diffusion transformer, trained across PPG, ECG, and blood pressure signals, handles denoising, imputation, and cross-modal synthesis in a single framework.

  15. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  16. MotionPro: A Precise Motion Controller for Image-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MotionPro uses region-wise trajectories and a motion mask to control object and camera motion in image-to-video generation, reporting improved trajectory alignment over prior methods.

  17. AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance

    cs.CV 2025-02 conditional novelty 6.0 of 10

    AnyCharV is a two-stage diffusion method that places a reference character into a target video scene using pose and mask guidance, with a self-boosting stage that improves identity preservation.

  18. Seeing World Dynamics in a Nutshell

    cs.CV 2025-02 conditional novelty 6.0 of 10

    NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.

  19. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  20. Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation

    cs.AI 2026-02 conditional novelty 5.0 of 10

    LASEV, a multi-agent LLM system that compiles structured 'executable video scripts' into educational videos, reports 92-96% expert-rated publishable quality and more than one million videos per day at 95% lower cost.

  21. The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

    cs.CV 2026-01 reject novelty 5.0 of 10

    An agentic dialogue-to-video pipeline (ScripterAgent, DirectorAgent, CriticAgent) claims to improve long-horizon cinematic coherence, but its supporting evaluation is partly self-referential.

  22. RewardDance: Reward Scaling in Visual Generation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    RewardDance reframes visual reward modeling as a yes/no judgment task in a VLM and reports consistent gains in text-to-image, text-to-video, and image-to-video generation as the reward model scales from 1B to 26B.

  23. Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A training-free prompt, image, and guidance enhancement framework improves face consistency and video quality for identity-preserving text-to-video generation, winning the ACM Multimedia 2025 IPVG challenge.

  24. InfinityHuman: Towards Long-Term Audio-Driven Human

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A coarse-to-fine audio-driven animation framework that uses pose-guided refinement and hand-specific reward learning to generate long, identity-stable talking videos.

  25. StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    StableAnimator++ combines learnable SVD-guided pose alignment, a distribution-aware ID Adapter, and an HJB-based inference-time face optimizer to preserve identity in human image animation under severe pose misalignment.

  26. Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Tora2 adds decoupled personalization embeddings, gated self-attention binding, and contrastive learning to Tora, enabling simultaneous appearance and trajectory customization for multiple entities in generated video.

  27. LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

    cs.CV 2025-06 reject novelty 5.0 of 10

    A hierarchical coarse-to-fine diffusion transformer with cross-granularity distillation improves long-term driving video prediction, but the reported gains may be inflated by future-derived text prompts and a selected...

  28. Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SCST reports the best perceptual quality (LPIPS/DISTS) on four synthetic benchmarks and the best no-reference quality scores on the real-world VideoLQ benchmark by adding spatio-temporal Mamba and contrastive ControlN...

  29. RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Metric-scale depth alignment plus scene-constrained noise shaping improves camera controllability and video quality for image-to-video generation on RealEstate10K.

  30. Goku: Flow Based Video Generative Foundation Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A joint image-video generation model family reports state-of-the-art benchmark scores using rectified flow transformers, with all key evidence self-reported and no artifacts released.

  31. Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

    cs.CV 2025-06 conditional novelty 4.0 of 10

    SmPO-Diffusion improves diffusion-model preference alignment with reward-model soft labels and ReNoise inversion, reporting higher human-preference scores and up to 26x lower training cost than Diffusion-KTO.

Pith tools