Pith. sign in

REVIEW 8 cited by

Goku: Flow Based Video Generative Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.04896 v2 pith:VO5PWDYV submitted 2025-02-07 cs.CV

classification cs.CV
keywords generationgokumodelsflowimage-and-videojointperformancetasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

    cs.CV 2025-09 conditional novelty 6.0 of 10

    HuMo uses a two-stage training scheme and a face-focus trick to generate human videos that follow text, keep a reference person's identity, and sync speech to audio, beating several single-task systems on benchmarks.

  2. AnyAni: An Interactive System with Generative AI for Animation Effect Creation and Code Understanding in Web Development

    cs.HC 2025-06 conditional novelty 6.0 of 10

    AnyAni combines LLM generation, a version tree, and video-based checking to help front-end developers create and understand web animations; a nine-person study reports usability gains over a chatbot baseline.

  3. Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Phantom-Data provides around one million cross-context, identity-consistent reference-video pairs for subject-to-video generation, and training on it improves prompt following and visual quality.

  4. M4V: Multimodal Mamba for Efficient Text-to-Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    M4V shows a Mamba-based text-to-video model can roughly match attention-based PyramidFlow on VBench while cutting mixer-layer FLOPs by 45% at 768x1280.

  5. Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Matching video diffusion transformer tokens to concatenated DINOv2 and SAM2 features improves FVD/FID and speeds convergence, e.g., 400K-step fusion beats 1M-step baseline on UCF-101.

  6. From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    An explainable model using BLEURT, CometKiwi, pause features, and Chinese phraseological diversity predicts human-rated quality dimensions in English-Chinese consecutive interpreting, with SHAP identifying the stronge...

  7. RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Metric-scale depth alignment plus scene-constrained noise shaping improves camera controllability and video quality for image-to-video generation on RealEstate10K.

  8. Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A hierarchical direct preference optimization with four alignment levels plus automated data selection improves physical plausibility of text-to-video models.

Pith tools