REVIEW 24 cited by
Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This technical report presents a cost-efficient strategy for training a video generation foundation model. We present a mid-sized research model with approximately 7 billion parameters (7B) called Seaweed-7B trained from scratch using 665,000 H100 GPU hours. Despite being trained with moderate computational resources, Seaweed-7B demonstrates highly competitive performance compared to contemporary video generation models of much larger size. Design choices are especially crucial in a resource-constrained setting. This technical report highlights the key design decisions that enhance the performance of the medium-sized diffusion model. Empirically, we make two observations: (1) Seaweed-7B achieves performance comparable to, or even surpasses, larger models trained on substantially greater GPU resources, and (2) our model, which exhibits strong generalization ability, can be effectively adapted across a wide range of downstream applications either by lightweight fine-tuning or continue training. See the project page at https://seaweed.video/
Forward citations
Cited by 24 Pith papers
-
FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers
A DiT-based portrait animation model transfers implicit facial expressions to one or more characters using a masked cross-attention mechanism, supported by a new multi-face dataset and benchmark.
-
Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations
A channel-wise reuse algorithm plus a reconfigurable systolic accelerator skips redundant vDiT attention and MLP computation, achieving up to 5.9x speedup and 16x energy savings.
-
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
PISCES post-trains text-to-video models using dual optimal-transport-aligned rewards (global quality plus token-level semantic) and outperforms annotation-based and annotation-free baselines on VBench and human evaluation.
-
Transition Matching Distillation for Fast Video Generation
Splitting a video diffusion model into a fixed feature extractor and a small recurrent flow head lets TMD generate videos in one to two effective steps with better VBench scores than prior distilled models.
-
Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
Timeripple cuts vDiT self-attention compute by up to 85% by reusing partial attention scores of spatially and temporally correlated tokens across channels, with VBench quality essentially unchanged.
-
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
HuMo uses a two-stage training scheme and a face-focus trick to generate human videos that follow text, keep a reference person's identity, and sync speech to audio, beating several single-task systems on benchmarks.
-
Waver: Wave Your Way to Lifelike Video Generation
Waver unifies text-to-video, image-to-video, and text-to-image generation in a single 12B-parameter DiT with a hybrid dual/single-stream architecture and a cascade refiner, claiming top-three public leaderboard performance.
-
X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents
A self-supervised framework encodes whole-body motion into four identity-agnostic latent tokens and uses them to animate reference images, outperforming skeleton-based baselines.
-
NeoBabel: A Multilingual Open Tower for Visual Generation
A 2B multilingual text-to-image model trained on 124M translated pairs matches or beats larger English-only baselines on English while scoring higher on the authors' multilingual benchmark extensions.
-
DreamArt: Generating Interactable Articulated Objects from a Single Image
From one image, DreamArt generates a textured 3D articulated object with segmented moving parts and plausible motion, using a mask-prompted video diffusion model and dual quaternion joint optimization.
-
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
Phantom-Data provides around one million cross-context, identity-consistent reference-video pairs for subject-to-video generation, and training on it improves prompt following and visual quality.
-
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.
-
PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms
PanoWan adapts the Wan 2.1 text-to-video model to generate seamless 360-degree videos by remapping initial noise, rotating the latent grid during denoising, and padding the latent before VAE decoding, trained on a new...
-
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.
-
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
NSG-VD detects AI-generated videos by measuring the ratio of spatial probability gradients to temporal density changes and comparing these 'NSG' features with a maximum mean discrepancy test.
-
RewardDance: Reward Scaling in Visual Generation
RewardDance reframes visual reward modeling as a yes/no judgment task in a VLM and reports consistent gains in text-to-image, text-to-video, and image-to-video generation as the reward model scales from 1B to 26B.
-
From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms
An explainable model using BLEURT, CometKiwi, pause features, and Chinese phraseological diversity predicts human-rated quality dimensions in English-Chinese consecutive interpreting, with SHAP identifying the stronge...
-
Captain Cinema: Towards Short Movie Generation
A two-stage text-to-movie system that plans keyframes for the story and then synthesizes video between them, using a compressed memory bank to keep long narratives consistent.
-
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
Timestep-segment preference optimization with separate motion and fidelity LoRAs improves audio-driven human animation quality and allows a 3.3x inference speedup.
-
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.
-
ATI: Any Trajectory Instruction for Controllable Video Generation
ATI injects user-drawn point trajectories as soft Gaussian feature masks into a pretrained image-to-video diffusion model, enabling unified camera, object, and local motion control.
-
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
Frame-level captions with per-frame cross-attention and parallel multi-window denoising reduce semantic confusion in long multi-scene video generation in the authors' internal evaluation.
-
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
A hierarchical direct preference optimization with four alignment levels plus automated data selection improves physical plausibility of text-to-video models.
-
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.
Discussion (0). Sign in to comment.