REVIEW 11 cited by
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses significant challenges, such as modeling the complex relationship between space and time, and the lack of large-scale text-video paired data. Existing text-video datasets suffer from limitations in both content quality and scale, or they are not open-source, rendering them inaccessible for study and use. For model design, previous approaches extend pretrained text-to-image generation models by adding temporal 1D convolution/attention modules for video generation. However, these approaches overlook the importance of jointly modeling space and time, inevitably leading to temporal distortions and misalignment between texts and videos. In this paper, we propose a novel approach that strengthens the interaction between spatial and temporal perceptions. In particular, we utilize a swapped cross-attention mechanism in 3D windows that alternates the "query" role between spatial and temporal blocks, enabling mutual reinforcement for each other. Moreover, to fully unlock model capabilities for high-quality video generation and promote the development of the field, we curate a large-scale and open-source video dataset called HD-VG-130M. This dataset comprises 130 million text-video pairs from the open-domain, ensuring high-definition, widescreen and watermark-free characters. A smaller-scale yet more meticulously cleaned subset further enhances the data quality, aiding models in achieving superior performance. Experimental quantitative and qualitative results demonstrate the superiority of our approach in terms of per-frame quality, temporal correlation, and text-video alignment, with clear margins.
Forward citations
Cited by 11 Pith papers
-
SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection
A hyperspherical prototype boundary with temporal-coherence losses improves continual AI-generated video detection by about 3 to 4 percentage points over prior methods.
-
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.
-
Detecting AI-Generated Videos with Spiking Neural Networks
MAST with spiking neural networks achieves 93.14% mean accuracy detecting AI-generated videos from 10 unseen generators by exploiting smoother pixel residuals and compact semantic trajectories.
-
AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
AEGIS is a large-scale benchmark for detecting AI-generated videos, with a hard test set of Sora and KLing clips that current vision-language models detect at near-chance accuracy.
-
FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion
FreeLong++ extends short-video diffusion models to 4x to 8x longer clips, without retraining, by fusing multiple windowed attention branches through frequency-domain filters and a spectral noise initialization.
-
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.
-
Retrieval-Driven Training-Free AI-Generated Video Attribution
A training-free retrieval pipeline using adaptive color transforms, multi-scale quantized residuals, and temporal aggregation attributes AI-generated videos to one of eight generators with 84.6% Rank-1 and 78.3% mAP o...
-
ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency
ObjFiller3D jointly optimizes a dense 360-degree ring of views to inpaint 3D objects with cross-view-consistent textures, reporting higher PSNR and LPIPS than per-view baselines at much lower runtime.
-
TextMesh4D: Zero-shot Text-to-4D Mesh Generation
TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.
-
RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images
Reversing bit-plane weights ('bit-reversed image') plus a gradient-selected 32×32 patch lets a small ResNet detect AI-generated images with state-of-the-art accuracy on many benchmarks.
-
A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality
A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.
Discussion (0). Continue with ORCID to comment.