Pith. sign in

REVIEW 11 cited by

Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.10874 v4 pith:LZP5ZLB5 submitted 2023-05-18 cs.CV

classification cs.CV
keywords generationtemporaltext-videovideoattentionqualityapproachapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses significant challenges, such as modeling the complex relationship between space and time, and the lack of large-scale text-video paired data. Existing text-video datasets suffer from limitations in both content quality and scale, or they are not open-source, rendering them inaccessible for study and use. For model design, previous approaches extend pretrained text-to-image generation models by adding temporal 1D convolution/attention modules for video generation. However, these approaches overlook the importance of jointly modeling space and time, inevitably leading to temporal distortions and misalignment between texts and videos. In this paper, we propose a novel approach that strengthens the interaction between spatial and temporal perceptions. In particular, we utilize a swapped cross-attention mechanism in 3D windows that alternates the "query" role between spatial and temporal blocks, enabling mutual reinforcement for each other. Moreover, to fully unlock model capabilities for high-quality video generation and promote the development of the field, we curate a large-scale and open-source video dataset called HD-VG-130M. This dataset comprises 130 million text-video pairs from the open-domain, ensuring high-definition, widescreen and watermark-free characters. A smaller-scale yet more meticulously cleaned subset further enhances the data quality, aiding models in achieving superior performance. Experimental quantitative and qualitative results demonstrate the superiority of our approach in terms of per-frame quality, temporal correlation, and text-video alignment, with clear margins.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A hyperspherical prototype boundary with temporal-coherence losses improves continual AI-generated video detection by about 3 to 4 percentage points over prior methods.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. Detecting AI-Generated Videos with Spiking Neural Networks

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    MAST with spiking neural networks achieves 93.14% mean accuracy detecting AI-generated videos from 10 unseen generators by exploiting smoother pixel residuals and compact semantic trajectories.

  4. AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences

    cs.CV 2025-08 conditional novelty 6.0 of 10

    AEGIS is a large-scale benchmark for detecting AI-generated videos, with a hard test set of Sora and KLing clips that current vision-language models detect at near-chance accuracy.

  5. FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FreeLong++ extends short-video diffusion models to 4x to 8x longer clips, without retraining, by fusing multiple windowed attention branches through frequency-domain filters and a spectral noise initialization.

  6. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  7. Retrieval-Driven Training-Free AI-Generated Video Attribution

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free retrieval pipeline using adaptive color transforms, multi-scale quantized residuals, and temporal aggregation attributes AI-generated videos to one of eight generators with 84.6% Rank-1 and 78.3% mAP o...

  8. ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ObjFiller3D jointly optimizes a dense 360-degree ring of views to inpaint 3D objects with cross-view-consistent textures, reporting higher PSNR and LPIPS than per-view baselines at much lower runtime.

  9. TextMesh4D: Zero-shot Text-to-4D Mesh Generation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.

  10. RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Reversing bit-plane weights ('bit-reversed image') plus a gradient-selected 32×32 patch lets a small ResNet detect AI-generated images with state-of-the-art accuracy on many benchmarks.

  11. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.

Pith tools