REVIEW 6 cited by
From Sora What We Can See: A Survey of Text-to-Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
With impressive achievements made, artificial intelligence is on the path forward to artificial general intelligence. Sora, developed by OpenAI, which is capable of minute-level world-simulative abilities can be considered as a milestone on this developmental path. However, despite its notable successes, Sora still encounters various obstacles that need to be resolved. In this survey, we embark from the perspective of disassembling Sora in text-to-video generation, and conducting a comprehensive review of literature, trying to answer the question, \textit{From Sora What We Can See}. Specifically, after basic preliminaries regarding the general algorithms are introduced, the literature is categorized from three mutually perpendicular dimensions: evolutionary generators, excellent pursuit, and realistic panorama. Subsequently, the widely used datasets and metrics are organized in detail. Last but more importantly, we identify several challenges and open problems in this domain and propose potential future directions for research and development.
Forward citations
Cited by 6 Pith papers
-
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
An agentic two-stage optimizer (LLM-generated question scoring plus Bayesian hyperparameter search) improves image-to-video prompt adherence, winning human preference tests up to 69% over random search.
-
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
Parallel, role-specialized prompt agents improve cultural relevance in text-to-video generation, with a new cross-cultural benchmark showing the largest gains for location cues.
-
Seeing World Dynamics in a Nutshell
NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.
-
Speaking images. A novel framework for the automated self-description of artworks
A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.
-
Medical Video Generation for Disease Progression Simulation
MVG generates synthetic disease-progression videos from one medical image and a text prompt, using repeated diffusion editing plus video interpolation, evaluated on chest X-ray, retina, and skin images.
-
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
Scene Copilot is a training-free pipeline that combines an LLM with retrieval over Infinigen's codebase and human-in-the-loop Blender editing to generate customized 3D scenes and videos from text prompts.
Discussion (0). Continue with ORCID to comment.