REVIEW 7 cited by
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this important ability for evaluation. In this work, we conduct the first systematic study on compositional text-to-video generation. We propose T2V-CompBench, the first benchmark tailored for compositional text-to-video generation. T2V-CompBench encompasses diverse aspects of compositionality, including consistent attribute binding, dynamic attribute binding, spatial relationships, motion binding, action binding, object interactions, and generative numeracy. We further carefully design evaluation metrics of multimodal large language model (MLLM)-based, detection-based, and tracking-based metrics, which can better reflect the compositional text-to-video generation quality of seven proposed categories with 1400 text prompts. The effectiveness of the proposed metrics is verified by correlation with human evaluations. We also benchmark various text-to-video generative models and conduct in-depth analysis across different models and various compositional categories. We find that compositional text-to-video generation is highly challenging for current models, and we hope our attempt could shed light on future research in this direction.
Forward citations
Cited by 7 Pith papers
-
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
KeyFrame-Compass tests nine video generators on 386 keyframe-sequence tasks and finds a consistent trade-off between keyframe fidelity and natural video quality, with control degrading under dense keyframes and open-s...
-
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.
-
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
A diffusion model fuses per-person masks with pose keypoints to generate identity-preserving, two-person interactive videos from a single reference image, outperforming prior single-person-animation pipelines.
-
Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model
A new large benchmark for AI-generated human-centric video quality with pairwise preferences, plus a Mixture-of-Experts MLLM that outperforms prior methods on rating, comparison, and Q&A.
-
VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios
VISTA is a video synthesis framework that creates controllable egocentric videos of daily tasks with reactive and proactive agent intervention modes via causal reverse reasoning scripts.
-
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.
-
FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
FaceAnonyMixer claims a cancelable face generation method that irreversibly mixes real latent codes with key-derived synthetic codes for privacy-preserving face recognition.
Discussion (0). Continue with ORCID to comment.