Pith. sign in

REVIEW 7 cited by

T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14505 v2 pith:APUXNAOC submitted 2024-07-19 cs.CV

classification cs.CV
keywords text-to-videocompositionalgenerationbindingmodelsbenchmarkgenerativemetrics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this important ability for evaluation. In this work, we conduct the first systematic study on compositional text-to-video generation. We propose T2V-CompBench, the first benchmark tailored for compositional text-to-video generation. T2V-CompBench encompasses diverse aspects of compositionality, including consistent attribute binding, dynamic attribute binding, spatial relationships, motion binding, action binding, object interactions, and generative numeracy. We further carefully design evaluation metrics of multimodal large language model (MLLM)-based, detection-based, and tracking-based metrics, which can better reflect the compositional text-to-video generation quality of seven proposed categories with 1400 text prompts. The effectiveness of the proposed metrics is verified by correlation with human evaluations. We also benchmark various text-to-video generative models and conduct in-depth analysis across different models and various compositional categories. We find that compositional text-to-video generation is highly challenging for current models, and we hope our attempt could shed light on future research in this direction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    KeyFrame-Compass tests nine video generators on 386 keyframe-sequence tasks and finds a consistent trade-off between keyframe fidelity and natural video quality, with control degrading under dense keyframes and open-s...

  2. "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.

  3. DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model fuses per-person masks with pose keypoints to generate identity-preserving, two-person interactive videos from a single reference image, outperforming prior single-person-animation pipelines.

  4. Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A new large benchmark for AI-generated human-centric video quality with pairwise preferences, plus a Mixture-of-Experts MLLM that outperforms prior methods on rating, comparison, and Q&A.

  5. VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    VISTA is a video synthesis framework that creates controllable egocentric videos of daily tasks with reactive and proactive agent intervention modes via causal reverse reasoning scripts.

  6. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

  7. FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    FaceAnonyMixer claims a cancelable face generation method that irreversibly mixes real latent codes with key-derived synthetic codes for privacy-preserving face recognition.

Pith tools