Pith. sign in

REVIEW 8 cited by

Fr\'echet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.16124 v1 pith:QV6JLE26 submitted 2024-07-23 cs.CV

classification cs.CV
keywords videomotionconsistencyqualitydistanceechetevaluatingfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Significant advancements have been made in video generative models recently. Unlike image generation, video generation presents greater challenges, requiring not only generating high-quality frames but also ensuring temporal consistency across these frames. Despite the impressive progress, research on metrics for evaluating the quality of generated videos, especially concerning temporal and motion consistency, remains underexplored. To bridge this research gap, we propose Fr\'echet Video Motion Distance (FVMD) metric, which focuses on evaluating motion consistency in video generation. Specifically, we design explicit motion features based on key point tracking, and then measure the similarity between these features via the Fr\'echet distance. We conduct sensitivity analysis by injecting noise into real videos to verify the effectiveness of FVMD. Further, we carry out a large-scale human study, demonstrating that our metric effectively detects temporal noise and aligns better with human perceptions of generated video quality than existing metrics. Additionally, our motion features can consistently improve the performance of Video Quality Assessment (VQA) models, indicating that our approach is also applicable to unary video quality evaluation. Code is available at https://github.com/ljh0v0/FMD-frechet-motion-distance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Explicit instance-level physical parameter conditioning with routing attention improves physical-law consistency in image-to-video generation, as measured on the authors' new simulator-based benchmark.

  2. Inference-Time Concept Suppression and Video-Centric Evaluation for Text-to-Video Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SIRUS is a training-free, inference-time framework that suppresses target concepts in text-to-video diffusion models via subspace-informed prompt projection and residual subtraction, with a new video-centric unlearnin...

  3. Generative Relightable Avatars

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    GRA combines UV-space material optimization and physics rendering with feed-forward texture refinement and a fine-tuned video-to-video diffusion model to achieve controllable, high-detail relighting of full-body avatars.

  4. MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos

    cs.CV 2026-01 conditional novelty 6.0 of 10

    MorphGS retargets motion from a monocular video onto a rigged 3D character by optimizing target morphology and pose with image-space losses, without 3D source reconstruction or parametric templates.

  5. DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model fuses per-person masks with pose keypoints to generate identity-preserving, two-person interactive videos from a single reference image, outperforming prior single-person-animation pipelines.

  6. HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A closed-loop image-to-video quality-control system reports 87–97% expert agreement on internal clips, with no released data, code, or thresholds.

  7. DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion

    cs.CV 2025-09 conditional novelty 5.0 of 10

    DynaMind reconstructs videos from EEG by combining region-aware semantic mapping, a temporal blueprint, and dual-guidance diffusion, outperforming EEG2Video on SEED-DV in most comparisons.

  8. Unveiling Audio Deepfake Origins: A Deep Metric learning And Conformer Network Approach With Ensemble Fusion

    cs.SD 2025-06 conditional novelty 4.0 of 10

    An XLSR-Conformer system trained with Real Emphasis, Fake Dispersion, and multi-class N-pair loss reaches 95.6% in-domain and up to 44.8% out-of-domain source tracing accuracy on MLAAD, versus 83.4% and 26.5% for the ...

Pith tools