Pith. sign in

REVIEW 10 cited by

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.02492 v2 pith:LNS3Y4ZI submitted 2025-02-04 cs.CV

classification cs.CV
keywords motionvideomodelmodelsvideojamcoherencegenerationappearance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model's own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Project website: https://hila-chefer.github.io/videojam-paper.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Olaf-World: Orienting Latent Actions for Video World Modeling

    cs.CV 2026-02 conditional novelty 7.0 of 10

    Latent actions become transferable across visual contexts when aligned to temporal feature differences from a frozen video encoder (SeqΔ-REPA), improving zero-shot action transfer and data-efficient adaptation of vide...

  2. DreamWAM: Beyond RGB Future Prediction for World Action Models

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Adding motion, depth, and semantic supervision to a world action model's future prediction during training, then removing it at inference, improves robot manipulation robustness under visual perturbations.

  3. Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A hierarchical global-keyframe then local-refinement diffusion pipeline produces stable 720p/30fps music-to-dance videos longer than one minute across five genres.

  4. REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Nonlinear multi-layer compression of frozen VFM patch semantics, jointly denoised with VAE latents, improves ImageNet 256x256 FID (12.9 vs 15.2 for REG at SiT-B/2, 400K) and accelerates convergence over REPA/ReDi/REG.

  5. LuxDiT: Lighting Estimation with Video Diffusion Transformer

    cs.GR 2025-09 conditional novelty 6.0 of 10

    A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.

  6. HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly

    cs.CV 2025-07 reject novelty 6.0 of 10

    A dual-branch video classifier uses depth and spatiotemporal features plus rank-weighted losses to categorize human-centric AI forgeries into spatial, appearance, and motion anomaly types on a new auto-labeled benchmark.

  7. UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Jointly predicting albedo and relit appearance with one video-diffusion pass improves relighting fidelity and generalization over two-stage inverse-plus-forward pipelines.

  8. LumosFlow: Motion-Guided Long Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    LumosFlow generates long videos by combining large-motion key frame generation, latent optical flow diffusion, and a ControlNet-style refinement module.

  9. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.

  10. Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Matching video diffusion transformer tokens to concatenated DINOv2 and SAM2 features improves FVD/FID and speeds convergence, e.g., 400K-step fusion beats 1M-step baseline on UCF-101.

Pith tools