REVIEW 10 cited by
VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model's own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Project website: https://hila-chefer.github.io/videojam-paper.github.io/
Forward citations
Cited by 10 Pith papers
-
Olaf-World: Orienting Latent Actions for Video World Modeling
Latent actions become transferable across visual contexts when aligned to temporal feature differences from a frozen video encoder (SeqΔ-REPA), improving zero-shot action transfer and data-efficient adaptation of vide...
-
DreamWAM: Beyond RGB Future Prediction for World Action Models
Adding motion, depth, and semantic supervision to a world action model's future prediction during training, then removing it at inference, improves robot manipulation robustness under visual perturbations.
-
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
A hierarchical global-keyframe then local-refinement diffusion pipeline produces stable 720p/30fps music-to-dance videos longer than one minute across five genres.
-
REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
Nonlinear multi-layer compression of frozen VFM patch semantics, jointly denoised with VAE latents, improves ImageNet 256x256 FID (12.9 vs 15.2 for REG at SiT-B/2, 400K) and accelerates convergence over REPA/ReDi/REG.
-
LuxDiT: Lighting Estimation with Video Diffusion Transformer
A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.
-
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
A dual-branch video classifier uses depth and spatiotemporal features plus rank-weighted losses to categorize human-centric AI forgeries into spatial, appearance, and motion anomaly types on a new auto-labeled benchmark.
-
UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting
Jointly predicting albedo and relit appearance with one video-diffusion pass improves relighting fidelity and generalization over two-stage inverse-plus-forward pipelines.
-
LumosFlow: Motion-Guided Long Video Generation
LumosFlow generates long videos by combining large-motion key frame generation, latent optical flow diffusion, and a ControlNet-style refinement module.
-
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.
-
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
Matching video diffusion transformer tokens to concatenated DINOv2 and SAM2 features improves FVD/FID and speeds convergence, e.g., 400K-step fusion beats 1M-step baseline on UCF-101.
Discussion (0). Continue with ORCID to comment.