Pith. sign in

Paper Citation Record · LEDGER

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

As of 13 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 30 inbound Pith citation observations for arXiv:2502.02492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02492 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:04:47.390095Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:12.520467Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.389947Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb2fc609-8a88-4ff2-8ce5-9f40dc1f3ad7 · outbound

This paper cites write newline.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.955949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.955949Z digest=sha256:1fb949cc48e5e393033504323ad82f4100a10f41c6f7f7ef2587a365e55d2ffc

Observation 3b54b379-bdad-499d-bb19-005e50082bb9 · outbound

This paper cites write newline.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.965726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.965726Z digest=sha256:c33ac8af19590bbf06f55373d30294a68cde4c1de953741480981980c3a97671

Observation b3fd3627-3c77-41da-982a-ef8173818147 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.973669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.973669Z digest=sha256:6cf1cafc6c101154f00de0d5816af10045b5990e96a3ec48d6d3c520f520c6e6

Observation 20ae7674-3293-4492-952c-a667b5745596 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.982926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.982926Z digest=sha256:44aa30fb186d8fe4bcaeff8f5d2806c303c9ea878cca458c16cb9005e55e3c4d

Observation 6fc56778-e1f4-4097-9f39-c269b6c784d0 · outbound

This paper cites FLUX , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models FLUX , 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.731711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:46.990934Z digest=sha256:bea81e2b4f87fe9a5a507d964cd367fed0852090b4c78582f9688ddf4caa5d0d

Observation 0e6ff5d8-b059-49b1-841d-617ded0a920d · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.998289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.998289Z digest=sha256:023536d48453264733364e79125762cdbc57fb63f281046f7dc299f5d2778449

Observation 1e97bc72-4b71-4d3e-8980-5c0f012ab992 · outbound

This paper cites W., Fidler, S., and Kreis, K.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models W., Fidler, S., and Kreis, K

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.705637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.005852Z digest=sha256:ca4cf487c6ce8067bd5d9d37a24d05e57813d7d73682556c72a2ddd25ba89cfb

Observation e9bce902-5ef0-48d9-a70a-da36699859a4 · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.685580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.012846Z digest=sha256:63ca717a8d61a5e60dd65084f674277ce25c3a9ed20f2ce402f330bf29488093

Observation 00c48ed4-15f6-4282-a5d7-53efe0339952 · outbound

This paper cites Video generation models as world simulators.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Video generation models as world simulators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.020437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.020437Z digest=sha256:b60e19eb13fc51e41593a478373470d05b320c7515ad9616c5209597152b9672

Observation b11d7e93-109e-4b49-823f-d1063822c9a0 · outbound

This paper cites The hidden language of diffusion models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models The hidden language of diffusion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.652827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.031598Z digest=sha256:21c38fe0e0804ef744aebeb8bb1fc1717ee9d35a4eeceba55d21ff23a10a5b40

Observation c8df70af-e99a-4a45-b005-41b3576f1b01 · outbound

This paper cites Still-Moving: Customized Video Generation without Customized Video Data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Still-Moving: Customized Video Generation without Customized Video Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.038836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.038836Z digest=sha256:8575036c76550a39bdf3897f2134cb42b7e8a7458d846605f3c5ba1af082df75

Observation c7b0eaa4-99aa-4dd1-b524-df3d3eb0f7d8 · outbound

This paper cites FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.045946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.045946Z digest=sha256:bfcba5bb3f771ee24347d8b81181596e8cadce4c5c839490ed7d71973495b12e

Observation 9b813077-7f04-4bf5-b216-b3b651c9f26a · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.053445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.053445Z digest=sha256:df34a4c475cfa0e45fc9faa9cac9de0bf07c3c9287bf9583aea5b3a2d6a85021

Observation c92ee6ee-8c57-42d4-b1f2-2cdbaca38738 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Diffusion Models Beat GANs on Image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.062406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.062406Z digest=sha256:fc2495b5bc1c27f34dbbc8b2911a8b0d90d2cfd74ff3b8d30add0fe02fb465e9

Observation a9690037-55bb-4b8e-8a0a-34d7ee51fc23 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.069229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.069229Z digest=sha256:2071e1740261f831dceaae475667ff27faff32e5b61dd7022267a1fbf3e92942

Observation 5802ce2b-5c00-4135-aaa4-ef23014dfc7a · outbound

This paper cites Motion prompting: Controlling video generation with motion trajectories, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motion prompting: Controlling video generation with motion trajectories, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.634312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.076608Z digest=sha256:c107b49a1c4d5fd24be401034026987378ff9a533de4b1e7e7dd8c4db0d95898

Observation 462eb414-94be-4d97-a5f6-0265445f844f · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.614770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.082935Z digest=sha256:f82a1acb9ad899e12623ea0245e5e7570543a26c82ef2084fe0078aa3e6a341b

Observation 0410430b-8358-4279-bfd1-a170355415da · outbound

This paper cites S., Shah, A., Yin, X., Parikh, D., and Misra, I.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models S., Shah, A., Yin, X., Parikh, D., and Misra, I

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.593746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.089154Z digest=sha256:331802e054e45438abb9779d3230ae9f0b108db67919295cf95b0775085ed54b

Observation c19d4bc7-9270-48c2-b27d-4dd49618a665 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.102846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.102846Z digest=sha256:b7fb466b92c79350171482f743b7afb1e78b05743db9105ff224904096feff84

Observation be521d57-e6eb-4338-be8a-ce4316659fd5 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models LTX-Video: Realtime Video Latent Diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.109845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.109845Z digest=sha256:90cbdaa10e19d36691678299fabeb782e23d4de043c03423404c65fbfc17a4f8

Observation 886536c6-0bdb-44ae-9568-19d34c0669f9 · outbound

This paper cites Classifier-Free Diffusion Guidance.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Classifier-Free Diffusion Guidance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.117338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.117338Z digest=sha256:2a2c6c834cc0df70f15fb50a96ed6053f9c966d9c52ec96d964826782a24afca

Observation 0f8b8cf5-c4e6-4738-ab4c-f8674d54989d · outbound

This paper cites Denoising diffusion probabilistic models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Denoising diffusion probabilistic models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.124758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.124758Z digest=sha256:7f062801709c978558b1fb65844d0a2bd71ed8aaaafd1aa900e408941b734722

Observation 0fc08a37-7298-4ac8-a968-fc873aaa90a3 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Imagen Video: High Definition Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.137721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.137721Z digest=sha256:3b565bd42bbc01fb3caed8fe9d7ab418d9337a3d404271ac133641688f86bd9a

Observation b0b13fc9-5d76-49f8-b1d9-fdcd9f4291c7 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.144778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.144778Z digest=sha256:783452e60dbd436c6177b1532513cb5a0bc17d5383e27bd797146d465a883928

Observation c31c91d1-ba7b-480c-947c-a1554dde9248 · outbound

This paper cites VBench : Comprehensive benchmark suite for video generative models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models VBench : Comprehensive benchmark suite for video generative models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.537112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.151679Z digest=sha256:08b622d8b2397e8ee8fc6d6c6e90db616d9d5b0797e42ff8241de19cff57d4db

Observation 09de8e98-60fe-4c79-b13a-4608e640291c · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling, 2024 a.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Pyramidal flow matching for efficient video generative modeling, 2024 a

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.517911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.157464Z digest=sha256:f89e0f9d78fd317f82baab00345aa60faa1679d1f71267892ae397c5774913d8

Observation 7c0fc3d1-122b-444c-a288-31572fa68350 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.493106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.163568Z digest=sha256:eb2df5b603cbde7a84bb4b7391057040dfcdad2f3c7a4bf2e00007abf9c364b2

Observation e4639c7f-dfac-475d-9c2e-ab49c6876d3a · outbound

This paper cites How far is video generation from world model: A physical law perspective, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models How far is video generation from world model: A physical law perspective, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.474814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.170092Z digest=sha256:96ef7e19d9af763dc3f8e1626b6685b67d8eda9b290f3d83d307e1f2464c0c91

Observation 053c16d2-de68-4394-8ecb-67bed69fee26 · outbound

This paper cites Kling AI , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Kling AI , 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.452135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.182859Z digest=sha256:079ca1bee8377c75259ab96d985f649d52922b9917888f73f99a24b1f5129fe7

Observation 8905310e-95d7-464a-896a-19c5a7fbfadc · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.433408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.189116Z digest=sha256:c3f06de52e0b3b4ca6c8c18b379f4b1f5806ebbb6965c671c63ef365a6086670

Observation 2987bede-615a-42d4-961b-8fbf7cedf034 · outbound

This paper cites Compositional Visual Generation with Composable Diffusion Models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Compositional Visual Generation with Composable Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.195759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.195759Z digest=sha256:b0d4cb2ef288f8fafd9b99077ced8a359a34c0223df4731f6132e8707b20f0ab

Observation 38d61803-4c6c-487b-81f4-5287cdb3ea3c · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Physgen: Rigid-body physics-grounded image-to-video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.410728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.202883Z digest=sha256:b3ae57cb6da0607be126029a0b7075c3f3f58fc32d1b877d992132677276ba48

Observation 9ffc3dfe-8fb2-4db8-86d8-2563d60760aa · outbound

This paper cites HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.210037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.210037Z digest=sha256:6f6cf990d1fca1fb6b016538dd73648da8ebc4a033469c0fe176f8c7a2a9e588

Observation 8a81af78-bd11-479b-8c02-c890c89495bb · outbound

This paper cites Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024 b.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024 b

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.387761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.217395Z digest=sha256:200aac52da07332786e88d5441e4724d97cb9b3dd67248be985872606079680d

Observation 6cf3b0bc-69a1-4c8f-b36e-cdb42c97fcdb · outbound

This paper cites K., Lewis, J.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models K., Lewis, J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.362582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.224980Z digest=sha256:a7de8219ad3df06f80abe876549c134bc75994920467d478a993db3fa71245f4

Observation 3284c59f-af94-4184-ab4a-99fdd23085b9 · outbound

This paper cites SDEdit : Guided image synthesis and editing with stochastic differential equations.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models SDEdit : Guided image synthesis and editing with stochastic differential equations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.338774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.231751Z digest=sha256:d58f7bd3db6495db13f65c73bc74cddced6c8d4809e7855a0819481780c53a3e

Observation 15175173-fb54-4991-95fd-78223bc2ae7e · outbound

This paper cites Motioncraft: Physics-based zero-shot video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motioncraft: Physics-based zero-shot video generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.304086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.239604Z digest=sha256:e9ca3ce061a1184f75cf037bac2a3f4daeabfcfee95d1026e324820b6b6a198b

Observation b42784c2-2786-44d7-8846-830f7ee23092 · outbound

This paper cites Dall-E 3 , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Dall-E 3 , 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.282485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.248404Z digest=sha256:ae38b7c902a374c6fa178c3ae77962bbf905c02e74971c37073e329cc51d307a

Observation 6837dce2-b85b-4ccc-99a3-f563cb951282 · outbound

This paper cites and Xie, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models and Xie, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.261801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.258880Z digest=sha256:6c477e495c733337b7968e9e564aa7aa8bb081b03f84049c86c66425aa11eb25

Observation 6a24c95c-e5d0-48b4-9362-644121881fda · outbound

This paper cites K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.238350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.267406Z digest=sha256:230ec81704eafde7205b4b8e9d6ae1321d472aa5b93fe86c417bb2cd045f46f7

Observation 95fc7424-e68e-4e7b-ae4d-9a6ddb5f337d · outbound

This paper cites Hierarchical spatio-temporal decoupling for text-to-video generation, 2023.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Hierarchical spatio-temporal decoupling for text-to-video generation, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.218282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.277593Z digest=sha256:b46f97695cac95b078975ec1539a63f2670fff30060a51020dc2fde0c725e3f4

Observation d528d5f8-aa23-4da5-bab4-9eaedcc3b43d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models High-resolution image synthesis with latent diffusion models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.285025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.285025Z digest=sha256:f1a67a1cab88e1f09e5eb4c752090fc65ba6b7cddf6de9f62bd75714e68007dd

Observation 185e09bf-0d55-4417-8560-1237961dad62 · outbound

This paper cites Enhancing motion in text-to-video generation with decomposed encoding and conditioning, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Enhancing motion in text-to-video generation with decomposed encoding and conditioning, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.184817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.292878Z digest=sha256:87b4b913ac1f71fd2c40c94e6d94c9d7c2d6059df7e0cbd942aefa67ddcb6533

Observation faf01b81-852b-4b01-9ed1-9ee4d869b5a8 · outbound

This paper cites DreamBooth : Fine tuning text-to-image diffusion models for subject-driven generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models DreamBooth : Fine tuning text-to-image diffusion models for subject-driven generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.165218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.299580Z digest=sha256:78105ee7d635d82b60de17695e5bb8327f966d13f8a7df239765a456935cf1bd

Observation 6d07b195-a7f5-4c05-a50b-6954504c5d32 · outbound

This paper cites Gen-3 Alpha , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Gen-3 Alpha , 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.145960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.307142Z digest=sha256:4257ce0dc3dc4635cb64cd997e1783bc79ee326d7f0bfe2db8bd0c3c8afc877e

Observation c07ad25f-1522-46f4-b83b-7d1a27265bff · outbound

This paper cites Decouple content and motion for conditional image-to-video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Decouple content and motion for conditional image-to-video generation

Reference 47

Resolution
verified exact
doi, observed 2026-08-09T12:04:47.449872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.313539Z digest=sha256:c27edce12a42693837e1d6d94b08c9bfa1a6ad8636cffe9750e0d3e07d2e42f2

Observation 8f39b790-6f7a-422b-8b89-4ebaf84d6829 · outbound

This paper cites C., See, S., Qin, H., Dai, J., and Li, H.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models C., See, S., Qin, H., Dai, J., and Li, H

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.126001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.320663Z digest=sha256:68c79ec06ef61eb21c4b64933807a5f780014dff1105556c9654624119384408

Observation 2f689545-a1a7-49e8-9b73-7538fd782d5e · outbound

This paper cites Make-A-Video : Text-to-video generation without text-video data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Make-A-Video : Text-to-video generation without text-video data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.107102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.327444Z digest=sha256:83f60b461a2e27617e0b3dee7ee14ef28b014b34ab5a25407afdd8ff5429a624

Observation 1727dd0f-d530-4044-8fe9-fed7c92c0a7f · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models UL2: Unifying Language Learning Paradigms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.336723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.336723Z digest=sha256:8f9f1e9a4fa68e8003339eb2e0cbe8220b01eab304a6dcd4cd7346303bd677bc

Observation c1a8b496-3a6c-42aa-8748-c751e71c2ecf · outbound

This paper cites and Deng, J.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models and Deng, J

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.087835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.344122Z digest=sha256:98673b98e1169032d8c071aac16538deab1e2c8d33ef8216e73a6e986b478e49

Observation 5f4f0caa-2d7f-4bc1-bdb7-33721262f9c5 · outbound

This paper cites MoCoGAN : Decomposing motion and content for video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models MoCoGAN : Decomposing motion and content for video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.070826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.350819Z digest=sha256:68d2edc85c3a9a7f37b54046647306ffd3406c77f9389ac7126f2a18cbfa3cb1

Observation d3275e23-66b0-4ff6-8e8b-84d242941f6c · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models ModelScope Text-to-Video Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.359552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.359552Z digest=sha256:ac3190f694b46be6eeb958c03abf30f50614cef3bd6f0b309b7bc15878ddf293

Observation 013abcbc-723c-424e-b43b-44ae7e72fdce · outbound

This paper cites Motif: Making text count in image animation with motion focal loss, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motif: Making text count in image animation with motion focal loss, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.049634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.368235Z digest=sha256:697fdc8bbb967b10e9950f2c50342957e3358336669918eeefb9ba561fae53cd

Observation 7394f5d3-49a2-497e-aff7-3604e1924e09 · outbound

This paper cites Z., Ge, Y., Wang, X., Lei, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Z., Ge, Y., Wang, X., Lei, S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.027826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.375593Z digest=sha256:6de6cf6bb2ba6ef05dd21e583c02a0b44b6b2204a2bff5b9591398d3ea53e80f

Observation 3104c1c0-3a9a-4f8d-905e-c33383543a60 · outbound

This paper cites Demystifying CLIP Data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Demystifying CLIP Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.381741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.381741Z digest=sha256:23aaa06ccc030d4d56e8656126252214b94b0f45aeca3eab772e3f19ea7ac42c

Observation da6c45d6-f6a4-4101-912f-577d24941d56 · outbound

This paper cites ByT5 : Towards a token-free future with pre-trained byte-to-byte models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models ByT5 : Towards a token-free future with pre-trained byte-to-byte models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.007015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.390095Z digest=sha256:dc48f27d9e1c067c84c99b01be897808e2ab54198774da1c2adc40013efb07de

Pith citing papers

Observation 3564ce24-f399-4a15-af4e-d0d840e1adc3 · inbound

MusicInfuser: Making Video Diffusion Listen and Dance cites this paper.

MusicInfuser: Making Video Diffusion Listen and Dance VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:32:15.635560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T23:28:51.767845Z digest=sha256:3d406eb8379392c57cd1cd8e53283b5b7de75f8338e7cf274d31eb704fbf6f17

Observation 8ba1652b-0765-49c2-8246-f68df6ddb1f5 · inbound

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation cites this paper.

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:12.520467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:12.520467Z digest=sha256:9e1f607152bdbd476db631d34f09e720570b57a3b58ec3779992ae66de9a1115

Observation ab1affd8-b904-4a21-afde-00c6dd068898 · inbound

LumosFlow: Motion-Guided Long Video Generation cites this paper.

LumosFlow: Motion-Guided Long Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:12.386394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:12.386394Z digest=sha256:3c6c251a70fd43a4d0f3fd127232cec3a730c0424fa0f4fe512e6dfd68719cfe

Observation f96657f5-9e9c-4fdd-a597-dc00e1d115ed · inbound

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting cites this paper.

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:33.701446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:33.701446Z digest=sha256:769208891623ddd0f54e9d0c8c31cde30ac2ce46f2b0f2867a85ff76f33e6088

Observation 0280e3ab-442d-4bc2-b5e5-35c36b402062 · inbound

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly cites this paper.

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:04.394521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:04.394521Z digest=sha256:adbfcfd39dbe34723322bd030aebd5251c6f26e825047c7bf51dadcfe2e62760

Observation 77f362a1-7732-49d8-b983-322f2ad649a1 · inbound

LuxDiT: Lighting Estimation with Video Diffusion Transformer cites this paper.

LuxDiT: Lighting Estimation with Video Diffusion Transformer VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:02.115858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:02.115858Z digest=sha256:1d66c15fc2f535584568e4d01e839e58de32fb601c9142f2cdaaaf6d13e40bb3

Observation 38f8dc10-0612-4b39-8cff-48529173ae1b · inbound

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders cites this paper.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.883059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.883059Z digest=sha256:1d2b24a9aab81e8af3e46ee1e35c894005a7b98dd9b02681e727da8fd3d21dcd

Observation e2de9bea-2e30-4ed5-b17b-54075e236bf4 · inbound

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation cites this paper.

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:39:54.294494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T22:39:53.995700Z digest=sha256:e644d0cf74f89cef1e6882dce47ef686a702507ea8022957e9df8c64cd628dbe

Observation 1af90540-05d7-4c30-ad29-e85fa9ad3195 · inbound

Motus: A Unified Latent Action World Model cites this paper.

Motus: A Unified Latent Action World Model VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:44:36.734675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T18:44:36.636455Z digest=sha256:ae624cf6d848ef4b34d6bb099daf26de043a93808a088ffbee3d1af6ba2b4485

Observation 0b855597-dbdd-4373-9433-b6114998bd80 · inbound

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion cites this paper.

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T15:34:54.195615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:34:54.195615Z digest=sha256:7d5270afbada29810a86d7f60c23a1ca26975ac56b5ff2229fce1aff29acf3c2

Observation 8d6fb0af-9848-4bb3-8cf2-efad3a500ad8 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.542476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:9f26b1a840272c4164b430c5fea47886c08070f9809276efe0d2e64bcb606b2d

Observation b241a86c-5a20-4bb5-aeee-dda2a521b46d · inbound

Reward-Forcing: Autoregressive Video Generation with Reward Feedback cites this paper.

Reward-Forcing: Autoregressive Video Generation with Reward Feedback VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:52:49.288157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T11:52:43.943226Z digest=sha256:9197514f8ba495a3895553c86c08feb71c4e03f73e11813802ffb7b5996ca695

Observation a5237902-803f-4c63-88f0-1a1ce5e7505b · inbound

Olaf-World: Orienting Latent Actions for Video World Modeling cites this paper.

Olaf-World: Orienting Latent Actions for Video World Modeling VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:20:04.762038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:20:04.762038Z digest=sha256:df44ae64c5da9bd54f110118be3535cbf958681f0f45be27b1ca32bbe72469cb

Observation feae8ad8-bf67-48e7-881c-dfd64b3de699 · inbound

HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation cites this paper.

HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.155549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T20:22:02.186361Z digest=sha256:ab3b45eb1853034b1628be7f41ad2e0a78bf1c2d039173c74c2e77835fe73d5f

Observation 7e30ce3f-34d1-4348-b70f-8a9c59eec9d9 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:04.207907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:910368cb91384b971af636dd309d21f2e9e6676c2cd4f2349717d70592e77804

Observation 6109990c-7d5a-48fa-9cb1-1a0c1e99444d · inbound

From Priors to Perception: Grounding Video-LLMs in Physical Reality cites this paper.

From Priors to Perception: Grounding Video-LLMs in Physical Reality VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:08.267707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T17:41:23.233366Z digest=sha256:eaf2f5040f6fb18d508e86af61a385a2fb9c0f599384fa52d314f2d1ea097314

Observation 17d67360-8438-420c-bfa0-88bde63bf99f · inbound

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking cites this paper.

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.858619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:28:12.151547Z digest=sha256:3eccb05bc2648043d418abe55341c42ccc84982f5ec19c84491f040f28e7abd0

Observation df95399c-099d-4deb-a662-963f0d237a07 · inbound

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation cites this paper.

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:49:41.385833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T02:49:21.291716Z digest=sha256:0938ad890a31b6431a117241c0d986c722021d6e5a6ff67715e6199f2340be11

Observation 85714818-bb9f-4531-ab90-c05f19eef78c · inbound

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes cites this paper.

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:10:22.241236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-25T05:07:14.227821Z digest=sha256:d2efa7d3c1ff045c4e6e0f0c0739724acab2ceed8ec736f44e60d398a04ee6ff

Observation e1b303de-5490-4d8a-8431-4d70e41e240a · inbound

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation cites this paper.

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.469846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-25T04:42:32.717968Z digest=sha256:dac90d2f596e02c2d6c26a7290a60e49d0718d260e0bd02871057e203b64297c

Observation 6fbf81cd-6226-4dbd-bf60-a242f417d914 · inbound

Tempered Self-Similarity Alignment for Physically Plausible Video Generation cites this paper.

Tempered Self-Similarity Alignment for Physically Plausible Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:38.411893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T11:39:06.597513Z digest=sha256:fa586f8014379e9afa885c8dddd8a6bc4cb174f547dfce4cb28d43cac2869721

Observation 5e1d9312-98de-4133-b35b-3fab77002416 · inbound

OptiWorld: Optimal Control for Video World Generation under Physical Constraints cites this paper.

OptiWorld: Optimal Control for Video World Generation under Physical Constraints VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.541970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T19:02:51.848742Z digest=sha256:0cedbf157d5c70db7a7d4fbc39e6448bca819c6cd73c43adcbdb8afcd0166c17

Observation 96b52b38-a148-4b51-af97-2297a40c1add · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.508999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:921a8c8f610c425303153b9e120f4574875ee597f5e76c49b9f21f4dc6c4fab7

Observation 2a65feb2-1926-42da-ad60-7c563ff866e8 · inbound

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation cites this paper.

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.807908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:00:35.608696Z digest=sha256:490d2f65732c8213242219f8c394e4a7ab9388e1de76b2528b6fb123e4fb1393

Observation 5d54411b-d156-4d00-9759-015cd3cdd3d0 · inbound

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics cites this paper.

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:43.391827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T04:25:27.067362Z digest=sha256:2591225c39510c6bb85a6b19606441126cd83d72ec70b993c8e53f1422c1420e

Observation 480316a2-93a5-4890-b86b-3749654d193a · inbound

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers cites this paper.

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:24:32.115411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T09:24:04.427234Z digest=sha256:6ed0a51e6183a3627efd2e44ecd8e45bd682b41e30da67adb1c973dad9561c29

Observation 2dc13058-177f-4fdd-9c28-8145f41e7066 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:1d9435b015a37c71df382f222b9e9da4047053753e452db32793bdc4f3fb0332

Observation eb4a1a26-ee33-49dc-a82f-f77a0ff436c8 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:dd18dc195741c75fbf8389d73fefcb73db80c592371282d409c5cbd58caf0f9a

Observation 819d11f0-4f0b-48fd-a282-f70213438dfb · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.332436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:35:22.332436Z digest=sha256:8824006d750e997d75163da880039434ba27fdbd342eb67efd542c3afde56de1

Observation be199bd7-a337-4f11-a865-fa77c59a3271 · inbound

DreamWAM: Beyond RGB Future Prediction for World Action Models cites this paper.

DreamWAM: Beyond RGB Future Prediction for World Action Models VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:28.004817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:59:28.004817Z digest=sha256:feb7c05d86ebd3c3b2794ac56b52c4602bb477b8e1a13a47544cc926eafe7de3