Pith. sign in

Paper Citation Record · LEDGER

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

As of 13 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 30 inbound Pith citation observations for arXiv:2502.02492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02492 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:04:47.390095Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:12.520467Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.389947Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb2fc609-8a88-4ff2-8ce5-9f40dc1f3ad7 · outbound

This paper cites write newline.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.955949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.955949Z digest=sha256:d1e45c0516cb276a636619557886faea5c9946f1b771a3324398aa65d7dd784a

Observation 3b54b379-bdad-499d-bb19-005e50082bb9 · outbound

This paper cites write newline.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.965726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.965726Z digest=sha256:07d70ac689fd16f58eb39f20556d1657b40430d19b20d674f19995d06945a748

Observation b3fd3627-3c77-41da-982a-ef8173818147 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.973669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.973669Z digest=sha256:aeaecc14949a7cb3a1307d6e4feaca4d7a579e2aeb015680864afc968a0330d7

Observation 20ae7674-3293-4492-952c-a667b5745596 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.982926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.982926Z digest=sha256:058815e39e6740815b7da5d30c44d6f3117b175178dea6bedf691fb25a653c4b

Observation 6fc56778-e1f4-4097-9f39-c269b6c784d0 · outbound

This paper cites FLUX , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models FLUX , 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.731711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:46.990934Z digest=sha256:eca6e4caf6f86cb4691c0d1426ecb88629f27ccc67d5a225ad2e78a87da79d4a

Observation 0e6ff5d8-b059-49b1-841d-617ded0a920d · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:46.998289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:46.998289Z digest=sha256:1092bf5ea5b79d7301db6908fe6cd7ab21ef7e7a7332e79336d826956bfc2c25

Observation 1e97bc72-4b71-4d3e-8980-5c0f012ab992 · outbound

This paper cites W., Fidler, S., and Kreis, K.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models W., Fidler, S., and Kreis, K

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.705637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.005852Z digest=sha256:04bff7cdc4d2d91dbc965a6935f45f6392b927cb8acfc8bc9a8faa78d06d144e

Observation e9bce902-5ef0-48d9-a70a-da36699859a4 · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.685580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.012846Z digest=sha256:c7518297193b265ea181da776f62396d277f95e1bb752ba7e92bff998adbbcf5

Observation 00c48ed4-15f6-4282-a5d7-53efe0339952 · outbound

This paper cites Video generation models as world simulators.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Video generation models as world simulators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.020437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.020437Z digest=sha256:d4e9a34ee341279815c559f8f10ba0d9d1f68b9d1239d6da18d76300446f5c81

Observation b11d7e93-109e-4b49-823f-d1063822c9a0 · outbound

This paper cites The hidden language of diffusion models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models The hidden language of diffusion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.652827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.031598Z digest=sha256:0bd0719985a56af0bd10eddf1197af42f3040a58eb265f2c3557045e23c23f81

Observation c8df70af-e99a-4a45-b005-41b3576f1b01 · outbound

This paper cites Still-Moving: Customized Video Generation without Customized Video Data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Still-Moving: Customized Video Generation without Customized Video Data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.038836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.038836Z digest=sha256:af2c86a67f1d315cf78c514ae902b37f95a1de2d5a0a088d29f6e49c5a3f775c

Observation c7b0eaa4-99aa-4dd1-b524-df3d3eb0f7d8 · outbound

This paper cites FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.045946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.045946Z digest=sha256:b1e754393945927749fec567b0332032cb574481b24ed39dcf7ae9372981074b

Observation 9b813077-7f04-4bf5-b216-b3b651c9f26a · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.053445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.053445Z digest=sha256:6f4fed4778c99beb2ff17397de3cdddeeb12c30c8ee98855ffa7724b1fe96633

Observation c92ee6ee-8c57-42d4-b1f2-2cdbaca38738 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Diffusion Models Beat GANs on Image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.062406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.062406Z digest=sha256:4515a75a5b3f2dee448b99e285018cd2fa1b28611a060ac59bc21dbb9d2427a4

Observation a9690037-55bb-4b8e-8a0a-34d7ee51fc23 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.069229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.069229Z digest=sha256:6af15c776d10e37ff9448703c7e6d8b0b1a1c026db2006ec4441a3b3cdb662c6

Observation 5802ce2b-5c00-4135-aaa4-ef23014dfc7a · outbound

This paper cites Motion prompting: Controlling video generation with motion trajectories, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motion prompting: Controlling video generation with motion trajectories, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.634312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.076608Z digest=sha256:a38c3df46608d76f4299f05c7092be7db473d6d5af331e2647afb457c10a8728

Observation 462eb414-94be-4d97-a5f6-0265445f844f · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.614770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.082935Z digest=sha256:22976d05179794fc4b24206ec0796f6543b1ae2190e50d0ce9b54e6a0751c64e

Observation 0410430b-8358-4279-bfd1-a170355415da · outbound

This paper cites S., Shah, A., Yin, X., Parikh, D., and Misra, I.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models S., Shah, A., Yin, X., Parikh, D., and Misra, I

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.593746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.089154Z digest=sha256:0daf33a397928659ef40e0c9c22b801372d89e79902aba4b994f210f593ee7b5

Observation c19d4bc7-9270-48c2-b27d-4dd49618a665 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.102846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.102846Z digest=sha256:48a0865db81c43fdab847c5108a9ab8beae3f83ee9ebec34e779dab370350dd4

Observation be521d57-e6eb-4338-be8a-ce4316659fd5 · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models LTX-Video: Realtime Video Latent Diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.109845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.109845Z digest=sha256:7c7854d66f160943420da1acec3126c571ea935e2f5889c6ca7eb184363954e6

Observation 886536c6-0bdb-44ae-9568-19d34c0669f9 · outbound

This paper cites Classifier-Free Diffusion Guidance.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Classifier-Free Diffusion Guidance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.117338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.117338Z digest=sha256:4ed5a00a8d8b3b1f28619b6d85403ec9804e704ab2125d1445f8732ce9ba4864

Observation 0f8b8cf5-c4e6-4738-ab4c-f8674d54989d · outbound

This paper cites Denoising diffusion probabilistic models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Denoising diffusion probabilistic models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.124758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.124758Z digest=sha256:cd542d0c9579d1dc71d75eb80f87446d356dfcaef22a945cc0a64746573e7615

Observation 0fc08a37-7298-4ac8-a968-fc873aaa90a3 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Imagen Video: High Definition Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.137721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.137721Z digest=sha256:44b040eb344411387ff8f342705344177caf12474801b10492ba7f97aa0e7212

Observation b0b13fc9-5d76-49f8-b1d9-fdcd9f4291c7 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.144778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.144778Z digest=sha256:0544c76cc32af83c0c7f6f9ec74bc570ffb7429756828e42611424f452337b75

Observation c31c91d1-ba7b-480c-947c-a1554dde9248 · outbound

This paper cites VBench : Comprehensive benchmark suite for video generative models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models VBench : Comprehensive benchmark suite for video generative models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.537112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.151679Z digest=sha256:acf5c0ce626fb013e8a59009a5543ac21365a7e53c36a61f294f7df63c4b972d

Observation 09de8e98-60fe-4c79-b13a-4608e640291c · outbound

This paper cites Pyramidal flow matching for efficient video generative modeling, 2024 a.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Pyramidal flow matching for efficient video generative modeling, 2024 a

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.517911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.157464Z digest=sha256:6220be05862312f0b7561a9db84ffc7a216e148aa638b94224d4a6e879d79cf7

Observation 7c0fc3d1-122b-444c-a288-31572fa68350 · outbound

This paper cites Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.493106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.163568Z digest=sha256:a31c4eba5f4c441c6b49b8085beee2a34afddaa1ba52f506780e0372baf4ecd6

Observation e4639c7f-dfac-475d-9c2e-ab49c6876d3a · outbound

This paper cites How far is video generation from world model: A physical law perspective, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models How far is video generation from world model: A physical law perspective, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.474814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.170092Z digest=sha256:17444182e50c28e7e0d86ed4038609eb9f1074e710dbf1b7d491fdababa2e64d

Observation 053c16d2-de68-4394-8ecb-67bed69fee26 · outbound

This paper cites Kling AI , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Kling AI , 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.452135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.182859Z digest=sha256:ea3f72407b515d1ef72bb48799c1d2c0a2f4d491aab3bfa9979032b33e0572d3

Observation 8905310e-95d7-464a-896a-19c5a7fbfadc · outbound

This paper cites an unresolved cited work.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:04:48.433408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.189116Z digest=sha256:8ec234b483fea81c87b0ec2631ffce97b5ef67d99158dd76357f1e5fd70baf4d

Observation 2987bede-615a-42d4-961b-8fbf7cedf034 · outbound

This paper cites Compositional Visual Generation with Composable Diffusion Models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Compositional Visual Generation with Composable Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.195759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.195759Z digest=sha256:8d1535d3aa7ee886fde64462d965b5e6127299fabfd2278d2f34ef422d377357

Observation 38d61803-4c6c-487b-81f4-5287cdb3ea3c · outbound

This paper cites Physgen: Rigid-body physics-grounded image-to-video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Physgen: Rigid-body physics-grounded image-to-video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.410728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.202883Z digest=sha256:f3f2caa587c12c2960016b4aa7526573d978c44927ca0a0d9c36c1df9ad07107

Observation 9ffc3dfe-8fb2-4db8-86d8-2563d60760aa · outbound

This paper cites HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.210037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.210037Z digest=sha256:0e37f73a451ec7f6a0cff498c37ee11fe03ba77d4801a55997ab1e7a458b92f2

Observation 8a81af78-bd11-479b-8c02-c890c89495bb · outbound

This paper cites Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024 b.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Sora: A review on background, technology, limitations, and opportunities of large vision models, 2024 b

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.387761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.217395Z digest=sha256:45b4332fdb31bf4eeb9c6d56b9657e4c2680677b5cb3135df98671c04fcc4e14

Observation 6cf3b0bc-69a1-4c8f-b36e-cdb42c97fcdb · outbound

This paper cites K., Lewis, J.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models K., Lewis, J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.362582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.224980Z digest=sha256:cf2bb114af349374d16a566815166cf8ec7bbe2de19eddfc13e36bfd1e8f5484

Observation 3284c59f-af94-4184-ab4a-99fdd23085b9 · outbound

This paper cites SDEdit : Guided image synthesis and editing with stochastic differential equations.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models SDEdit : Guided image synthesis and editing with stochastic differential equations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.338774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.231751Z digest=sha256:56eb0fb9e9d26789bea0aaaecdc6232faa86651d3f6dcc94d55a8e4795f0aec5

Observation 15175173-fb54-4991-95fd-78223bc2ae7e · outbound

This paper cites Motioncraft: Physics-based zero-shot video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motioncraft: Physics-based zero-shot video generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.304086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.239604Z digest=sha256:14287f8dd17a66bb6f41673c1fa36df505e2df50abb9290e8fc920e4f139a7f3

Observation b42784c2-2786-44d7-8846-830f7ee23092 · outbound

This paper cites Dall-E 3 , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Dall-E 3 , 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.282485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.248404Z digest=sha256:0751f36d662f23b78a1437ce13a233d0a7278be816ed50219280ca710cc50e83

Observation 6837dce2-b85b-4ccc-99a3-f563cb951282 · outbound

This paper cites and Xie, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models and Xie, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.261801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.258880Z digest=sha256:4ec9b49f7fc19a76737b8d1da299321826f4f833581cee38f0764eb566c01f16

Observation 6a24c95c-e5d0-48b4-9362-644121881fda · outbound

This paper cites K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models K., Zhang, P., Vajda, P., Duval, Q., Girdhar, R., Sumbaly, R., Rambhatla, S

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.238350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.267406Z digest=sha256:c7c83cda8bdcd5cc003ba9245c4e9c1b6b5fcde05cfcda5bfe56dc56817576d3

Observation 95fc7424-e68e-4e7b-ae4d-9a6ddb5f337d · outbound

This paper cites Hierarchical spatio-temporal decoupling for text-to-video generation, 2023.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Hierarchical spatio-temporal decoupling for text-to-video generation, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.218282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.277593Z digest=sha256:e578cf4fc8caddab8adefb06aea997b435728e4877d0c53517708921d864b86d

Observation d528d5f8-aa23-4da5-bab4-9eaedcc3b43d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models High-resolution image synthesis with latent diffusion models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.285025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.285025Z digest=sha256:30c05ef9f65c4c25b0917a356dbdc1b2fd223ef983dba8f359307e42c38a8733

Observation 185e09bf-0d55-4417-8560-1237961dad62 · outbound

This paper cites Enhancing motion in text-to-video generation with decomposed encoding and conditioning, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Enhancing motion in text-to-video generation with decomposed encoding and conditioning, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.184817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.292878Z digest=sha256:1bbb3b299cd52ec26d56a9783050dc940be050c9840b566d3b3460b47cd6cf9b

Observation faf01b81-852b-4b01-9ed1-9ee4d869b5a8 · outbound

This paper cites DreamBooth : Fine tuning text-to-image diffusion models for subject-driven generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models DreamBooth : Fine tuning text-to-image diffusion models for subject-driven generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.165218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.299580Z digest=sha256:637502ae9c5f3b811c1e746d5ca909b9ab830afe51e1927fc1cf1eea838eeca6

Observation 6d07b195-a7f5-4c05-a50b-6954504c5d32 · outbound

This paper cites Gen-3 Alpha , 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Gen-3 Alpha , 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.145960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.307142Z digest=sha256:8617ce42c46209de633d1792b4d743220d522e229586a6dfd8a466cd69a6f544

Observation c07ad25f-1522-46f4-b83b-7d1a27265bff · outbound

This paper cites Decouple content and motion for conditional image-to-video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Decouple content and motion for conditional image-to-video generation

Reference 47

Resolution
verified exact
doi, observed 2026-08-09T12:04:47.449872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.313539Z digest=sha256:492cab05c8c0dca3a6ea74adeedf919cfd57538bd41e1e1d274a8092f98820ad

Observation 8f39b790-6f7a-422b-8b89-4ebaf84d6829 · outbound

This paper cites C., See, S., Qin, H., Dai, J., and Li, H.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models C., See, S., Qin, H., Dai, J., and Li, H

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.126001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.320663Z digest=sha256:fc5082957696d037d04792250e0685ce06ebf8bd53a6b7178c88c369cbfbf2bf

Observation 2f689545-a1a7-49e8-9b73-7538fd782d5e · outbound

This paper cites Make-A-Video : Text-to-video generation without text-video data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Make-A-Video : Text-to-video generation without text-video data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.107102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.327444Z digest=sha256:17edfa0fd62eab36f88947f6886a5d6d5506134afdc21b4368752d62e75ab322

Observation 1727dd0f-d530-4044-8fe9-fed7c92c0a7f · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models UL2: Unifying Language Learning Paradigms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.336723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.336723Z digest=sha256:a98435c7114d4e8a2e8fee2a5be37a9eb249852ccd54ea2ea4577aa5e1b8deac

Observation c1a8b496-3a6c-42aa-8748-c751e71c2ecf · outbound

This paper cites and Deng, J.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models and Deng, J

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.087835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.344122Z digest=sha256:a8ca70f2d3c41ce0ca530f9a4f1e8cb1be4b5a9367fedea6db8fcdb70938cd91

Observation 5f4f0caa-2d7f-4bc1-bdb7-33721262f9c5 · outbound

This paper cites MoCoGAN : Decomposing motion and content for video generation.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models MoCoGAN : Decomposing motion and content for video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.070826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.350819Z digest=sha256:1e944386ae0e051f50644b8c05124890ff3d7d79e8081107c108dcbcb76912e1

Observation d3275e23-66b0-4ff6-8e8b-84d242941f6c · outbound

This paper cites ModelScope Text-to-Video Technical Report.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models ModelScope Text-to-Video Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.359552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.359552Z digest=sha256:102be3f2b504eae75015566a2fe9b6ece8b748ce14c22a103523615e6538ee4b

Observation 013abcbc-723c-424e-b43b-44ae7e72fdce · outbound

This paper cites Motif: Making text count in image animation with motion focal loss, 2024.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Motif: Making text count in image animation with motion focal loss, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.049634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.368235Z digest=sha256:dc1fb1d68d893e40cb4fd35f703f99518b6785568bb159772968ed51993e93f3

Observation 7394f5d3-49a2-497e-aff7-3604e1924e09 · outbound

This paper cites Z., Ge, Y., Wang, X., Lei, S.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Z., Ge, Y., Wang, X., Lei, S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.027826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.375593Z digest=sha256:e7819cc0eabc89038ced2eb1e9ddd8ef95d37594fef6a481f1324c5ca5a37590

Observation 3104c1c0-3a9a-4f8d-905e-c33383543a60 · outbound

This paper cites Demystifying CLIP Data.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models Demystifying CLIP Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T12:04:47.381741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:04:47.381741Z digest=sha256:0f291834da6954a5fa8b497a874960ba2e6e783f6108b25ad8f5c100a2bc66f8

Observation da6c45d6-f6a4-4101-912f-577d24941d56 · outbound

This paper cites ByT5 : Towards a token-free future with pre-trained byte-to-byte models.

VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models ByT5 : Towards a token-free future with pre-trained byte-to-byte models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:04:48.007015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T12:04:47.390095Z digest=sha256:278b0d7de4a8656e715bbb14fbd451ce82e98444a25b060196361f12ddac4bc5

Pith citing papers

Observation 3564ce24-f399-4a15-af4e-d0d840e1adc3 · inbound

MusicInfuser: Making Video Diffusion Listen and Dance cites this paper.

MusicInfuser: Making Video Diffusion Listen and Dance VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:32:15.635560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T23:28:51.767845Z digest=sha256:6e497fc69381b2f478025393fd4a56192f079ac09fdce8b3456ffc04a4dab6c7

Observation 8ba1652b-0765-49c2-8246-f68df6ddb1f5 · inbound

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation cites this paper.

FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:12.520467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:12.520467Z digest=sha256:97330661c8a91402879243079f5db76414b46d8bc84e9c81edc1a6eeb41b0b46

Observation ab1affd8-b904-4a21-afde-00c6dd068898 · inbound

LumosFlow: Motion-Guided Long Video Generation cites this paper.

LumosFlow: Motion-Guided Long Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:12.386394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:12.386394Z digest=sha256:b2f85eba254a37a45e6072ca97247ae87f29aa644b7ff1dc7bd31a2dfcaf9fb6

Observation f96657f5-9e9c-4fdd-a597-dc00e1d115ed · inbound

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting cites this paper.

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:33.701446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:33.701446Z digest=sha256:530a2187440187969ddf33b04faa90b44eafe0a110230c81d964d924d330a657

Observation 0280e3ab-442d-4bc2-b5e5-35c36b402062 · inbound

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly cites this paper.

HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:04.394521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:04.394521Z digest=sha256:6ba0928a033ed4e5f2beacc89cb8cd6739cc7f1080ed296579d43bfbcf306223

Observation 77f362a1-7732-49d8-b983-322f2ad649a1 · inbound

LuxDiT: Lighting Estimation with Video Diffusion Transformer cites this paper.

LuxDiT: Lighting Estimation with Video Diffusion Transformer VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:02.115858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:02.115858Z digest=sha256:6d997769b79a0365cf25d54871e29c1de32d5b014ed0e017a73d3d8a434532fb

Observation 38f8dc10-0612-4b39-8cff-48529173ae1b · inbound

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders cites this paper.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.883059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.883059Z digest=sha256:1f8cdeddfcb4792ae01d4fb2fb21cccbbb269e74c4cdb803620f3ba5a1d50f8f

Observation e2de9bea-2e30-4ed5-b17b-54075e236bf4 · inbound

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation cites this paper.

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:39:54.294494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T22:39:53.995700Z digest=sha256:0a65e517927c8d118e41a9e10772548c76759d5ee0eb4159f7f0d89b07879f0d

Observation 1af90540-05d7-4c30-ad29-e85fa9ad3195 · inbound

Motus: A Unified Latent Action World Model cites this paper.

Motus: A Unified Latent Action World Model VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:44:36.734675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T18:44:36.636455Z digest=sha256:61d14383f805cc45ecbd8be87fcca2ccedc17e0083e44802c81f94b35f8b849e

Observation 0b855597-dbdd-4373-9433-b6114998bd80 · inbound

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion cites this paper.

REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T15:34:54.195615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:34:54.195615Z digest=sha256:9e86bd3bf728bb0043d3128e8eba24eb377f9c4337574d98b2a32293f4b92e08

Observation 8d6fb0af-9848-4bb3-8cf2-efad3a500ad8 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.542476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:92346726d7ad2cc30243ff15a01cb3617827d517bd8422986d8df36a212a73ef

Observation b241a86c-5a20-4bb5-aeee-dda2a521b46d · inbound

Reward-Forcing: Autoregressive Video Generation with Reward Feedback cites this paper.

Reward-Forcing: Autoregressive Video Generation with Reward Feedback VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:52:49.288157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T11:52:43.943226Z digest=sha256:49f8e12bad4a70bbde441ced59215ab72a571693ec45e3d03f818079b0ec4cf9

Observation a5237902-803f-4c63-88f0-1a1ce5e7505b · inbound

Olaf-World: Orienting Latent Actions for Video World Modeling cites this paper.

Olaf-World: Orienting Latent Actions for Video World Modeling VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:20:04.762038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:20:04.762038Z digest=sha256:40d846bac445d22c2537ef7647d77ba3261beab575f3148540d435db748c502c

Observation feae8ad8-bf67-48e7-881c-dfd64b3de699 · inbound

HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation cites this paper.

HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.155549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T20:22:02.186361Z digest=sha256:fe1bea1f65a7471989878abd717c331d52d0fbce34bfe12529d651a632f29e22

Observation 7e30ce3f-34d1-4348-b70f-8a9c59eec9d9 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:04.207907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:6820250b528f3bf58ecd890508d3266eb2eafc8f876859979c5d2c51efea164a

Observation 6109990c-7d5a-48fa-9cb1-1a0c1e99444d · inbound

From Priors to Perception: Grounding Video-LLMs in Physical Reality cites this paper.

From Priors to Perception: Grounding Video-LLMs in Physical Reality VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:08.267707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T17:41:23.233366Z digest=sha256:ca3a4a89fd8b1eb7fc6043cd24e9df42ce4ac9d93d1735afdca5f794f3b62e61

Observation 17d67360-8438-420c-bfa0-88bde63bf99f · inbound

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking cites this paper.

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.858619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T21:28:12.151547Z digest=sha256:22dfc17b991274f1f7aaed37bfbad3871c6042addfcf74f0ea39a08664973cd7

Observation df95399c-099d-4deb-a662-963f0d237a07 · inbound

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation cites this paper.

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:49:41.385833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T02:49:21.291716Z digest=sha256:d70bd49e98918c9d263a237a924d9b852aca509283cde66901b36460598d56af

Observation 85714818-bb9f-4531-ab90-c05f19eef78c · inbound

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes cites this paper.

Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:10:22.241236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-25T05:07:14.227821Z digest=sha256:f29aae7acc02d4243119f0b49b114726c125cbab8784e30c116d5273e6c74192

Observation e1b303de-5490-4d8a-8431-4d70e41e240a · inbound

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation cites this paper.

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:45:20.469846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T04:42:32.717968Z digest=sha256:32f7d695869e439c8e3b8bee4a1cec7b6e0adc83e7183b9911c82bc5ea4c2d59

Observation 6fbf81cd-6226-4dbd-bf60-a242f417d914 · inbound

Tempered Self-Similarity Alignment for Physically Plausible Video Generation cites this paper.

Tempered Self-Similarity Alignment for Physically Plausible Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:38.411893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T11:39:06.597513Z digest=sha256:3412ff620de9605b39d183c4df2c544a40a737657f1774a258e736f4a0817c67

Observation 5e1d9312-98de-4133-b35b-3fab77002416 · inbound

OptiWorld: Optimal Control for Video World Generation under Physical Constraints cites this paper.

OptiWorld: Optimal Control for Video World Generation under Physical Constraints VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.541970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T19:02:51.848742Z digest=sha256:d091d7d199d27ad10199f26a922ecb483aa9312ab69449f75241fe533cdeed7a

Observation 96b52b38-a148-4b51-af97-2297a40c1add · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.508999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:fde94f1d6d9ff76da768ee72603af215f336a17b7d25b9a84efad06b0bd63be6

Observation 2a65feb2-1926-42da-ad60-7c563ff866e8 · inbound

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation cites this paper.

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.807908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T10:00:35.608696Z digest=sha256:ee60d55d541082abe1ebc5df12c01dc59ce20f3e5a0e6c2df405a325c1d83c5c

Observation 5d54411b-d156-4d00-9759-015cd3cdd3d0 · inbound

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics cites this paper.

NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:43.391827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T04:25:27.067362Z digest=sha256:123a1a793b8846b4357c763e80c3d15aa928ac6af97ce5a0f8e4096c35061a94

Observation 480316a2-93a5-4890-b86b-3749654d193a · inbound

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers cites this paper.

HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:24:32.115411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T09:24:04.427234Z digest=sha256:f91d46650fc52ca968bf69ba82db8594d2117d16314dadcd81f6d6af1ba5df7e

Observation 2dc13058-177f-4fdd-9c28-8145f41e7066 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:418c232a83098ac562ec21bdb62ad184cf7b5f908e4b39f5584ce6272e5878d4

Observation eb4a1a26-ee33-49dc-a82f-f77a0ff436c8 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:f21344fd548bd9a297a700e849ed77af2a4995a345b5f524bf0862b349d49498

Observation 819d11f0-4f0b-48fd-a282-f70213438dfb · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.332436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:35:22.332436Z digest=sha256:02cdb3ca526093d8612048d6615c1e8014a0c011bd71c7f2999b1edbe42dc856

Observation be199bd7-a337-4f11-a865-fa77c59a3271 · inbound

DreamWAM: Beyond RGB Future Prediction for World Action Models cites this paper.

DreamWAM: Beyond RGB Future Prediction for World Action Models VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:28.004817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:59:28.004817Z digest=sha256:df477407b972933131cf265e918d40a3c5aa630ff5afc13791ed106661919680