Pith. sign in

Paper Citation Record · LEDGER

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

As of 10 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 16 inbound Pith citation observations for arXiv:2502.04847.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04847 v5

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:16:16.877500Z

measured 93 of 93 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:27:36.393512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:27:09.332551Z

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0834b76e-0f24-4c5b-ae55-287b2553c5bc · outbound

This paper cites Conditional gan with discrimi- native filter generation for text-to-video synthesis.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Conditional gan with discrimi- native filter generation for text-to-video synthesis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.203410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.390131Z digest=sha256:8739804f9f9fd7d6d581f244cdc5e9bda5cbf86f05986b5d6c838fba234f4915

Observation 2c869da1-5d43-4b8a-8e5e-ae3029a27fd7 · outbound

This paper cites Multidiffusion: Fusing diffusion paths for controlled image generation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Multidiffusion: Fusing diffusion paths for controlled image generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.396032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.396032Z digest=sha256:3bec06e19385de933ab51d8abcbe31305b4bbd029cfffff24c22c6f7a5ec03d0

Observation efeb2e9a-f12b-4b51-8b1b-bc4548846815 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.401695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.401695Z digest=sha256:b61859ee841f38234cb7acb9fdb2b947e06c2ef6028170934367f5093380c374

Observation 5e867c90-832c-4d76-b588-9e933755b13a · outbound

This paper cites an unresolved cited work.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T21:16:18.177660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.406915Z digest=sha256:441b69ad9cbcfd3ac7abb748f326fd8d0a107e3c4fd2780d7bf0b179951ec2a8

Observation 3d33765d-5839-4cce-859a-b6d478ef8e19 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.411483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.411483Z digest=sha256:dab3f695c1c4d52c058b6ec3bc5f4d48790c84f42a57805b5165d70b4fbe6cbb

Observation 2f672520-901d-472b-8d6f-af74cccc060a · outbound

This paper cites Realtime multi-person 2d pose estimation using part affinity fields.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Realtime multi-person 2d pose estimation using part affinity fields

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.163718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.416429Z digest=sha256:baf0d7028876f749e47a34417d1ee1c0ac6e938aab27ce4ab9be762051c48e2f

Observation c498ad1a-e33d-4853-abbe-c37987dc16f6 · outbound

This paper cites Everybody dance now.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Everybody dance now

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.148978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.421709Z digest=sha256:b7972af9751ef9078496c87c5d228833daf55f10a41dc870988deabae2003b90

Observation a7639469-4c26-4c8b-91b4-555d6c4612f2 · outbound

This paper cites MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.428062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.428062Z digest=sha256:ea93a7fd6eb38ab76b717461ed4d9c42cdc1e1c3b7adc76680886e420c0f0eb9

Observation aa6ddcd7-b2fa-40d7-b52b-ca67bd2e72b4 · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Long video generation with time-agnostic vqgan and time- sensitive transformer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.134097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.433617Z digest=sha256:d8ecb1d9eb4e50790f00331449a76b822750fc715f4c82af1fe564431520efa1

Observation eca8bde4-c838-4f29-a199-83c04fef5c19 · outbound

This paper cites Learning individual styles of conversational gesture.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Learning individual styles of conversational gesture

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.118288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.438120Z digest=sha256:f37171f3cbf80b00aea2d340cfe1cb2f6ce087baefbd106ccd29e53daa2bfa35

Observation c20ca445-e2d8-46ea-8f40-bb6342833873 · outbound

This paper cites Generative adversarial nets.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Generative adversarial nets

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.442724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.442724Z digest=sha256:cb2eefaccd992cb3a476c1bd01a7b3e054e4d35fffb4389309cc4637a5a3a3a2

Observation 98ede29c-9ae8-4ec6-b8a0-7e943aa7e93e · outbound

This paper cites Talk-act: Enhance textural- awareness for 2d speaking avatar reenactment with diffusion model.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Talk-act: Enhance textural- awareness for 2d speaking avatar reenactment with diffusion model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.094178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.447391Z digest=sha256:8c9f86ad8978f4d050b49f5111fe7c92d7ccf62f99a103cfd2d1207c4257e0ca

Observation 16d9b75e-2cd6-48fd-aa2c-d6c220f1d6c3 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.452043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.452043Z digest=sha256:6ffd5e54fa36c85eab2deb5d5d7d3288b38bf2cc6130ff2ad008dcfb8d88d0da

Observation 3431fe6f-626a-4d1c-b21e-0696711b0326 · outbound

This paper cites Deep residual learning for image recognition.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Deep residual learning for image recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.079523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.456898Z digest=sha256:f21a53b8c3c271d88acb42dfe0ed2b9ff7e3c4c4d127fa49d5eaadddc417bae0

Observation fc46ad49-c553-4da6-9e06-a105509aa82e · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.461518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.461518Z digest=sha256:4d73fc60d5b2d8e0729283cc64936d269a054fdbad08ad373a6db67c7860eff8

Observation ea606338-560e-4248-b141-d20c78b9cf6a · outbound

This paper cites Denoising dif- fusion probabilistic models.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Denoising dif- fusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.466599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.466599Z digest=sha256:69e55a4c3dc0b05d904c7b803a3a8f5fd614a943086b9cc159a9226cab4f9ccf

Observation 18a105c2-6925-4da7-8873-a1211c7f9486 · outbound

This paper cites Image quality metrics: Psnr vs.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Image quality metrics: Psnr vs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.045336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.471480Z digest=sha256:163fda84c363bf8155b5f3381153c7e46b4ef50136f4e0c5fa76673a325ba1d9

Observation 549d5fbc-a99c-45ff-8cbd-55a3a9cd6411 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.028672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.476156Z digest=sha256:4ba98b7bfcb47b46a47ec5c64e7a4b70e2e06cdc6585b357738ba9dfc121938f

Observation f203d0e7-d3a2-49ec-b015-f4888087e0f0 · outbound

This paper cites Learning high fidelity depths of dressed humans by watching social media dance videos.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Learning high fidelity depths of dressed humans by watching social media dance videos

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:18.013348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.480789Z digest=sha256:c5d3f0b69bb69b7c76ecda2e3670c7e4c81810630a7644983212d34d3a5ad98a

Observation fbe126c7-0ba4-47d9-a498-0c300d1eed96 · outbound

This paper cites Ultralyt- ics yolo.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Ultralyt- ics yolo

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.997949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.485493Z digest=sha256:c653bf16c448fb116024879c1c065971eb237f2d57207f3dede981fb5116a571

Observation b8c3bbd4-7563-4444-9d6f-4d9273549119 · outbound

This paper cites CoTracker: It is Better to Track Together.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation CoTracker: It is Better to Track Together

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.490082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.490082Z digest=sha256:e5a138aa4a9044c8dcc1d0a6140baeaeff1b443625651c2c9e465a6d6ad71815

Observation caa24bc2-58d8-48b4-b717-e5cceb5bdbca · outbound

This paper cites Sapiens: Foundation for Human Vision Models.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Sapiens: Foundation for Human Vision Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.496425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.496425Z digest=sha256:0c4426d2dedcc4c5175df706b2ea689f68eb4f42a760af91ea991b1b8c71e6df

Observation 606a2cb4-fb2a-4153-a29e-930d68f22793 · outbound

This paper cites Auto-Encoding Variational Bayes.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Auto-Encoding Variational Bayes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.502978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.502978Z digest=sha256:76ec97276a7ea169730d3d11eba33b337a2c9c2a7ecc2413a27190cc6efd633c

Observation 9e654174-9035-441a-afb5-2289664958c6 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Sequence Parallelism: Long Sequence Training from System Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.508011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.508011Z digest=sha256:dbd71e6a84fad17be0b2a2572ff35f86d95a77161fe5f56f085b50c008dd049d

Observation 38b59f59-1114-4173-b6b6-5c56b655cdb6 · outbound

This paper cites Speech2video synthesis with 3d skeleton regularization and expressive body poses.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Speech2video synthesis with 3d skeleton regularization and expressive body poses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.983827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.513558Z digest=sha256:93cb167d56e50eb4ae449f45e834e3327ebe673c65149be39289cd210fc2e089

Observation 95f0d104-63d7-4cfe-b340-b57b9e71ad3d · outbound

This paper cites CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.518388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.518388Z digest=sha256:cb7772ea368dadf3b008cc4d5af7fe908d5204814b295eca74c3c6e7b7536d4a

Observation eb1ee01c-a96d-4f98-97e6-e1d5a1088a54 · outbound

This paper cites TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.523335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.523335Z digest=sha256:d9243bbbd146e46f12220e805850a41cddf06f215a33d94820840f7f9b0f8503

Observation 89a2f01c-ce7b-454c-96ea-dbb90475d0cf · outbound

This paper cites Multi-task deep model with margin ranking loss for lung nodule analysis.IEEE transactions on medical imaging, 39(3):718–728, 2019.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Multi-task deep model with margin ranking loss for lung nodule analysis.IEEE transactions on medical imaging, 39(3):718–728, 2019

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.969500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.528457Z digest=sha256:14f899e601f0cf4c600cc59b7673dc6a14123ce08d2c3202d1d6e3ab134e274e

Observation c5a8a129-ed4e-4935-b30d-d81f3f6f7038 · outbound

This paper cites Smpl: A skinned multi- person linear model.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Smpl: A skinned multi- person linear model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.533470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.533470Z digest=sha256:9d0c9a33e400d472b066c63144a5376fe30e577e52dd95bd9a3505d2ad4164e1

Observation 52b64872-ff26-4240-b71f-97b34095901c · outbound

This paper cites Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.946311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.537970Z digest=sha256:4d41c0ddeb6f3f3357c48035f3aa648cb58a9984b66de8035870cde67a7397ca

Observation 5aa0e527-979d-4b22-b76c-0616eaee2232 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- thesis.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Nerf: Representing scenes as neural radiance fields for view syn- thesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.542698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.542698Z digest=sha256:1a6d2edefdb4ea52b8c99f4e23d61cabbbbbd15010d1baf1ef941cddb50eacc5

Observation 2496d139-34fc-46a5-b581-446bd63ead28 · outbound

This paper cites Moore-animateanyone.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Moore-animateanyone

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.921316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.547379Z digest=sha256:be90f7e73eb57e4c0d8f5bca155f09818a5f9d16c590f1f7eb93b5395c4d48fa

Observation 233041cb-075e-4173-9bfa-5f6ccd144ef5 · outbound

This paper cites Conditional image-to-video gener- ation with latent flow diffusion models.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Conditional image-to-video gener- ation with latent flow diffusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.905179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.552150Z digest=sha256:4b044c2cd0b434359b695bbe85aa7e4bd496538bc93f2404f4e6b2145beb4736

Observation 9f915bc5-4a7f-44ad-8975-4549ea24517f · outbound

This paper cites Sora: Creating video from text.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Sora: Creating video from text

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.889890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.557548Z digest=sha256:45eccd7d37fcafab0b495b2a0bd686459e4e3335c453720017f3f1f6acf02197

Observation 88ea8bda-db1c-48b2-aef0-7a104e4ea5a4 · outbound

This paper cites Paddleocr.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Paddleocr

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.873815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.563034Z digest=sha256:f14ee81f7076db070330c1db2e17db21e4568117c64935ae1047e2be8c2b63be

Observation adfd2fdb-0242-4786-a0fa-7f47f3d2fe36 · outbound

This paper cites Scalable diffusion models with transformers.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Scalable diffusion models with transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.859729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.567873Z digest=sha256:2a115b20bfa09b3bee31048bf12caf64f80a189e4d5d34337c74b863620b18d9

Observation 09d26619-aab7-4ffd-8d26-9fda7e80d414 · outbound

This paper cites Deep spatial transformation for pose-guided person image generation and animation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Deep spatial transformation for pose-guided person image generation and animation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.844555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.572483Z digest=sha256:6c9eea7170dc0e3068400968797bdad3438c3eaedbe4d0eaa1811cabcb16b12d

Observation 24224dcf-4068-418c-8e1a-117086c20b20 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation High-resolution image syn- thesis with latent diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.577335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.577335Z digest=sha256:07e75c21b0628a78ce6b2205a78a21b94488be87d79f877b8cd04e96cb98c75c

Observation 33fb1b53-ee0f-4e6e-80e6-e70dbbb0045c · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation U- net: Convolutional networks for biomedical image segmen- tation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.582200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.582200Z digest=sha256:d6f0b9cd2a03fcdf217781946dc8c596842a3677c21576285d97db9f629b6d78

Observation a1c91d50-98b4-44bd-86c0-0c9c894cd653 · outbound

This paper cites Human4dit: 360-degree human video gen- eration with 4d diffusion transformer.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Human4dit: 360-degree human video gen- eration with 4d diffusion transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.810583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.586781Z digest=sha256:1489bb977b127070416ac4e7266a73d75df67258cab0627627a411957bd743f6

Observation 9dd8795f-b24a-43b9-9fe1-47d12aaaec3b · outbound

This paper cites Animating arbitrary objects via deep motion transfer.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Animating arbitrary objects via deep motion transfer

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.795373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.591807Z digest=sha256:9aa663fa9b29115f6a17cba287c3438a9fa988940544314e9c3941b2403b8c28

Observation 172ff92d-adf6-4301-98bc-71555fc6afc7 · outbound

This paper cites Motion representations for ar- ticulated animation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Motion representations for ar- ticulated animation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.778921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.596535Z digest=sha256:c9160486fa83a653cbe1d110879d5f4149220f2471eac2c35d3b31ac53d87be8

Observation 2103c47e-0a16-4176-b17b-25aa5352f454 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.602005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.602005Z digest=sha256:17d84803560da1881a645521b8743f7a67a27d5038341d467f71628747a9ac9a

Observation d98cf96c-4cc7-44fa-92d9-a43133b90b35 · outbound

This paper cites Denoising Diffusion Implicit Models.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Denoising Diffusion Implicit Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.607106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.607106Z digest=sha256:36062a0d11489aa83194ca7a7f5d20925f2da810ee65fe510fd3e79f48db2df4

Observation 6e965526-5176-4704-94f5-9d875645fc68 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Roformer: Enhanced transformer with rotary position embedding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.611957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.611957Z digest=sha256:66acf71bc6a9f4adc4d3fdee61e96b1fcc81c58aa4a8b850be03662efebcdfac

Observation 782dfa92-6bcd-488e-b40c-904810c8543a · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.617683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.617683Z digest=sha256:651b699a6ef555683cf178053a0378f207b7e4cd9c15eedf809993c7b5ded546

Observation 7baf9bc0-3533-4b64-9ab5-cce85c8a5e5e · outbound

This paper cites Animate-X: Universal Character Image Animation with Enhanced Motion Representation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Animate-X: Universal Character Image Animation with Enhanced Motion Representation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.622832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.622832Z digest=sha256:58800e463d6a533a3ae7af65a33baaf4ede6ba13bfa9faa32744487a600a3759

Observation 373bfee0-0b59-44d6-aee2-a60a77e9bf8d · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.628222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.628222Z digest=sha256:eed41dee14b3bf5fce2ffbb0c456bf7ba8e2ef31e9e45b88629b9c7a98719fd3

Observation 9072b334-5323-40ea-89c2-7011a2b3a55d · outbound

This paper cites Attention is all you need.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.755583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.633095Z digest=sha256:2ec4544811a5504b9e271c5258e20ee5f7a854528a32ae11c8d2e96eecd4d397

Observation 9d46e001-4bc7-480c-8ff6-046dcf66f649 · outbound

This paper cites Generating videos with scene dynamics.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Generating videos with scene dynamics

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.741557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.637631Z digest=sha256:27ad3f5ebc3fe3dfee9b71df81cb2286f3a2ed0bb2ee460995aaab2484ae58bc

Observation dd1837a2-7651-4dd2-ab15-f431a74f2edb · outbound

This paper cites DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.642095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.642095Z digest=sha256:8a1df4c6f9390872ecc4adb4fd83a6a7ac1df729d0c7201046c4856e4f774de9

Observation c50bc728-9176-4194-9a9e-54e4ec6930aa · outbound

This paper cites Disco: Disentangled control for realistic human dance generation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Disco: Disentangled control for realistic human dance generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.726467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.647776Z digest=sha256:ac3dc25dd4cc425c4d3fd9b349baa939e640ee14f3241f348c1f2a5fe1535276

Observation 2fb4275f-5ada-40e4-a547-0ca1db18da9e · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferenc- ing.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation One-shot free-view neural talking-head synthesis for video conferenc- ing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.711730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.652251Z digest=sha256:00a4e3f91733e0d69f48efd640a5f3b2c4a230dadeb98ea3c24cc22c897fed1f

Observation 416fd5ef-2c35-41e3-8b4a-535d8e32da04 · outbound

This paper cites UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.657225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.657225Z digest=sha256:bd025e444bc051cf87314bb7dc6c377c9b88e4af1efc24bc4858432656ebbd53

Observation e575bc99-aac1-4a4c-944b-0434b22fc388 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Image quality assessment: from error visibility to structural similarity

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.661873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.661873Z digest=sha256:1180dd48264098b0028bf64b5a9ed1af56f2e4d10d32379ae06a175521c8a39e

Observation 24e9cf9d-bea3-4070-953d-868b2bae6971 · outbound

This paper cites Hu- mannerf: Free-viewpoint rendering of moving people from monocular video.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Hu- mannerf: Free-viewpoint rendering of moving people from monocular video

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.685282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.666305Z digest=sha256:b734d88a73563a2dd5c5c958bf2cc47d7db9e147238a574ba3413c641152c07a

Observation 2f038d34-90d9-498f-a2e1-6fd9ad51f590 · outbound

This paper cites Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.670898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.670898Z digest=sha256:25c8f2eef57decc3eaaa6e6679c431ef69ed8ea5a3162273b7ef732c069d5ae0

Observation a55463b4-cb3a-4126-83c6-86b4632aa584 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Msr-vtt: A large video description dataset for bridging video and language

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.668483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.675811Z digest=sha256:a45a2fec6db0e501a8b8a1a813612a0ae4058a50b693e5271fb1f018fb20e2ea

Observation de0f70ff-e413-4a15-94a6-c9c1cc532d16 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Easyanimate: A high-performance long video generation method based on transformer architecture

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.680352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.680352Z digest=sha256:0bf3de15cbf0eb10c6e5ff951d920c2a8a9bbeef25a86cdbe1826890b25fc1d6

Observation ac3f24bd-effe-46ad-a0a0-0c2e50252347 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Magicanimate: Temporally consistent human image animation using diffusion model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.653090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.685334Z digest=sha256:d80e491c9f0a02b7621239e7ed551864e4e4b4b2e2e46779073ac1f4f6b7a8da

Observation e3cb2976-eb72-4c37-9028-95cd3dc88840 · outbound

This paper cites Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.689671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.689671Z digest=sha256:a9f9a4979b648ac7fe9fa3267aa0b5f197e7c48449dddb4fc7ddfe795c2397ae

Observation 555e26cb-a4ef-4315-815e-6292691ca29c · outbound

This paper cites Showmaker: Creating high-fidelity 2d human video via fine-grained diffusion mod- eling.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Showmaker: Creating high-fidelity 2d human video via fine-grained diffusion mod- eling

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.637641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.695401Z digest=sha256:e44d5e2c1e15f9f6a202d3cb0304b0640b9925d67db69956867fb29fbd07af07

Observation 5fff86d7-cb6c-44db-a67d-26601bd47015 · outbound

This paper cites Effec- tive whole-body pose estimation with two-stages distillation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Effec- tive whole-body pose estimation with two-stages distillation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.622111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.700622Z digest=sha256:daad2e29084a9729c4e949df5fd1593d25516e1c13c4c8ee80d4824abdd77fbc

Observation 18f37ecd-5158-45f5-9842-0e1ff2df2a33 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.705812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.705812Z digest=sha256:d2cd4e39d9c5d4b454d6625200ea6daee6df0bfb87fd01b9955a761c84440f19

Observation fc77aae5-26e4-4c1b-b632-b2b5a52ef798 · outbound

This paper cites GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.710740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.710740Z digest=sha256:a749584b605f59048d9e57ec495e5ce617958fa6e8126d83ead237966cc35e21

Observation fad4119a-6fa2-4c5e-957a-98147c8bfe6d · outbound

This paper cites Make pixels dance: High- dynamic video generation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Make pixels dance: High- dynamic video generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.606749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.715531Z digest=sha256:9e934c67a7d5c1e949938a51c724d1ec6704a5175868378f7a1c64071babeca0

Observation cd43f7a3-c811-407c-b541-57871757afc1 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Adding conditional control to text-to-image diffusion models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.720041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.720041Z digest=sha256:0ae806012a02790776638d3741225ceb0d69a91d13b232eb89e66ee424666b4d

Observation 61ab8a81-7c20-4c62-86c6-16f684085ff2 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation The unreasonable effectiveness of deep features as a perceptual metric

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.834912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.834912Z digest=sha256:a60b1a65f1cd5ed3d5b835078a87c7b773e47ffe67734fb212bee60c1a4f9147

Observation 21d3a4cb-102b-424d-a8fc-06c0e4785f05 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.839806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.839806Z digest=sha256:6cb2d6fecf846aefe4c080b837c46dd033592afb288ab9b71601cc9bbaf2074e

Observation 105b8940-7c80-4d26-a950-286a12b94711 · outbound

This paper cites Tora: Trajectory-oriented Diffusion Transformer for Video Generation.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Tora: Trajectory-oriented Diffusion Transformer for Video Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.844580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.844580Z digest=sha256:fe6644605cc9220c81e06cfd5759fd0ed8baf5b3d19efa38d8d248b184469f1b

Observation 175a91c8-fe74-4340-9fdc-bbb350425cac · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.849374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.849374Z digest=sha256:564a853a6e7d1bb0384e5bc3e215856dc260808e5014f8a32149a1e6d6855e7b

Observation 480f3328-3a0e-45b0-b6cf-a78765179e6a · outbound

This paper cites RealisDance: Equip controllable character animation with realistic hands.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation RealisDance: Equip controllable character animation with realistic hands

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.854235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.854235Z digest=sha256:c47ad5b35358a5b8db44f6a6d65d1edc30cbe764d1912e9b464211a6dc7243de

Observation 16637732-4e53-428d-81ef-8d4ac8e1e2b4 · outbound

This paper cites Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.858895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.858895Z digest=sha256:66501626b750114a74559506a501c93706748b3fee5c90db4758fbc1274158aa

Observation cf959b04-8e84-4ff4-bc71-8419ed68adbf · outbound

This paper cites Sapiens [22] is employed to obtain pose keypoints, providing robust human pose detection for each frame.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Sapiens [22] is employed to obtain pose keypoints, providing robust human pose detection for each frame

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.555041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.868609Z digest=sha256:4d025ce9259dc7365e1dea2dd0139c757db2d044079c0844d1bda8f03a4c9ca8

Observation bdb5b3d6-f602-403e-8425-878904bb3167 · outbound

This paper cites an unresolved cited work.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T21:16:17.537946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.873147Z digest=sha256:42e9ba680211928a681bbd336f46dac7f97a469d439515fddb6c029bfaa086a4

Observation e3058091-ac8c-49d6-977b-1dfa8a8c378a · outbound

This paper cites PaddleOCR [35] is used to identify and mark text regions in each frame, mitigating the potential interference of text artifacts with the generated data.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation PaddleOCR [35] is used to identify and mark text regions in each frame, mitigating the potential interference of text artifacts with the generated data

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.522331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.877500Z digest=sha256:c08881f1b825bcdc183cefb2470edd2ee2d8994a6d1aa8036a6bf018db30111e

Observation 1259710e-9eec-40fe-a124-9612c5edb2d9 · outbound

This paper cites More Details for Data Collection A.1.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation More Details for Data Collection A.1

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T21:16:17.571500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T21:16:16.863814Z digest=sha256:d9676dbf19b2b6643f0a9e8681c8b50173cedb5dbf9949c62478595a952796f3

Pith citing papers

Observation 07cca7ad-aab0-4be6-895e-9914d2213c04 · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:36.393512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:36.393512Z digest=sha256:dd97ec433aca60359f10eb299d2874a70203aeec9ce1dd661ece20eaf1c3c86f

Observation 7ac5ae84-74b1-4c0d-8aae-4c50a97ad4c1 · inbound

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer cites this paper.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:57.359048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:57.359048Z digest=sha256:90a6b04fae4433b93db48d25091c1368a6b8c9db666c184cb84b9d86f9658a81

Observation c953dc74-b63c-44c1-9bdf-5a01e206b2bc · inbound

FramePrompt: In-context Controllable Animation with Zero Structural Changes cites this paper.

FramePrompt: In-context Controllable Animation with Zero Structural Changes HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:04.478492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:04.478492Z digest=sha256:72044c520eb1c067c591f4afd2ea59da2a1c35f460ec8f919d94718b1c6f6777

Observation 864de853-98ee-47f8-9c56-56039a9a2dc7 · inbound

CharacterShot: Controllable and Consistent 4D Character Animation cites this paper.

CharacterShot: Controllable and Consistent 4D Character Animation HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:13:13.531008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:13:13.531008Z digest=sha256:7cff0a32e1ae261402e8d819b8c1bdd27159119817294f1af96312a76f3e9a60

Observation 2a2ed21f-c783-4b36-8408-d28c9f114b26 · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.115944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.115944Z digest=sha256:22297880f7f62da1a674b88a210922209f68cd618646d1523c87a22d60e6bc6a

Observation 230ce555-70fb-495c-a420-ec49c1992237 · inbound

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer cites this paper.

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:25:28.387712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T18:25:22.486891Z digest=sha256:9ae2d75a281e584806d79834ffdbf70315f194eec3c05bf8d4e5b508c06ce413

Observation 27613b73-29e8-4d03-b74f-6f1dc30001e1 · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:47:57.469405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:74b3c779f256dca9fbe6b2967d62fc74cbb905c9d8e810a3b639ea5cdb6087b8

Observation 67e18852-6877-4312-8590-388f7765649b · inbound

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors cites this paper.

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T22:49:03.259461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:49:03.259461Z digest=sha256:4dabd15e33d6075139194c7e49f1df5501434e8879ce23aa7d3cb054550ac054

Observation d78da58c-5a2f-425c-a00f-78ab574cdf1f · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.606070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:fd192bdf0b346c73e600e5c51d1eea107e84416703afc4e16595689f724f80a9

Observation a587e507-2fe7-47b8-a21b-f4ce67d6f45c · inbound

Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting cites this paper.

Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.902718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T23:00:45.496971Z digest=sha256:052fe083e541ac4935624713a5e95023006887c9a581b6ef10bd4097865d7d52

Observation 6dfe205b-da8b-4496-a966-c1f14f0859a3 · inbound

EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration cites this paper.

EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:25:04.928437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:21:58.123630Z digest=sha256:f59f886c9740352921f341246218c32ff117582d470980e7a68449f7b7832da0

Observation 5881a7e2-485c-4bac-a8df-f1ade38fdddb · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:25.067557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:3ca027dd596084f6186b21a45de3cc375bed1d10422aeef7ab1bc2eb4b965075

Observation 7d27c7c8-caf6-4dfc-b1e9-bc628558a0d5 · inbound

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization cites this paper.

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:46:19.837955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:02:10.901341Z digest=sha256:6a2bb3c8ae4dc6f72a6bee818d4ba04fcf8daf9e35b90393bbf50aa6b8f0d0c6

Observation 2ac62087-02d9-4145-92ab-920eeae36bf1 · inbound

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy cites this paper.

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.334025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T22:41:06.600948Z digest=sha256:83f73392163b2eb46923d5c14590bd09e573a98e7ecde882e134f3a2e65221d5

Observation 7dadeb01-d7c9-4a57-a229-86b334d4821b · inbound

Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis cites this paper.

Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:24:32.499642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T09:19:50.748415Z digest=sha256:47ebb6954244ea55683c0363b043313c94fab3d1e667b6907e6f182273bf6db9

Observation 34bcd504-bca9-4d4c-92d5-282b38a8f4ba · inbound

ViDS: Video Diffusion Shader using 3D Face Tracking cites this paper.

ViDS: Video Diffusion Shader using 3D Face Tracking HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T23:01:24.315729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:01:24.315729Z digest=sha256:fb95aa5a71d74f1c424a8be357d1b037e1a2e06c69bcaa80baae1d45ad7fc0bf