Pith. sign in

Paper Citation Record · LEDGER

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

As of 14 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2412.14531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14531 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:12:57.852352Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:07.730833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:01:15.907000Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b3fa548-3614-4917-972a-201c47a0543c · outbound

This paper cites write newline.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.467958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.467958Z digest=sha256:03ab5ce5e047c867f2a82ec5dbeb2f65802c3cae848be849d15dd9f4a456402b

Observation 7c98eda0-28c9-4fb1-b6d4-8f0687717af6 · outbound

This paper cites Conditional gan with discriminative filter generation for text-to-video synthesis.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Conditional gan with discriminative filter generation for text-to-video synthesis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.949517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.474392Z digest=sha256:7a99320188ec125b26952c25ed28bdd244de9ff7eefbe25e3e5fe0cf8eab9c89

Observation 2e544546-efc6-4a21-ac5b-7542beb296fd · outbound

This paper cites Person image synthesis via denoising diffusion model.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Person image synthesis via denoising diffusion model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.934811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.479824Z digest=sha256:c79e989d55f29b41e063552c492c8eb82e94d056b89c7eab91870109c50ed67e

Observation 9284ef0c-32c4-4670-a428-b1302cb584d7 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.484538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.484538Z digest=sha256:842e788d9879445b4bb034a65ec40f887517b71316f5a806af7c7f996b4def5c

Observation cc85ddcb-6443-4253-b2a3-3e8b7b6ba01f · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Align your latents: High-resolution video synthesis with latent diffusion models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.489686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.489686Z digest=sha256:7292e1f1fa1adfd2652471af0f80d7c691789b46df3b8dea372d64a407252a8e

Observation 1f26b371-e4cc-4f7a-86b9-d59785843cf7 · outbound

This paper cites Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.908486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.494623Z digest=sha256:3a7790ca4362a74352cb212b983dc528b8065acbd573ae104f838a9306be9165

Observation 59bc1fd1-ad4f-419a-a1fc-89303724cb57 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Emerging properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.499471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.499471Z digest=sha256:ece834693ee402ec8cd8acf06e18e8df083c04ffa149c7d5ae4f4f373791d914

Observation 2d63673b-2c57-4917-807a-535f1073b804 · outbound

This paper cites Pix2video: Video editing using image diffusion.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Pix2video: Video editing using image diffusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.882431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.504915Z digest=sha256:2a751bc23c71da1f2f8f36a2ee4cfb50e29c70363fc8bd2cbcd364e0b9c71cd3

Observation 5e465f6e-f97a-42ce-a0a6-b3eeca7a5784 · outbound

This paper cites Everybody dance now.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Everybody dance now

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.866637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.510099Z digest=sha256:0e8c4927d3e072320b770ab5ce01f0257db9e940eb54754ee079c76462668d34

Observation 3f060444-2005-487e-ab23-b1aa34d35a10 · outbound

This paper cites MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.514833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.514833Z digest=sha256:e1cc167cfce9443c62a344ae2e34c61a84f633103417199161821f04230dae59

Observation bb6a13b1-01ba-4697-9709-12e96132e3fd · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion AnyDoor: Zero-shot Object-level Image Customization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.520283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.520283Z digest=sha256:2c46f9ed15254a8e008ebe16d6c3db4888b889d7c3b5ed1ee36ffd10c9a5abe9

Observation 9d1bc5b7-2916-4ab2-9cb2-551fbda36462 · outbound

This paper cites Viton-hd: High-resolution virtual try-on via misalignment-aware normalization.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Viton-hd: High-resolution virtual try-on via misalignment-aware normalization

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.852005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.525417Z digest=sha256:ee338b88d40d76a34693039ab2a82f52c9476d4141c70dae4ea2d04100e11144

Observation e59961be-d43d-4767-8a46-756c8cf9943c · outbound

This paper cites Improving Diffusion Models for Authentic Virtual Try-on in the Wild.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Improving Diffusion Models for Authentic Virtual Try-on in the Wild

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.530350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.530350Z digest=sha256:9c91481138e5fbd8cb0b9c09f6b8c314e896d2a27cd39682310bf2a781ad4a52

Observation 33df9dfc-58cc-4215-92d3-ff04f2c1a4af · outbound

This paper cites Diffusion models beat gans on image synthesis.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Diffusion models beat gans on image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.535753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.535753Z digest=sha256:d06688db86b9340edee34ca4dc75bc0cc1299c89e517ce85e8147f00324e943b

Observation febfb330-8dd7-4f7b-ac35-33629bbfa28e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Scaling rectified flow transformers for high-resolution image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.540253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.540253Z digest=sha256:620e9abb8739d8f46843ba11a22fc255176b63b9088d81e5d0c6c6b3ca7d757c

Observation 920ae04d-f9bf-4923-87c7-4f4490c4943a · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.544680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.544680Z digest=sha256:f6d1a7787a55ad0bb09821130f59ea0dc8b996f4fb6ccad0388cefe1fdeb5b25

Observation 247acfed-e251-4b67-a595-f25f4b98737d · outbound

This paper cites Parser-free virtual try-on via distilling appearance flows.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Parser-free virtual try-on via distilling appearance flows

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.817264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.549676Z digest=sha256:f3045bcd2d507c2a69ace3228512ad2bd5d83a4a8399563d08f65bb4cd2c4380

Observation 3c84103f-ab7f-4212-ba04-e4eb6c7b62cf · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.554355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.554355Z digest=sha256:212da00d36ac26ebfdcbc0c0f183fe241cefb18c7bb017f2182b5938e8e99b7e

Observation 4a382498-32f8-4547-ade3-c8d14f7f5b31 · outbound

This paper cites Taming the power of diffusion models for high-quality virtual try-on with appearance flow.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Taming the power of diffusion models for high-quality virtual try-on with appearance flow

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.802187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.559452Z digest=sha256:da553e7bad559c9bc8e86f6208b6fa770b7e4848c88dd1adbaa153cd790fb45c

Observation 495cf92f-cb20-4126-a7a2-3faa42e65290 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.563920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.563920Z digest=sha256:449b169381449773af9ad3bfc29806c19394131ba61447b056e1a5e5f215b0d8

Observation d4a7cb1f-2c3b-473b-aa1c-a495f6db01d7 · outbound

This paper cites Viton: An image-based virtual try-on network.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Viton: An image-based virtual try-on network

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.786630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.569603Z digest=sha256:3418dc16058b7c8dcb2b5edc7cb5fd49eb251caa37d84c635b4b8cb4fa37baa5

Observation 8f3f0f78-ccf0-4bfe-8d29-37bae6c4152a · outbound

This paper cites Deep residual learning for image recognition.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Deep residual learning for image recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.574374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.574374Z digest=sha256:94917bce3895b5a9becfc643d0f7cad3ca84bd2b816c744cdd3f09ddb0aa81d5

Observation 15e4f354-1d7c-4f4a-9c5a-bddcca664c73 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.579033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.579033Z digest=sha256:52bb35b66180dfd074c2b5d59611fa09208bb481e892ffa95e8f78c846109b2d

Observation 78893b7b-5877-4fb2-8663-04698922644d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.583700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.583700Z digest=sha256:362ea7e585da303df01dddc384eba8c3daf246c7ffaeba3cd0f58c1ffb1e0d9b

Observation 893d594a-69c2-4dca-af3d-84a146890037 · outbound

This paper cites Denoising diffusion probabilistic models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Denoising diffusion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.588391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.588391Z digest=sha256:e4b99e5c074a31955a2bcb1d86aa194c0c2e003d093fcb88d24f445a76dc1535

Observation 6285f522-8412-4009-bdc6-220567277c30 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Imagen Video: High Definition Video Generation with Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.593305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.593305Z digest=sha256:2280d2bf7fdb49015ca686accbc1fd2009909a2391436e0845f75647feaf4392

Observation e530495d-3c5d-4eaf-a767-943909f66949 · outbound

This paper cites Video diffusion models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Video diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.741472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.598657Z digest=sha256:65b793d12441fbdff555e051e62a535348c6327bbb733487b87dfec144c4a9aa

Observation 6a7a033d-01d3-42c5-b50d-fa2b06a9d8bf · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.603149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.603149Z digest=sha256:379d69225ab364ddf9ebc3041d7b5b0057126605602ae8f4f3a71049beb1cf47

Observation 0415a9bb-8845-4659-85d5-17525280c221 · outbound

This paper cites Image quality metrics: Psnr vs.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Image quality metrics: Psnr vs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.608066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.608066Z digest=sha256:8ceeb723a37e0abd4cd702d3f909df69bf9d2e45532d27a59026b7ac280c755b

Observation cdd3e95c-ba26-4816-a8cf-9e7880d0b6e7 · outbound

This paper cites Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.612552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.612552Z digest=sha256:21d7b93c5453d45b638fad7da96ea42a6fc1e6cae1ea1cd9b2242ea45b237625

Observation 507dae02-1804-4a8e-b978-955e2121e947 · outbound

This paper cites Composer: Creative and Controllable Image Synthesis with Composable Conditions.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Composer: Creative and Controllable Image Synthesis with Composable Conditions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.617425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.617425Z digest=sha256:5f7c3b68d2bab58f5a69cb579e5201c575152bf74ce10262e910b3401681e6f4

Observation fe8da206-8d16-472a-9e73-f26f2fefa1d0 · outbound

This paper cites Learning high fidelity depths of dressed humans by watching social media dance videos.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Learning high fidelity depths of dressed humans by watching social media dance videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.716017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.623296Z digest=sha256:903f6daf27549e6bd2895f68d0b7eaafacc658b4ff72ff918ac1a2000d627e2d

Observation ce5a20c5-a804-4e23-b3ad-2282e97de88e · outbound

This paper cites Dreampose: Fashion video synthesis with stable diffusion.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Dreampose: Fashion video synthesis with stable diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.628294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.628294Z digest=sha256:97ef9ffb52c6b41dada4c2fac4990b3b86c019df86248404f212213b74f944e0

Observation 54352305-1cb9-4753-9ec1-823a7b9e7791 · outbound

This paper cites Text2video-zero: Text-to-image diffusion models are zero-shot video generators.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Text2video-zero: Text-to-image diffusion models are zero-shot video generators

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.691538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.632961Z digest=sha256:41c48edebad6ed12b0acbbe5b851316c3c5864f628f4970669b0d31422a3e0d3

Observation 91fc68b5-3002-4b29-89ed-5ca9d9a15a1c · outbound

This paper cites Stableviton: Learning semantic correspondence with latent diffusion model for virtual try-on, 2023.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Stableviton: Learning semantic correspondence with latent diffusion model for virtual try-on, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.638109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.638109Z digest=sha256:09c4f0dba5efe47a67a726adc43fcf875e7dfdcf3de7590956c4c200c24f95b0

Observation 6984ef33-bcd5-4397-a105-361951b5b224 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Adam: A Method for Stochastic Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.642602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.642602Z digest=sha256:8a8f3f0f072455b87672866b2507f7afba436a5fe74bf3512b369a2581e6f615

Observation a7d9edc9-c832-4005-a374-ff8164259c21 · outbound

This paper cites Video-P2P: Video Editing with Cross-attention Control.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Video-P2P: Video Editing with Cross-attention Control

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.647242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.647242Z digest=sha256:071486521eb83cd49db3a11a3754b01a29ade402c5015c7e018b01526911c559

Observation 39a68eda-6480-4580-9ce2-0cb4580ed316 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.651926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.651926Z digest=sha256:732fb15c07d18bfc591967bf9899df5ce618078d4e6080178fd78a7b20cc01b3

Observation 77b979fb-1be9-4432-976b-67252a0a1805 · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.656352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.656352Z digest=sha256:4fac4dd22ca508e76223415aefa0f4558b75d8a64167cea36481189c5d628716

Observation 878e0f60-c6cc-4474-be09-fa9d874ebe6a · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.661285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.661285Z digest=sha256:208d5e02d8c259b820d33ca5040395ddb01c137682a456fa488a04c034a41c49

Observation 82382d19-4df6-4cb4-8068-8065770118be · outbound

This paper cites Improved denoising diffusion probabilistic models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Improved denoising diffusion probabilistic models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.666046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.666046Z digest=sha256:26280185c4c363cde08793b589dbd4dc064278233c7c07445d1b93c685f0433d

Observation 34ba8370-e297-4687-a611-6c12ea3d6726 · outbound

This paper cites Scalable diffusion models with transformers.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Scalable diffusion models with transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.670382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.670382Z digest=sha256:1b6856cf6befcb8648d4ac3ac56c8a1a333d308cacefeec1bd624caee2c43aa8

Observation 81ecc2d2-5f1b-499e-85ec-ea6a9cba16e2 · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Fatezero: Fusing attentions for zero-shot text-based video editing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.636770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.674645Z digest=sha256:978d3565f601ff864d77e41f355b1d9a5994ae4164fd6bf4a342e50284b99255

Observation 0da1412c-edf4-41b4-bf48-e13b74b7c3bf · outbound

This paper cites Learning transferable visual models from natural language supervision.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Learning transferable visual models from natural language supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.679420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.679420Z digest=sha256:489370e857d422742fb5f5f6e8819cb8c2bdacbc7d15f0913b76e36b789bed7d

Observation b99669c5-297c-4ab3-a89e-d473e1c94238 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.683695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.683695Z digest=sha256:72ebcdbd546c4c2df905c8ae50ec695785357a78f7513d635ed4bba2bcaa270a

Observation e43fd788-83f5-4a2e-8542-489ada561da5 · outbound

This paper cites Deep spatial transformation for pose-guided person image generation and animation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Deep spatial transformation for pose-guided person image generation and animation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.612197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.688340Z digest=sha256:a0d2940ab7ca6edc94201ad62a979a2fc406f71be3ec0c8aa7022f11f5908d33

Observation 8c5cf574-b9b5-41c2-be72-539878c5b752 · outbound

This paper cites Neural texture extraction and distribution for controllable person image synthesis.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Neural texture extraction and distribution for controllable person image synthesis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.598115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.692808Z digest=sha256:02213a610c15e2f1bdf80be0b3f53d73ac35ecf80d2db70fbb9e45fb0ac5d9a2

Observation 98839c3e-314e-46c4-a760-e6343dcf86b8 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion High-resolution image synthesis with latent diffusion models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.698600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.698600Z digest=sha256:130f3898b98a53ba478b69cfd2cc191c6ad45eadc349354aaa6a37767570ae02

Observation c00f82e4-5c13-48cc-9283-eb10903ee149 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion U-net: Convolutional networks for biomedical image segmentation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.704068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.704068Z digest=sha256:db9fc5078f3157faaf3d1d9a01878d630f89f853538030bcadabee74a178618c

Observation d7fd9756-8343-41cd-b307-cdbcf5e2c030 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.708942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.708942Z digest=sha256:84f3b1013736816a9f8093ca68cf2ae5cff506bdbbb5ec7af177758557e109ae

Observation 28105dc5-41b0-4fd7-9f48-4b6d421dfa47 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Photorealistic text-to-image diffusion models with deep language understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.713812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.713812Z digest=sha256:50d4ded0f64f1aaaaecb94c20383aa3879198b6a20d4524fd7689cc4497d6356

Observation fbcf799a-1d54-4a31-83dc-91dffd3f9066 · outbound

This paper cites First order motion model for image animation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion First order motion model for image animation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.722465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.722465Z digest=sha256:91752657e4ea74f358a9b0423ab35ce4ca756af3da8b721f42621869cfc112eb

Observation 979e70fe-843e-462d-9aeb-db6ee3ab1eb7 · outbound

This paper cites Motion representations for articulated animation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Motion representations for articulated animation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.534814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.730806Z digest=sha256:4ef9fe038167114143c02bacd9ae8e94428faf1938f28922a4f281b2eff1f249

Observation 90436190-86b3-4a4c-b451-f289d56d31ba · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.735222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.735222Z digest=sha256:17ee318edc7bbfcb07a86da89e5a6352a4dd927a8c9ebe983ee8383fd8aee0af

Observation 63c74cc8-d089-431b-90fd-ff93b7bca78c · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Deep unsupervised learning using nonequilibrium thermodynamics

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.740338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.740338Z digest=sha256:097970ca9da9fb9c505376e7275f0014683a3452150ce51b1e37bb74766801eb

Observation 798b14cf-d1f4-4380-8557-6ccf65eafef0 · outbound

This paper cites Denoising Diffusion Implicit Models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Denoising Diffusion Implicit Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.745035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.745035Z digest=sha256:a373b77af8c896664cad84488de6bab8ed5243b9485c4efe3078f1c65bf99bc1

Observation c2ea87c6-ef73-4ca9-956a-3adae20946d5 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Generative modeling by estimating gradients of the data distribution

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.749422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.749422Z digest=sha256:b04c08bd4c8bcdd6611ae17131bf8d724d1f9b66d2d97b9286d494bfe29fd313

Observation 47617989-b043-464c-9c12-94d9d3a03da5 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Score-Based Generative Modeling through Stochastic Differential Equations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.753568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.753568Z digest=sha256:f263003dccfb409521bb8b8cfd1d5fc071eb1d662470dc8df7e9028e7bc9934d

Observation f2de290f-4692-4ff9-adf3-49e0841c6bf0 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Plug-and-play diffusion features for text-driven image-to-image translation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.499443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.758257Z digest=sha256:86e0ea2768b8821309c7f7516e6ee63f54a5fbb5f49fe687e1d9eee96d925950

Observation f82cc421-4a6a-44f6-ad4c-a4bbe997f9e5 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.762919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.762919Z digest=sha256:c8df0ff937827bf4c46148473290f4d0c1626777551f5a60fdbf874582cda454

Observation 520b5c83-66fb-416a-93c0-2f2acee2def9 · outbound

This paper cites Attention is all you need.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Attention is all you need

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.767501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.767501Z digest=sha256:fd906afa1e28a4baa281fa6543a5e61142700562d399ff70b77d4972360fe4f6

Observation 1da7e844-c39b-41c5-9236-60e4d8c6d779 · outbound

This paper cites Disco: Disentangled control for referring human dance generation in real world.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Disco: Disentangled control for referring human dance generation in real world

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.474769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.771508Z digest=sha256:207a957d23ddf1e013771cae064bce3c3b0f0c163f1e4539a9d2bf9151814c00

Observation 1df3eb50-5478-487a-859e-7e2fdd2f591c · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Image quality assessment: from error visibility to structural similarity

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.775664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.775664Z digest=sha256:258a9d4595d4f70eddaaf202e17ed6b0aae2bd60d9f25089779c225a07ca9ff7

Observation dde073e1-dcca-46b6-9ce8-cce938358fe9 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.779895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.779895Z digest=sha256:85ce8e9e472d38e1f57809af5663682ab961824605bb386bb223af8ffe92f8f5

Observation cc8fff43-ae8f-4f7b-a5c1-a8079069c88b · outbound

This paper cites Gp-vton: Towards general purpose virtual try-on via collaborative local-flow global-parsing learning.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Gp-vton: Towards general purpose virtual try-on via collaborative local-flow global-parsing learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.440246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.783835Z digest=sha256:4b158d679cb5fec04608d531b842039fbefb038417420a22d7b2a08cf787da92

Observation 32802769-e48f-40e4-8d71-35848afc2341 · outbound

This paper cites OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.787677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.787677Z digest=sha256:e03f324435d41af9b34584cb8ba3b1292994acdc8ada97fa34dfdf55df26b15b

Observation 21957b3a-71d8-4df2-8cb6-e36e75c55c8b · outbound

This paper cites MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.792023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.792023Z digest=sha256:7014f99564a5ac98093be2ce20ef305b45b3b4c186397459865d7d1c56544a74

Observation b88cfc2a-420d-4882-9951-42d1d44b46a5 · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Paint by example: Exemplar-based image editing with diffusion models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.424739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.796678Z digest=sha256:14a9fd418932adf24bf8973fb9dbdc6697526a8e6bc14a6f9ad5c085b3f2bb34

Observation 04272eb9-d3c1-479f-b938-e19e3b4ad0ba · outbound

This paper cites Towards photo-realistic virtual try-on by adaptively generating-preserving image content.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Towards photo-realistic virtual try-on by adaptively generating-preserving image content

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.410149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.800617Z digest=sha256:82ce5e7374000cc892c8f8c72e468f9a83d69b2112845ff489f9a4051376b8c8

Observation 095834e5-fb0d-48ec-9eac-3becbc903bc7 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Rerender a video: Zero-shot text-guided video-to-video translation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.395631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.804656Z digest=sha256:bb893d01e2c1212cebdd5ab05a74a3798ef49d1fb4c5ca723329f87ce8df9f1b

Observation 51b9105c-ad59-4f89-b252-6b93de91467c · outbound

This paper cites DwNet: Dense warp-based network for pose-guided human video generation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion DwNet: Dense warp-based network for pose-guided human video generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.808616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.808616Z digest=sha256:f814904db0cc5da82a070b5e9a07687afc95d6fd68ea0ed138d7b0014f0c6432

Observation 4a12ff71-7252-439d-bfdb-d8c85ca9b1de · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Adding conditional control to text-to-image diffusion models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.813042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.813042Z digest=sha256:4d491dfd30ea7d7090832494e9ee86779c2009cd433c932761f81a814c4752f9

Observation 23b6bab8-b333-434c-b300-33f6f420d40a · outbound

This paper cites Exploring dual-task correlation for pose guided person image generation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Exploring dual-task correlation for pose guided person image generation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.370540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.817271Z digest=sha256:8790782237193acc9bc455032e76ba2dda753ffd2e0bad19c0cd78f99f01face

Observation 5a4ecd48-1c69-4ac7-9785-839c1dbf13fb · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion The unreasonable effectiveness of deep features as a perceptual metric

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.821441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.821441Z digest=sha256:725daf2a3376310ef284f97b419b3dac8c39bcfc2dd8e1d510d63e183caf87e6

Observation 49bc4dd5-06a3-4844-8b0c-213aa504b8e5 · outbound

This paper cites Thin-plate spline motion model for image animation.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Thin-plate spline motion model for image animation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.345517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.825942Z digest=sha256:57284c48ae76de166ea67ced8c60282d7c934ba26af81b47a0cafc6c2a79564b

Observation a2850227-b6f3-48eb-ada3-6542146a758e · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Uni-controlnet: All-in-one control to text-to-image diffusion models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.830349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.830349Z digest=sha256:74aaddc6d9ab587ddac37aeaa6c04ada10a06ceb89e89824d8823077c2c11115

Observation a58c1b5a-d90c-4243-8bb2-b6c5add81572 · outbound

This paper cites Tryondiffusion: A tale of two unets.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Tryondiffusion: A tale of two unets

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:12:58.319981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T12:12:57.834736Z digest=sha256:6448fe01a64f4c936102aa44a7ae00743c917ce8da42bebfb102b8cdf03e46be

Observation 03b7daaa-e2e1-4c9f-bdb2-77230aa35ede · outbound

This paper cites Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.838913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.838913Z digest=sha256:60f94055b9585c480ccd4daee503d016351dae9178df570b08db6e5c2d5afb99

Observation 2a2c51e2-e288-4626-b14f-3e5bdb5b7dfe · outbound

This paper cites @esa (Ref.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion @esa (Ref

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.843378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.843378Z digest=sha256:998eb7a080e87413ca67d985c6b82e70317741ea1b4fa46d21d9be1809baa824

Observation 64206801-0635-4b0b-b049-fa65514305ab · outbound

This paper cites an unresolved cited work.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.847782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.847782Z digest=sha256:99a54d66456dc5bb6c92c4803f5c37093e5a5d7d11b6f62aacbff1b46bca40fe

Observation d581d0bb-41f0-4116-b867-6573d15a43c4 · outbound

This paper cites an unresolved cited work.

Consistent Human Image and Video Generation with Spatially Conditioned Diffusion Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T12:12:57.852352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:12:57.852352Z digest=sha256:2eebc8005154597abd7ae1039f4ea572be4079a7ed81f3146731ec9b147e21c5

Pith citing papers

Observation 0f079b54-9d7d-445d-87a1-b775d3881250 · inbound

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models cites this paper.

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models Consistent Human Image and Video Generation with Spatially Conditioned Diffusion

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:01:15.969353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:01:07.730833Z digest=sha256:9a1922de6c12b79cd0dae9aab052c5bd7d286cc9dd0fed0b29549e8b0a796b08