Pith. sign in

Paper Citation Record · LEDGER

Populate-A-Scene: Affordance-Aware Human Video Generation

As of 12 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2507.00334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00334 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:23:12.762375Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2c00502-cd33-45ef-9bfe-2ffe5bd25e96 · outbound

This paper cites Fouhey, Ivan Laptev, Josef Sivic, Abhinav Gupta, and Alexei A.

Populate-A-Scene: Affordance-Aware Human Video Generation Fouhey, Ivan Laptev, Josef Sivic, Abhinav Gupta, and Alexei A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:13.063341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T21:23:12.687060Z digest=sha256:0a6ae98b7100196550a3b31155dcd3520ef8ef547ecc31c980abb248f8d035cc

Observation 49c61981-868b-4297-82cc-18e96dff7a58 · outbound

This paper cites Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.697903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.697903Z digest=sha256:4902cc6ebb95d0d7bba3598dd0649bb9cedefb140ea6edd5db92914295afd843

Observation c25ede42-3767-4ac1-8629-05fa90181633 · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

Populate-A-Scene: Affordance-Aware Human Video Generation Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.708601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.708601Z digest=sha256:d9f95c7007e5467821d284ef1f158efb94093eb0ec0a7d32b90e6cbf3102610f

Observation f64611c1-989b-4065-a6af-33dbb54fbe1c · outbound

This paper cites DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.712364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.712364Z digest=sha256:0d1a8ce75c3977fe580de92f4df9ba9e2df279da2a69f73378d1b28ae61e6ecb

Observation 62a50608-7b31-4e50-a1fd-98ee4c5747bd · outbound

This paper cites Flow Matching for Generative Modeling.

Populate-A-Scene: Affordance-Aware Human Video Generation Flow Matching for Generative Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.721211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.721211Z digest=sha256:f2e5b6231409bd204efa3efbb8a4e4e7b366293725c3bdfcb0a3d37a07df93f4

Observation 257b1446-f133-4fbe-9d3a-a6d3b4b6b768 · outbound

This paper cites Synthesizing Environment-Specific People in Photographs.

Populate-A-Scene: Affordance-Aware Human Video Generation Synthesizing Environment-Specific People in Photographs

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:23:12.941665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T21:23:12.728400Z digest=sha256:1ec9c9297594acd653df65ada1f6daa5d69d918d0d9341cb7edb83fd7fe7ea36

Observation 93484021-023b-4bf8-b5db-e9eec2f6ad75 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.731808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.731808Z digest=sha256:3f95da7e2deba390c00c98399e049fa77f52effa4ecab0ac7728f8b21a31d8c9

Observation a2aed091-cead-4adc-91d6-3e620802ca67 · outbound

This paper cites ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.735107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.735107Z digest=sha256:5b100334c8dc4bb26ede5f15ba2b6b14c58f3c6587e04a56a51137bce6b15334

Observation 94434dd0-0ad6-458c-94ec-7aad763aed84 · outbound

This paper cites InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.738455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.738455Z digest=sha256:fcc47261be8a4d778ab122763a06e5d20f1eb826f0c34e240afe33987cbd78ba

Observation 75d06842-c7b2-4987-8b31-20e8e90525cf · outbound

This paper cites Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman.

Populate-A-Scene: Affordance-Aware Human Video Generation Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.741893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.741893Z digest=sha256:a8fe1318d372d3ed78a766e1643c340129045b3217d94ba4b951a5eb0ca6a15f

Observation 4d6b4ed2-1c92-4488-b1ec-a1dd95fee2c0 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

Populate-A-Scene: Affordance-Aware Human Video Generation UL2: Unifying Language Learning Paradigms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.745181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.745181Z digest=sha256:6540a51e4a7a838b26d84bdd6d957b1b33cfa12c27c22b4f88f91e07b429ad4e

Observation 0ab42c9e-6dae-4de6-bf30-21f072092da2 · outbound

This paper cites DAT++: Spatially Dynamic Vision Transformer with Deformable Attention.

Populate-A-Scene: Affordance-Aware Human Video Generation DAT++: Spatially Dynamic Vision Transformer with Deformable Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.752065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.752065Z digest=sha256:4ba643a8531743b44b2d47a7ad9a919467a99a486966bb5d0869fedbef038122

Observation 9cea3e29-550c-4d6b-8a8e-b62f6ee3c1b6 · outbound

This paper cites AMG: Avatar Motion Guided Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation AMG: Avatar Motion Guided Video Generation

Reference 24

Resolution
malformed identifier
local_arxiv, observed 2026-08-06T21:23:12.814830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T21:23:12.756044Z digest=sha256:fcd512c554efa267b555da424d74d6da94c9fd5c8d510c41b8bcf26bd0fb9abf

Observation afc3823d-ff8f-4963-a9ce-6039fa11997e · outbound

This paper cites Make Pixels Dance: High-Dynamic Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation Make Pixels Dance: High-Dynamic Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.759230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.759230Z digest=sha256:3869ef859f75c9ce14d60513c194fbd10bb8b04d16ce1ef4de9e7d336757e099

Observation ceb6a3c0-c822-45b2-a3fc-e5bdec43865e · outbound

This paper cites Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu.

Populate-A-Scene: Affordance-Aware Human Video Generation Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:13.054043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T21:23:12.762375Z digest=sha256:199501619a6a2d8b5223d2aefee6ca91bc1d797a01a32cd59ad680c3779df55f

Observation e35b201a-a990-4add-a86a-7cc744d679ba · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Populate-A-Scene: Affordance-Aware Human Video Generation AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.716870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.716870Z digest=sha256:62514768b234c472422812c663ad49d265f4a7d70d8dda6bef79a031a71b8bfd

Observation 1909a206-97fe-472a-ac8f-167bc2ca06ff · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Photorealistic Video Generation with Diffusion Models

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.701371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.701371Z digest=sha256:49c1367055572036af1df29159853c3f2b4d66fe96e11d2c7f777f876f722b2a

Observation 620ed033-fd70-4614-88da-2ba92398777f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Populate-A-Scene: Affordance-Aware Human Video Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.690412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.690412Z digest=sha256:d097c87c41c601ba56d814220e5b4bab66a0c351642aab2c0d642739d488044a

Observation b7dd9eae-fa44-4463-9bda-7f23d84e4b93 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Populate-A-Scene: Affordance-Aware Human Video Generation ModelScope Text-to-Video Technical Report

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.748660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.748660Z digest=sha256:14b548af839f18a025014a23a691a824b67fec86681cea1f54e78a51ae082fd6

Observation 74a42b92-080c-404c-a25d-ed8b1516d290 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Populate-A-Scene: Affordance-Aware Human Video Generation Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.682484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.682484Z digest=sha256:062092372efdd38e9f8611df4a28295a77a39b59263226b26e5c89cc5ccdece6

Observation 980abd31-064a-4b46-bc7d-3dfcd4079937 · outbound

This paper cites The Llama 3 Herd of Models.

Populate-A-Scene: Affordance-Aware Human Video Generation The Llama 3 Herd of Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.693914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.693914Z digest=sha256:2ca40465bf8e578df0a7a5876cf689687c73ccf699357af70c5e6e9f6996443e

Observation 8d11e3cc-a032-4197-bf32-4e8ae0c94b9a · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.725035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.725035Z digest=sha256:27f48819738e8a8352e15e47158090ae4e775f0a03de229e3f39dbc250498a48

Observation 3dbc78da-5f3a-41b8-b3be-4d77d77789c1 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.705127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.705127Z digest=sha256:5126415d7e26c38b24b096ef0949a1815c8d0f26197b2f4f11f434f98362b715

Observation ffc8f9e7-cb9e-492a-af0e-0bccb47cd800 · outbound

This paper cites Flow map matching with stochastic interpolants: A mathematical framework for consistency models.

Populate-A-Scene: Affordance-Aware Human Video Generation Flow map matching with stochastic interpolants: A mathematical framework for consistency models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.674110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.674110Z digest=sha256:49b9910b6b6b2e3e6207c82dd289132c72c15118f8027e5648b41568798da67c

Observation 38463b39-f0da-4dd0-bc6e-d479b6d341fd · outbound

This paper cites an unresolved cited work.

Populate-A-Scene: Affordance-Aware Human Video Generation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:13.072739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T21:23:12.669420Z digest=sha256:1eea01df617fd703501222dc307a951022dc97d2cc05311feac098bc200f6c5a

Observation c9db72b1-ff8b-4e03-9fa0-216c54c3cc74 · outbound

This paper cites doi: 10.1145/3715140.https://doi.org/10.1145/3715140.

Populate-A-Scene: Affordance-Aware Human Video Generation doi: 10.1145/3715140.https://doi.org/10.1145/3715140

Reference 2025

Resolution
verified exact
doi, observed 2026-08-06T21:23:12.791955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T21:23:12.678507Z digest=sha256:872b00ae03033cf8b780fae06ed704c2ee9e5184399a7351b888bc68e7133623

Pith citing papers

No inbound Pith citation observations are available.