Pith. sign in

Paper Citation Record · LEDGER

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

As of 19 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 4 inbound Pith citation observations for arXiv:2506.12847.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12847 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:41:07.918981Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:23:09.117957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy67
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d78cc45-a367-4c45-9b1d-ed589175f3ab · outbound

This paper cites Person image synthesis via de- noising diffusion model.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Person image synthesis via de- noising diffusion model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:55.849134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:55.849134Z digest=sha256:0797bd1951afb6a4d8fccc2c7ec90f5adefe219de03de1bbaa39710c222be498

Observation 87e4ccb9-97a3-4d1d-88c0-96a38758b037 · outbound

This paper cites Videopainter: Any- length video inpainting and editing with plug-and-play con- text control.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Videopainter: Any- length video inpainting and editing with plug-and-play con- text control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.026603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.026603Z digest=sha256:a937ef030e13fcdb4c4020c512f7bdf8c41e603244f7f97f3565290a07b67b93

Observation 3c87dd3b-a552-442f-a769-09586745fb9e · outbound

This paper cites Smpler-x: Scaling up expressive human pose and shape estimation.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Smpler-x: Scaling up expressive human pose and shape estimation.NeurIPS, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.147395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.147395Z digest=sha256:61a3617aa87dc8a8024bb76c89df67b1bdfc8ccbee97473f59e1623fd34f97ab

Observation b7b6547e-4d76-439d-8174-ab27619128e3 · outbound

This paper cites Pix2video: Video editing using image diffusion.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Pix2video: Video editing using image diffusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.276084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.276084Z digest=sha256:aca884b7441a45023e9b41ae97c6bead60fdde1fef26e0544036d0d1aa5f991f

Observation 3c527140-e1f0-462f-936b-7748adb26f9f · outbound

This paper cites Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.407455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.407455Z digest=sha256:a019ec0a687bc423d1f7ad8d4dc2270e50b5763fea0ba0475150a4a0551e733d

Observation d0cd001b-8d41-4423-92d4-90dcb07eaa0c · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.552104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.552104Z digest=sha256:77586ed6bef26247b9420ed7f7b7c8e43aff5b747fa406d41e6cdc09834e5ec0

Observation 8eec9d24-25f8-4151-9226-3683a7196b86 · outbound

This paper cites Control-a-video: Controllable text-to-video generation with diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Control-a-video: Controllable text-to-video generation with diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.697928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.697928Z digest=sha256:27716544cdfadb080218751a80bba154e77d58cd19f542a14e2e5c82796395f8

Observation d4bfc890-7a8e-4d38-a338-c94ad7667d51 · outbound

This paper cites Cg-hoi: Contact-guided 3d human-object interaction generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Cg-hoi: Contact-guided 3d human-object interaction generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.736762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:56.763263Z digest=sha256:19a4bbed7547cf5f5f7275622a4b468b5a7de29412618fb1094edf08b5cf71e5

Observation dd4570e3-0316-4bbc-8397-25a22af4e0c0 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Structure and content-guided video synthesis with diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.717139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:56.841966Z digest=sha256:703e05d83a5c4d9741f7c071279994eb4645877ff546f3a594ee34149e609fa0

Observation 28603426-b44b-423b-825c-2a8a5bbee419 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:57.009521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:57.009521Z digest=sha256:6b201b6068fc46dea0e316c9544c8d8f6dd34a30d386ec47f7a13152d2c9be4c

Observation 946c9b23-1f5d-4fc3-805c-9019e4e3eaf0 · outbound

This paper cites Re-hold: Video hand object interaction reen- actment via adaptive layout-instructed diffusion model.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Re-hold: Video hand object interaction reen- actment via adaptive layout-instructed diffusion model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.682177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:57.115303Z digest=sha256:fa1c30bd631b9af4f585969e43699961daef1e1653cba9f93a0e8da62522d5b2

Observation 444ee6dc-b79b-48c1-b1ee-f76fd6ff1325 · outbound

This paper cites Ccedit: Creative and controllable video editing via diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Ccedit: Creative and controllable video editing via diffusion models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.665051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:57.204044Z digest=sha256:ae32c5c6287c3310f268e43e164f2e72f3f6a3803d9f61558d81abc1b0f86e79

Observation 7ac5ae84-74b1-4c0d-8aae-4c50a97ad4c1 · outbound

This paper cites HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:57.359048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:57.359048Z digest=sha256:ffb99172a95266dd342e6a3d54213c5dd12b6731b1676d0b65848f6bc13706a6

Observation 375e4770-abf2-456f-8dc5-8c9e8339deaa · outbound

This paper cites Tokenflow: Consistent diffusion features for consistent video editing.arXiv, 2023.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Tokenflow: Consistent diffusion features for consistent video editing.arXiv, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.646670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:57.530289Z digest=sha256:54a5868469f8a8d86c941349b2b1eee68aba36c4a0afd388e953f94ab9532c6e

Observation 891ee0a2-a6b1-4f00-96c4-38ad8a68f67b · outbound

This paper cites Emu video: Factoriz- ing text-to-video generation by explicit image conditioning.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Emu video: Factoriz- ing text-to-video generation by explicit image conditioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.626921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:57.712846Z digest=sha256:82b4e9a12c09224f87b12e17189cc056381bfbc54fa76f7472c77ca0f0722c89

Observation a74efc51-aa1f-442b-b944-fa215e95add4 · outbound

This paper cites Videoswap: Customized video subject swapping with interactive seman- tic point correspondence.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Videoswap: Customized video subject swapping with interactive seman- tic point correspondence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.606275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:57.888052Z digest=sha256:d208eb7479775a0eec187350fb325cabfe425eb2a19db760b2e413ce97d28cbe

Observation ba21eacd-fd29-4ff8-b905-621052d961e4 · outbound

This paper cites Talk-act: Enhance textural- awareness for 2d speaking avatar reenactment with diffusion model.SIGGRAPH Asia, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Talk-act: Enhance textural- awareness for 2d speaking avatar reenactment with diffusion model.SIGGRAPH Asia, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.588021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.033983Z digest=sha256:a40ac31ba6fe0cf516a7ae362ae69080ff4baa41588acb49e9e2fc822747f8ac

Observation 8a25d911-f505-423f-8ea9-2127cd6a0025 · outbound

This paper cites Livepor- trait: Efficient portrait animation with stitching and retarget- ing control.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Livepor- trait: Efficient portrait animation with stitching and retarget- ing control.arXiv, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.569441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.141288Z digest=sha256:87be190b02a690598f71bff55e2f3a8d9528b5038743d75e98ec581f8b7d2aab

Observation 6c07f26e-efc0-4834-a2f8-88d5c99d8dd6 · outbound

This paper cites Resolving 3d human pose ambigui- ties with 3d scene constraints.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Resolving 3d human pose ambigui- ties with 3d scene constraints

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.551542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.247226Z digest=sha256:d4f4578fa7b31bbe171adb42002997aa51752517abfc7d844ab7d16e508a49fe

Observation 7768c4ab-da41-48e1-a27a-a94e565f3f9d · outbound

This paper cites Populating 3d scenes by learning human-scene interaction.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Populating 3d scenes by learning human-scene interaction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.528089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.327537Z digest=sha256:23302b4f219ca46def1922ad46ac305a161fae3bb1411bfe102e26882a97311c

Observation f652b488-9eb5-4420-8e2a-b8c3201b3c65 · outbound

This paper cites Synthesizing physi- cal character-scene interactions.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Synthesizing physi- cal character-scene interactions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.443031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.419397Z digest=sha256:14b837c0cb1c3a9954fb4106828130d9846295dd553b42cab941c7da85ca53f8

Observation ee101561-8130-4d4f-b99d-e6d58a8674e8 · outbound

This paper cites Hand-object interaction image generation.NeurIPS,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hand-object interaction image generation.NeurIPS,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.272435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.505914Z digest=sha256:6bde14e18b71157cbbf092d5685de8d3289e370bd1cf3358c3e6ab5e0ddfee5d

Observation c58761c5-83cf-436a-97f4-c6afe99a70db · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.062117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.587901Z digest=sha256:77ad902d62dff7d24e195da1414ce3acb86ddb80af0c4d8cb7117252993f04dc

Observation 42031114-248d-4ed7-9e35-685b36367ebc · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:58.645841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:58.645841Z digest=sha256:e878186ad61bb257d0666d5c7404feb5fdbd6fa9c48a38756984347f942b70b8

Observation b1c244ba-1ebc-459c-ab1f-38d2afd9f69a · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer VACE: All-in-One Video Creation and Editing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:58.747297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:58.747297Z digest=sha256:bdf331dff78dc94c7a4a64627bb05a39787efc7c47a1faf5e400703906b4e59e

Observation 1b1427ab-33b7-4321-9ca6-48409fc65980 · outbound

This paper cites Dreampose: Fashion image-to-video synthesis via stable diffusion.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Dreampose: Fashion image-to-video synthesis via stable diffusion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.923114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:58.861552Z digest=sha256:4bec52e167a742e5cc7a9a1baa5bc98dd241df256a4121a341a9e2ceeb486a1f

Observation b7ea3f57-8e41-466c-861b-fa63dce0ba4b · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:58.985574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:58.985574Z digest=sha256:023ffd3112af480d0f7f7ce0a09a3fadb8ee34cbaa6cedce19556836579412f9

Observation 3f2c233f-983d-4237-8e3d-570a41636637 · outbound

This paper cites Anyv2v: A plug-and-play framework for any video- to-video editing tasks.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Anyv2v: A plug-and-play framework for any video- to-video editing tasks.arXiv, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.845182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.071513Z digest=sha256:6587eee59971f465d232119b7f1a308ab6138cb28d57b7fa9b4a656726ebdce4

Observation 5bffd73d-dcee-4468-8fb5-52609dddfbf6 · outbound

This paper cites Flux.https://github.com/ black-forest-labs/flux, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Flux.https://github.com/ black-forest-labs/flux, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.650595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.129725Z digest=sha256:4e98c2a0af0c5230b0bb565734ae181597650f80e7fa7b7052a17412c17db751

Observation 46857b48-d2cc-43fd-8860-b9fb80935fb5 · outbound

This paper cites Lego: Learning egocentric ac- tion frame generation via visual instruction tuning.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Lego: Learning egocentric ac- tion frame generation via visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.470230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.225938Z digest=sha256:b69368c25418a2aabc462ab5596e751d222f4231b7e597d5fb66e79fffb67389

Observation 5ac1839c-76f6-4fac-92e1-f41c9f028dd3 · outbound

This paper cites Shape-aware text-driven lay- ered video editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Shape-aware text-driven lay- ered video editing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.322349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.323533Z digest=sha256:ade4d2d34b3c2691c20b66738bebf4f33cf90f1d3f7eeb4ace7eadb5646d8479

Observation 5f05962b-104c-477a-aaed-7848f3fd0afb · outbound

This paper cites Generative om- nimatte: Learning to decompose video into layers.arXiv,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Generative om- nimatte: Learning to decompose video into layers.arXiv,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.230606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.406104Z digest=sha256:2acf1126895667f6093e66c23bda4dedc9ef087b3e78d2552274ae7e49c9b00d

Observation 20f1775c-f3c3-45ef-a6b6-5ea631e9b977 · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Vidtome: Video token merging for zero-shot video editing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.036408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.468695Z digest=sha256:a49822bb7633ad88da653816737d2120081d37492fb122544d306a04a9581c69

Observation 133a9f74-e380-4fa8-84b3-bb139730b507 · outbound

This paper cites Video generation from text.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Video generation from text

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.739371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.569258Z digest=sha256:77791cb9951badb179776b93f6ad43bf64ad590574aad9720f42718aeb3ae590

Observation 4158dac0-cec8-4cd3-bf99-534b47953d49 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:59.660776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:59.660776Z digest=sha256:0d9442be8a57560ec9faa0b6eba8f52f4154d1f0ca4a3a9e7c76c4e3a31646b9

Observation 237fda55-aa45-4e5b-acbc-6a79a7f5f4bf · outbound

This paper cites Flowvid: Taming imperfect opti- cal flows for consistent video-to-video synthesis.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Flowvid: Taming imperfect opti- cal flows for consistent video-to-video synthesis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.482657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.767048Z digest=sha256:20808c70c40b8282c08578ad46e8a3c2a8a046e88faa2b5b4157fd1e26c41979

Observation ccd59e8f-5f78-471c-b09d-8b6dd3b0cb1e · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Video-p2p: Video editing with cross-attention control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.274883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.852408Z digest=sha256:e4e0a99561acdd88c3b5cda71255beb591d494ed3fea5c20078993c3ec23dd75

Observation 3a6392ac-1bf7-4d7c-bc07-7e55bdc26d36 · outbound

This paper cites Iterative ensemble training with anti-gradient control for mitigating memorization in diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Iterative ensemble training with anti-gradient control for mitigating memorization in diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.035930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:40:59.952831Z digest=sha256:712ba7c52cb64c7937ba738613ff7a8c21690f664ee460a2b24f1d4c5ad95320

Observation 21c8291f-9cfd-47cd-a041-c7b69fe23212 · outbound

This paper cites Live speech por- traits: real-time photorealistic talking-head animation.TOG,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Live speech por- traits: real-time photorealistic talking-head animation.TOG,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.816820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:00.025488Z digest=sha256:38cb1fa5b29c0581e328e51bd6f8dc22c77af9bf9b4dace9f2181439f01985dd

Observation d59c705e-ab3b-49b6-9c01-69b88252813f · outbound

This paper cites Dreamactor-m1: Holistic, expressive and robust human image animation with hybrid guidance, 2025.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Dreamactor-m1: Holistic, expressive and robust human image animation with hybrid guidance, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.614113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:00.125115Z digest=sha256:a0c7a8ad49e85b034045577cd45da9c44f678b0353a500856420d2d439900f65

Observation e418d60b-fe81-41ae-b4c5-e3e8adca9598 · outbound

This paper cites ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:00.228866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:00.228866Z digest=sha256:ac00014bfc5a8b22a4ff0a9ece626d2692c7e95cf679185d624b06666766ee5e

Observation e5ab77f8-b365-4af5-b0b8-fbc15c25910f · outbound

This paper cites Mimo: Controllable character video synthesis with spatial decomposed modeling.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Mimo: Controllable character video synthesis with spatial decomposed modeling

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.304385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:00.689236Z digest=sha256:e29b4a464da936ce653cd796d74d1077f28bbbdfa7495520279ed8db61623195

Observation 4d4c7e16-5d70-4c67-8d23-969a26fe7c26 · outbound

This paper cites Sora: Creating video from text.https:// openai.com/sora/, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Sora: Creating video from text.https:// openai.com/sora/, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.006145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:02.570292Z digest=sha256:29fa72fc2138aae027d8089d25a2e1c4838d2b028e0ef04e53e35054ba2e81ae

Observation ee9e51ab-dfc4-4569-8929-8ce1c90d0b96 · outbound

This paper cites Reconstruct- ing hands in 3d with transformers.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Reconstruct- ing hands in 3d with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.877718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:02.794703Z digest=sha256:e3d2a757bb4ad1166dd423c91ec041bb58ddc0413ce168428b0cdca1fae3a0b8

Observation 93e9b51d-433f-4654-a869-9ff43e1212b1 · outbound

This paper cites Vase: Object-centric appearance and shape manipulation of real videos.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Vase: Object-centric appearance and shape manipulation of real videos.arXiv, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.710627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:04.395040Z digest=sha256:e5e73a85cbae541a2362b4d5bb7eff55d6f2f46fdbcc918e3a0b55c8025d089d

Observation 04adb67f-2efb-4d33-af6a-4a884744d84c · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer A lip sync expert is all you need for speech to lip generation in the wild

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.536687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:04.643291Z digest=sha256:b3c07131c1fd839c54d088aaacbd42580eeb3edc9e55cf26be466a14fe394581

Observation f8e562a0-d805-4eed-97e3-6bea44ee846b · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.425726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:04.745650Z digest=sha256:50556fbb67fd5e2aec981c6f24c87def8e860430b68370afa01096de094d30df

Observation 6dfce370-292e-4c4c-805d-fa5807744f9f · outbound

This paper cites Stable diffu- sion 2 inpainting.https : / / huggingface.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Stable diffu- sion 2 inpainting.https : / / huggingface

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.274527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:04.868749Z digest=sha256:d67b02c714c0105549086d03ca7ccecbe0d452c3f8b037985de70cdb5ae39618

Observation db8cc9c2-7702-4c71-abc2-e3b9ec5bb039 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer High-resolution image syn- thesis with latent diffusion models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:04.923402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:04.923402Z digest=sha256:a4f107442d51d4bf7e56e597fa4fe135f3708a0ac878747cbd2537ab53b3e5ce

Observation 79bedfab-3e06-479d-90e1-5b4817113664 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer U-net: Convolutional networks for biomedical image segmentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:04.929104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:04.929104Z digest=sha256:4c398c7799acd69f0beb840e3915da7e3f1ffec02f99a27de9f8e08c02626900

Observation 6256689b-ee15-4038-8bda-8e49319670bb · outbound

This paper cites Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:04.992398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:04.992398Z digest=sha256:18ab34c9b4e7ae6ebf1e97cd56020fb35d86684c5eab02adb1065b48b5d064dc

Observation 1531149b-5832-43c2-a3da-32aad6c2706d · outbound

This paper cites Edit-a-video: Single video editing with object-aware consistency.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Edit-a-video: Single video editing with object-aware consistency

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.118624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.054638Z digest=sha256:a10b6afd820fa8ca7f3cd4b3578044cff3c11d5a218442efe00682c8dd987aab

Observation 76a50cca-916c-46f7-8a5d-960b83b9e323 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.arXiv, 2022.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Make-a-video: Text-to-video generation without text-video data.arXiv, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.967877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.127368Z digest=sha256:016749146a69efe6d4685df652e4529e59cd8251f03cc0e9ee44c9048e82aa97

Observation 4fcbaeb9-bbc8-4323-936a-98225a940114 · outbound

This paper cites Emo: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Emo: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.844013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.215658Z digest=sha256:5845df24892bab6112c27930699f0d28fc1e962e8f994f8ef8dec5b5dab032cf

Observation 1833d2a1-9209-449d-a166-9147a385e97e · outbound

This paper cites Attention is all you need.NeurIPS, 2017.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Attention is all you need.NeurIPS, 2017

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.725532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.282748Z digest=sha256:975a639478c6e2e06794d2a036af631d8d6acbacf281b6fc80c8fe1773a4a54b

Observation 84e4ab52-6e9c-41ae-ae7f-98caa501a480 · outbound

This paper cites Wan: Open and advanced large-scale video gen- erative models, 2025.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Wan: Open and advanced large-scale video gen- erative models, 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.590511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.351828Z digest=sha256:0f154d2fadfb8eda1874d4f5217468d6b68f28adc7bcc13b047c2f155dfb1543

Observation fca04e0f-4ebc-4160-b035-cb2e592d659f · outbound

This paper cites Robust video portrait reenact- ment via personalized representation quantization.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Robust video portrait reenact- ment via personalized representation quantization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.445126Z digest=sha256:3f728d6dfc2b8b8f191066977c5faa26e27f3a3cd580fcc1cc1c53567aa9275e

Observation 37e5e87b-32f7-40a8-9e2f-4f006ad5689c · outbound

This paper cites Efficient video portrait reenactment via grid-based codebook.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Efficient video portrait reenactment via grid-based codebook

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.264151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.540143Z digest=sha256:35aebd0e421c19b17bbbe7da154f47cc66553811068581319015013f24ce2efd

Observation 54a117d9-1ed8-4d65-927d-c2d7e24f1d09 · outbound

This paper cites Diffusion in diffusion: Cyclic one-way diffusion for text-vision-conditioned generation.ICLR, 2023.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Diffusion in diffusion: Cyclic one-way diffusion for text-vision-conditioned generation.ICLR, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.133190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.609950Z digest=sha256:c662d8de3a64f2421b8788fdb3ebb2507eae677c382d361a6c0bc42ded08dfa0

Observation a34be035-95e6-4c89-b858-84db5d41886d · outbound

This paper cites Disco: Disentangled control for realistic human dance generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Disco: Disentangled control for realistic human dance generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.936533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.693536Z digest=sha256:1ec1d183cbc60d003ff514eecf6d336e0566c27309843f29758561ebcd40174e

Observation 3fbab669-8ebb-455d-8e72-fd981d2c52ff · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferenc- ing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer One-shot free-view neural talking-head synthesis for video conferenc- ing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.730327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.791022Z digest=sha256:19a0e32153aebb87f76a112b3f1978d2b5918558233a65cb2935e4eefd9d749d

Observation 9b960a2d-fb93-481d-845f-7143d44f64e8 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Videocomposer: Compositional video synthesis with motion controllability.NeurIPS, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.534289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:05.885683Z digest=sha256:38ddb65db081670bf3ba3d647d090b3920c216ccdeb860febe00a362d290e95a

Observation 98c38943-a261-4899-8dcb-9822c011902e · outbound

This paper cites Unianimate-dit: Human image animation with large-scale video diffusion transformer.arXiv preprint arXiv:2504.11289, 2025.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Unianimate-dit: Human image animation with large-scale video diffusion transformer.arXiv preprint arXiv:2504.11289, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:05.956615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:05.956615Z digest=sha256:051816c01a6d5e69353e601d61319b49e6c59753e1f16f7217a62dcb38e56bf8

Observation f9d141e7-259d-49fe-94be-f4cdaf127bd8 · outbound

This paper cites Latent image animator: Learning to animate im- ages via latent space navigation.arXiv, 2022.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Latent image animator: Learning to animate im- ages via latent space navigation.arXiv, 2022

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.358387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.058039Z digest=sha256:f0b1d36e5babebc3df459875198a8424fcfffd7d808f442037b71bc3002a4e3b

Observation 88f8bc3d-fdd1-41cb-abf3-75b1f792f7bd · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.187867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.158274Z digest=sha256:7a57f8d8c32c4d9aad0f6c950906e925404b8c0611f2122f126dd96e6752951b

Observation 93e611b8-95ff-4542-a080-9dbed1fe7a5d · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.988082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.227531Z digest=sha256:e82bad5587252ccc545cd4261fb95675f145cfb10a6b20e1bdbb52333ba756bf

Observation d884100c-e331-4ee7-b609-eb18e98646e3 · outbound

This paper cites Hallo: Hierarchical audio-driven visual synthesis for portrait image animation.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hallo: Hierarchical audio-driven visual synthesis for portrait image animation.arXiv, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.830962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.307583Z digest=sha256:fdd024ad1f0d91a8c968c4242230416440a9d7a610602f35c6466b56e242aacd

Observation 3ea10905-0ca3-4cd4-a13b-ac1908019815 · outbound

This paper cites Hoi-swap: Swapping objects in videos with hand-object in- teraction awareness.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hoi-swap: Swapping objects in videos with hand-object in- teraction awareness.NeurIPS, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.670573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.377089Z digest=sha256:0fd15797b838d7737cc4ced502324cfa8a3ddc009dd537bdee47e73874408c99

Observation 180c060e-094b-4a4b-a15d-ad6630d416b7 · outbound

This paper cites Showmaker: Creating high-fidelity 2d human video via fine-grained diffusion mod- eling.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Showmaker: Creating high-fidelity 2d human video via fine-grained diffusion mod- eling.NeurIPS, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.431258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.445388Z digest=sha256:b8bc41b81f05065bd007f317323ec0903ce5291cf183254e9f141357b8c62330

Observation c6a856d5-2818-4056-95cd-7c779519471e · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Rerender a video: Zero-shot text-guided video-to-video translation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.304565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.520664Z digest=sha256:cc880705a30d73a489d07128ae6fe12fcc60ae1e87594a3eba39a1a21e20ca8e

Observation ffe09eaf-ae67-474f-a342-7cac1ae98853 · outbound

This paper cites Effec- tive whole-body pose estimation with two-stages distillation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Effec- tive whole-body pose estimation with two-stages distillation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.126525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.594525Z digest=sha256:a0fce787ae481529143569c41726206db2482e27e9d6958c64d9c3c70b60b7a0

Observation dcc5f97e-bee4-4cb1-983c-c7613a11b563 · outbound

This paper cites Space-time diffusion features for zero-shot text-driven motion transfer.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Space-time diffusion features for zero-shot text-driven motion transfer

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.950453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.674585Z digest=sha256:bc4095f4a43b6f0a388968f6adead432980d962592ea6b4b705afa993a7d0a1f

Observation b47d8aa3-a7b1-4180-8aa7-15ad68ed279f · outbound

This paper cites Diffusion-guided reconstruction of everyday hand- object interaction clips.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Diffusion-guided reconstruction of everyday hand- object interaction clips

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.782808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.773556Z digest=sha256:b387438cfbe76b948d1317d25016d6fa8f662a53b6bed4878f24014af777f1a5

Observation d42757ab-d2b5-405e-b095-279a5de21aab · outbound

This paper cites Affordance diffusion: Synthesizing hand-object inter- actions.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Affordance diffusion: Synthesizing hand-object inter- actions

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.615828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.846684Z digest=sha256:bf11c65d8b22b9105a772d8f2d70628c2fb07b4a154bcfc5cbe61845ba9d32a6

Observation 2a5871f5-458d-4221-823f-9bf05474adcf · outbound

This paper cites Moonshot: To- wards controllable video generation and editing with multi- modal conditions.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Moonshot: To- wards controllable video generation and editing with multi- modal conditions.arXiv, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.436269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.915945Z digest=sha256:577cefcc6457b4304b256c815220e620c3dc73b030aa1bdb45f075e2b59965be

Observation 4ffb5d07-6954-4cfe-9af1-3d347d2e6729 · outbound

This paper cites Graspxl: Generating grasping motions for di- verse objects at scale.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Graspxl: Generating grasping motions for di- verse objects at scale

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.220306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:06.985295Z digest=sha256:0819a9ba11791f650da5a9faa8068802b5dbd11a9f30f0c0ceff9a107fc0cd77

Observation 18df7205-b067-46c6-b67d-2504aa2089dd · outbound

This paper cites Hoidiffusion: Generating realistic 3d hand-object interaction data.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hoidiffusion: Generating realistic 3d hand-object interaction data

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.041631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.056644Z digest=sha256:a7f4409eb119171271384776162e11408b316288a94d966e5f441ff1ceb0f1e4

Observation ee2636fd-962e-40aa-b2af-a097141604fa · outbound

This paper cites Place: Proximity learning of articulation and con- tact in 3d environments.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Place: Proximity learning of articulation and con- tact in 3d environments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.852017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.125844Z digest=sha256:cefcbc2ebba3139f6c35ff6b868483acd9d7c6ec5e9285b9a19520765062bc31

Observation 05fe2724-05d6-488f-83ef-168f9d183dec · outbound

This paper cites Generating 3d people in scenes with- out people.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Generating 3d people in scenes with- out people

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.673946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.207170Z digest=sha256:a3ce6b6d891ae4bc0059216ca7ded7de01f814241a451aa34fab5f9b00319460

Observation 072c37cb-e248-4725-ba45-2e8b6c1da1bc · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:07.314086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:07.314086Z digest=sha256:41c57f707628902707e88fefb9b1f47eda3cf384cb542a9550f061cd37fd762b

Observation 789bac7d-8272-4bab-8e76-4cdd1d12f094 · outbound

This paper cites Avid: Any-length video inpainting with diffusion model.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Avid: Any-length video inpainting with diffusion model

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.464243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.430241Z digest=sha256:9c1a84a35aca406fcd18b57b9a492bc450e5ba928bc2f8d11f0ec61223f8fd61

Observation 3cbbf123-e7e0-49fa-8dac-c73266f3dae7 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.247269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.513431Z digest=sha256:2e27f3313a8a772d033e02870a61ad227b7b6469ae26a568ec511044b10b89f8

Observation ae7da5e1-6bca-4c2c-9f02-5fcb4cafb4b8 · outbound

This paper cites Realisdance: Equip controllable character ani- mation with realistic hands.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Realisdance: Equip controllable character ani- mation with realistic hands.arXiv, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.007311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.594471Z digest=sha256:b17de4e91fe452802d9d119ca8ee28c1e4d29d0db06e0cf52db479a0cf72a7d4

Observation a4e8ed7d-5a25-419c-ac8b-949aafbc6019 · outbound

This paper cites Discrete contrastive diffusion for cross-modal music and image generation.ICLR, 2022.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Discrete contrastive diffusion for cross-modal music and image generation.ICLR, 2022

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:08.827167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.671543Z digest=sha256:870d52d29609bf0b1c87bd6908653c1a30651666907692b8e7adb9793fb32467

Observation 9d0c6260-f658-4548-b34f-d47e83a80825 · outbound

This paper cites Cococo: Improving text-guided video inpainting for better consistency, controllability and compatibility.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Cococo: Improving text-guided video inpainting for better consistency, controllability and compatibility

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:08.662632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.779574Z digest=sha256:a435a4610c5fae544e7e747415b5276fec56bb838a86472e9313bd485adab20d

Observation 500b1a36-b8f8-4578-88ea-446576a18f7f · outbound

This paper cites Cut-and-paste: Subject- driven video editing with attention control.Neural Networks,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Cut-and-paste: Subject- driven video editing with attention control.Neural Networks,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:08.442439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:41:07.918981Z digest=sha256:b313dcd83db071ea0b65fe2b0228d30d2616c716134d836757eaec371ba44168

Pith citing papers

Observation 7addece3-57dc-4b59-ad74-400167705ffc · inbound

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation cites this paper.

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T10:37:54.999164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:37:54.999164Z digest=sha256:be5a3a4adcad349e32ea5f38b0f2673dd0735ba3d0c4dd82275451ef7bc849da

Observation 7100c40b-5eb5-4d71-b27d-1189285c9be0 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:18.894578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:18.894578Z digest=sha256:68aa28a29c2da5d540d17ad26f139fae342aa65ddf4c6b8b6b594a6bf0dfb72d

Observation f587cc87-3ff3-4bb0-8e72-3fd690d0778d · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 205

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:27.095194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:27.095194Z digest=sha256:3e80f4453213b64becfb45383ebe5be81adbaae7178b655e9ba93dba44a55782

Observation 30d0897e-a33f-44be-8a63-bbebb729aaee · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 187

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:09.117957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:09.117957Z digest=sha256:04c127d592828ed650dcd21788c6036e9cbaba44edc20db52442e4534e4a4ae4