Pith. sign in

Paper Citation Record · LEDGER

GenHSI: Controllable Generation of Human-Scene Interaction Videos

As of 19 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 2 inbound Pith citation observations for arXiv:2506.19840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19840 v2

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T07:25:41.219753Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T23:30:14.969895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-16T23:31:22.000482Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact37
  • verified fuzzy52
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2d686a18-97ad-4803-ad5c-e21ec9abcca0 · outbound

This paper cites https://huggingface.

GenHSI: Controllable Generation of Human-Scene Interaction Videos https://huggingface

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.391897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d681a05f5118103c9baf2fe9b936a8df01a8dacb8c150631275583c9373c8fb1

Observation 6aef5841-5d96-4d7f-bf92-0beee8b0ee9b · outbound

This paper cites https : / / klingai.

GenHSI: Controllable Generation of Human-Scene Interaction Videos https : / / klingai

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.383113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d4302028699208cb42e45c95b79cb8f2d158efa948b6eb00ccbb7362ef9e90e9

Observation ae8cea70-f067-4eee-a58e-543e57456759 · outbound

This paper cites https : / / klingai.

GenHSI: Controllable Generation of Human-Scene Interaction Videos https : / / klingai

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.370562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7aa9b6b6667f0b94a3963efcb9ea236d9bd18db164990feb634b860008954b9e

Observation e94bc09f-f9df-4fa8-9b1d-fe8372b631f3 · outbound

This paper cites Circle: Capture in rich contextual environments.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Circle: Capture in rich contextual environments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.387795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:77d11a8f0791b312831f4467235c98a490ffb94563526bebc9849bf94dc4a8cf

Observation 79e7acd1-7de7-44d2-a56f-9471e8513031 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:09.017753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1ca4a5828b06a19f5966860cdeeaccc7ba07e42f73dd951258e68939c2a64a75

Observation 0b34be10-717e-434d-b030-c5ffc30ce52b · outbound

This paper cites Align your latents: High-resolution video syn- thesis with latent diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Align your latents: High-resolution video syn- thesis with latent diffusion models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.380862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:5b98607604c263cba25e26f96990bad6f4dd7d57532fec6ea697ddba19574a3b

Observation 8e7dfc2e-fee8-46e4-b10a-53b6c422f84d · outbound

This paper cites Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.987121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:10f22c3e7c43330dd7d22b8ef4d1ec1cc910dd0192862104852c19f9acafcad1

Observation 64831998-ba09-4e44-9bf8-46a485de7063 · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:08.888669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1d92f3981ba61d3946f7fefe9c5e036e34924174ad5d486ba7f2c11e56851f68

Observation 18ccce00-371a-4fde-b02b-54c13375c4f3 · outbound

This paper cites Wang, and Gordon Wet- zstein.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Wang, and Gordon Wet- zstein

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.372865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:135ae0a74cfcca7148b83068da19fdef4c35d3c4672976e5d887e89d98d8a4ec

Observation 01b88b08-d983-4942-ab93-41a85ebec5fc · outbound

This paper cites Gen- erating human motion in 3d scenes from text descriptions.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Gen- erating human motion in 3d scenes from text descriptions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.432964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d6827a63e4c7cc95b3d1410ba6a597aaa0636e5bfd818459ee39f67532997767

Observation d1813d2b-2b25-413a-98d9-b96938e3bb6f · outbound

This paper cites Videocrafter2: Overcoming data limitations for high- quality video diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Videocrafter2: Overcoming data limitations for high- quality video diffusion models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.434985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a8ccbba830eaa98c86d43a7649e6be89f83d3a1ef2df5d8ebba123130f06cbcf

Observation d34ff940-edb5-4b82-ba0b-7606b5f9523b · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

GenHSI: Controllable Generation of Human-Scene Interaction Videos PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T07:27:09.031158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:9dec6c10af8ad9b5b64310f7f6b6e324079d3ad52a8365799d4bf094bc849331

Observation 8980f562-7d33-4e49-909f-b8303dd47069 · outbound

This paper cites FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.950681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:595a71c5b67863a874be0f38c92e9e9cf95a8e6af00a07378282d819dc17b222

Observation a41d63e8-a50f-4d6d-b081-8964572aa162 · outbound

This paper cites Multi-subject Open-set Personalization in Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Multi-subject Open-set Personalization in Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.962058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:33a98ffb5f11a4b344da3b3dfea1a6c8b947475c69da7d72e5be399d7aba996d

Observation 95f81163-0c62-48c3-b556-f280a720f4a5 · outbound

This paper cites DreamCinema: Cinematic Transfer with Free Camera and 3D Character.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DreamCinema: Cinematic Transfer with Free Camera and 3D Character

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.924583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d7b37d625886bc568c453a9b7ed840ce485b88aeb381e6cdba52e522279e667a

Observation 6d017889-c582-4163-9e44-3fa9aad3cf04 · outbound

This paper cites Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.407942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:125008039d165a68c7b75a8beb7393b903a886049abda8e1f679ca53b7773421

Observation ea15e6f4-ff2a-49bb-a810-e591ee266e7d · outbound

This paper cites LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.928143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7f68f3af13737defc0d461d3663bd375ccf4a38a4a1c543816fd35177c80e5c4

Observation c5853e52-9dc2-49d6-9516-9bf42fe391bc · outbound

This paper cites Dragvideo: Interactive drag-style video editing.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dragvideo: Interactive drag-style video editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.415913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:fabe63787c348ee4dc28f96f5178c3d0e3269b2e239fb90122df3b2117aa9d9f

Observation 4735d37b-4b3c-40b7-ac2b-de5f38da1b2b · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.460875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:565627f078d27f3fce3a536c623522a2911238c17c8fe54dcc1b644e1a393239

Observation 38402460-8972-4a89-b6c1-797260cc1575 · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.402967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2363405a889c77684ad3e0f065f46a55a07a79dcb755896984df0d414e816fc2

Observation 44297a95-be3f-42aa-a795-984e0ec19c5f · outbound

This paper cites Motioncharacter: Identity-preserving and motion controllable human video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motioncharacter: Identity-preserving and motion controllable human video generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.943609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:71464e97aed0f94d1bcd03b46d3d3fd845b0c89585ab06d0db493e83c25ea685

Observation 5833cf9c-ea0e-4167-b952-ddf3442a58a5 · outbound

This paper cites DreaMoving: A Human Video Generation Framework based on Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DreaMoving: A Human Video Generation Framework based on Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.903461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:9fdc4f1c8ceb9209a1e9a76594b275745336816090d4f64581710044a27cf544

Observation d4a87c14-ff9c-4fb5-992a-95da6d00b502 · outbound

This paper cites Hu- mandit: Pose-guided diffusion transformer for long-form human motion video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Hu- mandit: Pose-guided diffusion transformer for long-form human motion video generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.424764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:daca86b4419db29df7c06d796a3a8a2a08b7fd5a0776e642c5e18019f07a9442

Observation bde0305e-a6aa-4411-b26e-5ab1ff159543 · outbound

This paper cites Preserve your own cor- relation: A noise prior for video diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Preserve your own cor- relation: A noise prior for video diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.395903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:5c538612d6b369f534b49f5ca8f02173b3e4c77366ec60a3669d0fcb81c27115

Observation 1fe104f2-9d30-490c-a21c-dbc76fe11627 · outbound

This paper cites Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:08.870018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:323f7dc17c0901ca673cf39b74c58655a37ae23baaac7fb85ffe56b1faa0ae7a

Observation 6de8377e-6e7e-471a-99e9-7365c9a8aed0 · outbound

This paper cites I2v-adapter: A general image-to-video adapter for diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos I2v-adapter: A general image-to-video adapter for diffusion models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.470482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:617cda7ed55c5f28be3334a0a4f7b31a91aa1ab2ebeed48bbc911d39655492c1

Observation 57fbd5a5-c7e7-4861-98c4-c4c292d21c2f · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

GenHSI: Controllable Generation of Human-Scene Interaction Videos AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T07:27:08.982745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:943c05ba46a7f8a9123c517b82d828426384c0e8292a4d6781e76092ab25f858

Observation ec21d92d-03ff-446c-978f-9944d80f78c4 · outbound

This paper cites Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.448646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:aac2b619ce394ed3251fe3de09f745810b1c51355967463df48a51229b14a22e

Observation 6eb35fcf-5690-483d-b0ad-2b267534a11b · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LTX-Video: Realtime Video Latent Diffusion

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.881798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:23704577e24a7ab65a0aa66d29e00d590e0fc7cfab962f178fb6a2316bb36dd4

Observation f880363c-9778-4dc1-a132-03e48c0fac0c · outbound

This paper cites Resolving 3d human pose ambiguities with 3d scene constraints.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Resolving 3d human pose ambiguities with 3d scene constraints

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.442835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d01e65a4d88fb0dffdd874985346402a34d19a09dd7c4f819ef4e037d2100cca

Observation 29faf151-4e34-443f-a81c-c327a84bd195 · outbound

This paper cites Cameractrl: En- abling camera control for text-to-video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Cameractrl: En- abling camera control for text-to-video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.438862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3987c976b420c108b1a6581a6a8fa54e6a282378026095362c53c06a1a4f86e3

Observation 0ee1e086-8eb5-41b6-bc25-094c13d7ab0e · outbound

This paper cites Denoising dif- fusion probabilistic models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Denoising dif- fusion probabilistic models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.352457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:ac53be13904ed54f45ee3871957bdb8fbcb78bb223a49d8a852a4308a3e0ee0f

Observation 72adb8c4-cf57-42a5-950b-9c94279a883e · outbound

This paper cites Video dif- fusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Video dif- fusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.378506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3cfb8951316d859559be5dfc2da281f3ee7aa1376cd00ad5886de96e5f55fbb3

Observation d8dfb9d3-1e8f-49aa-b4a1-0c68a32fd565 · outbound

This paper cites StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration.

GenHSI: Controllable Generation of Human-Scene Interaction Videos StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:09.047774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:aeb5097740558e116012ba2959e01ade897d80d1dd1cbd67d3e19ab69eb5a9aa

Observation d8d95c5e-a9f6-46e1-8879-6bd7835672af · outbound

This paper cites Move-in-2D: 2D-Conditioned Human Motion Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Move-in-2D: 2D-Conditioned Human Motion Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.026206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:23f4547b1f0bcbc26ed3bf8198ead1867f1fe9bb270dbec8c62e69dc5b0ddc48

Observation 0a703d01-3746-4922-a8fa-52d7f51a99b3 · outbound

This paper cites Diffusion- based generation, optimization, and planning in 3d scenes.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Diffusion- based generation, optimization, and planning in 3d scenes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.400573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8e8612c4da9215a0baa06a220c38d37fe0c63b9edd770c0fb3332b93d852c555

Observation ab80ed95-0cf1-4260-bb39-cf4bdd53d804 · outbound

This paper cites Owl-1: Omni World Model for Consistent Long Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Owl-1: Omni World Model for Consistent Long Video Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.914073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:ccd3755b698599e142f4a802e6c2bb0866ff0a1c36a1035f93fbdb2618b346cd

Observation 5aef4116-47de-4fd1-8ffb-eab323b57762 · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative mod- els.

GenHSI: Controllable Generation of Human-Scene Interaction Videos VBench: Comprehensive benchmark suite for video generative mod- els

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.338032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:9f7d492ea88696d48d6be34f5ff154bbc959d5527ac4642e26f21c4adf3aaf79

Observation 8c7ea417-d225-4159-9b03-a54064538929 · outbound

This paper cites Peekaboo: Interactive video generation via masked- diffusion.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Peekaboo: Interactive video generation via masked- diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.355593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1958086f046ac27451b713385c0eb0938e1ccac150b1f5254858568a0ecd371d

Observation 1abb434e-c6e7-4f68-9e21-90ea0fc7fec9 · outbound

This paper cites Scaling up dynamic human-scene interaction mod- eling.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Scaling up dynamic human-scene interaction mod- eling

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.349819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:473afa7e884d7850a15ae9c2d073013cf9c4e1b39b9150802870a594b0316ae7

Observation d75ab0ff-7002-442f-8baa-d0364867cf24 · outbound

This paper cites Story-adapter: A training-free iterative framework for long story visualization.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Story-adapter: A training-free iterative framework for long story visualization

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.969859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1573d150479da66a9b84d1292abd57a990c0a3c13fa20287542c44059687359f

Observation 7fb8259d-5cb8-4452-a262-6136b00325dd · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.347205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a45e3794fc88487d11edf40fa9041bef76c79392a5373720772a5cf51c3f26b9

Observation 070fbb13-03a5-4e37-aef3-d79af46ed906 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

GenHSI: Controllable Generation of Human-Scene Interaction Videos 3d gaussian splatting for real-time radiance field rendering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.430996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:e02932ada489c0701b408e3a84350a04d262b86aa6a426caf34c5c9c87015eed

Observation 0536bf5d-27ba-4e47-b7e2-bf5ad2f02654 · outbound

This paper cites DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.899329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:79c0d586be0064f13d853ca7fdc0f0a2b0172a192414dbc4949f4eeec435e1f4

Observation 03e928e1-e6fa-47f9-b330-463f389b43dc · outbound

This paper cites Segment anything.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Segment anything

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.444712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2c7ff45ba5d764100c32f22e18bea6dee71a5f39160423ad38ff2e29b6658ff5

Observation e4bc4dcd-92d7-45d4-b02a-4937a5878675 · outbound

This paper cites Putting people in their place: Affordance-aware hu- man insertion into scenes.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Putting people in their place: Affordance-aware hu- man insertion into scenes

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.440806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:4396f68418e8263a49f4f409f78c82c12b65fdab110757ba903522880af74048

Observation 296e44be-638a-4080-b07d-a60c071a34bc · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.335703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1dd7e949088fba1229adcc3a9cbdf53e981de42d259a47fb3ef7c0c670e1c981

Observation d12c0e47-faf6-40eb-b8fd-ec341830cc80 · outbound

This paper cites ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.005765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:9df8f9b2ab22d8fc1ef1d4bef1347f858a06d29749922b77c6470e785da60de8

Observation 8535c6c0-ccc4-4ca8-92e7-d29bd5452c65 · outbound

This paper cites Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.342596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a909b94f473441561be79c34cbdfa59985230ff4d0a3344e1207e1ba78462bc7

Observation 1849a30c-4223-4bbd-8554-76fc0c4a368f · outbound

This paper cites HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery.

GenHSI: Controllable Generation of Human-Scene Interaction Videos HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.966145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:75f80357700e8c103d458472ba6db9cfa189f7bfd5bb9657fc050f857d5530e8

Observation 4a4af4a0-1bae-4dbf-8d40-ea9a1363cc3b · outbound

This paper cites Genzi: Zero-shot 3d human-scene interaction generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Genzi: Zero-shot 3d human-scene interaction generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.364710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a30aca37e030aeeb7b14dd8f10ef2b41a40ff49ba9f645fca910ee83ef751a11

Observation f588f632-89b7-4004-80a6-2383245b6ef4 · outbound

This paper cites Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.458737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3c3112584d5321979ad9a9cd931c84766f011620311a72d5d74bb1df8e9b3217

Observation 115bb29e-130f-4e5b-bd3f-2b358d905c90 · outbound

This paper cites Intergen: Diffusion-based multi-human motion generation under complex interactions.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Intergen: Diffusion-based multi-human motion generation under complex interactions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.456774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:78bba4ce623c64eb7a4f1294073e0e88a6be131557f9cfe99cf6068ca6ba8f7e

Observation 34ca8adf-9162-44e3-8901-b29d53285c1a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Open-Sora Plan: Open-Source Large Video Generation Model

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:09.022234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d021fa7bde55df9e2c2d6f17c7878f87fec8f60594a388029a277994b433373a

Observation 189a6fe0-a962-4f45-b246-e815895e0b69 · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.462806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8bf1a99fa05a6aff818af6e2ca89a9d23771fc39cc6e1a1f05a8a58e0e0fbdda

Observation 5f29e41d-f228-456c-88b1-cfec2d8fa30f · outbound

This paper cites Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.001999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:624398ccaa1b8599244861f1e524c322a1ae6fb515d9db553edaa9b58dc70f93

Observation 1be80dee-8757-4f06-a69e-df212fd69564 · outbound

This paper cites Phantom: Subject- consistent video generation via cross-modal alignment.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Phantom: Subject- consistent video generation via cross-modal alignment

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.452798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:523bde4053d1fe36a28cc77ccdc656a49b5207cc640b01dbfb41f7fe01a58677

Observation 2d7f362f-8694-427e-b6e4-b59211684046 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.454779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:b84bdcdb5493627384d86847bd471c4e30b93bef92c5399338bd59610c9b8dac

Observation 39061465-73aa-4e5a-b3a5-af5efd170eed · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.358337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7b3300fa58fb56707337b64fb3c5a37f4c0641a2cd93820aafc60241c8acacaa

Observation 96cf5621-27a3-4696-bea5-df3c0801f1de · outbound

This paper cites Smpl: A skinned multi-person linear model.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Smpl: A skinned multi-person linear model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.446574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:dee53849a0a43592c453629f62f6312b0bf8a53b15d6bf2fae04df61d9beaeaa

Observation 0d4a097e-ebf6-431b-9307-90e7333825a0 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a5a0cc4563f3ae5fbcbb844334b5eb8867e85d8db508b9aed2bab0e352f6f6c9

Observation 1d152e8b-37c2-4b57-bcc9-ad345fa439e4 · outbound

This paper cites Trailblazer: Trajectory control for diffusion-based video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Trailblazer: Trajectory control for diffusion-based video generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.464690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:6cb01e8e279bef511f9967b6b3399c37dc44ccb46f061bf296e86664a5cbf2c6

Observation 16e914aa-bf57-4099-b717-a41c5f3bbeab · outbound

This paper cites Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.974069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7f0ad9666770107f966afea7cb896e26fcc61f70dc3cc2ce0f8a400750908e23

Observation 4bfec488-8c34-49e8-b044-e6fbfcdd13ee · outbound

This paper cites GenHeld: Generating and Editing Handheld Objects.

GenHSI: Controllable Generation of Human-Scene Interaction Videos GenHeld: Generating and Editing Handheld Objects

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.990992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8e45b71d319c8032d481bc74b066f9b77f4f46397150a6b828e082eceb1587f3

Observation 300ce8fe-0613-4839-9536-e790d8bb5566 · outbound

This paper cites Chatgpt-4o.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Chatgpt-4o

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.466611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:544c7f06ac6e40d77871012dfe83e0deca35cc980e5aaeecf148d734c14cf031

Observation c4e4a262-85db-4ee2-8300-f1c4f750e3ca · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.873264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1b7332bd753b1d9925be5ae5df15e2ae5577631664e7ce33108786d878a815a7

Observation 0aa829ac-d6d5-4bb9-a7ca-3fedc685819d · outbound

This paper cites Text2place: Affordance-aware text guided human placement.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Text2place: Affordance-aware text guided human placement

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.389822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:0cf2f219b05d7a63fb8e302252d24cda8b17027bddd9f4ebac542e2e96c8acd8

Observation 71e92500-4790-482e-923f-fd7e378ded66 · outbound

This paper cites Expressive body capture: 3d hands, face, and body from a single image.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Expressive body capture: 3d hands, face, and body from a single image

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.450669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8bd3573b5d3f3b8bebfb2754465bdac5b57f01b61ba8281b00d78f6424389ae2

Observation b0877edb-f0a7-4c78-947a-deec8120059e · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.393722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:b318b1842654e318add55502ebbca4e92214e0ceb692778545a888693cbd09c1

Observation 5591942f-fc0b-4541-b148-251f5b228a7d · outbound

This paper cites HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.035768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:0337ad6890ac0e735e07d743b6d7d7fe435d5a7cc99e0555c41088f3d5314277

Observation bf2fba20-6a2a-449f-b352-c2ea3df051f5 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos High-resolution image synthesis with latent diffusion models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.385354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:42bcbc37953e3ba2db03e15249b2d9dcb23e5f4a555b0065a5ac41ca0dd210bd

Observation e902d4ca-924e-4236-aabc-db5fc03a42e6 · outbound

This paper cites Dreambooth: Fine 11 tuning text-to-image diffusion models for subject-driven generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dreambooth: Fine 11 tuning text-to-image diffusion models for subject-driven generation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.375906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1e5eeb286fb5be02926635d10c6eba43caa7ef90e546f17543190d26fef330bd

Observation 22bc4474-9940-4e22-83ba-eb914b642d16 · outbound

This paper cites Magic Insert: Style-Aware Drag-and-Drop.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Magic Insert: Style-Aware Drag-and-Drop

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.009377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:e115bcef5efccf6350c99f7fed21df5b3b77922d71f23ca2eac2601c7e825236

Observation 81b99bce-39b2-49de-8bae-4f071099b6c4 · outbound

This paper cites GeoDiffuser: Geometry-Based Image Editing with Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos GeoDiffuser: Geometry-Based Image Editing with Diffusion Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.043662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:df5c19409bea9cf7eb06d5c2ae2aa44ed26d3653983f5c517144915dc44ce985

Observation 7bacf2e7-fe99-4746-8a9e-19a8c7e2727c · outbound

This paper cites Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.917643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d48548a45d409c3296a8e077fdf2251fdc6bcfeaa2d0d8e01c02830f48761239

Observation 7efac2ba-12ec-46e9-b146-7442b8da72ea · outbound

This paper cites Dragdiffusion: Harnessing diffusion models for interac- tive point-based image editing.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dragdiffusion: Harnessing diffusion models for interac- tive point-based image editing

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.367844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1adf1a8b736cfb0e7165c05deec270d741b8ff14997aa04888cc6810d48540a0

Observation f7b39e52-d76b-492b-9d47-b0adf1258b4a · outbound

This paper cites Denoising Diffusion Implicit Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Denoising Diffusion Implicit Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.920941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:07db6f80ba4b360c3d489127f34f5fd946ca627b05fda95037fe700e424f63cd

Observation 4620c837-2869-44c6-af38-f5cd775c838c · outbound

This paper cites Sound to visual scene gener- ation by audio-to-visual latent alignment.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Sound to visual scene gener- ation by audio-to-visual latent alignment

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.340214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:808140ea964b3cb799ae9ee349033c1f8c25fdbfae1295021db113d75a9988f9

Observation e1b32098-9779-423d-82a5-4c6ce7bf9df8 · outbound

This paper cites Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:09.013480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:05e248cbd26894f1ca4ab722857b63ba11e6b8ce300591a16ae119246fbcaff3

Observation d2d1e529-6b84-4fb0-92c2-28dede2602f8 · outbound

This paper cites LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.877210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:0363d945366b710abe5c72b5f59daceb2340deb5349a249669e5defe73349ca9

Observation c3f32616-caa4-40b8-a3aa-2f27f420d4f9 · outbound

This paper cites Motion Inversion for Video Customization.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motion Inversion for Video Customization

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.946885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:77b4134f08c698c03480aace9a3652b23bad5a3f59bf03c611999821f09c9e58

Observation 925e6259-f10e-4e01-9422-6020090b9be7 · outbound

This paper cites MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision.

GenHSI: Controllable Generation of Human-Scene Interaction Videos MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.895161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:bad75683c78f071c1445b902db7f6214d65c5b55dd301b4726ddeefbfebee4c3

Observation 50ce6ba8-6eed-4fe4-a2ae-a9a976fec222 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.958137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:683e796ff01aa07ee2ca1abc476112d70234788c22c9766e84e8836120a3f82b

Observation 27d965a2-5d41-472a-8f55-7d87a2629c64 · outbound

This paper cites Humanise: Language-conditioned hu- man motion generation in 3d scenes.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Humanise: Language-conditioned hu- man motion generation in 3d scenes

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.420514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:069e9b499b7ebaa591d2b1143aaea96db83cb5f223d1e5a48d5211336d5482b8

Observation c249ca28-d251-4837-b1b9-5beb53033a29 · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motionctrl: A unified and flexible motion controller for video generation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.468603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:e4ae710278c679c6adaafa8da0bc50c6a4146459eca96b988cf89eb4383e646b

Observation f9bb6f49-3570-46f7-8380-9ba05f467d6f · outbound

This paper cites Move as you say interact as you can: Language-guided human motion generation with scene af- fordance.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Move as you say interact as you can: Language-guided human motion generation with scene af- fordance

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.426869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:b22175bdfd26de179a126a38af568396fd61875a7cdc931eb99c2d0f40bc622d

Observation e9760be3-30a2-4784-9777-75f367f9a1bb · outbound

This paper cites Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.891900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:47bee853538d0ac5dda1089c866eb4eb6ad3f18f0a9c9d527e6c477b39d0e837

Observation 88e19441-c22e-44dc-858e-1a9cb996cf0d · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dreamvideo: Composing your dream videos with customized subject and motion

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.429085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:276ab7d2cdde7a70e719283b7b230b28bd1caab02284d8937742c5dfecafd28a

Observation 64cbe6db-137f-44c0-9de4-4844d727dc9f · outbound

This paper cites Detectron2.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Detectron2

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.436871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:be30011dee06571d8c01d24f0fb2a83740f3cdb0a0e4106847a8d1a2c3ad09f2

Observation 43fc0743-f735-4ecd-87f2-75b0dc7b0efb · outbound

This paper cites Mind the Time: Temporally-Controlled Multi-Event Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Mind the Time: Temporally-Controlled Multi-Event Video Generation

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.907092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2f901936e6661f310f232c47f323c5b87a11edd1274652935cd9f34084fa6c93

Observation 8290e019-f9b5-49cc-afa5-b322e83843f7 · outbound

This paper cites Structured 3D Latents for Scalable and Versatile 3D Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Structured 3D Latents for Scalable and Versatile 3D Generation

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.978037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:6c0c63a1c8161f50943e3d1033305b05c03bb93411e6fb8353ec1a679d65252a

Observation 49fb1186-151a-4121-86a4-0fe4be1d0868 · outbound

This paper cites VideoAuteur: Towards Long Narrative Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos VideoAuteur: Towards Long Narrative Video Generation

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.885606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:4b74a5b6d27677cc3795d55319ec53fb4892ccc0f94cf2a73b2a2bacc85bca54

Observation 43bde310-09d2-4678-88eb-0cfe5a0906bf · outbound

This paper cites DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.040021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2b32074b242f9cc2bec5642933211bf265182595d320cd749e6d9a54011bee7d

Observation ebf16d2e-4af2-4012-8a34-6ea564663f14 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.422556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:de4d514401c9fc391faf81fa2476a249e92cd0149751db55b401d04d40ac5d76

Observation 37457f85-25e3-4d6e-956d-77a39d7f0d06 · outbound

This paper cites AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.931807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:18d6e739928bf0119cbcbc144d5b050be0ed89d2da1cce97a5eacd209329f8d4

Observation 88d53b2d-cad6-438f-a4b7-61efa9a77901 · outbound

This paper cites Person in place: Generating associa- tive skeleton-guidance maps for human-object interaction image editing.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Person in place: Generating associa- tive skeleton-guidance maps for human-object interaction image editing

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.410778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7e2dd184de449d98007ef60e936e96558447b770cf7e46d674b8c837cc5f5e0d

Observation fb632a70-f0a9-4d66-b20e-c4f0989dbe53 · outbound

This paper cites Generating human interaction motions in scenes with text control.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Generating human interaction motions in scenes with text control

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.418152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:17d1fc5c8b3dfb6d1c061a159a9d8a827eb74b877a5d7c37d680e0c3a626bb9e

Observation b675014c-3183-4d0e-9b43-feb2333148a4 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Adding conditional control to text-to-image diffusion models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.413583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3215997084a48ddf85da05ac5e63c55e1466c963ab5dc8d106d38b48fc16ad86

Observation 04faad48-d1f5-4bff-8960-98b4b22362f1 · outbound

This paper cites Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.055448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:75a073cf80fa280ca690b02b0aacc406bb22ddf88208506c10b818ac34ee1869

Observation 171c67e9-8747-4036-b9ea-b0410b11eb38 · outbound

This paper cites Generating 3d people in scenes with- out people.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Generating 3d people in scenes with- out people

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.405656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:dae3983d68d907059e71bc7ed8f96c7a71b60676f864b36dd951d0e9c4eff7c4

Pith citing papers

Observation 4a07035f-929a-442a-9a64-d6dbb5889d01 · inbound

VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification cites this paper.

VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification GenHSI: Controllable Generation of Human-Scene Interaction Videos

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:31:22.002931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T23:30:14.969895Z digest=sha256:e134a456df9a41aeec6ac45733c416ab980db4aaa287fb7432c9acc8e76f99a8

Observation fce67bb8-d112-45d1-9e60-fdf9e3d2b5ba · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos GenHSI: Controllable Generation of Human-Scene Interaction Videos

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:47:57.383574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:c647128dea5331135b108b5ce36034662cb5423519b7d92c935911a7c31b2b08