Pith. sign in

Paper Citation Record · LEDGER

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

As of 18 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2412.04189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04189 v5

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:45:10.247547Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:16:44.469541Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T18:16:44.762071Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f642428d-8390-4af5-9e2e-5c12478cec54 · outbound

This paper cites The mug facial expression database.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation The mug facial expression database

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.848461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.848461Z digest=sha256:0cd193ec3dda2aa4f4b5f210c65ceb015cb4e4fe86b147de49f18dad94f8e39d

Observation 3729c403-f6e6-4ebd-be85-0bfa87f6a5b4 · outbound

This paper cites Detours for navigating instructional videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Detours for navigating instructional videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.853490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.853490Z digest=sha256:a2d1eea5b4bc84f9bc11234c275164d4af05dbbaea64e4792cfa1cfd1a394a1d

Observation 0d056796-4471-4d98-8e3a-c9b3035531ba · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.857437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.857437Z digest=sha256:6324923e49025c7f0f5955fa23f0a5d411100c2c0e9826ad550d66a666519349

Observation ef9e9c3f-5e94-4e05-ae3d-8469909202c7 · outbound

This paper cites Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.512917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.860935Z digest=sha256:44f2793eb5bbc72e4fb37fe0e678dd8934e474d6b4dbf375b096eb258b1e10ac

Observation 081a8a1c-63a7-44b8-916d-1cb2173a61ee · outbound

This paper cites EditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation EditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.864718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.864718Z digest=sha256:b1912239f02935bb8f6931ae43c038a674ebb10e7eac1ffe0c41c91418c92631

Observation ea11565d-4198-43ef-8e98-275e2cda8b11 · outbound

This paper cites Brooks, A.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Brooks, A

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.496580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.868784Z digest=sha256:a927df314eda92894fe06c932bdbfc603057a8bdeeb8756fccc1344a48d18191

Observation 31a2ed05-e1eb-4f98-a350-6f0915c6de41 · outbound

This paper cites Gener- ating human motion in 3d scenes from text descriptions.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Gener- ating human motion in 3d scenes from text descriptions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.483499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.873338Z digest=sha256:883e6194682aa173acab14532ecbb8bdd78787965787be78ed73ed519ccf190e

Observation 164c07f7-1778-4547-9ded-584d90d528d5 · outbound

This paper cites Ceylan, C.-H.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ceylan, C.-H

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.469388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.876682Z digest=sha256:21ceb931097aafb2cb8bba3adc68cfa4dc782f7953eceaa4d39ba6db8872449a

Observation 653a53e4-930a-455c-9b5c-2321bfceb89d · outbound

This paper cites Cognitive load theory and the format of instruction.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Cognitive load theory and the format of instruction

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.457298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.883210Z digest=sha256:4fe0ffa95e7cac028ef9addb1dfde0fa04d6b6020b8943291307249b96a0dc35

Observation 2446aa98-4e4f-453a-84e2-436b1a2a8a33 · outbound

This paper cites Learning video-conditioned policies for unseen manipula- tion tasks.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning video-conditioned policies for unseen manipula- tion tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.439261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.887050Z digest=sha256:a8ccc13c4063053a996b7c33f1b4c5ff3577e9ea76d7b8ebf145be0bc747d561

Observation 3bf0f642-0624-4c18-b885-de977ce05134 · outbound

This paper cites Cheikh Youssef, A.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Cheikh Youssef, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.425456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.892699Z digest=sha256:ac72562d5397daacfcc6710c08d776b1efdf87ef5ff8866916d9508cb57d3e07

Observation f1dc980f-22a6-483d-b76c-4f6416783538 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.897586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.897586Z digest=sha256:d234b8ce8ed99a449f5f2c09794f558d18a8e1be4c003149e4719ac94a93e1f5

Observation d422d952-13fb-484c-8976-fac8f8d821a2 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.902375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.902375Z digest=sha256:f8a3da492c40bd1bd634a7132eb28f66de6211332365e9207b2a0ad78e9ab7ce

Observation b93ae10d-05e5-4231-aa50-6a683b3c535d · outbound

This paper cites AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.906734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.906734Z digest=sha256:027cd674aaab290a4de5ff66307ed6a573e7fa085b8c9c9951a6625593be001c

Observation a078f0a4-2dce-4c56-95c3-d08364ef75ad · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.384219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.914009Z digest=sha256:7068b2af185ad10f0106c4289c4393e837f3b0169e51187be840d0fccf330fea

Observation af9eabd8-b54e-4094-8547-fa15793678a9 · outbound

This paper cites Learning universal policies via text-guided video genera- tion.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning universal policies via text-guided video genera- tion

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.367378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.919008Z digest=sha256:85a78a2c87e74ed12cad68cc0f17b485a04105bebb6210b69e583d05cdabfd3a

Observation 66797e2b-787c-43e0-ba36-820b7eeb3cd5 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Structure and content-guided video synthesis with diffusion models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.924083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.924083Z digest=sha256:7554ba71bc1bccefb8522909a15a04448943d2d4f0a17cb8983bfde9c0e1ebef

Observation c68ac1b6-f258-4cdc-a606-cdcb743c14bb · outbound

This paper cites Handrawer: Lever- aging spatial information to render realistic hands using a conditional diffusion model in single stage, 2025.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Handrawer: Lever- aging spatial information to render realistic hands using a conditional diffusion model in single stage, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.342122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.927920Z digest=sha256:f0bccca2162eb9d2f05462c52a8a8c1e580400cec18e6baebe9395be83bd4afd

Observation 5be0a0e2-f208-4406-9f00-0b9e5df9ebd9 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ego4d: Around the world in 3,000 hours of egocentric video

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.328972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.932703Z digest=sha256:48fd5b199e4233a5a8f5d8145164f144f75bdb9928ec18cd8e6019bde4590e36

Observation 300d4442-7dcd-420d-9ed7-4b57719550eb · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.315625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.937048Z digest=sha256:f025c3fcd90f711d14c4ff1d939f8271aa877a5f32c75bd0c92309d7d1e590a4

Observation c172d233-926b-4cbf-be3a-8421f482e350 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.941469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.941469Z digest=sha256:756627b82f88362f49eb1427dfedce9f6154b9996814c1762154c27fbd0b3ce8

Observation ec3f5388-06bb-44f9-8e5e-8ffcae4ad229 · outbound

This paper cites Denoising dif- fusion probabilistic models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Denoising dif- fusion probabilistic models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.947071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.947071Z digest=sha256:7cd941a5f868e97c7f138243174894436c6609e3434edeca7199707824956e56

Observation 0c3099ee-0f63-4835-b28c-08df527f70b3 · outbound

This paper cites Video dif- fusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Video dif- fusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.278384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.952593Z digest=sha256:3ee34a090825980c4d429253cc6c0e9a6285ec686ce2833c1e55c2574a3de0e5

Observation a0aa761a-7ed0-40bf-9aec-66ccbc44ea45 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.261520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.957306Z digest=sha256:4b6e824e7b5be39ce1cd44e25dc7b04f8f8f53d08895c79c9d7d94cbabc6f318

Observation 5824e7d4-8696-4c6c-a8a4-8a053a525956 · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Vbench: Comprehensive bench- mark suite for video generative models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.246127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.961560Z digest=sha256:27d64797124d65b37169976bca736d483b4d8bf81fe10deb63439f28b106b461

Observation 8f37f44a-9c29-430d-85bf-2d5840cbbf56 · outbound

This paper cites VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.965337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.965337Z digest=sha256:11413f07b598435b0f3ab8bfb8dfa5da8eb65623988af3524896048521cfcc04

Observation c817a730-3f91-4b55-8beb-f5f1a6508de5 · outbound

This paper cites Vid2robot: End-to- end video-conditioned policy learning with cross-attention transformers, 2024.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Vid2robot: End-to- end video-conditioned policy learning with cross-attention transformers, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.230494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.970349Z digest=sha256:cdf80fe889faaaac43dc2ae468ef5d2f6819fb9d14b3a9eab5f6c32708310a13

Observation eb93ff5f-db30-46f1-9b59-6d6aebd881da · outbound

This paper cites LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.974729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.974729Z digest=sha256:75e546246fdd68ced01eb4801a46a650b147e71df6bde8c7f27b7b39a972c43e

Observation 680d6f3b-8e02-4a57-b672-610afca8998b · outbound

This paper cites Temporal convolutional networks for ac- tion segmentation and detection.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Temporal convolutional networks for ac- tion segmentation and detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.217323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.979486Z digest=sha256:d63058f68eeaf29dd16946b6ba09b004be3239442edea7d707f3c5c3652a1329

Observation e39195db-2bee-4e88-8eb2-6b62179a1336 · outbound

This paper cites Gradient-based learning applied to document recog- nition.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Gradient-based learning applied to document recog- nition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.203636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.983851Z digest=sha256:594783eab32c44e2a917c03eb90786b01ccca85cccbc7fce3c8b711501db144a

Observation 7b190191-8e12-4d35-ae9e-5f8621dff8e6 · outbound

This paper cites Holis- tic evaluation of text-to-image models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Holis- tic evaluation of text-to-image models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:09.988588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:09.988588Z digest=sha256:08030833638a8d33fa679a9b392bd3ba109dbe8f72be2a0370d6f757f78eda3d

Observation ca687851-383c-4a8d-984f-15c92b7ccb87 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.175203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.993512Z digest=sha256:26f2385bdcc5add0f585eb9665487a6edbfb3053bcc2f931b01a27b814311f05

Observation f94d3f16-8793-46ab-8654-9082073dac07 · outbound

This paper cites Egocentric video-language pretraining.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Egocentric video-language pretraining

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.160981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:09.997890Z digest=sha256:6c5687ed9698ea1a6ca48bbc32cb6d538c1d5bca06458b7e4a9b2295e5f57f94

Observation 98c6eb90-b37a-4244-8bf4-5efb8029c341 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.002519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.002519Z digest=sha256:9763ac2ef526a8ecae161195a708714cd35d94432abb7c7409466d31369851c4

Observation 8d20e3a2-82c3-40c9-824d-5b1d9fec4651 · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.007032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.007032Z digest=sha256:bd6131e94e1afafe70fcc74ebf9cd6636e66cc43fff2f04eed7237bc5b6a8486

Observation 50bc3a62-629e-481a-8aa1-fbb9cc8b7afb · outbound

This paper cites Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Handrefiner: Refining malformed hands in generated images by diffusion-based conditional inpainting

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.132941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.011270Z digest=sha256:897ea786b5245bdd098d265e4562d2cd9c05454138a649332e46c4205f52e960

Observation df996d96-a8ef-47e7-a982-d9058f18d572 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation MediaPipe: A Framework for Building Perception Pipelines

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.017361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.017361Z digest=sha256:2285a2256603fd9e4ebe05a7cdbb9d73eeaf676b1793c8f1c0278fcea49db4ce

Observation 49015ec3-e604-4884-8243-cf92068f8b29 · outbound

This paper cites Dexvip: Learning dexterous grasping with human hand pose priors from video.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Dexvip: Learning dexterous grasping with human hand pose priors from video

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.117322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.021924Z digest=sha256:2d4b2d121a8d9dd053b3ef1e66fe3f69055c7918af12e0542b37996513e1f7df

Observation bf88c734-6f0a-4a9f-aa4c-b6d1cf86dbce · outbound

This paper cites Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Recipe1M+: A Dataset for Learning Cross-Modal Embeddings for Cooking Recipes and Food Images

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.025263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.025263Z digest=sha256:3553af68734a57cf0ae6588b77318131537b2060aa8baa079e99cc40febd9c43

Observation 86664d79-ab7f-40d2-b01b-9fddadd615b9 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.028678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.028678Z digest=sha256:76f40ad03fd9d2ed0da0643bfbaf5d551075e04e50beb9ea806f9c496a04cac5

Observation fa91b37d-daa1-4561-938b-4830fd64619d · outbound

This paper cites Han- diffuser: Text-to-image generation with realistic hand ap- pearances.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Han- diffuser: Text-to-image generation with realistic hand ap- pearances

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.095541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.032120Z digest=sha256:f0e1886858014992ec97a91361aa1417b1c30740e21059b30cdaebedfe80b91d

Observation 0af6a45f-b2e8-4f19-9e17-be6b9a04bc55 · outbound

This paper cites Conditional image-to-video gener- ation with latent flow diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Conditional image-to-video gener- ation with latent flow diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.078521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.036469Z digest=sha256:a6cc1638ba1601823461fee8dc39f91c64c8753ac350a9def1e922c4afe24e92

Observation f6768da8-c1d6-4744-bbc8-cf575ce84005 · outbound

This paper cites Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:11.061746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.040254Z digest=sha256:8f42c6b9776674c89be444a450cf22e2393fc42240d7a90cdc49924a5bd4fc31

Observation 78ee2f4e-1f22-4847-8472-d2e9bcb3e993 · outbound

This paper cites Improved denoising diffusion probabilistic models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Improved denoising diffusion probabilistic models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.043511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.043511Z digest=sha256:074211e6cc7a0e04b762bf76e75eac4ce72c9c79da0335cceca886edfde7550f

Observation a4102e6b-ab42-4357-a7fb-06d1fcc38034 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning transferable visual models from natural language supervi- sion

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.046644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.046644Z digest=sha256:28174b8ddac59f5917d4410a8e4ee8855c8eaa65fb93c90ab6e61265e3b85bf8

Observation e52dd587-8d4e-46dc-b45b-69cbee120526 · outbound

This paper cites The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.050808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.050808Z digest=sha256:39463b3bb81e86dfa07ad27bab2d2cef6ae4014bfca9fb38980d63805eb6f48e

Observation ef77aa70-7a2d-415a-9c0f-a102233f68e6 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.991089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.055175Z digest=sha256:1a109c26eab5900e067a6ecda69e0c953e323a5dcb912582ec87cc81929a7234

Observation bc4e86e4-e2c6-4ed0-89e9-ecd873b53b03 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.058737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.058737Z digest=sha256:c313be77241a2cdbfa57c69e1852eff2b57ea1cd67936eb4710bac9c36b5a4d1

Observation 6c642cb2-8f57-4ffd-9513-3129aa1ebdb9 · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.063091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.063091Z digest=sha256:a094b43765e6b7c880f69441cc727a5dbb99365c63cd98c1a018890f98fd3156

Observation 53feb2f4-ee4d-4333-943d-3dbf01ed72a6 · outbound

This paper cites As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.067305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.067305Z digest=sha256:1a7bbdf672dc614a47947eb443320b9b09db7fc9c42ff33097cd990f97cacbfb

Observation db523071-473f-4964-9366-d7e124177089 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.071463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.071463Z digest=sha256:d911b3817e2ce7ef3d8507356a4d2298aabae3937bd1939a98dc478a0f74585c

Observation c4df0f93-d0e2-4c0e-b988-69f5664eb36d · outbound

This paper cites Many turn to youtube for children’s content, news, how-to lessons.Pew Research Center, 7, 2018.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Many turn to youtube for children’s content, news, how-to lessons.Pew Research Center, 7, 2018

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.959083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.079620Z digest=sha256:392c0fa909e5587b71e8d16dc9776ef89b201e8f43c7c5804a88717e591627ae

Observation 66cdd9d8-2910-4151-a238-2f0e5b04c7f3 · outbound

This paper cites Denoising Diffusion Implicit Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Denoising Diffusion Implicit Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.084373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.084373Z digest=sha256:6836dd482e70d920db1e1c1b4c65aa2ac27a8483a4279df7bf9f99257b87278f

Observation a897df8d-f6e3-4e7a-8dfe-13e88d87d733 · outbound

This paper cites VideoAgent: Self-Improving Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation VideoAgent: Self-Improving Video Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.088431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.088431Z digest=sha256:2e04899854191a1420b2d5522fe3e6ba0bfdb0b5cb7981cc2464099b5e3c452f

Observation c95c5cdd-8d6b-448d-83a9-3522f8228f47 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.944483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.092424Z digest=sha256:b95ba40582724319916484e12cad5955e59746515b1937ce23fd9532c99f4cb7

Observation c43752f6-ee06-4057-b6f4-c3d3a596b5f7 · outbound

This paper cites Fvd: A new metric for video generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Fvd: A new metric for video generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.929599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.097624Z digest=sha256:94129df575e886ad44be4f1a7aae6849e19e62b4e4535ac5d12e3bb2c8f6abca

Observation 1b24bb18-76bf-4fe2-9906-07aeb04a19af · outbound

This paper cites Long-term temporal convolutions for action recognition.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Long-term temporal convolutions for action recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.911202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.101769Z digest=sha256:847b60dea8bc42bad5b316d51bcc0b878981be0d29dd5b563f23ccaffab52f9f

Observation 31ee753b-b15c-464d-88bb-54425c944b60 · outbound

This paper cites Attention is all you need.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Attention is all you need

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.108198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.108198Z digest=sha256:060d286645e1d4548248d7d0441743d730946f7ffe1ba28065d713562479786e

Observation 0b87055f-1e6c-4f93-8fb3-8744c03b4455 · outbound

This paper cites Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.113909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.113909Z digest=sha256:96924c349ff443009222847125d209d68de28e6176587360a60fafae110c84b6

Observation 57f8ec96-e070-4281-9377-522da41f2d4e · outbound

This paper cites Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world supple- mentary material.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world supple- mentary material

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.878103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.118962Z digest=sha256:07c1d0b847ae82c85ce62221093ef46ec594993dcb1cfef316df405fa5c73b23

Observation 31f1c55b-a807-4f1e-80fe-b475490bfc37 · outbound

This paper cites Multi- modal augmented-reality assembly guidance based on bare- hand interface.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Multi- modal augmented-reality assembly guidance based on bare- hand interface

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.859282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.123297Z digest=sha256:bcffeb914607586efe7f2e50ee18ea608d4eec397d81e5f9be9cac172714cdfd

Observation 01d2e141-519f-4a36-bfba-906867de73c2 · outbound

This paper cites Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.127496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.127496Z digest=sha256:0f606fad17639605a4613b7984326ef1dbe55a394d004f066eda21ef7f6c4802

Observation d2f5d3fa-986b-484a-b3ef-a2dd0c6c8b6b · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Videocomposer: Compositional video synthesis with motion controllability

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.132365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.132365Z digest=sha256:23335da8ba251740facf9bc61348491b4342c5d26394ef5fbe9f4a47c46aea31

Observation 6f5f1db1-2755-4779-a441-58553b189d1e · outbound

This paper cites EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.136640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.136640Z digest=sha256:379cf92938b9cab4817bd81c2b08a4884cd9344b2efe4ae78db44a638819ac27

Observation 95626d99-19c3-4dac-a75d-04c1eacc7e6c · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Lavie: High-quality video generation with cascaded latent diffusion models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.825113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.141927Z digest=sha256:e27cd1d49e7d8d95b3d0031f45b8154b042da2cf984dc3eebc48696ca50e86f3

Observation c3a86d49-5613-41b7-80cf-553565c35d36 · outbound

This paper cites Towards A Better Metric for Text-to-Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Towards A Better Metric for Text-to-Video Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.147322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.147322Z digest=sha256:6bb96c964d15fa0fd7fd9339b748cb1b5d25a5fdbbacf338c2dab80ada124434

Observation c2d1a377-f961-4b2d-9d36-2c0252e97de2 · outbound

This paper cites Freeinit: Bridging initialization gap in video dif- fusion models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Freeinit: Bridging initialization gap in video dif- fusion models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.809310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.152539Z digest=sha256:6de9171edbbdab019b5742794ba952442d0cb0977aab4aa328e4bd06a7faa07f

Observation a319beee-3af4-4f41-ba47-b358f3743d64 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.791674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.156459Z digest=sha256:0e97dbc4bb8cc2c062e057fefa1233340989aebef7ff157478806866445c8125

Observation 69cbdec1-6388-4ac8-b1f9-b99551f26838 · outbound

This paper cites X-gen: Ego-centric video prediction by watching exo-centric videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation X-gen: Ego-centric video prediction by watching exo-centric videos

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.775728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.161433Z digest=sha256:dc7a585bc2bf02a16d22c51837ed90387e30351bfac4d7ffc8376c09cd9edcc9

Observation 1ed1cdca-47e3-4de5-aa48-68d8ba40aacb · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.761706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.166491Z digest=sha256:af4d5d5636a4938888b2760ab48ac61e3ba19bf91cd7e38f4a442523238fbf00

Observation f616454d-f686-4b37-812b-d70680323439 · outbound

This paper cites Stat: Spatial-temporal attention mechanism for video cap- 11 tioning.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Stat: Spatial-temporal attention mechanism for video cap- 11 tioning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.739904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.172257Z digest=sha256:3a2f357bfb7c07fb6f10d7bea8cae23dbc2a0042a2d80b1269c6b8e137747cd8

Observation 7fe80857-1b75-4ccd-bb03-8787d860ba96 · outbound

This paper cites Learning Interactive Real-World Simulators.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning Interactive Real-World Simulators

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.177014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.177014Z digest=sha256:7baea06b5df88969a09b159b75df75fe90d216f2ea969f6361d55ecb0f3cc755

Observation d4d2ce0d-5e36-412d-9e67-848dc9fd2a39 · outbound

This paper cites Annotated Hands for Generative Models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Annotated Hands for Generative Models

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:45:10.372031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.184737Z digest=sha256:fee316d19059c2d8e2b22718e908eda7586317c1abba5a305f72e52db022d38b

Observation 3b0d8d95-bca7-4ccd-9b57-b725e623dc4f · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.192103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.192103Z digest=sha256:4e81d47b50b075a37fdf49ae13346da9ce33a4ff2265c47de878a15ecfc7aad3

Observation eb9af77c-8673-4a21-9319-b37c2a7c5d11 · outbound

This paper cites Learning universal policies via text-guided video generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Learning universal policies via text-guided video generation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.718568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.200996Z digest=sha256:90d875a53077f8b97c3127d8940e62868d6c73e41dbb59a5b4810579c7d061fc

Observation 833fe562-249e-434f-8fb5-ff90f3d117ad · outbound

This paper cites DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.209618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.209618Z digest=sha256:582eda1bea23410bb0486010360926c219c87373fc960590856779949a9fad75

Observation 34151f3a-bd05-4fbf-9a39-1978b2539da7 · outbound

This paper cites an unresolved cited work.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:45:10.703802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.215010Z digest=sha256:29e8eaacf4284a76ca74f95271ef297b31e36e3e1aaca237bf40098d4d06bcef

Observation 1800aa40-4ab1-4cab-85ee-30f1711d8c9e · outbound

This paper cites MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.224664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.224664Z digest=sha256:b0947ab107da6ca9da29d5378c080b12121106d4a4e532136628d6115ae52246

Observation def0720d-1a4a-4ebb-835e-c95ad5a8c746 · outbound

This paper cites Pia: Your personalized image animator via plug-and-play modules in text-to-image models.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Pia: Your personalized image animator via plug-and-play modules in text-to-image models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.683823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.231845Z digest=sha256:886ab0431cf2623a5f84c6e79218c6ff47373c11d4298d667b846f1ebfa242d6

Observation 503325e1-b18b-4b60-b170-f8db274132fb · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Open-sora: Democratizing efficient video production for all, 2024

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.668299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.237158Z digest=sha256:2d19c29b2c5a039446f8eef7cc0b7d7b4015ee2f5f4a9f8bfbd682864b2ecf5f

Observation 180a3c7e-14d0-45b2-bee9-7b78ef2a8c7f · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Towards automatic learning of procedures from web instructional videos

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:45:10.645763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T21:45:10.243509Z digest=sha256:eb8540fc655440af4ceb5ff8ad5c6cb6477596570af80a901df2e2f3bae1c0fb

Observation e119dd37-e116-46a1-90e5-68fbe57e7a6b · outbound

This paper cites Motion Control for Enhanced Complex Action Video Generation.

HANDI: Hand-Centric Text-and-Image Conditioned Video Generation Motion Control for Enhanced Complex Action Video Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T21:45:10.247547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:45:10.247547Z digest=sha256:e67e481abfcd7776b104c023025886660379d2843e56703b0e44808a0598e202

Pith citing papers

Observation 8d701ebc-adbb-4fda-a20e-b6f577056511 · inbound

Towards Effective Human-in-the-Loop Assistive AI Agents cites this paper.

Towards Effective Human-in-the-Loop Assistive AI Agents HANDI: Hand-Centric Text-and-Image Conditioned Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:16:44.766696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:16:44.469541Z digest=sha256:df492a81dd100993133625e3e34adf1714c95a8a6aeba31d8ba128fab138fd1b