Pith. sign in

Paper Citation Record · LEDGER

Precise Action-to-Video Generation Through Visual Action Prompts

As of 9 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 4 inbound Pith citation observations for arXiv:2508.13104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13104 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:15:34.751856Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:36:01.254174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:49:53.218336Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 60256d56-1918-4bb7-b7ce-f244880d90be · outbound

This paper cites an unresolved cited work.

Precise Action-to-Video Generation Through Visual Action Prompts Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:13:33.172988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:13:33.172988Z digest=sha256:8b7655fb49f7204e65303c1b3d11c7a9bf8f6a857f1c244f0bc754c2e8e5e4b8

Observation 1daa2b1b-c9ed-47ff-9302-d7b07c4255a8 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Precise Action-to-Video Generation Through Visual Action Prompts Cosmos World Foundation Model Platform for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:14:34.361533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:14:34.361533Z digest=sha256:9487afdfa37cfaeec33c92963803cff67a12878927cf8a92ee822e04c80683d4

Observation 9a4cb098-bed9-4131-b721-3a24e55fcf06 · outbound

This paper cites InterDyn: Controllable Interactive Dynamics with Video Diffusion Models.

Precise Action-to-Video Generation Through Visual Action Prompts InterDyn: Controllable Interactive Dynamics with Video Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T19:14:51.841623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:14:51.841623Z digest=sha256:48849ed1d4086cdd2e4c5e4a5688636b4e3b67a782f8fb506ca387726015ad61

Observation 2cfdc77f-8959-4e24-b4a0-2b4519aab69d · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Precise Action-to-Video Generation Through Visual Action Prompts Diffusion for world modeling: Visual details matter in atari

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:04.622150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:04.622150Z digest=sha256:e2858b3845e4b4898b9be6425a9bf8ed3be0611657def51507a9d4e8d0b997e2

Observation 5075b4ea-386a-4733-be38-7dd996117d1d · outbound

This paper cites Qwen2.5-VL Technical Report.

Precise Action-to-Video Generation Through Visual Action Prompts Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:30.984449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:30.984449Z digest=sha256:600c3f24d70cbe25de7f39778103c7b7eb2c4606aebec1df8d69ca2a03a80576

Observation 70878d32-88ff-450a-9db4-193333146606 · outbound

This paper cites Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.262239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.262239Z digest=sha256:9feda626ea588dea72a6f2ccf0745069a8f2ab044db987240d5da392c569da08

Observation a755468e-3c01-4254-b336-fec4b6e284bb · outbound

This paper cites Lumiere: A space-time diffusion model for video generation.

Precise Action-to-Video Generation Through Visual Action Prompts Lumiere: A space-time diffusion model for video generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.315633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.315633Z digest=sha256:d2e212f7fb9ef09d2a0a6ad99f6af728acd22be709e176611ed210eb2d0410a8

Observation 8d3787ea-af80-409d-af2e-01179bc12842 · outbound

This paper cites Automatic rigging and anima- tion of 3d characters.

Precise Action-to-Video Generation Through Visual Action Prompts Automatic rigging and anima- tion of 3d characters

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.348171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.348171Z digest=sha256:34670cdf60a35b58c2daf45e29a3452535b839dfe4a0a32a0be95e30ecd3935f

Observation 78af8e13-775a-43df-af28-0c2d8a3144e8 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

Precise Action-to-Video Generation Through Visual Action Prompts Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.357011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.357011Z digest=sha256:a45747350f83de1e2db6bf6802c5b2ac07476cbf3739878be18ccc3a6fb8bb66

Observation a4738a33-af55-4a7b-b4a4-8dc1d9898485 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Precise Action-to-Video Generation Through Visual Action Prompts Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.367241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.367241Z digest=sha256:bc8c11d3103948e4efe10f88d42364532fac47effe5cd69f69ca7792437dc21e

Observation 2c495f89-e5a7-439f-8d27-ec80194f2e00 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Precise Action-to-Video Generation Through Visual Action Prompts RT-1: Robotics Transformer for Real-World Control at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.378635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.378635Z digest=sha256:ef535ff2657e6f680a37441e6c7b9b5333a49231977416c711676238bfaf46f7

Observation c203145b-f52a-4401-9386-5b8a5b1a257a · outbound

This paper cites Genie: Generative interactive environments.

Precise Action-to-Video Generation Through Visual Action Prompts Genie: Generative interactive environments

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.407516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.407516Z digest=sha256:41b19c5fecb99ca2dbb69573b8f224d0095c466fd3cffc7f10c8ba9b6787a895

Observation 8d5a3834-d60e-4088-bd2b-457991f717b8 · outbound

This paper cites 1C filter: a simple speed-based low-pass filter for noisy input in interac- tive systems.

Precise Action-to-Video Generation Through Visual Action Prompts 1C filter: a simple speed-based low-pass filter for noisy input in interac- tive systems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.425634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.425634Z digest=sha256:15e4616e0e0a9fbb14c105ca1cfa2ae7bf2c216c1cbf3a04a5c5f272a1a25bff

Observation b456ac54-02b6-492e-a7f8-ee9f91467e62 · outbound

This paper cites GameGen-X: Interactive Open-world Game Video Generation.

Precise Action-to-Video Generation Through Visual Action Prompts GameGen-X: Interactive Open-world Game Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.439193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.439193Z digest=sha256:358a37f054e39bff117f3746b34367456342c8a7c4f0867d5c4711bb538c49c1

Observation dfc4a088-9b51-4689-8f1e-d636e08ec80b · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

Precise Action-to-Video Generation Through Visual Action Prompts Scaling egocentric vision: The epic-kitchens dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.474949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.474949Z digest=sha256:fe287e863d3fb876f6a019232973072000ba011a4f105d39a8479dbbe7a16ffe

Observation 2da638c3-b55e-4aac-ade5-58b95b99c301 · outbound

This paper cites Oasis: A universe in a transformer,.

Precise Action-to-Video Generation Through Visual Action Prompts Oasis: A universe in a transformer,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.503297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.503297Z digest=sha256:49106f37620969ff6019de8a4e21e90eda61e4ef3171553385fa72b1eaefc098

Observation f3df8d26-b67a-4672-97e2-625ed288367b · outbound

This paper cites Genie 2: A large-scale foundation world model,.

Precise Action-to-Video Generation Through Visual Action Prompts Genie 2: A large-scale foundation world model,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.527442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.527442Z digest=sha256:2a9c9f48407e43ac78dd1f8b75a01ea06d82517c41ce701ecc84eacfe2b125ee

Observation 0596b06a-5d90-485a-89fc-5d833fc7f272 · outbound

This paper cites Motion capture from internet videos.

Precise Action-to-Video Generation Through Visual Action Prompts Motion capture from internet videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.553318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.553318Z digest=sha256:bfd3a90ab12d3671026dc3f8e9e01024a964359c14e68502552e509e35d55f13

Observation 745f3d2c-3b99-4f3f-b138-c0e89ca35bb0 · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

Precise Action-to-Video Generation Through Visual Action Prompts Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.584748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.584748Z digest=sha256:db555900c40f6b861c482d209ac15c86229d8e2865de3998fcd12a09fc3209e7

Observation b7c33008-0a3d-4f81-9981-1bf6483dec1d · outbound

This paper cites Arctic: A dataset for dexterous bimanual hand- object manipulation.

Precise Action-to-Video Generation Through Visual Action Prompts Arctic: A dataset for dexterous bimanual hand- object manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.613428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.613428Z digest=sha256:1e1edc47bcff2d3c60a4f9fd215c149062e9cb7149f838f5890bb6f70a238b3c

Observation 9187ef7e-f0f3-4c0b-bb64-a00df3bfad4d · outbound

This paper cites The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control.

Precise Action-to-Video Generation Through Visual Action Prompts The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.619881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.619881Z digest=sha256:a4849f21bf849514a4627a70aa2b636e995f3ca713e87c7440f496b187ed5865

Observation 3bc7d8dc-607f-4fc9-a95c-df317c42ff33 · outbound

This paper cites Gigahands: A massive annotated dataset of bimanual hand activities, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Gigahands: A massive annotated dataset of bimanual hand activities, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.634652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.634652Z digest=sha256:d4d19c69a8b7da585dba73139209078b69746147f6f1fadc92a4957d5021a982

Observation 2d6592d7-74cf-42d5-9acb-ce5335339722 · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.

Precise Action-to-Video Generation Through Visual Action Prompts Vista: A generalizable driving world model with high fidelity and versatile controllability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.650167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.650167Z digest=sha256:92a5e1f19459bb996498e5ae274819b5ac8e5c8173e39481ce7eb8e1d0c60bc9

Observation e67a4cd9-5fa8-426e-9919-e6c2f4ce8062 · outbound

This paper cites Motion Prompting: Controlling Video Generation with Motion Trajectories.

Precise Action-to-Video Generation Through Visual Action Prompts Motion Prompting: Controlling Video Generation with Motion Trajectories

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.659153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.659153Z digest=sha256:b52ea0f9faab00eccabc104e615dfa934b3ff45a080b3d2ea1bec3cec895b92d

Observation 0ea2ced9-fdde-4c0d-88cc-59a759ccb189 · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense, 2017.

Precise Action-to-Video Generation Through Visual Action Prompts The ”something something” video database for learning and evaluating visual common sense, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.670489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.670489Z digest=sha256:b945569470f78d0b2b3127582cfb0b6286985d6d6953b186b9b38f939db8e434

Observation 2243be9b-abab-4a12-abb2-846fb1109606 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Precise Action-to-Video Generation Through Visual Action Prompts Ego4d: Around the world in 3,000 hours of egocentric video

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.686672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.686672Z digest=sha256:4f1ebb4695eb82046d7bb3f5df36068f99f93d1601ef750a9494de3060574c8c

Observation b01cae2c-f5d8-4898-a66b-b83f3949bae1 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives.

Precise Action-to-Video Generation Through Visual Action Prompts Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.701792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.701792Z digest=sha256:74f40a4795b7234d4276b9496961bda5d0dc45633d5c8ae627a1e56d1f8b5935

Observation 8d956fe0-0779-45b0-a5fc-97ce2918c6dd · outbound

This paper cites MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training.

Precise Action-to-Video Generation Through Visual Action Prompts MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.712338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.712338Z digest=sha256:be40dd4d9be970496097adb3fbb1ae28325795420e0bc80282776b4a2f8ff322

Observation 7c1a7c24-223f-41cc-aef4-50d0f1328e73 · outbound

This paper cites Hand-eye calibration.

Precise Action-to-Video Generation Through Visual Action Prompts Hand-eye calibration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.729237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.729237Z digest=sha256:bab693ccd2e739ab874582deb40dff30f43975eee8926e042b71b9d3cc4ef043

Observation b25704f9-b7e5-4ec6-9c08-b1ef38a1f23c · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Precise Action-to-Video Generation Through Visual Action Prompts LoRA: Low-rank adaptation of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.734702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.734702Z digest=sha256:f07e0a2eee988d25830ac0c491102eba965489b153164d9ae338793883ed6fda

Observation 3ff48e37-f7db-43c7-8cc9-9007c51c5501 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

Precise Action-to-Video Generation Through Visual Action Prompts Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.739938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.739938Z digest=sha256:417c7df4f8ecfc627df50a08c21374b5b5e501c5c83cbfbc9b16b2298c1a3e1a

Observation 1aa7adb4-b970-4a1b-b514-7fbeeba0fa7d · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

Precise Action-to-Video Generation Through Visual Action Prompts Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.750331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.750331Z digest=sha256:86d6866c0040bb1b66066dd04472456c9d04b855eb6eecb2882409b0e75e9fbc

Observation 813318ee-f29c-464b-922a-6c6e56ba610f · outbound

This paper cites CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos.

Precise Action-to-Video Generation Through Visual Action Prompts CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.768262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.768262Z digest=sha256:2d0fe6bf6066ce0129f8f25260908796ce4279a23ed370fe460854cda7ef4f63

Observation 9132b549-d6a0-4de2-88d8-b5841df3d597 · outbound

This paper cites The kinetics human action video dataset, 2017.

Precise Action-to-Video Generation Through Visual Action Prompts The kinetics human action video dataset, 2017

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.778318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.778318Z digest=sha256:78b141d530532adf42ec195b4bc32ca9e016de4befd785f5ef1c86c271adf306

Observation d50df9b7-850b-4f6a-bcc9-30dccab68493 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Precise Action-to-Video Generation Through Visual Action Prompts DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.792222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.792222Z digest=sha256:5f539d003d12af0728fb1075f17a8f07ee5276aa20c723536c7776dd6b25ab3a

Observation a2725ab3-fc2f-4225-aedf-0470386e666b · outbound

This paper cites Sapiens: Foundation for human vision mod- els.

Precise Action-to-Video Generation Through Visual Action Prompts Sapiens: Foundation for human vision mod- els

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.854660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.854660Z digest=sha256:fdde291a9b804d67f544b35b71f501419e638e6a8f55c7e44a7d5adf1ed5cc1e

Observation f1065572-00dc-42ae-aabf-adf95f136f96 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

Precise Action-to-Video Generation Through Visual Action Prompts Learning to Act from Actionless Videos through Dense Correspondences

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.867862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.867862Z digest=sha256:fa681e871cec3d6eb3713da6737e56fe07f99beb9a685af4c1161a039026d617

Observation 0685dd8e-9c77-4bfd-9ec1-7256dccc8706 · outbound

This paper cites Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation.

Precise Action-to-Video Generation Through Visual Action Prompts Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.881776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.881776Z digest=sha256:ecc3103b04fa052088fc4a0dc3330156069c265023ddcf2abd3165004a1cdeaa

Observation b3e93b1b-f504-47ef-9f7e-fb46ff0a4d09 · outbound

This paper cites Wonderland: Navigating 3D Scenes from a Single Image.

Precise Action-to-Video Generation Through Visual Action Prompts Wonderland: Navigating 3D Scenes from a Single Image

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.963426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.963426Z digest=sha256:af3eaa7ed5c41e858e04917cc4f2407af337d1c639c33b6227223905934970db

Observation 0742e22e-6796-4822-8fdd-5499e550eeab · outbound

This paper cites Taco: Benchmarking general- izable bimanual tool-action-object understanding.

Precise Action-to-Video Generation Through Visual Action Prompts Taco: Benchmarking general- izable bimanual tool-action-object understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.124734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.124734Z digest=sha256:80688f956a29c56c5b4931962e735ff9fcd1b5a6b421e7d236b3c00afdeb3d22

Observation cce376c2-119d-4b1b-bb34-8dbc05b709fb · outbound

This paper cites The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences.

Precise Action-to-Video Generation Through Visual Action Prompts The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.211938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.211938Z digest=sha256:a68d0d8f7412612c582dc605208a81d1579b8f993b72cfa610c08896d9560585

Observation d10c0c7f-277f-4683-866b-e7f5739a9199 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Precise Action-to-Video Generation Through Visual Action Prompts MediaPipe: A Framework for Building Perception Pipelines

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.347462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.347462Z digest=sha256:585785dbceab89591e04e8ba22a9e0c3dc846bd865901a469189471df9cefd78

Observation 1c9e602f-c4d6-4f7f-befa-3be7ee9e27df · outbound

This paper cites MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling.

Precise Action-to-Video Generation Through Visual Action Prompts MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.454826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.454826Z digest=sha256:c561cee6de0cefbe3072eeb3a875ac73fe8d8a82874bfeb6f35b95a405666b79

Observation 60ca9129-a756-4b67-9b64-e36f1bd4e604 · outbound

This paper cites Do generative video models understand physical principles?.

Precise Action-to-Video Generation Through Visual Action Prompts Do generative video models understand physical principles?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.503772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.503772Z digest=sha256:b26ae27be2640abd4b7b557b7b1f5e2cbd56021760e6a76fb38609b51ffa05b2

Observation 81e2e1f7-a9e4-4e20-a133-5215c25c9563 · outbound

This paper cites A survey on deep learning for skeleton-based human animation.

Precise Action-to-Video Generation Through Visual Action Prompts A survey on deep learning for skeleton-based human animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.584211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.584211Z digest=sha256:fc1a47ca23c356afa7c7fee4607f799e2c7a1f178c8916eb0931ecbd73b8dc1e

Observation b6d7e9f7-ad75-4f5d-9473-5e51d7c31b0c · outbound

This paper cites Openai sora, 2023.

Precise Action-to-Video Generation Through Visual Action Prompts Openai sora, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.671158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.671158Z digest=sha256:9220e733ffa33a74209edc47bb0cc5395086ae4d04e3ce5811813071cceafe22

Observation 17bc754a-3a3e-4632-99da-c975150821bb · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Precise Action-to-Video Generation Through Visual Action Prompts Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.734752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.734752Z digest=sha256:085a88fc2e5c25340c52530a89be6512425801582a7578f94a59229e4e4dc557

Observation 5880a6a6-32bb-4b1c-9441-42314804c31a · outbound

This paper cites Computer animation: algorithms and techniques.

Precise Action-to-Video Generation Through Visual Action Prompts Computer animation: algorithms and techniques

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.828586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.828586Z digest=sha256:eb15fc69fd84abe47e40916108b1d78ca415bc09cb0c39d5a07d7a4c4feeeb37

Observation 66b4011d-db36-4198-9a9a-cb93a57612ff · outbound

This paper cites Scalable diffusion models with transformers.

Precise Action-to-Video Generation Through Visual Action Prompts Scalable diffusion models with transformers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.903369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.903369Z digest=sha256:c6e06de8c4ce73cf2df5169b57dd3eaf7f580885af82b197550e5b5be77b0ea4

Observation 8ad72c32-8f59-4344-b9e1-ce4d8a19e61a · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du.

Precise Action-to-Video Generation Through Visual Action Prompts Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.964980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.964980Z digest=sha256:f4fc3c19f44652394f743b864d7c8c9854bb4fe1b808009577936e00fc847f26

Observation 333eddb6-26d6-48b0-97e6-d63793da210e · outbound

This paper cites WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild.

Precise Action-to-Video Generation Through Visual Action Prompts WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.986642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.986642Z digest=sha256:6e712208c80c7f23608de3e956b9844bde025f357ea23c740a67518a0e108248

Observation c7867f6f-f8ec-4e64-a853-463835f4df2c · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Precise Action-to-Video Generation Through Visual Action Prompts SAM 2: Segment Anything in Images and Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.991524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.991524Z digest=sha256:4e33c6a77cb85f916f0613cc46d683e86dcfb03dd522b8935ad02c659d6e4294

Observation e1b1493d-301d-40c6-b942-fbfa0484ab2b · outbound

This paper cites World-grounded human motion recovery via gravity-view co- ordinates.

Precise Action-to-Video Generation Through Visual Action Prompts World-grounded human motion recovery via gravity-view co- ordinates

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.055594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.055594Z digest=sha256:f4bae054ce774c702201a2da730fb3208a3f5d886dbf5e6bed4e2402504482b8

Observation 56c0324b-d41f-4071-9e1f-4162b3135e57 · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video gen- eration with explicit motion modeling, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Motion-i2v: Consistent and controllable image-to-video gen- eration with explicit motion modeling, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.169774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.169774Z digest=sha256:9cf6259bcaf09539918a30370bbf9081648ab17a0c3e868dc4fb394533c66111

Observation 0a760d21-6dbc-4a72-b1c9-c4d977d0d174 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

Precise Action-to-Video Generation Through Visual Action Prompts Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.330539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.330539Z digest=sha256:b71316e37e0470d223da1cb25527916ffabddb53603ad80239b7ef473ede0404

Observation a7ddd85e-195c-4d2d-9b7a-6795245cb9b7 · outbound

This paper cites Optimal hand-eye cali- bration.

Precise Action-to-Video Generation Through Visual Action Prompts Optimal hand-eye cali- bration

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.437010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.437010Z digest=sha256:6e4cdc0d351be65723a93f30f410ad3a061ffc0d2187b75c610cc715a5276055

Observation 17874761-d284-4e65-ba06-b254102ef742 · outbound

This paper cites Controlling the world by sleight of hand.

Precise Action-to-Video Generation Through Visual Action Prompts Controlling the world by sleight of hand

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.534183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.534183Z digest=sha256:a11a7906ec7ac86b3640c13e54e87cdef0735857100562d994f48ac92ff915df

Observation 55ca3260-b9bd-4bf6-8d98-38558533cbfd · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Precise Action-to-Video Generation Through Visual Action Prompts Raft: Recurrent all-pairs field transforms for optical flow

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.604874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.604874Z digest=sha256:97a3959605865245aa1a6df280550ecdcece876667a0f6458f848cf4c5e261d6

Observation a59459fc-d4f8-468d-a52a-4ba3aedb1085 · outbound

This paper cites A new technique for fully autonomous and efficient 3 d robotics hand/eye calibration.

Precise Action-to-Video Generation Through Visual Action Prompts A new technique for fully autonomous and efficient 3 d robotics hand/eye calibration

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.719708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.719708Z digest=sha256:43c81b8b0a5763061f89b1e86dde1d3629ade30b92f650aab7392a50980ecde9

Observation caf6a23d-d957-4466-a80c-918dc681d3d6 · outbound

This paper cites Fvd: A new metric for video generation.

Precise Action-to-Video Generation Through Visual Action Prompts Fvd: A new metric for video generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.765796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.765796Z digest=sha256:d840a4566e0dfb9729c3f6098e3aab34bcb1309e1cc5bee99fd879bac0f7d0c9

Observation 69492a8f-750b-4b60-bbb9-21a373efdea6 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Precise Action-to-Video Generation Through Visual Action Prompts Diffusion Models Are Real-Time Game Engines

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.850238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.850238Z digest=sha256:839e19b731be1727bac5597df5af2e2a87ea894d9e47e8b8062d9cc781545b89

Observation 7e07eb07-a377-4dd4-b077-448780dd9982 · outbound

This paper cites Boximator: Generating Rich and Controllable Motions for Video Synthesis.

Precise Action-to-Video Generation Through Visual Action Prompts Boximator: Generating Rich and Controllable Motions for Video Synthesis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.008515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.008515Z digest=sha256:4e2f92b283c790d0fdce5980563aa6b96d592fdf1f9bb5152c7b502dd17eec31

Observation 21d70202-8869-4e07-a7eb-a46f1a171958 · outbound

This paper cites Motion Inversion for Video Customization.

Precise Action-to-Video Generation Through Visual Action Prompts Motion Inversion for Video Customization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.140230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.140230Z digest=sha256:72ba802eea3bcc9d0eb158d24943a91add1bbb7669651c2d7fafd6629c3782cb

Observation 8d4ee067-2a01-470c-a0e8-d8cbe31222ec · outbound

This paper cites EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation.

Precise Action-to-Video Generation Through Visual Action Prompts EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.292314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.292314Z digest=sha256:247c284503f2e472b2e5b2ed232c3382194423ea70f5bfac6c2c6642921f73de

Observation a9e55f6f-982f-4e2a-91d3-5804930451cb · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Precise Action-to-Video Generation Through Visual Action Prompts Image quality assessment: from error visibility to structural similarity

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.451746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.451746Z digest=sha256:a829a0528a4e0736d8f07803664e6e930cb413b76d6f63f1e9a058b93269d902

Observation 60b236e7-d8f2-49a4-9292-c167d04d171d · outbound

This paper cites Mo- tionctrl: A unified and flexible motion controller for video generation.

Precise Action-to-Video Generation Through Visual Action Prompts Mo- tionctrl: A unified and flexible motion controller for video generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.548947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.548947Z digest=sha256:3c71784902d3f9401e4870c59c6bb7d07233515d791ebba52b12ae3be2ac7410

Observation 8435d1d5-5001-42f7-a240-8e5b2d7f2baa · outbound

This paper cites ivideogpt: Interactive videogpts are scalable world models.

Precise Action-to-Video Generation Through Visual Action Prompts ivideogpt: Interactive videogpts are scalable world models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.655334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.655334Z digest=sha256:9f63ff6e8f7bddd6c3f2905aa09f083cf6bc2f1a5cc2a0c77c9178021c970b69

Observation 3efa7fda-a3a9-4916-9a96-170102a4f36a · outbound

This paper cites Samurai: Adapt- ing segment anything model for zero-shot visual tracking with motion-aware memory, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Samurai: Adapt- ing segment anything model for zero-shot visual tracking with motion-aware memory, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.751856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.751856Z digest=sha256:ed6db5ada14fa162c1582e9f97fbcfeef636d823f88a128247521307e1cc3723

Pith citing papers

Observation 5146c375-6962-4624-8511-7a7893229b71 · inbound

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints cites this paper.

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints Precise Action-to-Video Generation Through Visual Action Prompts

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:20:00.625747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:18:45.538658Z digest=sha256:c83223ca66e25ffa231080ec79b8a5f428af978cefe0d8f8d7e5cb0a0312180c

Observation 92200f97-2258-4608-8c8a-e82fe275d20b · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models Precise Action-to-Video Generation Through Visual Action Prompts

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:49:53.220072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:46:21.198781Z digest=sha256:b3b736d09d52c3bc719de34fe8024a0bb6d2b3004defd8a15e76fbd83690f15b

Observation 63f2e8db-c8e4-4c18-b917-fd3e1bf5c50d · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models Precise Action-to-Video Generation Through Visual Action Prompts

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:04:39.281373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T10:19:06.268547Z digest=sha256:29ed32d97a44e365bc8ae85b81aa17cfcc663b48aac9e392aba2be8661b1211e

Observation 58c0452d-2ab4-4841-a0c3-ca847f61293a · inbound

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models cites this paper.

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models Precise Action-to-Video Generation Through Visual Action Prompts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T17:36:01.254174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:36:01.254174Z digest=sha256:db387d96860f6bfe5029e2acab0feae18dc4a14b1df382a09b410203223c33f8