Pith. sign in

Paper Citation Record · LEDGER

Precise Action-to-Video Generation Through Visual Action Prompts

As of 19 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 4 inbound Pith citation observations for arXiv:2508.13104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13104 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:15:34.751856Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:36:01.254174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:49:53.218336Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 60256d56-1918-4bb7-b7ce-f244880d90be · outbound

This paper cites an unresolved cited work.

Precise Action-to-Video Generation Through Visual Action Prompts Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T19:13:33.172988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:13:33.172988Z digest=sha256:798ec1c6a17778a399492ad8a091916d74ed04cfe4046179c6e1ad73949828a4

Observation 1daa2b1b-c9ed-47ff-9302-d7b07c4255a8 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Precise Action-to-Video Generation Through Visual Action Prompts Cosmos World Foundation Model Platform for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:14:34.361533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:14:34.361533Z digest=sha256:6c635fe94716c4b0a3bc5e800cb6e8b23ab851020d0f3a5451277afe0bed5e9a

Observation 9a4cb098-bed9-4131-b721-3a24e55fcf06 · outbound

This paper cites InterDyn: Controllable Interactive Dynamics with Video Diffusion Models.

Precise Action-to-Video Generation Through Visual Action Prompts InterDyn: Controllable Interactive Dynamics with Video Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T19:14:51.841623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:14:51.841623Z digest=sha256:fff17e70234f8ba6f3ef9f2ac13602e4c0623218939712fd983e8afa2145cd39

Observation 2cfdc77f-8959-4e24-b4a0-2b4519aab69d · outbound

This paper cites Diffusion for world modeling: Visual details matter in atari.

Precise Action-to-Video Generation Through Visual Action Prompts Diffusion for world modeling: Visual details matter in atari

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:04.622150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:04.622150Z digest=sha256:7eb7756ba945b219ebf2030b1db215c7a427708a73b28206efd2150492fa01b1

Observation 5075b4ea-386a-4733-be38-7dd996117d1d · outbound

This paper cites Qwen2.5-VL Technical Report.

Precise Action-to-Video Generation Through Visual Action Prompts Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:30.984449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:30.984449Z digest=sha256:c3d3cc648e02f5612bf40a6b3c19678992e26ca8c5d394f6efcf173942e6ef4a

Observation 70878d32-88ff-450a-9db4-193333146606 · outbound

This paper cites Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Introducing hot3d: An egocentric dataset for 3d hand and object tracking, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.262239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.262239Z digest=sha256:547af5b1e4a60505f5b2487893bf2255eb43b9d73654dcd366ecb581e55fdf8b

Observation a755468e-3c01-4254-b336-fec4b6e284bb · outbound

This paper cites Lumiere: A space-time diffusion model for video generation.

Precise Action-to-Video Generation Through Visual Action Prompts Lumiere: A space-time diffusion model for video generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.315633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.315633Z digest=sha256:2a348eec7063519e9a745b6e60b82b0ea9bd68068ebc70c00dac9ee76b09271a

Observation 8d3787ea-af80-409d-af2e-01179bc12842 · outbound

This paper cites Automatic rigging and anima- tion of 3d characters.

Precise Action-to-Video Generation Through Visual Action Prompts Automatic rigging and anima- tion of 3d characters

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.348171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.348171Z digest=sha256:06ac337dff017623a92590b6af617f43b61675a8644c5cd879f3cd3542579cb3

Observation 78af8e13-775a-43df-af28-0c2d8a3144e8 · outbound

This paper cites Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation.

Precise Action-to-Video Generation Through Visual Action Prompts Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.357011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.357011Z digest=sha256:d5c9bad0abdff7ca65083d5301abf6be551e54d6ac40e71569f9c41380e2e91a

Observation a4738a33-af55-4a7b-b4a4-8dc1d9898485 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Precise Action-to-Video Generation Through Visual Action Prompts Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.367241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.367241Z digest=sha256:e262f84d3839cabfb6648fb6d39bfc6411f5209ad8e38c214b1b7ca8b0a7df54

Observation 2c495f89-e5a7-439f-8d27-ec80194f2e00 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Precise Action-to-Video Generation Through Visual Action Prompts RT-1: Robotics Transformer for Real-World Control at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.378635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.378635Z digest=sha256:a4d04b946835041751616820d83dce987de0712615ad03ae9df103df23d5f6cb

Observation c203145b-f52a-4401-9386-5b8a5b1a257a · outbound

This paper cites Genie: Generative interactive environments.

Precise Action-to-Video Generation Through Visual Action Prompts Genie: Generative interactive environments

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.407516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.407516Z digest=sha256:973c6587c8682879c446c4be23f4159ae3b105e5543c9b81f22cd2fd217a3101

Observation 8d5a3834-d60e-4088-bd2b-457991f717b8 · outbound

This paper cites 1C filter: a simple speed-based low-pass filter for noisy input in interac- tive systems.

Precise Action-to-Video Generation Through Visual Action Prompts 1C filter: a simple speed-based low-pass filter for noisy input in interac- tive systems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.425634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.425634Z digest=sha256:7e7f97d229ed3d4842fc0252c9553dc47707668cddd7baa64c06ea9dd5edd939

Observation b456ac54-02b6-492e-a7f8-ee9f91467e62 · outbound

This paper cites GameGen-X: Interactive Open-world Game Video Generation.

Precise Action-to-Video Generation Through Visual Action Prompts GameGen-X: Interactive Open-world Game Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.439193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.439193Z digest=sha256:e6c9dc091d7072edf3b1956b2b71897873eb7fd071628faaabe4463f41467d70

Observation dfc4a088-9b51-4689-8f1e-d636e08ec80b · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

Precise Action-to-Video Generation Through Visual Action Prompts Scaling egocentric vision: The epic-kitchens dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.474949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.474949Z digest=sha256:aaf0069becd1ee742127d24c460fd21ebfe807efff125e23c7cdd1e2953b3662

Observation 2da638c3-b55e-4aac-ade5-58b95b99c301 · outbound

This paper cites Oasis: A universe in a transformer,.

Precise Action-to-Video Generation Through Visual Action Prompts Oasis: A universe in a transformer,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.503297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.503297Z digest=sha256:bf1baa52bcf690755783b6092f77e326c55a680ae972a2e0a62af485774967ec

Observation f3df8d26-b67a-4672-97e2-625ed288367b · outbound

This paper cites Genie 2: A large-scale foundation world model,.

Precise Action-to-Video Generation Through Visual Action Prompts Genie 2: A large-scale foundation world model,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.527442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.527442Z digest=sha256:46875863eb67d643e3452596629e83481a14fd2439b7f50d9eefb5914366790b

Observation 0596b06a-5d90-485a-89fc-5d833fc7f272 · outbound

This paper cites Motion capture from internet videos.

Precise Action-to-Video Generation Through Visual Action Prompts Motion capture from internet videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.553318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.553318Z digest=sha256:68f4486d13d450583f8106c2589697221664faeedd87b89f11dc6e017f896a85

Observation 745f3d2c-3b99-4f3f-b138-c0e89ca35bb0 · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

Precise Action-to-Video Generation Through Visual Action Prompts Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.584748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.584748Z digest=sha256:0151fae663ef4514ca47a92560c0081d3421d931e0ea9cca7967dcc42e5df855

Observation b7c33008-0a3d-4f81-9981-1bf6483dec1d · outbound

This paper cites Arctic: A dataset for dexterous bimanual hand- object manipulation.

Precise Action-to-Video Generation Through Visual Action Prompts Arctic: A dataset for dexterous bimanual hand- object manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.613428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.613428Z digest=sha256:fe6e7c54ce66d60c43513dc26a176f15a6606ef85855688d991e0e76c3fb19ba

Observation 9187ef7e-f0f3-4c0b-bb64-a00df3bfad4d · outbound

This paper cites The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control.

Precise Action-to-Video Generation Through Visual Action Prompts The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.619881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.619881Z digest=sha256:f49182517425f31ed362f26cc4a2874e39d7fa28197e5abc8ab30f4231fb8904

Observation 3bc7d8dc-607f-4fc9-a95c-df317c42ff33 · outbound

This paper cites Gigahands: A massive annotated dataset of bimanual hand activities, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Gigahands: A massive annotated dataset of bimanual hand activities, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.634652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.634652Z digest=sha256:2ff8cf6468bdd675ea42b2cdaae98b236e6e9c786a58ea60bf650b690e171a76

Observation 2d6592d7-74cf-42d5-9acb-ce5335339722 · outbound

This paper cites Vista: A generalizable driving world model with high fidelity and versatile controllability.

Precise Action-to-Video Generation Through Visual Action Prompts Vista: A generalizable driving world model with high fidelity and versatile controllability

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.650167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.650167Z digest=sha256:e26da60d934210c5f53186310df9f07b77b06764eb5a7cb19ff3c96ffc246a6d

Observation e67a4cd9-5fa8-426e-9919-e6c2f4ce8062 · outbound

This paper cites Motion Prompting: Controlling Video Generation with Motion Trajectories.

Precise Action-to-Video Generation Through Visual Action Prompts Motion Prompting: Controlling Video Generation with Motion Trajectories

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.659153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.659153Z digest=sha256:f7f66630db7dda0059aff107cc1274354205a4788a3e7f1df6994ec2dff3ac07

Observation 0ea2ced9-fdde-4c0d-88cc-59a759ccb189 · outbound

This paper cites The ”something something” video database for learning and evaluating visual common sense, 2017.

Precise Action-to-Video Generation Through Visual Action Prompts The ”something something” video database for learning and evaluating visual common sense, 2017

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.670489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.670489Z digest=sha256:9dcfe8321ea42f58498495a33da085119d0094d346f75aff899467ffff8035e7

Observation 2243be9b-abab-4a12-abb2-846fb1109606 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Precise Action-to-Video Generation Through Visual Action Prompts Ego4d: Around the world in 3,000 hours of egocentric video

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.686672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.686672Z digest=sha256:f07a1d5024b9b9e297f910f10816bcf7cb5707300cc7474803c38d636f6e105c

Observation b01cae2c-f5d8-4898-a66b-b83f3949bae1 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives.

Precise Action-to-Video Generation Through Visual Action Prompts Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.701792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.701792Z digest=sha256:ee61446e225425debe6c30c1dc4b23127de0700c5fbbc5be515f8352f04dd6dc

Observation 8d956fe0-0779-45b0-a5fc-97ce2918c6dd · outbound

This paper cites MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training.

Precise Action-to-Video Generation Through Visual Action Prompts MatchAnything: Universal Cross-Modality Image Matching with Large-Scale Pre-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.712338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.712338Z digest=sha256:447187e8535ad7262ac081eefdf6828811526e570cace5340245f243a29c3d01

Observation 7c1a7c24-223f-41cc-aef4-50d0f1328e73 · outbound

This paper cites Hand-eye calibration.

Precise Action-to-Video Generation Through Visual Action Prompts Hand-eye calibration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.729237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.729237Z digest=sha256:922793c4cb2dd9533af03ec4dc509cd8b907b503cbbd7898af60fecd19ec8eba

Observation b25704f9-b7e5-4ec6-9c08-b1ef38a1f23c · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Precise Action-to-Video Generation Through Visual Action Prompts LoRA: Low-rank adaptation of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.734702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.734702Z digest=sha256:8e6697fefe9bc07e184e37c2c97b401721336dc612b865f5a5ff4ddd32c45d24

Observation 3ff48e37-f7db-43c7-8cc9-9007c51c5501 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

Precise Action-to-Video Generation Through Visual Action Prompts Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.739938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.739938Z digest=sha256:aa737a5c47d6d53c443a06bd96c162a62582141ef1abea624040f070017da5dc

Observation 1aa7adb4-b970-4a1b-b514-7fbeeba0fa7d · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

Precise Action-to-Video Generation Through Visual Action Prompts Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.750331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.750331Z digest=sha256:93bf7e5bc271ad12992264fbc800eec6f2fcff32b89074113d56f36bb392e1c5

Observation 813318ee-f29c-464b-922a-6c6e56ba610f · outbound

This paper cites CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos.

Precise Action-to-Video Generation Through Visual Action Prompts CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.768262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.768262Z digest=sha256:a649e0a6fe517fa575eeba05b0efa42adce6baf566e4fd3c637e5c9706843957

Observation 9132b549-d6a0-4de2-88d8-b5841df3d597 · outbound

This paper cites The kinetics human action video dataset, 2017.

Precise Action-to-Video Generation Through Visual Action Prompts The kinetics human action video dataset, 2017

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.778318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.778318Z digest=sha256:8b3064a1f2f802fa8342d3f5d2c0a6d3dc5cbb62379216b8c4d970f5fb751795

Observation d50df9b7-850b-4f6a-bcc9-30dccab68493 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Precise Action-to-Video Generation Through Visual Action Prompts DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.792222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.792222Z digest=sha256:8837959f50afcb7c1e8a285a12f2f6f3186e8b8c40dd9a5ee823f0225a3a09e8

Observation a2725ab3-fc2f-4225-aedf-0470386e666b · outbound

This paper cites Sapiens: Foundation for human vision mod- els.

Precise Action-to-Video Generation Through Visual Action Prompts Sapiens: Foundation for human vision mod- els

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.854660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.854660Z digest=sha256:995ee9003a053e96458bbca8dcd0834945f5351fdf36e040c3044c0c5281a3fd

Observation f1065572-00dc-42ae-aabf-adf95f136f96 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

Precise Action-to-Video Generation Through Visual Action Prompts Learning to Act from Actionless Videos through Dense Correspondences

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.867862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.867862Z digest=sha256:dfa88168a0e9871aa783f7e0066cc51ea6882341a948ff77f92a0bb3d7a7036d

Observation 0685dd8e-9c77-4bfd-9ec1-7256dccc8706 · outbound

This paper cites Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation.

Precise Action-to-Video Generation Through Visual Action Prompts Pose space deformation: a unified approach to shape interpolation and skeleton-driven deformation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.881776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.881776Z digest=sha256:808b60b78aad388c000b48fe43ebf52351abff8f39004457645afea3283b3e23

Observation b3e93b1b-f504-47ef-9f7e-fb46ff0a4d09 · outbound

This paper cites Wonderland: Navigating 3D Scenes from a Single Image.

Precise Action-to-Video Generation Through Visual Action Prompts Wonderland: Navigating 3D Scenes from a Single Image

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.963426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.963426Z digest=sha256:12489a577bf5eb2e5f0c71473a4a047c84c980c0fc128c598e6e034b6dd4080e

Observation 0742e22e-6796-4822-8fdd-5499e550eeab · outbound

This paper cites Taco: Benchmarking general- izable bimanual tool-action-object understanding.

Precise Action-to-Video Generation Through Visual Action Prompts Taco: Benchmarking general- izable bimanual tool-action-object understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.124734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.124734Z digest=sha256:7260a77df6d22d191e06d72db9933b9e619cfd97fec9edb2d8d2f0a5cf614c4d

Observation cce376c2-119d-4b1b-bb34-8dbc05b709fb · outbound

This paper cites The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences.

Precise Action-to-Video Generation Through Visual Action Prompts The BabyView dataset: High-resolution egocentric videos of infants' and young children's everyday experiences

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.211938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.211938Z digest=sha256:2d18c1a57088365463daffdd3a8976a565e3953cc58e69ab837da054b0d39c01

Observation d10c0c7f-277f-4683-866b-e7f5739a9199 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Precise Action-to-Video Generation Through Visual Action Prompts MediaPipe: A Framework for Building Perception Pipelines

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.347462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.347462Z digest=sha256:6d26a14689ab0adac962139ac70a75009891f3599765ff518d7858fabfb39a8f

Observation 1c9e602f-c4d6-4f7f-befa-3be7ee9e27df · outbound

This paper cites MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling.

Precise Action-to-Video Generation Through Visual Action Prompts MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.454826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.454826Z digest=sha256:3260974a19d4beda2dc0485959c176a2d7626abce0d6ae542f5acf4fc4054f20

Observation 60ca9129-a756-4b67-9b64-e36f1bd4e604 · outbound

This paper cites Do generative video models understand physical principles?.

Precise Action-to-Video Generation Through Visual Action Prompts Do generative video models understand physical principles?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.503772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.503772Z digest=sha256:598f25ee23360f3d286a28f109437993672562ad5cbdef950cc9d6a5f9bb8b1e

Observation 81e2e1f7-a9e4-4e20-a133-5215c25c9563 · outbound

This paper cites A survey on deep learning for skeleton-based human animation.

Precise Action-to-Video Generation Through Visual Action Prompts A survey on deep learning for skeleton-based human animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.584211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.584211Z digest=sha256:dab3f24d7dd68aa6ca7340ba8d2e7e620024218ef2b961839dc61e4c72a26544

Observation b6d7e9f7-ad75-4f5d-9473-5e51d7c31b0c · outbound

This paper cites Openai sora, 2023.

Precise Action-to-Video Generation Through Visual Action Prompts Openai sora, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.671158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.671158Z digest=sha256:f8f47f0eb64d61e4025b57f961c4978b90285f15e4cc3c9d242090034f564ce8

Observation 17bc754a-3a3e-4632-99da-c975150821bb · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Precise Action-to-Video Generation Through Visual Action Prompts Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.734752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.734752Z digest=sha256:cf205b31e4e20b3e5d56595b9fbb715a5704037d2ccdd7b059c17db8c6f329a3

Observation 5880a6a6-32bb-4b1c-9441-42314804c31a · outbound

This paper cites Computer animation: algorithms and techniques.

Precise Action-to-Video Generation Through Visual Action Prompts Computer animation: algorithms and techniques

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.828586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.828586Z digest=sha256:65ef29a0db072f57f2c05418c0d26ad7b46d492a10d6c186daffd097cf925dbd

Observation 66b4011d-db36-4198-9a9a-cb93a57612ff · outbound

This paper cites Scalable diffusion models with transformers.

Precise Action-to-Video Generation Through Visual Action Prompts Scalable diffusion models with transformers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.903369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.903369Z digest=sha256:0e9036614455ef3b4ce5f171268f5f915905bba55b202749c5fe6894c501afdd

Observation 8ad72c32-8f59-4344-b9e1-ce4d8a19e61a · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du.

Precise Action-to-Video Generation Through Visual Action Prompts Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.964980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.964980Z digest=sha256:6a05e3c1017aaf3261671275602fefe1287288db3ee0f539e60b80527b0df728

Observation 333eddb6-26d6-48b0-97e6-d63793da210e · outbound

This paper cites WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild.

Precise Action-to-Video Generation Through Visual Action Prompts WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.986642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.986642Z digest=sha256:d5478f0d825186e3e1f3869889381d09b4f24b9333ecc58d11fe3c20971022b8

Observation c7867f6f-f8ec-4e64-a853-463835f4df2c · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Precise Action-to-Video Generation Through Visual Action Prompts SAM 2: Segment Anything in Images and Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:32.991524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:32.991524Z digest=sha256:8f07733c17115c50e747fc9a4e0b7123b547e61ca870a2013e82da5571b157e3

Observation e1b1493d-301d-40c6-b942-fbfa0484ab2b · outbound

This paper cites World-grounded human motion recovery via gravity-view co- ordinates.

Precise Action-to-Video Generation Through Visual Action Prompts World-grounded human motion recovery via gravity-view co- ordinates

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.055594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.055594Z digest=sha256:27a8c1bfd283e2709f810d6c0739ffc11fd8ff2111d806bae124e0864f4a89b4

Observation 56c0324b-d41f-4071-9e1f-4162b3135e57 · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video gen- eration with explicit motion modeling, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Motion-i2v: Consistent and controllable image-to-video gen- eration with explicit motion modeling, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.169774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.169774Z digest=sha256:daec8658477a8e4f1de7b40c3bab989a0a5b0d3fa9528f4d6090ebad1258c4f4

Observation 0a760d21-6dbc-4a72-b1c9-c4d977d0d174 · outbound

This paper cites Genhowto: Learning to generate actions and state transformations from instructional videos.

Precise Action-to-Video Generation Through Visual Action Prompts Genhowto: Learning to generate actions and state transformations from instructional videos

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.330539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.330539Z digest=sha256:bf4dcb48ddb4d01e58f830bcc85cee5f3162dbd62985954d810141364003b3c3

Observation a7ddd85e-195c-4d2d-9b7a-6795245cb9b7 · outbound

This paper cites Optimal hand-eye cali- bration.

Precise Action-to-Video Generation Through Visual Action Prompts Optimal hand-eye cali- bration

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.437010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.437010Z digest=sha256:491f95335d4dcbc0259d0a6613e99425abf6aa0888a3445729087f11d0d35634

Observation 17874761-d284-4e65-ba06-b254102ef742 · outbound

This paper cites Controlling the world by sleight of hand.

Precise Action-to-Video Generation Through Visual Action Prompts Controlling the world by sleight of hand

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.534183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.534183Z digest=sha256:16b5aa3ff5c9fda32e9cac2ab265b700fa25e5b7109d88059afd847cacf548c1

Observation 55ca3260-b9bd-4bf6-8d98-38558533cbfd · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Precise Action-to-Video Generation Through Visual Action Prompts Raft: Recurrent all-pairs field transforms for optical flow

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.604874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.604874Z digest=sha256:be9e4cc0476ecb65a411e0eeb83cf26084bae5cd18e734d6abe927107b87dba2

Observation a59459fc-d4f8-468d-a52a-4ba3aedb1085 · outbound

This paper cites A new technique for fully autonomous and efficient 3 d robotics hand/eye calibration.

Precise Action-to-Video Generation Through Visual Action Prompts A new technique for fully autonomous and efficient 3 d robotics hand/eye calibration

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.719708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.719708Z digest=sha256:a9d52a7978893bef499b548a516b718e81d9a36a2c5a881ab98f8908d989f340

Observation caf6a23d-d957-4466-a80c-918dc681d3d6 · outbound

This paper cites Fvd: A new metric for video generation.

Precise Action-to-Video Generation Through Visual Action Prompts Fvd: A new metric for video generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.765796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.765796Z digest=sha256:52f5484fbe045aff490fd5fef0602d97414592a00301bf971e4d058c233cf12c

Observation 69492a8f-750b-4b60-bbb9-21a373efdea6 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Precise Action-to-Video Generation Through Visual Action Prompts Diffusion Models Are Real-Time Game Engines

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:33.850238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:33.850238Z digest=sha256:e1d3b4848a67545f68a61536b08f1212ad2436cba98c5d1343f0ba73e9172fbc

Observation 7e07eb07-a377-4dd4-b077-448780dd9982 · outbound

This paper cites Boximator: Generating Rich and Controllable Motions for Video Synthesis.

Precise Action-to-Video Generation Through Visual Action Prompts Boximator: Generating Rich and Controllable Motions for Video Synthesis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.008515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.008515Z digest=sha256:94d33a58ef84d82d3d26552ea5b6785a58e12aded8ded7735b5bdad31403edd7

Observation 21d70202-8869-4e07-a7eb-a46f1a171958 · outbound

This paper cites Motion Inversion for Video Customization.

Precise Action-to-Video Generation Through Visual Action Prompts Motion Inversion for Video Customization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.140230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.140230Z digest=sha256:2fc9f367692f2b4f45c2cf151d1ce705c792a2936092383367dfcfcf5c361b68

Observation 8d4ee067-2a01-470c-a0e8-d8cbe31222ec · outbound

This paper cites EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation.

Precise Action-to-Video Generation Through Visual Action Prompts EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.292314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.292314Z digest=sha256:e35df57f2fcfef9369beb1413a559145d6afbaf8a738f644b972446f6244811a

Observation a9e55f6f-982f-4e2a-91d3-5804930451cb · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Precise Action-to-Video Generation Through Visual Action Prompts Image quality assessment: from error visibility to structural similarity

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.451746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.451746Z digest=sha256:712a156992d261c3fa8f98ac0f671b62c10ac2bf30dcaec8094dc414d2ecd5bf

Observation 60b236e7-d8f2-49a4-9292-c167d04d171d · outbound

This paper cites Mo- tionctrl: A unified and flexible motion controller for video generation.

Precise Action-to-Video Generation Through Visual Action Prompts Mo- tionctrl: A unified and flexible motion controller for video generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.548947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.548947Z digest=sha256:15160b9440c6a06a7322f27d5b233001fd0c2668fd14566e885569f118c5bd7c

Observation 8435d1d5-5001-42f7-a240-8e5b2d7f2baa · outbound

This paper cites ivideogpt: Interactive videogpts are scalable world models.

Precise Action-to-Video Generation Through Visual Action Prompts ivideogpt: Interactive videogpts are scalable world models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.655334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.655334Z digest=sha256:fc88c61b09987df6dfc41cdd38f32daab7b7b48da0513ab630006c044a725621

Observation 3efa7fda-a3a9-4916-9a96-170102a4f36a · outbound

This paper cites Samurai: Adapt- ing segment anything model for zero-shot visual tracking with motion-aware memory, 2024.

Precise Action-to-Video Generation Through Visual Action Prompts Samurai: Adapt- ing segment anything model for zero-shot visual tracking with motion-aware memory, 2024

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:34.751856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:34.751856Z digest=sha256:4bf3152c2700dec3f118b19b01e145ad978ce2ff9af5c156b43d9c60dda3872b

Pith citing papers

Observation 5146c375-6962-4624-8511-7a7893229b71 · inbound

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints cites this paper.

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints Precise Action-to-Video Generation Through Visual Action Prompts

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:20:00.625747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T12:18:45.538658Z digest=sha256:9b283e2a24f331e9db52ce821ce4e0f735939afb60db18c23213ad19eb0ab221

Observation 92200f97-2258-4608-8c8a-e82fe275d20b · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models Precise Action-to-Video Generation Through Visual Action Prompts

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:49:53.220072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T05:46:21.198781Z digest=sha256:169ce4a1c9a721cd5f3b478c3438cdf00ff34a96162b248e4ca896785b98b35a

Observation 63f2e8db-c8e4-4c18-b917-fd3e1bf5c50d · inbound

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models cites this paper.

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models Precise Action-to-Video Generation Through Visual Action Prompts

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:04:39.281373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T10:19:06.268547Z digest=sha256:4f11518b0356015b2799bbe784036e3cfebb8a36ca0dc798bf4b0ab946bb92d3

Observation 58c0452d-2ab4-4841-a0c3-ca847f61293a · inbound

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models cites this paper.

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models Precise Action-to-Video Generation Through Visual Action Prompts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T17:36:01.254174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:36:01.254174Z digest=sha256:f17516033001da3c41626dfd1add2c037ff3e314dff048ba4fb605edeb45dacd