Pith. sign in

Paper Citation Record · LEDGER

Learning to Act from Actionless Videos through Dense Correspondences

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2310.08576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.08576 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:10:13.371215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:09:35.261339Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3aeba27b-6d61-45e8-823f-61b7b3fc3d7e · inbound

Any-point Trajectory Modeling for Policy Learning cites this paper.

Any-point Trajectory Modeling for Policy Learning Learning to Act from Actionless Videos through Dense Correspondences

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:32:50.976352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:32:50.916085Z digest=sha256:d40c78d490f2626e9954e014ad8c0a286e638b21b344fef5b649abc879feac29

Observation 8c4d3cbd-e360-410f-8c90-b25e27309042 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Learning to Act from Actionless Videos through Dense Correspondences

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:25:59.387306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:f8e93edc5a923c33d4910359c8a8c860d8fd07f699a2e07d63fa795b6a026ee8

Observation bb09e5b1-9a2a-4aa0-a4e6-37be6eeaa1ea · inbound

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning cites this paper.

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning Learning to Act from Actionless Videos through Dense Correspondences

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:06:09.770224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T16:06:09.448517Z digest=sha256:3b0e3f2c3a23fd5130456387c11353554df5a4d430016594290c5b22c08e29b6

Observation b85776dd-69b6-4ec7-b660-b452a7b2d027 · inbound

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation cites this paper.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Learning to Act from Actionless Videos through Dense Correspondences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.371215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.371215Z digest=sha256:b05ca382b30616591c538bf31a989ad714cb5cec109c70dc65f91bd44e84fa85

Observation 24052ae1-99a5-4399-beeb-237e10d3fe62 · inbound

Self-Consistent Model-based Adaptation for Visual Reinforcement Learning cites this paper.

Self-Consistent Model-based Adaptation for Visual Reinforcement Learning Learning to Act from Actionless Videos through Dense Correspondences

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T20:08:49.465586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:08:49.465586Z digest=sha256:1d9a837d53c3d90c4669bfde2792428ee714172fc9f90d2bf8f9c6011d2c50cf

Observation d561e5f9-dc15-4e00-ab53-803e0186867c · inbound

Unified Video Action Model cites this paper.

Unified Video Action Model Learning to Act from Actionless Videos through Dense Correspondences

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:50:29.752426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:50:29.675358Z digest=sha256:817be22c5fe0743cfaa4e5fa19b03c89b0243a5ebe333afb179d418316795f80

Observation 038e05b4-7263-4fee-b217-7f22183a200c · inbound

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model cites this paper.

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model Learning to Act from Actionless Videos through Dense Correspondences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:03:33.248255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:03:33.248255Z digest=sha256:50127a8a3428da40562aea8561df3f0f53d9d66693e9496ca6091b5db88fe30f

Observation a8b07aa3-b76d-49fb-8b1d-8c119778b33f · inbound

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos cites this paper.

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos Learning to Act from Actionless Videos through Dense Correspondences

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:28.228345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:28.228345Z digest=sha256:694ed519c3f91b518aaa3e197ea259c721e1fc37508dc938533d9a9a3dabb874

Observation b68a66b7-ebec-4d1c-b25d-4d87d2dafbc4 · inbound

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation cites this paper.

RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation Learning to Act from Actionless Videos through Dense Correspondences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:17:13.262031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:17:13.262031Z digest=sha256:7a68cee23715d9124bff0884cfa2efe6be15c8c5a2b09b9b55eb3e18cacf67e8

Observation 8a768ac1-4892-4d08-9cdf-6ec5cdfac315 · inbound

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations cites this paper.

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations Learning to Act from Actionless Videos through Dense Correspondences

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:37:07.533841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T06:36:13.144868Z digest=sha256:afe0c2dee56a5aabc57e44cacb559cac4af6eed0cbfeef6a6c2b2ea4856172cc

Observation f1065572-00dc-42ae-aabf-adf95f136f96 · inbound

Precise Action-to-Video Generation Through Visual Action Prompts cites this paper.

Precise Action-to-Video Generation Through Visual Action Prompts Learning to Act from Actionless Videos through Dense Correspondences

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:15:31.867862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:15:31.867862Z digest=sha256:fa681e871cec3d6eb3713da6737e56fe07f99beb9a685af4c1161a039026d617

Observation 5de35d0b-5b26-46dd-a4a7-6f8f71c7f439 · inbound

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation cites this paper.

Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation Learning to Act from Actionless Videos through Dense Correspondences

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T13:46:41.955554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:46:41.955554Z digest=sha256:10a47bfa2dc38955215d4d083d413f4659ea38f255cb6e3c8df939cdb19f876a

Observation be557eca-5c0d-4c5a-90cd-45f52f3f0f7c · inbound

IGen: Scalable Data Generation for Robot Learning from Open-World Images cites this paper.

IGen: Scalable Data Generation for Robot Learning from Open-World Images Learning to Act from Actionless Videos through Dense Correspondences

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:58:54.982338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:58:36.214948Z digest=sha256:9a0a4fc6a3b01f2d6a48e854b55381e8d34ba192cf7d1ea22ff93f799a6dffee

Observation f627f937-feab-4a91-a22a-6683e1f63285 · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment Learning to Act from Actionless Videos through Dense Correspondences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:49.386975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:49.386975Z digest=sha256:6b3fc10486173d7093703180f17d0536fca1019be2d0bfdef6152b195580ab1a

Observation eaa3dc8e-31ce-4fb1-9ffb-ea2304b89dfd · inbound

Large Video Planner Enables Generalizable Robot Control cites this paper.

Large Video Planner Enables Generalizable Robot Control Learning to Act from Actionless Videos through Dense Correspondences

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:33.873937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:26:32.048309Z digest=sha256:a64d4328b47083b3478acc248907b14b71ebbfefd2fee714d2d4625d44c2e625

Observation fff2440d-9bb9-4e7c-b318-be2b2652f23c · inbound

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning cites this paper.

On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning Learning to Act from Actionless Videos through Dense Correspondences

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T05:21:36.377047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:21:36.377047Z digest=sha256:71d9774fd0073a8f9930d769364792f69eb09555559f136f66209664b797f57f

Observation 5a2ec42a-bc32-4d94-940e-41693131aded · inbound

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model cites this paper.

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model Learning to Act from Actionless Videos through Dense Correspondences

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:58:08.919862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T18:54:07.081457Z digest=sha256:59c54efaf3f44c6bac758b2a268a71ca2e847c26a147214d903402a8ed582562

Observation 9c0a6822-3700-4022-8b99-9a83c66f8544 · inbound

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data cites this paper.

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data Learning to Act from Actionless Videos through Dense Correspondences

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:03:01.081202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T17:02:18.358675Z digest=sha256:bf8714e8c2c156a2be75dd8528397e076e40a4767ab5ee2e0e9433ec2fcb22ad

Observation 8a90e2a2-5dcc-41ab-b71c-37bfb2ec9c51 · inbound

Action Images: End-to-End Policy Learning via Multiview Video Generation cites this paper.

Action Images: End-to-End Policy Learning via Multiview Video Generation Learning to Act from Actionless Videos through Dense Correspondences

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:50:53.642876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:51:05.206602Z digest=sha256:9174b5d212f67bbff9f2681b0f8b792dedb50f07a5953b778c100cb6517417f4

Observation fb396572-8724-4147-8c5f-50bbdb92c5a6 · inbound

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation cites this paper.

ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation Learning to Act from Actionless Videos through Dense Correspondences

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:11:01.576249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:11:29.558490Z digest=sha256:19856ec191b06866c899583d991e4bceef6ce50c424d220e175cd05275bf41d6

Observation ecd53afa-9db9-4640-934c-4e50eb86db95 · inbound

VADF: Vision-Adaptive Diffusion Policy Framework for Efficient Robotic Manipulation cites this paper.

VADF: Vision-Adaptive Diffusion Policy Framework for Efficient Robotic Manipulation Learning to Act from Actionless Videos through Dense Correspondences

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:32:52.223880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:31:13.882275Z digest=sha256:543be8299f51ef464602fbfdd9996b5347533911b4145193fb10ecb2f3073b2b

Observation bfa1d90c-26cf-42e0-ade1-c92096b052c7 · inbound

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling cites this paper.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Learning to Act from Actionless Videos through Dense Correspondences

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.801872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:7421ae9239b2c3df1329a825ad02d17123343900bd69324d9960ba889a195670

Observation 9bf3b0d6-1ab6-4762-80b9-821ca37e7586 · inbound

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing cites this paper.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing Learning to Act from Actionless Videos through Dense Correspondences

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.562222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:c2231fd3360d226254332089cf89a134ce50f0cb4c500ba965da5f764e3ba98d

Observation 19837ab5-291c-43bd-9549-e7b647ed0adf · inbound

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations cites this paper.

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations Learning to Act from Actionless Videos through Dense Correspondences

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.452734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:51:28.068087Z digest=sha256:b21951406bd7be5f22230c5bc9bd89a09529a0eb65908c9434102e44bdeba570

Observation 610ccdbc-e91b-4615-b336-d6e68066a720 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Learning to Act from Actionless Videos through Dense Correspondences

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:02:17.821723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:3288178e451ea6cf498e6a28659bc024fa00906332c28a0fda76cca99c762cf9

Observation 14845897-427a-4a84-836f-04c454d00f3a · inbound

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking cites this paper.

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking Learning to Act from Actionless Videos through Dense Correspondences

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:29:28.946912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:28:12.151547Z digest=sha256:d1851981db9d66a8353c964ce1acc362f42ef8d0afc707a2aea624739990f6f5

Observation 27f43554-28ff-4d10-9a90-9533a7886926 · inbound

World Action Models: A Survey cites this paper.

World Action Models: A Survey Learning to Act from Actionless Videos through Dense Correspondences

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:35.263071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:11:12.686936Z digest=sha256:a556c0d0a146b1f515ccf4079e347d5f8d4f752fe3333948ee79ca79600da9a0

Observation e39ff17c-3763-4afc-996a-7e299882002b · inbound

From World Models to World Action Models: A Concise Tutorial for Robotics cites this paper.

From World Models to World Action Models: A Concise Tutorial for Robotics Learning to Act from Actionless Videos through Dense Correspondences

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:26:53.592737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T11:26:48.626947Z digest=sha256:c8349a3b2012826727fbb9dc7c01511791db16e37d40a07b4d4b273c8acfbd12

Observation 2575fae4-6302-4315-b8a9-346cb712e54e · inbound

From World Models to World Action Models: A Concise Tutorial for Robotics cites this paper.

From World Models to World Action Models: A Concise Tutorial for Robotics Learning to Act from Actionless Videos through Dense Correspondences

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.067037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T20:35:21.512564Z digest=sha256:b7243f3756b4e530f2baf89ebdc95154de035addc5027b1e67a47bb853f0572c

Observation b9c137da-9f8d-486f-8ef3-1faab3442251 · inbound

From World Models to World Action Models: A Concise Tutorial for Robotics cites this paper.

From World Models to World Action Models: A Concise Tutorial for Robotics Learning to Act from Actionless Videos through Dense Correspondences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T09:15:00.615740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:15:00.615740Z digest=sha256:c655cf76bb0c832970cd970a70146d71391c97efaf4f491e9d4c68ddbceed163

Observation 40bfbb25-6c19-4d74-a68b-b17e16312023 · inbound

From World Models to World Action Models: A Concise Tutorial for Robotics cites this paper.

From World Models to World Action Models: A Concise Tutorial for Robotics Learning to Act from Actionless Videos through Dense Correspondences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T09:18:01.832054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:18:01.832054Z digest=sha256:04e98b9f25c99e3d9221cc5e149dde80471be3d9e9a6c145677553a2c7ba3d33

Observation 1c28cbb1-fc72-4ba6-8969-a804d8b5f6f5 · inbound

Structured 4D Latent Predictive Model for Robot Planning cites this paper.

Structured 4D Latent Predictive Model for Robot Planning Learning to Act from Actionless Videos through Dense Correspondences

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:06:52.499232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T11:02:42.282246Z digest=sha256:a681514fa37f387149c20913e4a24a7b103925561d714296c9517e518c416b9c

Observation d97b8c95-b61e-4c8d-854f-b13625b305e5 · inbound

Masked Visual Actions for Unified World Modeling cites this paper.

Masked Visual Actions for Unified World Modeling Learning to Act from Actionless Videos through Dense Correspondences

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:03.044032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:03.044032Z digest=sha256:a3ed1270ad08bba317dfbe9559d04eb9704988f44a90f127ae8d9b0d56b005bb