Pith. sign in

Paper Citation Record · LEDGER

Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2403.12943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12943 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T01:01:49.881280Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:16:36.381522Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f54d54d-777c-4a64-807e-a2aad2cf201b · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.425532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:6bd76eb4089de86a42ccb09b54bab6e3f60c9961a348e65afa1b1d191739bb46

Observation 1b1f8b56-47d3-4cda-b77d-f2bcf6e103a9 · inbound

Action-Free Reasoning for Policy Generalization cites this paper.

Action-Free Reasoning for Policy Generalization Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T01:01:49.881280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:01:49.881280Z digest=sha256:84b48b2d845049d15eacf968c9328ff9f1657b137452f2ec30498920044a5f00

Observation 3a7c7053-cb28-479e-8e60-d78e5b54088f · inbound

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt cites this paper.

Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:25.743860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:51:25.743860Z digest=sha256:5a6c9c5451b5929a9f101fe8d1a8122c83dd2d82d0e6af6b43994999010f92be

Observation 10acbeeb-755b-47a5-817c-206621baabf8 · inbound

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos cites this paper.

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:36.054289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:51:36.054289Z digest=sha256:e95523fb5a4abef09a722f23e9990256d9877cf0c31e8058508051b0731dca17

Observation 6d471f07-c92a-459a-99ea-cfb68383ce45 · inbound

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations cites this paper.

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:33.758626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T00:37:12.170711Z digest=sha256:800296be9f831d5e7430d60c03a7b080316d945db0eabb540628decbdba2895a

Observation 9822926b-549b-4efd-8df1-35e56a64c3a1 · inbound

MonoDuo: Using One Robot Arm to Learn Bimanual Policies cites this paper.

MonoDuo: Using One Robot Arm to Learn Bimanual Policies Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.504103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:22:17.679021Z digest=sha256:9b0c775e3157f7600c5f00362e98f63ee1302d90468fae04c7e9f29ec6ba078f

Observation 22f8a5fe-0331-4c21-bd76-2aad2605fb70 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:27.849572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:980598b6d785c0c056a7b272d10f9f00b51b8f09328d22ed6935f5f8773fe03b

Observation a5ae3bd1-d1c1-412e-9acc-fa82ad88fd8c · inbound

SynthICL: Scalable In-context Imitation Learning with Synthetic Data cites this paper.

SynthICL: Scalable In-context Imitation Learning with Synthetic Data Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.875163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T19:23:45.402022Z digest=sha256:efedaffbeb3299582a05558c220fb7671d85b88c8855bae46732f8ed4a02b402

Observation bb398fcd-6270-4c20-b429-1c078dc459b9 · inbound

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning cites this paper.

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:56.423229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:35:25.740110Z digest=sha256:5e5d31548652c46a2ded637ada57472f2e9e3e63b44d7e4d2d92e25fd98c2c10

Observation 3dbc32e8-6aaa-4d65-bf1b-f936ee1834f3 · inbound

In-Context World Modeling for Robotic Control cites this paper.

In-Context World Modeling for Robotic Control Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:09.584646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T19:12:22.513577Z digest=sha256:be44d7bdb7365ca7436ae7c04fc86087d0ee4b98507ad193cd6b1bebacc567cb

Observation 984a43b8-6ba8-40d4-b44f-acfa9417ff3f · inbound

In-Context World Modeling for Robotic Control cites this paper.

In-Context World Modeling for Robotic Control Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:51.616019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:11:07.089829Z digest=sha256:f4907efcb3d8e32e7c8206496e0916d408b3b64f1823f8aea16c9e671ede5668

Observation a3d179df-d985-4252-8814-069ac488d1a8 · inbound

In-Context World Modeling for Robotic Control cites this paper.

In-Context World Modeling for Robotic Control Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T12:05:57.682386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:05:57.682386Z digest=sha256:43c3ce5833593e54596e5f51c3e6c539a4447fd28ce8efe38b7bbd0245deafbd

Observation 0d801c73-4984-416d-8b0f-34d423ede1ed · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:16:36.382745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T22:16:31.529359Z digest=sha256:fe7ecbb437dfe667d55afedf07adc83cb29274d9c4ce3743d19eec937687e469

Observation fcbc64e7-e5f8-4c8e-b8bd-09bcc44480ca · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:996ca71fd8eb1524c2570df00e44580d582cf967c39f57014693602497953e00