Pith. sign in

Paper Citation Record · LEDGER

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2503.07135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07135 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:01:16.616474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T11:01:02.081938Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3122f98a-e85a-494a-94b6-ed1f2baa0346 · inbound

OpenTie: Open-vocabulary Sequential Rebar Tying System cites this paper.

OpenTie: Open-vocabulary Sequential Rebar Tying System VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:19:17.903642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:19:17.903642Z digest=sha256:c7bbb0a5582288715f89388ec611f6a92649cee8ce76d8a9b674eadbf9753d3f

Observation 2be61ed7-57b6-4d4a-bee5-4819472138c9 · inbound

Calib3R: Hand-Eye Calibration and 3D Metric-Scaled Scene Reconstruction with 3D Foundation Models cites this paper.

Calib3R: Hand-Eye Calibration and 3D Metric-Scaled Scene Reconstruction with 3D Foundation Models VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:47.067346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:47.067346Z digest=sha256:747e11c4cd4fce5f6252324c36942ee5dd4da9605dab9141e3cf45356952b03a

Observation f72c2d64-4e90-42a6-9ae8-083e281d9cf0 · inbound

3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos cites this paper.

3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-15T12:36:50.019492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:36:50.019492Z digest=sha256:ab4323f22d5284be540ad0cddcdc7d522e62ab8ead874bb93925aa84757d034d

Observation 6fe2d1e5-deb1-4a79-a5dc-1aa4945546b0 · inbound

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations cites this paper.

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:02.088600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:14:24.932972Z digest=sha256:e372511d8d1ceb945e629b92e6a98777b1bde5cfb253a895308bed8eae4b5951

Observation 1a14e57d-602e-4b17-9612-b82b3bf0bf1c · inbound

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances cites this paper.

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:01:16.616474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:01:16.616474Z digest=sha256:cd7e92c97984911b1478d72f929868480f9339ce1e9d1f5ef2c01374c2362532