Pith. sign in

Paper Citation Record · LEDGER

CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2212.05711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.05711 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:04:38.255102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T06:17:42.050023Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ccee9b44-a069-4530-ba15-426c82564d49 · inbound

Scaling Robot Learning with Semantically Imagined Experience cites this paper.

Scaling Robot Learning with Semantically Imagined Experience CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:59:10.467994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:59:10.352342Z digest=sha256:d6b198e58c9fc64e99657d1f1e2ab55425d5eb57f66358217cce1c2e4bc000f2

Observation 1438a73a-8b2d-4ccc-b535-d6a738a3bedd · inbound

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models cites this paper.

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:57:22.445056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T08:57:22.299028Z digest=sha256:c9a1947a8998362d70f2a37123fc606cb1093c4a3254ea6214e201a094f307d0

Observation faae6cf2-2bfb-4cbd-85c7-166429e7c1f8 · inbound

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models cites this paper.

Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:54:59.037935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:54:58.940428Z digest=sha256:ae1838764cead924dfc7462ea943a3cf96d78ae805fc4d1daea7e7187618e894

Observation 5ba7c19e-629c-431e-a009-ffc71f5ca888 · inbound

MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations cites this paper.

MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:47:55.246932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T09:47:54.977716Z digest=sha256:4632bc2e7d016967b27335a0892b1188ff7dd4ce2d9e1874146847bd8ab86a8d

Observation 46a2aa4f-ce57-45bf-809b-1a9cc977d222 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:43:11.241612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:ce29e4ffeb9f939c21e1f42729ad142903387a82d3b479b8b4b24c641e22e48e

Observation d4b4610d-f874-4e4d-8bde-0ad1a9df96eb · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:17:01.409154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:7b2b114ee4370955eae64672e1a82210bd26e5d024ceb03c81a06da49bd83d50

Observation 94c6acde-dc36-4f39-a18a-8cee3b614b53 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T19:09:10.236670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:c11700a3924d850ccc78e5fe0ef4aa4d699c788f493ce5f72e08d6d6a06c96e0

Observation 5c4b2a0c-d593-4718-a61e-237dc2e49515 · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:50:45.479464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:4dfbfe030a60e3491d7ed758d2cfa30a1a3968409c08a30b70e6d91a35e6f184

Observation eec1cdd2-a854-4cb0-8a0b-7e5119b578db · inbound

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot cites this paper.

RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:04:38.255102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:04:38.255102Z digest=sha256:dd36accc21497eaedaddab2afc5085bcf8db8298d1bd8b2a203596420fdbff4a

Observation 242ca6d9-c30b-4d47-843e-1ce637131503 · inbound

VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation cites this paper.

VLM-TDP: VLM-guided Trajectory-conditioned Diffusion Policy for Robust Long-Horizon Manipulation CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:08.462855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:08.462855Z digest=sha256:99945aeaa6c9c06e636390b0bb5e763efbfce1b4d19483ece0621f5f280849b1

Observation 9ca5a907-6541-4edd-be98-86e26c94919c · inbound

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents cites this paper.

ERMV: Editing 4D Robotic Multi-view images to enhance embodied agents CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:31.294528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:31.294528Z digest=sha256:674ebe1b4c29106b8532dc48fff9164b4a28f0de36568ef3642843447be0b427

Observation fa351bf5-fd5a-4939-b1c8-8ef92aeee40c · inbound

OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction cites this paper.

OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T13:33:46.495390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:33:46.495390Z digest=sha256:b52b13c5765f83f288d3242e12b23fe4a3c6f99baa99d75bfee432013736dc5f

Observation 2de5ce1b-6051-45ce-b39f-4ad8f6005dfc · inbound

RoboLight: A Dataset with Linearly Composable Illumination for Robotic Manipulation cites this paper.

RoboLight: A Dataset with Linearly Composable Illumination for Robotic Manipulation CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T18:56:44.946861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:56:44.946861Z digest=sha256:edecd82c102dfd39afbff149f7d416a173958848d2aa6ff88bca5a1ea35a0908

Observation c0ef0fc2-307a-41ce-9fd5-e5f940d02a43 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:45:21.599029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:a3c5b0a84160bdf83faf002d9be61197ee1e7fe46e31f8d88775c27d2e611166

Observation c8d6dc5e-7358-4aac-80a2-eccc76018574 · inbound

Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation cites this paper.

Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:08.700462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T19:48:57.065807Z digest=sha256:d4f1d1d4eedd9781cac0a0204692a76d45a43857b09bba996c4884432ca00cda

Observation 32221a29-b551-41e3-87d5-5b20e190cd69 · inbound

Task Robustness via Re-Labelling Vision-Action Robot Data cites this paper.

Task Robustness via Re-Labelling Vision-Action Robot Data CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T06:17:42.052484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T12:46:56.049470Z digest=sha256:fea930b8e2667db0b64af1392cb0ef2f54668bac15b79d67c379a8a169b7c5da

Observation 2410ff9d-bb4a-43ea-8be0-7ac1c302dd34 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills CACTI: A Framework for Scalable Multi-Task Multi-Scene Visual Imitation Learning

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.019264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.019264Z digest=sha256:acc53d6c7970ce4293e91c539b9adf33905bc58aa9e6fbc64e4b4235b1b1c5af