Pith. sign in

Paper Citation Record · LEDGER

PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2402.07872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07872 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:03:16.378325Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.793749Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2991d99c-93e6-4bb8-a16b-cb2a4873dcc4 · inbound

RT-H: Action Hierarchies Using Language cites this paper.

RT-H: Action Hierarchies Using Language PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:53:27.756324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T06:53:27.642020Z digest=sha256:966ed5bc7bd65f7c0a22eda3485e7e4f1a463d226b1dfe5240e083fe71f4e00c

Observation 8658726a-81ab-4629-9493-68cb57c47fea · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:17.946537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:5ce5a56b9e5e553176466b109784ecad500a80cc7b295ac470535bbe82811c3f

Observation 9712e9ff-b380-4a6d-aeb8-f7cbc3a4af07 · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:27:22.974511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:e619cecf74190e628d37c98385bf2232e95a6201773b855b0f912094d80988df

Observation 5b4a0c47-36b8-4475-a554-5451989a4704 · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.378325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.378325Z digest=sha256:0be1512444680a795adac9466e6c458a07f4617e5e6a892a4a65da77ab87ec4d

Observation bac584d5-64b5-4b8d-85e2-2279e245f1fa · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.226372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:9b224134613f1a210724c764f302e5d0bf38fb3034b4ca568206c7aac195305c

Observation 395c7c37-590c-48b4-bf69-aaf3dd4bf8a8 · inbound

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning cites this paper.

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:27:17.800649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T00:26:58.273861Z digest=sha256:869528cd16c5a5e9108535c989209159abd200c6ff3ac52341c16536fd9b64cf

Observation 753df86c-10a5-4030-87b2-d0e9b3d2b2d4 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.962878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:f1bf7026b569882e7146671b3f98e279dde9f45603711e8544cc77b7c8b76873

Observation af1f007e-283e-4536-b106-f94e0420e065 · inbound

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models cites this paper.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.833021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.833021Z digest=sha256:2f536afdc516851e566162f8353d7fdf98af48d4ae8bc9a80a7b0de4872e0d08

Observation 19d7957c-ead7-4c6a-af9c-1d95c5e5b4f9 · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:52.663287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:52.663287Z digest=sha256:4c3c7ea9c358c28c7463c75f6c7d651617cd86fdea375312ce5945b12bce0bda

Observation 2016e4c1-d9bc-4f68-bb0d-2fade28129d1 · inbound

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis cites this paper.

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:48:02.509867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:48:02.509867Z digest=sha256:dacb1773b90f49ea68cf8e36eba4cb36baee7fe5069b9abc8d61c0da7139fdbc

Observation e5c1427f-7675-42e3-8e7c-fe077f9071ec · inbound

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation cites this paper.

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.856037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:53.856037Z digest=sha256:26129cfd754a70a28e70c935c27bc16b9ad7c56c647e510fe1943dfa8502a54b

Observation f898e174-1805-4323-9a7f-b500b241d761 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.306305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.306305Z digest=sha256:3783ff480669186b709e921cdcb145f196cd6b7eda0ec81f177a9c6bcd9f3be2

Observation 6d45fc77-94b5-4864-a124-9f09338f21e3 · inbound

Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins cites this paper.

Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:00.413992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:00.413992Z digest=sha256:3b0e95ac05bba7f2d790672bead3572240c010ea730b11cbbc9b8caafe9b1f35

Observation facb93a0-e113-4668-a409-e49abe72a247 · inbound

Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models cites this paper.

Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:15.277552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:15.277552Z digest=sha256:1996ae7e4f1a1d3d0aea90e086bbca5aa161a1d1c8786aa3dbb245b23e546bb9

Observation 9593aab6-c73b-4d81-bdf0-6cc766bb02c1 · inbound

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity cites this paper.

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:43:15.132492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:43:15.132492Z digest=sha256:1bc3701c53d0ce5a3d2e6670a63c63422c826234622fc4bc15ca29d349ae47c9

Observation 8da321ad-6616-4dbe-a41e-20a59c43a8cc · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.661869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.661869Z digest=sha256:b13f9647e29ea5b83f01b188dc88cf8e28e45ecf5f3e5464224e9f2e56e86aa8

Observation ed9989cc-d9e3-4f85-a6fd-7b067971dd23 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.709761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.709761Z digest=sha256:c4354a3a5c8a837d0c4d6d0eb4c359eb47401bd1cc7e8ef902bb6c123765083b

Observation 0b29dbe9-3ca1-4c05-96b9-a4b48a22266c · inbound

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding cites this paper.

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:16.210236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:16.210236Z digest=sha256:2cc512dc2f916683d49478104693c4251130130ba7be1b0de31bdcee7b4d2c1b

Observation a08e6030-d0b3-4472-95c8-74a160fd17ec · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.282950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.282950Z digest=sha256:33e455579a05ea8d7215c95baf638a20a223f5cc9110c6d15f9406e94804190b

Observation 63494fdd-680f-4243-a7d3-231435e93d5c · inbound

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping cites this paper.

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:25.133699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:25.133699Z digest=sha256:843da0aebc51f2535823e3d23921f878d5952237fcd440bda0c9432133802c6c

Observation 844b8e2d-5cc2-4da6-a348-74db30750c5b · inbound

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies cites this paper.

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T21:44:58.996315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:44:58.996315Z digest=sha256:ca1be8437edef96f08f61ec80afa914a4f965e90f693534de03ad8c1f3577081

Observation 74a96285-36ca-470f-9429-f7ac7f3740d6 · inbound

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT cites this paper.

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:23:18.991949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:23:18.991949Z digest=sha256:1738c3b5751bab13c15e27dd3b1939ba33df42ad964a6bd04a33426a6d476d3c

Observation 2494ea49-5bac-4ac1-8540-65685f94adef · inbound

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models cites this paper.

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:28.691845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:28.691845Z digest=sha256:6a25ced71a9c9f7b8ef2fe6e45b74fef216fce71dee77f9fe0668b8d80a21983

Observation 3052474a-8db0-4b40-ab2a-d602a08bbd40 · inbound

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals cites this paper.

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:06.276723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:21:06.276723Z digest=sha256:b0ea4a77ad054772a095ff78185fac922cb273d0e9b0ef635676a850cfc19875

Observation 2ace252f-c4e0-4843-b1d3-24b59dcc3d43 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:52.414897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:52.414897Z digest=sha256:506dded81786d5d81c935499a7bc1111efd41117406c4075729b06644146d97c

Observation dd755236-743f-4464-9e78-88eb3b1ba4a9 · inbound

EVE: A Generator-Verifier System for Generative Policies cites this paper.

EVE: A Generator-Verifier System for Generative Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:36.225511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:36.225511Z digest=sha256:00b64edd5cfa2c0a02a2583e6418e4a5521536448ad301d307a3c875366c6d0a

Observation 6693fff6-e3fc-4240-9e81-4f5b1e23c063 · inbound

Visual-Language-Guided Task Planning for Horticultural Robots cites this paper.

Visual-Language-Guided Task Planning for Horticultural Robots PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:58:53.900823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:58:53.900823Z digest=sha256:89423b8d7bcc518af816d8ca460cd2662a31c425bd637fefa169acef5e6d83a9

Observation 3047f114-df0f-444d-9136-f0d4615c164b · inbound

Sem-NaVAE: Semantically-Guided Outdoor Mapless Navigation via Generative Trajectory Priors cites this paper.

Sem-NaVAE: Semantically-Guided Outdoor Mapless Navigation via Generative Trajectory Priors PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T05:43:21.189276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:43:21.189276Z digest=sha256:ed527663aed9714f5c4315c943e7f491429337faea517bdffa8455531fd93ab4

Observation b07d05f7-0c8d-40d3-97d8-29f0648ac513 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:33d850ee5d02ec7f86084600d0caee1e0c8623621c4990cf12339beb34ea5b1f

Observation 33c2c122-6f5b-4a7f-815d-bf60d59310c0 · inbound

JailWAM: Jailbreaking World Action Models in Robot Control cites this paper.

JailWAM: Jailbreaking World Action Models in Robot Control PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:48.143404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:22:13.890495Z digest=sha256:25a5b9ebe34ee4833df31a55baf2e0f812079c14dcea9321125f6969eb66e3f7

Observation a876af4c-f548-405c-9b3d-d3731957eca7 · inbound

Visually-grounded Humanoid Agents cites this paper.

Visually-grounded Humanoid Agents PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:04.606662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:10:32.710853Z digest=sha256:2fe377bde849377f34e3e5055ae0eaa61601bf4b4ed3c25e0fd3dbee41484e24

Observation 4942f695-e5c0-4837-bbda-d22b9cf0164e · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.824129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:eb351db1dddebd422191d5abb240d5f209c19596eeef7af710bbd9ebd6e29c3e

Observation 6b6447d3-9169-4cee-89e4-3107a2ef1ef0 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.256553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:4d30deb116b6b89023d2eac3e3bf62f6ce32c49fc18617e7fe2f58e112eb9350

Observation 656704d4-258c-4f0b-99f6-d44d53412430 · inbound

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning cites this paper.

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.795525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:48:49.071281Z digest=sha256:c73d831c2ae01d3871c298b38d830ad893de014bcfcde1090859414991e299b1

Observation bec7434e-6f77-4818-ba85-a7f72b2950db · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:56.957400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:25:21.796778Z digest=sha256:a2aafe226c5c6901ee61669a8fdf5ee90cb41407b4b887fe804baca30aaa609c

Observation c2a296ef-5bae-432c-a4be-7fbb039a02d8 · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.671107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T00:54:50.393828Z digest=sha256:8ae2e42a798a3b7dcbc4704b49bca5b8cdf9de92dcc65ef354f04e8d342c9ba7

Observation b385d68e-4993-401d-b552-27f589b8085f · inbound

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models cites this paper.

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.638351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:10:55.016700Z digest=sha256:3d8cf834d9206c17be2ee3d39a154be9bfe9f07ffc567c9d630496e5604587f1

Observation de5719f8-13bf-4504-889a-2a8122df22d1 · inbound

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance cites this paper.

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:41.952355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T05:17:25.534715Z digest=sha256:53acacf478f606e69a544e0d8d8bb2245e7e25d145dcca46ec37deb6fc98af84

Observation 5f1c91ea-0129-49e2-9af2-8b921cd5b863 · inbound

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance cites this paper.

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.682933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T19:15:49.118807Z digest=sha256:c5f2caf6dd7f81941eb375b6679cf9ccb1caa276274d514b88ab62955952372c

Observation cd4561a9-4371-447c-8fda-5928d69a8e9c · inbound

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies cites this paper.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:6fcab3cfbeab895b7b85900087e8f7a8b73f879458304a946a2487612bf7573e

Observation b67453e3-438b-4aeb-9b8c-403755cc4fee · inbound

IMBench: A Benchmark for Intuitive Robotic Manipulation cites this paper.

IMBench: A Benchmark for Intuitive Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:13.878543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:13.878543Z digest=sha256:fec84b736b9de2006a84c741b170772fd5605253b0ba31be999cf8b8f740dccb

Observation b2cec11d-36f2-49d3-9adf-56765bcaf9fe · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:50.899469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:50.899469Z digest=sha256:09bb7672ef81dffa32e51ebbeb7caa2f6d7bcbe67d91d2d971f41effa2e20d3b

Observation 9bb385f8-e7c9-409c-9286-d8a5d9ebee03 · inbound

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models cites this paper.

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T04:59:36.361796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:59:36.361796Z digest=sha256:f442286ce676a0c3eca6854e85204ef7eaec948b6644a46b54815751b031abe5

Observation ce1420cd-c1cb-421e-9729-6b2f8b0ce335 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.094779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.094779Z digest=sha256:b292d98805f4392d0c2bc51fc9c5a565ccb175961bef30ad5bddc028d5a7b639