Pith. sign in

Paper Citation Record · LEDGER

PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2402.07872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07872 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:15:05.982335Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.793749Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2991d99c-93e6-4bb8-a16b-cb2a4873dcc4 · inbound

RT-H: Action Hierarchies Using Language cites this paper.

RT-H: Action Hierarchies Using Language PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:53:27.756324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T06:53:27.642020Z digest=sha256:ccbd3edd19384479b64d80b1455ae3325ba6e59079ab69801d31be10e0ec425e

Observation 8658726a-81ab-4629-9493-68cb57c47fea · inbound

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation cites this paper.

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:25:17.946537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T08:25:17.847571Z digest=sha256:df96ed34278b698d914c70f2dbd3abe08a8b1316f35183de5005bd59952b9ed1

Observation 9712e9ff-b380-4a6d-aeb8-f7cbc3a4af07 · inbound

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies cites this paper.

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:27:22.974511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T18:27:22.760982Z digest=sha256:457ff96ea72219535d21965df6861fa663b0cbb07d745046631fa5eb3a96f7f8

Observation f906147b-9869-40ed-bbc1-58d8f34ed453 · inbound

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints cites this paper.

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:50:41.277102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:50:41.277102Z digest=sha256:080052a6b400ab172e4cc25377d825133a59c10620464afde0d88484a0b70ca8

Observation 1aa3c740-11bd-45c0-94c3-fa33bedd3349 · inbound

VLM-driven Behavior Tree for Context-aware Task Planning cites this paper.

VLM-driven Behavior Tree for Context-aware Task Planning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:45:22.615279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:45:22.615279Z digest=sha256:93ed9f7b4649943f6535fe74782e46babcdeeab3a73f6688fc40c8221e838b23

Observation a2476a74-ef01-4a2f-9994-a2999e1149e1 · inbound

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models cites this paper.

Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:46:17.569537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:46:17.569537Z digest=sha256:075be72fef3f2d37367770863f3c5691d0928cd3959cbaf476ba9d760519a0a9

Observation 5b4a0c47-36b8-4475-a554-5451989a4704 · inbound

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards cites this paper.

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:03:16.378325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:03:16.378325Z digest=sha256:80a9dd390882db7265cdc754dc5681fd44f826ca80300735ecbfd56d3d52f069

Observation bac584d5-64b5-4b8d-85e2-2279e245f1fa · inbound

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models cites this paper.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.226372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:fba5d2ad51668ca9d3042e846b0bf0da82b12047f409d9190b4cea769da2fd12

Observation 395c7c37-590c-48b4-bf69-aaf3dd4bf8a8 · inbound

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning cites this paper.

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:27:17.800649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T00:26:58.273861Z digest=sha256:d5bc92b57072cd1af86c2e4cac8a96f809cb2bd4ea4eb3ee888ea28da9ad9614

Observation ca246feb-1907-4498-abb9-91b4e55c323f · inbound

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models cites this paper.

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:15:05.982335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:15:05.982335Z digest=sha256:8cefd49d997c0ecff5e67baeb9b72c308b9552e18d425f3de5cc720db8318632

Observation 753df86c-10a5-4030-87b2-d0e9b3d2b2d4 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:05:00.962878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:994d03ed5d96063f63a50d4dfc032e435673a8d81deb3b66dd0247a06d250536

Observation 0144ab4d-97db-4be1-ab66-10c4e27daf41 · inbound

Visual Test-time Scaling for GUI Agent Grounding cites this paper.

Visual Test-time Scaling for GUI Agent Grounding PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:37.602363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:37.602363Z digest=sha256:c16c947421942199a0d2ba421c888a259b7226097daf3fedbe6c88be09d46d34

Observation 5d6588cb-0c37-4de2-ab3c-f02932a10a9e · inbound

CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation cites this paper.

CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T01:07:35.318012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:07:35.318012Z digest=sha256:44ddd112ae9245f0b99b478c8487bf50b2d04c8410ea397d94b7fa6bde1c32b2

Observation 3560dda1-e017-497b-97c2-a83a1851eb77 · inbound

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects cites this paper.

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:10:22.999721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:10:22.999721Z digest=sha256:84f676182e521229f9b9882e04277d19203d50ef94666bd3de6a17c06d5da5f2

Observation 1f9fd103-4eba-4333-8dca-36bfa5ca5c4f · inbound

Efficient Sensorimotor Learning for Open-world Robot Manipulation cites this paper.

Efficient Sensorimotor Learning for Open-world Robot Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:51.275182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:27:51.275182Z digest=sha256:8e29ba7cc3f12fcd5d30348d86b36b2b2415beeb5faed184822907f6ec633713

Observation bf1a7f22-853c-491a-8107-e48a09ac636b · inbound

Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare cites this paper.

Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:08:54.129790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:08:54.129790Z digest=sha256:939bf85ed7ed3628143df2695eaa2a2f0e023bbed4099d4c576ea1b3f27bc421

Observation ec0e5f45-3049-49ce-b7bd-a866edb56ba4 · inbound

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing cites this paper.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.090778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.090778Z digest=sha256:fc64f5d9d44d73d3d88d6140a120b3bc867c02b801952699d37c10d0c0d9fe7d

Observation 763f7a6a-ea56-4eae-8381-24fa40878154 · inbound

Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning cites this paper.

Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:51.502500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:51.502500Z digest=sha256:2350606685c1611859c524ac716f9c55a2b71f06fa5d836846ae7280975e4282

Observation af1f007e-283e-4536-b106-f94e0420e065 · inbound

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models cites this paper.

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:40.833021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:40.833021Z digest=sha256:da736d86eb28b5d027863d4911474c17b3cfd3f1366e528a044390eab62f7d34

Observation 19d7957c-ead7-4c6a-af9c-1d95c5e5b4f9 · inbound

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks cites this paper.

LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:52.663287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:52.663287Z digest=sha256:c17c31cf11529b4e0df0e8f72767bfc899574a4d123e73e7aff626427337cdbb

Observation 2016e4c1-d9bc-4f68-bb0d-2fade28129d1 · inbound

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis cites this paper.

OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:48:02.509867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:48:02.509867Z digest=sha256:ed9df73c1b970112f767c5891e6093c9a5153fcb37e5ba529187ae4474331c59

Observation e5c1427f-7675-42e3-8e7c-fe077f9071ec · inbound

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation cites this paper.

UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:53.856037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:53.856037Z digest=sha256:05259e79fb38165cabd1d9b4a73d381f35cbe7da9895d386c6f670cf9ed0aac2

Observation f898e174-1805-4323-9a7f-b500b241d761 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:26.306305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:26.306305Z digest=sha256:51b2eb7ccd7d052c4012a44716e14f6c3e2cd9de9471c9683f3d77314a519b6f

Observation 6d45fc77-94b5-4864-a124-9f09338f21e3 · inbound

Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins cites this paper.

Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:00.413992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:00.413992Z digest=sha256:cc24290d0ae73e03d1095683d57557abd329afb76f5ac93cfbf88e21e7f84006

Observation facb93a0-e113-4668-a409-e49abe72a247 · inbound

Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models cites this paper.

Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:15.277552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:15.277552Z digest=sha256:664f27d12c24eeb1914449f4a3520ce113b3d828224cf8554fe15083f01dbfe0

Observation 7b6a019f-9107-4a46-8b67-74a7316ad964 · inbound

DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory cites this paper.

DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:11.811469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:48:11.811469Z digest=sha256:c24f16e0963ccfa3de0ec614a12a6a2b3caf91a6f4c83f8d332eb68a70513c71

Observation 17ac6a62-d29c-4476-afc3-184a100b89bc · inbound

Grounding Language Models with Semantic Digital Twins for Robotic Planning cites this paper.

Grounding Language Models with Semantic Digital Twins for Robotic Planning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:28:36.425673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:28:36.425673Z digest=sha256:c2f0e69d0e0d64ae0e71abbc2f456c05aaa7af3c630e1920d1c3b8d10449371c

Observation 1fbc5949-a91c-4f0a-a408-94bd1d2310af · inbound

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity cites this paper.

CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T19:31:53.455936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:31:53.455936Z digest=sha256:2082ecc5435475b21cfdc64ab9d262d2894860c58b99b6c8b5494e32890ffeab

Observation dba92367-874c-441b-b2c1-f972801ef107 · inbound

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models cites this paper.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.298511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.298511Z digest=sha256:314deeb6b644186208cc383ae80165427ef2a31fa948e6fe064f31f1bb2954de

Observation 8da321ad-6616-4dbe-a41e-20a59c43a8cc · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:07.661869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:07.661869Z digest=sha256:15a8daded9bdec2d7d2d9d5b3529f8f4788edcd5ecf3dd2913e6aff81a5bd18f

Observation ed9989cc-d9e3-4f85-a6fd-7b067971dd23 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.709761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.709761Z digest=sha256:245c86bf6efd02eafa90724030cf6161a1aef1b285d6319f40dcc1cb92c04cbe

Observation 0b29dbe9-3ca1-4c05-96b9-a4b48a22266c · inbound

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding cites this paper.

MOSU: Autonomous Long-range Robot Navigation with Multi-modal Scene Understanding PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:16.210236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:47:16.210236Z digest=sha256:7c500805aa60b192e8114c9326705eb21a63d644595b665866cf96f5e32e3c7e

Observation a08e6030-d0b3-4472-95c8-74a160fd17ec · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.282950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.282950Z digest=sha256:d356642381977f7a739d92d5437aa529f5b2ad7de95a9826bb5220729342c132

Observation 63494fdd-680f-4243-a7d3-231435e93d5c · inbound

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping cites this paper.

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:25.133699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:25.133699Z digest=sha256:1fa382e847421a888ae0f946cd81825a5d69a8b54336b09cd677a2f4aacc7b65

Observation 844b8e2d-5cc2-4da6-a348-74db30750c5b · inbound

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies cites this paper.

AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T21:44:58.996315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:44:58.996315Z digest=sha256:00ca5e976d4eb22e02f6cbe024e13ae968659cb61e9ae4f3021c8ee84458fae9

Observation 74a96285-36ca-470f-9429-f7ac7f3740d6 · inbound

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT cites this paper.

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:23:18.991949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:23:18.991949Z digest=sha256:3a58c9bcc70f32fe59dc97eadf59ed74913633ae37f826ea4a72661aa672bced

Observation db4bb1ff-e23a-4481-8c5a-cb60b7602da8 · inbound

EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos cites this paper.

EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:46.057999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:46.057999Z digest=sha256:faf44885b4c6526323854afe4db9daae87e196c4da095c334b1335d8f2ecc9eb

Observation 2494ea49-5bac-4ac1-8540-65685f94adef · inbound

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models cites this paper.

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T19:06:28.691845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:06:28.691845Z digest=sha256:6a24f41c4d1d14f53b4df9c85d27dc89e5a1fa80b7e9a4fe9ddc9f5932834996

Observation 3052474a-8db0-4b40-ab2a-d602a08bbd40 · inbound

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals cites this paper.

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T20:21:06.276723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:21:06.276723Z digest=sha256:126fda475c9a392269b75f9d303520168fe10b14e458d4cc76743ed212163d02

Observation 4d0442ee-7f3e-4452-8e7e-353bd32c35f6 · inbound

SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation cites this paper.

SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:09:09.473481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:09:09.473481Z digest=sha256:e08c54e740f76334ee9e0c65d94c0798cbf8bc57c69b52856c61ecea26a5bf18

Observation 2ace252f-c4e0-4843-b1d3-24b59dcc3d43 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:52.414897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:52.414897Z digest=sha256:db312d79ff815ba748fc959d8c57588ab68ffcd2a509bd669673f9f07062dc29

Observation dd755236-743f-4464-9e78-88eb3b1ba4a9 · inbound

EVE: A Generator-Verifier System for Generative Policies cites this paper.

EVE: A Generator-Verifier System for Generative Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T14:10:36.225511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:10:36.225511Z digest=sha256:1f44f34713ae971176fe50df599d69636518d86c6f8425cd8a2dd00f6961d54e

Observation 6693fff6-e3fc-4240-9e81-4f5b1e23c063 · inbound

Visual-Language-Guided Task Planning for Horticultural Robots cites this paper.

Visual-Language-Guided Task Planning for Horticultural Robots PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:58:53.900823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:58:53.900823Z digest=sha256:16ae8840097a05191d13c2d2427933423fd66ede6aa640a26a3b5d0e66da7d9e

Observation 3047f114-df0f-444d-9136-f0d4615c164b · inbound

Sem-NaVAE: Semantically-Guided Outdoor Mapless Navigation via Generative Trajectory Priors cites this paper.

Sem-NaVAE: Semantically-Guided Outdoor Mapless Navigation via Generative Trajectory Priors PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T05:43:21.189276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:43:21.189276Z digest=sha256:7c26c317500576d05964d885bc7c17dd774f2fed61a2d5d665a05cd2aad1accd

Observation b07d05f7-0c8d-40d3-97d8-29f0648ac513 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:6f96c7494dc6175769f5f059521bbd2f78276cd7966c746bad5f0098fce11a53

Observation 33c2c122-6f5b-4a7f-815d-bf60d59310c0 · inbound

JailWAM: Jailbreaking World Action Models in Robot Control cites this paper.

JailWAM: Jailbreaking World Action Models in Robot Control PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:48.143404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:22:13.890495Z digest=sha256:5e58c1dc88d5dd408b3215e74570957ae25e55189d0ab0e1b621b5b66dd56f43

Observation a876af4c-f548-405c-9b3d-d3731957eca7 · inbound

Visually-grounded Humanoid Agents cites this paper.

Visually-grounded Humanoid Agents PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:04.606662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:10:32.710853Z digest=sha256:fb13400d011abf130e58f1127e7243da4688c557afed36a95a72b7563fb3aa73

Observation 4942f695-e5c0-4837-bbda-d22b9cf0164e · inbound

Improving Robotic Generalist Policies via Flow Reversal Steering cites this paper.

Improving Robotic Generalist Policies via Flow Reversal Steering PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:48:35.824129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:20:19.209180Z digest=sha256:fe9c593694cb3214f3e98983b37fe58821972b3cacef01ef1c7c15b52db55238

Observation 6b6447d3-9169-4cee-89e4-3107a2ef1ef0 · inbound

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning cites this paper.

ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:20.256553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T20:49:01.269513Z digest=sha256:bdc903057f4f0c5bdb90e6de259aec85067ded0896cf70910d71d2675d9a2f5d

Observation 656704d4-258c-4f0b-99f6-d44d53412430 · inbound

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning cites this paper.

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.795525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T20:48:49.071281Z digest=sha256:507d04a3933af64bf5f09232a8fc2714b0eb0dc5613d7e7037a3afde97ba0416

Observation bec7434e-6f77-4818-ba85-a7f72b2950db · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:56.957400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T01:25:21.796778Z digest=sha256:f68fe1e7060628fa017306c4f69cb8575181555124d247c1ad0c329e50707880

Observation c2a296ef-5bae-432c-a4be-7fbb039a02d8 · inbound

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation cites this paper.

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.671107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T00:54:50.393828Z digest=sha256:bcf189902176eb9a6d636f86ebf3f710b71d2b7550f3ad34aa44735d94ed8ec0

Observation b385d68e-4993-401d-b552-27f589b8085f · inbound

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models cites this paper.

TAP-VLA: Tactile Annotation Prompting for Vision Language Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.638351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T09:10:55.016700Z digest=sha256:55808f88e57f3101fabd552e1c6c46bdd9ef80daecd7a3e176931df632f2e385

Observation de5719f8-13bf-4504-889a-2a8122df22d1 · inbound

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance cites this paper.

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:41.952355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:17:25.534715Z digest=sha256:ecd1cf6c4f10780f85c9da596d3a149d8fbacfeed1c802446d32a5044796a3f2

Observation 5f1c91ea-0129-49e2-9af2-8b921cd5b863 · inbound

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance cites this paper.

3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.682933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T19:15:49.118807Z digest=sha256:1cc1934d0ac2bbca5a6698951c4b2c81e6eefa2ba7fdd6a0ff486aaeeb7a5752

Observation cd4561a9-4371-447c-8fda-5928d69a8e9c · inbound

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies cites this paper.

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T08:29:01.477958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:29:01.477958Z digest=sha256:8ca6f32315bfbbff99b98aba4f158ae1656ce545c9766210720aa221a4a34826

Observation b67453e3-438b-4aeb-9b8c-403755cc4fee · inbound

IMBench: A Benchmark for Intuitive Robotic Manipulation cites this paper.

IMBench: A Benchmark for Intuitive Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:13.878543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:13.878543Z digest=sha256:ad6ab59e1a0c5aa7b4a9b70ca4f1e760679d67662fb052f3233759c81f628a79

Observation b2cec11d-36f2-49d3-9adf-56765bcaf9fe · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:50.899469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:50.899469Z digest=sha256:3ce706d2eced7c12886c3bb5b4cf4e6768c48b29515e8e5c6991b82cdbccc3ae

Observation 9bb385f8-e7c9-409c-9286-d8a5d9ebee03 · inbound

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models cites this paper.

World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T04:59:36.361796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:59:36.361796Z digest=sha256:3fb5157ebac17b7422e68e8e6e7a818d03a44620476b83b3eac70089219b9399

Observation f861b24e-85a7-4831-8b27-43d507a2decd · inbound

Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions cites this paper.

Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:16:52.244668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:16:52.244668Z digest=sha256:ddeb8c904af72aab8583c9f054b5bdc5e8e5b014464fe3fbd3fbf4b05ea09744

Observation ce1420cd-c1cb-421e-9729-6b2f8b0ce335 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.094779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.094779Z digest=sha256:edb69ff9a0e7f8469f034577461cfd9ad233837a3b1fd6be437fb2dabe4a8971