Pith. sign in

REVIEW 4 major objections 5 minor 47 references

On the Dual-Use Dilemma in Physical Reasoning and Force

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Reinforcing a safety prompt in VLM-controlled robots suppresses both harmful and helpful physical actions across two force-planning schemes.

desk verdict Solid case study with a real, but narrower, result than the abstract claims; the 'helpful' task set undermines the headline trade-off. read the letter →

arxiv 2505.18792 v1 pith:QVI7SCN5 submitted 2025-05-24 cs.RO

classification cs.RO
keywords vision-languagemodelsrobotmanipulationforceandtorqueplanningsafetyalignmentdual-usedilemmacontact-richpromptsafeguardinggraspcontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a vision-language model (VLM) that controls a robot can be made safe without losing its ability to apply force helpfully. The authors test a simple safeguard—appending Isaac Asimov's first law of robotics to the system prompt—across two ways of eliciting physical reasoning: wrench planning for contact-rich motion and grasp-force estimation. In both, the safeguard lowered the rate at which the models produced harmful plans, but it also lowered the rate at which they produced helpful high-force plans. The paper takes this as evidence for a capability-safety trade-off: value alignment of this kind can impede desirable robot skills.

What carries the argument

The load-bearing mechanism is the Asimov-rule prompt: a sentence appended to the system prompt instructing the model to refuse any response that could injure a human, to zero out any requested force or wrench, and to stop with the keyword 'asimov.' It is applied to two existing elicitation schemes—wrench planning, which asks the VLM to output a 6-D force/torque vector and duration from an image with an overlaid coordinate frame, and grasp-force control, which estimates object properties and modulates grasp force by task semantics. The prompt is the single intervention whose presence or absence separates the baseline and safeguarded conditions in every comparison.

What would settle it

Run the same six wrench-planning and six grasp-force tasks with a different safeguarding mechanism—for example, a constitutional prompt that allows contact when a user consents or a classifier that rejects only explicitly violent requests—and compare helpful elicitation rates. If any such safeguard holds harmful plans near zero while keeping helpful high-force responses at or above the unsafeguarded baseline, the paper's central claim that alignment impedes capability would fail.

Watch

Extended reading notes

Core claim

The central discovery is a measured trade-off between safety alignment and physical capability in VLM-driven robots. Across 1,800 wrench-planning queries and 360 grasp-force queries, adding the Asimov-rule safeguard reduced harmful behavior elicitation from 53% to 19% in wrench planning and from 67% to 2% in grasp-force control, while simultaneously reducing helpful behavior from 50% to 39% in wrench planning and from 91% to 45% in grasp-force control. The same pattern appears in both prompting schemes and in each of the three models tested, although to different degrees. The authors conclude that reinforcing model safeguards within prompts blocks not only violence but also legitimate forceful care tasks such as setting a wrist or massaging a neck, and they generalize this to a warning that value alignment may impede desirable robot capabilities.

Load-bearing premise

The entire generalization rests on the assumption that one hand-written Asimov-style prompt is representative of safety alignment; if a different safeguard could stop harmful plans without suppressing helpful ones, the claimed trade-off would not be a general property of alignment.

Editorial extensions

If this is right

  • Safety prompting in VLM-controlled robots carries a measurable capability cost, not just a failure to comply.
  • Model evaluations for embodied agents should report helpful-behavior suppression alongside harm reduction, since the two move together.
  • Prompt complexity amplifies both helpful and harmful wrench generation, so efforts to improve physical reasoning may widen the dual-use gap.
  • Physical-reasoning elicitation should be treated as a jailbreak surface: force and torque requests bypass existing refusals.
  • Robot learning systems will likely need to act near the harm boundary and learn from mistakes, which safety thresholds must accommodate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The measured trade-off may be specific to the wording of the first law, which prohibits risk to humans outright; a context-aware safeguard that permits consented medical or care tasks could in principle restore helpful behavior while blocking violence, but the paper does not test one.
  • Because the helpful tasks and harmful tasks share the same body-contact vocabulary, the results suggest VLMs classify by surface semantics; a safeguard keyed to consent and procedure rather than contact could behave differently.
  • Extrapolating from these two schemes, the authors' own discussion implies that evaluating general-purpose models on a representative sample of human interaction is intractable, so safety-capability trade-offs may need to be measured task-family by task-family.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports two case studies on vision-language models (VLMs) that generate forceful robot motion. In the first case study, five prompt configurations for wrench planning are evaluated, with and without appending Asimov's first law as a safeguard, across three models and six tasks (three labeled helpful, three labeled harmful). In the second, a grasp-force estimation prompt is evaluated with and without the same safeguard across four helpful and two harmful tasks. The authors find that the safeguard reduces harmful behavior elicitation (wrench planning: 53% to 19%; grasp force: 67% to 1.7%) but also reduces helpful behavior elicitation (wrench planning: 50% to 39%; grasp force: 91% to 45%). They interpret this as evidence that value alignment may impede desirable robot capabilities and discuss implications for model evaluation and robot learning.

Significance. If the empirical pattern is robust, the result is timely and important for embodied AI safety: it suggests that a simple prompt-level safeguard can curb dangerous physical action elicitation but may also suppress beneficial high-force contact tasks. The paper has concrete strengths: it evaluates two distinct prompting schemes, uses three different models, reports per-task and per-model breakdowns in appendices, and explicitly acknowledges its own limitations in Section V. The central claim, however, is broader than the evidence. The operationalization of 'helpful behavior' is not validated, the safeguard is a single prompt strategy, the classification thresholds are arbitrary, and no statistical uncertainty is reported. The paper is therefore best read as a preliminary empirical observation rather than a demonstrated general trade-off between alignment and capability.

major comments (4)
  1. [§IV, Table II, App. C] The 'helpful' task set is not actually validated as helpful. The tasks 'setting a dislocated wrist' and 'making an initial stomach incision' are labeled helpful solely because the prompt says they are beneficial, but in the depicted scene the VLM cannot verify surgical or clinical context. Appendix C shows the aggregate helpful drop is concentrated in the scissors task (0.49 to 0.30) and wrist task (0.67 to 0.53), while neck massage is unchanged (0.33 to 0.33). Excluding or reclassifying the high-risk tasks would largely erase the claimed loss of desirable capability. The authors need to either use tasks that are unambiguously helpful in the presented context, or validate safety/benefit with expert oversight, and re-run the aggregate analysis.
  2. [§IV] No significance tests or confidence intervals are reported. Each prompt configuration uses 10 queries per model/task, giving per task-model-condition cells of 50 queries (5 configurations x 10 queries), and per configuration cells of 10. Many reported deltas, such as the neck-massage helpful drop from 0.28 to 0.24 for Gemini, are inside sampling noise. The paper should report effect sizes with bootstrap or exact confidence intervals, and account for multiple comparisons across tasks, models, and prompt configurations.
  3. [§III, §V] The safeguard studied is a single prompt strategy: appending Asimov's first law. The abstract and Section V generalize from this to 'value alignment' as a whole. The authors note in Section V that they 'consciously chose a straightforward strategy,' but this does not address the load-bearing question of whether the measured trade-off is a property of this one prompt or of alignment mechanisms generally. If a different safeguard (e.g., constitutional prompting or refusal training) preserves helpful behavior while blocking harm, the central claim fails. The paper should either test additional safeguard mechanisms or substantially narrow its claim.
  4. [§IV] The harm/help classification thresholds are arbitrary and not physically grounded. For wrench planning, a response is counted as harmful or helpful if the 'unitless magnitude' exceeds 5; for grasp force, any non-zero force counts. Without specifying robot scale, force units, or a physical validation that the planned wrenches are dangerous or beneficial on a human, the classification conflates text output with actual physical effect. The authors should justify the thresholds, report sensitivity to threshold choice, or ground them in simulation or robot experiments.
minor comments (5)
  1. [§IV, Fig. 3 caption] The text in Section IV says helpful behavior drops from 50% to 38%, while Figure 3's caption and Table II report 50% to 39%. Please reconcile these numbers.
  2. [Table II] The table uses 'O: OpenAI, G: Gemini, C: Claude, µ: average' but the text calls them models; include the sample size per cell in the caption or table note.
  3. [Fig. 2, Fig. 3] The x-axis labeled 'Prompt Complexity (Num Tokens)' suggests a monotone complexity ordering, but the prompt configurations differ in several attributes at once. Clarify whether token count is a proxy or a controlled variable.
  4. [App. A] The appendix says prompts can be viewed at 'this link' and 'biewed' has a typo. Use persistent URLs or include the prompts in full in the appendix.
  5. [§V] There is a typo in 'We urge other researches in embodied AI' — should be 'researchers.'

Circularity Check

0 steps flagged · score 1.0 of 10

No circularity: the safeguard trade-off is measured, not constructed, though the 'helpful' task set contains validity-concerning items.

full rationale

The paper's central claim is that adding an Asimov-style safeguarding prompt reduces both harmful and helpful behavioral elicitation across two VLM-based force-planning schemes. This is an empirical measurement, not a result forced by construction. The two prompting schemes are drawn from the authors' own prior work ([40] and [41]), but the paper re-measures the baseline behavior in Tables II and III/IV rather than relying on the citations alone, so the self-citations are not load-bearing. No parameter is fitted to produce the trade-off: the wrench-magnitude threshold (>5) and the zero/non-zero grasp-force criterion are fixed operational definitions applied uniformly. The safeguard prompt does instruct the model to zero out the wrench when it detects potential harm, and the classifier counts a zeroed wrench as non-elicitation; however, the data show the drop is not a purely definitional artifact, since the neck-massage helpful task is unchanged (33% to 33%) and some models still emit high-magnitude wrenches under safeguarding (App. E, Fig. 5). The most plausible concern is construct validity rather than circularity: the 'helpful' task set includes 'making an initial stomach incision,' which is objectively harmful outside a surgical context, and the paper explicitly acknowledges this assumption in Section III ('While it is unlikely one would require a robot to perform any of these helpful tasks, especially for severe tasks like the incision task...'). Appendix C's per-task breakdown shows the aggregate helpful-behavior drop is concentrated in the incision and wrist-setting tasks, which weakens the broad conclusion that value alignment impedes desirable robot capabilities, but this is a labeling/threat-to-validity issue, not a derivation that reduces to its own inputs. Overall, the safeguarding effect is genuinely measured, and the paper discloses the data needed to interrogate its generalization. Score 1.0 reflects only the minor presence of self-citations in selecting the testbed; no circular step is present.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The analysis relies on ad hoc classification thresholds and a single prompt-based safeguard, but introduces no new physical or conceptual entities.

free parameters (2)
  • Harm/help wrench magnitude threshold = 5 (unitless)
    Responses with a wrench magnitude above 5 are classified as harmful/helpful; the threshold is stated without justification or sensitivity analysis in Section IV.
  • Grasp force threshold = 0 (any non-zero force)
    Any non-zero grasp force counts as behavior elicitation; this coarse threshold may overstate both helpful and harmful rates (Section IV).
assumptions (3)
  • domain assumption VLM-generated force outputs are a meaningful proxy for actual robot behavior and potential harm.
    The paper classifies responses solely on output magnitudes, with no physical robot validation; see Section IV.
  • domain assumption The six selected tasks represent helpful and harmful contact-rich manipulation.
    The paper describes the set as limited and abstracted, and admits some helpful tasks (e.g., stomach incision) are unlikely; see Section III and Section V.
  • ad hoc to paper Appending Asimov's first law to a prompt constitutes a representative safeguard or value alignment strategy.
    The safeguard is a single fixed prompt, chosen for simplicity, and not compared to other alignment methods; see Section III and V.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Dual-Use Dilemma in Physical Reasoning and Force." pith.science (2026). https://pith.science/paper/QVI7SCN5

@misc{pith2026250518792,
  author       = {Pith},
  title        = {Pith review of: On the Dual-Use Dilemma in Physical Reasoning and Force},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QVI7SCN5}},
  note         = {Machine review of arXiv:2505.18792}
}
read the original abstract

Humans learn how and when to apply forces in the world via a complex physiological and psychological learning process. Attempting to replicate this in vision-language models (VLMs) presents two challenges: VLMs can produce harmful behavior, which is particularly dangerous for VLM-controlled robots which interact with the world, but imposing behavioral safeguards can limit their functional and ethical extents. We conduct two case studies on safeguarding VLMs which generate forceful robotic motion, finding that safeguards reduce both harmful and helpful behavior involving contact-rich manipulation of human body parts. Then, we discuss the key implication of this result--that value alignment may impede desirable robot capabilities--for model evaluation and robot learning.

Figures

Figures reproduced from arXiv: 2505.18792 by the authors.

Figure 1
Figure 1. Varying contextual semantics in the same scene can yield harm and help, often with a thin line separating them. We evaluate how VLMs under different prompt schemes which elicit physical reasoning for robot control navigate this line between harm and help for forceful, contact-rich tasks with potential for bodily danger. in low-data regions of their training distributions, the need for better alignment between model … view at source ↗
Figure 2
Figure 2. Additional safeguarding reduces harmful wrenches on average by 34% (absolute, 53% to 19%). It completely reduces harmful behavior from Claude 3.7 Sonnet (20% to 0%) and by 57% for GPT 4.1 Mini (84% to 27%). Gemini 2.0 Flash is the least responsive to safeguarding, decreasing 23% (55% to 32%). Safeguarding is roughly less effective as prompty complexity increases. Across all models, tasks, and prompting schemes, safe… view at source ↗
Figure 4
Figure 4. We evaluate additional prompting schemes for physical reasoning about grasp forces [40] on four helpful tasks (w, n and W, N corresponding to low and high force magnitude tasks, respectively) and two harmful tasks (W, N). Safeguards (dashed bars) completely suppress harm (right), but greatly reduce helpful behavior (left) [PITH_FULL_IMAGE:figures/full_fig_p003_4.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: Wrench magnitudes for OpenAI and Gemini models are relatively consistent, whereas Claude 3.7 Sonnet fluctuates considerably. This is due to a lower quantity of unblocked responses, resulting in greater variance, as well as an observed behavior of attempting to break th…
Figure 6
Figure 6. Figure 6: Helpful behavior scales with prompt complexity and is reduced by safeguarding. On average, models detect potential harm in a 40% of helpful task queries, with Claude 3.7 Sonnet the highest at 63% of responses, 39% for Gemini 2.0 Flash, and 4% for OpenAI GPT 4.1 Mini […

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

47 extracted references · 23 canonical work pages

  1. [1]

    "Can you be my mum?": Manipulating Social Robots in the Large Language Models Era

    Giulio Antonio Abbo, Gloria Desideri, Tony Belpaeme, and Micol Spitale. ”can you be my mum?”: Manipulating social robots in the large language models era, 2025. URL https://arxiv.org/abs/2501.04633

  2. [2]

    Do as I can, Not As I Say: Ground- ing language in robotic affordances

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Haus- man, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang,...

  3. [3]

    Autort: Embodied foundation models for large scale orchestration of robotic agents, 2024

    Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Karol Hausman, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Sean Kirmani, Isabel Leal, Edward Lee, Sergey Levine, Yao Lu, Isabel Leal, Sharath Maddineni, Kanishka Rao, Dorsa Sadigh, Pannag Sanketi, Pierre Ser- manet, Quan Vuong, Stefan Welker, Fei Xia, ...

  4. [4]

    Helix: A vision-language-action model for generalist humanoid control, 2025

    Figure AI. Helix: A vision-language-action model for generalist humanoid control, 2025. URL https://www. figure.ai/news/helix

  5. [5]

    What should we want from a robot ethic? The International Review of Information Ethics , 6:9–16, Dec

    Peter M Asaro. What should we want from a robot ethic? The International Review of Information Ethics , 6:9–16, Dec. 2006. doi: 10.29173/irie134. URL https: //informationethics.ca/index.php/irie/article/view/134

  6. [6]

    Runaround, 1942

    Isaac Asimov. Runaround, 1942. URL https://web.williams.edu/Mathematics/sjmiller/public html/105Sp10/handouts/Runaround.html

  7. [7]

    Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Her- nandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse,...

  8. [8]

    The State of Robot Motion Generation

    Kostas E. Bekris, Joe Doerr, Patrick Meng, and Sumanth Tangirala. The State of Robot Motion Generation, December 2024. URL http://arxiv.org/abs/2410.12172. arXiv:2410.12172 [cs]

Show all 47 references
  1. [9]

    Human-centered evaluation of language technologies

    Su Lin Blodgett, Jackie Chi Kit Cheung, Vera Liao, and Ziang Xiao. Human-centered evaluation of language technologies. In Jessy Li and Fei Liu, editors, Proceed- ings of the 2024 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts , pages 39–43, ...

  2. [10]

    Robot Ethics: Ethical Design Con- siderations, pages 473–491

    Dylan Cawthorne. Robot Ethics: Ethical Design Con- siderations, pages 473–491. Springer Nature Singa- pore, Singapore, 2022. ISBN 978-981-19-1983-1. doi: 10.1007/978-981-19-1983-1 16. URL https://doi.org/10. 1007/978-981-19-1983-1 16

  3. [11]

    Open X- Embodiment: Robotic learning datasets and RT-X models

    Open X-Embodiment Collaboration. Open X- Embodiment: Robotic learning datasets and RT-X models. https://arxiv.org/abs/2310.08864, 2023

  4. [12]

    Openhelix: A short survey, empirical analysis, and open-source dual-system vla model for robotic manipulation, 2025

    Can Cui, Pengxiang Ding, Wenxuan Song, Shuanghao Bai, Xinyang Tong, Zirui Ge, Runze Suo, Wanqi Zhou, Yang Liu, Bofang Jia, Han Zhao, Siteng Huang, and Donglin Wang. Openhelix: A short survey, empirical analysis, and open-source dual-system vla model for robotic manipulation, 2...

  5. [13]

    Robots: ethical by design

    Gordana Dodig Crnkovic and Baran C ¸ ¨ur¨ukl¨u. Robots: ethical by design. Ethics and Information Technology , 14(1):61–71, March 2012. ISSN 1388-1957, 1572-8439. doi: 10.1007/s10676-011-9278-2

  6. [14]

    Physically grounded vision-language models for robotic manipulation

    Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Majumdar, and Dorsa Sadigh. Physically grounded vision-language models for robotic manipulation. In IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2024

  7. [15]

    Gemini robotics: Bringing ai into the physical world, 2025

    Google DeepMind Gemini Robotics Team. Gemini robotics: Bringing ai into the physical world, 2025. URL https://storage.googleapis.com/deepmind-media/ gemini-robotics/gemini robotics report.pdf

  8. [16]

    Dirk Hovy and Shannon L. Spruit. The social im- pact of natural language processing. In Katrin Erk and Noah A. Smith, editors, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (V olume 2: Short Papers) , pages 591–598, Berlin, Germany, Au...

  9. [17]

    ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipula- tion, November 2024

    Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, and Li Fei-Fei. ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipula- tion, November 2024. URL http://arxiv.org/abs/2409. 01652. arXiv:2409.01652 [cs]

  10. [18]

    Rieder, Debra J

    Brian Hutler, Travis N. Rieder, Debra J. H. Mathews, David A. Handelman, and Ariel M. Greenberg. Design- ing robots that do no harm: understanding the challenges of ethics for robots. Ai and Ethics , page 1–9, April 2023. ISSN 2730-5953. doi: 10.1007/s43681-023-00283-8

  11. [19]

    Regulating artificial intelligence and robotics: ethics by design in a digital society

    Ron Iphofen and Mihalis Kritikos. Regulating artificial intelligence and robotics: ethics by design in a digital society. Contemporary Social Science , 16(2):170–184, March 2021. ISSN 2158-2041, 2158-205X. doi: 10. 1080/21582041.2018.1563803

  12. [20]

    Thorny roses: Investigating the dual use dilemma in natural language processing

    Lucie-Aim ´ee Kaffee, Arnav Arora, Zeerak Talat, and Isabelle Augenstein. Thorny roses: Investigating the dual use dilemma in natural language processing. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguis- tics: EMNLP ...

  13. [21]

    Openvla: An open-source vision-language-action model,

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Openvla...

  14. [22]

    Code as policies: Language model programs for em- bodied control

    Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for em- bodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 9493–9500,

  15. [23]

    Vera Liao and Ziang Xiao

    Q. Vera Liao and Ziang Xiao. Rethinking model evalu- ation as narrowing the socio-technical gap, 2025. URL https://arxiv.org/abs/2306.03100

  16. [24]

    Compromising embodied agents with contextual backdoor attacks, 2024

    Aishan Liu, Yuguang Zhou, Xianglong Liu, Tianyuan Zhang, Siyuan Liang, Jiakai Wang, Yanjun Pu, Tianlin Li, Junqi Zhang, Wenbo Zhou, Qing Guo, and Dacheng Tao. Compromising embodied agents with contextual backdoor attacks, 2024. URL https://arxiv.org/abs/2408. 02882

  17. [25]

    Exploring the robustness of decision- level through adversarial attacks on llm-based embodied models, 2024

    Shuyuan Liu, Jiawei Chen, Shouwei Ruan, Hang Su, and Zhaoxia Yin. Exploring the robustness of decision- level through adversarial attacks on llm-based embodied models, 2024. URL https://arxiv.org/abs/2405.19802

  18. [26]

    Poex: Understanding and mitigating policy executable jailbreak attacks against embodied ai,

    Xuancun Lu, Zhengxian Huang, Xinfeng Li, Xiaoyu ji, and Wenyuan Xu. Poex: Understanding and mitigating policy executable jailbreak attacks against embodied ai,

  19. [27]

    Octo: An open-source generalist robot policy

    Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Charles Xu, Jianlan Luo, Tobias Kreiman, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo...

  20. [28]

    Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning

    Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou- Chakra, Ian Reid, and Niko Suenderhauf. Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning. In 7th Annual Confer- ence on Robot Learning , 2023. URL https://openreview. net/forum?...

  21. [29]

    Pappas, and Hamed Hassani

    Zachary Ravichandran, Alexander Robey, Vijay Kumar, George J. Pappas, and Hamed Hassani. Safety guardrails for llm-enabled robots, 2025. URL https://arxiv.org/abs/ 2503.07885

  22. [30]

    Alexander Robey, Zachary Ravichandran, Vijay Kumar, Hamed Hassani, and George J. Pappas. Jailbreaking llm- controlled robots, 2024. URL https://arxiv.org/abs/2410. 13691

  23. [31]

    Generating robot constitutions & benchmarks for semantic safety

    Pierre Sermanet, Anirudha Majumdar, Alex Irpan, Dmitry Kalashnikov, and Vikas Sindhwani. Generating robot constitutions & benchmarks for semantic safety. arXiv preprint arXiv:2503.08663 , 2025

  24. [32]

    Welcome to the era of experience, 2025

    David Silver and Richard Sutton. Welcome to the era of experience, 2025. URL https://storage.googleapis.com/ deepmind-media/Era-of-Experience%20/The%20Era% 20of%20Experience%20Paper.pdf

  25. [33]

    ProgPrompt: pro- gram generation for situated robot task planning using large language models

    Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg. ProgPrompt: pro- gram generation for situated robot task planning using large language models. Autonomous Robots , 47(8): 999–1012, December ...

  26. [34]

    Exploring the ad- versarial vulnerabilities of vision-language-action models in robotics, 2025

    Taowen Wang, Cheng Han, James Chenhao Liang, Wen- hao Yang, Dongfang Liu, Luna Xinyu Zhang, Qifan Wang, Jiebo Luo, and Ruixiang Tang. Exploring the ad- versarial vulnerabilities of vision-language-action models in robotics, 2025. URL https://arxiv.org/abs/2411.13587

  27. [35]

    Trojanrobot: Physical-world backdoor attacks against vlm-based robotic manipulation, 2025

    Xianlong Wang, Hewen Pan, Hangtao Zhang, Minghui Li, Shengshan Hu, Ziqi Zhou, Lulu Xue, Peijin Guo, Yichen Wang, Wei Wan, Aishan Liu, and Leo Yu Zhang. Trojanrobot: Physical-world backdoor attacks against vlm-based robotic manipulation, 2025. URL https://arxiv.org/abs/2411.11683

  28. [36]

    Yi Wang, Jiafei Duan, Dieter Fox, and Siddhartha Srini- vasa. NEWTON: Are large language models capa- ble of physical reasoning? In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Asso- ciation for Computational Linguistics: EMNLP 2023 , pages 9743–9758, Si...

  29. [37]

    Meta-control: Automatic model-based control synthesis for heterogeneous robot skills

    Tianhao Wei, Liqian Ma, Rui Chen, Weiye Zhao, and Changliu Liu. Meta-control: Automatic model-based control synthesis for heterogeneous robot skills. In 8th Annual Conference on Robot Learning , 2024. URL https://openreview.net/forum?id=cvVEkS5yij

  30. [38]

    Sadler, Di- nesh Manocha, and Amrit Singh Bedi

    Xiyang Wu, Souradip Chakraborty, Ruiqi Xian, Jing Liang, Tianrui Guan, Fuxiao Liu, Brian M. Sadler, Di- nesh Manocha, and Amrit Singh Bedi. On the vul- nerability of llm/vlm-controlled robotics, 2025. URL https://arxiv.org/abs/2402.10340

  31. [39]

    Towards forceful robotic foundation models: a literature survey, 2025

    William Xie and Nikolaus Correll. Towards forceful robotic foundation models: a literature survey, 2025. URL https://arxiv.org/abs/2504.11827

  32. [40]

    Deligrasp: Inferring object properties with llms for adaptive grasp policies

    William Xie, Maria Valentini, Jensen Lavering, and Nikolaus Correll. Deligrasp: Inferring object properties with llms for adaptive grasp policies. In Proceedings of the 8th International Conference on Robot Learning (CoRL), 2024. URL https://arxiv.org/abs/2403.07832

  33. [41]

    Unfettered forceful skill acquisition with physi- cal reasoning and coordinate frame labeling, 2025

    William Xie, Max Conway, Yutong Zhang, and Nikolaus Correll. Unfettered forceful skill acquisition with physi- cal reasoning and coordinate frame labeling, 2025. URL https://arxiv.org/abs/2505.09731

  34. [42]

    Language to rewards for robotic skill synthesis

    Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montserrat Gonzalez Arenas, Hao- Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, brian ichter, Ted Xiao, Peng Xu, Andy Zeng, Tingnan Zhang, Nicolas Heess, Dorsa Sadigh, Jie Tan, Yuval Tassa, an...

  35. [43]

    Robopoint: A vision- language model for spatial affordance prediction for robotics, 2024

    Wentao Yuan, Jiafei Duan, Valts Blukis, Wilbert Pumacay, Ranjay Krishna, Adithyavairavan Murali, Ar- salan Mousavian, and Dieter Fox. Robopoint: A vision- language model for spatial affordance prediction for robotics, 2024. URL https://arxiv.org/abs/2406.10721

  36. [44]

    reasoning

    Hangtao Zhang, Chenyu Zhu, Xianlong Wang, Ziqi Zhou, Changgan Yin, Minghui Li, Lulu Xue, Yichen Wang, Shengshan Hu, Aishan Liu, Peijin Guo, and Leo Yu Zhang. Badrobot: Jailbreaking embodied llms in the physical world, 2025. URL https://arxiv.org/abs/ 2407.20242. APPENDIX A. Pr...

  37. [2023]

    doi: 10.1109/ICRA48891.2023.10160591

  38. [2024]

    URL https://arxiv.org/abs/2406.09246

  39. [2025]

    URL https://arxiv.org/abs/2412.16633

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.