Pith. sign in

REVIEW 15 cited by

Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.01928 v2 pith:WS5JKJ7H submitted 2023-07-04 cs.RO cs.AIstat.AP

classification cs.ROcs.AIstat.AP
keywords knownohelpuncertaintycapabilitieshumaninvolveknowlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) exhibit a wide range of promising capabilities -- from step-by-step planning to commonsense reasoning -- that may provide utility for robots, but remain prone to confidently hallucinated predictions. In this work, we present KnowNo, which is a framework for measuring and aligning the uncertainty of LLM-based planners such that they know when they don't know and ask for help when needed. KnowNo builds on the theory of conformal prediction to provide statistical guarantees on task completion while minimizing human help in complex multi-step planning settings. Experiments across a variety of simulated and real robot setups that involve tasks with different modes of ambiguity (e.g., from spatial to numeric uncertainties, from human preferences to Winograd schemas) show that KnowNo performs favorably over modern baselines (which may involve ensembles or extensive prompt tuning) in terms of improving efficiency and autonomy, while providing formal assurances. KnowNo can be used with LLMs out of the box without model-finetuning, and suggests a promising lightweight approach to modeling uncertainty that can complement and scale with the growing capabilities of foundation models. Website: https://robot-help.github.io

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. MADE: Belief-Driven Dual-Agent Coordination for Autonomous Model Deployment

    cs.SE 2026-08 conditional novelty 6.0 of 10

    A belief-driven two-agent system automatically serves 84 of 122 open-source model repositories as working APIs, and the authors release a benchmark for model-to-API deployment.

  3. From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Coding each mechanism for evidence of lineage and implementation depth, the review closes one candidate gap (GraSP) and isolates five residual control bundles for language agents.

  4. PhyAgentOS: A Self-Evolving Operating System for Embodied Agents with Decoupled Cognitive Planning and Physical Execution

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A file-based operating-system layer with a session verifier and persistent memory improves embodied-agent task completion on game, simulated, and real-robot platforms without retraining policies.

  5. VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs

    cs.RO 2025-12 conditional novelty 6.0 of 10

    VL-LN Bench turns instance-goal navigation into an interactive dialog task, contributes a 41k-trajectory house-scale benchmark with a GPT-4o oracle, and shows active questioning improves embodied agents' success.

  6. GhostShell: Streaming LLM Function Calls for Concurrent Embodied Programming

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    A streaming XML function-token interface with multi-channel scheduling lets robots execute concurrent speech and motion while the LLM is still generating, reportedly beating native function calling 15/15 vs 6/15 on co...

  7. Casper: Inferring Diverse Intents for Assistive Teleoperation with Vision Language Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A VLM-powered assistive teleoperation system infers diverse user intents from teleoperation snippets and executes them with a skill library, outperforming baselines on real-world mobile manipulation tasks.

  8. When Two LLMs Debate, Both Think They'll Win

    cs.CL 2025-05 reject novelty 6.0 of 10

    Frontier LLMs asked to bet on their own win probability in adversarial debates start overconfident and grow more confident each round, even when debating identical copies of themselves.

  9. A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

    cs.RO 2025-02 conditional novelty 6.0 of 10

    IKER uses VLM-generated keypoint rewards to train manipulation policies in simulation that transfer to a real robot, enabling multi-step tasks and replanning.

  10. Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A four-layer systems framework and T0–T5 hierarchy for grading and maintaining bounded trustworthiness claims in embodied AI systems.

  11. SAFER: A Calibrated Risk-Aware Multimodal Recommendation Model for Dynamic Treatment Regimes

    cs.LG 2025-06 reject novelty 5.0 of 10

    SAFER combines tabular EHR and clinical notes to make treatment recommendations with a claimed conformal FDR guarantee, but the proof and evaluation do not support the formal assurances.

  12. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

    cs.RO 2025-08 reject novelty 4.0 of 10

    A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

  13. LLM-based ambiguity detection in natural language instructions for collaborative surgical robots

    cs.RO 2025-07 conditional novelty 4.0 of 10

    An ensemble of five LLM evaluators plus conformal prediction labeled surgical instructions as ambiguous or clear with 70% (Llama 3.2 11B) and 82.5% (Gemma 3 12B) accuracy, measured in-sample on the 40-instruction cali...

  14. WQLCP: Weighted Adaptive Conformal Prediction for Robust Uncertainty Quantification Under Distribution Shifts

    cs.LG 2025-05 reject novelty 4.0 of 10

    WQLCP weights calibration samples by VAE reconstruction losses and scales test scores by a test-loss quantile to improve conformal prediction under shifts, but the algorithm is ill-defined and the empirical support is weak.

  15. Human-Centered Shared Autonomy for Motor Planning, Learning, and Control Applications

    cs.HC 2025-06 unverdicted novelty 3.0 of 10

    A review chapter that organizes BCI, rehabilitation, and assistive robotics under a single adaptive-arbitration framework and illustrates it with the authors' own prior systems.

Pith tools