Pith. sign in

REVIEW 2 cited by

DeLLMa: Decision Making Under Uncertainty with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02392 v3 pith:KRRWJPDD submitted 2024-02-04 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords decision-makingdellmalanguagedecisionlargemodelsaccuracyenhance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The potential of large language models (LLMs) as decision support tools is increasingly being explored in fields such as business, engineering, and medicine, which often face challenging tasks of decision-making under uncertainty. In this paper, we show that directly prompting LLMs on these types of decision-making problems can yield poor results, especially as the problem complexity increases. To aid in these tasks, we propose DeLLMa (Decision-making Large Language Model assistant), a framework designed to enhance decision-making accuracy in uncertain environments. DeLLMa involves a multi-step reasoning procedure that integrates recent best practices in scaling inference-time reasoning, drawing upon principles from decision theory and utility theory, to provide an accurate and human-auditable decision-making process. We validate our procedure on multiple realistic decision-making environments, demonstrating that DeLLMa can consistently enhance the decision-making performance of leading language models, and achieve up to a 40% increase in accuracy over competing methods. Additionally, we show how performance improves when scaling compute at test time, and carry out human evaluations to benchmark components of DeLLMa.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

    cs.AI 2026-07 conditional novelty 5.5 of 10

    Code-owned provisional ranking of typed strategic routes under delayed ground truth shows preliminary retrospective AUC 0.756 on 21 venture cases, with no decomposition advantage and clear leakage risk for LLM judges.

  2. The World As Large Language Models See It: Exploring the reliability of LLMs in representing geographical features

    cs.CY 2025-05 conditional novelty 4.0 of 10

    GPT-4o and Gemini 2.0 Flash approximate the geography of Austria but show systematic biases in coordinates and elevations and frequent errors in assigning federal states.

Pith tools