REVIEW 2 cited by
DeLLMa: Decision Making Under Uncertainty with Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The potential of large language models (LLMs) as decision support tools is increasingly being explored in fields such as business, engineering, and medicine, which often face challenging tasks of decision-making under uncertainty. In this paper, we show that directly prompting LLMs on these types of decision-making problems can yield poor results, especially as the problem complexity increases. To aid in these tasks, we propose DeLLMa (Decision-making Large Language Model assistant), a framework designed to enhance decision-making accuracy in uncertain environments. DeLLMa involves a multi-step reasoning procedure that integrates recent best practices in scaling inference-time reasoning, drawing upon principles from decision theory and utility theory, to provide an accurate and human-auditable decision-making process. We validate our procedure on multiple realistic decision-making environments, demonstrating that DeLLMa can consistently enhance the decision-making performance of leading language models, and achieve up to a 40% increase in accuracy over competing methods. Additionally, we show how performance improves when scaling compute at test time, and carry out human evaluations to benchmark components of DeLLMa.
Forward citations
Cited by 2 Pith papers
-
From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth
Code-owned provisional ranking of typed strategic routes under delayed ground truth shows preliminary retrospective AUC 0.756 on 21 venture cases, with no decomposition advantage and clear leakage risk for LLM judges.
-
The World As Large Language Models See It: Exploring the reliability of LLMs in representing geographical features
GPT-4o and Gemini 2.0 Flash approximate the geography of Austria but show systematic biases in coordinates and elevations and frequent errors in assigning federal states.
Discussion (0). Continue with ORCID to comment.