Pith. sign in

REVIEW 2 cited by

Humans vs Large Language Models: Judgmental Forecasting in an Era of Advanced AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.06941 v2 pith:QCF3DEO5 submitted 2023-12-12 cs.LG cs.CY

classification cs.LGcs.CY
keywords forecastingllmsforecastershumanadvancedmodelsaccuracyduring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study investigates the forecasting accuracy of human experts versus Large Language Models (LLMs) in the retail sector, particularly during standard and promotional sales periods. Utilizing a controlled experimental setup with 123 human forecasters and five LLMs, including ChatGPT4, ChatGPT3.5, Bard, Bing, and Llama2, we evaluated forecasting precision through Mean Absolute Percentage Error. Our analysis centered on the effect of the following factors on forecasters performance: the supporting statistical model (baseline and advanced), whether the product was on promotion, and the nature of external impact. The findings indicate that LLMs do not consistently outperform humans in forecasting accuracy and that advanced statistical forecasting models do not uniformly enhance the performance of either human forecasters or LLMs. Both human and LLM forecasters exhibited increased forecasting errors, particularly during promotional periods and under the influence of positive external impacts. Our findings call for careful consideration when integrating LLMs into practical forecasting processes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting Empirical AI Research Outcomes with Language Models

    cs.AI 2025-06 conditional novelty 7.0 of 10

    A fine-tuned, retrieval-augmented language model predicts which of two AI research ideas will perform better empirically, beating human experts and frontier models on a new contamination-controlled benchmark.

  2. Argumentatively Coherent Judgmental Forecasting

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Filtering forecasts that are incoherent with their argumentative reasoning improved accuracy in human and LLM experiments, though users do not naturally follow the proposed coherence notion.

Pith tools