REVIEW 3 major objections 5 minor 1 cited by
ChatGPT and Deepseek: Can They Predict the Stock Market and Macroeconomy?
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read ChatGPT's good-news ratio, the fraction of Wall Street Journal headlines it labels positive, predicts up to six months of future stock-market returns, with an out-of-sample R2 of 1.17%.
desk verdict Original and careful empirical work, but the headline out-of-sample result has an unresolved standardization detail that could be a look-ahead leak; the paper deserves refereeing, not desk rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the ChatGPT good-news ratio, NRG, computed each month as the share of Wall Street Journal front-page headlines and alerts that GPT-3.5 labels GOING UP under a fixed zero-shot prompt. The machinery is the predictive regression of future average market excess returns on NRG, with overlapping-horizon standard errors for in-sample inference and the standard out-of-sample $R^2_{OS}$ test against the historical average. The ratio carries the argument because, in the paper's account, the large model captures context-sensitive meaning in headlines—words like 'bounce' and 'buoy'—that word lists and smaller models miss.
What would settle it
Recompute the out-of-sample $R^2_{OS}$ using only expanding-window standardization of the good-news ratio, or raw counts, and compare with the reported 1.17%; if it collapses, the reported gain comes from full-sample normalization rather than real-time forecasting. Then rerun the 1996-2005 in-sample regressions with a language model whose training data ends before 1996; if predictability disappears in the earlier window and survives only after 2006, look-ahead bias rather than underreaction is the cause.
Extended reading notes
Core claim
Using front-page headlines and alerts from the Wall Street Journal for January 1996 to December 2022, the paper asks GPT-3.5, with a single zero-shot prompt, to label each headline as GOING UP, GOING DOWN, or UNKNOWN for U.S. stock prices. The monthly good-news ratio NRG is the fraction of headlines labeled GOING UP. The paper reports that a one-standard-deviation increase in NRG is followed by 0.53% higher average excess market return over the next month, and the predictive regression $R^2$ rises from 1.37% at the one-month horizon to 8.52% at the twelve-month horizon. Recursive out-of-sample forecasts from January 2006 to December 2022 produce $R^2_{OS} = 1.17\%$ against the historical-average benchmark, and a mean-variance investor with risk aversion 3 earns a certainty-equivalent gain of 4.92% per year (3.55% after 50 basis point transaction costs). The paper interprets the asymmetry—good news predicts, bad news does not—as investors' underreaction to good news, consistent with ambiguity aversion and limited attention, and supports it with interaction tests showing stronger predictability in downturns, high economic-policy uncertainty, and novel news.
Load-bearing premise
The load-bearing assumption is that the out-of-sample forecasts use only information available at the forecast date, and that a model trained through September 2021 can classify 1996-2005 headlines without encoding later market outcomes into its labels.
Editorial extensions
If this is right
- If the central claim is right, the ChatGPT good-news ratio joins the short list of predictors that beat the historical average out of sample, and it does so over horizons investors care about: one to six months.
- The absence of predictive power in the bad-news ratio implies that the market's fast reaction to bad news is not an artifact of the model; the asymmetry itself is the economic content.
- Because predictability concentrates in downturns, high-policy-uncertainty periods, and novel news, a real-time trading rule would tilt toward good-news signals in exactly those states rather than applying the ratio uniformly.
- The estimated economic value—4.92% annual certainty-equivalent gain at risk aversion 3, still 3.55% after 50 basis point transaction costs—means a mean-variance investor should be willing to pay a substantial fee for the forecast.
Reading between the lines
- If the underreaction story is right, the signal should decay as LLM-based sentiment tools become standard investment inputs; the paper's own cumulative forecast-error plot already appears to flatten after 2021, consistent with a publication effect.
- The DeepSeek comparison raises a testable model-specific question: a replication with other large English-trained models should show whether the predictive edge is about English-language training intensity or about something specific to the GPT architecture.
- The macro results imply a sharper, untested prediction: the good-news ratio should lead not just the equity premium but also real activity indicators such as payroll growth and industrial production at horizons beyond one month, and that lead should survive controls for lagged macro announcements.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether large language models can extract predictive information from Wall Street Journal front-page headlines for the aggregate U.S. stock market and the macroeconomy. Using ChatGPT-3.5 and DeepSeek-R1 to classify headlines from 1996 to 2022 into good, bad, and neutral news, the authors construct monthly good-news and bad-news ratios and test their in-sample and out-of-sample predictive power for S&P 500 excess returns. The central finding is that ChatGPT's good-news ratio is positively correlated with contemporaneous returns and predicts subsequent returns for up to six months, with a one-month in-sample R2 of 1.37% and an out-of-sample R2_OS of 1.17%. The bad-news ratio, DeepSeek, word-list methods, BERT, and RoBERTa do not predict future returns. The paper also connects the predictability to underreaction to good news during economic downturns, high economic policy uncertainty, and novel news, and shows that ChatGPT's news ratios forecast macroeconomic conditions.
Significance. If the central claim holds, the paper makes a valuable contribution by showing that an LLM-based news measure contains real-time information about the equity risk premium that traditional predictors and dictionary methods miss. The empirical design has clear strengths: the regression framework is standard, Hodrick (1992) standard errors are used for overlapping horizons, the out-of-sample evaluation uses Campbell and Thompson (2008) and Clark and West (2007), and the paper includes multiple robustness checks with alternative prompts, fine-tuned models, and ChatGPT-4. The economic-mechanism tests in Section 5 are thoughtful and falsifiable. The main quantitative results, however, are modest in magnitude and depend on the details of the out-of-sample protocol, especially the standardization of the predictor and the training-data cutoff of the language model. Because the paper does not report code or machine-checked proofs, its contribution rests on the credibility of the empirical procedure; that procedure is currently not fully specified in two load-bearing respects.
major comments (3)
- [Section 4.5, Table 6]
- [Section 4.7]
- [Tables 2-5]
minor comments (5)
- [Section 4.5, Equation (4)]
- [Section 3.5, Figure 1 discussion]
- [Section 4.5, Equations (5) and (6)]
- [Throughout]
- [Section 3.3]
Circularity Check
No circularity: the ChatGPT news-ratio predictor is an external LLM classification, and the out-of-sample evaluation re-estimates only alpha and beta recursively.
full rationale
The derivation chain is: a fixed prompt asks ChatGPT-3.5 to classify WSJ headlines as GOING UP, GOING DOWN, or UNKNOWN; the monthly good-news ratio NRG is then used as a regressor for future S&P 500 excess returns. The predictor is produced outside the return regression and is not fitted to the return series being predicted. The out-of-sample procedure in Section 4.5 estimates only alpha_t and beta_t recursively and compares against the historical-average benchmark, so R2_OS of 1.17% is an independent forecast-evaluation result rather than a restatement of the in-sample fit. Robustness comparisons to word lists, BERT, macro controls, and the post-September-2021 weekly forecasts provide external grounding. Citations to coauthor work, such as Lin, Wu, and Zhou (2018) and Rapach and Zhou (2022), are methodological or contextual and are not load-bearing for the central claim. A possible concern is that the standardization described in table notes ('All the forecast variables are standardized to have a zero mean and unit variance') could be applied with full-sample moments, which would be look-ahead leakage; however, this is a data-handling ambiguity rather than a circular construction, and the paper does not specify the OOS standardization rule. No step in the paper's own equations makes the prediction equivalent to its inputs by construction.
Assumptions & free parameters
free parameters (4)
- beta_NRG in one-month predictive regression =
0.53% per standardized unit
- Recursive OLS coefficients in Equation (4) =
estimated monthly
- Risk aversion coefficient and transaction cost in economic value section =
gamma=1,3,5; TC=50bp
- Variance forecast window for asset allocation =
5-year moving window
assumptions (4)
- domain assumption GPT-3.5's judgments of whether a headline is good or bad news are stable across time and unaffected by its training cutoff (September 2021).
- domain assumption The WSJ front-page corpus from the Dow Jones archive is complete and correctly matched to months.
- domain assumption Monthly good and bad news ratios linearly relate to future excess returns after standardization.
- standard math Statistical inference with Hodrick (1992) and Clark-West (2007) is valid for the small sample and persistent regressors.
Cite this review
Pith. "Pith review of ChatGPT and Deepseek: Can They Predict the Stock Market and Macroeconomy?." pith.science (2026). https://pith.science/paper/GSSL23GN
@misc{pith2026250210008,
author = {Pith},
title = {Pith review of: ChatGPT and Deepseek: Can They Predict the Stock Market and Macroeconomy?},
year = {2026},
howpublished = {\url{https://pith.science/paper/GSSL23GN}},
note = {Machine review of arXiv:2502.10008}
}
read the original abstract
We study whether ChatGPT and DeepSeek can extract information from the Wall Street Journal to predict the stock market and the macroeconomy. We find that ChatGPT has predictive power. DeepSeek underperforms ChatGPT, which is trained more extensively in English. Other large language models also underperform. Consistent with financial theories, the predictability is driven by investors' underreaction to positive news, especially during periods of economic downturn and high information uncertainty. Negative news correlates with returns but lacks predictive value. At present, ChatGPT appears to be the only model capable of capturing economic news that links to the market risk premium.
Forward citations
Cited by 1 Pith paper
-
Electricity Market Predictability: Virtues of Machine Learning and Links to the Macroeconomy
Machine learning models, especially boosted trees and GLM, predict Singapore daily electricity returns out-of-sample, and a correlation-penalized ensemble beats all individual models, with predictability concentrated ...
Reference graph
Works this paper leans on
-
[1]
Anastassia, F. (2023). Front-page news: The effect of news positioning on financial markets. Forthcoming inThe Journal of Finance. Ang,A.andG.Bekaert(2007). Stockreturnpredictability: Isitthere? TheReviewofFinancial Studies 20, 651–707. Antweiler, W. and M. Z. Frank (2004). Is all that talk just noise? The information content of internet stock message boa...
arXiv 2023
-
[2]
This setting aligns with a contemporaneous regression framework
Forecasting Results for RoBERTa This table reports estimation results of the following regression, 𝑅𝑡+ℎ = 𝛼+𝛽𝑁𝑅𝐾 𝑡 +𝜀𝑡+ℎ , 𝐾 =𝐵 or𝐺, where 𝑅𝑡 denotes the current market excess return on the S&P 500 index at time𝑡 for "ℎ = 0". This setting aligns with a contemporaneous regression framework. For scenarios whereℎ > 0,𝑅𝑡+ℎ represents the average excess return...
work page 1992
-
[3]
Comparisons with Economic Variables This table reports estimation results of the following regression, 𝑅𝑡+ℎ = 𝛼+𝛽1𝑁𝑅𝐺 𝑡 +𝛽2𝑃𝐶 1𝑡+𝛽3𝑃𝐶 2𝑡+𝛽4𝑃𝐶 3𝑡+𝛽5𝑃𝐶 4𝑡++𝛽6𝑃𝐶 5𝑡+𝜀𝑡+ℎ , where𝑅𝑡+ℎ represents the average excess returns of the market portfolio from𝑡+ 1 to𝑡+ℎ (with ℎ being 1, 3, 6, 9, and 12 months);𝑁𝑅𝐺 represents the proportion of good news identified by Cha...
work page 2008
-
[4]
𝑁𝑅𝐺 isthemonthlyproportionofgoodnewsidentifiedbyChatGPT-3.5
Control for Lagged Return This table reports estimation results of the following regression, 𝑅𝑡+ℎ = 𝛼+𝛽𝑁𝑅𝐺 𝑡 +𝜓𝑅𝑡+𝜀𝑡+ℎ , where𝑅𝑡+ℎ is the average excess returns of market portfolio from𝑡+ 1 to𝑡+ℎ, ℎ = 1, 3, 6, 9, and 12 months. 𝑁𝑅𝐺 isthemonthlyproportionofgoodnewsidentifiedbyChatGPT-3.5. Itanswerswhether the input news meansGOING UPfor the stock market.𝑅𝑡...
work page 1992
-
[5]
Mean Std. Dev. Skew. Median Min. Max. No. of News Total News 260.91 116.12 −0.22 288 47 577 84535 Panel A: Results for ChatGPT Bad News 32.76 23.93 1.15 32 0 142 10613 NeutralNews 181.75 74.34 −0.07 190 43 422 58888 Good News 46.40 28.25 0.10 47 0 121 15034 Panel B: Results for DeepSeek Bad News 37.51 24.43 0.93 38 1 137 12153 NeutralNews 180.46 77.40 −0....
work page 1992
-
[6]
This setting aligns with a contemporaneous regression framework
Bad News Ratio (𝑁𝑅𝐵) GoodNews Ratio (𝑁𝑅𝐺) 𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) 𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) Panel A: Results for ChatGPT ℎ=0 −1.03∗∗∗ −2.70 5.17 1.02 ∗∗∗ 5.30 5.07 ℎ=1 0.05 0.14 0.01 0.53 ∗∗ 2.22 1.37 ℎ=3 0.06 0.18 0.05 0.56 ∗∗ 2.34 4.60 ℎ=6 0.14 0.47 0.48 0.51 ∗ 1.84 6.91 ℎ=9 0.21 0.76 1.57 0.44 1.51 7.46 ℎ=12 0.25 0.94 2.65 0.42 1.48 8.52 Panel B: Resul...
work page 2011
-
[7]
This setting aligns with a contemporaneous regression framework
Panel A: Results for𝑁𝑅𝐵 Panel B: Results for𝑁𝑅𝐺 𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) 𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) Word Lists ℎ= 0 −0.77∗∗∗ −2.68 2.90 0.24 1.05 0.27 ℎ= 1 −0.17 −0.47 0.14 −0.23 −0.87 0.27 ℎ= 3 −0.14 −0.56 0.28 −0.14 −0.77 0.27 ℎ= 6 0.05 0.27 0.06 −0.09 −0.74 0.20 ℎ= 9 0.09 0.46 0.31 −0.14 −1.38 0.70 ℎ= 12 0.11 0.49 0.56 −0.11 −1.36 0.60 Bert ℎ= 0 −0.38 −1...
work page 1992
-
[8]
This setting aligns with a contemporaneous regression framework
𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) 𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) Panel A:PessimisticNews Panel B:OptimisticNews ℎ= 0 −0.84∗∗ −2.34 3.46 1.09 ∗∗∗ 5.55 5.76 ℎ= 1 0.03 0.10 0.01 0.43 ∗ 1.92 0.88 ℎ= 3 0.07 0.20 0.06 0.54 ∗∗ 2.52 4.42 ℎ= 6 0.11 0.36 0.29 0.52 ∗∗ 2.07 7.37 ℎ= 9 0.18 0.62 1.13 0.49 ∗ 1.80 9.30 ℎ= 12 0.25 0.90 2.80 0.47 ∗ 1.71 10.99 Panel C:NegativeNews Panel D...
work page 1992
Show all 23 references
-
[9]
𝑁𝑅𝐺 (𝑁𝑅𝐵) is the monthly proportion of good (bad) news identified by ChatGPT-3.5
Panel A: Results for𝑁𝑅𝐵 Panel B: Results for𝑁𝑅𝐺 𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) 𝛽 (%) Hodrick- 𝑡 𝑅 2 (%) ChatGPT-3.5 Fine–tuning ℎ= 0 −1.38∗∗∗ −4.69 9.20 0.80 ∗∗∗ 4.04 3.11 ℎ= 1 0.14 0.48 0.09 0.50 ∗∗ 2.25 1.23 ℎ= 3 0.26 0.82 0.93 0.43 ∗∗ 2.44 2.71 ℎ= 6 0.21 0.71 1.05 0.46 ∗∗ 2.57 5....
2018
-
[10]
The stock market return forecasts are generated by the regression model based on news ratios (𝑁𝑅𝐵 or 𝑁𝑅𝐺)
𝑅2 𝑂𝑆 (%) MSFE-adjusted 𝑝-value 𝑁𝑅𝐺 1.17∗∗ 2.17 0.01 𝑁𝑅𝐵 −2.55 −1.07 0.86 MC 0.39 1.03 0.15 IMC 1.36 ∗∗ 1.70 0.04 IWC 2.51 ∗∗∗ 2.33 0.01 Economic Variables −0.41 0.07 0.47 60 Table 7: Economic Values This table presents the CER gains in percentage points and annualized Sharpe ...
2006
-
[11]
𝑁𝑅𝐾 repre- sents news ratios
CER gain (%) CER gain (%), TC=50bp Sharpe Ratio Panel A: Results for𝛾= 1 𝑁𝑅𝐵 −1.49 −2.29 0.15 𝑁𝑅𝐺 5.63 4.73 0.51 Panel B: Results for𝛾= 3 𝑁𝑅𝐵 −3.23 −3.90 −0.07 𝑁𝑅𝐺 4.92 3.55 0.51 Panel C: Results for𝛾= 5 𝑁𝑅𝐵 −3.21 −3.67 −0.15 𝑁𝑅𝐺 2.97 1.51 0.53 61 Table 8: Forecasting Macroeco...
1987
-
[12]
𝑁𝑅𝐾 repre- sents news ratios
Panel A: Bad News Panel B: Good News 𝛽 NW-𝑡 𝑅 2 (%) 𝛽 NW-𝑡 𝑅 2 (%) IPG −0.18∗∗ −1.97 3.33 0.09 ∗∗ 2.35 0.89 VIX 0.33 ∗∗∗ 3.80 11.06 −0.14∗∗ −2.55 2.05 CFNAI −0.17∗ −1.69 2.84 0.12 ∗∗∗ 4.30 1.38 ADSI −0.13 −1.37 1.56 0.10 ∗∗∗ 3.66 1.00 KCFSI 0.42 ∗∗∗ 3.30 17.79 −0.31∗∗∗ −5.96 9...
2006
-
[13]
In each quarter, SPF provides the forecasts for current quarter (nowcasting) and subsequent four quarters
Bad News Ratio (𝑁𝑅𝐵) Good News Ratio (𝑁𝑅𝐺) 𝛽 NW-𝑡 𝑅 2 (%) 𝛽 NW-𝑡 𝑅 2 (%) Panel A: Forecasting Macroeconomy IPG −0.22∗∗ −1.96 4.67 0.04 0.89 0.17 VIX 0.39 ∗∗∗ 4.76 14.89 −0.04 −0.62 0.14 CFNAI −0.21∗ −1.75 4.46 0.05 1.30 0.26 ADSI −0.16 −1.58 2.69 0.05 1.38 0.22 KCFSI 0.46 ∗∗∗ ...
1987
-
[14]
Panel A: Bad News Panel B: Good News 𝛽 NW-𝑡 𝜓 NW-𝑡 𝑅 2 (%) 𝛽 NW-𝑡 𝜓 NW-𝑡 𝑅 2 (%) Nowcasting −0.20∗∗ −2.46 −0.14 −0.60 9.50 0.04 0.92 0.06 0.25 0.92 1 quarter −0.09∗∗∗ −3.10 0.54 ∗∗∗ 8.06 38.04 0.06 ∗ 1.70 0.59 ∗∗∗ 4.54 36.69 2 quarter −0.09∗∗∗ −4.10 0.56 ∗∗∗ 6.84 42.21 0.03 0....
1992
-
[15]
𝑁𝑅𝐺 isthemonthlyproportion of good news identified by ChatGPT-3.5
ℎ= 1 ℎ= 3 ℎ= 6 ℎ= 9 ℎ= 12 𝐼𝐻𝑖𝑔ℎ ×𝑁𝑅𝐺 0.12 0.17 −0.03 −0.08 −0.06 [0.44] [0.68] [ −0.12] [ −0.38] [ −0.38] 𝐼𝐿𝑜𝑤×𝑁𝑅𝐺 0.71∗ 0.83∗∗∗ 0.97∗∗∗ 0.93∗∗∗ 0.84∗∗∗ [1.86] [2.91] [3.35] [3.08] [3.08] 𝐼𝐻𝑖𝑔ℎ 0.90 0.55 0.53 ∗ 0.42 0.44 ∗ [1.58] [1.49] [1.73] [1.40] [1.74] 𝑅2 (%) 1.79 6.36 14...
1992
-
[16]
ℎ= 1 ℎ= 3 ℎ= 6 ℎ= 9 ℎ= 12 𝑈𝐻𝑖𝑔ℎ ×𝑁𝑅𝐺 0.78∗∗ 0.81∗∗∗ 0.84∗∗∗ 0.89∗∗∗ 0.85∗∗∗ [2.20] [3.11] [2.90] [3.02] [2.90] 𝑈𝐿𝑜𝑤×𝑁𝑅𝐺 0.26 0.29 0.15 −0.05 −0.05 [0.99] [1.07] [0.59] [ −0.23] [ −0.25] 𝑈𝐻𝑖𝑔ℎ 0.13 0.26 0.02 −0.07 0.09 [0.28] [0.72] [0.05] [ −0.23] [0.39] 𝑅2 (%) 0.79 4.95 9.35 ...
1992
-
[17]
It answers whether the input news meansGOING UPor GOING Down for stock market
Weekly Out-of-sample Forecasting Errors This figure plots the weekly difference between the cumulative squared forecast error (CSFE) gener- atedbythehistoricalaveragebenchmarkforecastandtheCSFEderivedfromforecastsbasedongood news ratio (𝑁𝑅𝐺) or bad news ratio (𝑁𝑅𝐵), which are ...
2021
-
[18]
This setting aligns with a contemporaneous regression framework
Results for Alternative Statistical Standard Errors This table reports estimation results of the following regression, 𝑅𝑡+ℎ = 𝛼+𝛽𝑁𝑅𝐾 𝑡 +𝜀𝑡+ℎ , 𝐾 =𝐵 or𝐺, where 𝑅𝑡 denotes the current market excess return on the S&P 500 index at time𝑡 for "ℎ = 0". This setting aligns with a cont...
1987
-
[22]
This setting aligns with a contemporaneous regression framework
Additional Results of Alternative Prompts This table reports estimation results of the following regression, 𝑅𝑡+ℎ = 𝛼+𝛽𝑁𝑅𝐾 𝑡 +𝜀𝑡+ℎ , 𝐾 =𝐵 or𝐺, where 𝑅𝑡 denotes the current market excess return on the S&P 500 index at time𝑡 for "ℎ = 0". This setting aligns with a contemporaneou...
1992
-
[23]
𝑁𝑅𝐾 represents news ratios
Forecasting Macroeconomic Conditions This table presents the results of following regression, 𝑌𝑡+1= 𝛼+𝛽𝑁𝑅𝐾 𝑡 +𝜀𝑡+1, 𝐾 =𝐵 or𝐺, where𝑌𝑡+1 represents macroeconomic condition variables at future time𝑡+ 1, including the indus- try production growth (IPG), the VIX of CBOE, the CFNAI...
2011
-
[30]
Veronesi, P. (1999). Stock market overreactions to bad news in good times: A rational expectations equilibrium model.The Review of Financial Studies 12, 975–1007. Wei, J., Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. ...
1999 arXiv
-
[2022]
2006 2009 2012 2015 2018 2021 0 0.005 0.01 0.015 0.02 0.025 54 Table 1: Summary Statistics ofWall Street Journal News This table reports the mean, standard deviation (Std
Grey shadow bars denote NBER recessions. 2006 2009 2012 2015 2018 2021 0 0.005 0.01 0.015 0.02 0.025 54 Table 1: Summary Statistics ofWall Street Journal News This table reports the mean, standard deviation (Std. Dev.), skewness (Skew.), median, minimum (min.) and maximum (Max...
2006
-
[3557]
Guo, D., D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. (2025). Deepseek-r1: Incentivizingreasoningcapabilityinllmsviareinforcementlearning. arXiv preprint arXiv:2501.12948. Hansen, P. R. and A. Timmermann (2012). Choice of sample split in o...
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.