{"id":"14f9d2fb-8bfb-4d02-a22e-bc8ddab91bbc","arxiv_id":"2506.10546","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Reddit-derived LLM sentiment, augmented with comment voting signals, improves out-of-sample nowcasts of euro area inflation and unemployment, with gains up to 13 percentage points.","lead":"The authors use a large language model to score millions of Reddit posts and comments about inflation and unemployment, then build daily indicators that improve month-ahead euro area nowcasts compared with newspaper sentiment, market swaps, and a simple benchmark. The paper matters because it tests whether social media noise, filtered with AI, can give central banks a cheaper, more timely read on prices and jobs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Selection of the best of 120 Reddit specifications on the same out-of-sample period used for evaluation inflates the reported nowcasting gains; the headline improvements over sentiment and financial variables are not identified without a holdout or multiple-testing correction.","rationale":"The reader's weakest assumption identifies exactly the load-bearing concern: the 120 Reddit specifications are selected on the same out-of-sample data used to evaluate them, so the reported gains are optimistic and no multiple-testing correction is applied. My independent reading confirms this and finds no additional concern of comparable weight. The paper is otherwise carefully executed: the LLM classification is validated against human labels with F1 scores around 0.71-0.75 (Section 2.3); the comment-timing check in Section 2.4 reduces look-ahead risk; the Giacomini-Rossi tests in Appendix B.3 show that gains are concentrated in the COVID/high-inflation period, which is consistent with the stated claim; and the appendix reports results for the full grid, allowing readers to see the dispersion. Those strengths do not, however, address the selection issue: choosing the best of 120 configurations by the same RMSFE criterion that defines the reported performance is a textbook case of selection on the evaluation sample. The magnitude of the advantage over newspaper sentiment and financial variables is therefore not identified without a holdout period, a pre-registered specification, or a formal model confidence set procedure. The reader's CONDITIONAL verdict is appropriate; I would not move it in either direction.","tokens_in":19730,"tokens_out":3668,"duration_ms":46288,"concrete_test":"Run a permutation null test: independently permute the winning daily Reddit indicator values within the out-of-sample period (2018-2023), re-run the full selection among the 120 specifications and the MIDAS horse-race, and record the best RMSFE per target. Repeat at least 500 times. If the observed best Reddit RMSFE (e.g., 0.679 for food inflation) is not below the 5th percentile of this null distribution, the headline gains are attributable to selection overfitting. A complementary check is to split the out-of-sample period: select the best specification on 2018-2020 and evaluate only on 2021-2023; if the selected configurations change materially and the RMSFE gains shrink or disappear, the original claim is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.4 constructs 120 Reddit series per target as the product of 2 comment sets, 2 scoring rules, 5 thresholds, and 6 MA windows. Table 3.2 then reports the single best Reddit specification per target, chosen by RMSFE computed over January 2018 to December 2023 — the same out-of-sample sample on which the horse-race is evaluated. Selecting the best of 120 candidates on the evaluation sample means that even an indicator with no true predictive content will produce an apparent RMSFE improvement: the minimum of 120 exchangeable statistics lies far below the median of the null distribution. The headline claims ('up to 13 percentage points for food price inflation', 'at least 5 percentage points for youth unemployment') compare this selected maximum against a much smaller set of competitors (3 sentiment indices, 5 swaps, oil), so the comparison is tilted. The winning configurations are also heterogeneous across targets — com 60/0.3 for HICP, com 365/0.1 for core, com 365/0.7 for food, com 90/0.3 for all unemployment measures — which is more consistent with selection noise than with a stable economic signal. The LLM F1 validation against human labels (Section 2.3) supports the classification step, but it does not establish that the selected indicator's out-of-sample gains are free of selection overfitting. This is the load-bearing weakness for the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs daily Reddit-based indicators for euro-area inflation and unemployment by using LLaMa-3-70B to classify submissions and comments into UP/DOWN/NEUTRAL signals, then applies combinations of moving-average smoothing, thresholded comment voting, optional upvote/downvote weighting, and two comment sets, yielding 120 indicator specifications per target. These are evaluated one at a time as daily predictors in MIDAS-AR nowcasting regressions against a monthly AR(1) benchmark, newspaper sentiment indices, inflation/output swaps, and oil prices over January 2018 to December 2023. The authors report consistent RMSFE, MAFE, and CRPS gains, with improvements of up to 13 percentage points for food price inflation and at least 5 percentage points for youth unemployment, and they attribute the gains primarily to the social-interaction layer that reclassifies submissions using comment signals.","tokens_in":20012,"tokens_out":5946,"duration_ms":67366,"significance":"If the reported gains are genuine, the paper makes a useful contribution: it introduces a novel high-frequency, LLM-processed social media information set for euro-area nowcasting, and the social-interaction layer is an interesting and transferable idea. The authors are transparent about the LLM prompt, report F1 accuracy against human labels, and use standard Diebold-Mariano/Harvey and Giacomini-Rossi tests. However, the headline empirical claim is not yet identified because the best Reddit specification is selected on the same out-of-sample period used for evaluation, so the reported RMSFE improvements are minima over 120 candidates rather than the performance of a pre-specified model. This selection-overfitting concern is load-bearing for the paper's central claim and must be addressed before the empirical conclusions can be accepted.","major_comments":[{"comment":"The 'winning' Reddit specification is chosen by RMSFE computed on the January 2018 to December 2023 out-of-sample period, which is the same period on which the horse-race performance is reported. This makes the headline ratios (e.g., 0.679 for food, 0.723 for HICP) the minimum over 120 highly correlated statistics rather than the performance of a pre-specified rule; under the null of no predictive content, the minimum of 120 exchangeable statistics lies far below the median, so the reported gains are inflated. The heterogeneity of the winning configurations across targets (com 60/0.3 for HICP, com 365/0.1 for core, com 365/0.7 for food, com 90/0.3 for all unemployment rows) is more consistent with selection noise than with a stable economic signal. The F1 validation in Section 2.3 supports classification quality but does not establish that the selected indicator's time-series gains are free of selection overfitting. Please address this by reporting the full distribution of RMSFE over all 120 specifications, applying a multiple-testing correction such as White's reality check or a model confidence set, and/or selecting specifications on a pre-2018 subsample or a pre-registered rule and evaluating only those on 2018-2023.","section":"Section 2.4 and Table 3.2"},{"comment":"The RMSFE gains for unemployment are not statistically significant against the AR(1) benchmark. For example, for the unemployment rate under 25, the RMSFE ratio is 0.906 with no asterisk and the CRPS ratio is 0.909 with no asterisk; only the MAFE ratio (0.872) is significant at the 5% level. The abstract and conclusion nevertheless claim 'consistent gains' and a minimum gain of 5 percentage points for youth unemployment, where the 5 percentage points is an RMSFE comparison against the best sentiment indicator, not a significant improvement over the benchmark. The wording should be qualified, or the evaluation should be extended (for example by pooling across targets or using a test with higher power), before claiming consistent unemployment gains.","section":"Table 3.2, unemployment rows"},{"comment":"The comparison is asymmetric in candidate set size. Best Reddit is selected from 120 specifications, whereas 'Best Sentiment' is the best of 3 newspaper indices and 'Best Swap' is the best of 5 swap maturities. Even if the selection were performed honestly, comparing the best of 120 Reddit series against the best of a much smaller set tilts the comparison in Reddit's favor. The paper should either compare a fixed Reddit specification (or the distribution of all Reddit specifications) against the full set of competitors, or at least report how many of the 120 Reddit specifications beat the best sentiment and best swap series.","section":"Section 3.1 and Table 3.2"}],"minor_comments":[{"comment":"The text contains a typo: 'Redddit' should be 'Reddit'.","section":"Section 2.1"},{"comment":"The sentence listing the availability of newspaper sentiment indices says 'Germany, France, Italy, and Germany'; the final country should presumably be Spain.","section":"Section 3.1"},{"comment":"The shorthand in Table 3.2 (e.g., 'com 60 0.3 0 1' interpreted as 'firstlevel') is not clearly mapped to the figure legends, which use 'filter' and 'nofilter' labels such as 'comments_60_0.3_noscore_filter' and 'comments_365_0.1_noscore_nofilter'; please state explicitly which comment set corresponds to each label.","section":"Table 3.2 and Figures 3.6, B.11"}],"recommendation":"major_revision","confidential_remarks":"The core empirical claim rests on selecting the best of 120 specifications on the same out-of-sample period used for evaluation; without a holdout or a multiple-testing correction, the abstract's 'consistent gains' cannot be verified. I would support publication after the authors report the full grid distribution and/or a validation split, and after they qualify the unemployment claims in line with the significance results. The paper is otherwise well-executed and the social-interaction layer is a genuinely interesting contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The construction of daily Reddit inflation/unemployment indicators is genuinely new and carefully done. Using LLaMa-3-70b with a forward-looking prompt, validating against human labels with F1 scores (0.71-0.75 vs 0.33-0.34 for dictionaries), and then adding a comment-vote reclassification step that acts as a regularizer is a sensible contribution. The paper shows many specifications beat an AR(1) benchmark, and the Giacomini-Rossi stability tests, information-cutoff checks, and smoothing sensitivity analysis are useful. The authors are transparent about the 120-series grid and report the full spread in Appendix B.6.\n\nThe soft spot is exactly what the stress-test flags. The \"best\" Reddit series per target is chosen by RMSFE over January 2018-December 2023, the same out-of-sample period used for the horse-race. Selecting the minimum of 120 exchangeable statistics guarantees an apparent improvement over the AR(1) even if every series were pure noise, and the comparison to the few non-selected competitors (3 sentiment indices, 5 swaps, oil) is tilted. The winning configurations are heterogeneous across targets (MA 60/0.3 for HICP, 365/0.1 for core, 365/0.7 for food, 90/0.3 for all unemployment measures), which is more consistent with selection noise than with a stable economic signal. The DM p-values are not corrected for the grid search, so the asterisks overstate significance. I also note the unemployment gains are significant only for MAFE, not RMSFE or CRPS, and there is no code, data, or replication package.\n\nStill, I would not call the central claim hollow. Appendix B.6 shows the average RMSFE ratio across all 120 specs is around 0.80 for HICP and 0.93-0.97 for unemployment, meaning even a randomly chosen Reddit indicator usually beats the AR(1). The signal is likely real; the reported 5-13 percentage-point gains over sentiment and swaps are probably optimistic because of the selection.\n\nThis paper is for applied nowcasters in policy institutions and anyone working with social media as a macroeconomic data source. It deserves a serious referee, but the referee should demand a pre-specified indicator or a holdout (e.g., train on 2012-2020, test on 2021-2023), a model confidence set or multiple-testing correction for the 120-spec search, and a replication package with LLM outputs or at least the aggregated series. With those fixes, the magnitude claims would be credible. My vote is to engage with it, not to desk-reject.","headline":"A well-built Reddit-based nowcasting pipeline whose headline gains are inflated by selecting the best of 120 specifications on the evaluation sample; the economic signal looks real, but the magnitudes need a holdout or multiple-testing correction before they can be trusted.","tokens_in":20562,"tokens_out":1711,"would_cite":true,"duration_ms":22987,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Reddit comment-voting signals, scored by a large language model, improve out-of-sample nowcasts of euro-area inflation and unemployment.","keywords":["nowcasting","euro area","Reddit","large language models","sentiment analysis","inflation","unemployment","MIDAS"],"falsifier":"Run the specification search only on data through 2020 and then evaluate the chosen Reddit indicator on 2021-2023; if it no longer beats the AR(1) benchmark and the newspaper or swap indicators, the claimed out-of-sample gains are an artifact of in-sample selection.","tokens_in":1595,"feed_emoji":"📈","tokens_out":1703,"duration_ms":91546,"temperature":0.7,"pith_summary":"The paper sets out to show that social media, and Reddit in particular, carries timely information about euro-area inflation and unemployment that official statistics capture only with a lag. It builds daily indicators by asking a large language model to classify millions of Reddit submissions and comments as signaling that prices or unemployment will move up, down, or stay neutral, and then lets the comments vote on each submission's original label. These indicators are tested in an out-of-sample nowcasting exercise against newspaper sentiment, financial variables, and a standard AR(1) benchmark. The reported outcome is consistent gains for the Reddit indicators, with the largest improvements during the COVID-19 recession and the 2021-2023 inflation surge. A sympathetic reader would take the paper's contribution to be evidence that community discussion, not just news supply, contains economically meaningful forward-looking signals.","feed_headline":"Reddit chatter sharpens euro-area inflation and job forecasts","feed_subtitle":"LLM-scored Reddit posts plus comment votes beat newspaper sentiment and swaps, cutting nowcast errors by up to 13 points.","key_machinery":"The device that carries the argument is the comment-weighted vote score $L_i = (S_i + \\sum_{j=1}^{J} C_{i,j})/(J+1)$, where $S_i$ is the LLM's UP, DOWN, or NEUTRAL label for submission $i$ and $C_{i,j}$ are the LLM labels of its comments, optionally weighted by each comment's upvote-minus-downvote net score. A threshold $\\tau$ maps $L_i$ back to a revised label $\\bar{S}_i$: UP if $L_i>\\tau$, DOWN if $L_i<-\\tau$, NEUTRAL otherwise, and daily signals $\\bar{X}_t = \\sum_i \\bar{S}_i$ are smoothed with backward-looking moving averages. The paper varies the comment set, the upvote weighting, the threshold, and the smoothing window to produce 120 candidate series, and evaluates each in a mixed-frequency MIDAS-AR regression with Almon-polynomial weights against a monthly AR(1) benchmark.","core_discovery":"The central claim is that a comment-voting scheme applied to LLM-classified Reddit posts yields daily indicators that beat daily newspaper sentiment, inflation swaps, and oil prices in nowcasting eight euro-area series: overall HICP and its core, energy, food, and services components, plus total, over-25, and under-25 unemployment. The paper reports relative RMSFE improvements over the AR(1) benchmark ranging up to 13 percentage points for food price inflation and at least 5 percentage points for youth unemployment. The authors argue that the LLM's forward-looking, context-sensitive classification of informal Reddit text is what makes the raw signal work, and that letting comment votes revise a submission's original UP, DOWN, or NEUTRAL label regularizes the signal, reclassifying about 11 percent of inflation posts and 15 percent of unemployment posts.","pith_inferences":["(Editorial inference) If the gains survive a genuine holdout, the same pipeline should transfer to other European subreddits and to targets such as GDP or housing, giving country-level daily indicators; the paper itself only tests the euro area as a whole through r/europe.","(Editorial inference) The comment-voting scheme acts like a cheap ensemble regularizer, but comment votes are not independent observations because later commenters have already read earlier comments; weighting comments by depth, time lag, or author diversity could either strengthen the signal or reveal that a few active users drive it.","(Editorial inference) A direct testable extension would compare the Reddit signal against demographic-specific survey expectations, since the paper's larger gains for food inflation and youth unemployment fit the idea that personally experienced prices shape expectations more than abstract index numbers."],"forward_implications":["Reddit-derived daily signals can be added to the nowcaster's toolkit for euro-area prices and labor markets, with the largest reported gains for food price inflation and meaningful gains for youth unemployment.","Including comments rather than submissions alone improves nowcasting accuracy: the winning specification for every target variable uses the comment-voting scheme.","The gains are concentrated in unusual periods, namely the COVID-19 recession and the high-inflation episode of 2021-2023, consistent with high-frequency information mattering most during turmoil.","The LLM's classification of informal Reddit text reaches F1 scores around 0.71 to 0.75, well above dictionary-based baselines, so the signal extraction is reproducible with a general-purpose large language model.","The information value persists when the daily information cutoff is moved earlier in the month, with gains clear up to about 14 days before the end of the nowcast period."],"supporting_citations":[{"why":"Constructs the daily newspaper sentiment indices used as the main textual competitor and one of the dictionary benchmarks.","marker":"Barbaglia et al. (2024)"},{"why":"Provides the MIDAS regression machinery that mixes daily indicators with monthly targets.","marker":"Ghysels et al. (2016)"},{"why":"Supplies the euro-area target series for inflation and unemployment used in the horse race.","marker":"Barigozzi and Lissona (2024)"},{"why":"Demonstrates LLM extraction of economic sentiment from news, the closest methodological precedent the paper extends to social media.","marker":"Bybee (2023)"},{"why":"Shows LLMs can simulate economic beliefs, supporting the claim that LLM labels capture forward-looking expectations.","marker":"Horton (2023)"},{"why":"Documents the informal, slang-heavy language of Reddit that motivates using an LLM instead of dictionaries.","marker":"Long et al. (2023)"},{"why":"Provides the dictionary-based sentiment benchmark used to compare against the LLM.","marker":"Granziera et al. (2025)"},{"why":"Justifies the one-sided Diebold-Mariano tests used despite the models being nested.","marker":"Clark and Ravazzolo (2015)"},{"why":"Provides the fluctuation test used to show that Reddit gains concentrate in the COVID-19 and high-inflation episodes.","marker":"Giacomini and Rossi (2010)"}],"fun_headline_variants":["Reddit posts sharpen euro inflation and jobless forecasts","LLM-scored Reddit beats news for euro-area nowcasts","Comment votes on Reddit refine euro inflation signals","Reddit chatter nowcasts euro inflation, beating swaps","Euro nowcasts get a Reddit boost in unusual times"],"cache_read_input_tokens":22656,"weakest_assumption_plain":"The reported gains assume that picking the best of 120 Reddit indicator designs on the same 2018-2023 period used to score them did not overfit that period.","fun_headline_variants_meta":{"raw":{"variants":["Reddit posts sharpen euro inflation and jobless forecasts","LLM-scored Reddit beats news for euro-area nowcasts","Comment votes on Reddit refine euro inflation signals","Reddit chatter nowcasts euro inflation, beating swaps","Euro nowcasts get a Reddit boost in unusual times"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1535,"prompt_tokens":826,"completion_tokens":709,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":630}},"tokens_in":442,"tokens_out":709,"duration_ms":9557,"temperature":1.0,"reasoning_tokens":630,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:23:42.007710+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the specification search only on data through 2020 and then evaluate the chosen Reddit indicator on 2021-2023; if it no longer beats the AR(1) benchmark and the newspaper or swap indicators, the claimed out-of-sample gains are an artifact of in-sample selection.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Constructs the daily newspaper sentiment indices used as the main textual competitor and one of the dictionary benchmarks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the MIDAS regression machinery that mixes daily indicators with monthly targets."},{"cited_title":"and Lissona, C","cited_arxiv_id":null,"evidence_quote":"Supplies the euro-area target series for inflation and unemployment used in the horse race."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows LLMs can simulate economic beliefs, supporting the claim that LLM labels capture forward-looking expectations."},{"cited_title":"i just like the stock","cited_arxiv_id":null,"evidence_quote":"Documents the informal, slang-heavy language of Reddit that motivates using an LLM instead of dictionaries."},{"cited_title":"H., Meggiorini, G., and Melosi, L","cited_arxiv_id":null,"evidence_quote":"Provides the dictionary-based sentiment benchmark used to compare against the LLM."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies the one-sided Diebold-Mariano tests used despite the models being nested."},{"cited_title":"and Rossi, B","cited_arxiv_id":null,"evidence_quote":"Provides the fluctuation test used to show that Reddit gains concentrate in the COVID-19 and high-inflation episodes."}],"review_version":1}