{"id":"e93855e9-a632-4ecb-a926-1649d5045bfa","arxiv_id":"2607.24072","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LLM-derived multidimensional sentiment shows stronger but still asset-dependent statistical links to meme-stock returns than VADER, with no stable forecasting advantage.","lead":"This paper compares two ways of reading Reddit chatter about meme stocks—a simple word-list method (VADER) and a large language model—to see which better anticipates wild price moves in GME, AMC, and NOK. It finds that the LLM yields a richer emotional and bullishness profile, but neither method reliably forecasts next-day returns, so richer language does not mean better prediction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AMC LMI—the sole significant support for the central claim—may not survive a proper pre-close cutoff and HAC-robust errors; both are acknowledged limitations in §3.7–§3.8.","rationale":"The reader's weakest assumption—unvalidated LLM labels—is real and worth flagging, but I see an even more direct threat to the central claim. The paper's positive case reduces to one regression: AMC LMI with p = 0.000005 and R² = 0.0831. The paper itself concedes the two conditions that could make that p-value spurious: no intraday market-close cutoff (§3.7) and conventional standard errors unadjusted for serial dependence (§3.8). These are not external speculations; they are stated limitations. If a proper cutoff removes post-close posts, the LMI_t variable changes meaning, and if HAC standard errors inflate the p-value, the single significant result may disappear. In that case the abstract's 'stronger asset-specific statistical structure' would be unsupported, leaving only a robust negative finding about VADER and heterogeneity across assets—still interesting, but not the claimed advantage for LLM multidimensionality. I do not think this warrants rejection: the paper is appropriately cautious, and the checks are straightforward. Nor do I think it moves the verdict from the reader's CONDITIONAL. The appropriate response is unchanged: conditional acceptance pending robustness checks. The label-validation issue is secondary for the empirical comparison because the market-association claim can survive even if LLM outputs are noisy proxies, though the semantic interpretation would weaken. Hence partial agreement with the reader.","tokens_in":11859,"tokens_out":7854,"duration_ms":74656,"concrete_test":"Re-estimate the AMC LMI row of Table 3 with (i) LMI recomputed using only entries timestamped strictly before 16:00 U.S. Eastern on each calendar day and (ii) Newey-West HAC standard errors (e.g., lag 5), and report the p-value. Also apply a Bonferroni correction across the 9 asset–indicator regressions. If the pre-close/HAC p-value rises above 0.05 or the coefficient changes sign, the paper's single significant support for 'stronger statistical structure' is an artifact of alignment/inference; if it remains < 0.05, that support survives and the claim is materially strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"§4.2 reports exactly one significant LLM-comparison result: AMC LMI, β = −0.000041, p = 0.000005, R² = 0.0831. The abstract's 'stronger asset-specific statistical structure' is carried by this single row. Two self-acknowledged design choices threaten it. First, §3.7 states no exchange-time cutoff or timezone conversion was applied: LMI_t aggregates entries from the full calendar day, including posts/comments made after the U.S. market close on day t. Those entries fall inside the close-to-close interval that defines R_{t+1}, so the regression is not strictly ex ante. Figure 1's pronounced contemporaneous AMC correlation followed by next-lag reversal suggests the negative β may be mean reversion to same-day discourse reaction, not predictive signal. Second, §3.8 reports conventional OLS standard errors with no adjustment for heteroskedasticity, serial dependence, or multiple testing. With roughly 400 daily observations, autocorrelated returns/sentiment can materially inflate t-statistics; the reported p-value is not trustworthy as is. Multiple testing alone is less decisive—Bonferroni across the 9 asset–indicator regressions would still leave p < 0.05—but the combination of alignment and unadjusted inference is the real danger. The unvalidated LLM labels (§3.5, §5.2) are a genuine but distinct worry: noisy labels would weaken the semantic interpretation of 'sentiment,' but a robust pre-close association would still be an empirical finding. The load-bearing question is whether the one significant coefficient survives its own acknowledged caveats.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares VADER lexicon-based sentiment with a Gemini 2.5 Flash-Lite-derived multidimensional sentiment framework on r/WallStreetBets posts and comments for GME, AMC, and NOK over November 2020 to May 2022. Daily indicators (VADER average and score-weighted; LLM bullishness, weighted bullishness, and composite LMI) are aligned to next-day returns in OLS regressions, ROC-AUC directional classification, lead/lag correlations, and a rolling 95th-percentile early-warning backtest. The main finding is that VADER indicators are nearly uninformative, while LLM indicators—especially the LMI—show stronger asset-specific statistical structure, most notably for AMC (β = −0.000041, p = 0.000005, R² = 0.0831). The paper is careful to note that results are heterogeneous, partly reactive, and not a stable standalone forecasting signal.","tokens_in":12252,"tokens_out":5759,"duration_ms":54437,"significance":"If the AMC result held under stricter inference, the paper would be a useful, honest contribution to the literature on LLM-based financial text signals: it compares a modern LLM pipeline to a standard lexicon baseline on public data, reports null results for most specifications, and reproduces the full classification prompt in Appendix A. The LMI is defined a priori rather than fitted to the return series, and the paper explicitly discloses its main limitations in §3.7, §3.8, and §5.2. However, the central claim of 'stronger asset-specific statistical structure' rests heavily on one regression row, and that row is vulnerable to two self-acknowledged design choices: the absence of an exchange-time cutoff and the use of conventional OLS standard errors. The unvalidated LLM labels are an additional, distinct concern for the semantic interpretation of the dimensions. As submitted, the evidence is not yet sufficient to support the abstract's comparative claim.","major_comments":[{"comment":"The next-day design is not strictly ex ante. The paper states that no exchange-time cutoff or timezone conversion was applied, so entries posted after the U.S. close on day t remain in the calendar-day aggregate X_t, even though they occur after the start of the close-to-close interval that defines R_{t+1}. The headline AMC LMI result may therefore reflect discourse that is contemporaneous with the return window, not predictive of it. The negative coefficient and the lead/lag reversal in Figure 1 are consistent with same-day discourse reacting to price action. Please re-estimate with a pre-close cutoff (e.g., 15:59 ET) or, if timestamps do not allow this, explicitly reframe the result as an association that includes post-close information.","section":"§3.7, Eq. (9)"},{"comment":"The reported p-value for the AMC LMI regression is based on conventional OLS standard errors, with no adjustment for heteroskedasticity, serial dependence, or multiple testing. With roughly 400 daily observations for AMC, autocorrelation in returns and sentiment can materially inflate the t-statistic. Although the p = 0.000005 would survive a Bonferroni correction across the nine asset–indicator regressions, the combination of inference misspecification and the alignment problem makes the single significant row unreliable as the sole support for 'stronger asset-specific statistical structure.' Please report Newey–West or block-bootstrap p-values, and preferably also a multiple-testing adjustment across the full set of specifications.","section":"§3.8, Table 3"},{"comment":"The four LLM outputs—sentiment, bullishness, sarcasm, and relevance—are treated as valid semantic measurements, but the paper explicitly states that there is no manually validated ground-truth evaluation and that the manual checks were exploratory. If the model's outputs are noisy or systematically biased by linguistic style, the comparison with VADER is a comparison of two noisy proxies, and the claim that LLM representations are 'richer' becomes harder to interpret. This is acknowledged as a limitation, but it is load-bearing for the semantic dimension of the central claim. A small human-annotated validation set, or at least a transparent error analysis on a random sample, would considerably strengthen the paper.","section":"§3.5 and §5.2"},{"comment":"The early-warning results are based on very few true positives: AMC has 8 TP out of 33 events, GME 3 out of 31, and NOK 12 out of 23. The precision differences (e.g., AMC 20.0% vs. 8.33% base rate) are not accompanied by confidence intervals, permutation tests, or any uncertainty quantification. With this sample size, the conclusion that the LMI 'can contain tail-risk-related information in selected cases' is fragile. At minimum, report exact binomial confidence intervals for precision and recall, or a permutation test of the signal-return association.","section":"§4.4, Table 4"}],"minor_comments":[{"comment":"The AMC LMI row reports p < 0.0001 in the table while the text gives p = 0.000005. Please make the table and text consistent.","section":"Table 3"},{"comment":"The LMI clips non-positive Reddit scores at zero, while the weighted VADER and bullishness indices retain negative scores for entries with score < −1. This asymmetry is acknowledged in §5.2, but it would be helpful to restate it directly at the point of the LMI definition.","section":"§3.6, Eq. (5)"},{"comment":"The lead/lag correlations are described as 'more pronounced' for the LMI, but no confidence bands or significance tests are provided. Adding a brief note that these are descriptive would help prevent over-interpretation.","section":"§4.3, Figure 1"},{"comment":"Consider adding a data availability statement with the exact versions of the Kaggle and Figshare datasets used, since reproducibility depends on these external sources.","section":"§3.2"},{"comment":"The limitations paragraph on construct non-equivalence is good, but it could be moved earlier or referenced near the definition of the LMI so readers evaluate the indicators with this caveat in mind.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent and the research question is reasonable, but the central comparative claim depends on a single significant regression whose validity is threatened by the acknowledged lack of a pre-close cutoff and by unadjusted OLS inference. I recommend inviting a revision in which the authors re-estimate after applying an exchange-time cutoff and robust/multiple-testing corrections. If the AMC LMI result does not survive, the paper should be substantially reframed as a descriptive comparison with a null or mixed finding. I do not see this as a reject: the limitations are disclosed, the data are public, and the prompt is provided, so the analysis is reproducible and fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a straightforward empirical comparison: VADER versus a zero-shot Gemini 2.5 Flash-Lite prompt on r/WallStreetBets discourse for GME, AMC, and NOK, with a composite LMI and a rolling-quantile tail-risk backtest. What's new is the application, not the concept. The multidimensional prompt (sentiment, bullishness, sarcasm, relevance) and the hand-built LMI are legitimate concrete pieces, and the authors set the LMI weights, lookback, and quantile thresholds a priori rather than fitting them to returns. Credit also for stating the limitations openly — §3.7 and §3.8 admit the alignment problem and the unadjusted inference, and §5.2 lists the missing ground-truth labels. That is better behavior than most.\n\nThe central negative finding is probably robust: VADER polarity has essentially no relationship to next-day meme-stock returns, and even the LLM indicators don't produce stable forecasts. That is worth knowing. But the abstract's claim of \"stronger asset-specific statistical structure\" rests on exactly one significant coefficient: AMC LMI, β = −0.000041, p = 0.000005, R² = 0.083. The stress-test note is right. No exchange-time cutoff means entries posted after the U.S. close on day t are still in the day-t indicator, so the regression of R_{t+1} on X_t is not strictly ex ante — part of the close-to-close interval is leaking into the predictor. Figure 1's contemporaneous AMC correlation followed by next-lag reversal suggests the negative β may be mean reversion to same-day discourse, not prediction. And with roughly 400 daily observations, conventional OLS p-values without HAC adjustment or any multiple-testing correction are not trustworthy; the reported p is far below 0.05, so Bonferroni across the nine asset–indicator regressions wouldn't kill it, but the alignment issue alone is enough to doubt it. The unvalidated LLM labels are a real but secondary worry: noisy semantic measurements would weaken the interpretation of \"sentiment,\" but a robust pre-close association would still be an empirical fact.\n\nSo the letter is conditional. The negative result stands; the positive one is fragile and needs a proper cutoff, HAC-robust errors, robustness scans on the 30-day/95% thresholds, and ideally code or processed data release. None of that is fatal — it's addressable. The paper deserves a serious referee: the questions are sensible, the authors know what they didn't do, and a revision could make the claims match the evidence. I would not cite it as it stands, but I'd bring it to a reading group to talk about why ex-ante alignment and robust inference matter so much in this literature. Send it to peer review, yes, but expect heavy revision.","headline":"An honest, clearly-scoped comparison, but the one significant positive result (AMC LMI) may not survive a proper pre-close cutoff and robust errors; the negative finding—lexicon and even LLM sentiment don't forecast meme stocks—is more solid.","tokens_in":856,"tokens_out":1011,"would_cite":false,"duration_ms":23355,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A composite, LLM-derived discourse index extracts asset-specific market signal from meme-stock Reddit chatter, but the gain is uneven across tickers and falls short of stable forecasting.","keywords":["sentiment analysis","large language models","VADER","meme stocks","Reddit","WallStreetBets","tail risk","return prediction"],"falsifier":"Human-annotate a random sample of the Reddit entries used here, a few hundred per asset, on the same four scales; if the LLM's ratings show near-zero or systematically biased agreement with raters—say, on sarcasm or hype posts—then the claimed market signal is measuring linguistic style rather than sentiment. Alternatively, rerun the AMC LMI regression on dates shuffled to break the alignment; if the p≈0.000005 result survives data permutation, the OLS finding is an alignment artifact.","tokens_in":11699,"feed_emoji":"📈","tokens_out":4708,"duration_ms":41017,"temperature":0.7,"pith_summary":"The paper sets out to test whether a large language model can extract more market-relevant signal from Reddit meme-stock chatter than a standard word-list sentiment baseline. Its claim is that multidimensional LLM indicators—especially a composite \"Language-based Market Index\" that combines emotional tone, trading intent, relevance, sarcasm penalties, and social visibility—capture asset-specific statistical structure that simple polarity scores miss. The main evidence is one strong association: for AMC, a rise in this composite index is linked to a next-day return drop with p≈0.000005 and R²≈8%, and in the tail-risk backtest its warnings hit 20% precision against an 8.33% base tail rate. The paper is careful to say this does not amount to stable forecasting: across GME, AMC, and NOK the relationships are heterogeneous, and directional classification stays near chance. The contribution is therefore a demonstration of where richer linguistic representation helps, paired with a warning that it does not automatically translate into reliable prediction.","feed_headline":"LLM sentiment beats word lists at spotting meme-stock risk","feed_subtitle":"A composite Reddit-chatter score flagged next-day AMC drops, but gains were asset-specific and not forecast-ready.","key_machinery":"The carrying object is the LMI, a composite entry-level score defined as (sentiment + bullishness) × relevance × (1 − sarcasm) × ln(1 + max(Reddit score, 0)), then averaged per asset per day. Sentiment and bullishness add when tone and intent agree; relevance filters off-topic chatter; the sarcasm term down-weights ironic posts; and the log of social visibility (with negative scores clipped to zero) limits the influence of a few viral entries. The paper uses this composite to test whether a multidimensional discourse summary has information about next-day returns and extreme upper-tail return days that a one-dimensional polarity score does not.","core_discovery":"The paper argues that a zero-shot LLM, prompted to rate each Reddit entry on sentiment, bullishness, sarcasm, and relevance, yields daily discourse indicators that are structurally richer and statistically more asset-specific than the one-dimensional VADER compound score. Concretely, the composite LMI—defined as the daily average of (sentiment + bullishness) × relevance × (1 − sarcasm) × ln(1 + max(Reddit score, 0))—shows a statistically significant negative association with next-day AMC returns in the main regression specification, and in the tail-risk backtest it warns of upper-tail return days at twice the base rate for AMC. For GME and NOK, however, the signal is weaker and inconsistent,","pith_inferences":["Editorial inference: If the AMC regression is read causally—which the paper itself warns against—the negative sign suggests that unusually intense, relevant, non-sarcastic bullish chatter tends to precede a next-day pullback, consistent with a retail buying climax; a direct test would split the LMI into sentiment versus bullishness components.","Editorial inference: The high recall but low precision for NOK suggests the LMI may be acting as a volatility-regime indicator rather than a direction predictor; one could test whether the signal merely proxies for trading volume or return dispersion.","Editorial inference: The paper's stated absence of manually annotated ground truth implies a targeted audit: if human raters disagree with LLM labels on sarcastic or hype-driven posts, the 'richer representation' claim may reflect surface linguistic features, making the AMC result a measurement artifact rather than a market signal.","Editorial inference: The zero-weighting of non-positive Reddit scores in the LMI, but not in the weighted indices, creates an acknowledged design asymmetry; re-running the AMC regression with continuous score weighting would reveal how sensitive the headline result is to that choice."],"forward_implications":["VADER-style polarity is nearly useless for next-day meme-stock return direction; the paper's regressions and AUCs stay near zero and near random.","Multidimensional LLM indicators can capture asset-specific discourse-market structure that polarity misses, shown most clearly for AMC.","The LMI can serve as an exploratory early-warning signal for extreme positive-return days, with precision above base rate in one of the three assets and higher recall in another.","Because effects vary across assets, a single global sentiment feature is unlikely to be reliable; asset-specific calibration is needed.","The framework creates a testable template: applying the same pipeline to longer windows or other tickers would show whether the AMC effect is stable or idiosyncratic."],"fun_headline_variants":["LLM sentiment outperforms VADER in meme-stock tail-risk signal","Composite LLM score doubles warning rate for AMC tail-risk days","AI sentiment richer but asset-specific for meme-stock crash prediction","LLM Reddit scoring beats lexicon for AMC next-day drop forecast","Meme-stock risk: LLM signals stronger, but gains vary by asset"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The LLM's four numbers (sentiment, bullishness, sarcasm, relevance) are treated as valid measurements of what they claim to measure, even though the study reports no manually annotated ground-truth evaluation.","fun_headline_variants_meta":{"raw":{"variants":["LLM sentiment outperforms VADER in meme-stock tail-risk signal","Composite LLM score doubles warning rate for AMC tail-risk days","AI sentiment richer but asset-specific for meme-stock crash prediction","LLM Reddit scoring beats lexicon for AMC next-day drop forecast","Meme-stock risk: LLM signals stronger, but gains vary by asset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1412,"prompt_tokens":741,"completion_tokens":671,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":575}},"tokens_in":485,"tokens_out":671,"duration_ms":6362,"temperature":1.0,"reasoning_tokens":575,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:04:12.516291+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Human-annotate a random sample of the Reddit entries used here, a few hundred per asset, on the same four scales; if the LLM's ratings show near-zero or systematically biased agreement with raters—say, on sarcasm or hype posts—then the claimed market signal is measuring linguistic style rather than sentiment. Alternatively, rerun the AMC LMI regression on dates shuffled to break the alignment; if the p≈0.000005 result survives data permutation, the OLS finding is an alignment artifact.","supporting_citations":[],"review_version":1}