{"id":"44624d21-18cd-469c-9524-3831e1b46fac","arxiv_id":"2608.09089","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Tail heat signals crash risk when the firm's disclosed business network is losing relationships, and potential winners when the network is intact or growing, but only in densely disclosing sectors.","lead":"A new paper uses LLM-extracted networks from corporate filings to separate two kinds of stock return tails: those that predict crashes and those that may hide future winners. The crash-side signal is supported in a small pilot, while the winner side awaits a pre-registered confirmatory test.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The death-side discriminator depends on un-audited LLM extraction of edge deaths; the pending cutoff-matched extractor audit is the load-bearing test.","rationale":"The reader's weakest assumption is exactly the validity of the disclosure-measured network state as a point-in-time representation of economic configuration, and my concern is that specific assumption. The paper is unusually honest about its limitations: it discloses the single-pass extraction, the recall audit failures, the pending cutoff-matched audit, the panel-interior restriction, and the failed pre-registered replication. The central claim, however, is the sign pattern involving D, and D is the least-secure measured variable in the paper. The headline interaction t-statistics (-2.9 monthly, -3.9 firm-vintage, wild-cluster p=0.04) are computed from a D that has not yet survived an audit designed to test exactly whether the measured deaths are real. Without that audit, the possibility remains that the interaction is an artifact of extraction error correlated with tail heat. The failed replication is consistent with this concern as much as with the sparsity-boundary story, and the dense-subset test cannot discriminate between the two. I therefore do not change the reader's conditional verdict: the paper's confirmatory design is the right next step, but the crash-side claim should not be treated as established until the extractor audit is completed and the pilot interaction re-estimated on audited death events. The paper's own pre-committed consequence clause—demoting P1 to estimation-with-interval if the density pre-test pushes the minimum detectable effect above the pilot anchor—shows a healthy willingness to let evidence decide; the same discipline should apply to the measurement audit before any risk-monitoring use is asserted.","tokens_in":14750,"tokens_out":3141,"duration_ms":39738,"concrete_test":"Execute the pending cutoff-matched extractor audit on a stratified sample of the 24-firm pilot corpus (or, if feasibility requires, on the Stage-B0 replication filings) using the archived multi-agent protocol, and re-estimate the Table 2 firm-vintage interaction (collapsed, N=200) using D re-constructed only from death events that survive the audited extraction. Additionally report the death-classification error rate in high-tail-heat firm-vintages versus low-tail-heat firm-vintages. If the interaction coefficient falls outside the original confidence interval, the wild-cluster p rises above 0.10, or the error rate is materially higher during hot-tail periods, the pilot crash-side claim is not robust to measurement error and the confirmatory P1 gate should be preceded by a positive audit outcome rather than assumed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central sign pattern asserts that tail heat interacted with death mass D predicts negative forward abnormal returns, while tail heat with birth mass does not. Every version of this claim depends on D being a faithful, point-in-time measure of actual economic relationship dissolution. That measurement is not yet established. Section 4.2 and Table 1 disclose that the pilot used single-pass extraction, that the cutoff-matched extractor audit is pending, and that the recall audit (Section 6(i)) found 2 of 12 zero-edge filings were extraction failures. The recall audit checks whether absent edges are truly absent; it does not validate the harder distinction that the interaction needs: whether an edge that disappears between vintages is a real death or an extraction false negative. If extraction false negatives are more likely in periods when a firm is under stress—when 10-K language is restructured, counterparty names change, or the LLM's span-grounded prompt fails on unusual text—then D will be spuriously elevated exactly when tail heat is high, mechanically generating a negative interaction with no economic content. The panel-interior restriction (Section 7, Limitations) compounds this: deaths of outside-panel counterparties are unmeasured, so D is a partial death count whose coverage varies by firm and sector. The failed replication (Section 5.5) is honestly reported, but its dense-subset test is uninformative (point estimates span zero with standard errors an order of magnitude larger than the pilot effect), so the interpretation that the failure is purely a sparsity boundary rather than a measurement failure is not independently supported. The pilot's robustness battery is helpful but does not address measurement error in D; leave-one-firm-out still relies on the same audited-but-retracted variable. The confirmatory design gates on density, but density does not cure extraction noise.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes separating heavy upper-tail return processes into \"lottery tails\" and \"structural tails\" using an LLM-extracted 10-K disclosure network. It defines a tail-heat signal from the Hill index and decomposes network rewiring into birth, death, and drift. The central empirical claim is a sign pattern: the interaction of tail heat with death mass predicts negative forward abnormal returns (crash side), while tail heat with birth mass or an intact network is where structural winners live. The pilot uses 24 technology firms (2014--2025) and reports a significant negative death-side interaction with descriptive, not confirmatory, status. A pre-registered replication on 50 random S&P 500 firms failed, with 83% of firm-vintages having zero death mass, which the paper interprets as a scope condition: the discriminator is only defined where firms densely document counterparties. The manuscript then presents a fully pre-specified confirmatory design with a density gate, a gatekept primary pair, archived power simulations, and pending audits, and it explicitly labels the alpha side as still unresolved.","tokens_in":15077,"tokens_out":4572,"duration_ms":53860,"significance":"If the mechanism is confirmed, the contribution is substantial: it would condition the MAX anomaly on an observable, text-measured network state and provide a crash-risk flag that return-based screens cannot produce. The paper is unusually careful in several ways: pilot results are labeled descriptive; the pre-registered replication failure is reported verbatim; inference includes wild-cluster bootstrap and leave-one-firm-out checks; and the confirmatory design freezes thresholds, winsorization, and power analyses before estimation. These features are real strengths and should be preserved. However, the current evidence for the central sign pattern rests on one pilot coefficient in a 24-firm convenience panel, the death-mass measurement is not yet validated for the specific quantity the interaction needs, and the abstract's risk-monitoring claim goes beyond what the pilot can support. The paper is best read as a credible, well-scoped research design with promising descriptive results, not as an established empirical finding.","major_comments":[{"comment":"The death-side interaction is load-bearing, but the death mass D is measured with single-pass extraction and the cutoff-matched extractor audit is still pending. The executed recall audit (2 of 12 zero-edge filings gained edges under re-extraction) validates absent edges, not edge disappearances. Since D is computed from edge-weight decreases between vintages, an extraction false negative at the later vintage is algebraically indistinguishable from a real death. If extraction quality deteriorates when a firm is under stress, or when filing language changes, D will be spuriously elevated precisely in the periods where the interaction is claimed to matter. This is not a minor measurement concern; it is a plausible mechanical generator of the paper's headline interaction. The confirmatory stage must include an explicit \"death recall\" audit on a sample of edges present in one vintage and absent in the next, and the authors should report whether extraction false negatives are correlated with tail heat or with disclosure-language change. This audit should be completed before any claim of practical risk-monitoring utility is made.","section":"§4.2, Table 1, §6(i)"},{"comment":"The evidential basis for the abstract's claim that the death-side signal is \"immediately useful as a risk-monitoring danger flag\" is one interaction coefficient in a 24-firm pilot (wild-cluster p = 0.04), with the leave-one-firm-out analysis showing that the coefficient roughly halves when NVIDIA is dropped. The pre-registered replication failed under its frozen criterion, and the dense-subset check inside the replication is explicitly uninformative (coefficient -0.001 with SE 0.057). The paper is honest about these facts, but Section 7 nevertheless asserts that the death-side flag \"already meets the risk-pool standard on the pilot record.\" Given the failed replication and the pending extractor audit, this practical-readiness claim is not proportionate to the evidence. The revised manuscript should either weaken the risk-pool claim to a hypothesis-generation statement, or support it with an out-of-sample validation that does not rely on the pilot firms.","section":"§5.3, §5.5, §7"},{"comment":"The confirmatory P1 endpoint is powered using the pilot's collapsed coefficient (-0.028) as the effect-size anchor, and the archived simulation claims near-certain detection at magnitudes down to 0.010. But the pilot anchor is estimated on 200 firm-vintages with pooled standardizations, overlapping horizons, and a graph state that varies only annually. The design should specify how the firm-vintage collapsed regression will handle vintage overlap, staleness, and the panel-interior restriction acknowledged in Section 7, where deaths of counterparties outside the universe are unmeasured. Without such specifications, P1 may reproduce a partial-death-mass artifact rather than test the economic mechanism. I ask the authors to add explicit measurement-error language to the P1 specification and to the archived power simulation.","section":"§6(d), §6(g)"},{"comment":"The two-component tail decomposition in Definition 3.1 is definitional rather than identified from returns, and Proposition 3.2 is explicitly described as a pricing argument rather than a theorem. This is acceptable for hypothesis generation, but the paper should be more careful in the summary and abstract to distinguish the definitional decomposition from the empirical discriminability claim. As written, the abstract presents the decomposition as an established fact (\"decomposes exactly into edge birth, death, and drift\") when the exactness is only about the accounting identity; the economic claim that network state discriminates lottery from structural tails remains to be established.","section":"§3.2, Remark 3.1"}],"minor_comments":[{"comment":"The acronym LSG is used in the figure and caption but is not defined in the text; please define it at first use, and clarify whether it denotes total rewiring, birth mass, or a normalized graph-change statistic.","section":"Figure 3"},{"comment":"The collapsed regression omits momentum and volatility coefficients for space, and the MAX coefficient changes sign on collapsing. Since the collapsed regression is the stated honest unit of identification, please include the full coefficient table in an appendix or online supplement.","section":"§5.3, Table 2"},{"comment":"The pilot reports 73% of firm-vintages with D > 0 and the replication reports 83% with zero death mass; the two statements are consistent but easy to misread. Please state the pilot's zero-death share directly and use identical denominators in both sections.","section":"§5.3, §5.5"},{"comment":"The companion protocol and its archived deviations are cited as \"companion v0.5\" and \"arXiv:2607.15640,\" but no direct link to the archived hashes or to the deviation log is provided. Please include an availability statement with the archive location for the frozen prompts, the SHA-256 hash, and the v0.5 deviation log.","section":"§4.2"},{"comment":"Figure 1 would benefit from explicit shading of the 2018 and 2022 drawdown windows and a marked scale for the panel Q80 series, since the text refers to specific threshold-crossing months that are hard to read from the figure.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is highly transparent and the confirmatory design is unusually disciplined, but the central empirical claim is not yet established: the pilot is descriptive, the pre-registered replication failed, and the most load-bearing measurement audit (death-event extraction) is pending. The abstract and Section 7's practical-use claim for the danger flag should be softened until that audit is completed. I would also encourage the editor to require independent validation of the companion measurement pipeline, since the current paper relies heavily on a self-cited, not-yet-independent companion article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the rare empirical asset-pricing paper that reports a failed pre-registered replication up front and uses it to bound its claims. Second, the central sign pattern—tail heat times network death predicting negative returns—lives or dies on the quality of LLM extraction of edge deaths, and that audit is still pending.\n\nWhat's genuinely new: the two-kind tail decomposition (lottery vs. structural) and the use of disclosure-network rewiring (birth/death/drift) to condition the MAX anomaly. The measurement pipeline comes from the companion paper, but applying it to separate crash risk from winner tails is a real idea. The pilot battery is serious: wild-cluster bootstrap, two-way clustering, leave-one-out, and they honestly report that dropping NVIDIA halves the coefficient. The pre-specified confirmatory design with a density gate and archived power simulations is exactly the right next step, and the failed replication is informative in defining a scope condition—outside dense disclosure ecosystems the graph nearly vanishes.\n\nThe soft spots are real. The stress-test concern holds up on reading: the recall audit checked whether zero-edge filings were extraction failures, but it did not validate the harder distinction—whether an edge disappearing between vintages is a true death or a false negative. If extraction misses edges more often during stress (when 10-K language gets restructured), D could be spuriously elevated exactly when tail heat is high, mechanically producing the negative interaction. The panel-interior restriction also makes D a partial count, and the failed replication's dense-subset test is uninformative because standard errors are an order of magnitude wider than the pilot effect. So the pilot result is one coefficient in a 24-firm panel, with one firm carrying a disproportionate share, and the measurement layer isn't yet validated.\n\nThe paper is honest about all of this—limitations are disclosed, pilot results are labeled descriptive, and the confirmatory design is gated. It's a serious paper for people working on lottery stocks, tail risk, or text-based asset pricing, and it deserves a serious referee. I'd send it to peer review, with the expectation that referees should push hard on measurement validity, not just the regression tables. The central claim is plausible but not established; the pending cutoff-matched extractor audit is the test that will decide it.","headline":"Honest, well-structured paper with a plausible conditional tail signal, but the pending extractor audit is the load-bearing test and the alpha side is not yet supported.","tokens_in":15658,"tokens_out":2081,"would_cite":false,"duration_ms":22522,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G32","91G70","62P05","62M10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The same hot return tail can signal a crash or a structural winner; this paper separates the two using the firm's disclosure-measured network, with tail heat plus network death predicting negative forward returns and tail heat with an…","keywords":["extreme value theory","lottery stocks","MAX anomaly","crash risk","economic networks","corporate disclosures","large language models","tail risk"],"falsifier":"Run the frozen confirmatory P1 endpoint on the ecosystem-coherent universe after the density gate passes: if the test-period coefficient on z(ξ̂+) × z(D) in the firm-vintage collapsed regression is non-negative or has a one-sided wild-cluster p ≥ 0.05, the paper's central separation claim is refuted. A cheaper observation is the density gate itself: if a seeded 200-filing pre-test shows fewer than 30% of firm-vintages with D > 0, the universe rule fails and no confirmatory claim can be made.","tokens_in":14530,"feed_emoji":"📉","tokens_out":6282,"duration_ms":57153,"temperature":0.7,"pith_summary":"The paper asks whether a stock's heavy upper return tail is a lottery tail—transient jump risk that investors overpay for—or a structural tail, the statistical shadow of an economic reconfiguration that can precede extreme winners. Returns alone cannot tell the two apart, so the paper proposes the firm's disclosure-measured network, a directed graph extracted from 10-K filings by an auditable LLM pipeline, as the discriminator. Its central sign pattern is that tail heat interacted with network death predicts negative forward abnormal returns, while tail heat with an intact or forming network is where structural tails and historical winners live. The pilot's crash side is statistically supported; the alpha side is directionally positive but not significant and awaits a pre-registered confirmatory test. A failed replication shows the discriminator only exists where firms document counterparties densely, which bounds all claims.","feed_headline":"Hot tails are crash flags when disclosure networks die","feed_subtitle":"Same return tail can precede winners or drawdowns; network rewiring from 10-K filings tells them apart.","key_machinery":"The central object is the disclosure-measured economic network: a directed, weighted, point-in-time graph of documented counterparty relationships extracted from 10-K filings by a frozen LLM prompt, with every edge grounded in quoted text. Its vintage-to-vintage rewiring decomposes exactly into birth mass B, death mass D, and continuing drift C via B + D + C = Σ|w_e,t − w_e,t−1| over directed pairs. The paper's argument is carried by the interaction between tail heat (the rolling Hill estimator ξ̂+ on the top 5% of trailing two-year daily excess returns) and death mass D; the sign of that interaction is the discriminator that separates crash tails from structural winner tails.","core_discovery":"The paper claims that heavy tails come in two kinds: lottery tails, which earn the MAX discount, and structural tails, which are persistent repricing processes. The same observable tail heat sits on top of both, and only the rewiring of the firm's disclosure-measured network separates them. Concretely, in the pilot the interaction of the rolling Hill tail index with death mass—the sum of lapsed edge weights between filing vintages—predicts forward six-month abnormal returns negatively (monthly t ≈ -2.9; collapsed to firm-vintages t ≈ -3.9; wild-cluster p = 0.04), while the interaction with birth mass is statistically indistinguishable from zero. The paper presents this as support for the crash side and as a danger flag today, while the winner side—heat with births, without deaths—is held to the stricter confirmatory standard.","pith_inferences":["As an extension of the paper's mechanism, the same measurement could be applied to more frequent disclosures such as 8-K filings or earnings-call transcripts to sharpen the timing of the danger flag, which the annual filing grid currently blurs.","If the slow-diffusion mechanism is right, the alpha from reading filings before the category error corrects should shrink over time as machine-readable extraction becomes widespread; this could be tested by estimating whether the interaction's predictive power declines across later vintages.","The density-boundary finding suggests the discriminator may generalize beyond technology into other hub-centered ecosystems such as aerospace-defense primes and telecom infrastructure, where the replication map shows unexpected pockets of dense disclosure."],"forward_implications":["If the confirmatory endpoint P1 confirms, the MAX anomaly is revealed as a pooled price on two objects: correctly applied to lottery tails, misapplied to structural tails.","The death-side interaction is immediately usable as a risk-monitoring danger flag even before confirmation, because a missed crash costs more than a false alarm.","If the confirmatory endpoint P2 confirms, a screen of tail heat plus birth-dominated rewiring would identify forward extreme winners at higher precision than MAX, momentum, or volatility screens.","Both effects should concentrate where tail heat is high and decay as configuration information becomes cheap to observe.","All claims are bounded by the scope condition: the discriminator exists only where disclosure is dense, such as coherent supply-chain and platform ecosystems."],"supporting_citations":[{"why":"Supplies the MAX anomaly—stocks with the highest recent maximum daily returns underperform—which the paper conditions rather than disputes.","marker":"Bali et al., 2011"},{"why":"Provides the probability-weighting mechanism by which skewness-preferring investors overpay for lottery-like payoffs.","marker":"Barberis and Huang, 2008"},{"why":"Establishes that crashes are forecastable from trading-based variables, the baseline the crash-side claim extends.","marker":"Chen et al., 2001"},{"why":"Links crash risk to opaque financial reporting, which the paper sharpens into documented disintegration under a hot tail.","marker":"Hutton et al., 2009"},{"why":"Shows filing-text information diffuses slowly into prices, the friction that generates both the alpha and the crash-side mispricing.","marker":"Cohen et al., 2020"},{"why":"Provides the companion pipeline that extracts span-grounded directed economic graphs from 10-K filings, which the paper inherits as its measurement layer.","marker":"Yang and Zhang, 2026"},{"why":"Supplies the prior text-based firm-network construction from 10-K product descriptions, anchoring the network approach.","marker":"Hoberg and Phillips, 2016"},{"why":"Demonstrates predictable returns through economic customer links, motivating why counterparty structure should carry pricing information.","marker":"Cohen and Frazzini, 2008"}],"fun_headline_variants":["Network death flips tail heat into crash signal","Hot tails + dying disclosure networks = crash flag","Tail heat is a crash warning when disclosure networks die","When disclosure networks die, hot tails predict drawdowns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the disclosure-measured network is a valid, point-in-time representation of the firm's actual economic configuration, so edge births and deaths correctly capture real relationship changes. That requires accurate LLM extraction and complete counterparty coverage; the paper's own audits show 2 of 12 zero-edge filings were extraction failures, a cutoff-matched extractor audit is still pending, and the panel-interior restriction attenuates death detection for counterparts outside the universe.","fun_headline_variants_meta":{"raw":{"variants":["Network death flips tail heat into crash signal","Hot tails + dying disclosure networks = crash flag","Tail heat is a crash warning when disclosure networks die","When disclosure networks die, hot tails predict drawdowns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000236,"raw_usage":{"total_tokens":1580,"prompt_tokens":1095,"completion_tokens":485,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":711,"completion_tokens_details":{"reasoning_tokens":423}},"tokens_in":711,"tokens_out":485,"duration_ms":5140,"temperature":1.0,"reasoning_tokens":423,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:45:09.219239+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the frozen confirmatory P1 endpoint on the ecosystem-coherent universe after the density gate passes: if the test-period coefficient on z(ξ̂+) × z(D) in the firm-vintage collapsed regression is non-negative or has a one-sided wild-cluster p ≥ 0.05, the paper's central separation claim is refuted. A cheaper observation is the density gate itself: if a seeded 200-filing pre-test shows fewer than 30% of firm-vintages with D > 0, the universe rule fails and no confirmatory claim can be made.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MAX anomaly—stocks with the highest recent maximum daily returns underperform—which the paper conditions rather than disputes."},{"cited_title":"Huang (2008)","cited_arxiv_id":null,"evidence_quote":"Provides the probability-weighting mechanism by which skewness-preferring investors overpay for lottery-like payoffs."},{"cited_title":"Hong, and J","cited_arxiv_id":null,"evidence_quote":"Establishes that crashes are forecastable from trading-based variables, the baseline the crash-side claim extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Links crash risk to opaque financial reporting, which the paper sharpens into documented disintegration under a hot tail."},{"cited_title":"Malloy, and Q","cited_arxiv_id":null,"evidence_quote":"Shows filing-text information diffuses slowly into prices, the friction that generates both the alpha and the crash-side mispricing."},{"cited_title":"LLM Latent Edge Measurement: Point-in-Time Economic Graphs for Quantitative Investing from Corporate Disclosures","cited_arxiv_id":"2607.15640","evidence_quote":"Provides the companion pipeline that extracts span-grounded directed economic graphs from 10-K filings, which the paper inherits as its measurement layer."},{"cited_title":"Phillips (2016)","cited_arxiv_id":null,"evidence_quote":"Supplies the prior text-based firm-network construction from 10-K product descriptions, anchoring the network approach."},{"cited_title":"Frazzini (2008)","cited_arxiv_id":null,"evidence_quote":"Demonstrates predictable returns through economic customer links, motivating why counterparty structure should carry pricing information."}],"review_version":1}