{"id":"d0a7d3fa-0499-4672-b792-3f547153b1a1","arxiv_id":"2601.20285","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Using LLM-extracted newspaper records, the authors identify 3,421 U.S. bank runs from 1863-1934 and find that runs lead to failure and severe local economic damage mainly when banks have weak fundamentals.","lead":"This paper builds a new database of 3,421 U.S. bank runs from 1863 to 1934 by applying large language models to historical newspapers, and uses it to show that bank runs only rarely cause failure or severe economic damage unless the bank already has weak fundamentals. A smart generalist should read it because it directly tests whether self-fulfilling panics can topple healthy banks, a question at the center of financial crisis theory and policy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hidden insolvency among banks classified as 'strong' threatens the claim that weak fundamentals are necessary for run-induced failure; test with OCC examiner data.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: observable annual balance-sheet fundamentals may not capture hidden insolvency. My stress test agrees and proposes a concrete falsification using existing OCC examiner and recovery-rate data. The paper is otherwise careful: the run database is validated against official failure counts, deposit outflows, and independent human audits, and the central descriptive patterns are robust to restricting to large banks. The run-count inconsistency between the abstract (3,984) and the text (3,421) is a data-reporting issue that should be fixed but does not threaten the substantive conclusion. The out-of-sample AUC concern is secondary because the paper's claims about predictability are descriptive and the in-sample AUCs are not used to argue for a structural model. The hidden-insolvency concern, however, directly bears on whether the observed non-failure of 'strong' banks is evidence that runs do not cause solvent banks to fail, or merely that the proxy for solvency is mismeasured. The proposed test—reclassifying top-tercile failures using examiner assessments and recovery rates—can settle this with data the authors already possess. Until that check is run, the conditional verdict is appropriate: the paper's headline claim should be interpreted as conditional on observable fundamentals accurately reflecting solvency.","tokens_in":52655,"tokens_out":3871,"duration_ms":40504,"concrete_test":"For every national bank failure in the sample that was preceded by a run and classified in the top tercile (and separately the top decile) of Fundamentals, use the OCC receiver records and examiner assessments digitized in Correia, Luck and Verner (2025) to classify the bank as 'actually insolvent' if the examiner's asset-quality rating or the depositor recovery rate implies losses exceeding the bank's pre-run capital. Re-estimate Figure 5 and Table 5 after reclassifying those banks into the weak-fundamentals category. If the failure probability conditional on a run for the genuinely solvent top decile remains close to zero and statistically insignificant, the concern is resolved. If it rises materially (e.g., above 10%), the paper's claim that weak fundamentals are necessary for run-induced failure is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that weak bank fundamentals are necessary for a run to be associated with bank failure and severe real effects—depends on the fundamentals measure derived from annual balance-sheet data accurately capturing true solvency at the time of the run. The authors acknowledge in Section 6.1 that 'publicly available financial statements cannot capture episodes where a bank was insolvent but had not yet recognized losses.' If such hidden insolvency is common among banks classified in the top decile of Fundamentals, then the headline evidence—a 4% failure probability for top-decile runs that is not statistically significant—would not show that runs cannot topple solvent banks; it would show only that annually observed balance sheets miss some insolvent banks. This measurement concern is structurally distinct from the run-count discrepancy: it directly conditions the inference that weak fundamentals are necessary for failure. The robustness check restricting to large banks (Table 5, columns 5–6) addresses underreporting of runs but does not address misclassification of solvency. Likewise, the Appendix A.1 event studies show failing banks have weak observables on average, but they do not examine whether the few top-decile banks that fail after runs were actually secretly insolvent. Absent such a check, the paper's strongest claim is only as strong as the annual call-report solvency proxy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs a new dataset of 3,421 bank runs in U.S. newspapers from 1863 to 1934 using large language models, links these runs to national-bank balance sheets and official OCC receivership records, and studies the determinants and consequences of runs. The authors find that runs occur disproportionately in weak banks but also occur in strong banks, especially after negative aggregate or local news; that conditional on a run, failure is far more likely in weak banks, with a top-decile failure probability of about 4 percent that is not statistically significant; that strong banks typically survive runs through suspension, signaling, equity injections, and interbank support; and that local declines in deposits, lending, and manufacturing are concentrated in runs on weak banks or in runs with failure. The paper interprets these patterns as evidence that poor fundamentals are necessary for runs to translate into failure and severe real effects, tempering the view that small shocks can trigger costly self-fulfilling panics.","tokens_in":52988,"tokens_out":7403,"duration_ms":68635,"significance":"If the results hold, this is a major empirical contribution. The new run database is validated against official OCC failure records, narrative crisis chronologies, and an independent human audit with a 95 percent match rate; the authors are transparent about underreporting of small-bank runs and publicly document the episodes. The finding that most runs do not result in failure, and that failure pass-through is strongly graded in observable fundamentals, directly informs the debate between fundamentals-based and multiple-equilibrium views of banking panics. The paper also provides novel descriptive evidence on how strong banks survive runs and on the local real effects of runs with and without failure. The main risks are that the headline 'necessity' claim is stronger than the statistical tests support, and that the observational local projections cannot fully rule out confounding by local economic trends.","major_comments":[{"comment":"The central claim that weak fundamentals are necessary for a run to be associated with failure is identified off the annual call-report fundamentals index. The paper itself states in Section 6.1 that 'publicly available financial statements cannot capture episodes where a bank was insolvent but had not yet recognized losses'; this is exactly the alternative reading of the 4 percent top-decile failure probability: those runs may have hit banks that were already insolvent on an economic basis but looked strong on observable book capital. The large-bank robustness (Table 5, columns 5-6) addresses underreporting of runs, not mismeasurement of solvency, and the Appendix A.1 event studies show failing banks have weak observables on average but do not examine the few top-decile banks that fail after runs. Please use OCC examiner assessments, as in Correia, Luck and Verner (2025), for the strong-fundamentals banks that fail following runs, or otherwise bound the extent of hidden insolvency, and state explicitly how the 'necessary' conclusion is affected.","section":"Section 6.1, Figure 5, Table 5"},{"comment":"The claim that weak bank fundamentals are 'necessary' for runs to cause failure is not supported by the reported test. The top-decile conditional failure probability is 4 percent and statistically insignificant; failing to reject zero is not evidence that the probability is exactly zero. The abstract and conclusion state necessity, but the evidence can at most support 'runs on the healthiest banks are rarely followed by failure.' Please report the confidence interval for the top-decile estimate, use language consistent with the statistical precision, or conduct an explicit equivalence or bounding exercise that justifies the necessity wording.","section":"Section 6.1, Figure 5, and Section 8"},{"comment":"The local projections that distinguish runs on weak versus strong banks are observational, and the paper's own Table 7 shows that runs are more likely after local business failures, so cities experiencing weak-bank runs may be on differentially declining trends. The non-fundamental-run design (Section 6.3, based on 44 runs) is a useful step, but the city-level non-fundamental-run indicators are extremely sparse, so the claim that pure liquidity runs have no significant local effects may be underpowered. Please add a formal discussion of pre-trends, alternative control groups, or sensitivity bounds, and temper the causal language in Section 8 accordingly.","section":"Section 7.2, Equations (11)-(16), Figures 9-10"}],"minor_comments":[{"comment":"The abstract in the article header reports 3,984 runs, while the full-text abstract and Section 4 report 3,421 runs; please reconcile the two numbers.","section":"Abstract (arXiv metadata)"},{"comment":"The suspension total in Table 1 Panel A (13,358) does not reconcile with the episode-type counts implied by Figure 1 (run only 1,325, suspension only 2,158, run-suspension-reopening 581, suspension-failure 8,815, run-suspension-failure 1,515, which sum to 13,069 suspensions). Please check the counts or clarify the discrepancy.","section":"Table 1 and Figure 1"},{"comment":"The units of the dependent variable appear inconsistent across tables: Table 3 reports a mean dependent variable of 0.33 (consistent with a rate expressed in percent), while Table 5 reports 0.0085 (a decimal proportion). Please make the units uniform or label them clearly in the table notes.","section":"Tables 3 and 5"},{"comment":"The sentence describing the manufacturing index says 'every month from 1933 to 1935' and later 'monthly for the remainder of 1933 through 1933'; the latter appears to be a typo for 1935. Please also clarify how the forward fill from monthly to weekly frequencies is implemented.","section":"Section 7.2.2"},{"comment":"The AUC values in Table 4 are described as in-sample. Please note explicitly in the text or table notes that these are in-sample fits, since out-of-sample predictive power may be lower, particularly for the run-without-failure panel where the AUC is close to 0.6 in column 1.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"This is a strong empirical paper with a valuable public dataset, and the main results are well validated descriptively. The required revisions are about aligning the strength of the claims with the identification actually achieved: the hidden-insolvency measurement concern directly affects the 'necessity' claim, and the language in the abstract and conclusion overstates what the top-decile test can establish. These are fixable with additional analysis or more careful framing; I do not see a load-bearing error that would require rejection. The citation practice and data documentation are appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the dataset is the real deal: 3,421 newspaper-identified runs linked to bank balance sheets, validated against OCC receivership records, deposit outflows, and a human audit that caught 95% of hand-collected runs. That is a serious contribution, and it lets them show something new: runs without failure are common, strong banks are run but rarely fail, and the local real effects are concentrated in runs on weak banks or runs with failure. Second, the headline claim—that weak fundamentals are necessary for a run to cause failure—is descriptively solid but causally fragile in exactly the way the stress-test note says. The fundamentals index is fit to failure outcomes, the AUCs are in-sample, and the top-decile failure probability of 4% is not statistically distinguishable from zero. That last point could mean strong banks don't fail in runs, or it could mean the annual call-report proxy misses hidden insolvency. The paper acknowledges this in Section 6.1, and the Sedalia example—a 'non-fundamental' run on a bank with an 18% recovery rate—shows the concern is not hypothetical. The robustness checks on large banks address underreporting, not misclassification of solvency. So the 'necessary' language is too strong; 'necessary conditional on observable balance-sheet fundamentals' would be exact.\n\nThe paper does well in several places. The event studies in Appendix A.1 are thoughtful: failing banks look weak on observables five years before, whether or not they fail with a run. The distinction between runs with and without failure, and the local projection results showing small effects for non-fundamental runs, are genuinely informative. The narrative-based classification of bank responses (suspension, clearinghouse, equity injection) is a nice use of the newspaper data.\n\nSoft spots, in order of importance. The run count: abstract says 3,984, main text says 3,421. That needs fixing. The in-sample AUCs: before describing runs as 'predictable,' give cross-validated or out-of-sample numbers. The non-fundamental runs sample is 44 episodes—fine as a robustness check, but the confidence intervals around the bank-level effects are wide. And there is no replication package mentioned anywhere; for a dataset this large, that is a requirement, not a nicety.\n\nThe hidden-insolvency concern is the one that could change the conclusion, and I don't think the current draft fully addresses it. A check using OCC examiner data on asset quality or recovery rates for the few top-decile banks that fail after runs would be the natural fix. Without that, the paper should present the 'weak fundamentals are necessary' claim as a statement about observable fundamentals, not about true solvency.\n\nWho is this for? Financial historians and banking scholars will use the dataset for years. Macro-finance empiricists will read the real effects sections closely. This deserves a serious referee—send it out. I would condition acceptance on the run-count fix, cross-validated AUCs, and a replication package, and I would ask the authors to soften the necessity claim or add the examiner check.","headline":"A genuinely new bank-run dataset that supports the fundamentals-matter-more-than-self-fulfilling-panics view, but the paper's strongest causal claim rests on a solvency proxy that the authors admit is incomplete.","tokens_in":53379,"tokens_out":1490,"would_cite":true,"duration_ms":17957,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Runs rarely kill healthy banks: 3,421 U.S. bank runs from 1863 to 1934 show failure concentrates in weak institutions.","keywords":["bank runs","bank failure","financial crises","large language models","historical newspapers","bank fundamentals","National Banking Era","Great Depression"],"falsifier":"Find a set of banks in the top decile of the fundamentals index that failed after a run and whose receivership records show high recovery rates and examiner-assessed asset quality, meaning they were genuinely solvent when they failed. If such cases are numerous, well above the estimated 4 percent failure probability for that decile, the necessity claim fails; if examiner and recovery data confirm that top-decile failure cases were actually insolvent, the run-induced-failure channel for healthy banks is essentially empty.","tokens_in":52434,"feed_emoji":"🏦","tokens_out":5790,"duration_ms":51494,"temperature":0.7,"pith_summary":"This paper builds a newspaper-based database of 3,421 bank runs in the United States between 1863 and 1934 and uses it to ask when runs matter. It argues that runs are most likely in banks with weak balance sheets, but that runs also hit healthy banks, particularly after bad aggregate or local news. The central claim is that weak bank fundamentals are necessary for a run to end in failure: the weakest decile of banks fails in about 63 percent of runs, while the strongest decile shows no statistically detectable failure after a run. Runs on weak banks translate into large local contractions in deposits, lending, and manufacturing, whereas runs on strong banks and non-fundamental runs have limited consequences. If right, this tempers the view that small shocks routinely trigger severe crises through self-fulfilling runs on solvent banks.","feed_headline":"Runs rarely kill healthy banks in 3,421 U.S. cases","feed_subtitle":"Weak balance sheets decide which runs end in failure and which dent the local economy.","key_machinery":"The central object is a new bank-distress-episodes database built by applying large language models to historical newspaper scans, classifying each episode into run only, run-suspension-reopening, or run with failure, and matching banks to annual national-bank balance sheets. The argument's load-bearing measure is a fundamentals index: the negative of a recursively estimated predicted failure probability from a regression of next-year receivership on balance-sheet ratios such as surplus-to-equity, noncore funding, liquid assets, deposits-to-assets, and asset growth. This index separates banks into weak and strong terciles and deciles; the paper then estimates pass-through regressions of failure on runs interacted with the index, local projections of deposits and loans, and city-level impulse responses for manufacturing. A secondary mechanism is textual classification of bank responses, including accommodation of withdrawals, equity injection, borrowing, partial or full suspension, and clearinghouse examination, which explains why strong banks survive.","core_discovery":"On its own terms, the paper establishes that bank runs in the pre-FDIC United States were common but rarely fatal to sound institutions. Using text classifications of contemporary newspapers, it documents 3,421 runs, about 1,515 of which ended in failure. Conditional on a run, banks in the lowest decile of a fundamentals index fail with probability around 63 percent; for banks in the top decile the failure probability is estimated at about 4 percent and statistically indistinguishable from zero. Runs that newspapers identify as stemming from misinformation or confusion (44 national-bank cases) carry an 11 percent failure probability, and failures after such runs occur only in weak or fraudulent banks. At the city level, runs on weak banks are followed by declines of roughly 40 to 60 percent in local deposits and loans and a 5 percent fall in manufacturing activity, while runs on strong banks and non-fundamental runs show small or no effects. The paper reads this as evidence that poor fundamentals, not liquidity panics themselves, are what turn runs into failures and into real economic damage.","pith_inferences":["An implicit extension is that the 1863-1934 U.S. institutional environment, with no deposit insurance, limited lender of last resort, and branching restrictions, may make runs more frequent than today, but the fundamental-contingent failure pattern could persist in modern runs; testing it on recent uninsured-deposit runs would be a natural extension.","The paper's reliance on observable annual balance sheets means hidden insolvency such as fraud or unrecognized losses could move some 'strong-bank' failures into the fundamentals camp; linking run outcomes to ex-post recovery rates and examiner asset-quality ratings would test this.","The non-fundamental-run sample is small, only 44 national-bank cases, so the 11 percent estimate is imprecise; a larger corpus or cross-country historical newspapers could sharpen it.","City-level reallocation effects, where deposits flee a run bank to other local banks, may explain why non-fundamental runs have no local effect; bank-level and city-level results together suggest redistribution rather than destruction."],"forward_implications":["If the central claim holds, bank runs should be modeled as symptoms or amplifiers of underlying insolvency rather than as exogenous shocks that independently determine bank failure.","Policy attention should focus on detecting and repairing weak bank fundamentals; pure liquidity support is not enough to prevent failure when fundamentals are poor, but suspension and interbank assistance can protect sound banks.","Empirical studies of banking crises that measure distress only through bank failures will miss most runs and will overstate the relationship between runs and failure.","Expect runs without failure to be substantially more common during panics, and local deposit and loan contractions to be concentrated where weak banks are run.","The 11 percent failure probability of non-fundamental runs gives a concrete upper bound on how often pure panic can kill a bank in this historical environment."],"supporting_citations":[{"why":"Supplies the canonical theory of self-fulfilling runs on demandable deposits that the paper tests against empirical patterns.","marker":"Diamond and Dybvig (1983)"},{"why":"Provides the global-games framework predicting that panic runs concentrate in banks with weak fundamentals.","marker":"Morris and Shin (2000)"},{"why":"Formalizes fundamental-based panic runs with a unique equilibrium threshold, giving the prediction that runs should occur in weak banks.","marker":"Goldstein and Pauzner (2005)"},{"why":"Models suspension of convertibility as a costly signal that lets solvent banks survive runs, explaining the paper's survival mechanisms.","marker":"Gorton (1985a)"},{"why":"Establishes the information-based theory where runs follow adverse public signals, supporting the finding that strong banks can be run after bad news.","marker":"Gorton (1988)"},{"why":"Provides the narrative and theoretical background that banking panics are predictable responses to fundamental shocks, which the paper's evidence refines.","marker":"Calomiris and Gorton (1991)"},{"why":"Supplies the crisis chronology and the claim that panics are not necessary for banking crises, used for validating the new run database.","marker":"Baron, Verner and Xiong (2021)"},{"why":"Provides the national-bank balance-sheet and receivership data and the evidence that failed banks were fundamentally insolvent, on which the fundamentals index is built.","marker":"Correia, Luck and Verner (2025)"},{"why":"Offers a regional banking-panic chronology used to validate that newspaper runs line up with known panics.","marker":"Jalil (2015)"},{"why":"Provides the American Stories dataset of segmented historical newspaper articles that the LLM pipeline processes.","marker":"Dell et al. (2023)"}],"fun_headline_variants":["Weak fundamentals, not panic, doom banks in 3,421 runs","Bank runs rarely kill healthy banks in historic data","3,421 runs: strong banks survive, weak ones fail","Fundamentals decide bank run outcomes, study finds","Healthy banks survive runs; weak ones fail, 1863-1934"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that annual balance-sheet data reveal which banks were truly weak before a run; if many banks classified as strong were secretly insolvent or hid losses, the claim that strong banks almost never fail in runs would be overstated.","fun_headline_variants_meta":{"raw":{"variants":["Weak fundamentals, not panic, doom banks in 3,421 runs","Bank runs rarely kill healthy banks in historic data","3,421 runs: strong banks survive, weak ones fail","Fundamentals decide bank run outcomes, study finds","Healthy banks survive runs; weak ones fail, 1863-1934"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000535,"raw_usage":{"total_tokens":2560,"prompt_tokens":920,"completion_tokens":1640,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":1554}},"tokens_in":536,"tokens_out":1640,"duration_ms":9418,"temperature":1.0,"reasoning_tokens":1554,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:39:43.947460+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a set of banks in the top decile of the fundamentals index that failed after a run and whose receivership records show high recovery rates and examiner-assessed asset quality, meaning they were genuinely solvent when they failed. If such cases are numerous, well above the estimated 4 percent failure probability for that decile, the necessity claim fails; if examiner and recovery data confirm that top-decile failure cases were actually insolvent, the run-induced-failure channel for healthy banks is essentially empty.","supporting_citations":[],"review_version":2}