{"id":"663850f2-6ee0-4fc5-8383-982d5366044f","arxiv_id":"2604.19699","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new LLM-derived metric for evidence versus intuition in political speech correlates positively with deliberative democracy and law transparency across 15 million speech segments from 1946-2025.","lead":"The paper introduces the Evidence-Minus-Intuition (EMI) score, derived from LLM ratings and semantic embeddings, to quantify evidence-based versus intuition-based language in parliamentary speeches. It reports that higher EMI scores are positively associated with measures of deliberative democracy and governance quality across seven countries over eight decades.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"LLM-derived EMI score validity is the load-bearing assumption; confounds by topic/speaker/model could artifactually drive the reported associations","rationale":"The reader's weakest_assumption directly identifies the measurement validity issue that must hold for any downstream association to be interpretable. Because the full text was unavailable to the reader, the current UNVERDICTED status remains appropriate until the concrete validation test above is performed; no stronger internal inconsistency is visible from the provided abstract and claim description.","tokens_in":1654,"tokens_out":388,"duration_ms":16621,"concrete_test":"Sample 1,000 speech segments stratified by country, decade, and topic; obtain independent human ratings of evidence-vs-intuition on a 1-5 scale; compute Pearson correlation with the paper's EMI scores and re-run the main within-country regression after residualizing EMI on topic and speaker fixed effects. If the human correlation falls below 0.55 or the democracy coefficient drops >30% after residualization, the association is not robust to measurement artifacts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that EMI (LLM ratings of evidence vs. intuition plus embedding similarity) is a stable, unbiased proxy for epistemic orientation across 15M segments, countries, and 80 years. Parliamentary speech is heavily topic-dependent (e.g., budget debates vs. moral issues differ systematically in verifiable claims), speaker style correlates with party and era, and LLMs can embed training-data priors on what counts as 'evidence.' If these factors are not fully orthogonalized from the EMI signal, the within-country temporal correlation with deliberative-democracy indices and rule-of-law measures could be spurious rather than causal or even reflective of the intended construct. The abstract and methods summary provide no indication of human validation, cross-model robustness, or topic-fixed-effects checks that would secure this step.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces the Evidence-Minus-Intuition (EMI) score, derived from LLM ratings of evidence versus intuition in speech segments combined with embedding-based semantic similarity, as a scalable measure of epistemic orientation. Applying this to 15 million parliamentary speech segments across seven countries from 1946 to 2025, the authors report positive within-country temporal associations between EMI and indices of deliberative democracy (including in lagged specifications) as well as with the transparency and predictable implementation of laws as a governance dimension.","tokens_in":1830,"tokens_out":610,"duration_ms":30418,"significance":"If the EMI score proves to be a valid, stable, and unconfounded proxy for epistemic orientation, the work would deliver large-scale empirical evidence linking evidence-based parliamentary discourse to democratic quality and governance outcomes over decades and across countries. The scale (15M segments), multi-country coverage, long temporal window, and use of both contemporaneous and lagged within-country designs are notable strengths that could advance computational approaches to political discourse analysis.","major_comments":[{"comment":"The manuscript provides no reported validation of the LLM-based EMI ratings against human judgments, inter-rater reliability metrics, or cross-model robustness checks. This is load-bearing because LLM-as-judge outputs are known to be sensitive to prompt wording, model choice, and training-data priors on what constitutes 'evidence,' directly affecting interpretation of the positive associations with deliberative democracy and rule-of-law measures.","section":"Methods (EMI score construction)"},{"comment":"No topic-fixed effects, speaker-style controls, or checks for topic-dependent language use (e.g., budget debates vs. moral-issue debates) are described in the association analyses. Parliamentary speech topics correlate with both verifiable claims and democratic indices, raising the possibility that the reported within-country temporal correlations are partly artifactual rather than reflective of epistemic orientation.","section":"Results (temporal and lagged analyses)"},{"comment":"The abstract and methods summary supply no information on statistical controls for confounders, robustness to alternative embedding models, or falsification tests (e.g., placebo outcomes). These omissions leave the central claim that EMI is 'positively associated' vulnerable to alternative explanations.","section":"Abstract and Results"}],"minor_comments":[{"comment":"Clarify the exact data cutoff for the 1946-2025 span and whether any segments post-2023 rely on imputation or different collection methods.","section":"Data description"},{"comment":"The notation for the EMI score (difference of ratings plus embedding similarity) should be formalized with an explicit equation to aid reproducibility.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":"The paper sits at the intersection of computational linguistics and political science; the absence of basic validation steps for the core measure is a standard expectation in this venue and should be addressed before acceptance."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive and detailed comments, which highlight important areas for strengthening the validity and robustness of our EMI score and its reported associations. We address each major comment point by point below, with clear indications of revisions to be incorporated in the next version of the manuscript.","responses":[{"response":"We agree that explicit validation is essential for interpreting the EMI score, particularly given documented sensitivities of LLM judgments. Although the initial submission did not include these details, we have conducted a post-submission human validation on a stratified sample of 500 speech segments independently rated by two expert coders, yielding Cohen's kappa of 0.68 for evidence vs. intuition classification. We will also add cross-model robustness checks using GPT-4o and an open-source model (Llama-3-70B). These results, along with prompt sensitivity analyses, will be added to a new subsection in the Methods section of the revised manuscript, with discussion of how they support the main findings.","revision_made":"yes","referee_comment":"The manuscript provides no reported validation of the LLM-based EMI ratings against human judgments, inter-rater reliability metrics, or cross-model robustness checks. This is load-bearing because LLM-as-judge outputs are known to be sensitive to prompt wording, model choice, and training-data priors on what constitutes 'evidence,' directly affecting interpretation of the positive associations with deliberative democracy and rule-of-law measures."},{"response":"The referee correctly identifies a potential source of confounding, as topic composition may correlate with both EMI and the outcome measures. Our primary specifications rely on within-country temporal variation with year fixed effects and lagged EMI to reduce reverse causality and stable confounders, but we did not explicitly model topic dependence. In the revision, we will incorporate topic-fixed effects using LDA-derived topics from the full 15M-segment corpus and add supplementary analyses with speaker fixed effects (where speaker identifiers are available). These additions will be reported in the Results section to demonstrate that the associations hold after accounting for topic and speaker variation.","revision_made":"yes","referee_comment":"No topic-fixed effects, speaker-style controls, or checks for topic-dependent language use (e.g., budget debates vs. moral-issue debates) are described in the association analyses. Parliamentary speech topics correlate with both verifiable claims and democratic indices, raising the possibility that the reported within-country temporal correlations are partly artifactual rather than reflective of epistemic orientation."},{"response":"We acknowledge that the original abstract and methods did not sufficiently detail these elements, leaving the associations open to alternative interpretations. In the revised manuscript, we will update the abstract to reference key robustness features and expand the Results section with: (i) additional controls for time-varying confounders such as GDP per capita and government ideology; (ii) robustness checks using alternative embedding models (e.g., different sentence-transformer variants); and (iii) falsification tests with placebo outcomes (e.g., unrelated governance indicators like infrastructure spending). These will be presented in new tables, with the abstract revised accordingly.","revision_made":"yes","referee_comment":"The abstract and methods summary supply no information on statistical controls for confounders, robustness to alternative embedding models, or falsification tests (e.g., placebo outcomes). These omissions leave the central claim that EMI is 'positively associated' vulnerable to alternative explanations."}],"tokens_in":1400,"tokens_out":715,"duration_ms":51388,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core result is that higher EMI scores—measuring evidence-based over intuition-based language in speeches—track with better deliberative democracy indices and law transparency/predictability, both contemporaneously and in lagged models across seven countries from 1946 onward. This uses a fresh operationalization on a genuinely large corpus of 15 million segments, which is the main advance over prior discourse work that stayed smaller or more qualitative. The within-country temporal design is a reasonable way to look at change rather than just cross-sectional differences. The scale and the lagged checks are the parts that stand out as useful data points for anyone tracking how political talk relates to institutional quality. The soft spot is the EMI construction itself. It rests on LLM ratings plus embeddings with no reported human validation, no cross-model checks, and no obvious handling for topic or speaker confounds that could easily bleed into the scores. Parliamentary debates on budgets versus social issues differ systematically in verifiable content, and LLMs carry their own priors on what counts as evidence. Without those controls or robustness tests, the associations risk being driven by artifacts rather than the intended construct. The abstract gives little on statistical details or alternative specifications, so the strength of the evidence is hard to judge from what's here. This is the kind of work that computational political scientists or democracy researchers might want to see in a reading group for the dataset and the basic correlation pattern, even if they end up skeptical on the measurement. It deserves peer review because the data effort is substantial and the question is substantive; a referee could push on validation and confounds without the paper being a non-starter.","headline":"The paper's new EMI score from LLM ratings on 15M parliamentary segments shows positive links to deliberative democracy and rule-of-law measures within countries over time, but the score's validity as a clean measure of epistemic orientation is untested.","tokens_in":2301,"tokens_out":413,"would_cite":false,"duration_ms":22579,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Parliamentary speeches favoring evidence over intuition track higher deliberative democracy across countries and decades.","keywords":["epistemic orientation","parliamentary discourse","deliberative democracy","evidence-based reasoning","large language models","political speech analysis","governance indicators"],"falsifier":"A replication that rates the same speech segments with human coders and finds no association between those human EMI scores and deliberative democracy indices.","tokens_in":2559,"feed_emoji":"📊","tokens_out":594,"duration_ms":26844,"temperature":0.7,"pith_summary":"The paper introduces a measure called the Evidence-Minus-Intuition score to quantify whether political speech relies on verifiable information or on subjective beliefs. It applies this score to 15 million parliamentary speech segments from seven countries spanning 1946 to 2025. The central finding is a consistent positive association between higher EMI scores and indices of deliberative democracy within each country over time, as well as with the transparency and predictable enforcement of laws. A sympathetic reader would care because the result suggests that the epistemic style of elite discourse is not merely stylistic but tied to measurable qualities of democratic governance.","feed_headline":"Evidence-based parliamentary talk tracks stronger democracy","feed_subtitle":"EMI scores from 15 million speeches rise with deliberative quality and law transparency within countries over time.","key_machinery":"The Evidence-Minus-Intuition (EMI) score, obtained from LLM ratings of evidence versus intuition in speech segments combined with embedding-based semantic similarity to quantify epistemic orientation.","core_discovery":"Using large language models to rate speech segments for evidence-based versus intuition-based reasoning, the authors derive an EMI score for each segment and aggregate it at the country-year level. They report that EMI is positively associated with deliberative democracy in both contemporaneous and lagged within-country analyses, and that EMI also correlates positively with the governance dimension of transparent and predictable law implementation.","pith_inferences":["If the association is causal, interventions that shift parliamentary language toward evidence could produce measurable improvements in democracy scores.","The method could be extended to other text corpora such as social media or court rulings to test whether the same epistemic pattern appears outside formal legislatures.","Countries that already score high on deliberative democracy may sustain those scores partly by maintaining higher EMI in ongoing debate."],"forward_implications":["Higher average EMI in a country's parliament in one year predicts higher deliberative democracy scores in subsequent years.","The association appears for the specific governance sub-component of transparent and predictable laws.","The pattern holds across multiple countries and remains stable when speech is examined at the level of individual segments.","Temporal changes in parliamentary discourse style can be tracked quantitatively over eight decades."],"fun_headline_variants":["Evidence orientation in parliaments links to deliberative democracy","EMI scores link evidence talk to democracy and law transparency","Parliamentary speeches EMI associates with deliberative democracy","Evidence minus intuition tracks with deliberative quality over time"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That automated LLM ratings of evidence versus intuition in political speech produce a valid and unbiased measure of epistemic orientation that is not driven by topic, speaker style, or model artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Evidence orientation in parliaments links to deliberative democracy","EMI scores link evidence talk to democracy and law transparency","Parliamentary speeches EMI associates with deliberative democracy","Evidence minus intuition tracks with deliberative quality over time"]},"model":"grok-4.3","cost_usd":0.007804,"raw_usage":{"total_tokens":3446,"prompt_tokens":595,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":78040500,"prompt_tokens_details":{"text_tokens":595,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2791,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":595,"tokens_out":60,"duration_ms":25639,"temperature":1.0,"reasoning_tokens":2791,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T02:38:13.750399+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication that rates the same speech segments with human coders and finds no association between those human EMI scores and deliberative democracy indices.","supporting_citations":[],"review_version":1}