{"id":"85933620-e6b5-4ca9-821d-396c5bd56efb","arxiv_id":"2504.17052","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM political stances are topic-dependent and vary in stability under argumentative pressure, with unstable stances more easily reversed by prompting or fine-tuning.","lead":"This paper introduces PReSS, a stress-test framework that measures whether LLMs hold their political stance when confronted with supporting and counter arguments. Across 12 models and 19 economic topics, it finds that stance stability varies by topic and that unstable stances are more likely to flip under ideology-reversal interventions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stance typology contradicted: Eq. 3 makes UR shift left and UL shift right, while Table 1 states the opposite, so the four-class labels behind the controllability claim are ambiguous.","rationale":"The paper proposes a genuinely useful framework, and the empirical pattern—topic-dependent stability and greater malleability of unstable stances under fine-tuning—is plausible. The reader's conditional verdict is reasonable. In stress-testing, I looked first at the weakest link in the chain from raw responses to the headline claim: the four-class taxonomy. The formal definitions in Eq. 1-3 are explicit: IB=1 marks an original right-aligned stance; UR is the unstable version of that, so the stance must move left under pressure. Table 1 describes the opposite movement for UR and UL. This is not a subtle interpretation issue: the two sources cannot both be true. Since the Section 6.2 claim depends on the exact UR and UL labels, the paper's central quantitative evidence is currently ambiguous. I considered the reader's weaker assumption, unidimensionality of the 19 items. That is a real empirical threat, but the authors explicitly scope the framework to the economic axis and the limitation section acknowledges multi-dimensionality; a second factor (25.66% variance after rotation) is concerning but does not by itself show the polarity labels are wrong. The typology contradiction is a direct internal inconsistency in the measurement instrument, so it takes priority. If the authors release the classification code and the full transition table, the ambiguity can be resolved; if the pattern survives both labelings, the central claim is much stronger.","tokens_in":14545,"tokens_out":9237,"duration_ms":84871,"concrete_test":"Obtain the authors' classification script or raw response data (the paper promises release). Recompute all stance labels strictly from Eq. 3 and also from the Table 1 verbal descriptions. Compare the resulting four-type labels and the Section 6.2 fine-tuning transition matrix. The central claim survives only if, under both labeling conventions, the probability of an initially unstable topic (UR union UL) becoming SR is significantly greater than the SL->SR probability; if the UR/UL swap changes which topics contribute to the 61.1%/57.8% figures, the reported numbers are artifacts of an inconsistent definition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equation 3 and Table 1 define the four stance types inconsistently. IB=1 iff the original stance is right-aligned (b=+1 and o=+1, or b=-1 and o=+1), so UR (delta=0, IB=1) is a right-aligned original stance that changes under argumentative pressure; since delta=0 forces a != o, the post-pressure stance is left. UR thus shifts left, not right. Conversely, UL (delta=0, IB=0) is an original left stance that shifts right. Table 1 states the opposite for both. The formalization in Section 3.2 and the interpretive table cannot both be correct. Every downstream use of UR/UL—Figure 8, the Section 5.1 distribution, and especially the Section 6.2 transition probabilities UR->SR (61.1%) and UL->SR (57.8%)—depends on which rule was implemented. If Table 1 was followed, the code swapped the two unstable categories relative to Eq. 3; if Eq. 3 was followed, the paper's interpretation is wrong and the reported transition rates are attached to mislabeled types. Without a code release or a correction, the central claim that stability labels predict controllability is not reproducible from the text.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PReSS, a black-box framework for measuring whether LLM political stances are stable or unstable under argumentative pressure, classifying each model-topic pair as stable-left, unstable-left, stable-right, or unstable-right. Applying PReSS to 12 instruction-tuned LLMs across 19 Political Compass economic statements, the authors report substantial topic-level variation in stability, a correlation between instability and semantic-entropy uncertainty (AUROC up to 0.78), and evidence that stable stances resist ideology reversal by prompting and DPO fine-tuning while unstable stances are more readily converted to stable-right. The paper argues that stability should be treated as a distinct dimension of political bias in LLM evaluation.","tokens_in":14747,"tokens_out":3340,"duration_ms":33371,"significance":"If the framework and its empirical claims withstand scrutiny, PReSS would be a useful and practical contribution: it offers a low-cost black-box protocol, uses open-ended rather than forced-choice prompts, and connects an interpretable stability taxonomy to both white-box uncertainty metrics and downstream controllability. The proposed distinction between stable and unstable stances is potentially valuable for alignment and safety evaluations. However, the paper's central claims currently rest on an internal inconsistency in the stance typology, an unvalidated unidimensionality assumption, and outcome definitions that are partly circular with respect to the predictor labels. Each of these issues is load-bearing for the headline conclusion that stability labels predict controllability, so the contribution cannot be accepted in its present form.","major_comments":[{"comment":"The formal definition and the interpretive table define the unstable categories in opposite directions. Eq. (3) sets IB=1 exactly when the original stance o is right-aligned (for b=+1, o=+1; for b=-1, o=+1). For UR, δ=0 means a≠o, so a post-pressure stance is necessarily left. Thus UR denotes an original right stance that shifts left under pressure, and UL denotes an original left stance that shifts right. Table 1 states the reverse for both UR and UL. Since the labels UR and UL feed every downstream analysis, including Table 2, Fig. 8, and the §6.2 transition probabilities UR→SR (61.1%) and UL→SR (57.8%), the paper must specify which rule was implemented and correct either the equations or the interpretation. Without a code release or a worked example, the central controllability claims are not reproducible from the text.","section":"§3.2, Eq. (3) and Table 1"},{"comment":"The factor analysis does not support the claim of a single dominant economic dimension. Both the unrotated solution (eigenvalues 5.72 and 1.63) and the rotated solution retain two factors under the Kaiser criterion, with the rotated second factor explaining 25.66% of the variance. Several items load substantially on Factor 2 (e.g., Q8 0.72, Q14 0.54, Q7 0.53, Q15 0.44), so the items are not unidimensional in the sample used. Since the left/right polarity labels assigned to the 19 statements depend on the single-axis assumption, this gap could confound the stance-direction classifications and, in turn, the stability labels. A parallel-analysis or eigenvalue-based justification is needed, or the analysis should be restricted to items that are demonstrably unidimensional.","section":"§3.1 and Appendix B (Tables 5–7)"},{"comment":"The SF/SU/ID outcomes used to demonstrate controllability are defined from the same PReSS stability labels that are the predictor of interest. SF is defined as stable stances of only one direction appearing across conditions, SU as stable stances of both orientations, and ID as no stable stance appearing; these are essentially restatements of the topic-level SL/SR/UL/UR classification under different prompts. Consequently, the high PSF rates in Figs. 5–6 may reflect the definition rather than independent evidence that stability predicts controllability. The paper should either define an outcome measure that does not presuppose the four-class taxonomy, or explicitly justify why the taxonomy-based outcome is not circular. In addition, the headline transition probabilities UR→SR≈61.1% and UL→SR≈57.8% and the claim that SL→SR is significantly less frequent are asserted in §6.2 without a table, confidence intervals, or per-model breakdowns, with the full table promised only for a later version. These numbers are the central evidence for the controllability claim and must be presented and accompanied by uncertainty estimates.","section":"§6 and §6.2"},{"comment":"The abstract states that PReSS was applied to 9 LLMs, while the full text, Fig. 2, and Section 4 consistently report 12 models. This inconsistency should be resolved. Additionally, Section 5.1 reports Table 2 as using 'a single set of supporting- and counter-arguments,' while Section 5.2 and Section 6 use three independent argument sets; the paper should clarify which argument set was used for the baseline distribution and whether the reported percentages and confidence intervals account for variation across the three sets.","section":"Abstract and Section 5.1"}],"minor_comments":[{"comment":"There are several typographical errors, including 'tyree' in Appendix A and 'Y ue Dong' in the author block; these should be corrected.","section":"Throughout"},{"comment":"The AUROC values are reported only as printed on the figure, with no confidence intervals or significance tests across the 12 models; a small table with bootstrap intervals would strengthen the validation claim.","section":"Fig. 4 and §5.3"},{"comment":"The rotated loadings are not discussed item by item; the paper should identify which items are not clearly aligned with the intended polarity and explain their impact on the stance labels.","section":"Appendix B, Table 7"},{"comment":"The limitations section is honest and useful, but it should be expanded to mention the internal typology inconsistency and the factor-analysis issue, rather than only topic coverage and axis dimensionality.","section":"Section 8 (Limitations)"}],"recommendation":"major_revision","confidential_remarks":"The typology contradiction in §3.2 is not a matter of wording: Eq. (3) and Table 1 disagree on which direction UR and UL shift, and every downstream analysis inherits one of the two interpretations. Until the authors specify which rule was implemented and provide code or worked examples, the empirical claims should be treated with caution. The circularity of the SF/SU/ID outcomes is also serious because it may inflate the apparent predictive power of stability labels. These issues are fixable in principle, but they require substantial revision rather than local copyediting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper proposes PReSS, a black-box framework that classifies LLM political responses into four stance types—stable-right, stable-left, and their unstable counterparts—by stress-testing models with supportive and counter arguments. The core contribution, investigating stance stability rather than a static left-right position, is genuinely new and worth building on. The open-ended prompting protocol, three argument sets, and the semantic-entropy validation (AUROC up to 0.78) are solid methodological choices.\n\nThe major problem is an internal contradiction in the formalization. Equation (3) defines UR as δ=0, IB=1, which means the original stance is right-aligned and the response changes under pressure, so the post-pressure stance is left. Table 1, however, says UR \"shifts toward right,\" and the description for UL is similarly inverted. These cannot both be true. Every result that uses UR/UL—Figure 8, the distribution in Table 2, and especially the Section 6.2 transition probabilities (UR→SR 61.1%, UL→SR 57.8%)—depends on which labeling was actually implemented. Without code release, this is not reproducible from the text. This needs a correction and a rerun of the affected analyses.\n\nA second issue is the factor analysis used to support unidimensionality. The rotated solution retains a second factor with eigenvalue 2.16 and 25.66% of the variance. That is not a clean unidimensional result. The authors claim it supports a single-dimension reading, but it suggests a meaningful second component that could confound the left/right polarity labels.\n\nSmaller problems: the NLI stance classifier has no human validation; the paper says \"single set\" in Section 5.1 but \"three sets\" elsewhere; and the key transition table in Section 6.2 is deferred to the final version. The circularity concern is milder than it looks—the fine-tuning intervention is distinct from the original stress test, so the stability label isn't purely predicting itself.\n\nThe citation coverage is reasonable; the work engages with prior static-leaning measurements and uncertainty literature.\n\nBottom line: the framework deserves serious referee time. The revision must resolve the UR/UL contradiction, release the transition table, and either fix the factor analysis or soften the unidimensionality claim. As presented, the specific numbers are not trustworthy.\n\nRecommendation: send to peer review.","headline":"A useful framework for measuring LLM stance stability, but a load-bearing contradiction in the UR/UL definitions and an overclaimed factor analysis undermine the current version.","tokens_in":15279,"tokens_out":4881,"would_cite":true,"duration_ms":43876,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes PReSS, a black-box stress test that labels each LLM's topic-level political stance stable or unstable, and shows the labels predict which stances survive fine-tuning.","keywords":["political stance stability","large language models","political bias","ideology reversal","fine-tuning","argumentative pressure","semantic entropy","black-box evaluation"],"falsifier":"Re-run PReSS on the two factor-homogeneous subsets of the 19 statements (items loading primarily on the first versus the second rotated factor). If stability scores and the UR→SR / UL→SR transition rates shift sharply between subsets, the single-axis assumption is load-bearing and the reported controllability numbers are partly an artifact of item selection. A cleaner test: run the mirror-image experiment with left-alignment DPO and check whether SL topics convert to stable-left at comparable rates — if not, the claimed stability–controllability link is specific to one intervention direction.","tokens_in":14332,"feed_emoji":"🗳️","tokens_out":8331,"duration_ms":66771,"temperature":0.7,"pith_summary":"Existing evaluations of political bias in large language models put a model on a left–right spectrum. This paper argues that a second property matters just as much: stance stability, whether a model holds its position on a specific topic when confronted with supporting or counter-arguments. The authors build PReSS, a black-box framework that probes each of 19 economic statements in three forms — neutral, with a supporting argument, with a counter-argument — and labels every model–topic pair as stable-left, unstable-left, stable-right, or unstable-right. Applied to 12 instruction-tuned models, the framework shows that global ideology does not determine topic-level behavior: left-leaning models take right stances on 27.6% of topics, and right-leaning models take left stances on 34.3%. The load-bearing result is that stability predicts controllability: under right-alignment fine-tuning, unstable topics convert to stable-right roughly 58–61% of the time, while reversing a stable-left stance is significantly rarer.","feed_headline":"Unstable political stances in LLMs are the ones fine-tuning can flip","feed_subtitle":"A black-box stress test labels each model-topic stance stable or unstable, predicting which political biases resist alignment.","key_machinery":"The carrying object is the four-class stance typology built from two binary signals: the persistence indicator $\\delta(o,a)$, which is 1 when the model's stance survives both supportive and counter-argumentative prompting, and the bias-alignment indicator $I_B(o,b)$, which is 1 when the original stance matches the annotated rightward (or leftward) polarity of the statement. Their combination yields stable-left, unstable-left, stable-right, and unstable-right for every model–topic pair. This typology does the work: direction alone (left vs right) does not predict fine-tuning outcomes, but the stable/unstable split does, which is why the paper can treat stability as a moderating factor for debiasing and alignment.","core_discovery":"On its own terms, the paper establishes that stance stability is a measurable, topic-specific property of LLMs and that PReSS labels carry predictive force for alignment interventions. Each response under the three argumentative conditions is coded by whether the stance persists ($\\delta = 1$) and whether it aligns with the statement's right-bias direction ($I_B = 1$), producing the four classes. A model that is left-leaning overall can hold a stable-right stance on one topic and an unstable-left stance on another; the topic-wise stability score $S_{t,m}$ across three argument sets ranges from near zero to one within a single model. The decisive empirical claim concerns fine-tuning: transitions UR→SR and UL→SR occur with probabilities around 61.1% and 57.8% under right-alignment DPO, whereas SL→SR is significantly less frequent, so the framework's stability labels identify where fine-tuning has leverage and where a model preserves prior commitments. The paper further validates the black-box labels against semantic entropy, a white-box uncertainty measure, reaching AUROC 0.78, and reports that prompting with opposite-persona instructions rarely overturns a stable stance.","pith_inferences":["If stability is a genuine model property, fine-tuning on a stable topic should require proportionally more preference data than on an unstable topic; the paper does not vary dataset size per topic, so this is a direct testable extension.","The factor analysis retains a second factor (25.66% of rotated variance) by the Kaiser criterion, so part of what the paper labels left-right polarity may be a second ideological axis; re-running the typology within factor-homogeneous statement subsets would show whether the stability and transition results survive.","The reported asymmetry — both unstable-right and unstable-left topics convert to stable-right under right-alignment — could reflect the intervention direction rather than an intrinsic property; a mirror experiment with left-alignment DPO should show whether SL conversions become the frequent ones.","Because semantic entropy correlates with instability at AUROC 0.78, unstable stances may be cases of genuine model uncertainty; one could check whether unstable topics are also those with highest lexical diversity across unpressured paraphrases, connecting stance stability to the broader distinction between epistemic uncertainty and sycophantic compliance."],"forward_implications":["Debiasing or ideology-reversal pipelines should be targeted at topics labeled unstable, where the reported transition rates (roughly 58–61%) show fine-tuning has leverage, rather than applied uniformly across topics.","Model-level left/right classifications should be replaced or supplemented by topic-level stance maps; a single model can be stable-right on some topics and unstable-left on others.","Three black-box probes per topic (neutral, supporting, counter) can substitute for roughly 20 generations of white-box uncertainty estimation when the goal is to locate malleable political stances, since PReSS labels reach AUROC 0.78 against semantic entropy.","Any application that requires the model to hold a consistent political position — tutoring, multi-agent debate, AI-powered persuasion — should audit the relevant topics with a stability probe before deployment.","Because stability and instability are properties of model–topic pairs rather than of whole models, safety evaluations that sample only a few topics could miss either the most entrenched or the most flippable behaviors of the same system."],"supporting_citations":[{"why":"Supplies the open-ended stance elicitation protocol that PReSS adopts to avoid forced-choice bias.","marker":"Feng et al., 2023"},{"why":"Provides the right-partisan preference data used to construct the right-leaning candidate models and the fine-tuning intervention.","marker":"Agiza et al., 2024"},{"why":"The DPO algorithm behind the ideology-reversal fine-tuning whose transition rates the stability labels predict.","marker":"Rafailov et al., 2023"},{"why":"Semantic entropy, the white-box uncertainty measure used to validate PReSS labels with AUROC 0.78.","marker":"Farquhar et al., 2024"},{"why":"The NLI model that classifies each generated response as agreement or disagreement with the statement.","marker":"Lewis et al., 2019"},{"why":"Documents prompt-based susceptibility to influence, the comparison baseline for the prompting version of ideology reversal.","marker":"Anagnostidis and Bulian, 2024"},{"why":"Provides the predictive entropy formulation reused as a secondary uncertainty baseline.","marker":"Kadavath et al., 2022"}],"fun_headline_variants":["Stable political stances resist fine-tuning in LLMs","Black-box test spots which LLM political stances flip","LLM political stance stability predicts alignment success","PReSS labels: topic-dependent political stability in LLMs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 19 Political Compass statements are treated as measuring a single left-right economic axis, yet the paper's own factor analysis keeps a second factor above the retention threshold (eigenvalue 1.62; 25.66% of variance after rotation) — if that second dimension is real, the left/right polarity labels that drive the whole typology could be confounded.","fun_headline_variants_meta":{"raw":{"variants":["Stable political stances resist fine-tuning in LLMs","Black-box test spots which LLM political stances flip","LLM political stance stability predicts alignment success","PReSS labels: topic-dependent political stability in LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1446,"prompt_tokens":1021,"completion_tokens":425,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":637,"completion_tokens_details":{"reasoning_tokens":360}},"tokens_in":637,"tokens_out":425,"duration_ms":4361,"temperature":1.0,"reasoning_tokens":360,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:50:38.441808+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run PReSS on the two factor-homogeneous subsets of the 19 statements (items loading primarily on the first versus the second rotated factor). If stability scores and the UR→SR / UL→SR transition rates shift sharply between subsets, the single-axis assumption is load-bearing and the reported controllability numbers are partly an artifact of item selection. A cleaner test: run the mirror-image experiment with left-alignment DPO and check whether SL topics convert to stable-left at comparable rates — if not, the claimed stability–controllability link is specific to one intervention direction.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the right-partisan preference data used to construct the right-leaning candidate models and the fine-tuning intervention."},{"cited_title":"However, we find that they are significantly more resilient on topics with stable stances","cited_arxiv_id":null,"evidence_quote":"The DPO algorithm behind the ideology-reversal fine-tuning whose transition rates the stability labels predict."}],"review_version":1}