{"id":"a97b62f1-e658-465d-a1bf-84774423946e","arxiv_id":"2505.00268","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This review of 2019-2025 research on language model consistency concludes that the field needs standardized definitions, comprehensive multilingual benchmarks, and mixed human-automatic evaluation.","lead":"This paper surveys how consistency of language models is defined, measured, and improved in the research literature. It argues that the field lacks standard terminology, benchmarks, and evaluation protocols, and calls for interdisciplinary efforts to fix that.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey corpus is undocumented, so the core diagnosis ('no standard approaches', proportions like 'slim majority') rests on an unauditable sample that cannot be reconstructed from the paper.","rationale":"The reader's weakest_assumption identifies the same point, and I agree it is the load-bearing one. My read did not surface a separate fatal flaw. The paper's qualitative contribution—a taxonomy of consistency types and a map of tasks, metrics, and challenges—does not require a perfect sample to be informative, and the call for standardized benchmarks is reasonable even if the precise proportions shift. However, the abstract and introduction make strong field-level claims ('no standard approaches', 'most advanced LLMs ... frequently demonstrate inconsistent behavior') that are supported by the survey rather than by new experiments. Without a documented and reproducible corpus, those claims cannot be audited, which is exactly the correctness risk. The internal tension in Section 2 (calling perturbation-based test construction 'one standard approach' while the abstract denies standard approaches) is a wording issue; the more substantive problem is the undocumented denominator. A conditional accept with a request for a methods appendix and a full study table is the right outcome; no change to the reader's verdict is needed.","tokens_in":11928,"tokens_out":4753,"duration_ms":48644,"concrete_test":"Obtain from the authors the complete list of surveyed studies with unique identifiers, or reconstruct it by marking every reference in the paper that is explicitly used for a landscape claim and deduplicating the preprints (Asai & Hajishirzi 2020a/b; Liu et al. 2023/2024b; Raj et al. 2022/2025). Independently code each study for task type, model architecture, proprietary status, and consistency notion, then recompute all five reported proportions. If any proportion changes by more than 10 percentage points from the stated value, or if the total number of surveyed studies is too small to support the stated percentages, the field-level diagnosis should be weakened and the review needs a documented screening protocol (e.g., PRISMA-style) before publication.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of this review is a field-level diagnosis: consistency research lacks standardized definitions and benchmarks, and 'most advanced LLMs' are inconsistent. For a survey, that claim is only as strong as the corpus reviewed. Section 2 reports quantitative landscape statements—'a slim majority of studies' use standard NLP tasks, 'approximately a third' use custom tasks, 'more than two-thirds' use generative transformers, 'slightly more than half' test proprietary models, 'about a quarter' focus on encoder-only models—but the paper never defines the denominator. There is no search strategy, no inclusion/exclusion criteria, no screening process, and no table of included studies. The reference list mixes surveyed papers with background citations and contains duplicates (Asai & Hajishirzi 2020a/b; Liu et al. 2023/2024b; Raj et al. 2022/2025), so an independent reader cannot reproduce any of the percentages. If the selection is skewed, the headline finding 'there are no standard approaches to assessing model consistency' could be an artifact of the sample rather than a property of the field. This is the load-bearing weakness; the qualitative taxonomy and cited examples are useful but do not by themselves establish the field-level proportions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a narrative review of research on consistency in text-only language models, covering literature from 2019 to 2025. It proposes a taxonomy that separates logical/formal consistency (negational, symmetric, transitive, additive, semantic) from nonlogical/informal consistency (moral, norm, informational/factual), and it surveys which tasks, datasets, models, and evaluation metrics are used in the field. The paper's central claims are that most advanced LLMs exhibit inconsistent behavior, that there is no standard terminology or benchmark for consistency, and that the field needs standardized definitions, comprehensive multilingual benchmarks, and evaluation protocols combining automatic and human assessment. It closes with a call to action and a short appendix on multimodal consistency.","tokens_in":12275,"tokens_out":3891,"duration_ms":39828,"significance":"If established, the paper's central diagnosis would be valuable: it identifies a core reliability property of LMs that currently lacks shared definitions, metrics, and benchmarks, and it connects this to downstream risks in high-stakes and multilingual settings. The taxonomy of consistency types is a useful organizing device, and the discussion of challenges such as decoding-time bias, adversarial consistency attacks, and multilingual gaps is sensible. The paper also gives concrete recommendations that a research community could act on. However, the quantitative landscape claims—the proportions of studies using certain tasks, models, and evaluation styles—are presented without a documented corpus, so the field-level diagnosis is not yet reproducible. The qualitative examples and taxonomy are useful, but they alone do not establish the strength of the headline claims.","major_comments":[{"comment":"The quantitative statements in §2 are the backbone of the claim that consistency research lacks standard approaches: the paper reports that 'a slim majority of studies' use standard NLP tasks, 'approximately a third' use custom tasks, 'more than two-thirds' use transformer-based generative LMs, 'slightly more than half' test proprietary models, and 'about a quarter' focus on encoder-only models. Yet no denominator is defined. There is no search strategy, no inclusion or exclusion criteria, no screening protocol, no coding scheme, and no table of the included studies. The reference list mixes surveyed primary studies with background citations and contains duplicate entries (Asai & Hajishirzi 2020a and 2020b are the same paper; Liu et al. 2023 and Liu et al. 2024b are versions of the same work), so an independent reader cannot reconstruct the sample or verify any of the percentages. This undermines the reproducibility of the central field-level diagnosis. Please add a documented search and screening protocol, a table of included studies with coding, and counts for each reported proportion.","section":"Abstract and §2 (Challenges)"},{"comment":"The abstract states that 'most advanced large language models (LLMs) struggle with consistency and frequently demonstrate inconsistent behavior,' and the introduction repeats this as a field-level fact. The evidence cited is Elazar et al. (2021) and Raj et al. (2025), which demonstrate inconsistency on specific tasks and model families, not a systematic survey of all advanced LLMs. As written, this overstates the strength of the evidence. Please either support the claim with a systematic synthesis of the reviewed corpus or soften it to a claim such as 'many state-of-the-art LLMs have been shown to exhibit inconsistent behavior in a range of tasks and settings.' This distinction matters because the paper's call to action rests on the severity and generality of the inconsistency problem.","section":"Abstract and §2 (Challenges)"},{"comment":"Several other quantitative or comparative claims are asserted without supporting counts. In §2, the paper says 'very few studies explore how inconspicuous or subtle manipulation of prompts can lead to inconsistent LLM responses,' and in §3 it says 'There are surprisingly few approaches that actually increase the consistency of LMs.' These are the kinds of landscape claims that require a denominator and an explicit search. Without a documented corpus, these statements are not auditable. Please provide the number of studies reviewed and the number classified into each category, or rephrase as qualitative observations with representative examples.","section":"§2 (Challenges) and §3 (Improving Consistency)"}],"minor_comments":[{"comment":"The percentages for task categories are presented without stating whether the categories are mutually exclusive; a study using both a standard task and a custom task would need a coding rule. Please clarify the coding procedure or acknowledge overlaps.","section":"§2 (Analyzed Tasks)"},{"comment":"The sentence 'Definitions of factual consistency are often not clearly specified, and instead are replaced with human annotations' is unclear: human annotations can define ground truth, but they do not replace a definition. Please rephrase to clarify the intended contrast.","section":"§2 (Terminology)"},{"comment":"The list of approaches that improve consistency is confusingly phrased: 'Elazar et al. (2021) used a custom loss function, Raj et al. (2025) used knowledge distillation from more consistent teacher models, and Raj et al. (2025); Zhao et al. (2024) used synthetic datasets...' The repetition of Raj et al. (2025) and the semicolon are likely formatting errors. Please revise for clarity and ensure each cited work is credited for the correct contribution.","section":"§3 (Improving Consistency)"},{"comment":"The opening sentence of Appendix A, 'Until 2022, every consistency study was analyzing robustness of LMs to various text perturbations or to semantically equivalent texts only,' is a universal negative claim over all prior studies. This is not supported by the review's documented corpus and could be easily falsified. Please qualify the claim with respect to the reviewed literature.","section":"Appendix A"},{"comment":"There are minor typographical and formatting errors, for example 'parahrasing' in §2 (Dataset Size and Availability), 'Enteprise' in the affiliation block, and inconsistent use of 'LM' versus 'LLM' in places. A careful proofreading pass would improve readability.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a position-style survey rather than a systematic review. For a venue expecting rigorous survey methodology, the missing corpus protocol is a blocking issue, but it is readily fixable within the manuscript's scope by adding a search protocol, inclusion criteria, and a table of coded studies. The authors' own prior works are cited as part of the landscape, but the core diagnosis does not derive from those works, so I do not see a circularity problem. I recommend major revision rather than rejection because the qualitative taxonomy and recommendations are useful and the quantitative gaps can be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you work on consistency or evaluation of LLMs. The paper is a workshop review that organizes the scattered literature into a sensible two-part taxonomy: logical/formal consistency against nonlogical/informal, with the Jang et al. categories and the newer moral/norm/factual variants mapping cleanly into that frame. It also gives a digestible summary of metrics (pairwise similarity, semantic entropy, etc.) and tasks (QA, summarization, NLI, reasoning). The writing is clear, and the call for standardized definitions and holistic benchmarks is reasonable. As a synthesis, it earns its place.\n\nThe load-bearing weakness is the survey methodology, or rather the lack of one. The paper reports proportions—a slim majority of studies use standard NLP tasks, about a third use custom tasks, more than two-thirds use generative transformers, etc.—but never defines the denominator. There is no search strategy, no inclusion/exclusion criteria, no screening process, and no table of included studies. The reference list compounds the problem: Asai & Hajishirzi appears twice (2020a and 2020b), Liu et al. twice (2023 and 2024b), and Raj et al. twice (2022 and 2025), so an independent reader cannot even count the corpus reliably. This matters because the paper's core empirical claim—that the field is fragmented and lacks standard approaches—rests partly on these proportions. If the sample is skewed, the diagnosis could be an artifact of the selection. I would push back on any suggestion that this flaw is minor; it is the difference between a measured landscape and an informed opinion. The author's own citations (Raj et al.) appear as part of the landscape, not as premises for the diagnosis, so I do not see a circularity problem.\n\nMinor points: the multimodal appendix is thin, but the paper explicitly scopes itself to text-only, so that is not a real flaw. Deeper treatment of the trade-off between consistency and creativity is promised but only sketched; again, fine for a workshop position paper.\n\nVerdict: conditional acceptance. A serious referee should ask for the screening protocol and a table of surveyed papers so the percentages are auditable. Without that, the quantitative claims should be softened. The taxonomy and recommendations are worth keeping.","headline":"A clear, useful taxonomy of consistency research, but its field-level percentages rest on an undocumented survey corpus and should not be taken as measured fact.","tokens_in":12664,"tokens_out":1398,"would_cite":false,"duration_ms":15112,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that large language models are frequently inconsistent and that the field lacks any standard definition, metric, or benchmark for consistency, making reliability hard to measure.","keywords":["consistency","large language models","evaluation benchmarks","logical consistency","semantic consistency","cross-lingual consistency","self-consistency","reliability"],"falsifier":"A reader could settle the landscape claim by running a preregistered systematic search over the same 2019–2025 window with explicit inclusion criteria; if it surfaced many peer-reviewed studies on encoder-only models, consistency-oriented pretraining, or adversarial consistency attacks that the review omitted, the claim that those directions are underexplored would need revision. The central claim that no standard evaluation exists would be falsified by the emergence of a benchmark that the field widely adopts as the default consistency test.","tokens_in":11747,"feed_emoji":"📏","tokens_out":6180,"duration_ms":60793,"temperature":0.7,"pith_summary":"This review argues that large language models frequently give different answers to equivalent questions, and that researchers cannot yet reliably measure this because there is no shared definition, metric, or benchmark for consistency. The paper surveys consistency research from 2019 to 2025, organizes the field into logical/formal and nonlogical/informal types of consistency, and finds that a slim majority of studies use standard NLP tasks while most evaluations focus on English, generative transformer models. It also finds that very few methods actually improve consistency, and that most existing approaches address symptoms rather than root causes. If the paper is right, progress on model reliability is currently hard to measure, compare, or verify.","feed_headline":"No standard way exists to measure language-model consistency","feed_subtitle":"A survey of 2019–2025 research finds fragmented definitions, mostly English tests, and few improvement methods.","key_machinery":"The organizing device is a behavioral-consistency taxonomy that splits logical/formal consistency from nonlogical/informal consistency, anchored in the negational, symmetric, transitive, and additive classification and in the semantic equivalence property that equivalent inputs should produce equivalent outputs. The measurement machinery is pairwise evaluation: input-based sampling creates paraphrases or perturbed prompts, output-based sampling generates multiple decodings from identical inputs, and base metrics such as BERTScore, ROUGE, and entailment or contradiction scores are aggregated over pairs, sometimes through semantic entropy. This taxonomy and sampling-and-metric pipeline is what lets the paper compare otherwise disparate studies and identify the gaps in the landscape.","core_discovery":"On the paper's own terms, the central discovery is that consistency research is fragmented: authors define consistency differently or not at all, evaluation relies on ad hoc pairwise similarity metrics applied to paraphrased or repeated prompts, and no standard benchmarks have emerged. The authors organize existing work into logical/formal consistency, which includes negational, symmetric, transitive, additive, and semantic consistency, and nonlogical/informal consistency, which includes moral, norm, and informational/factual consistency. They find that most studies use standard NLP tasks in English with decoder-only or encoder-decoder transformer models, and that current improvement methods, such as fine-tuning on consistent input-output pairs or encouraging self-consistency in chain-of-thought reasoning, mainly address symptoms. The paper concludes with a call for standardized definitions, comprehensive multilingual benchmarks, combined automatic and human evaluation, and research into the structural basis of consistency.","pith_inferences":["If the taxonomy were adopted widely, a natural next step is a consistency report card that scores a model separately on logical, semantic, factual, and moral consistency rather than reporting one aggregate number.","The review's emphasis on root causes suggests a testable hypothesis: if inconsistency is not merely a decoding artifact, then meaning-preserving interpolations in a model's internal representation space should produce consistent outputs, and checking this would localize the source of inconsistency.","A low-cost extension would be to build a multilingual paraphrase set from the existing benchmarks the paper reviews, which would directly quantify the cross-lingual consistency gap the paper identifies.","The review's observation that adversarial consistency attacks are underexplored implies that subtle prompt perturbations could serve as a practical stress test for consistency before deployment."],"forward_implications":["Until definitions and metrics are standardized, consistency scores reported across studies are not comparable, and the performance of state-of-the-art models can be overestimated.","Evaluation protocols that use high-temperature output sampling can inflate apparent inconsistency, so consistency results need to be reported with their sampling setup.","Current improvement methods, such as fine-tuning on consistent input-output pairs and self-consistency decoding, mainly treat symptoms; lasting progress likely requires consistency-oriented pretraining or architectural changes.","Cross-lingual deployment is riskier than monolingual deployment because safety guardrails, factual associations, and even political positions can shift with the input language.","Consistency should be evaluated together with factuality, safety, and helpfulness, since a model can be internally consistent yet factually wrong, or it may sacrifice consistency to maintain safety."],"supporting_citations":[{"why":"Supplies the original logical-consistency taxonomy (negational, symmetric, transitive, additive) and a benchmark for consistency evaluation.","marker":"Jang et al. (2022)"},{"why":"Introduces measuring and improving semantic consistency in pretrained language models, forming the basis for pairwise paraphrase evaluation.","marker":"Elazar et al. (2021)"},{"why":"Defines and measures semantic consistency as a reliability property of large language models.","marker":"Raj et al. (2022)"},{"why":"Provides SelfCheckGPT, a black-box method for detecting factual inconsistency through sampled generations.","marker":"Manakul et al. (2023)"},{"why":"Separates faithfulness from self-consistency when evaluating natural language explanations.","marker":"Parcalabescu & Frank (2024)"},{"why":"Documents self-contradictory hallucinations and uses sequential aggregation of contradiction scores for factual consistency.","marker":"Mündler et al. (2024)"},{"why":"Introduces semantic entropy over an entire set of outputs, used to measure consistency and uncertainty.","marker":"Kuhn et al. (2023)"},{"why":"Shows that self-consistency across chain-of-thought samples improves reasoning, a key improvement approach the review identifies.","marker":"Wang et al. (2023)"},{"why":"Uses knowledge distillation from more consistent teacher models to improve consistency in large language models.","marker":"Raj et al. (2025)"},{"why":"Introduces the SaGE score for evaluating moral consistency across semantically equivalent scenarios.","marker":"Bonagiri et al. (2024)"}],"fun_headline_variants":["Consistency in LMs: Fragmented definitions, no benchmarks","Survey finds LM consistency research lacks unified metrics","No standard benchmarks exist for LM consistency","LM consistency: No unified definitions or benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's landscape description rests on the assumption that the peer-reviewed publications and influential preprints it selected from 2019 to 2025 are representative of consistency research as a whole.","fun_headline_variants_meta":{"raw":{"variants":["Consistency in LMs: Fragmented definitions, no benchmarks","Survey finds LM consistency research lacks unified metrics","No standard benchmarks exist for LM consistency","LM consistency: No unified definitions or benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000402,"raw_usage":{"total_tokens":2017,"prompt_tokens":787,"completion_tokens":1230,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":403,"completion_tokens_details":{"reasoning_tokens":1172}},"tokens_in":403,"tokens_out":1230,"duration_ms":9294,"temperature":1.0,"reasoning_tokens":1172,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:45:24.411636+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could settle the landscape claim by running a preregistered systematic search over the same 2019–2025 window with explicit inclusion criteria; if it surfaced many peer-reviewed studies on encoder-only models, consistency-oriented pretraining, or adversarial consistency attacks that the review omitted, the claim that those directions are underexplored would need revision. The central claim that no standard evaluation exists would be falsified by the emergence of a benchmark that the field widely adopts as the default consistency test.","supporting_citations":[],"review_version":1}