{"id":"3a0fa778-fdc0-4372-8bef-93b76ad487a8","arxiv_id":"2509.11208","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Order-induced prediction error in binary question answering grows logarithmically with evidence length, and a pre-specified information-sufficiency gate abstains on uncertain items to hold hallucination near zero.","lead":"This paper claims that language models hallucinate because evidence order changes their predictions, and proposes an information-budget rule that tells the model when to abstain. The authors turn a measured uncertainty signal into an answer/refuse decision using bits-to-trust math.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central 'Bayesian in expectation' claim rests on unproven training/architecture premise: standard next-token training is never shown to minimize Eπ[ℓ(Y|Γπ(X))]; Theorem 3's averaging head is non-standard and no convergence-to-I-projection result is given.","rationale":"I agree with the reader that the weakest assumption is the unverified training/closure premise. The central claim is a conjunction: transformers minimize Eπ[ℓ], therefore dispersion is a compression failure, therefore ISR gates hallucination. The first conjunct is the root; if it fails, the explanatory mechanism for the empirical laws is absent. Theorem 3 is an existence result for a non-standard architecture, not a statement about standard transformers or about gradient descent; Theorem 4's infimum over the convex hull does not equal the infimum over the standard parameter class unless closure is shown. The paper's validation experiments—dispersion scaling, Jensen gains, dose-response—test consequences of order sensitivity but do not test whether the trained model attains the order-averaged optimum; indeed the fixed-order model's CE is higher than the uniform mixture's, which is at least consistent with the premise failing. Because the empirical contributions may stand on their own, I would not reject the paper outright, but the theory's key premise is unsupported; the reader's CONDITIONAL verdict is appropriate. I would keep it unchanged. (A secondary concern about the ISR gate's one-sided bound and the unspecified reference distribution P in Algorithm 1 reinforces the need for revision.)","tokens_in":12144,"tokens_out":21260,"duration_ms":251746,"concrete_test":"Train a small transformer from scratch on a synthetic exchangeable binary-adjudication task (i.i.d. evidence chunks, label from a known sufficient statistic) with only the canonical order per item—no permutation augmentation. After convergence, for held-out items compute (i) the order-averaged loss Eπ[-log pθ(Y|Γπ(X))], (ii) the loss of the same model trained with uniform random permutation augmentation, and (iii) the loss of the tied-weight averaging-head ensemble from Theorem 3. If either (ii) or (iii) achieves strictly lower order-averaged loss, or if the fixed-order model's predictions differ materially from the exchangeable Bayes predictor, then standard training does not minimize Eπ and the central premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's resolution of the Bayesian-in-expectation paradox rests on the premise that standard transformers minimize Eπ[ℓ(Y|Γπ(X))] (Section 3.3, Theorems 3–4). This premise is load-bearing: without it, the QMV bound is merely a descriptive property of order-sensitive models, and the claim that hallucinations are predictable compression failures loses its mechanism. Theorems 3–4 do not establish the premise. Theorem 3 constructs a non-standard tied-weight multi-branch ensemble with an averaging head; it shows only that the closure of the model family contains permutation mixtures, not that a standard transformer contains such a head or that its training objective drives it to the order-averaged I-projection. Theorem 4's infimum is taken over that convex hull, not over the parameter class of an ordinary decoder, and no training experiment demonstrates convergence to the exchangeable target. Standard next-token training on naturally ordered evidence optimizes E_{(X,Y)}[-log pθ(Y|X)]; equality with the order-averaged objective requires an exchangeable population and does not imply per-instance minimization of Eπ[ℓ(Y|Γπ(x))]. The paper's own Experiment 1 shows a fixed-order model's CE (0.2666 nats/token) exceeds the uniform mixture CE (same table), consistent with the model not being the order-averaged minimizer. The premise must be tested, not assumed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that transformers with positional encodings minimize expected conditional description length over orderings, E_π[ℓ(Y|Γ_π(X))], rather than the permutation-invariant ℓ(Y|X), making them 'Bayesian in expectation, not in realization.' It derives a Quantified Martingale Violation (QMV) bound claiming O(log n) permutation-induced dispersion, an Expectation-level Decompression Law (EDFL) for Bernoulli predicates, and operational planners (B2T, RoH, ISR) for answer/abstain decisions. Experiments on 3,059 evidence-grounded QA items report logarithmic dispersion, positive Jensen gains from permutation mixtures, a causal dose-response of hallucination to information budget, and a pre-specified audit in which an analytically fixed ISR=1 gate yields near-0% hallucination with ~24% abstention.","tokens_in":12549,"tokens_out":9196,"duration_ms":111770,"significance":"The intended contribution is substantial: a quantitative, information-theoretic account of why order-sensitive models can appear Bayesian on average yet fail on fixed inputs, together with deployable abstention rules. The EDFL lower bound and the QMV upper bound, where proved, are clean and potentially useful. The pre-specified audit is a genuine strength, as is the use of external observables (ground-truth likelihood, accuracy, hallucination rates) rather than the proposed metrics themselves for validation. However, the central explanatory claim — that standard transformers minimize the order-averaged objective — is not established, and the operational gate appears to rely on reversing a necessary-condition inequality. If these gaps can be repaired, the framework would be an important step; as it stands, the paper offers a set of interesting bounds and empirical regularities rather than the resolution claimed in the abstract.","major_comments":[{"comment":"The paper's central premise — that standard transformers minimize E_π[ℓ(Y|Γ_π(X))] over orderings — is not derived. Theorem 3 constructs a tied-weight multi-branch ensemble with a linear averaging head and shows only that this non-standard architecture can realize permutation mixtures. Theorem 4 takes the infimum over the convex hull of such mixtures, not over the parameter class of an ordinary decoder-only transformer. No training theorem shows that gradient-based next-token training converges to the order-averaged I-projection, and no training experiment tests the premise. Without this, the claim that hallucinations are predictable compression failures lacks its stated mechanism. The authors should either prove the minimization claim under explicit assumptions, or reframe the contribution as a conditional framework and validate the premise empirically.","section":"Section 3.3, Theorems 3–4"},{"comment":"Theorem 1, the general QMV bound under Assumption 1, is stated without proof. Appendix A.2 proves only Theorem 2, which relies on the stronger Assumption 2. Since Theorem 1 is the foundational result behind the O(log n) scaling claim, a proof (or a precise proof sketch with all steps) is required. In particular, the step from coordinate total variation to the expectation over pairs of random permutations needs justification; the constant 1/4 and the handling of the sum of coordinate variations are not self-evident.","section":"Section 3.2, Theorem 1"},{"comment":"Lemma 3 states E[HD] = H_n - 3/2 + O(1/n), but a direct calculation for D=|U-V| with U,V i.i.d. uniform on {1,...,n} gives E[H_D] = H_n - 3/2 + (H_n + 1/2)/n - 1/n^2 = H_n - 3/2 + O((log n)/n). The stated O(1/n) is therefore incorrect. This does not change the leading log n behavior in Theorem 2, but the claimed 'explicit constants' and o(1) term need correction. Please also define HD/H_D explicitly, as the notation is ambiguous.","section":"Appendix A.1, Lemma 3"},{"comment":"EDFL as stated is a lower bound: for any event A with posterior mass p and prior mass \\bar q, the expected budget satisfies \\bar\\Delta \\ge KL(Ber(p)||Ber(\\bar q)). This is a necessary condition on the budget needed for a given reliability level. The operational planners, however, treat the inequality as if exceeding the lower bound were sufficient: ISR ≥ 1 is used as a license to answer, and Box 2 reports a 'maximum achievable success' at a given budget. No theorem establishes that a model whose budget exceeds KL(Ber(p)||Ber(\\bar q)) can actually achieve reliability p; the I-projection P* is a constructed distribution, not the model's predictive distribution. The audit results may still be a useful calibration check, but they do not follow from EDFL alone. Please either provide a sufficiency guarantee under additional assumptions or explicitly downgrade the planners to heuristics.","section":"Section 3.4 and Box 1 (B2T/RoH/ISR)"}],"minor_comments":[{"comment":"The statement K_U(y|Γ_π(x)) = L_θ(y|Γ_π(x)) + O(1) is not the coding theorem. The coding theorem relates Kolmogorov complexity to a universal prefix code, not to the log-loss of an arbitrary trained model. This section needs substantial qualification or removal.","section":"Appendix A.6"},{"comment":"The text says 'content is held fixed across dose arms' but the design varies the number of support vs. non-support chunks. Please clarify what is held fixed (question, evidence pool, prompt length) and discuss the potential direct effect of dose on answerability, which the IV strategy does not automatically exclude.","section":"Appendix D"},{"comment":"The notation E[HD] is undefined; it should be E[H_D] where H_D = ∑_{t=1}^{D} 1/t. The current text is confusing.","section":"Appendix A.1"},{"comment":"Model names are inconsistent: 'Qwen2-7B' vs 'Qwen-2-7B-Instruct', 'Llama-3.1-8B' vs 'Llama-3.1-8B-Instruct'. Please standardize.","section":"Throughout"},{"comment":"The claim that a fitted line a + b ln n 'confirms Theorem 1' is overstrong: Theorem 1 gives an upper bound, not an equality, and the Llama R² of 0.515 leaves substantial unexplained variance. Please soften the language and report the comparison of the fitted slope to the theoretical constant rather than only the fit.","section":"Table 1, Section 4.1.1"}],"recommendation":"major_revision","confidential_remarks":"The paper has several strong components, but the central architectural premise is currently a conditional statement about a non-standard ensemble, not a property of standard transformers. Given the paper's framing as a resolution of the Bayesian-in-expectation paradox, this is a load-bearing gap. The EDFL-to-ISR inference also needs a careful reworking. I would not reject the paper outright because the bounds and empirical protocol are potentially salvageable, but the authors need to substantially revise the theory or explicitly change the claims' scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: the paper gives a practical, pre-generation abstention rule for binary adjudication and a clean O(log n) bound on permutation dispersion under a harmonic positional-sensitivity assumption. The empirical effort is real. But the headline explanation — that standard transformers minimize expected description length over orderings, making them 'Bayesian in expectation' — is not established by the theorems, and the validation is weaker than it first appears.\n\nWhat's genuinely new: the QMV bound with explicit constants, the EDFL specialization to Bernoulli events, and the B2T/RoH/ISR planners. The authors are unusually transparent: the ISR=1 threshold is fixed in advance, the audit is pre-specified, and they acknowledge the binary-only scope. The 3,059-item dispersion study and the randomized dose-response are real measurements, and the code/data are pinned.\n\nWhere it frays: the central premise is assumed. Theorem 3 shows that an ensemble-within-the-network with tied weights and an averaging head can realize the permutation mixture. That's a non-standard architecture, and Theorem 4 takes the infimum over that convex hull, not over the parameter class of an ordinary decoder. Nothing shows that next-token training of a standard transformer converges to the order-averaged I-projection. Without that, the 'Bayesian in expectation' claim is conditional on a closure that isn't demonstrated. The O(log n) prediction is tested by fitting a slope and intercept; that can't distinguish logarithmic from other monotone curves. The Jensen-gap experiment confirms Jensen's inequality — a mathematical certainty — rather than a model-specific prediction, so it tells you little about whether the architecture actually minimizes the averaged objective. The audit's reference distribution P for the information budget isn't specified, and there are no baseline detectors, though the authors explicitly say they're doing calibration, not benchmarking — fair enough.\n\nThe paper is honestly written and the limitations are stated. But the load-bearing claim needs either a training experiment or a clear reframing as a property of a specific closed model family. For a reading group, it's a good sparring partner; I'd send it to peer review with major revision expected.","headline":"Useful abstention toolkit; the Bayesian-in-expectation story is not proven for standard transformers.","tokens_in":13000,"tokens_out":3432,"would_cite":false,"duration_ms":40730,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07","68Q30","62B10","94A17"],"pacs":[],"model":"deepseek-v4-flash","headline":"Transformers are Bayesian in expectation but order-sensitive in realization, making hallucination a predictable compression failure.","keywords":["hallucination","permutation invariance","information budget","compression","cross-entropy","abstention","transformer interpretability","calibration"],"falsifier":"Compute, on held-out items, both the measured expected KL budget Eπ KL(P∥Sπ) and the EDFL lower bound KL(Ber(p)∥Ber(q̄)); if a nontrivial fraction of items satisfy the reverse inequality at the claimed reliability, the law is falsified. Alternatively, train a standard transformer on exchangeable data and check whether the dispersion slope b exceeds the O(log n) prediction (e.g., super-logarithmic growth) or whether permutation mixtures fail to improve ground-truth cross-entropy.","tokens_in":12026,"feed_emoji":"🤖","tokens_out":4414,"duration_ms":51266,"temperature":0.7,"pith_summary":"This paper argues that transformers trained with next-token prediction minimize expected conditional description length averaged over evidence orderings, not the permutation-invariant loss that Bayesian symmetry would require. That distinction explains why models can look well-calibrated on average yet flip their answers when evidence chunks are reordered. The authors derive a bound showing order-induced prediction deviations grow logarithmically with context length, and an information-theoretic lower bound (the Expectation-level Decompression Law) showing how much evidence information is needed to reach a target reliability for rare events. They build these into a decision rule that abstains when the measured information budget falls short, and report that a fixed threshold achieves near-zero hallucination at about 24% abstention on a held-out audit. A sympathetic reader would care because the framework turns hallucination from an unpredictable failure into a quantity that can be computed before generation.","feed_headline":"Hallucination risk becomes computable from evidence order","feed_subtitle":"A fixed information-sufficiency gate cuts hallucination to near zero by abstaining on one in four unsure answers.","key_machinery":"The Quantified Martingale Violation (QMV) bound is the first pillar: it converts adjacent-rank positional sensitivity into an expected-absolute-deviation bound using a logistic Lipschitz lemma and a harmonic-distance identity, yielding the O(log n) law. The second pillar is the Expectation-level Decompression Law (EDFL), which applies convexity of KL divergence plus data processing to Bernoulli predicates, giving a closed-form lower bound on the information budget needed for reliability. The third is the Information Sufficiency Ratio (ISR), the ratio of the measured information budget to the bits-to-trust target; the gate ISR≥1 triggers answering, otherwise abstention. The theory also includ","core_discovery":"The central claim is that the apparent paradox of LLMs being both Bayesian and permutation-violating is resolved by recognizing the training objective: models minimize Eπ[ℓ(Y|Γπ(X))], the cross-entropy averaged over all orderings of the evidence, rather than the permutation-invariant ℓ(Y|X). As a result the predictive distribution is a uniform mixture over permutations, which is exchangeable in expectation but not in any fixed realization. The paper proves a Quantified Martingale Violation bound showing that, under bounded total variation of logit changes under adjacent swaps, the expected absolute residual grows as O(log n) in the harmonic regime. It then proves the Expectation-level Decomp","pith_inferences":["If the minimization-over-orderings premise holds for standard transformers, the same EDFL bound should apply to any verifiable binary predicate (unit tests, rubric checks), not just support/refute; that would give a generic safety layer for structured generation.","The O(log n) law suggests that position-sensitivity will grow only slowly with context length, but the constant b may differ by pretraining objective; comparing b across next-token vs. permutation-robust training objectives would test whether the law is universal or architecture-specific.","A training-time regularizer that penalizes the variance of logits across permutations is proposed; a natural extension is to test whether this regularizer improves the calibration of the ISR gate on long-context tasks.","The paper leans on binary adjudication because the bounds are tightest there; extending EDFL to multi-class via one-vs-rest is mentioned but not evaluated, leaving open whether the same information-budget thresholds transfer."],"forward_implications":["Order-induced dispersion in binary adjudication should scale as a+b log n across model families, with the slope constant capturing architecture-specific positional sensitivity.","Uniform permutation mixtures are near-optimal in cross-entropy, so averaging predictions over reorderings improves ground-truth likelihood without learning new weights.","Randomized variation in the number of supporting evidence chunks should causally move hallucination rates by roughly 0.13 per additional nat of information budget.","A pre-specified ISR=1 abstention rule can hold hallucination to near-zero at moderate abstention rates, making the threshold an ex-ante operating point rather than a tuned hyperparameter."],"fun_headline_variants":["Evidence order predicts hallucination risk in verifiers","Proven bound: order dispersion drives verifier hallucination","Fixed gate from evidence order: <1% hallucination on audit","Order-sensitive LLM trust: new bound yields abstain gate","Order-derived gate: <1% hallucination, 1-in-4 abstain"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes that standard next-token training actually drives the model toward minimizing expected conditional description length averaged over evidence orderings; if gradient descent does not converge to that permutation-mixture optimum, the expectation-realization gap is not explained by the theory.","fun_headline_variants_meta":{"raw":{"variants":["Evidence order predicts hallucination risk in verifiers","Proven bound: order dispersion drives verifier hallucination","Fixed gate from evidence order: <1% hallucination on audit","Order-sensitive LLM trust: new bound yields abstain gate","Order-derived gate: <1% hallucination, 1-in-4 abstain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001145,"raw_usage":{"total_tokens":4617,"prompt_tokens":808,"completion_tokens":3809,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":3721}},"tokens_in":552,"tokens_out":3809,"duration_ms":40700,"temperature":1.0,"reasoning_tokens":3721,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:56:13.386391+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute, on held-out items, both the measured expected KL budget Eπ KL(P∥Sπ) and the EDFL lower bound KL(Ber(p)∥Ber(q̄)); if a nontrivial fraction of items satisfy the reverse inequality at the claimed reliability, the law is falsified. Alternatively, train a standard transformer on exchangeable data and check whether the dispersion slope b exceeds the O(log n) prediction (e.g., super-logarithmic growth) or whether permutation mixtures fail to improve ground-truth cross-entropy.","supporting_citations":[],"review_version":1}