Pith. sign in

REVIEW 2 major objections 4 minor

Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

T0 review · 2 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read A language model's probabilities over phrases can be turned into calibrated, auditable posterior estimates over meaningful states when a semantic map and held-out calibration are fixed in advance.

desk verdict A serious, unusually honest attempt to turn LLM continuation probabilities into calibrated posteriors over declared states, but the theoretical recovery guarantee is not certified by the experiments — the empirical claims rest on a constrained grammar and finite-design diagnostics. read the letter →

arxiv 2607.17447 v2 pith:5VISVFQG submitted 2026-07-20 stat.ME cs.CLmath.STstat.MLstat.TH

classification stat.MEcs.CLmath.STstat.MLstat.TH MSC 62F1562G0562H12
keywords semanticuncertaintylanguage-modelprobabilitiesposteriorcalibrationinverseproblemscompositionaldatasemiparametricinferenceconformalpredictiongrouping
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a language model's probabilities over complete verbal continuations—rather than the numbers it prints when asked for confidence—can be turned into a calibrated posterior over meaningful states, if one fixes in advance how phrases map to states and validates the map against a reference posterior. The method is a 'semantic map': a prespecified grouping of phrase probabilities into a lexical composition, followed by an inverse calibration fitted on held-out scenarios and expressed in additive log-ratio coordinates. The paper derives error bounds, identifiability conditions, presentation-stability bounds, and a sequential-filtering bound for this recovery. Empirically, language-derived probabilities outperform numerical elicitation, recover exact held-out posteriors with nominal conformal coverage, are stable under paraphrase, and move in the right direction when evidence changes. What is at stake is replacing unverified confidence statements with an auditable statistical measurement.

What carries the argument

The carrying object is the semantic map, defined as a prespecified measurable coarsening φ_u of the continuation space into K declared semantic states plus an unexpressed symbol ⊥, followed by a calibrated inverse ψ_u fitted in additive log-ratio (alr) coordinates. The estimator is the composition bπ_u = alr^{-1}∘ψ_u∘alr∘p_u∘C_u, where p_u is the normalized pushforward of the language law under φ_u. The map's load-bearing property is the lower modulus c_u: the minimal separation the language channel preserves between different reference posterior log-ratios. It converts a bounded observation error δ into a bounded recovery error 2δ/c_u, and its reciprocal controls how strongly inversion ampl

What would settle it

Find a pair of evidence values inside the declared operating domain whose exact reference posteriors differ materially but whose lexical compositions under the same prompt and semantic map are (nearly) identical, so the fitted affine map's smallest singular value is effectively zero on that pair. Alternatively, apply the published semantic map and calibration to a different language-model checkpoint or to unrestricted free-text responses and check whether held-out conformal coverage falls below the nominal level or held-out error exceeds the reported bounds.

Watch

Extended reading notes

Core claim

The central claim is that the observable, prompt-dependent distribution over verbal responses can serve as a measurement of a reference posterior over a finite set of declared states. The construction fixes a semantic coarsening φ_u (complete continuations mapped to states or to an unexpressed category), forms the normalized lexical composition p_u = pushforward of Q_u under φ_u, and calibrates an inverse map ψ_u in alr coordinates to obtain bπ_u = alr^{-1}∘ψ_u∘alr∘p_u∘C_u. Under a positive lower-modulus condition the recovery error is bounded by 2δ/c_u (or δ/κ_u in the affine case), so observability of the language channel is exactly what protects against amplified noise. The paper's experi

Load-bearing premise

The load-bearing premise is that the forward map from reference posterior log-odds to lexical log-odds is identifiable and stable—has a positive lower modulus and is well approximated by the fitted affine class—on the intended operating domain, and the experiments certify this only on a sampled design for a constrained three-code response grammar rather than for unrestricted free text.

Editorial extensions

If this is right

  • Language-model continuation probabilities can be treated as a measurement channel with a declared estimand, making posterior estimates auditable rather than ad hoc.
  • Uncertainty sets with nominal coverage can be attached to language-derived state probabilities when calibration and conformal radii are computed on held-out scenarios.
  • Information-preserving rewording can be validated as a nuisance factor, while reordering that preserves the facts is shown to be consequential.
  • Sequential Bayesian updating becomes defensible when the reference filter is contractive and one-step recursion defects are observable.
  • The same declared-map-plus-calibration template extends to auditing classifications and recommendations beyond probability reporting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The constrained three-code response grammar used in the controlled experiment may be much easier to calibrate than unrestricted free text; using the method on open-ended prose will likely require richer semantic partitions and explicit handling of unexpressed mass, and the reported results do not certify that setting.
  • Because the forward map is certified only as affine on the sampled design, posterior estimates near the simplex boundary or outside the calibrated region could carry larger inversion error than the average held-out metrics suggest.
  • The dependence on a frozen fitted language distribution implies that model updates or prompt-service changes invalidate the calibration; deployed systems would need ongoing recalibration rather than a one-time semantic map.
  • The entropy contrast (raw word entropy nearly uncorrelated with posterior entropy, calibrated entropy correlated around 0.74) suggests that using raw token entropy as an uncertainty measure is unsupported; the semantic map plus calibration is the effective bridge.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper introduces a 'semantic map': a prespecified statistical bridge from an LLM's observable continuation probabilities over verbal responses to a reference posterior over a finite set of declared states. The construction separates the reference experiment from the language experiment, defines a lexical composition p_u through semantic coarsening φ_u and an unexpressed-mass category ⊥, and proposes held-out calibration of an inverse map ψ_u to estimate the reference posterior. The theoretical part derives an approximation decomposition (Theorem 1), a presentation-stability bound (Theorem 2), an inverse-recovery bound based on a lower-modulus observability condition (Theorem 3), and a sequential-filtering error bound (Theorem 4). The empirical part compares continuation-derived lexical probabilities with numerically elicited probabilities on market text, and evaluates posterior recovery and conformal coverage in a controlled three-state experiment with two fitted LLMs. The paper is unusually careful about prespecification, scenario clustering, and limitations, but the central empirical certification is confined to a constrained enum response grammar and does not certify the key observability condition of Theorem 3.

Significance. If the central claim held as stated, the paper would be a substantial contribution: it provides a declared estimand, a decomposition of semantic, probability-mass, and inverse-recovery error, and a held-out validation protocol for converting LLM token probabilities into auditable state posteriors. The strengths are genuine: the semantic coarsening is prespecified, calibration and testing use disjoint scenarios, uncertainty is clustered by scenario, the experiments are replicated across two fitted models, and the limitation statements are unusually explicit. However, the significance is currently limited by a gap between the advertised theorem-driven recovery guarantee and the empirical evidence, which establishes only finite-design, average-case recovery under a constrained grammar. The contribution is therefore promising but needs reframing before the abstract-level claims are supported.

major comments (2)
  1. [§3.2, 'Observability interpretation'; §2.2.3, Theorem 3] The central recovery guarantee of Theorem 3 is conditional on a positive lower modulus c_u. The paper's own diagnostic reports that the empirical pairwise lower modulus is 'effectively zero' for GPT-4.1-mini and 'zero at machine precision' for GPT-4o-mini, with affine residual RMSEs of 7.62 and 7.35 in lexical log-ratio units. If the pairwise modulus on the sampled design is zero, there exist reference posteriors that are indistinguishable from the mean lexical measurement, and Theorem 3's bound 2δ/c_u is vacuous. The positive bootstrap lower bounds on the smallest singular value (1.686 and 1.600) are properties of the fitted affine matrix, not of the true forward regression h_u, and the large residuals show the affine approximation is poor. Consequently, the abstract's claim that the method 'recover[s] held-out posteriors with valid uncertainty coverage' as an 'auditable posterior estim
  2. [§2.1.1 (constrained law Q^G_u); §3.2; Abstract] The general theorems are stated for the unrestricted free-text law Q_u, but the experiments evaluate only the constrained enum grammar Q^G_u with three enum codes plus an explicit residual, under a completeness threshold of 0.999999. The paper acknowledges this in Section 2.1.1 and Section 3.2.1, but the abstract and several conclusions state without qualification that 'language-derived probabilities outperform printed numerical probabilities, recover held-out posteriors with valid uncertainty coverage.' This overstates the scope of the empirical certification. The abstract and conclusion should explicitly state that the recovery results are for the constrained response grammar used in the controlled experiment and do not extend to unrestricted free-text continuations.
minor comments (4)
  1. [§2.2.3, Theorem 3] The theorem assumes existence of the argmin b̂ℓ but does not state a compactness or attainment condition. The appendix proof mentions compactness of L; this assumption should be stated in the theorem itself.
  2. [Table 7] The split-conformal coverage values 0.94 and 0.90 are reported on 100 test scenarios without a confidence interval. With 100 scenarios, the binomial standard error is about 0.024–0.030, so nominal coverage is plausible but not tightly certified; this should be noted or an interval provided.
  3. [§3.2, entropy association] The raw entropy correlation r ≈ −0.001 and calibrated posterior entropy r = 0.742 are reported without confidence intervals. Given that these are descriptive and based on a modest number of scenarios, an interval or at least a statement of uncertainty would avoid overinterpretation.
  4. [Throughout] The distinction between Q_u and Q^G_u is central, but the notation is introduced only in Section 2.1.1 and then not always reused. For readability, consider using Q^G_u consistently in the experimental sections and in Table 4 to remind readers which law is being certified.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: held-out calibration breaks the main reduction; theorems are standard conditional bounds with explicit assumptions.

full rationale

The paper's derivation chain is not circular. The reference posterior π⋆ is defined by a declared reference model, the lexical composition p_u is derived from the LLM's observable continuation probabilities through a prespecified semantic map φ_u, and the inverse ψ_u is fitted on calibration scenarios and tested on untouched scenarios. This held-out design prevents the central 'recovery' claim from reducing to a fitted input: the test posteriors are not used to fit the inverse, and the reported held-out Jensen–Shannon errors and split-conformal coverage are genuine out-of-sample evaluations. The theorems are standard deterministic bounds (inverse-problem stability, Birkhoff contraction) with assumptions stated explicitly; they are not imported from the authors' prior work, and the paper contains no self-citations. The semantic grouping is credited to the external literature (Farquhar et al., Kuhn et al.) and is explicitly a prespecified design choice rather than an ansatz smuggled in via citation. The paper's own limitations—'Finitely many observations cannot determine this uniform infimum' and the empirical pairwise lower modulus being effectively zero—are validity caveats about whether Theorem 3's condition is certified, not evidence that a prediction is equivalent to its input by construction. The observability diagnostic is honestly framed as a property of the fitted affine map on the sampled design, not as a uniform identifiability result. Thus no specific circular step can be exhibited.

Assumptions & free parameters 4 free parameters · 7 assumptions · 2 invented entities

The central claim leans on a user-supplied reference model, a fixed semantic partition, scenario-level exchangeability, and an affine calibration class with positive fitted singular values. The paper is transparent about most of these, but they are assumptions the reader must accept rather than results the paper proves.

free parameters (4)
  • Affine calibration map parameters B, d = not fully reported; σ_min ≈ 1.89 (GPT-4.1-mini), 1.85 (GPT-4o-mini); affine residual RMS ≈ 7.62 and 7.35
    The inverse map ψ_u is restricted to affine maps H_aff and fitted on 240 calibration scenarios per model (Section 2.2.3, Table 4). Recovery claims are for this fitted class.
  • Conformal error radius = 0.1371 (GPT-4.1-mini), 0.0922 (GPT-4o-mini)
    The 90% split-conformal radius is calibrated on a separate 100-scenario conformal-calibration partition per model (Table 7).
  • Semantic partition and candidate phrase set = none (hand-specified before evaluation)
    φ_u and the finite candidate expressions are fixed before evaluation (Definition 4, Section 2.3); different ontologies define different lexical compositions, so the map is not unique across semantic partitions.
  • Completeness threshold for constrained enum law = 0.999999
    A hand-chosen acceptance threshold for the constrained response grammar; both models pass (Section 3.2.1, Table 5).
assumptions (7)
  • domain assumption Reference validity: the evidence generator, ontology, prior, likelihoods, and reference-inference procedure are declared and frozen; the posterior is exact conditional on this model (Assumption 1).
    The target posterior is defined by an evaluator-chosen probability model, not by an external ground truth.
  • domain assumption Observable measurement: only presented evidence, service-supplied token probabilities, and complete-continuation probabilities enter the estimator; if the service hides needed probabilities, the measurement is marked unsupported (Assumption 2).
    In the experiments this reduces to a constrained grammar Q^G_u with three enum codes and a residual bin, not unrestricted free text.
  • domain assumption Design separation: ontology discovery, semantic validation, calibration, model selection, conformal calibration, and final testing use disjoint scenario partitions (Assumption 3).
    Prevents look-ahead, but depends on the authors' self-reported discipline.
  • domain assumption Sampling units and exchangeability: scenarios, not repeated responses, are the sampling units; calibration and future test nonconformity scores are exchangeable within strata (Assumption 4).
    Split-conformal coverage and bootstrap intervals rely on this; the paper acknowledges that distribution shift or inter-cluster dependence would invalidate them.
  • ad hoc to paper In the experiments h_u is assumed affine (H_aff) with a fitted positive smallest singular value; uniform nonlinear observability is not certified.
    Section 2.2.3 and Theorem 3. The paper explicitly states that finite observations cannot determine the uniform lower modulus c_u, and the empirical pairwise lower modulus is zero.
  • standard math Markov-kernel measurability and pushforward existence for the semantic coarsening (Lemma 1).
    Standard measure theory from Kallenberg; not a paper-specific contribution.
  • standard math Filtering theorem assumptions: strictly positive transition, positive likelihoods, and Birkhoff contraction coefficient ρ<1 (Theorem 4).
    Standard filter-stability assumptions from van Handel, Birkhoff, Bushell; not empirically verified in this paper.
invented entities (2)
  • Semantic map (φ_u, ψ_u) independent evidence
    purpose: Bridges LLM continuation probabilities to a reference posterior over declared states.
    Core introduced construction; testable via held-out posterior recovery, conformal coverage, and paraphrase stability. It is a statistical object, not a physical entity.
  • Unexpressed-mass category ⊥
    purpose: Retains probability mass outside the declared semantic partition so it is not silently discarded during normalization.
    Bookkeeping category introduced in Definition 4; useful for diagnostics but not independently falsifiable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Calibrating Semantic Uncertainty from Observable Language-Model Probabilities." pith.science (2026). https://pith.science/paper/5VISVFQG

@misc{pith2026260717447,
  author       = {Pith},
  title        = {Pith review of: Calibrating Semantic Uncertainty from Observable Language-Model Probabilities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5VISVFQG}},
  note         = {Machine review of arXiv:2607.17447}
}
read the original abstract

As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for inference and decision-making. Language models assign probabilities to words, whereas applications require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. We introduce a \emph{semantic map}: a prespecified, testable bridge from probabilities over verbal responses to a posterior over declared finite states. The language distribution remains unrestricted; held-out calibration connects it to a reference posterior. We derive posterior-error bounds and conditions for existence, conditional uniqueness, presentation stability and stable inverse recovery. This distinction matters because language probabilities depend on prompt wording, while the target posterior should not change under information-equivalent rewording. Experiments use professional market text compiled from Federal Reserve economic and financial series, together with controlled simulations having exact posteriors. Across two fitted language models, language-derived probabilities outperform printed numerical confidence, recover held-out posteriors with valid uncertainty coverage, remain largely stable under paraphrase and respond appropriately to altered evidence. \textbf{Prompt engineering optimises a wording-dependent response; robust scientific use requires validated stability of application-relevant meaning.} The proposed map turns semantic uncertainty in generative systems into an identifiable and testable statistical measurement problem and, when its acceptance conditions hold, yields an auditable posterior estimate.

Figures

Figures reproduced from arXiv: 2607.17447 by the authors.

Figure 1
Figure 1. The proposed bridge in the patient-diagnosis example. The evaluator supplies evidence and a prompt (green); the [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Observable lexical measurement. The annotations follow one realization through the construc [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Quantitative observability and inverse recovery. The lower modulus [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Approximate commutation of lexical measurement and Bayesian filtering. The upper path is the [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: Two views of measurement quality in the semantic study. The calibration panels assess whether [PITH_FULL_IMAGE:figures/full_fig_p023_5.png]
Figure 6
Figure 6. Figure 6: Evidence for inverse recovery and finite-sample uncertainty. The left panel tests whether lexical [PITH_FULL_IMAGE:figures/full_fig_p026_6.png]

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.