REVIEW 4 major objections 4 minor 21 references
Beyond Procedure: Substantive Fairness in Conformal Prediction
T0 review · 4 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Conformal prediction should be tuned to equalize prediction-set sizes across groups, not coverage: across four tasks, set-size parity tracks downstream equity while coverage parity tracks disparity, and label-clustered sets deliver the best
desk verdict The theory and the LLM-in-the-loop evaluator are worth engaging, but the headline empirical claim about equalized set size rests on a GEE that conditions on a post-treatment mediator, and the paper's own Table 12 shows how fragile that result is. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three pieces carry the argument. (1) maxROR, the substantive-fairness metric: the maximum, over group pairs, of the ratio of odds ratios comparing each group's downstream decision accuracy under a conformal treatment to a no-set control; values near zero mean the decision aid benefits groups equally. (2) The LLM-in-the-loop evaluator: a large language model plays the role of the human decision-maker, and a logistic generalized estimating equation with an 'adoption' covariate estimates the treatment effects; the paper validates it by reproducing the qualitative ordering of a prior human-subject trial. (3) Theorem 4.1, which bounds the expected set-size disparity between groups by three interp
What would settle it
A human-subject randomized trial on a held-out dataset — not among the three used for validation — that gives participants conformal sets from marginal and coverage-equalizing Mondrian methods and measures per-group decision accuracy: if equalized coverage yields equal or smaller cross-group disparity than marginal sets, the central correlation claim fails. A second check is quantitative: re-running the three validation datasets with human raters and comparing per-group odds ratios to the LLM's maxROR values; large mismatches would show the proxy is not faithful.
Extended reading notes
Core claim
The central claim is that substantive fairness — equal benefit to decision-makers across groups — is driven by set-size parity, not coverage parity. The paper measures substantive fairness as maxROR, the maximum cross-group ratio of odds ratios of downstream decision accuracy under a treatment versus a control with no prediction set, estimated by a logistic GEE that adjusts for task difficulty and the decision-maker's adoption of the set. Across all four datasets, the coverage gap regresses negatively on maxROR and the set-size gap positively, so shrinking the coverage gap increases downstream unfairness while shrinking the set-size gap reduces it. The paper further claims that label-cluster
Load-bearing premise
The load-bearing premise is that the LLM-in-the-loop evaluator, after adjusting for adoption, faithfully reproduces how human decision-makers would use these prediction sets; the paper validates this only as a qualitative ordering on three datasets, and the adoption adjustment itself may remove part of the exact effect being measured.
Editorial extensions
If this is right
- Designers of conformal prediction systems for consequential decisions should monitor and minimize the set-size gap between groups, not the coverage gap; the paper finds these two procedural objectives are in direct tension.
- Equalized coverage — the standard fairness target achieved by group-conditional (Mondrian) calibration — is, on the paper's evidence, actively inequitable: in all four datasets, smaller coverage gaps are associated with larger downstream disparity as measured by maxROR.
- Label-clustered conformal prediction, which pools calibration data across groups inside difficulty-based label clusters, offers the best observed balance between utility and substantive fairness, with low maxROR and high helpfulness to the decision-maker.
- Theorem 4.1's decomposition gives a diagnostic tool: a measured set-size disparity can be attributed to label heterogeneity within clusters, spread across clusters, or intrinsic label-level group differences, pointing to where intervention is needed.
- The LLM-in-the-loop protocol scales substantive-fairness evaluation to datasets and modalities where human trials would be too expensive, enabling routine fairness audits of CP pipelines.
Reading between the lines
- Because 'adoption' — how often the decision-maker chooses a label inside the provided set — is likely a mediator rather than a confounder on the path from set properties to downstream correctness, adjusting for it may absorb part of the treatment effect the paper intends to measure; a mediation analysis treating adoption as an outcome of set design would test how sensitive maxROR is to this choice
- If the correlation generalizes beyond these four tasks, the natural policy consequence — which the paper only gestures at — is that regulators should audit downstream decision outcomes rather than certify per-group statistical guarantees on the prediction sets themselves.
- The V-shaped dependence of set-size disparity on the number of label clusters, with a minimum at K=2, implies a cheap pre-deployment heuristic: tune K over small values on calibration data before committing to a downstream evaluation, since both extremes (global calibration and per-label calibration) inflate disparity.
- A quantitative human calibration of the evaluator — matching per-group error rates, not just the Marginal-versus-Mondrian ordering — would let maxROR magnitudes be read as estimates of true human disparate impact; the paper only claims qualitative alignment, so its numbers are best treated as relative.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies substantive fairness of conformal prediction (CP) in downstream decision-making. It introduces an LLM-in-the-loop evaluator that approximates human use of prediction sets and measures downstream accuracy disparities via a GEE model, reporting a maxROR metric. Theoretically, it derives an upper bound (Theorem 4.1) decomposing the disparity in expected prediction-set size between protected groups into three interpretable components, and uses this to argue that Label-Clustered CP reduces set-size disparity. Empirically, it benchmarks Marginal, Mondrian, Label-Clustered, Group-Clustered, and Backward CP on four datasets (FACET, BiosBias, RAVDESS, ACSIncome). The central empirical claim is that equalized set sizes correlate with improved substantive fairness, whereas equalized coverage is actively inequitable (Section 6.3).
Significance. If substantiated, the paper's central claim would redirect practical design choices in conformal prediction from coverage parity toward set-size parity, a potentially important and actionable insight. The LLM-in-the-loop evaluation framework is a creative and scalable proposal, and the theoretical decomposition of set-size disparity is original and could be useful beyond this paper. The paper is also commendable for providing code, covering four modalities, and including bootstrap uncertainty for maxROR. However, the main empirical conclusion rests on a GEE specification that conditions on a post-treatment mediator (adoption), and on an LLM proxy validated only in qualitative ordering. The theorem statement also contains inconsistencies with its proof. These issues are load-bearing for the headline claims, but they are addressable with reanalysis and reframing, so the appropriate decision is major revision.
major comments (4)
- [§4.2, Eq. (15), Table 12, and Appendix E.1] The GEE model includes adoption_jt as a covariate. Adoption is not a confounder but a post-treatment mediator: whether the LLM's answer lies inside the provided set is itself affected by the CP method and the group, and it is strongly predictive of correctness (Table 11). Conditioning on it estimates a controlled direct effect, not the total effect of the CP method on downstream accuracy that the substantive-fairness definition (Eq. 14) requires. The fragility is visible in Table 12: omitting adoption flips the BiosBias ordering (maxROR_Marginal 27.4% vs. Mondrian 9.9%). The paper's defense that adoption should be held fixed removes exactly the mechanism by which set properties affect downstream equity. Please report unadjusted and adjusted analyses, or use mediation/causal methods, and reframe maxROR accordingly.
- [§6.3, Figure 4] The claim that equalized set size 'strongly correlates' with improved substantive fairness is based on regression lines through five method-level points per dataset, with no correlation coefficients, confidence intervals, or significance tests. With only five points, a single outlier can determine the slope. Bootstrap standard errors for maxROR (Appendix E.3) are not propagated into the correlation analysis. Provide correlation coefficients with uncertainty, or generate more points (e.g., varying K or score-function hyperparameters) to support the claim.
- [Theorem 4.1 vs. Appendix A] The theorem statement does not match the proof. The theorem's term (III) is |∑_y P(Y=y|A=b)(r_{y,a}-r_{y,b})|, but the proof bounds the analogous quantity as ∑_y P(Y=y|A=a)|r_{y,a}-r_{y,b}| (Equations A2 and A10). The absolute value is inside the sum in the proof but outside in the theorem, and the weighting differs. Similarly, term (I) in the theorem uses ϵ_{k,a}, while the proof's derivation (A4) uses ϵ_{k,b}. Term (II) also uses μ_{k,a} in the theorem but the proof derives it with group b. As stated, the theorem is not what is proven. This must be corrected, and the corrected bound should be verified in the numerical study (Figure 7).
- [§6.1, Table 2; §6.3] The LLM-in-the-loop is validated only by qualitative ordering of Marginal vs. Mondrian on three datasets, with large quantitative discrepancies (e.g., RAVDESS Marginal maxROR 11% in the LLM vs. 1.0% in the human study; Mondrian 79% vs. 28%). The subsequent correlation analysis uses maxROR as a quantitative dependent variable. If the LLM overstates or reorders disparities in a way not captured by the three validation points, the RQ3 correlations could be an artifact of the proxy. Please add a sensitivity analysis (e.g., using unadjusted maxROR, or rescaling/ranking-based correlations) and state clearly that the quantitative maxROR values are LLM-specific until further human validation.
minor comments (4)
- [Abstract and §6.2] The abstract says label-clustered CP 'consistently deliver[s] superior substantive fairness,' but Figure 1 and Table 13 show Backward CP outperforms Label-Clustered on FACET and BiosBias in maxROR. The text in §6.2 is more careful ('on average'); please align the abstract with the actual results.
- [§6.4, text before Equation (13)] The sentence 'we decomposed the set size disparity Δa,b (Equation (14))' should refer to Equation (9), where Δa,b is defined. Equation (14) is the group improvement disparity Δt.
- [Appendix A] The proof uses notation (I) and (II) for both the decomposition terms and sub-terms (I)_1, (I)_2, creating confusion with the main text's terms (I), (II), (III). Rename these for readability.
- [§5.2 / Table 1] The calibration-test split details for RAVDESS mention stratification by emotion and gender, but Table 5 shows equal counts; please clarify whether splits are also stratified by speaker identity to avoid data leakage.
Circularity Check
LLM-evaluator validation is tuned to the authors' prior human-ordering result; central theorem and RQ3 correlation retain independent content.
-
fitted input called prediction
[Appendix E.1 (GEE calibration, Table 12); Section 6.1 (RQ1 validation)]
"This counterexample motivates including “adoption” as a covariate in the GEE. ... To make our evaluator reflect the substantive-fairness pattern discovered by Cresswell et al. (2025) while respecting the different experimental design, we treat each task instance as the clustering unit in the GEE Equation (15) with adoption as a covariate."
RQ1's validity claim is that the LLM evaluator 'consistently reproduces the qualitative maxROR ordering' (Mondrian > Marginal) from Cresswell et al. (2025). The paper explicitly chose the GEE specification to produce that ordering: Table 12 shows the no-adoption GEE gives the opposite ordering on BiosBias (maxROR_Marginal 27.4% vs maxROR_Mondrian 9.9%), and the adoption covariate was added 'to make our evaluator reflect' the target pattern. Reporting the resulting match as validation is therefore a design outcome, not an independent confirmation. All downstream maxROR values in RQ2 and RQ3 come from this same evaluator, so the empirical fairness claims inherit this partially circular validation.
full rationale
The theoretical contribution (Theorem 4.1, Equation 13) is a self-contained inequality derived from the law of total expectation and triangle inequalities; it does not fit parameters to reproduce set-size disparity, and it holds for any label-clustered CP. The practical guidelines and RQ3 correlation are empirical: maxROR is measured from the LLM-in-the-loop GEE and correlated with coverage/set-size gaps, and those correlations are not identities. The main circular element is the RQ1 validation step: the evaluator's GEE specification was adjusted (adding adoption) precisely to match the authors' prior human-in-the-loop ordering from Cresswell et al. 2025, then cited as successfully reproducing that ordering. This is a partial fit-to-validation-anchor, and since the same evaluator supplies all maxROR values, it weakens but does not by itself force the Section 6.3 conclusion. I also weighed the adoption-mediator critique (Equation 15 conditions on adoption, a post-treatment variable tied to set membership; Appendix E.1/Table 12 shows specification sensitivity; Section 7 admits causal isolation requires controlling adoption). That is a correctness/estimand concern, not a circularity by construction. Apart from the RQ1 validation fitting, the paper is transparent about the sensitivity and the central correlation is not equivalent to its inputs.
Assumptions & free parameters
free parameters (5)
- Number of clusters K =
K=2 or 3 in main experiments
- Score-function hyperparameters (T, λ, k_reg) =
Listed in Table 7
- Split ratio γ for clustering =
0.3
- Backward CP offset =
Iteratively increased until coverage target is met
- M (number of LLM samples per task) =
Not specified
assumptions (4)
- domain assumption Lipschitz continuity of expected set size as a function of threshold (Equation A8)
- domain assumption Logistic GEE model is correctly specified
- domain assumption LLM is a valid proxy for human decision-makers
- domain assumption Adoption is an outcome-relevant covariate, not a mediator
Cite this review
Pith. "Pith review of Beyond Procedure: Substantive Fairness in Conformal Prediction." pith.science (2026). https://pith.science/paper/FFGV2JMR
@misc{pith2026260216794,
author = {Pith},
title = {Pith review of: Beyond Procedure: Substantive Fairness in Conformal Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/FFGV2JMR}},
note = {Machine review of arXiv:2602.16794}
}
read the original abstract
Conformal prediction (CP) offers distribution-free uncertainty quantification for machine learning models, yet its interplay with fairness in downstream decision-making remains underexplored. Moving beyond CP as a standalone operation (procedural fairness), we analyze the holistic decision-making pipeline to evaluate substantive fairness-the equity of downstream outcomes. Theoretically, we derive an upper bound that decomposes prediction-set size disparity into interpretable components, clarifying how label-clustered CP helps control method-driven contributions to unfairness. To facilitate scalable empirical analysis, we introduce an LLM-in-the-loop evaluator that approximates human assessment of substantive fairness across diverse modalities. Our experiments show that label-clustered CP often provides a favorable balance between utility and substantive fairness, while reducing set-size disparities in line with our theory. Finally, we empirically show that equalized set sizes, rather than coverage, strongly correlate with improved substantive fairness, enabling practitioners to design more fair CP systems. Our code is available at https://github.com/layer6ai-labs/llm-in-the-loop-conformal-fairness.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Bound|(I) 1|: We can rewrite(I) 1 as (I) 1 = KX k=1 P(h(Y) =k|A=a) X y∈Yk P(Y=y|h(Y) =k, A=a)−P(Y=y|h(Y) =k, A=b) E[|C(X)| |Y=y, A=b] (A3) First, fix a cluster k. Since P y∈Yk P(Y=y|h(Y) =k, A=a)−P(Y=y|h(Y) =k, A=b) = 0 , for any constant c, we have X y∈Yk P(Y=y|h(Y) =k, A=a)−P(Y=y|h(Y) =k, A=b) E[|C(X)| |Y=y, A=b] = X y∈Yk P(Y=y|h(Y) =k, A=a)−P(Y=y|h(Y) ...
-
[2]
Bound|(I) 2|: After combining the common terms, we get (I) 2 = KX k=1 P(h(Y) =k|A=a)−P(h(Y) =k|A=b) X y∈Yk P(Y=y|h(Y) =k, A=b)E[|C(X)| |Y=y, A=b] (A5) We can simplify the P y∈Yk P(Y=y|h(Y) =k, A=b)E[|C(X)| |Y=y, A=b]in Equation (A5) as follows. X y∈Yk P(Y=y|h(Y) =k, A=b)E[|C(X)| |Y=y, A=b] = X y∈Yk P(Y=y|h(Y) =k, A=b)E[|C(X)| |Y=y, h(Y) =k, A=b] (becauseY...
2005
-
[3]
adoption
How Label-Clustered CP helps to control|(II)|: Finally, we show how Label-Clustered CP can reduce |(II)| in Equation (A2) compared to group-conditional CP. The term (II) is a weighted sum over y∈ Yof the intra-label set size gap between group a and group b. The differences in expected set size across protected groups after conditioning on the true label d...
2023
-
[4]
Emotional expression in delivery Based on these cues in the audio, which emotion best matches the speaker's vocal expression? Control: You are an expert in emotion classification from audio. Instructions: - Focus on HOW they speak, not WHAT they say - Listen for vocal tone, pitch patterns, and energy - Classify which emotion does the audio convey Choose o...
-
[6]
FACET: Fairness in computer vision evaluation benchmark
3 Gustafson, L., Rolland, C., Ravi, N., Duval, Q., Adcock, A., Fu, C.-Y ., Hall, M., and Ross, C. FACET: Fairness in computer vision evaluation benchmark. InProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20370–20382, 2023. 5, 25 Hardt, M., Price, E., and Srebro, N. Equality of Oppor- tunity in Supervised Learning. InAdvances ...
arXiv 2023
-
[11]
doi: 10.52202/079017-3454. 3 Zwart, P. H. Probabilistic conformal coverage guarantees in small-data settings.arXiv:2509.15349, 2025. 2 11 Beyond Procedure: Substantive Fairness in Conformal Prediction Appendix Table of Contents A Justification that Label-Clustered CP Reduces Set Size Disparity 12 B Conformal Prediction Algorithms 19 B.1 Clustered Conforma...
arXiv 2025
-
[15]
Select exactly one class name from this list
FACET You are given an image of a person and a list of possible classes: [{all options}]. Select exactly one class name from this list. Respond with only the class name. Do not provide explanations or additional text. image attached The following shorter list contains classes that are more likely to be correct: [{options}] Based on the image, which class ...
-
[16]
Given a short biography, select the correct occupation from the provided list: [{all options}]
BiosBias You are an HR expert. Given a short biography, select the correct occupation from the provided list: [{all options}]. Respond with only the occupation name from this list. Do not explain your reasoning. Biography: {prompt} The following shorter list contains occupations that are more likely to be correct: [{options}] Based on the biography above,...
Show all 21 references
-
[17]
RA VDESS 28 Beyond Procedure: Substantive Fairness in Conformal Prediction You are an expert in emotion classification from audio. Instructions: - Focus on HOW they speak, not WHAT they say - Listen for vocal tone, pitch patterns, and energy - Classify which emotion does the a...
-
[20]
Pitch patterns (high/low, rising/falling)
-
[21]
Emotional expression in delivery Based on these cues in the audio, which emotion best matches the speaker's vocal expression? Choose from: [{all options}]
-
[22]
adoption
ACSIncome You are a labor economics expert. Given a structured demographic and employment profile from a census survey, select the correct income bracket from the provided list: [{all options}]. Respond with only the income bracket label. Do not explain your reasoning. Profile...
2025
-
[118]
1 10 Beyond Procedure: Substantive Fairness in Conformal Prediction Office of the Comptroller of the Currency
Springer Science & Business Media, 2012. 1 10 Beyond Procedure: Substantive Fairness in Conformal Prediction Office of the Comptroller of the Currency. Fair lending,
2012
-
[2016]
Fair prediction sets through multi-objective hyper- parameter optimization.Machine Learning, 114(1):27,
1 Garcia-Galindo, A., Lopez-De-Castro, M., and Armananzas, R. Fair prediction sets through multi-objective hyper- parameter optimization.Machine Learning, 114(1):27,
-
[2018]
and Guestrin, C
25 Chen, T. and Guestrin, C. XGBoost: A Scalable Tree Boost- ing System. InProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794, 2016. ISBN 9781450342322. doi: 10.1145/2939672.2939785. 6, 25 Cresswell, J. C. Trustworth...
2016
-
[2019]
6, 25 Ding, F., Hardt, M., Miller, J., and Schmidt, L
doi: 10.18653/v1/N19-1423. 6, 25 Ding, F., Hardt, M., Miller, J., and Schmidt, L. Retiring Adult: New Datasets for Fair Machine Learning. In Advances in Neural Information Processing Systems, vol- ume 34, pp. 6478–6490, 2021. 6, 25 Ding, T., Angelopoulos, A., Bates, S., Jordan...
2021 doi
-
[2020]
Qwen2.5-VL Technical Report
6, 25 Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y ., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y ., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J. Qwen2.5-VL...
2025 arXiv
-
[2022]
M.Bayesian learning for neural networks, volume
3 Neal, R. M.Bayesian learning for neural networks, volume
-
[2024]
Accessed: 2026-01-26
URL https://openai.com/index/gpt-4 o-mini-advancing-cost-efficient-intellige nce/. Accessed: 2026-01-26. 6 OpenAI. Gpt-4o audio model, 2026. URL https://plat form.openai.com/docs/models/gpt-4o-audio -preview. Accessed: 2026-01-26. 6 Radford, A., Kim, J. W., Hallacy, C., Ramesh...
2026 arXiv
-
[2025]
3 Gauthier, E., Bach, F., and Jordan, M. I. Backward con- formal prediction. InAdvances in Neural Information Processing Systems, volume 38, 2025. 3, 19 Gibbs, I., Cherian, J. J., and Cand `es, E. J. Conformal prediction with conditional guarantees.Journal of the Royal Statist...
2025 arXiv
-
[2026]
Accessed: 2026-01-25
URL https://www.occ.treas.gov/topics /consumers-and-communities/consumer-prote ction/fair-lending/index-fair-lending.htm l. Accessed: 2026-01-25. 2 OpenAI. GPT-4o mini: advancing cost-efficient intelligence,
2026
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.