{"id":"b31746ef-fc1e-4576-8f73-d53615d5e70e","arxiv_id":"2506.07469","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Valid prediction intervals for individual treatment effects from large RCTs are trivial unless response rates are extreme, and sharp pmf bounds are given by sums of Fréchet cell bounds.","lead":"This paper characterizes what a large randomized trial can and cannot reveal about how a treatment affects a specific individual. It shows that in a binary treatment and outcome setting, the observed data are often uninformative: the only valid prediction interval is the trivial one, unless response rates are extreme.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 12's sharpness proof is not valid as printed: the upper-bound construction uses misindexed residual marginals and an ill-defined denominator, so the central sharp-bounds claim is unproven unless repaired.","rationale":"The paper's advertised contributions are the sharp pmf bounds (Theorem 12) and the binary/continuous prediction-interval characterizations. The binary interval characterization is sound and clean, and the cdf bounds are taken from earlier work; the genuinely new piece is Theorem 12. The reader correctly identified that the lower-bound proof essentially works but the upper-bound construction contains an indexing error. I checked the construction and found that the residual-marginal formula uses the wrong cells: for row i, a J2-filled cell is at column i-delta, not i+delta; for column j, a J1-filled cell is at row j+delta, not j-delta. Additionally, J1 and J2 are not a partition when ties occur, making s potentially zero with a nonempty submatrix. These are concrete defects in the proof, not mere presentational slips. Because the theorem can be tested by finite-dimensional linear programming, I do not regard the statement as disproven; the concern is that the proof as printed does not establish sharpness. That is exactly what a conditional verdict should capture. I would not change the reader's CONDITIONAL verdict: the flaws are repairable, the statement is likely true (equivalently, the maximal and minimal coupling results for P(X=Z)), and the rest of the paper's interval results stand. The reader's stated weakest assumption, known marginals and no dependence information, is a modeling caveat rather than the proof defect, so my agreement is partial.","tokens_in":21619,"tokens_out":15915,"duration_ms":176702,"concrete_test":"Fix finite marginals (e.g., several random p,q vectors on 4-point supports) and solve the linear program over couplings: maximize and minimize P(Y1-Y0=delta) subject to row sums p and column sums q. Compare the extreme values to sum L_i and sum U_i in Eq. (5). If any gap appears, Theorem 12 is false. Independently, re-derive the submatrix filling formula with corrected indices: subtract P(Y0=i-delta) from row i when i-delta is in J2, and subtract P(Y1=j+delta) from column j when j+delta is in J1, using a disjoint partition of ties. Verify that the corrected formula yields a valid joint distribution for all tested supports; if it does, the theorem can be repaired and the conditional verdict stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is Theorem 12/Eq. (5): sharp bounds on P(Y1-Y0=delta) are obtained by summing per-cell Frechet bounds. The lower-bound half is essentially correct: at most one shifted-diagonal cell can have a positive Frechet lower bound, and Strassen's theorem gives a coupling avoiding all such cells. The upper-bound half is not established. In the construction after Figure 7, J2 is defined by j in J2 when the upper bound is achieved by P(Y0=j), i.e. the filled cell is (j+delta,j). But the submatrix filling formula subtracts I((i+delta) in J2) P(Y0=i+delta) from row i; a J2-filled cell in row i would lie at column i-delta, not i+delta. Symmetrically, the column residual subtracts I(j-delta in J1) P(Y1=j-delta), while a J1-filled cell in column j would lie at row j+delta. So the residual marginals are computed with the wrong cells, and the displayed product does not generally produce a table with marginals P(Y1), P(Y0). The denominator s = 1 - sum_{J1} P(Y1=i) - sum_{J2} P(Y0=j) double-counts cells where equality puts i in J1 and i-delta in J2; when such ties occur and a submatrix remains, s can be zero or negative, so the formula is undefined. The text also permutes rows for J2 although J2 is a set of Y0-column indices. Consequently, sharpness of the upper bound is unproven as printed. This matters because Theorem 12 is the paper's main new result and is used in Corollary 15 and Section 5; the flaw is in the proof, not obviously in the statement, and is likely repairable via a maximal-coupling argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies what can be learned about individual treatment effects (ITE) from a large randomized experiment when only the marginal distributions of the potential outcomes are identified. In the binary outcome model it characterizes the valid (1−α) prediction intervals for the ITE as a function of the two response probabilities, including conditions for a trivial interval, singleton intervals, and non-negative/non-positive intervals. For continuous and ordinal outcomes it derives conservative intervals, points that must be included in any valid interval, and conditions under which zero must be included, using the sharp cdf bounds of Fan–Park and Zhang–Richardson. Its main new result is Theorem 12, which claims sharp bounds on the pmf of the ITE for discrete outcomes: P(Y1−Y0=δ) is bounded by the sum of per-cell Fréchet bounds. The paper concludes with a discussion contrasting ATE confidence intervals with ITE prediction intervals and a synthetic example where the ATE is nonzero while {0} is the only valid 95% prediction interval.","tokens_in":22008,"tokens_out":18250,"duration_ms":168687,"significance":"The results are useful and mostly correct. The binary-outcome characterization is clean and parameter-free: it depends only on the two marginal response probabilities and the nominal level, and it makes the distinction between Fisher and Neyman nulls concrete. Theorem 12, if established, is a valuable contribution because it gives an extremely simple sharp bound on the ITE pmf from marginals alone, extending earlier cdf bounds. The proof strategy is appropriate: the lower bound uses Strassen's theorem, and the upper bound attempts an explicit coupling construction. There are no fitted parameters, and the sharp bounds are checked against external Fréchet/Strassen benchmarks. However, the upper-bound half of Theorem 12 is not proven as printed, and a second rigor issue appears in the quantile definitions of Section 3.2; these need repair before the paper can be accepted.","major_comments":[{"comment":"The upper-bound sharpness construction is invalid as printed. The set J2 is defined over columns j (with U_{j+δ}=P(Y0=j)), but the residual row factor subtracts I((i+δ)∈J2)P(Y0=i+δ); a J2-filled cell in row i lies at column i−δ, so the correct indicator is I(i−δ∈J2) and the correct subtracted mass is P(Y0=i−δ). The column factor has the symmetric error: I(j−δ∈J1)P(Y1=j−δ) should be I(j+δ∈J1)P(Y1=j+δ). The denominator s=1−∑_{J1}P(Y1=i)−∑_{J2}P(Y0=j) double-counts tied cells in which U_i=P(Y1=i)=P(Y0=i−δ); in such cases s is too small by the tied mass and can be zero or negative even when an (N1−n1)×(N2−n2) submatrix remains. Concretely, with p=(0.2,0.3,0.5), q=(0.5,0.3,0.2), δ=1, the printed row factor for i=0 equals 0.2−0.3<0. Since this construction is the only argument for sharpness of the upper bound, Theorem 12 is not proven as printed; the statement itself appears correct, and a residual-marginal argument (fill the diagonal cells with U_i, then complete the remaining transportation problem) would repair the proof.","section":"Section 5.1, Theorem 12, Eq. (5)"},{"comment":"The 'must include' claim is not precisely stated. The quantities R'_0 and R'_1 are defined as maxima of sets of the form {ℓ:P(Y0>ℓ)>α}, which need not have a maximum; for a two-point distribution with P(Y0=0)=0.6, P(Y0=1)=0.4 and α=0.3, the set {ℓ:P(Y0>ℓ)>α} is (−∞,1), so R'_0 would be undefined. The intended objects are suprema (or quantiles defined with ≤/≥), and the proof's assertion that the minimum of the two tail probabilities is 'greater than α' is not true in general. The claim is plausibly correct after replacing the definitions, but as written it is not a theorem.","section":"Section 3.2"}],"minor_comments":[{"comment":"The paper switches between p=P(Y=0|D=0), q=P(Y=0|D=1) in Section 4 and the response probabilities P(Y=1|D=j) used in Section 2 and Appendix A; because the two are complementary, the pmf formulas in Section 4 are easy to misread. It would help to state the convention explicitly once.","section":"Table 1 / Section 4"},{"comment":"The final displayed inequality in the proof is written as P((Y1−Y0)∈[L1,R1])≥1−α, but it should be P((Y1−Y0)∈[L1−R0,R1−L0])≥1−α.","section":"Section 3.1, proof of Eq. (1)"},{"comment":"The proof of Proposition 14 starts the contradiction with 'i≠j, k≠l'; the intended assumption should be i≠k and j≠l, matching the statement that at most one lower bound is nonzero when the row and column indices are both distinct.","section":"Proposition 14"},{"comment":"The sentence 'Let P(Y1=i,Y0=k)=0 for any i≠j or k≠j−δ' should read 'i≠j and k≠j−δ' (or equivalently, all cells outside row j and column j−δ); as written the sentence contradicts the construction that follows.","section":"Section 5.1, Case 2"},{"comment":"Figure 7 is difficult to interpret: the permuted row and column labels do not clearly implement the definitions of J1 and J2, and the caption does not state which entries are the filled U_i cells; the figure should be redrawn after the proof is repaired.","section":"Figure 7"},{"comment":"The synthetic example concludes a 95% prediction interval from estimated marginals; because the formal results require the true marginal distributions, the example should acknowledge that the conclusion is approximate or use a finite-sample construction.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"The main issue is the defective upper-bound proof of Theorem 12; if the authors repair the construction (e.g., by filling diagonal cells with U_i and then completing the residual marginals), the paper should be acceptable. The statement itself is likely correct, and the lower-bound proof is sound. There is no circularity concern: the bounds are derived from Fréchet/Strassen results and match external benchmarks. The scope is appropriate for a statistics journal; no code or data is required for this theoretical paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The binary-outcome prediction interval taxonomy in Sections 2.3–2.6 is the cleanest part of the paper and, as far as I know, new. The conditions for trivial, singleton, and one-sided intervals are correct, and the figures make the dependence on the two response probabilities clear. The Fisher-versus-Neyman discussion in Section 6 is also thought-provoking and worth engaging with. The lower-bound half of Theorem 12, via Strassen's theorem, is elegant and works.\n\nThe main problem is the upper-bound sharpness construction in Section 5.1. The stress-test note is right: the submatrix filling formula subtracts the wrong cells from the row and column residuals. J2 is a set of column indices, but the formula subtracts I((i+delta) in J2) P(Y0=i+delta) from row i, while a J2-filled cell in row i would sit at column i-delta. Symmetrically, the column residual subtracts based on J1 at the wrong offset. The denominator s also double-counts cells that belong to both J1 and J2, so it can be zero or negative; the product formula is not a valid joint distribution as written. There is also a typo saying J2 rows are permuted when J2 is a set of columns, and minor notation slips in Proposition 14 and Table 1. None of this makes me think Theorem 12 is false; the statement is plausible and likely provable by a maximal-coupling argument or a corrected explicit construction. But the paper's main new result currently lacks a valid proof of sharpness.\n\nThe continuous-outcome section is straightforward and honestly acknowledges its similarity to Lei and Candès. The strong assumption—known marginals, no dependence information—is clearly stated, so I do not count that as a flaw.\n\nThis is a paper for causal inference methodologists and statisticians working on partial identification. The binary taxonomy alone is a solid contribution, and the pmf bounds question is important. I would send it to peer review, but the authors need to fix the upper-bound construction before publication. The lower bound and the binary sections are publishable as is; Theorem 12 needs repair.","headline":"A genuinely useful binary-outcome taxonomy plus a plausible sharp pmf bound whose upper-bound proof has a real indexing error, so the central new claim is unproven as printed.","tokens_in":22562,"tokens_out":2597,"would_cite":false,"duration_ms":26899,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that in large randomized experiments the individual treatment effect is only partially identified, and characterizes exactly which prediction intervals can be guaranteed and how sharply the ITE distribution can be bounded…","keywords":["individual treatment effect","prediction interval","partial identification","sharp bounds","Fréchet bounds","potential outcomes","randomized experiment","treatment effect distribution"],"falsifier":"Take any two discrete marginal distributions and solve the linear program over the joint probability table to minimize and maximize $\\sum_i P(Y_1=i,Y_0=i-\\delta)$ subject to the marginals; if the optimum falls outside the interval in Equation (5), the sharpness claim fails. In the binary case, one can similarly enumerate $t\\in[\\max\\{0,p+q-1\\},\\min\\{p,q\\}]$ and check whether any of the six candidate intervals has coverage at least $1-\\alpha$ whenever $p$ and $q$ both lie between $\\alpha$ and $1-\\alpha$.","tokens_in":21405,"feed_emoji":"📊","tokens_out":6583,"duration_ms":69846,"temperature":0.7,"pith_summary":"This paper asks what a large randomized trial can and cannot tell us about a single individual's treatment effect, when the joint distribution of the two potential outcomes is allowed to be anything consistent with the observed outcomes in each treatment arm. It proves that in the binary-outcome setting, whenever both response probabilities lie strictly between $\\alpha$ and $1-\\alpha$, the only valid $(1-\\alpha)$ prediction interval for the individual treatment effect is the entire range $[-1,1]$. It also derives sharp bounds on the probability mass function of the ITE for discrete outcomes, expressed as sums of Fr\\'echet cell bounds, and shows these bounds are attainable. These results matter because they make precise the limits of individualized inference: an average treatment effect can be precisely estimated while the individual effect remains largely unknowable.","feed_headline":"Only the trivial interval can predict individual effects in typical RCTs","feed_subtitle":"When both arms' response rates sit between 5% and 95%, no narrower interval is guaranteed valid.","key_machinery":"The carrying object is the collection of all couplings of the two potential-outcome marginals $P(Y_1)$ and $P(Y_0)$, the only part of the joint distribution identified under randomization. The proofs use Fr\\'echet cell bounds on individual joint probabilities $P(Y_1=i,Y_0=j)$, sum these cell-wise bounds to bound $P(Y_1-Y_0=\\delta)$, and then use Strassen's theorem for finite sets to construct joint distributions that attain the summed lower and upper bounds. In the binary setting, all couplings are parameterized by a single variable $t=P(Y_0=1,Y_1=1)$, which makes the six possible prediction intervals easy to check.","core_discovery":"The central discovery is that with only the marginal outcome distributions identified from a randomized experiment, the ITE inference problem reduces to a coupling problem, and sharp answers are available. For discrete potential outcomes with fixed marginals, the sharp bounds on $P(Y_1-Y_0=\\delta)$ are\n$$\\left[\\sum_i \\max\\{P(Y_1=i)+P(Y_0=i-\\delta)-1,0\\},\\;\\sum_i \\min\\{P(Y_1=i),P(Y_0=i-\\delta)\\}\\right],$$\nwith both endpoints attainable by some joint distribution compatible with the marginals. In the binary case, the paper gives a complete characterization: the only valid prediction intervals can be trivial, a singleton, or one of the two unit-length intervals, depending on the two response probabilities, and when both response probabilities are in $(\\alpha,1-\\alpha)$ the only valid interval is $[-1,1]$. The paper also shows that the Fisher null and the Neyman null can appear to conflict: a prediction interval of $\\{0\\}$ can be valid even when the average treatment effect is nonzero and its confidence interval excludes zero.","pith_inferences":["The sum-of-Fr\\'echet-bounds formula is a general fact about the difference or sum of two discrete random variables with fixed marginals, so it can be reused outside causal inference whenever only marginal distributions are known.","In any realistic finite sample the marginals are estimated, so strictly valid finite-sample prediction intervals require accounting for estimation error; the large-sample results here set a lower envelope, not a plug-in recipe.","The results clarify what data would be needed to escape the trivial interval: any information about dependence between $Y_0$ and $Y_1$, such as a cross-over design or a rank-preserving assumption, since marginal data alone cannot provide it.","They also support a substantive policy point: aggregate evidence of a nonzero average effect does not imply that individualized treatment rules have detectable individual-level benefits, so decisions need to weigh external assumptions about effect heterogeneity."],"forward_implications":["In the binary-outcome case, if both response probabilities lie in $(\\alpha,1-\\alpha)$, the only valid $(1-\\alpha)$ prediction interval is $[-1,1]$, so the trial data alone carry no nontrivial information about the individual treatment effect.","The singleton $\\{0\\}$ is a valid prediction interval exactly when the sum of the two less-common observed outcome probabilities is at most $\\alpha$; this can happen even when the average treatment effect is nonzero, so Fisher and Neyman nulls can appear to disagree.","For continuous outcomes, any valid prediction interval must include the quantile-difference points $R'_1-L'_0$ and $L'_1-R'_0$, and the union-bound interval $[L_1-R_0,R_1-L_0]$ is always valid.","The sharp pmf bounds of Theorem 12 apply to any discrete potential outcomes, including ordinal outcomes, so marginal data alone identify an exact range for $P(ITE=\\delta)$ even though the ITE itself is not identified."],"supporting_citations":[{"why":"Introduces the four patient types (NR, HE, HU, AR) that organize the binary treatment and outcome model.","marker":"Copas (1973)"},{"why":"Supplies sharp bounds on the distribution of treatment effects using copulas, the starting point the paper extends and corrects.","marker":"Fan and Park (2010)"},{"why":"Provides the sharp cdf bounds for $Y_1-Y_0$ used in the continuous and ordinal sections, including the zero-inclusion conditions.","marker":"Zhang and Richardson (2024)"},{"why":"States Strassen's theorem for finite sets, which the paper uses to prove attainability of the lower bounds in Theorem 12.","marker":"Koperberg (2024)"},{"why":"Gives the conformal inference prediction interval for ITE that the paper compares with its conservative continuous-outcome interval.","marker":"Lei and Candès (2021)"}],"fun_headline_variants":["Sharp bounds on individual treatment effect prediction intervals","ITE prediction often trivial: sharp bounds from marginals","Individual effects: when can we predict? Sharp bounds answer","Fisher and Neyman nulls conflict in individualized inference","RCT data only yields trivial ITE intervals in common cases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The joint distribution of the potential outcomes is assumed completely unspecified beyond its marginals, and the large-sample analysis treats those marginals as exactly known.","fun_headline_variants_meta":{"raw":{"variants":["Sharp bounds on individual treatment effect prediction intervals","ITE prediction often trivial: sharp bounds from marginals","Individual effects: when can we predict? Sharp bounds answer","Fisher and Neyman nulls conflict in individualized inference","RCT data only yields trivial ITE intervals in common cases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000445,"raw_usage":{"total_tokens":2289,"prompt_tokens":1023,"completion_tokens":1266,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":1188}},"tokens_in":639,"tokens_out":1266,"duration_ms":11089,"temperature":1.0,"reasoning_tokens":1188,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:35:38.337073+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any two discrete marginal distributions and solve the linear program over the joint probability table to minimize and maximize $\\sum_i P(Y_1=i,Y_0=i-\\delta)$ subject to the marginals; if the optimum falls outside the interval in Equation (5), the sharpness claim fails. In the binary case, one can similarly enumerate $t\\in[\\max\\{0,p+q-1\\},\\min\\{p,q\\}]$ and check whether any of the six candidate intervals has coverage at least $1-\\alpha$ whenever $p$ and $q$ both lie between $\\alpha$ and $1-\\alpha$.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the four patient types (NR, HE, HU, AR) that organize the binary treatment and outcome model."},{"cited_title":"and Park, S","cited_arxiv_id":null,"evidence_quote":"Supplies sharp bounds on the distribution of treatment effects using copulas, the starting point the paper extends and corrects."},{"cited_title":"Bounds on the Distribution of a Sum of Two Random Variables: Revisiting a problem of Kolmogorov with application to Individual Treatment Effects","cited_arxiv_id":"2405.08806","evidence_quote":"Provides the sharp cdf bounds for $Y_1-Y_0$ used in the continuous and ordinal sections, including the zero-inclusion conditions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"States Strassen's theorem for finite sets, which the paper uses to prove attainability of the lower bounds in Theorem 12."},{"cited_title":"and Cand \\`e s, E","cited_arxiv_id":null,"evidence_quote":"Gives the conformal inference prediction interval for ITE that the paper compares with its conservative continuous-outcome interval."}],"review_version":1}