{"id":"527d378e-dbae-4d74-b9ba-b8a0ab9c627e","arxiv_id":"2510.15839","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Best-of-three rank data makes correlated probit reward models identifiable, with a near-optimal estimator, while pairwise data provably cannot.","lead":"Pairwise preference data cannot identify correlation structure in probit reward models; asking humans to rank three options can. The paper gives the first identifiability and near-optimal estimation guarantees for this setting.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorems require full rank-3 permutation data; the usual best-of-three (winner-only) data has only 2 independent probabilities for n=3, so the identifiability claim does not transfer.","rationale":"The reader's weakest assumption was that the true DGP is a single multivariate Gaussian; my concern is distinct but related: the paper's central claim is stated for full rank-3 permutation data, whereas the title and the usual RLHF interpretation of 'best-of-three' denote winner-only data. For n=3, winner-only data has only 2 independent probabilities while the model has 4 parameters, so the claimed identifiability cannot hold under that interpretation. This does not change the validity of Theorem 4.4 for the formal model, but it changes how the results can be advertised and used. The reader's CONDITIONAL verdict already flags scope and overclaim issues; my concern reinforces that a scope clarification is necessary, so I leave the verdict unchanged. I partially agree with the reader because both concerns center on a gap between the clean theoretical model and the practical data-collection setting.","tokens_in":46531,"tokens_out":18504,"duration_ms":165313,"concrete_test":"Derive or numerically simulate the winner-only map for n=3: for (μ,Σ) satisfying Assumption 3.1, compute p_i = P{X_i ≥ X_j, X_i ≥ X_k} for i=1,2,3. Search over the 4-dimensional parameter space for two distinct (μ,Σ) pairs whose winner-probability vectors match to 1e-6. If such a pair exists (as dimension counting predicts), the identifiability claim fails for winner-only best-of-three data. Alternatively, analytically exhibit a continuous family of (μ,Σ) with identical winner probabilities.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's formal object is a full ranking of three alternatives: observables are the six probabilities P{X_i ≥ X_j ≥ X_k} (Section 4, Figure 1, Theorem 4.4). The title/abstract and the RLHF reading of 'best-of-three' usually mean the annotator selects the best of three options, i.e., only the winner event is observed. These are not equivalent. For n=3, under Assumption 3.1 the model has 4 free parameters (μ lies on a 2D plane; Σ is a rank-2 PSD matrix on that plane with trace fixed, 2 parameters), while winner-only data gives three probabilities summing to 1, i.e., 2 independent degrees of freedom. A 4-parameter family cannot be identified from 2-dimensional data, so the analogous identifiability statement for winner-only data is false for n=3. The experiments do not close this gap: they use full rankings (sushi) or convert ratings to rankings and then sample sorted triples, not winner-only annotations. Thus the practical message—collect 'best-of-three' data—is unsupported unless that phrase is explicitly defined as full ranking of the three alternatives.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the correlated probit RUM X ∼ N(µ,Σ) on n alternatives, normalized as in Assumption 3.1 (mean on the hyperplane, Σ1=0, trace n, rank n−1). The main results are: (i) Theorem 3.2, pairwise comparison probabilities are not sufficient to identify the model; (ii) Theorems 4.1 and 4.4, the model is identifiable from three-way full-ranking probabilities P{X_i ≥ X_j ≥ X_k} for all permutations; (iii) Theorem 5.2, a polynomial-time estimator from rank-3 permutation data with N ≥ Cn² ε^{-2} γ^{-24} ... samples; (iv) Theorem 5.3, a minimax lower bound showing that Ω(n² ε^{-2}) queries are necessary. The paper also reports synthetic and real-data experiments comparing probit models fit from pairwise and three-way data, claiming improved personalization. The formal positive results are for full rankings of three alternatives, not for winner-only 'best-of-three' data, a distinction that matters for the paper's practical claims.","tokens_in":46807,"tokens_out":19425,"duration_ms":164403,"significance":"If the results are read as applying to full rank-3 permutation data, they are a valuable contribution: to my knowledge this is the first rigorous identifiability and finite-sample analysis for the correlated probit model from higher-order choice data, and the pairwise-insufficiency theorem explains a real gap in the literature. The lower-bound proof is a genuine minimax argument (Bretagnolle–Huber + Le Cam) and the upper-bound construction is non-circular: it builds estimators for 3-item triples and aggregates them via a sparse graph with logarithmic diameter. The sample-complexity dependence n² ε^{-2} is tight in n and ε. The paper also openly provides code. However, the matching of the formal object (full ranking) to the advertised object ('best-of-three') is not sound as written, and the experiments do not substantiate the 'great improvement' language. The theoretical core is defensible after revision.","major_comments":[{"comment":"The data model is stated inconsistently. §3 says observations are 'the choices arg max_{i∈R} X_i' for query sets R; for |R|=3 this is winner-only data. The positive results, however, are proved for the six three-way ranking probabilities P{X_i ≥ X_j ≥ X_k} (Figure 1, Theorem 4.4, Theorem 5.2). These are very different. For n=3 under Assumption 3.1 the model has four free parameters (two for μ on the plane, two for the rank-2 trace-normalized covariance), whereas winner-only best-of-three data has only three cell probabilities summing to one, i.e. two independent degrees of freedom. Identifiability from winner-only data is therefore impossible by dimension counting. Thus the paper's central slogan that 'best-of-three preference data provably overcomes' pairwise insufficiency is not supported unless 'best-of-three' is explicitly and prominently defined as a full ranking of the three altern","section":"§3 and §4 (Theorems 4.1, 4.4, 5.2)"},{"comment":"The experimental validation is overstated. In Table 2 the best-of-three probit differs from logit or pairwise probit by at most 0.01–0.03 in most rows, and many entries are identical at the reported precision (e.g., ml-1k-llm all 0.57; nf-10k-llm all 0.59; sushi-B-default 0.65 for logit and best-of-three). No error bars, confidence intervals, or repeated-seed results are given, and Table 1 reports only quantile levels. The text claims 'great improvement' and the abstract claims 'improved personalization'; these claims are not supported by the displayed effect sizes. Please provide variance estimates, report the actual differences, or soften the conclusions.","section":"§6, Table 2"},{"comment":"The proof of the key monotonicity step is incomplete for the zero-mean case. Claim E.3 states eξ'(θ) = −d1 − d2 cosθ / sin²θ < 0 for θ∈(0,π), but this is false when d1=d2=0, which is permitted by Assumption 3.1 and is in fact a case used in the experiments. Since Lemma 4.3 establishes identifiability of α12 from the monotone relation γ(θ), the proof as written does not cover μ=0. The lemma is true (for an isotropic Gaussian, γ12 directly determines arccos⟨v1,v2⟩), so this is a patchable gap, but it must be addressed explicitly.","section":"§E.1, Lemma 4.3 (and Claim E.3)"}],"minor_comments":[{"comment":"The normalization is inconsistent: DIFF-FORM sets Tr(Σ)=n−1, but in Case 2 the text says Tr(eΣ)=n+ν and rescales with t=(n+ν)/n. This appears to be a typo and should be corrected.","section":"§E.2, Theorem 3.2 proof"},{"comment":"The phrase 'by relaxing the scale constraint Tr ˙Σ = 1' is unclear because the preceding normalization is Tr(Σ)=n and, since P is orthogonal, Tr(˙Σ)=n. State the scale convention used in the construction.","section":"§4.1, Eq. (2)"},{"comment":"'given there is a polynomial time algorithm which when allowed N observations ... of its choice' should be rephrased and should specify whether the choice of triples is nonadaptive.","section":"Theorem 5.2"},{"comment":"The caption says 'two figures on the right'/'two figure on the left' but does not clearly identify which panels are pairwise, ground truth, and best-of-three; label the panels.","section":"Figure 2"},{"comment":"Column headings 'dds. var. feat.' are not defined; please spell out the dataset variants and feature encodings.","section":"§6, Table 2"},{"comment":"The γ-dependence γ^{-24} makes the finite-sample guarantee very sensitive to rare orderings; a sentence discussing when this is reasonable for preference data would help the reader.","section":"Assumption 5.1"}],"recommendation":"major_revision","confidential_remarks":"To the editor only: The theoretical results are, as far as I can judge, sound for the full-ranking model, and the pairwise-insufficiency theorem plus the minimax lower bound give the paper a solid core. The main risk is that the 'best-of-three' terminology will be read by the RLHF community as winner-only data, for which the identifiability claim is false. I would ask the authors to rename the data type (e.g., 'rank-3 permutation feedback') and to temper the experimental claims. The zero-mean gap in Lemma 4.3 is patchable but should be fixed before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the theoretical core is real: pairwise data provably cannot identify the correlated probit, while observing the full ranking of triples identifies (μ, Σ), and there is a near-optimal polynomial-time estimator with a matching lower bound. Second, the paper calls this \"best-of-three\" data, but the model and proofs use the full ordering—six permutation probabilities per triple, five independent degrees. Winner-only best-of-three, the standard RLHF annotation, gives two independent probabilities for three items and cannot identify the five-parameter model. The stress-test is right: the identifiability theorem does not transfer to winner-only data, and the experiments all use full rankings or ratings converted to rankings, so the practical message as written is unsupported.\n\nWhat is actually new and good: Theorem 3.2 is a clean formalization of the folklore that pairwise comparisons lose correlation information; the proof is elementary and self-contained. The geometric identifiability argument in Section 4—project onto the plane, recover the isotropic transformation, then read off the covariance—is neat and convincing. The estimator is a real contribution: the graph construction cuts the number of triples from O(n^3) to O(n^2) with only logarithmic diameter, and error propagation is handled explicitly. The proofs are detailed; I did not find a load-bearing gap.\n\nSoft spots. (1) The winner-only vs. full-ranking gap is the main issue. The abstract and title should say \"full ranking of triples\" unless the authors can prove something for winner-only data, which for n=3 they cannot. (2) The real-data experiments show mostly 0.01–0.03 accuracy gains over pairwise probit, with no error bars, yet the text claims \"great improvement.\" That overstates the evidence; the synthetic results, where best-of-three clearly recovers correlations that pairwise methods miss, are much stronger and should be foregrounded. (3) Assumption 5.1—every permutation of every triple has probability at least γ—is strong, and the γ^{-24} dependence in Theorem 5.2 means the finite-sample guarantee is mostly of theoretical interest. Worth stating in the limitations. (4) The lower-bound KL derivation is a bit loose in presentation but the bound itself checks out; a few lines of cleanup would help.\n\nWho is this for? Anyone working on choice models, preference learning, or RLHF data collection with an interest in what can be identified from ordinal feedback. The theory is solid, but the reader should not take the \"best-of-three\" label at face value. I would accept with major revision: clarify the data type, moderate the empirical claims, add error bars, and discuss the γ dependence and the homogeneous-Gaussian assumption.","headline":"Strong identifiability and estimation theory for correlated probit from full ranking triples; but the 'best-of-three' framing is misleading because winner-only data is a different, non-identifiable object, and the experiments overclaim.","tokens_in":47264,"tokens_out":4150,"would_cite":true,"duration_ms":36697,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F07","62H12","91B06"],"pacs":[],"model":"deepseek-v4-flash","headline":"Three-way choices reveal correlated preferences that pairwise comparisons cannot.","keywords":["random utility models","correlated probit","preference learning","RLHF","identifiability","best-of-three rankings","pairwise comparisons","sample complexity"],"falsifier":"Generate zero-mean correlated probit data from two normalized covariance matrices that differ only in one off-diagonal entry and have identical pairwise win probabilities; such pairs exist by the paper's construction. Feed best-of-three rankings to the proposed estimator and check that the recovered covariance separates the two models at the predicted O(epsilon^-2) sample rate, and feed pairwise rankings to any estimator and verify that the two models remain statistically indistinguishable at any sample size.","tokens_in":46452,"feed_emoji":"📊","tokens_out":5842,"duration_ms":54910,"temperature":0.7,"pith_summary":"Reward models built from human feedback typically assume independent utilities, which forces every user to share one universal utility and ignores correlations. This paper shows that with the correlated probit model, pairwise \"A or B?\" comparisons are fundamentally powerless to recover the covariance: infinitely many distinct models produce identical pairwise probabilities. It then proves that best-of-three ranking data is both necessary and sufficient to identify the mean and covariance, and gives a polynomial-time estimator achieving epsilon accuracy with about n^2 triple samples. Because triple-wise data is cheap to collect, the result makes correlated, personalized preference models practical in settings such as reinforcement learning from human feedback. Experiments confirm the theory: a probit trained on triples matches an oracle baseline and recovers interpretable correlations, while a pairwise-trained probit does not.","feed_headline":"Pairwise preferences can't learn correlated reward models; triples can.","feed_subtitle":"A new proof shows best-of-three ranking data is both necessary and sufficient to recover a correlated probit utility model.","key_machinery":"The central object is the correlated probit model X ~ N(mu, Sigma), and the load-bearing identity is that a three-way ranking probability equals a difference of pairwise and two-dimensional wedge probabilities, so after whitening every wedge mass is a monotone function of the angle between two unit vectors. Those angles, together with normal-CDF-derived projections, determine the covariance's shape. For the global n-item problem, the machinery is a graph whose vertices are pairs of items and whose edges are shared triples, chosen as a sparse O(n^2)-edge graph with logarithmic diameter via per-anchor binary trees. The lower bound uses two zero-mean Gaussians differing only in one off-diagonal","core_discovery":"The paper's central claim is that moving from pairs to triples removes a statistical barrier. Under a zero-sum normalization and full-rank condition, any Gaussian random utility model with at least three alternatives has infinitely many distinct mean-covariance pairs generating the same pairwise ranking probabilities, so pairwise data cannot identify correlation even before sampling noise. In contrast, three-way ranking probabilities determine the mean and covariance uniquely, and the paper constructs a polynomial-time estimator whose error in every coordinate is at most epsilon from N >= C n^2 epsilon^-2 gamma^-24 log(n/delta) log^6(n/(gamma epsilon)) independent triple rankings, with a mat","pith_inferences":["Editorial inference: because the identifiability argument only uses rank-3 probabilities, any higher-order query format that can be decomposed into triples, such as full rankings, should also be sufficient; the paper does not state this.","Editorial inference: the pairwise impossibility suggests that covariance estimates from existing pairwise-only pipelines are determined by regularization and initialization rather than by data; augmenting collection with triples is a cheap test of whether claimed correlations are real.","Editorial inference: the gamma dependence, requiring every triple ordering to have probability at least gamma, is likely improvable through adaptive sampling; the paper's lower bound does not address gamma.","Editorial inference: if real human preferences are a mixture of subpopulations, the single-Gaussian assumption is a misspecification; triple-order frequencies may still identify mixture components under a separation condition, an open direction the paper does not explore."],"forward_implications":["Pairwise comparison datasets, as collected in many RLHF pipelines, cannot support recovery of correlated reward structure; the impossibility is information-theoretic, not a sample-size issue.","Best-of-three preference data is sufficient: a polynomial-time estimator recovers mean and covariance to accuracy epsilon from about n^2 epsilon^-2 triple rankings, and the sample count is near-optimal in n, epsilon, and confidence.","Triples-trained probit models match the performance of an oracle matrix-completion baseline on synthetic and real preference data, while pairwise-trained probit learns spurious or missing correlations.","Correlation-aware welfare optimization changes concrete recommendations: menus chosen with a triple probit differ from logit or pairwise choices, and can serve more users' top preferences.","Learned correlations are interpretable: sequels and similar items cluster positively, while divisive items and mainstream blockbusters are negatively correlated."],"fun_headline_variants":["Pairwise data can't learn correlation; triples do","Best-of-three: the key to correlated reward models","Correlated preferences need triple comparisons","Why pairwise preference feedback isn't enough"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that all users' utilities come from one multivariate Gaussian distribution, normalized so utilities sum to zero, and that every triple ranking has probability at least gamma; if preferences are heterogeneous or multi-modal, or some rank-3 outcomes are nearly impossible, the identifiability and sample-complexity guarantees do not apply.","fun_headline_variants_meta":{"raw":{"variants":["Pairwise data can't learn correlation; triples do","Best-of-three: the key to correlated reward models","Correlated preferences need triple comparisons","Why pairwise preference feedback isn't enough"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000328,"raw_usage":{"total_tokens":1671,"prompt_tokens":749,"completion_tokens":922,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":865}},"tokens_in":493,"tokens_out":922,"duration_ms":7837,"temperature":1.0,"reasoning_tokens":865,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T09:19:00.356105+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate zero-mean correlated probit data from two normalized covariance matrices that differ only in one off-diagonal entry and have identical pairwise win probabilities; such pairs exist by the paper's construction. Feed best-of-three rankings to the proposed estimator and check that the recovered covariance separates the two models at the predicted O(epsilon^-2) sample rate, and feed pairwise rankings to any estimator and verify that the two models remain statistically indistinguishable at any sample size.","supporting_citations":[],"review_version":1}