{"id":"c5b8b731-274f-42e8-828e-01a457f17437","arxiv_id":"2606.11865","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Compares post-hoc vs. in-training strategies for conformal Bayes under label shift, finding regime-dependent efficiency gains with up to 43% narrower sets in high-dimensional cases.","lead":"The paper compares post-hoc calibration and in-training adaptation as two ways to achieve valid conformal prediction sets under label shift in Bayesian models. A smart generalist might read it to learn practical choices for maintaining coverage when training and deployment data have different label distributions.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Importance weight estimation accuracy under label shift in high-dimensional regime","rationale":"The reader's weakest assumption is precisely the load-bearing condition for the regime-specific claim. Because the review was performed on the abstract, the full text may contain additional diagnostics on weight estimation, but the concern itself is unchanged and remains the point that must be checked before the 43% figure can be treated as robust.","tokens_in":1671,"tokens_out":333,"duration_ms":28301,"concrete_test":"Recompute the high-dimensional experiment (the one reporting 43% reduction) after replacing the estimated weights with weights corrupted by additive Gaussian noise whose variance matches the empirical variance of the weight estimator on the given sample sizes; if coverage falls below 1-α or the width reduction drops below 20%, the claim does not survive realistic weight estimation error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline empirical claim (Strategy B yields up to 43% width reduction at nominal coverage) requires that the importance weights w(y) = p_T(y)/p_S(y) can be estimated accurately enough from finite source and target samples that the importance-weighted quantile still delivers exact finite-sample coverage. In the high-dimensional underdetermined regime the paper studies, the label-shift model is assumed to hold exactly and the weights are treated as known or perfectly recoverable; any estimation error (variance from small target sample size, model misspecification, or support mismatch) directly perturbs both the tilted posterior and the conformal threshold, so the reported coverage guarantee and width reduction become conditional on an unverified estimation step.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript studies conformal Bayes under label shift and identifies two strategies that restore target-domain coverage via importance weighting: post-hoc calibration (Strategy A), which tilts the posterior predictive and corrects the conformal threshold with an importance-weighted quantile while leaving the parameter posterior unchanged, and in-training adaptation (Strategy B), which tilts the parameter posterior itself to produce a target-domain predictive whose HPD region forms the prediction set. Controlled experiments show regime dependence: Strategy A yields the narrowest valid intervals in the low-dimensional well-estimated regime, while Strategy B achieves up to 43% width reduction at unchanged coverage in the high-dimensional underdetermined regime, under the stated source-sampling and label-shift assumptions.","tokens_in":1788,"tokens_out":543,"duration_ms":15769,"significance":"If the empirical claims hold under accurate weight estimation, the work supplies a unified perspective on post-hoc versus in-training mechanisms for conformal Bayes and concrete regime-dependent guidance on which yields narrower valid sets, which is useful for practitioners facing label shift in Bayesian predictive modeling.","major_comments":[{"comment":"The headline claim that Strategy B achieves up to 43% width reduction at nominal coverage in the high-dimensional regime is load-bearing for the central conclusion, yet the manuscript treats the importance weights w(y) = p_T(y)/p_S(y) as known or perfectly recoverable from finite source and target samples; any estimation error (variance, support mismatch, or misspecification) perturbs both the tilted posterior and the conformal threshold, so the reported coverage and width reduction become conditional on an unverified estimation step that is not shown to be accurate in the underdetermined regime studied.","section":"Experiments (high-dimensional regime)"},{"comment":"The finite-sample coverage guarantee for the importance-weighted quantile in Strategy A is stated to hold exactly under the label-shift model, but the paper does not provide a derivation or bound showing that the guarantee survives when the weights themselves must be estimated from the same finite target sample used for calibration.","section":"Post-hoc calibration section"}],"minor_comments":[{"comment":"The abstract refers to 'the stated source-sampling and label-shift assumptions' without enumerating them; a brief explicit list in the introduction would improve readability.","section":"Abstract"},{"comment":"Notation for the tilted posterior predictive and the HPD-based set in Strategy B should be introduced with an equation early in the methods section to avoid ambiguity when comparing the two strategies.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which highlight important practical considerations around importance weight estimation. We address each major point below and indicate planned revisions to clarify assumptions and limitations.","responses":[{"response":"We agree that the reported efficiency gains, including the 43% width reduction, are shown under the assumption of known or perfectly recoverable importance weights, as stated in the source-sampling and label-shift assumptions. The controlled experiments isolate the mechanistic difference between the two strategies by using simulated data with exact knowledge of the label shift. In the high-dimensional underdetermined regime, practical estimation of weights from finite samples is indeed challenging and can affect both coverage and set sizes. We will revise the manuscript to explicitly qualify the headline claim as conditional on accurate weight estimation and add a new discussion subsection on practical weight estimation (e.g., via separate unlabeled target samples or density-ratio methods) together with a brief sensitivity simulation under moderate estimation error.","revision_made":"partial","referee_comment":"[Experiments (high-dimensional regime)] The headline claim that Strategy B achieves up to 43% width reduction at nominal coverage in the high-dimensional regime is load-bearing for the central conclusion, yet the manuscript treats the importance weights w(y) = p_T(y)/p_S(y) as known or perfectly recoverable from finite source and target samples; any estimation error (variance, support mismatch, or misspecification) perturbs both the tilted posterior and the conformal threshold, so the reported coverage and width reduction become conditional on an unverified estimation step that is not shown to be accurate in the underdetermined regime studied."},{"response":"The exact finite-sample coverage guarantee for Strategy A is derived under the assumption that the importance weights are known exactly, consistent with the label-shift model as formulated. When weights must be estimated from the calibration sample, the guarantee becomes approximate rather than exact. We will revise the post-hoc calibration section to state this assumption clearly and add a remark explaining that estimation error can be reduced by using a hold-out set for weight estimation. A rigorous concentration bound quantifying the coverage deviation induced by weight estimation error would require additional technical work (e.g., empirical-process arguments), which we can sketch at a high level but are not in a position to derive fully in the current revision.","revision_made":"partial","referee_comment":"[Post-hoc calibration section] The finite-sample coverage guarantee for the importance-weighted quantile in Strategy A is stated to hold exactly under the label-shift model, but the paper does not provide a derivation or bound showing that the guarantee survives when the weights themselves must be estimated from the same finite target sample used for calibration."}],"tokens_in":1409,"tokens_out":592,"duration_ms":24421,"standing_objections":["Full derivation of a finite-sample coverage bound for Strategy A that accounts for weights estimated from the same finite target sample used for calibration."]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that the work unifies two routes to restore coverage under label shift—post-hoc tilting of the predictive plus weighted quantile versus tilting the parameter posterior itself during training—and runs controlled experiments showing the first wins in low dimensions while the second can deliver narrower sets in high-dimensional underdetermined regimes.\n\nIt does a solid job laying out the mechanisms without overclaiming and in using targeted experiments to map when each strategy is preferable. The distinction between leaving the posterior fixed and adapting it is useful for anyone who has to choose between quick post-processing and retraining under shift.\n\nThe soft spot is exactly the one the stress-test note flags: the importance weights w(y) = p_T(y)/p_S(y) must be estimated from finite samples, and any error in the high-dimensional case directly affects both the tilted predictive and the conformal threshold. The abstract gives no information on how the weights were obtained, what sample sizes were used for the target, or whether coverage holds under realistic estimation noise. If the full paper contains only the stated assumptions without checks, the coverage guarantee and width numbers become conditional rather than unconditional.\n\nThis is for researchers working on conformal methods and distribution shift. A reader who needs practical guidance on calibration choices under label shift will get value from the regime comparison.\n\nIt deserves peer review because the unified framing and the controlled experiments are worth detailed checking, even with the open question on weight estimation.","headline":"The paper cleanly separates post-hoc vs in-training adaptation for conformal Bayes under label shift and shows regime-dependent performance, but the headline 43% width claim rests on unexamined weight estimation.","tokens_in":2232,"tokens_out":370,"would_cite":false,"duration_ms":28765,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Under label shift, adapting the Bayesian posterior during training yields narrower valid prediction sets than post-hoc calibration in high-dimensional regimes.","keywords":["conformal prediction","label shift","Bayesian prediction sets","importance weighting","post-hoc calibration","in-training adaptation","high-dimensional regime"],"falsifier":"An experiment in the high-dimensional regime that uses deliberately inaccurate importance weights or violates the exact label-shift assumption and then checks whether nominal coverage fails or the reported width reduction vanishes.","tokens_in":2569,"feed_emoji":"📊","tokens_out":692,"duration_ms":36397,"temperature":0.7,"pith_summary":"The paper examines how to restore target-domain coverage for conformal Bayes predictors when labels shift between source and target data. It isolates two mechanisms: post-hoc calibration that tilts the predictive distribution and adjusts the conformal threshold with importance weights, versus in-training adaptation that shifts the parameter posterior itself to the target domain. Controlled experiments reveal clear regime dependence, with post-hoc calibration producing narrower valid sets in low-dimensional well-estimated settings and in-training adaptation delivering up to 43 percent width reduction at fixed coverage when the problem is high-dimensional and underdetermined.","feed_headline":"In-training adaptation narrows sets 43% under label shift in high dims","feed_subtitle":"Post-hoc calibration wins in low dimensions, but training-time posterior tilting produces narrower valid sets when data is scarce relative t","key_machinery":"Importance-weighted conformal calibration that either adjusts the predictive and quantile after training or adjusts the parameter posterior during training under an exact label-shift model.","core_discovery":"Conformal Bayes under label shift can be handled by post-hoc calibration, which leaves the parameter posterior unchanged while tilting the posterior predictive toward the target and correcting the conformal threshold via an importance-weighted quantile, or by in-training adaptation, which tilts the parameter posterior to the target domain so that the highest predictive density region under the fitted target predictive serves as the prediction set. In the low-dimensional well-estimated regime the first strategy produces the narrowest valid intervals; in the high-dimensional underdetermined regime the second strategy achieves up to 43 percent width reduction at unchanged coverage, under the st","pith_inferences":["For complex models trained on limited target-domain data the in-training route may be the practical default when the high-dimensional advantage appears.","The observed regime split suggests running both strategies on any new problem and selecting the narrower valid set.","Similar comparisons could be run for other shift types provided importance weights remain estimable."],"forward_implications":["In low-dimensional well-estimated regimes post-hoc calibration yields narrower valid prediction intervals than in-training adaptation.","In high-dimensional underdetermined regimes in-training adaptation reduces prediction-set width by up to 43 percent while preserving nominal coverage.","Efficiency gains from in-training adaptation remain model-dependent and carry no guarantee of finite-sample conditional optimality.","Both strategies assume the label-shift model holds exactly and that importance weights can be estimated accurately from the given samples."],"fun_headline_variants":["Post-hoc vs in-training for conformal Bayes under label shift","High dim label shift: in-training cuts widths 43% at same coverage","Low dim favors post-hoc calibration in conformal Bayes label shift","In-training adaptation tilts posterior for high dim efficiency gains"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The importance weights that correct for label shift can be accurately estimated from the available source and target samples, and the label-shift model itself holds exactly.","fun_headline_variants_meta":{"raw":{"variants":["Post-hoc vs in-training for conformal Bayes under label shift","High dim label shift: in-training cuts widths 43% at same coverage","Low dim favors post-hoc calibration in conformal Bayes label shift","In-training adaptation tilts posterior for high dim efficiency gains"]},"model":"grok-4.3","cost_usd":0.004384,"raw_usage":{"total_tokens":2206,"prompt_tokens":690,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":43837000,"prompt_tokens_details":{"text_tokens":690,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1447,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":690,"tokens_out":69,"duration_ms":18327,"temperature":1.0,"reasoning_tokens":1447,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T02:12:20.166887+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment in the high-dimensional regime that uses deliberately inaccurate importance weights or violates the exact label-shift assumption and then checks whether nominal coverage fails or the reported width reduction vanishes.","supporting_citations":[],"review_version":2}