{"id":"74184a67-7071-46b1-a868-0356755a46fd","arxiv_id":"2505.23196","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"JAPAN constructs conformal prediction sets by thresholding normalizing-flow density estimates, yielding valid, compact, possibly disjoint regions with lower area than residual-based baselines.","lead":"This paper proposes JAPAN, a conformal prediction method that builds prediction regions by thresholding the density estimated by normalizing flows, instead of using residuals or distances. The result is a flexible uncertainty-quantification tool that can produce compact, validity-guaranteed, and possibly disconnected regions that follow the shape of multimodal data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Area-efficiency claim is not yet supported: HPD-split, the direct density-based baseline, is absent and baseline area computation is unspecified.","rationale":"The reader's verdict is CONDITIONAL and I do not move it. The coverage guarantee is standard split-conformal and valid; the theory's rank-preservation/homogeneity assumptions are idealised, but the method would still be practically useful if empirical efficiency were firmly established. The decisive gap is that the efficiency claim has not been tested against HPD-split, the direct density-threshold predecessor, and the area metric lacks a common protocol across methods. Both gaps are fixable with one benchmarking pass. If HPD-split matches JAPAN, the novelty shrinks to 'use a flow instead of a kernel density estimator' rather than 'density scoring enables compact regions'; if JAPAN remains smallest under a shared protocol, the central claim stands. I therefore retain CONDITIONAL rather than ACCEPT.","tokens_in":30499,"tokens_out":17689,"duration_ms":212038,"concrete_test":"Implement HPD-split on the toy, multivariate, and time-series benchmarks and recompute all methods' areas with one shared protocol: a common test-point grid or the same number of Monte Carlo samples drawn from a fixed proposal distribution, with identical seeds and bounding volume. If HPD-split attains comparable or smaller areas than JAPAN, the method's incremental contribution above density-based scoring is not established; if JAPAN remains smallest under the shared protocol, the efficiency claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The main empirical claim is that JAPAN gives tighter prediction areas than existing baselines, but the comparison is not yet auditable. The most relevant density-based predecessor, HPD-split (Izbicki et al., 2021), is cited in Section 6 but never benchmarked; without it, the experiments only show JAPAN beats residual/geometric conformal methods, not that density-threshold scoring itself is what produces the gain. In addition, Appendix B.2 specifies a 3,000-sample Monte Carlo area estimator for flow-based methods (JAPAN, CONTRA, PCP) but gives no area-computation protocol for CQR, RCP, NLE, Dist-Split, or the copula baselines. If those areas are computed by exact formulas for boxes or ellipsoids while JAPAN uses a possibly low-variance Monte Carlo estimator on the flow's own density, the reported area gaps could be partly a measurement artifact. The split-conformal coverage validity of the main JAPAN score is standard and not in question; what is load-bearing is the efficiency half of the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes JAPAN, a split-conformal method for multivariate prediction regions. A normalising flow is trained to estimate the conditional density p(y|x); the log-density is used as the conformity score, and the prediction region is the superlevel set {y : log p̂(y|x) ≥ τ}, where τ is a calibration quantile. The authors state optimality propositions under rank preservation, give an importance-sampling formula for area estimation, and report experiments on toy densities, multivariate regression benchmarks, and time-series datasets, comparing against eleven baselines. Extensions cover unconditional, conditional-on-prediction, posterior, latent-space, and adaptive-threshold variants.","tokens_in":30715,"tokens_out":7696,"duration_ms":83058,"significance":"If the efficiency results hold, JAPAN is a practically useful instantiation of density-threshold conformal prediction, producing geometry-free, potentially disjoint, context-adaptive regions with finite-sample marginal coverage inherited from split conformal prediction. The calibration half of the method is standard and sound, and the latent-space importance-sampling area estimator (Proposition 3) is a useful implementation device. The theoretical efficiency claims, however, rest on strong assumptions that are not stated in the main text, and the empirical efficiency comparison is not yet auditable: the most direct density-based predecessor, HPD-split, is cited but never benchmarked, and the area-computation protocol is specified only for flow-based methods. With those gaps addressed, the paper would be a solid empirical contribution; as it stands, the central efficiency claim is not fully supported.","major_comments":[{"comment":"The proof asserts that uniform density approximation |fθ(y|x) − g(p(y|x))| ≤ δ implies |τ* − g^{-1}(τ)| ≤ ε(δ) with ε(δ)→0, but this is not a consequence of uniform density error unless additional regularity is assumed: the quantile map from density level to threshold is not automatically Lipschitz, and density error near the level set does not by itself control the threshold shift. The proof also silently uses the homogeneity assumption p(y|x)=φ(T_x(y)), which is absent from Proposition 2 in Section 3.1 and is violated in heteroscedastic or input-dependent multimodal settings. Please state the regularity conditions explicitly and prove the threshold-transfer step, or present the result as a heuristic motivation rather than a theorem.","section":"Appendix A, Theorem 2 (main-text Proposition 2)"},{"comment":"The proof concludes g^{-1}(τ)=τ* from the fact that both regions 'achieve the same coverage level', but the conformal region has a random calibration threshold and satisfies only P(y ∈ Γϵ(x)) ≥ 1−ϵ in finite samples; exact equality of regions therefore holds only in an idealized population/continuity sense, not for the actual split-conformal procedure. The fixed strictly increasing transformation g and the homogeneity assumption should be stated in the main-text Proposition 1, and the statement should be qualified to reflect the finite-sample, randomized nature of the conformal threshold.","section":"Appendix A, Theorem 1 (main-text Proposition 1)"},{"comment":"HPD-split (Izbicki et al., 2021), the most direct density-threshold conformal predecessor, is cited in Section 6 but is not included in any experiment. Without this baseline, the experiments show only that JAPAN improves on residual-based and geometric methods, not that density-threshold scoring itself drives the area gains. Please add HPD-split (and ideally CD-split) to the comparison, or explicitly restrict the empirical claim to the baselines actually evaluated.","section":"Section 4 and Tables 1, 3, 4"},{"comment":"Area computation is specified only for flow-based methods: Appendix B.2 states 3,000 Monte Carlo samples for JAPAN, CONTRA, and PCP, but no protocol is given for CQR, RCP, NLE, Dist-Split, or the copula baselines. If those areas are computed via exact box or ellipsoid formulas while JAPAN uses a stochastic estimator on its own fitted density, the reported area gaps could be partly a measurement artifact. Please report the area estimator used for every baseline, and include Monte Carlo standard errors or otherwise quantify estimator uncertainty.","section":"Appendix B.2 and Tables 1, 3, 4"},{"comment":"The misranking probability mass µ(x) is a probability with respect to p(y1|x)p(y2|x), but the proposition claims a bound on the Lebesgue measure of a symmetric difference of level sets. Small µ(x) does not imply small Lebesgue measure of the boundary strip unless the density is bounded away from zero on the relevant level sets and the level sets have controlled perimeter; these conditions are not stated. The proposition should be proved with explicit regularity assumptions, or removed from the set of theoretical efficiency claims.","section":"Appendix A, Proposition 4"}],"minor_comments":[{"comment":"The rank formula k = ⌊ϵ·(1−1/m)·m⌋ = ⌊ϵ(m−1)⌋ is nonstandard; with this k the test score at the k-th calibration value gives a p-value slightly below ϵ, so the procedure is conservative. Please reconcile the formula with the p-value definition in Section 2.1 and state the intended finite-sample convention.","section":"Algorithm 1, line 5"},{"comment":"The caption mentions 'except Drone', but Table 3 reports only tabular datasets; this appears to be a typo for 'SCM'.","section":"Table 3 caption"},{"comment":"The Particle-1 text says the dataset includes 5,000 samples, while Table 7 reports 2,000 + 500 + 500 = 3,000 samples; please reconcile the numbers.","section":"Appendix C.2 and Table 7"},{"comment":"The estimator uses exp(ϕ(z,x)) but ϕ is never defined; please either connect it to Φ in Eq. (1) or define ϕ as the log absolute determinant of the inverse Jacobian.","section":"Appendix A, Theorem 3"},{"comment":"The captions refer to 'the bottom panel', but the panels are arranged side by side; please update the captions to match the layout.","section":"Figures 10–15 captions"},{"comment":"The text refers to 'Propositon A'; this should be Proposition 4.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The finite-sample coverage guarantee is standard split conformal prediction, and the paper would benefit from stating this more prominently and positioning the contribution relative to HPD-split earlier. The absent HPD-split baseline and the unstated area-computation protocol for non-flow methods are the main obstacles to accepting the efficiency claim; these are fixable within the manuscript's scope, so I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read on the JAPAN paper. The core idea is simple and worth taking seriously: use a normalizing flow's log-density in data space as the conformity score, threshold at a global quantile, and get split-conformal coverage. The coverage argument is standard and sound. The new bits are the data-space full-likelihood score (as opposed to CONTRA's latent-space ball), an importance-sampled area estimator, and a tau-adaptive threshold extension. That is a legitimate subfield refinement, not a paradigm shift.\n\nWhat the paper does well: it evaluates on many datasets (toy, tabular, time series) and consistently reports smaller regions than residual/geometric baselines while maintaining coverage. The extensions are thoughtful, and framing CONTRA as a latent special case is a useful unification. The area estimator via importance sampling in latent space is a nice touch.\n\nThe soft spots are real, though not fatal. The biggest is the efficiency claim. Density-thresholding for conformal sets is known from Lei et al. 2013 and HPD-split (Izbicki et al. 2021), and the paper cites HPD-split but never benchmarks it. Without including the direct predecessor, you cannot tell whether the gains come from density-thresholding itself or from the flow's density estimate—and the paper's story is specifically about density-thresholding. That is a missing control and should be fixed. Second, area computation for non-flow baselines is underspecified; only flow methods get a 3,000-sample Monte Carlo estimator. If baselines use exact geometric formulas, part of the reported gap could be a measurement artifact. The authors need to state the protocol for every baseline. Third, the efficiency theory (Propositions 1, 2, 4) relies on strong assumptions—monotone transformation, homogeneity of p(y|x), and regularity of level sets—and the bounds are asserted rather than derived. The coverage guarantee does not need those, but the area-efficiency claim does. The paper would be stronger if it acknowledged how unrealistic the homogeneity assumption is and tested the flow's ranking quality directly.\n\nNo code or data is released, so the numbers are not independently auditable.\n\nWho is this for? Practitioners working on multivariate conformal prediction and researchers interested in density-based conformity scores. It deserves a serious referee; the core method is sensible, and the weaknesses are fixable with a fair baseline, transparent area computation, and more careful claims. I would send it to peer review with a request for major revision. I would not cite it yet until the HPD-split comparison exists, but I would bring it to a reading group.","headline":"Sound split-conformal validity with a sensible density-based score, but the area-efficiency claim is not yet proven because the direct density-based baseline is missing and baseline area computation is opaque.","tokens_in":31235,"tokens_out":2302,"would_cite":false,"duration_ms":24587,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"JAPAN constructs conformal prediction regions by thresholding a normalising-flow estimate of the predictive density, giving compact, possibly disconnected, context-adaptive sets with finite-sample coverage.","keywords":["conformal prediction","normalising flows","density-based conformity scores","prediction regions","multivariate regression","time series forecasting","uncertainty quantification","finite-sample validity"],"falsifier":"On a synthetic regression with a known two-mode density, train a deliberately restricted flow (for example, a shallow affine flow) so that its density ranking of the two modes is reversed, then compare JAPAN's region area at nominal coverage with the true level-set area. A near-true area despite the reversal would show the rank assumption is not load-bearing; a much larger area would confirm that misranking, not miscalibration, is what destroys efficiency.","tokens_in":30216,"feed_emoji":"🎯","tokens_out":10547,"duration_ms":102001,"temperature":0.7,"pith_summary":"JAPAN replaces residual-based conformity scores (absolute error, $\\ell^2$ balls, rectangles) with a normalising-flow estimate of the predictive density and builds the conformal region as $\\Gamma_\\epsilon(x) = \\{y : \\log \\hat{p}(y \\mid x) \\ge \\tau_\\epsilon\\}$, where $\\tau_\\epsilon$ is the $(1-\\epsilon)$-quantile of calibration scores. The set is geometry-free, can be disjoint and multimodal, and adapts to $x$, while conformal calibration keeps finite-sample marginal coverage. The theory states that if the flow ranks candidate outputs the way the true density does, the region matches the optimal density-thresholded region, and bounded misranking costs only bounded excess area. Empirically, JAPAN reports near-nominal coverage and smaller areas than residual- and latent-ball baselines on toy densities, multivariate regression, and time-series forecasting. The paper is explicit that, like all marginal conformal methods, JAPAN does not guarantee conditional coverage, and the efficiency theory assumes conditional densities have the same shape across $x$ up to a transformation.","feed_headline":"Flow-based density scores shrink conformal prediction areas","feed_subtitle":"Thresholding flow-estimated densities yields disjoint, context-shaped regions that are smaller than residual-based sets.","key_machinery":"The machine that carries the argument is the normalising-flow change-of-variables identity, which turns a bijection $z = h(y,x)$ into a tractable conditional log-density, $\\log \\hat{p}(y \\mid x) = \\log p_Z(h(y,x)) + \\Phi(y,x)$, where $\\Phi$ is the log-volume correction (a log-determinant Jacobian for discrete flows). This log-density is the conformity score, and the prediction region is its superlevel set above the calibrated threshold $\\tau_\\epsilon$. The same identity also gives an area estimator: sample $z$ from the base density $p_Z$, map back through $h^{-1}$, and reweight by $|\\det J_{h^{-1}}|/p_Z(z)$, so region volume is computed in latent space without integrating over the data space. The supporting propositions transfer density accuracy to area efficiency: exact monotone agreement with the true density gives identical regions, uniform density error gives error-controlled area, and bounded misranking gives bounded excess area.","core_discovery":"The central claim is that a conformal prediction region can be made both valid and near-minimal by thresholding an estimated conditional density instead of transforming residuals. After training a conditional normalising flow, the paper uses $\\log \\hat{p}(y_i \\mid x_i)$ as the calibration score for each calibration pair and defines the test-time set as all $y$ whose estimated log-density is at least the empirical $(1-\\epsilon)$-quantile. The finite-sample coverage guarantee follows from standard conformal calibration and does not depend on the flow being accurate; the efficiency claim rests on the flow's ranking being close to the true density's ranking. Under exact rank preservation the constructed region coincides with the true density level set, and under approximate rank preservation or bounded misranking the excess area is controlled and goes to zero as estimation error goes to zero. The empirical section demonstrates the mechanism on spiral, moons, and checkerboard densities, where only the density-thresholded region tracks all disconnected modes, and shows smaller reported areas on tabular and time-series benchmarks while coverage stays near the nominal level.","pith_inferences":["Since conformal calibration enforces coverage no matter how bad the score is, the practical bottleneck for JAPAN is density quality, not calibration: a better density estimator should shrink areas while leaving the conformal machinery unchanged.","The homogeneity assumption $p(y \\mid x) = \\varphi(T_x(y))$ in the optimality proofs is unlikely to hold in heteroscedastic or multi-scale data; an implied open problem is an area bound that depends on how much level-set geometry varies with $x$.","A cheap deployment check would be to run JAPAN on synthetic data with known level sets and compare its thresholded region with the oracle-density region; a large area gap would diagnose rank failure rather than calibration failure.","The $\\tau_\\epsilon(x)$ extension points toward conditional coverage, and combining it with local neighbourhood calibration is a natural testable follow-up that the paper leaves open."],"forward_implications":["Prediction sets for multimodal outputs can be split into several disconnected high-density pieces rather than a single ball or box centred on the mean.","The same conformal pipeline serves multivariate regression and multi-step time-series forecasting by changing only how the flow is conditioned on a context representation.","Region volume can be estimated efficiently by Monte Carlo sampling in the flow's latent space, avoiding expensive data-space integration.","JAPAN can be layered on top of an existing point predictor by modelling $p(y \\mid \\hat{y})$, $p(\\hat{y} \\mid y)$, or $p(y)$; its latent-space variant is shown to coincide with CONTRA.","The adaptive-threshold extension modulates $\\tau_\\epsilon(x)$ using the flow's volume-change term, which on a crescent-density example improves coverage at a small cost in area."],"supporting_citations":[{"why":"Supplies the result that density-thresholded prediction regions are compact and can be optimal, which motivates using densities as conformity scores.","marker":"(Lei et al., 2013)"},{"why":"Provides the discrete normalising flow (RealNVP) with exact log-likelihoods and an efficient Jacobian log-determinant, the density estimator used in experiments.","marker":"(Dinh et al., 2017)"},{"why":"Defines the normalising-flow framework and its tractable likelihoods, the modelling basis for $\\hat{p}(y \\mid x)$.","marker":"(Papamakarios et al., 2021)"},{"why":"Gives the conditional normalising-flow change-of-variables formula used to evaluate the log-density conformity score.","marker":"(Winkler et al., 2019)"},{"why":"Introduces CONTRA, a normalising-flow conformal baseline whose latent-space construction JAPAN's latent variant is shown to reduce to.","marker":"(Fang et al., 2024)"},{"why":"Supplies the TarFlow autoregressive flow architecture adapted for time-series forecasting with exact likelihoods.","marker":"(Zhai et al., 2024)"},{"why":"Lays out inductive conformal prediction and its finite-sample coverage guarantees, which JAPAN inherits via calibration.","marker":"(Vovk et al., 2005)"},{"why":"Provides the PCP baseline that forms regions from local balls over generated samples, against which JAPAN compares on disjoint-mode densities.","marker":"(Wang et al., 2023)"}],"fun_headline_variants":["Density thresholding makes conformal sets adaptive and tight","Flow-based conformal prediction yields tighter, disjoint regions","Conformal sets via density scores: more compact, still valid","Replace residuals with flow densities for smaller prediction areas","Normalizing flows refine conformal prediction regions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The area-efficiency claim rests on the trained flow ordering candidate outputs the way the true conditional density does; the theory also assumes the conditional densities are shape-invariant in $x$, $p(y \\mid x) = \\varphi(T_x(y))$, which real regression problems generally violate.","fun_headline_variants_meta":{"raw":{"variants":["Density thresholding makes conformal sets adaptive and tight","Flow-based conformal prediction yields tighter, disjoint regions","Conformal sets via density scores: more compact, still valid","Replace residuals with flow densities for smaller prediction areas","Normalizing flows refine conformal prediction regions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0005,"raw_usage":{"total_tokens":2445,"prompt_tokens":941,"completion_tokens":1504,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":1428}},"tokens_in":557,"tokens_out":1504,"duration_ms":10297,"temperature":1.0,"reasoning_tokens":1428,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:51:24.818285+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a synthetic regression with a known two-mode density, train a deliberately restricted flow (for example, a shallow affine flow) so that its density ranking of the two modes is reversed, then compare JAPAN's region area at nominal coverage with the true level-set area. A near-true area despite the reversal would show the rank assumption is not load-bearing; a much larger area would confirm that misranking, not miscalibration, is what destroys efficiency.","supporting_citations":[{"cited_title":"Density estimation using real nvp, 2017","cited_arxiv_id":null,"evidence_quote":"Provides the discrete normalising flow (RealNVP) with exact log-likelihoods and an efficient Jacobian log-determinant, the density estimator used in experiments."},{"cited_title":"Normalizing flows for probabilistic modeling and inference, 2021","cited_arxiv_id":null,"evidence_quote":"Defines the normalising-flow framework and its tractable likelihoods, the modelling basis for $\\hat{p}(y \\mid x)$."},{"cited_title":"Learning Likelihoods with Conditional Normalizing Flows","cited_arxiv_id":null,"evidence_quote":"Gives the conditional normalising-flow change-of-variables formula used to evaluate the log-density conformity score."},{"cited_title":"CONTRA : Conformal Prediction Region via Normalizing Flow Transformation","cited_arxiv_id":null,"evidence_quote":"Introduces CONTRA, a normalising-flow conformal baseline whose latent-space construction JAPAN's latent variant is shown to reduce to."},{"cited_title":"Normalizing Flows are Capable Generative Models , December 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the TarFlow autoregressive flow architecture adapted for time-series forecasting with exact likelihoods."},{"cited_title":"Probabilistic Conformal Prediction Using Conditional Random Samples","cited_arxiv_id":null,"evidence_quote":"Provides the PCP baseline that forms regions from local balls over generated samples, against which JAPAN compares on disjoint-mode densities."}],"review_version":1}