{"id":"de067dc7-73af-4a28-b34f-f2d9a6ce828f","arxiv_id":"2502.01495","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A quantum-fidelity proximity measure from QCML-trained quantum states gives lower k-NN prediction error than random forest proximity for high-yield corporate bond similarity.","lead":"This paper tests a 'quantum cognition' machine learning model, which represents data points as quantum states, for learning similarities between corporate bonds. It reports that the quantum-based similarity finds better substitutes for hard-to-trade high-yield bonds than random forest or Euclidean distances, with comparable results for investment-grade bonds.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HYG advantage may be an artifact of asymmetric ensembling: QCML distances averaged over 3 initializations, RF GAP over a single forest; test with RF GAP ensembled.","rationale":"The reader's weakest assumption already identifies the fairness of the comparison, including the QCML 3-initialization averaging. This is the most load-bearing issue because the primary evidence for the central claim (HYG outperformance) rests on a single-figure comparison that is not balanced. The proposed test directly addresses it. Other concerns (no code, hand-picked outlier) are secondary. The reader's CONDITIONAL verdict is appropriate; a targeted re-analysis could either confirm or refute the claim.","tokens_in":12468,"tokens_out":9702,"duration_ms":82702,"concrete_test":"Recompute the HYG unweighted and weighted KNN-MAPE (Figure 4a-b) using RF GAP proximities averaged over 3 (and ideally 10) independent random-forest seeds, matching the QCML 3-initialization averaging. If the QCML advantage decreases to within the plotted standard error bands or reverses at most k, the reported outperformance is explained by the ensembling asymmetry. Also run a paired Wilcoxon signed-rank test on the per-split MAPE differences between QCML and ensembled RF GAP at each k; report p-values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in Figure 4a-b for HYG compares a QCML distance matrix averaged over 3 independent initializations (Section 5.2) with RF GAP proximities computed from a single random forest. Averaging fidelities reduces variance and smooths the distance landscape; the paper itself shows QCML distances are graded (mostly <0.5) while RF GAP distances are near 1 for almost all pairs (Figure 3). This asymmetry is not an inherent property of the two methods alone: averaging RF GAP proximities over multiple seeds would also produce more distributed values. If the QCML advantage in unweighted KNN MAPE is driven by the variance reduction from ensembling rather than by the learned metric itself, the claim that QCML outperforms classical tree-based similarity in HY markets is not established. The paper acknowledges that proximity-weighting favors RF but does not address the initialization ensembling asymmetry, which favors QCML. No significance tests are reported, so the observed differences in mean MAPE over 10 splits could be within noise once the comparison is balanced.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a supervised distance metric for corporate bonds derived from Quantum Cognition Machine Learning (QCML), where each bond is mapped to a quantum state and distance is defined through quantum fidelity. The authors compare this QCML distance with random forest GAP proximity and Euclidean distance by evaluating the MAPE/MAE of a k-nearest-neighbors regressor on hold-out splits of HYG (high-yield) and IGSB (investment-grade) bond data, as well as on three public datasets. They report that QCML proximity outperforms the other metrics for HYG in both unweighted and proximity-weighted KNN, and performs comparably or slightly worse for IGSB in the weighted case. The paper also includes MDS visualizations of the learned distances and a comparison of QCML versus random forest regression performance.","tokens_in":12676,"tokens_out":3388,"duration_ms":29962,"significance":"If the comparison were fully controlled, the paper would make a useful contribution to supervised similarity learning for illiquid corporate bonds: it addresses a concrete practical problem (finding tradeable substitutes for illiquid bonds), benchmarks against the current state-of-the-art (RF GAP proximity), and reports public-dataset controls. The authors are transparent about one known confound, the proximity-weighting scheme favoring RF. However, the central empirical claim of a HYG advantage rests on an asymmetric ensemble setup and on mean differences without significance testing; these issues must be resolved before the claim can be considered established.","major_comments":[{"comment":"Section 5.2 states that the QCML distance matrix is \"ensembled 3 different distance matrices... by taking the average of each entry,\" while the Random Forest GAP proximities are computed from a single forest. This asymmetry is not controlled in Figure 4(a)-(b). Because averaging fidelities over independent initializations reduces variance and can smooth the distance distribution, the HYG advantage attributed to QCML may instead reflect ensembling rather than the learned metric itself. Please report results for a single QCML initialization and for RF GAP proximities averaged over multiple forests (e.g., three or more seeds), so that the comparison controls for ensembling.","section":"§5.2, Figure 4(a)-(b)"},{"comment":"The paper claims that QCML \"outperforms\" the other metrics, but no statistical significance tests are reported. In Table 2 the HYG metrics for QCML and RF overlap heavily at the level of one standard deviation (e.g., MAE .79±.06 versus .93±.08, RMSE 1.73±.17 versus 1.80±.19), and the bands in Figure 4 are standard errors of the mean only. Please report paired tests over the 10 splits (e.g., paired t-test or Wilcoxon signed-rank) for each value of k, with appropriate multiple-comparison correction, or otherwise quantify the effect size and its uncertainty.","section":"§6.2, Figure 4, Table 2"},{"comment":"The manuscript acknowledges that \"proximity-weighting inherently favors the RF-based proximities\" because RF distances are near the maximum for most pairs, while QCML distances are broadly distributed. This admission implies that the weighted KNN results in Figure 4(b,d) do not provide a neutral comparison of metric quality: the weighting function is chosen in a model-dependent way, and for RF GAP it coincides with the model's own prediction weights, an interpretation not available for the other metrics. The unweighted KNN results are therefore the primary evidence for metric quality. Please either analyze weighted results under a model-independent weighting function, or clearly state that the weighted comparison is an application-specific evaluation and not a fair head-to-head metric test.","section":"§5.2, §6.2"}],"minor_comments":[{"comment":"The introduction's organization paragraph states \"Section 8, the conclusion\" but the manuscript contains a Section 7 on MDS visualization; the outline should be updated to include Section 7.","section":"§1"},{"comment":"There is a typo: \"one-hot-econding\" should be \"one-hot encoding.\"","section":"§4"},{"comment":"The y-axis is logarithmic for the random forest panels but linear for the QCML panels; this makes the visual comparison of the distance distributions misleading. Please use a consistent scale or add a note explaining the different scales.","section":"Figure 3"},{"comment":"The caption for the Student Performance rows explains why MAPE diverges (target values include zero), but the main text in §6.1 does not mention this. Consider adding one sentence to the text to avoid confusion.","section":"Table 3"},{"comment":"Reference [15] (Lin and Jeon) and reference [20] (Rhodes, Cutler, Moon) are both cited in connection with the GAP proximity definition; please verify that the GAP formula in Eq. (4) is attributed to the correct source.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The authors include employees of BlackRock and Qognitive, the company commercializing QCML; the paper contains a disclaimer that views are not investment advice, but the dual affiliation suggests a potential conflict of interest in the empirical comparison. Additionally, no data or code availability statement is provided, which limits reproducibility of the bond dataset and the QCML implementation. These issues do not affect the technical soundness assessment but are worth considering for editorial policy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you should know: the paper's headline result—QCML fidelity proximity outperforms RF GAP proximity for high-yield bond KNN regression—is plausible but not actually established by the experiments as reported. The comparison is asymmetric: the QCML distance matrix is averaged over three independent initializations (Section 5.2), while the RF GAP proximities come from a single forest. Averaging smooths the distance landscape and reduces variance, and the paper itself shows QCML distances are graded while RF GAP distances saturate near 1 for almost all pairs. That could be a real property of the methods, but it could also be an artifact of ensembling. The stress-test note has it right: you need to run RF GAP over multiple seeds and average the proximities before claiming an inherent advantage. I checked the text; there is no mention of ensembling RF, and no significance tests anywhere. The bands in Figure 4 are standard errors over 10 splits, and several curves overlap within those bands. For IGSB the difference largely vanishes under proximity weighting, which they admit is favorable to RF. So the central superiority claim is conditional at best.\n\nWhat the paper does well: it is an honest, coherent empirical study. The authors explicitly concede that proximity weighting favors RF, they show the distance distributions side by side, and they test on public datasets where QCML sometimes underperforms (Diabetes). The idea of using quantum fidelity as a supervised similarity metric for sparse financial data is a legitimate extension of their prior QCML work, and the observation that RF proximities collapse to near-maximum distances in outlier-heavy high-yield data is worth taking seriously. The math is standard; no red flags there. Self-citation is heavy but the cited prior work is their own framework, so that is not a flaw per se.\n\nThe soft spots are real but fixable: the ensembling asymmetry, the lack of significance testing, the hand-picked MDS outlier example, and no released code. The paper knows about the weighting issue but not the ensembling issue, which is the more serious one. This is not a takedown; the central idea may survive a balanced comparison. But the paper as written overstates its case.\n\nWho it is for: readers interested in metric learning for illiquid securities and in QCML as a practical tool. They will get a clear hypothesis and an honest caveat-filled evaluation. It deserves a serious referee, but the referee should demand the balanced experiment before publication. My recommendation: send to peer review, with the ensembling and significance issues as required revisions.","headline":"Plausible empirical claim that QCML fidelity proximity beats RF GAP on HY bonds, but the comparison has an unaddressed ensembling asymmetry and no significance tests; worth refereeing but not yet convincing.","tokens_in":696,"tokens_out":915,"would_cite":false,"duration_ms":20096,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a supervised distance metric derived from quantum-state representations, one minus the squared overlap of learned quantum states, makes k-nearest-neighbor yield prediction more accurate than random-forest or…","keywords":["quantum cognition machine learning","supervised similarity learning","corporate bonds","high-yield bonds","random forest proximity","k-nearest neighbors","quantum fidelity","distance metric learning"],"falsifier":"Re-run the HYG comparison with matched resources: one QCML initialization versus one random forest, with both unweighted and proximity-weighted neighbor averaging on the same 10 held-out splits. If random-forest GAP achieves equal or lower mean absolute percentage error across most values of $k$, the paper's claim that QCML proximity is better for high-yield bonds collapses.","tokens_in":12304,"feed_emoji":"📉","tokens_out":9435,"duration_ms":79442,"temperature":0.7,"pith_summary":"This paper tries to establish that a supervised distance metric drawn from quantum cognition machine learning (QCML) is better than the current tree-based standard for finding similar corporate bonds. In QCML, each bond is encoded as a quantum state, the encoding is trained to predict the bond's yield or spread, and the distance between two bonds is defined as one minus the squared overlap of their quantum states. Using that distance as the metric in a k-nearest-neighbors regressor, the paper reports lower yield-prediction error than random-forest GAP proximity and Euclidean distance on high-yield bonds, and comparable or better performance on investment-grade bonds. This matters because corporate bonds trade sparsely, so a similarity measure that truly reflects the target variable can identify tradable substitutes, price bonds with few recent quotes, and explain model predictions. The paper's explanation is that QCML keeps sparse, outlier-heavy data in a compact representation, while random-forest proximity spreads most pairs to near-maximum distance.","feed_headline":"Quantum-state distances beat random forests on high-yield bonds","feed_subtitle":"A supervised quantum-fidelity distance finds closer yield neighbors for illiquid bonds than trees or Euclidean distance.","key_machinery":"The central object is the QCML distance $d_Q(x_t,x_{t'}) = 1 - |\\langle \\psi_t | \\psi_{t'} \\rangle|^2$, where each bond $x_t$ is mapped to the ground state $|\\psi_t\\rangle$ of the error Hamiltonian $H(x_t) = \\frac{1}{2}\\sum_k (A_k - x_{t,k} I)^2$ built from learned Hermitian feature observables $A_k$ and a target observable $B$ trained by minimizing mean absolute error. The quantum fidelity $|\\langle \\psi_t | \\psi_{t'} \\rangle|^2$ between two ground states is the similarity measure; it does the work of making closeness in feature space track closeness in yield or spread. The comparison baseline is random forest GAP proximity, which is the exact weight each training point contributes to a random forest prediction, and ordinary Euclidean distance.","core_discovery":"The central claim is that QCML proximity is a genuinely supervised similarity measure: the quantum states used to define it are learned to predict the target, so fidelity between states encodes target-relevant similarity. On the HYG high-yield cohort, k-nearest-neighbors regression with this metric achieves lower mean absolute percentage error than with random-forest GAP or Euclidean distances for both unweighted and proximity-weighted averaging, and on the IGSB investment-grade cohort it wins under unweighted averaging while being comparable or slightly worse when proximity-weighted. The paper attributes the edge to geometry: QCML produces a compact, coherent representation of a sparse manifold with many one-hot features and near-default outliers, whereas random-forest proximity places most pairs at maximum distance, so its neighbors-based predictions rest on very few points. The same pattern appears on public datasets: where QCML regression matches or beats random forests, the QCML metric matches or beats random-forest GAP in unweighted neighbor averaging.","pith_inferences":["A decisive test of the proposed mechanism is to compare single-initialization QCML with an un-ensembled random forest using both GAP and out-of-bag proximities; the paper's own comparison averages three QCML distance matrices but no random-forest ensembles.","Because the learned distance is target-specific, the same QCML pipeline could be trained on a multi-target objective such as yield, spread, duration, and rating to produce a similarity that balances several dimensions of bond risk, which a single-target metric cannot.","A practical extension is to measure realized trading costs rather than neighbor-averaging error: find the nearest substitute under each metric, attempt to trade it, and compare execution quality, since lower yield-prediction error does not by itself prove a better tradable alternative."],"forward_implications":["For high-yield portfolios, the QCML metric can directly identify tradable substitutes, because neighbors under it have yields much closer to a target bond than neighbors under random-forest GAP or Euclidean distance.","For investment-grade bonds, the QCML metric remains competitive: the paper finds it better under unweighted neighbor averaging and comparable or slightly worse under proximity-weighted averaging.","Where QCML regression matches or beats random forests on public datasets, the QCML metric also matches or beats random-forest GAP in unweighted neighbor averaging, suggesting the effect is not unique to bonds.","The compact representation property means QCML neighbor predictions draw on a broad set of nearby points, while random-forest proximity concentrates weight on a few near-maximum-distance points, which explains the high-yield advantage."],"supporting_citations":[{"why":"Introduces random forests and the original leaf-coincidence proximity that serves as a baseline.","marker":"[2]"},{"why":"Provides the QCML framework and the compact data-manifold representation that motivates the distance metric.","marker":"[7]"},{"why":"Establishes the supervised corporate-bond similarity application that this paper extends with QCML.","marker":"[11]"},{"why":"Shows tree ensembles act as adaptive weighted nearest neighbors, the conceptual basis for extracting proximities from forests.","marker":"[15]"},{"why":"Introduces the QCML paradigm from which the paper's proximity is derived.","marker":"[16]"},{"why":"Supplies the quantum fidelity formula and measurement formalism used to define the QCML distance.","marker":"[17]"},{"why":"Defines the GAP baseline proximity that the paper compares against.","marker":"[20]"},{"why":"Describes the supervised QCML regression training procedure that learns the feature and target observables.","marker":"[23]"}],"fun_headline_variants":["Quantum cognition metric beats random forests on high-yield bonds","QCML similarity tops trees for high-yield bonds","Supervised quantum distance outperforms trees on high yield","Quantum metric learning wins over forest proximity on HY bonds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that testing each distance metric by averaging the yields of nearby training bonds is a fair and unbiased measure of practical similarity quality, even though the two methods are compared with different averaging rules, ensemble sizes, and numbers of random starts.","fun_headline_variants_meta":{"raw":{"variants":["Quantum cognition metric beats random forests on high-yield bonds","QCML similarity tops trees for high-yield bonds","Supervised quantum distance outperforms trees on high yield","Quantum metric learning wins over forest proximity on HY bonds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000811,"raw_usage":{"total_tokens":3542,"prompt_tokens":913,"completion_tokens":2629,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":2565}},"tokens_in":529,"tokens_out":2629,"duration_ms":18101,"temperature":1.0,"reasoning_tokens":2565,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T15:06:15.368058+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the HYG comparison with matched resources: one QCML initialization versus one random forest, with both unweighted and proximity-weighted neighbor averaging on the same 10 held-out splits. If random-forest GAP achieves equal or lower mean absolute percentage error across most values of $k$, the paper's claim that QCML proximity is better for high-yield bonds collapses.","supporting_citations":[{"cited_title":"Random forests","cited_arxiv_id":null,"evidence_quote":"Introduces random forests and the original leaf-coincidence proximity that serves as a baseline."},{"cited_title":"Supervised similarity learning for corporate bonds using random forest proximities","cited_arxiv_id":null,"evidence_quote":"Establishes the supervised corporate-bond similarity application that this paper extends with QCML."},{"cited_title":"Random forests and adaptive nearest neighbors","cited_arxiv_id":null,"evidence_quote":"Shows tree ensembles act as adaptive weighted nearest neighbors, the conceptual basis for extracting proximities from forests."},{"cited_title":"Quantum cognition machine learning: Ai needs quantum, 2024","cited_arxiv_id":null,"evidence_quote":"Introduces the QCML paradigm from which the paper's proximity is derived."},{"cited_title":"Nielsen and Isaac L","cited_arxiv_id":null,"evidence_quote":"Supplies the quantum fidelity formula and measurement formalism used to define the QCML distance."},{"cited_title":"Rhodes, Adele Cutler, and Kevin R","cited_arxiv_id":null,"evidence_quote":"Defines the GAP baseline proximity that the paper compares against."},{"cited_title":"Quantum cognition machine learning: financial forecasting, 2024","cited_arxiv_id":null,"evidence_quote":"Describes the supervised QCML regression training procedure that learns the feature and target observables."}],"review_version":1}