{"id":"43af09a6-bcc1-42f7-b39e-6a9ee89c50ef","arxiv_id":"2607.02540","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Population statistics (diversity, freeze, flip, improvement) enable closed-loop adaptive control of simulated bifurcation, yielding lowest mean gap on 74.6% of G1–G81 MaxCut graphs.","lead":"The paper introduces AE-QSB, a population-state-driven adaptive layer for quantum-inspired simulated bifurcation that uses four runtime statistics to close the loop on step size, coupling mode, guidance, and restarts. On MaxCut G-set graphs it reports lower mean gaps than fixed-schedule SB, GSB, and Tabu-SB baselines on most instances, with three complementary variants trading speed against refinement quality.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Four-indicator sufficiency (D,F,Q,R) is asserted as non-redundant foundation but never directly ablated; only downstream mechanisms are toggled.","rationale":"The reader already isolates exactly this assumption as the weakest link and correctly assigns CONDITIONAL status because of the large hand-tuned surface and multi-start asymmetry. My concern is a sharper formulation of the same point: the paper never isolates the perception layer itself. Because the empirical win rates and ablations still stand, and because ME-BSB (single-start) alone already dominates many segments, the evidence remains strong enough that the verdict need not move. The proposed test is cheap, uses the paper’s own G22 protocol, and would cleanly confirm or refute whether the four-indicator story is load-bearing.","tokens_in":46081,"tokens_out":542,"duration_ms":17356,"concrete_test":"Re-run ME-BSB and SE-DSB on G22 (T=1000, B=256, 5 independent seeds) with decision rules driven solely by {F,R}: replace D-gated guidance (Eq. 13) by constant α_gbest, replace Q-dependent n_flip and early-stop by fixed maturity F, keep all other parameters identical. If mean gap stays within 0.1 % of the full four-indicator versions, the sufficiency claim is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract, Sec. 3.1, Sec. 5.2) is that the four population statistics D, F, Q, R (Eqs. 7–10) form a sufficient, non-redundant state representation that enables closed-loop adaptive control generalizing across G-set densities with fixed thresholds. Sec. 3.1 asserts that “relying on any single indicator makes it difficult to reliably distinguish evolutionary states,” yet Appendix A and Sec. 5.1 only ablate the mechanisms those indicators drive (exploration subpopulation, rescue, F-switch, density scheduling, etc.). No experiment disables or replaces subsets of the indicator set itself while holding the decision/execution layers fixed. Consequently it remains untested whether the reported gains (74.6 % lowest-gap, 84.5 % highest-AR) actually require the full quartet or whether a simpler adaptive schedule (e.g., F- and R-driven only) would recover essentially the same performance, which would shrink the claimed novelty of “population-state perception.”","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes AE-QSB, a population-state-aware adaptive enhancement framework for quantum-inspired simulated bifurcation (SB) on Ising/MaxCut problems. It defines four runtime statistics—diversity D, freeze rate F, flip rate Q, and improvement rate R (Eqs. 7–10)—and uses them in a perception–decision–execution loop to adapt step sizes, coupling mode (BSB/DSB/mixed), elite guidance, restarts, and early stopping. Three complementary instances are introduced: ME-BSB (F-driven hard BSB→DSB switch with a weak exploration subpopulation), SE-DSB (linear mixed coupling with gated rescue), and SG-DSB (density-aware scheduling plus velocity-momentum EMA). On G22 (T=1000, B=256, 10 repeats), SE-DSB and SG-DSB report mean gaps below 0.05%; ME-BSB reports 0.26% with the best single-run time–quality trade-off. Across 71 G-set graphs, AE-QSB variants are claimed to achieve the lowest mean gap on 74.6% of graphs and the highest average approximation ratio on 84.5%. A two-tier 30-variant ablation on G22 ranks the exploration subpopulation first and rescue second, with super-additive degradation when both are removed.","tokens_in":46546,"tokens_out":1489,"duration_ms":16330,"significance":"If the comparative claims hold under fair experimental conditions, the work is a solid empirical methods contribution to quantum-inspired combinatorial optimization: it replaces open-loop SB schedules with a lightweight, batch-compatible closed loop driven by computable population statistics, and it documents clear complementarity among three design points (extremum-seeking, smooth refinement, density-aware generalization) on the standard G-set. Strengths include multi-metric reporting (gap, AR, TTS), Welch tests, convergence and distribution figures, and a systematic ablation that isolates super-additive exploration–rescue coupling. The framing that runtime population statistics can ground adaptive control for SB-type solvers is useful and transferable in principle to other batch Ising frameworks. The contribution is primarily empirical and engineering-oriented rather than theoretical.","major_comments":[{"comment":"Sec. 4.3 and Table 4: SE-DSB and SG-DSB use a multi-start strategy (3–5 independent starts for T≥250, with lexicographic selection), while Standard, GSB, Tabu, and ME-BSB are single-start. The abstract and Sec. 4.5 headline claims—lowest mean gap on 74.6% of graphs and highest AR on 84.5%—aggregate these multi-start variants with single-start baselines. This asymmetry is load-bearing for the comparative superiority claim. Either re-run all methods under matched multi-start budgets (or report single-start SE/SG only), or restate the G1–G81 win rates with multi-start clearly excluded from the primary comparison and confined to a secondary reliability analysis.","section":null},{"comment":"Sec. 3.1 and Sec. 5.2 assert that D, F, Q, R form a “sufficient and non-redundant” state representation and that “relying on any single indicator makes it difficult to reliably distinguish evolutionary states,” yet Appendix A and Sec. 5.1 only ablate the mechanisms those indicators drive (exploration subpopulation, rescue, F-switch, density scheduling, momentum, etc.). No experiment disables or replaces subsets of the indicator set itself while holding the decision/execution layers fixed. Without such an indicator-level ablation (e.g., F+R only vs. full quartet), the central novelty claim that the four-indicator perception layer is necessary for the reported gains remains untested. A compact indicator-ablation table on G22 (and a dense/sparse pair) would close this gap.","section":null},{"comment":"Sec. 4.1–4.2 and Algorithms 1–2 list a large free-parameter set (F_switch, β_dense, τ_min, ρ_explore, α_gbest, D_thresh, r0/κτ/κF, stall/F/Q early-stop thresholds, γ, µ bounds, rescue parameters, etc.). Sec. 5.3 acknowledges redundancy and the need for graph-feature-based auto-tuning, but the main results use fixed thresholds claimed to generalize “without per-graph retuning.” Given that the weakest assumption of the paper is precisely this fixed-threshold generalization, the manuscript should either (i) report a sensitivity study over the main thresholds on a held-out graph subset, or (ii) clearly mark which parameters were tuned on G22 versus held fixed a priori, so that the 74.6%/84.5% figures cannot be read as fully parameter-free transfer.","section":null}],"minor_comments":[{"comment":"Eq. (3): C(s) = (2W_total + s^⊤Js)/4 is standard for J=−W, but the factor of 2 vs. the usual 1/4 form should be cross-checked against the reported G22 optimum 13359 so readers can reproduce cut values from spins without ambiguity.","section":null},{"comment":"Figure 1 is dense; the indicator→mechanism mapping box is hard to parse at print scale. Consider splitting perception vs. decision into two panels or moving the full equation block to the appendix.","section":null},{"comment":"Notation: Table 1 lists typical values but omits several symbols used later (ρ_explore, r_t, s density scale, m momentum). A short expanded symbol table would help.","section":null},{"comment":"Sec. 4.4: “GSB-BSB equation” appears to be a typo for GSB-BSB dynamics/collapse; please correct.","section":null},{"comment":"References: several arXiv-style and “for review” citations (e.g., free-energy machine, edge-of-chaos SB) should be updated to final venues where available before camera-ready.","section":null},{"comment":"Code availability is promised post-publication [50]; for reproducibility review, a frozen artifact (or anonymized repo) with seeds and G-set loaders would strengthen the empirical claims.","section":null}],"recommendation":"major_revision","confidential_remarks":"The multi-start asymmetry is the single most important fix; without it the abstract percentages overstate the method’s advantage relative to single-start baselines. The indicator-sufficiency gap is real but fixable with a modest extra ablation. Scope is appropriate for a methods/optimization venue; novelty is incremental over GSB/Tabu-SB rather than paradigm-shifting, which is fine if the experimental fairness issues are resolved. I would not reject on novelty grounds alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is straightforward: they take simulated bifurcation, wire four cheap population statistics (diversity, freeze, flip, improvement) into a real-time sense–decide–execute loop, and ship three complementary solvers that dominate most of the G-set on gap and approximation ratio. That is new enough for the SB/CIM niche and useful.\n\nWhat they actually deliver is concrete. ME-BSB, SE-DSB and SG-DSB are not just re-branded GSB or Tabu-SB; the F-driven hard switch, linear mixed coupling, density-aware schedules, gated elite guidance and the 15–18 % unguided exploration subgroup are specific and ablated. The G22 depth analysis (10 repeats, distributions, TTS trade-offs) and the 71-graph sweep are thorough. The two-tier 30-variant ablation is the best part: exploration subgroup is load-bearing (~7.5 pp gap hit when removed), rescue is second, and the super-additive interaction is cleanly shown. They also map which algorithm wins which density/frustration band, which is more honest than claiming universal superiority.\n\nSoft spots, in proportion. The stress-test note is right: they assert D,F,Q,R are sufficient and non-redundant but only ablate the mechanisms those indicators drive, never the indicator set itself. That is a real gap in the central claim, not a fatal one—the performance numbers still stand. Multi-start is used only for SE/SG, not the baselines, so some of the 0.04 % gap edge is inflated. Parameter surface is large and hand-tuned; ablations already show several knobs are redundant. Code is promised but not yet out. None of this overturns the empirical ranking.\n\nThis is for people who already run SB, CIM, or Ising machines on MaxCut/QUBO and want a practical adaptive layer rather than another fixed schedule. Math is standard Euler SB plus well-defined statistics; citations cover the right prior art. I would send it to referees. Engage if you work in the subfield; the ablation ranking alone is worth having on the shelf.","headline":"Solid engineering advance for SB solvers: closed-loop population sensing works in practice, even if the four indicators themselves were never ablated.","tokens_in":47114,"tokens_out":534,"would_cite":true,"duration_ms":13755,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Population statistics during search can close the loop on quantum-inspired simulated bifurcation, beating fixed schedules on most MaxCut benchmarks.","keywords":["Simulated bifurcation","MaxCut","Adaptive dynamics","Population diversity","Quantum-inspired optimization","Ising model","Closed-loop control"],"falsifier":"On a held-out suite of MaxCut instances whose density and frustration lie outside the G-set range, re-run the three AE-QSB variants with the published thresholds; if their mean-gap and approximation-ratio advantage over fixed-schedule BSB/DSB baselines disappears or reverses, the sufficiency claim for the four indicators fails.","tokens_in":46997,"feed_emoji":"🔄","tokens_out":645,"duration_ms":5800,"temperature":0.7,"pith_summary":"Simulated bifurcation solvers for combinatorial problems such as MaxCut usually run with fixed schedules for step size, coupling mode, and guidance strength. Those open-loop schedules often collapse population diversity or switch from exploration to refinement at the wrong moment. This paper argues that four simple statistics computed from the current batch of candidate solutions—symbol-space diversity, amplitude freeze rate, sign-flip activity, and recent objective improvement—are enough to sense the evolutionary stage and drive closed-loop decisions. Under that Adaptive Enhanced Quantum-inspired Simulated Bifurcation (AE-QSB) framework the authors instantiate three complementary algorithms that trade extreme-value speed against population-level refinement and density-aware generalization. On the classic G-set, the family records the lowest mean gap on roughly three-quarters of the graphs and the highest approximation ratio on more than four-fifths. The practical message is that runtime population statistics form a lightweight, computable feedback signal that lets quantum-inspired dynamics leave fixed schedules behind.","feed_headline":"Population stats close the loop on simulated bifurcation","feed_subtitle":"Four runtime indicators let quantum-inspired solvers beat fixed schedules on most G-set MaxCut graphs","key_machinery":"The four population-state indicators D (diversity), F (freeze rate), Q (flip rate) and R (improvement rate), together with the five adaptive mechanisms they drive—column-wise step size, BSB/DSB coupling selection, diversity-gated elite guidance, emergency/elite restart, and maturity-gated bit-flip plus early stopping—that close the loop inside every evaluation window.","core_discovery":"Population statistical information collected during batch simulated-bifurcation evolution supplies a sufficient and non-redundant foundation for adaptive control, allowing quantum-inspired Ising solvers to replace fixed open-loop schedules with a closed perception–decision–execution loop and thereby improve solution quality and cross-instance robustness on MaxCut benchmarks.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Population states close adaptive loop for quantum bifurcation","Four indicators let QSB beat fixed schedules on MaxCut graphs","Population stats enable closed-loop quantum-inspired solvers","AE-QSB adapts bifurcation via runtime population perception","Data-driven population control lifts QSB MaxCut robustness"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"That the four hand-crafted indicators and their fixed decision thresholds already capture every evolutionary state that matters across graphs of widely different density and frustration, without needing per-graph retuning.","fun_headline_variants_meta":{"raw":{"variants":["Population states close adaptive loop for quantum bifurcation","Four indicators let QSB beat fixed schedules on MaxCut graphs","Population stats enable closed-loop quantum-inspired solvers","AE-QSB adapts bifurcation via runtime population perception","Data-driven population control lifts QSB MaxCut robustness"]},"model":"grok-4.5","effort":"low","cost_usd":0.004784,"raw_usage":{"total_tokens":1436,"prompt_tokens":862,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":47840000,"prompt_tokens_details":{"text_tokens":862,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":513,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":862,"tokens_out":61,"duration_ms":5441,"temperature":1.0,"reasoning_tokens":513,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T12:34:54.154381+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a held-out suite of MaxCut instances whose density and frustration lie outside the G-set range, re-run the three AE-QSB variants with the published thresholds; if their mean-gap and approximation-ratio advantage over fixed-schedule BSB/DSB baselines disappears or reverses, the sufficiency claim for the four indicators fails.","supporting_citations":[],"review_version":1}