{"id":"332c9ea4-d099-4262-9e91-af464ed4711c","arxiv_id":"2505.18508","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A heuristic called Cosm achieves new best-known cuts on G72 (7008), G77 (9940), and G81 (14060), with reported speedups of 655x to 3560x over the previous best heuristic.","lead":"The author reports a new heuristic algorithm, Cosm, that finds higher Max-Cut values than any previously reported solver on three large Gset benchmarks, including a cut of 14060 on the 20,000-variable G81. The paper provides solution bitstrings for independent validation, but does not describe how the algorithm works.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The new-best-cut claims rest entirely on the three hex bitstrings in Appendices C-E, and the report supplies no checksum or validator; a single transcription or formatting error would invalidate the headline. Independent bit-exact evaluation should be a precondition for accepting the cut claims.","rationale":"The reader's conditional verdict already requires validation scripts and code release; this concern adds a concrete precondition (bit-exact bitstring verification) but does not move the verdict in a different direction. If the verification fails, the paper should be rejected or revised; if it passes, the numerical cut claims stand. I differ from the reader's weakest assumption: rather than Eq. (3)'s success-probability model being the most load-bearing, I judge the bitstrings themselves to be the least independently verified support for the headline cut claims. The speed claims, while uncertain, are robust at the reported 86/100 and 66/100 success rates even with pessimistic confidence bounds; the 3/100 G81=14060 projection is secondary and already flagged by the author as an expectation awaiting Gurobi confirmation. The bitstrings, by contrast, are the sole falsifiable evidence for the new best cuts, and they are presented without any machine-readable checksum or validation artifact.","tokens_in":13003,"tokens_out":11247,"duration_ms":104062,"concrete_test":"Download G72, G77, and G81 from the Stanford Gset page. Write a small validator that parses the exact hex strings from the arXiv source (not a PDF render), expands them to 10,000/14,000/20,000 bits, maps 0 to -1 and 1 to +1, evaluates Eq. (1) over the official edge lists, and asserts cuts 7008, 9940, and 14060 and Ising energies -14022, -19672, and -28086. If all assertions pass, the accuracy half of the central claim survives; if any fail, request corrected machine-readable bitstrings and a checksum before further review.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that Cosm produces cuts 7008 (G72), 9940 (G77), and 14060 (G81), is supported only by the bitstrings in Appendices C, D, and E plus the assertion that evaluating Eq. (1) on the public Gset instances yields these values. No verification code, checksum, or independently computed cut value is provided in the manuscript; the strings are embedded as wrapped text, which is exactly the kind of artifact that can be corrupted by transcription, line wrapping, or PDF/arXiv rendering. Every downstream quantity in Table I (success probabilities, sweeps-to-target, time-to-target, speedups, hardware projections) and the optimality discussion inherits these cut values. If any hex character in Appendix E is wrong, the G81 cut=14060 claim may fail even though the surrounding narrative remains internally consistent. This is a data-integrity concern, not an allegation of fabrication; the manuscript itself asks readers to do the verification but gives no machine-readable artifact.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports that a heuristic algorithm called Cosm, implemented as a MATLAB proof-of-concept, attains higher Max-Cut values than previously reported on the three largest Gset instances: G72 (cut 7008), G77 (cut 9940), and G81 (cut 14060). The evidence for these numerical claims consists of three hexadecimal solution bitstrings in Appendices C-E. The paper also reports success probabilities, sweeps-to-target, and time-to-target for reaching 99.9% and 100% of these best cuts, and compares these values with published GES-PR time-to-target results to claim speedups of 655x (G77) and 3560x (G81). The abstract and Section III further state that the new best solutions 'appear to be optimal,' pending confirmation from an unpublished Gurobi run referenced as [13]. The Cosm algorithm itself is not described; the manuscript states it 'is to be published shortly after this report.'","tokens_in":13190,"tokens_out":2409,"duration_ms":21335,"significance":"If the three reported cut values are correct, this is a significant empirical result for a benchmark family that has been widely studied for 25 years. The decision to provide solution bitstrings is commendable because it makes the central numerical claims externally checkable without trusting the implementation. However, the absence of a verification script or checksum, the lack of an algorithm description, and the uncontrolled hardware comparison in the speed claims substantially limit the reproducibility of the paper's performance claims. The paper is a performance report rather than a complete algorithmic study; its value as a standalone publication depends on whether the reviewer and editor accept this format. The manuscript itself explicitly asks readers to perform validation, which is appropriate, but it does not provide the tooling to do so.","major_comments":[{"comment":"The central cut claims (G72=7008, G77=9940, G81=14060) rest entirely on the three hex bitstrings, yet the manuscript provides no verification script, checksum, or a separate independently computed value of the objective function. A single transcription or line-wrap error would invalidate the headline result while leaving the surrounding narrative internally consistent. The paper should either (a) provide a machine-readable solution file and/or a short evaluation script, or (b) include an independently computed cut value or signature (e.g., a hash of the bitstring with a stated cut value). Without this, the principal numerical claims are not verifiable from the manuscript as submitted.","section":"Section III, Appendices C-E"},{"comment":"The Cosm algorithm is not described. The text says the 'algorithm is to be published shortly after this report,' but the abstract and Section III make strong empirical claims that depend on the algorithm's details: solver parameters, neighbor structure, termination criteria, and the 'single solver parameter' whose improved setting is credited for 5-6x better sweeps-to-target. Without at least a high-level algorithmic description or pseudocode, the reported speed and quality results cannot be reproduced or independently assessed. This is a load-bearing omission for a paper whose contribution is a heuristic performance report.","section":"Section III, Methodology"},{"comment":"The reported speedups against GES-PR (655x for G77, 3560x for G81) are not a controlled comparison. Table I compares Cosm-MATLAB TTT measured on an Intel Core Ultra 7 155H laptop to GES-PR TTT published for an Intel i7-3770 CPU, and the implementations differ in language, parallelization, and code maturity. The manuscript acknowledges this indirectly by calling Cosm-MATLAB a 'proof-of-concept,' but the abstract's 'orders of magnitude faster' claim is nonetheless presented as a headline result. A fair comparison would require rerunning both solvers on the same hardware or at least reporting normalized timings.","section":"Section III, Speed and Table I"},{"comment":"The sweeps-to-target and time-to-target estimates rely on Eq. (3), which assumes independent trials with a constant success probability p_s. The success probabilities in Table I are estimated from only 100 trials (e.g., 3/100 for G81 100% cut), so the 99%-success-level extrapolation (r = log(0.01)/log(1-p_s)) has very wide confidence intervals. For the 3/100 case, r is 151, but the 95% confidence interval for p_s ranges from about 0.006 to 0.085, giving r between roughly 53 and 765. The paper does not report such uncertainties, making the stated sweeps-to-target (454M) and projected TTT (910 ms at 2 ns/sweep) appear spuriously precise.","section":"Section III, Equation (3) and Table I"},{"comment":"The abstract states that the new best solutions 'appear to be optimal,' but the the only support is an informal expectation about an unpublished Gurobi run in Ref. [13], which itself has not released solutions. The manuscript says 'It is our understanding and expectation... awaits official confirmation.' This is not evidence of optimality. The claim should be either removed from the abstract or explicitly labeled as an unverified conjecture, especially since a 'second-best' cut of 14058 is 99.986% of 14060 and the paper cites a common threshold for optimality [15] that is not met by 14058.","section":"Section III, Optimality"}],"minor_comments":[{"comment":"The definitions of STT and TTT use a variable name 't_t_trial' that is not defined in the text; it should be clarified as the average execution time per trial, which is then used in Table I.","section":"Section III, Equations (4), (5)"},{"comment":"The 'State-of-the-art TTT from GES-PR' column says 'Not enough data' for G72 99.9% but then reports a speedup figure of 655x for G77 and 3560x for G81. For G72, no speedup is stated, which is fine, but the column header should be split or annotated to indicate that the 'Not enough data' entry means no comparison was made.","section":"Table I, G72 row"},{"comment":"The visualization of G81 is very low resolution in the provided text and does not convey meaningful structure beyond the toroidal grid; consider replacing it with a labeled diagram showing the 100x200 periodic grid and representative edge weights.","section":"Appendix A, Figure 2"},{"comment":"The references list mixes arXiv preprints and peer-reviewed papers without consistent citation dates or DOIs; for example, Ref. [22] is a book chapter but is cited in text simply as 'GESPR – 2017 [22]' without a chapter or page range.","section":"Section II, References"}],"recommendation":"major_revision","confidential_remarks":"The paper is an unusual format: a performance report with the algorithm deliberately withheld. Even if the bitstrings validate, the lack of algorithm description and verification tooling makes this more of an announcement than a peer-reviewable research paper. The editor may want to consider whether the journal's scope accommodates such a report or whether the authors should be encouraged to submit a full algorithmic paper with code. The optimality language in the abstract is also stronger than the evidence supports, which may create a reproducibility risk for future readers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper's headline claim is that a heuristic called Cosm found new best cuts on three of the largest Gset instances (G72 cut=7008, G77 cut=9940, G81 cut=14060), and it provides hex bitstrings so the cuts can be independently checked. That is the real contribution. If those bitstrings evaluate correctly, the cut values are new best-known results, which would matter to the Ising/Max-Cut benchmarking community. I would treat the cut claims as unverified but verifiable, not as something to take on faith.\n\nWhat the paper does well: it gives the bitstrings, which is more than many benchmark reports do. It also states plainly that the algorithm is not described in this report and that optimality is unconfirmed but expected. The parameters were tuned on G57, a different instance, before running on G72/77/81, so the specific new-best cuts are not directly fitted to the targets. That is a point in its favor.\n\nSoft spots, in order. First, the bitstrings are wrapped hex text with no checksum or validator. The stress-test note is right that a transcription error would sink the headline. Independent bit-exact evaluation is a precondition for accepting the cut claims. Second, the speed claims rest on success probabilities from 100 trials, e.g., 3/100 for G81 at 14060, and the comparison to GES-PR uses a different CPU and implementation. So the \"orders of magnitude faster\" part is underdetermined, whatever the algorithm turns out to be. Third, the algorithm is undisclosed, so we cannot evaluate novelty or reproduce the method; the future \"algorithm release\" is a promise, not part of this paper. Fourth, the optimality remark is speculation.\n\nI do not think the undisclosed algorithm alone is a fatal flaw — the paper is explicitly a performance report with data for validation. But the lack of a machine-readable validator is a simple fix and a real weakness.\n\nWho this is for: anyone tracking best-known solutions on Gset, or comparing heuristic speedups, would want to know if the bitstrings check out. The hardware-projection part (2 ns/sweep) is speculative and should be read as such.\n\nMy recommendation: yes, send it to peer review. A serious referee should verify the three bitstrings independently, then decide whether the speed claims are adequately supported given the small samples and uncontrolled comparison. The paper would be more useful with the algorithm description or at least code, but the checkable cut values justify referee time.","headline":"New best cuts on Gset G72/77/81 with verifiable bitstrings is a real, checkable claim; the speed claims are the soft underbelly.","tokens_in":13742,"tokens_out":2548,"would_cite":true,"duration_ms":21290,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C27","90C59"],"pacs":[],"model":"deepseek-v4-flash","headline":"A heuristic algorithm named Cosm reports record Max-Cut values on the three largest Gset instances — G72 at 7008, G77 at 9940, G81 at 14060 — reaching 99.9% quality hundreds to thousands of times faster than prior heuristics.","keywords":["Max-Cut","Ising model","Gset benchmarks","heuristic algorithm","Cosm","QUBO","algorithm discovery","time-to-target"],"falsifier":"Take the G81 bitstring from Appendix E, expand it to 20,000 bits, and evaluate the weighted Max-Cut objective on the G81 instance from the Gset collection: if the cut is not 14060, the record claim fails. For the speed claims, once the algorithm is published, run several thousand independent trials at 3 million sweeps per trial; a success rate for cut 14060 significantly below 3/100 would invalidate the extrapolated time-to-target. For the optimality claim, compare the three bitstring cuts with the Gurobi-certified optimal values when they are released.","tokens_in":12754,"feed_emoji":"⚡","tokens_out":15573,"duration_ms":111501,"temperature":0.7,"pith_summary":"This paper reports a heuristic algorithm, Cosm, that posts the highest Max-Cut values ever claimed on the three largest Gset benchmark problems: 7008 on G72, 9940 on G77, and 14060 on the 20,000-variable G81 — all higher than any cut previously reported for those instances. The same algorithm reaches 99.9% of the best-known solution quality far faster than the prior state of the art, with reported speedups of more than 600x on G77 and more than 3,000x on G81. The author argues the new best solutions are likely optimal, matching what an exact solver has reportedly certified, and publishes the solution bitstrings so any reader can evaluate the cuts directly. If the claims hold, the largest Gset instances become the first of their size to be effectively cracked by a heuristic solver after 25 years of challenge results.","feed_headline":"Record cut 14060 reported on 20,000-variable G81","feed_subtitle":"New bests on G72, G77 and G81 include bitstrings for checking; 99.9% quality up to 3,500x faster.","key_machinery":"The load-bearing object is Cosm itself, an iterative heuristic defined functionally here (its design is promised in a forthcoming publication), whose fundamental step is a single sweep; the target problems are weighted Max-Cut instances on a toroidal square grid with weights $\\{-1, +1\\}$, a structure known to be NP-hard. For the speed claims, the carrying identity is the repetition formula $r = \\max\\left(1,\\ \\frac{\\log(0.01)}{\\log(1-P_s)}\\right)$, which converts a per-trial success probability $P_s$ into the number of independent trials needed to reach a target cut with 99% probability; sweeps-to-target and time-to-target are this $r$ times the per-trial sweep count and per-trial runtime. The record cuts themselves are carried by the published solution bitstrings, which let any reader plug bits into the Max-Cut objective and check 7008, 9940, and 14060 directly.","core_discovery":"On the paper's own terms, the central discovery is that an iterative heuristic named Cosm, whose fundamental operation is a single 'sweep' over the variables, attains higher cuts than ever previously reported on the three largest Gset instances: G72 cut=7008, G77 cut=9940, and G81 cut=14060, with corresponding Ising energies −14022, −19672, and −28086. Cosm also reaches 99.9% of the best known solution quality on G77 in 39 seconds and on G81 in 78 seconds, against published GES-PR baselines of 7 hours and 77 hours, respectively. The author states that the best solutions appear to be optimal and expects them to match Gurobi-certified optimal solutions once those are published, and includes solution bitstrings in appendices so the cuts can be independently validated. For the speed claims, success probabilities measured over 100 trials per instance are converted through a geometric-trial formula into sweeps-to-target and time-to-target; with a projected 2 ns per sweep on parallel hardware, 99.9% quality would be reached in about a millisecond.","pith_inferences":["The speed comparison is not a controlled benchmark: the GES-PR times were published for a different CPU and code base, so the true speedup on identical hardware is untested; re-running GES-PR on the same machine would settle the gap.","Because a single parameter change yielded 5–6x better sweeps-to-target, the algorithm's behavior appears highly sensitive to tuning, so the reported record probabilities (3/100 on G81) may shift across implementations and machines until Cosm is fully specified and stable.","If the 14060 cut is confirmed optimal, the three largest Gset instances lose much of their power to separate heuristics from exact solvers, and the natural next test is whether Cosm transfers to larger, denser, or more heavily weighted Ising instances where optimality is not yet certified.","The closing suggestion that disruptive performance can come from hardware-centric algorithm discovery is a research program rather than a demonstrated result; a direct test would be whether Cosm-style search also wins on problem families unrelated to toroidal grids."],"forward_implications":["Independent verification of the three bitstrings would make G72, G77, and G81 the first instances of their size in the Gset family to be solved at or near optimality by a heuristic algorithm.","The reported speedups imply that 99.9% solution quality on the largest Gset instances is reachable in seconds on a laptop, where the previous best heuristic needed hours.","A parallelized Cosm implementation at a projected 2 ns per sweep would reach 99.9% quality in about a millisecond and the best-known cuts within about a second, a regime that would reset expectations for Ising-machine hardware.","The improved G72 and G77 results over the author's earlier report are attributed mainly to resetting a single solver parameter using a more representative 5,000-variable tuning problem, suggesting the tuning protocol matters as much as the search moves."],"supporting_citations":[{"why":"Defines the Gset benchmark instances, including G72, G77, and G81, whose coefficients are needed to evaluate the claimed cuts.","marker":"[1]"},{"why":"Reports the GES-PR heuristic results whose published time-to-target values (25,800 s on G77 and 276,000 s on G81) serve as the speed-comparison baseline.","marker":"[14]"},{"why":"Source for the expectation that Gurobi has certified optimal solutions of the toroidal Gset problems, against which the Cosm cuts are expected to match.","marker":"[13]"},{"why":"The author's previous Cosm report, supplying the earlier success probabilities that the current campaigns improve on by 5–6x in sweeps-to-target.","marker":"[10]"},{"why":"Gives the Simulated Bifurcation Machine's roughly 99.6% solution quality and more-than-10-hour time-to-target on the largest Gset problems, the prior hardware-solver context.","marker":"[9]"},{"why":"Reports the previously highest G81 cut of 14056, the record that the new 14058 and 14060 cuts surpass.","marker":"[22]"},{"why":"Shows quantum annealing-inspired heuristics and simulated Coherent Ising Machine variants fall well below the best G81 quality, supporting the claim that the new cuts are unprecedented.","marker":"[6]"}],"fun_headline_variants":["Cosm heuristic cuts G81 at 14060, new record","Speeds up 3500x to 99.9% best quality on Gset","New best cuts on G72, G77, G81 by Cosm","Record G81 cut 14060, likely optimal","Cosm reaches near-optimal G81 in 78 seconds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The speed and speedup claims, computed in Section III, rest on the premise that the success probabilities measured from 100 trials per instance (as low as 3/100 for the G81 record cut) describe a true, constant per-trial success rate with independent trials, against a comparison baseline taken from published results on a different CPU and implementation.","fun_headline_variants_meta":{"raw":{"variants":["Cosm heuristic cuts G81 at 14060, new record","Speeds up 3500x to 99.9% best quality on Gset","New best cuts on G72, G77, G81 by Cosm","Record G81 cut 14060, likely optimal","Cosm reaches near-optimal G81 in 78 seconds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000905,"raw_usage":{"total_tokens":3903,"prompt_tokens":962,"completion_tokens":2941,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":2848}},"tokens_in":578,"tokens_out":2941,"duration_ms":20160,"temperature":1.0,"reasoning_tokens":2848,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:29:23.243656+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the G81 bitstring from Appendix E, expand it to 20,000 bits, and evaluate the weighted Max-Cut objective on the G81 instance from the Gset collection: if the cut is not 14060, the record claim fails. For the speed claims, once the algorithm is published, run several thousand independent trials at 3 million sweeps per trial; a success rate for cut 14060 significantly below 3/100 would invalidate the extrapolated time-to-target. For the optimality claim, compare the three bitstring cuts with the Gurobi-certified optimal values when they are released.","supporting_citations":[{"cited_title":"Retrieved April 2025","cited_arxiv_id":null,"evidence_quote":"Defines the Gset benchmark instances, including G72, G77, and G81, whose coefficients are needed to evaluate the claimed cuts."},{"cited_title":"Team s of Global Equilibrium Search Algorithms for Solving the Weighted Maximum Cut Problem in Parallel,","cited_arxiv_id":null,"evidence_quote":"Reports the GES-PR heuristic results whose published time-to-target values (25,800 s on G77 and 276,000 s on G81) serve as the speed-comparison baseline."},{"cited_title":"Improved Sparse Ising Optimization","cited_arxiv_id":"2311.09275","evidence_quote":"The author's previous Cosm report, supplying the earlier success probabilities that the current campaigns improve on by 5–6x in sweeps-to-target."},{"cited_title":"High -performance combinatorial optimization based on classical mechanics,","cited_arxiv_id":null,"evidence_quote":"Gives the Simulated Bifurcation Machine's roughly 99.6% solution quality and more-than-10-hour time-to-target on the largest Gset problems, the prior hardware-solver context."},{"cited_title":"Shylo and O.V","cited_arxiv_id":null,"evidence_quote":"Reports the previously highest G81 cut of 14056, the record that the new 14058 and 14060 cuts surpass."},{"cited_title":"Performance of quantum annealing inspired algorithms for combinatorial optimization problems,","cited_arxiv_id":null,"evidence_quote":"Shows quantum annealing-inspired heuristics and simulated Coherent Ising Machine variants fall well below the best G81 quality, supporting the claim that the new cuts are unprecedented."}],"review_version":1}