{"id":"4a87376e-f833-4520-aef9-d0fcd3362263","arxiv_id":"2605.28003","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors release ResearchMath-14k, the largest dataset of research-level math problems, and demonstrate that agent-filtered reasoning trajectories from open models improve fine-tuned Qwen3 models by 9.2 points on average.","lead":"This paper creates ResearchMath-14k, a dataset of 14,056 research-level math problems curated from academic sources using a multi-agent pipeline, along with 220K reasoning trajectories and observations on model avoidance behaviors. A smart generalist might read it to see how large-scale data collection can help train AI systems on unsolved mathematical problems rather than textbook exercises.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"The multi-agent curation of genuinely research-level unsolved problems and the agentic filtering's retention of useful signals lack any reported validation or ablations.","rationale":"The reader's weakest assumption matches the load-bearing point exactly. The abstract alone already flags the pipeline as the unsupported link; nothing in the reported empirical result contradicts that assessment, so the UNVERDICTED / LOW verdict stands.","tokens_in":1723,"tokens_out":409,"duration_ms":26375,"concrete_test":"Draw a random sample of 100 problems from ResearchMath-14k; have two independent domain experts (not involved in the pipeline) label each as (a) research-level unsolved, (b) contest-level, or (c) previously solved. Separately, fine-tune the same Qwen3 models on the unfiltered ResearchMath-Reasoning trajectories using identical hyperparameters and report the average delta versus the filtered 9.2-point figure. If expert agreement on (a) is below 60% or the unfiltered delta is statistically indistinguishable, the central attribution weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (9.2-point average gain on Qwen3-4B to 30B after filtering ResearchMath-Reasoning) is interpreted as evidence that filtered open-problem attempts supply useful supervision. This interpretation requires that (1) the 14,056 problems are verifiably research-level and unsolved, and (2) the filtering step removes only avoidance behaviors while preserving genuine (even if incorrect) reasoning signals without introducing selection bias. The abstract describes a multi-agent pipeline operating on academic sources and notes avoidance behaviors plus increasing fake references in newer models, but supplies no human-expert validation of problem novelty, no overlap analysis with existing training corpora, and no control experiments (e.g., unfiltered trajectories or trajectories from standard contest problems). If either condition fails, the observed gains are compatible with ordinary data-augmentation effects rather than the claimed research-level supervision.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces ResearchMath-14K, a dataset of 14,056 research-level mathematical problems curated from academic sources via a multi-agent pipeline, claimed to be the largest such collection. It generates ResearchMath-Reasoning consisting of 220K teacher trajectories from two open models, documents avoidance behaviors (non-attempts, fabricated references) that increase in newer models, applies agentic filtering, and reports that fine-tuning Qwen3 models (4B–30B) on the filtered data yields an average 9.2-point improvement over base models. The work concludes that filtered open-problem attempts supply useful supervision even without fully correct traces and releases the dataset publicly.","tokens_in":1921,"tokens_out":636,"duration_ms":20569,"significance":"If the curation and filtering steps are shown to reliably isolate genuine research-level unsolved problems while preserving useful (even if imperfect) reasoning signals, the result would be significant for scaling mathematical reasoning beyond contest problems. The public release of ResearchMath-14K is a concrete strength that enables follow-on work. The observed scaling of avoidance behaviors across model generations is also a useful empirical observation.","major_comments":[{"comment":"Abstract: the central empirical claim of a 9.2-point average improvement after agentic filtering is presented without any description of the evaluation benchmarks, the number of test problems, statistical significance testing, or comparisons against unfiltered trajectories or standard contest-problem fine-tuning. This information is required to assess whether the gains exceed ordinary data-augmentation effects.","section":"Abstract"},{"comment":"Abstract and § on dataset construction: the claim that the 14,056 problems are verifiably research-level and unsolved rests on the multi-agent pipeline alone; no human-expert validation, novelty checks against existing corpora, or overlap analysis with training data are reported. If these problems overlap with standard training sets or are not genuinely open, the interpretation that the gains demonstrate utility of research-level supervision does not follow.","section":"Abstract"},{"comment":"Abstract: the filtering step is asserted to retain useful supervision signals while removing only avoidance behaviors, yet no ablation (e.g., performance with unfiltered trajectories or with trajectories from non-research problems) or human inspection of retained vs. discarded traces is described. This is load-bearing for the claim that imperfect open-problem attempts are beneficial.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract states the dataset size as 14,056 but does not clarify how many problems were initially collected versus retained after filtering; adding this breakdown would improve reproducibility.","section":"Abstract"},{"comment":"The observation that newer models produce 5.6× more references and 5.0× more fake references is interesting but would benefit from a precise definition of “fake reference” and inter-annotator agreement if any human labeling was involved.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback. We address each major comment below. We agree the abstract requires expansion and will revise it accordingly. For the dataset validation and filtering claims, we clarify our methodology while noting practical limitations on additional experiments.","responses":[{"response":"We agree that the abstract should provide more context on the evaluation. In the revised version, we will expand the abstract to name the evaluation benchmarks, indicate the number of test problems, note statistical significance testing, and reference comparisons against unfiltered trajectories and contest-problem fine-tuning.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central empirical claim of a 9.2-point average improvement after agentic filtering is presented without any description of the evaluation benchmarks, the number of test problems, statistical significance testing, or comparisons against unfiltered trajectories or standard contest-problem fine-tuning. This information is required to assess whether the gains exceed ordinary data-augmentation effects."},{"response":"The multi-agent pipeline sources problems from recent academic publications and applies agent-based verification to select likely unsolved research-level problems. We did not conduct or report human-expert validation, novelty checks, or overlap analysis with training data. We will add a limitations paragraph acknowledging this reliance on automated curation. The public release of the dataset enables independent verification by the community.","revision_made":"partial","referee_comment":"[Abstract] Abstract and § on dataset construction: the claim that the 14,056 problems are verifiably research-level and unsolved rests on the multi-agent pipeline alone; no human-expert validation, novelty checks against existing corpora, or overlap analysis with training data are reported. If these problems overlap with standard training sets or are not genuinely open, the interpretation that the gains demonstrate utility of research-level supervision does not follow."},{"response":"The manuscript reports gains from the filtered trajectories over base models but does not include the requested ablations on unfiltered trajectories, non-research problems, or human inspection of traces. These analyses were outside the scope of the current study. We will not add them in this revision and maintain that the observed improvements provide evidence for the utility of the filtered data even without perfect traces.","revision_made":"no","referee_comment":"[Abstract] Abstract: the filtering step is asserted to retain useful supervision signals while removing only avoidance behaviors, yet no ablation (e.g., performance with unfiltered trajectories or with trajectories from non-research problems) or human inspection of retained vs. discarded traces is described. This is load-bearing for the claim that imperfect open-problem attempts are beneficial."}],"tokens_in":1499,"tokens_out":570,"duration_ms":54864,"standing_objections":["Human-expert validation of the 14,056 problems as research-level and unsolved","Novelty checks against existing corpora and overlap analysis with training data","Ablations comparing filtered vs. unfiltered trajectories and human inspection of retained vs. discarded traces"]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main output is ResearchMath-14k, a public set of 14,056 problems drawn from academic sources with a multi-agent pipeline, plus 220k trajectories from open models. They also document that newer models generate more fabricated references.\n\nThe scale and the public release stand out as new. The trend on reference fabrication across model generations is a straightforward observation worth having on record.\n\nThe soft spots are in the validation. The abstract gives no human checks or overlap analysis to confirm the problems are actually unsolved research-level questions rather than contest or textbook material. The 9.2-point average gain after filtering is reported without benchmarks, baselines, or ablations, so it is hard to tell whether the filtering step adds anything beyond ordinary data effects.\n\nThe central interpretation—that filtered open-problem attempts supply useful supervision—depends on the curation being accurate and the filter preserving signal without bias. Neither is shown in the provided details.\n\nPeople building or testing math reasoning systems will want the dataset to experiment with. The work is honest about the data collection effort and the observed behaviors.\n\nIt deserves a serious referee to examine the curation pipeline and evaluation setup in full. I would send it for review.","headline":"The dataset release is the real contribution here, but the claims about curating unsolved research problems and the value of filtered trajectories rest on unverified assumptions.","tokens_in":2438,"tokens_out":324,"would_cite":false,"duration_ms":25184,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Agent-filtered attempts at unsolved math problems improve Qwen3 models by 9.2 points on average.","keywords":["research-level mathematics","agentic filtering","mathematical reasoning","dataset curation","language model fine-tuning","open problems","reasoning trajectories","Qwen3 models"],"falsifier":"Fine-tuned models show no gain over baselines when evaluated on a fresh set of unsolved problems that were never seen during curation or filtering.","tokens_in":2637,"feed_emoji":"📐","tokens_out":648,"duration_ms":25392,"temperature":0.7,"pith_summary":"The paper presents ResearchMath-14k, a collection of 14,056 research-level mathematical problems drawn from academic sources through a multi-agent curation process. It also produces ResearchMath-Reasoning, a set of 220,000 reasoning trajectories generated by open models, which exhibit patterns such as non-attempts and fabricated references that grow more common in newer models. After an agentic filtering step removes low-quality traces, fine-tuning Qwen3 models ranging from 4B to 30B parameters produces an average gain of 9.2 points over the untuned baselines. The result indicates that partial or incorrect attempts at open problems can still deliver effective training signals. The authors release the dataset to support further work on advanced mathematical reasoning.","feed_headline":"Filtered attempts at hard math problems lift models 9.2 points","feed_subtitle":"A 14,000-problem dataset from academic sources shows imperfect trajectories still train models effectively after agent filtering.","key_machinery":"The multi-agent pipeline that curates research-level problems from academic sources and filters teacher trajectories to retain useful supervision signals.","core_discovery":"A multi-agent pipeline can extract a large set of unsolved research problems and filter generated reasoning traces to yield useful supervision, delivering an average 9.2-point improvement when used to fine-tune Qwen3 models from 4B to 30B parameters even though the traces themselves are not fully correct.","pith_inferences":["The same curation and filtering method could be applied to open problems in other technical fields to generate training data.","Models might learn to produce more reliable attempts on frontier problems by repeated exposure to filtered traces.","Larger-scale versions of the dataset could test whether gains continue to grow with model size or data volume."],"forward_implications":["Filtered trajectories from open problems supply usable training data even without correct final answers.","Newer open models produce 5.6 times more references and 5.0 times more fabricated references per trace than earlier ones.","The approach scales across model sizes from 4B to 30B parameters with consistent average gains.","Public release of the 14k-problem set enables further experiments on research-level reasoning."],"fun_headline_variants":["14k research math problems improve Qwen3 models 9.2 points","Agent-filtered math traces improve open models 9.2 points","Research-level math dataset improves models by 9.2 points","Multi-agent pipeline creates math data improving models 9.2 points"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The multi-agent system accurately selects genuinely unsolved research problems and the filtering step keeps genuine learning value without adding bias or contamination.","fun_headline_variants_meta":{"raw":{"variants":["14k research math problems improve Qwen3 models 9.2 points","Agent-filtered math traces improve open models 9.2 points","Research-level math dataset improves models by 9.2 points","Multi-agent pipeline creates math data improving models 9.2 points"]},"model":"grok-4.3","cost_usd":0.008931,"raw_usage":{"total_tokens":4002,"prompt_tokens":644,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":89312000,"prompt_tokens_details":{"text_tokens":644,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3284,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":644,"tokens_out":74,"duration_ms":30169,"temperature":1.0,"reasoning_tokens":3284,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T13:19:04.925928+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Fine-tuned models show no gain over baselines when evaluated on a fresh set of unsolved problems that were never seen during curation or filtering.","supporting_citations":[],"review_version":1}