{"id":"87c410a7-a9c5-4cb3-b922-3109ba45ea55","arxiv_id":"2608.12841","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"AQuA claims that two sealed LLM research loops improve trading factors and models over time, but the headline evidence is partly a validation score and the code and data are withheld.","lead":"AQuA is a pair of separate language-model research systems for quantitative trading, one mining trading factors and one developing models, each meant to improve its own hypotheses from prior experiments. The paper reports strong simulated results, including a Sharpe ratio of about 2.5 on US equities, but key evidence is measured on validation data and neither code nor data is released.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Part I's headline 0.190 combined IC is the validation objective the loop optimizes, so the main evidence for recursive improvement may be adaptive overfitting rather than genuine research progress.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the fixed validation slice is treated as an unbiased measure after many iterations of selection on that slice, and Part I's headline combined IC is reported on that validation objective. My independent reading of Sections 3.1, 4.5, and 7 confirms that the paper never reports an untouched test-set number for Part I, so the central recursive-improvement claim is not established. The sealed-sandbox design and the candid failure-case appendix are genuine strengths, and Part II does report a held-out 2021-2025 test window with a positive Sharpe in every year, so the paper is not without merit. But the strongest claim, that both systems implement recursive self-improvement of the research process, relies on a metric that the system was allowed to optimize. A concrete test, scoring the final Part I signal on a never-returned slice, would settle whether the observed improvement is real or an artifact of adaptive validation selection. The reader's REJECT verdict is therefore appropriate; my concern reinforces it rather than changing it.","tokens_in":17085,"tokens_out":2640,"duration_ms":30018,"concrete_test":"Recompute the Part I combined IC on a data slice that was never returned to the agent during any iteration: freeze the final factor combination, score it once on a period excluded from all search, validation, and memory updates, and overlay that test IC on the Figure 3 curve. If the test IC does not rise with iterations or is materially below 0.190, the reported improvement is a validation-selection artifact rather than recursive self-improvement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that AQuA demonstrates recursive self-improvement rests mainly on Part I's iteration curve, but that curve is measured on the validation slice the search optimizes. Section 4.5 and Figure 3 report 'Combined Validation Spearman IC' reaching approximately 0.190, and Section 3.1 states that during search the harness returns to the agent only a score on a validation slice fixed in advance and uses that score to rank candidates. Thus the headline Part I number is the optimization objective itself, not an out-of-sample measure. The paper's defense against selection leakage is that the loop reports a metric it never optimizes; for Part I that defense is unavailable because no untouched test-set IC is reported anywhere. Section 7 concedes that test isolation rests on 'operator discipline rather than a hard technical barrier,' which is especially relevant if the reported number is merely the validation score. The alternative explanation, adaptive overfitting to a fixed validation slice through repeated selection, is well documented and would produce exactly the kind of monotone improvement shown in Figure 3 without any real improvement in research quality. Part II's test-window result is stronger, but it is a final model comparison, not an iteration curve, so it does not by itself establish recursive improvement. The load-bearing weakness is therefore that the principal quantitative evidence for the paper's central claim is the same metric the system was allowed to optimize.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces AQuA, a pair of independent LLM-driven research systems for quantitative investment research: one for symbolic factor discovery (Part I) and one for trainable model development (Part II). The central claim is that each system closes its own research loop by retaining validated evidence and using it to guide later proposals, thereby implementing recursive self-improvement at the level of the research process, while a sealed sandbox prevents data leakage by construction. Part I reports a combined validation Spearman IC of about 0.190 on a crypto universe; Part II reports a per-stock IC of +0.0843 on US equities and a threshold long/short held-out Sharpe of up to +2.50 at a two-leg cost, positive in every year from 2021 to 2025. The paper also reports a fully causal walk-forward Sharpe of about +2.0 for Part II.","tokens_in":17364,"tokens_out":5125,"duration_ms":53392,"significance":"If fully supported, the paper would make a useful conceptual contribution by separating leakage into a generation channel and a selection channel and by insisting that the metric used for ranking candidates differ from the metric reported as the result. The Part II evaluation is a genuine strength: the 2021–2025 test window is untouched by model selection, and the fully causal walk-forward procedure addresses a common failure mode in backtesting. The worked failure case in Appendix B and the candid Limitations section are also valuable. However, the principal quantitative evidence for Part I's self-improvement is the validation objective itself, and Part II's held-out result is a single final point rather than evidence of improvement over iterations. The central claim of recursive self-improvement is therefore not yet established with the evidence presented.","major_comments":[{"comment":"The Part I headline result, a combined validation IC of approximately 0.190, is measured on the validation slice that the search loop uses to rank candidates and select factors (Section 3.1). No untouched test-set IC for Part I is reported anywhere. The paper's own defense against selection leakage, 'reporting a metric the loop never optimizes against' (Section 6), does not apply to Part I because the reported 0.190 is the optimized validation metric itself. The monotone improvement in Figure 3 is exactly what adaptive overfitting to a fixed validation slice would produce. Please report a Part I evaluation on a test window that is frozen before the final combined signal is fixed and never returned to the agent, and disclose the number of candidate factors and selection rounds so the multiple-testing burden can be assessed.","section":"Section 4.5, Figure 3, Section 3.1"},{"comment":"Part II's untouched 2021–2025 test window is a genuine out-of-sample result, and the walk-forward Sharpe is a strength. However, a single final test point cannot by itself establish recursive self-improvement. The paper does not report how many configuration diffs were tried, how many autonomous iterations the loop ran, or how validation or test performance evolved over those iterations. Without this information, the result is compatible with a broad search that found one good configuration rather than an autonomous loop whose research process improves over time. Please report the iteration trajectory and the total number of configurations evaluated, and, if possible, the test-window performance of the final few iterations.","section":"Section 5.1, Section 5.6, Table 3"},{"comment":"The abstract and Section 6 claim that the systems are leakage-free 'by construction' and that reported results are out of sample, but Section 7 concedes that keeping the final test window out of selection 'rests on the sealed protocol and operator discipline rather than a hard technical barrier.' This caveat is especially important for Part I, where no final test window is designated at all. The unqualified language in the abstract and Section 6 should be revised to state that Part I's headline IC is a validation-slice number and that the isolation of any final test evaluation is a governance property rather than a cryptographic guarantee.","section":"Abstract, Section 7, Section 6"}],"minor_comments":[{"comment":"The legend includes a 'Fit-to-validation IC range' that is never defined in the text; please explain how this range is computed and why it is reported alongside the combined validation IC.","section":"Figure 3"},{"comment":"Several important experimental details are missing: the crypto universe is not specified (exchange, coin list, data vendor), the baseline implementations in Figure 3 and Table 2 are not accompanied by hyperparameters or adaptation details, and the exact feature set and convolutional configuration for Part II are explicitly left undisclosed. This limits reproducibility.","section":"Sections 4 and 5"},{"comment":"The 'expression: withheld' entries mean that no actual factor formula is disclosed for Part I, so an independent reader cannot inspect, reproduce, or falsify the specific factor claimed to carry the signal.","section":"Listing 1 and Appendix A"},{"comment":"The sentence describing the split as 'train on 2010–2019, leave 2020 as an embargo gap' is clear, but the same paragraph says 'the exact feature set, normalization, and label construction are part of that sandbox and are not disclosed'; please clarify whether the label is the 30-minute forward return mentioned in Section 5.3 or a different construction.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing about this paper. First, the sealed-sandbox design — where the agent can only emit DSL specifications and the evaluator is outside its reach — is a real contribution to agentic research integrity. The two-channel leakage split (generation vs. selection) is useful well beyond finance. Second, the paper's own evidence doesn't support the \"recursive self-improvement\" claim: Part I's headline IC of 0.190 is measured on the validation slice the loop optimizes, and the rise in that curve is exactly what adaptive overfitting would produce. The paper explicitly defends against selection leakage by reporting a metric the loop never optimizes, but for Part I no untouched test IC is reported anywhere, so that defense is unavailable.\n\nWhat the paper does well: the failure case in Appendix B is first-rate. An LLM reviewer approves a volume-participation feature whose denominator includes future bars, and the authors respond by making such an error inexpressible rather than relying on another LLM to catch it. That is honest and instructive. Part II also reports a genuinely untouched 2021–2025 test window, and the strategy being positive in 2022 is a point in its favor. The limitations section is candid about test isolation resting on operator discipline rather than a hard barrier.\n\nSoft spots, in proportion: Part I's circularity is the load-bearing issue. Part II gives a final model comparison but no iteration curve, no count of search iterations, no multiple-testing correction, no error bars on the Sharpe, and no code, data, or exact configurations. The feature set is withheld. These are not minor omissions; they make the quantitative claims hard to verify. But the design idea remains valuable despite the weak empirical support.\n\nWho this is for: anyone building autonomous research loops, especially with an eye on leakage and evaluation integrity. It deserves a serious referee, not a desk reject, because the design contribution is real and the authors are thoughtful. The referee should demand an untouched test-split IC for Part I, an iteration curve and search budget for Part II, and code or at least detailed configuration logs. I'd send it out with clear requests for those additions.","headline":"The sealed-sandbox design and honest failure case are genuinely valuable, but the headline evidence for recursive self-improvement is the validation metric the loop optimizes, so the central claim is not established as stated.","tokens_in":17879,"tokens_out":2413,"would_cite":true,"duration_ms":26002,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two autonomous loops recursively improve quantitative research without leaking data, reaching a held-out Sharpe of +2.50.","keywords":["recursive self-improvement","LLM agents","quantitative trading","factor discovery","model development","data leakage prevention","sealed sandbox","information coefficient"],"falsifier":"Run the Part II loop a second time with the same seeds but return the test-window score during search instead of validation only; if the reported 2021–2025 IC and Sharpe survive that change, the sealed test window was not doing the work, and if they collapse, the headline numbers depend on never seeing the test window. Separately, audit the operator's access logs to the test store during the search: any read before the configuration is frozen falsifies the test-isolation claim.","tokens_in":16893,"feed_emoji":"📈","tokens_out":8303,"duration_ms":67424,"temperature":0.7,"pith_summary":"This paper tries to establish that an autonomous system can improve its own quantitative-investment research when the definition of success is fixed and the agent cannot reach the data or the evaluator. It describes two separate loops: one discovers symbolic trading factors, the other develops a trainable model, and each retains validated evidence that guides its next proposal. In that bounded sense the paper claims recursive self-improvement happens at the level of the research process, not by rewriting the language model or the test. The reported payoffs are a combined factor information coefficient of about 0.190 on crypto data and a per-stock information coefficient of +0.0843 on US equities, which converts to a held-out long/short Sharpe of up to +2.50 at a two-leg cost.","feed_headline":"AI agents self-improve to a held-out Sharpe of +2.50","feed_subtitle":"Sealed sandboxes stop data leakage while two research loops — factors and models — each close their own loop.","key_machinery":"The load-bearing object is the sealed sandbox: a fixed tuple of data splits, feature and label definitions, and evaluator, paired with a domain-specific language whose expressions cannot reach those sealed components. The paper defines causality as closed under composition in the factor DSL, meaning the agent cannot express a future-looking normalizer or label even by accident. The second half of the mechanism is the split metric: the loop optimizes a validation score but reports a test metric computed once on an untouched window, so the number the agent is judged by is never the number it selects on.","core_discovery":"In AQuA's design, recursive self-improvement is made safe by asymmetric freedom: the agent can explore freely inside a restricted specification language, but every admissible action is compiled and scored by a sealed, human-authored sandbox whose data splits, features, labels, and evaluator the agent cannot modify. Because every time-series operator reads only a trailing window and every cross-sectional operator only the current timestamp, any expression the agent assembles is causal by construction. Selection leakage is handled separately: during search the agent receives only a fixed validation-score signal, while the designated test window is scored once after the configuration is frozen and never returned to the loop. The paper reports that this design lets the factor system reach a combined validation information coefficient of about 0.190 and lets the model system beat a GRU baseline (per-stock IC +0.0843 vs +0.0613) and produce a sector-neutral, volatility-targeted long/short book with a held-out Sharpe of +2.50, positive in every year from 2021 to 2025.","pith_inferences":["Beyond finance, the two-channel split (generation leakage versus selection leakage) is a general recipe: any LLM-driven discovery loop should make the evaluator unreachable and report a metric the loop never optimizes.","A testable extension is to track the gap between the validation IC the loop optimizes and the untouched test IC over successive iterations; a widening gap would reveal residual selection leakage even under a nominally sealed test window.","The paper itself flags coupling the two systems as the next step, but also notes that coupling introduces a new leakage channel; a concrete corollary is that the factor library must be frozen before model search begins if the model loop is not to steer factor selection.","The demonstrated autonomy is bounded by a human operator who sets goals, owns the sandbox, and supervises promotion, so the results do not imply unattended operation on new markets or frequencies without re-tuning."],"forward_implications":["If the reported results hold, an autonomous research loop can accumulate validated evidence across iterations and improve factor and model quality without human review of each candidate.","The test-window numbers (per-stock IC +0.0843, Sharpe +2.50, positive in every year from 2021 to 2025) would be out of sample with respect to the whole search, not just the final model fit.","The sealed-sandbox recipe would make leakage-inducing actions structurally unavailable in any autonomous research agent, rather than relying on a reviewer to catch subtle errors.","Causality by construction in the factor DSL means every expression the agent can emit is a legitimate candidate, so the loop can run without per-candidate human review within its fixed research contract.","The fully causal walk-forward evaluation retains a Sharpe of about +2.0, implying the headline result is not the artifact of one favorable configuration or parameter choice."],"supporting_citations":[{"why":"Quantifies the probability of backtest overfitting, motivating structural sealing of data splits.","marker":"Bailey et al. (2017)"},{"why":"Documents how backtest overfitting manufactures convincing but non-reproducible results, supporting the need for held-out isolation.","marker":"Bailey et al. (2014)"},{"why":"Shows that adaptive reuse of a fixed holdout can overfit it, motivating the split-metric design.","marker":"Blum and Hardt (2015)"},{"why":"Provides the standard skepticism toward unusually strong quantitative results that the sealed sandbox is designed to survive.","marker":"Harvey et al. (2016)"},{"why":"Supplies the formulaic-alpha operator vocabulary that underlies the Part I factor DSL.","marker":"Kakushadze (2016)"},{"why":"Establishes operator-based genetic factor search that Part I extends with language-model agents.","marker":"Cui et al. (2021)"},{"why":"Contributes the factor expression-tree representation and evolutionary search lineage Part I builds on.","marker":"Zhang et al. (2020)"},{"why":"Provides the LSTM sequence-modeling baseline against which the Part II hybrid model is compared.","marker":"Hochreiter and Schmidhuber (1997)"},{"why":"Introduces the GRU, the strongest baseline the Part II model must beat under the same evaluator.","marker":"Cho et al. (2014)"}],"fun_headline_variants":["Self-improving agents hit Sharpe 2.5 without leakage","AQuA: self-improving factor and model loops reach Sharpe 2.50","Sealed sandboxes let agents self-improve to 2.50 Sharpe","Causal-by-construction agents reach Sharpe 2.50 with no leakage"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The fixed validation slice that the harness returns to the agent during search is assumed to remain an honest measure of out-of-sample skill after many rounds of selection on that same slice, and the paper concedes that keeping the final test window untouched rests on operator discipline rather than a hard technical barrier.","fun_headline_variants_meta":{"raw":{"variants":["Self-improving agents hit Sharpe 2.5 without leakage","AQuA: self-improving factor and model loops reach Sharpe 2.50","Sealed sandboxes let agents self-improve to 2.50 Sharpe","Causal-by-construction agents reach Sharpe 2.50 with no leakage"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000715,"raw_usage":{"total_tokens":3247,"prompt_tokens":1011,"completion_tokens":2236,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":2154}},"tokens_in":627,"tokens_out":2236,"duration_ms":15893,"temperature":1.0,"reasoning_tokens":2154,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:09:58.252041+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Part II loop a second time with the same seeds but return the test-window score during search instead of validation only; if the reported 2021–2025 IC and Sharpe survive that change, the sealed test window was not doing the work, and if they collapse, the headline numbers depend on never seeing the test window. Separately, audit the operator's access logs to the test store during the search: any read before the configuration is frozen falsifies the test-isolation claim.","supporting_citations":[],"review_version":1}