{"id":"2f5c24f6-629d-4fbf-b5d3-6ced481d627d","arxiv_id":"2608.09988","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"OpenPM is an auditable point-in-time evaluation benchmark for LLM portfolio agents, whose 44-day case study shows analyst quality matters more than the constructor model and equal-weighting is a hard baseline.","lead":"OpenPM is an evaluation framework that makes LLM portfolio-management agents prove they only use information available at decision time, obey typed risk limits, and show cost sensitivity. In a 44-day S&P 500 case study, the paper finds analyst quality matters more than constructor choice, and equal-weighting is a hard baseline.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Contamination certificate's PASS rests on provider-stated knowledge cutoffs and externally supplied timestamps, neither independently verified; Opus-4.7's cutoff is only two months before the window, so model-memory leakage could invalidate the central PIT guarantee.","rationale":"The empirical case-study findings are fragile—one window, one decision date, INTC-driven, with gemini excluded—but the paper labels them as snapshot evidence and upper bounds, so that fragility is disclosed rather than hidden. The leakage guarantee is different: it is the reason the framework is proposed as 'auditable point-in-time evaluation.' The gate and certificate can only be as strong as the provenance timestamps and the model-memory assumption they stand on. No software gate can detect information that is in the model weights rather than in a gated record, and no provider-stated cutoff is a substitute for a behavioral test. This is not an accusation of leakage; it is a precise boundary on what the certificate can mean. The reader's CONDITIONAL verdict is the right one: the framework and its artifacts are reproducible, the authors are explicit about scope, and the benchmark is a useful contribution, but the headline guarantee should be accompanied by a temporal probe and an independent provenance audit, or the certificate status should be downgraded to UNKNOWN. Therefore I keep the reader's verdict unchanged.","tokens_in":15654,"tokens_out":9938,"duration_ms":104204,"concrete_test":"Run a blinded temporal probe on each constructor, especially Opus-4.7, using 100 multiple-choice questions about public market events dated between 2026-01-01 and 2026-03-01, i.e., after the stated cutoff and before the evaluation window, covering facts such as specific 8-K acceptance dates, index levels, and major headlines, with four plausible options each. Query each model through the same serving endpoint at temperature 0. If accuracy is significantly above chance under a binomial test, the model has post-cutoff knowledge and the §6.5 PASS cannot certify no leakage; if accuracy is at chance, the PASS claim still needs an independent audit of the external availability stamps before being reported as a guarantee.","verdict_should_be":"UNCHANGED","load_bearing_attack":"OpenPM's core guarantee is that no agent-visible information is used before it was available. The gate enforces this only for records whose availability stamps are supplied by third parties (EDGAR acceptanceDateTime, FRED/ALFRED vintages, GDELT publish+30 min), and the contamination certificate in §6.5 additionally takes each constructor's stated knowledge cutoff as ground truth. Neither anchor is independently tested. The sharpest case is Opus-4.7: its ~2026-01 cutoff is only about two months before the evaluation window opens, so if its weights or serving pipeline contain post-cutoff information about 2026-03, the certificate can return PASS while the model leaks the future through memory rather than through a gated record. The paper's Limitations correctly says leakage is 'reduced and bounded rather than proven,' but the abstract and §6.5 still advertise PASS certificates and 'prevent look-ahead leakage.' Because the framework's entire rationale is auditable point-in-time evaluation, this unverified trust anchor is the load-bearing assumption: a PASS certificate currently certifies that the pipeline's own bookkeeping is consistent, not that no post-cutoff information reached the agent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents OpenPM, an evaluation framework and benchmark for LLM portfolio-management agents. Its central design is a per-record availability gate: any evidence record enters the agent-visible state at decision time T only if its availability timestamp is before or at T. Natural-language risk mandates are compiled into typed constraints and enforced by a deterministic critic on the executed portfolio, and each run emits audit artifacts including a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report. A reference 'tiered allocator' separates six typed analyst LLMs, a constructor LLM, and a deterministic risk-projection layer. A 44-day S&P 500 case study with once-then-hold and daily rebalancing reports that constructor gains over equal-weighting are modest and model-dependent, that analyst quality is the larger lever, and that turnover is the main cost driver. The authors repeatedly and explicitly frame all returns as no-market-impact upper bounds on a single frozen window, not validated alpha.","tokens_in":15822,"tokens_out":10480,"duration_ms":107648,"significance":"If OpenPM delivers on its claims, it is a timely and useful contribution: it directly targets real evaluation pitfalls (look-ahead leakage, optimistic fills, unenforced mandates), ships reproducible code and data, and factors an agent into analyst, constructor, and enforcement stages whose contributions can be isolated. The deterministic risk critic and the availability gate are principled and the paper is unusually candid about the upper-bound nature of its results. The framework is the main deliverable; the case study is illustrative. However, the leakage guarantee currently depends on trust anchors that are not independently verified, and the empirical findings are sensitive to a single stock and a single window. These issues do not invalidate the framework, but they must be addressed before the auditable point-in-time claim is fully supported.","major_comments":[{"comment":"The contamination certificate is reported as PASS for all constructors, including claude-opus-4.7, whose stated knowledge cutoff (~2026-01) is only about two months before the evaluation window opens (2026-03-02). The certificate treats this cutoff and the external availability stamps (EDGAR acceptanceDateTime, FRED/ALFRED vintages, GDELT publish+30-min) as ground truth without independent verification. The paper's own Limitations correctly says the gate and certificate 'reduce and bound leakage risk rather than proving that every possible configuration is leakage-free,' but the abstract and §6.5 still advertise a PASS certificate and 'prevent look-ahead leakage.' As written, PASS certifies that the pipeline's bookkeeping is internally consistent, not that no post-cutoff information reached the model through training data or serving updates. Please either qualify the certificate and abstract accordingly, add a margin-based UNKNOWN status (e.g., a cutoff within X months of the window makes the certificate UNKNOWN), or include a validation protocol that adversarially tests the gate by injecting records with incorrect timestamps and checking that the certificate fails. This is load-bearing because auditable point-in-time evaluation is the paper's central claim.","section":"§3.2, §5, §6.5, Abstract"},{"comment":"The claim that same-pool equal-weighting is a 'hard baseline' depends on the uniform analyst aggregation weights wa=1, which the paper says were adopted after an IC-weighted variant 'failed out of sample' (§4). No details are given about the validation split, the IC weights attempted, or whether the uniform choice was made on data overlapping the evaluation window. Since the same-pool EW baseline is then used as the benchmark against which constructor skill is measured, this selection is a free parameter tuned on or near the evaluated data. Please report the IC-weighting experiment, the validation protocol, and the sensitivity of the EW baseline to the aggregation weights. Without this, 'equal-weighting ... is a hard baseline' is not the parameter-free statement it appears to be.","section":"§4, §5, §6.2"},{"comment":"The empirical conclusion that 'analyst quality matters more than constructor choice' is heavily concentrated in a single name: in the case study, INTC contributes $104k of $110k total P&L for gpt-5 on one capture (Table 6), and the analyst-capture swap in Table 7 flips gpt-5's INTC weight from 0% to 8.5% and net return from +3.1% to +11.5%. There are no error bars, bootstrap intervals, or leave-one-name-out checks, and only one 44-day window. The paper does hedge in Limitations and §7, but the abstract and §6.2 present the analyst-vs-constructor claim as a finding. If the empirical results are to remain in the abstract, please add a robustness check (e.g., leave-one-name-out attribution, bootstrap over bars/seeds, or a second window) or explicitly relabel them as illustrative single-window observations rather than findings.","section":"§6.1, §6.2, §D.1, Table 6"}],"minor_comments":[{"comment":"The citation (Song et al., 2024) for the claim that audits 'reduce and bound leakage risk' is a paper on cyber-threat monitoring and appears unrelated; please replace it with a relevant contamination or audit reference, or justify the connection.","section":"Limitations"},{"comment":"The statement that the constructor ordering is 'stable when pooled across K' is per-capture sensitive: Table 2 shows DeepSeek-V3.2 outperforming gpt-5 and Opus on the gpt-5 analyst capture. Please clarify that the ordering is pooled over six captures and three K values, not stable within each analyst capture.","section":"§6.1, Table 2"},{"comment":"The sentence in §5 says the data are 'after every evaluated backbone's knowledge cutoff.' Given the thin margin for Opus-4.7, please state the exact cutoff dates used and the source of those dates in the released data card.","section":"§5, §6.5"},{"comment":"All Sharpe ratios are annualized from a 44-day window, and the tables report three-seed means or single runs without standard errors; a brief note in the caption that these are high-variance point estimates would help prevent over-reading.","section":"Table 2, Table 3"},{"comment":"The table would benefit from stating whether the DeepSeek analyst capture's lower INTC rank is due to lower confidence or contradictory signals; this would help readers judge whether the analyst-quality finding is about information content or calibration.","section":"Appendix D, Table 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and well-scoped, and the framework is a genuine step forward. My main concern is that the contamination certificate's PASS status overstates what is actually verified; this is fixable by rewording, adding an UNKNOWN margin rule, or adding an adversarial test of the gate. The baseline-weight selection also needs transparency. The empirical section is thin but acceptable as a case study if the claims are appropriately relabeled. I would not reject, but I would require these changes before the auditable point-in-time guarantee is presented as established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nOpenPM is worth your attention as a benchmark design, not as evidence that LLM analysts beat markets. The framework is careful: per-record availability gating, a deterministic critic that enforces the mandate on the executed weights, byte-identical analyst captures replayed across constructors, and audit artifacts per run. The authors are honest about upper bounds and single-window limits. That puts it ahead of the typical trading-agent paper.\n\nThe soft spots are real. The contamination certificate rests on external availability stamps and on the stated knowledge cutoffs of the backbones. Neither is independently verified. Opus-4.7's cutoff is two months before the window, so the certificate can PASS while future information arrives through model memory. The paper concedes leakage is bounded, not proven, but the abstract still says 'prevent look-ahead leakage.' A careful referee should force language that matches the guarantee.\n\nThe empirical headline—analyst quality matters more than constructor—is fragile. It is essentially one stock, INTC, in one 44-day window, with no error bars. The case study itself shows INTC is 104k of 110k P&L. The three-seed structure helps, but the gemini-2.5-flash exclusion is post-hoc. That does not kill the framework, but it means the paper's main claim to a non-expert reader is over-stated.\n\nOne more thing: the uniform analyst weighting was chosen after the IC-weighted variant failed out of sample. That is model selection on the same window. The paper mentions it, but it should be treated as a free parameter.\n\nBottom line: the framework deserves a serious referee and probably publication after revision. The empirical claims need robustness work—longer window, more stocks, error bars, and ideally an independent check of availability stamps. I would cite this for the evaluation protocol, and I'd bring it to reading group for the design discussion.","headline":"Well-engineered benchmark framework; the empirical claim about analyst quality is fragile because it rests on one stock and a short window.","tokens_in":16462,"tokens_out":2021,"would_cite":true,"duration_ms":21577,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Analyst quality, not constructor, drives LLM portfolio gains","keywords":["point-in-time evaluation","LLM portfolio management","look-ahead leakage","contamination certificate","risk mandate enforcement","deterministic critic","equal-weight baseline","turnover costs"],"falsifier":"Plant a deliberate leak: at decision time $T$, feed the agent a news article whose body contains a fact first reported after $T$ but whose timestamp is $T$ minus one minute, and check whether the contamination certificate flips to FAIL; if it stays PASS, the gate is not actually binding. A second check is to ask each backbone about an event from late in its stated cutoff window and compare its answers to the certificate's cutoff audit.","tokens_in":15387,"feed_emoji":"📈","tokens_out":5651,"duration_ms":52214,"temperature":0.7,"pith_summary":"OpenPM is an evaluation framework built to stop LLM portfolio-management agents from looking better than they are. Its central claim is that three common failure modes—information that leaks from the future, fills priced more favorably than reality, and risk mandates that are described but never enforced—can be controlled by engineering: a per-record availability gate, side-aware execution pricing, and a deterministic critic that projects every proposed portfolio onto typed constraints. The paper also argues, from a case study that replays byte-identical analyst evidence across five constructor models, that the quality of the upstream analyst evidence moves returns more than the choice of the constructor, and that equal-weighting the same candidate pool is a hard baseline. The claimed payoff is that every reported number ships with audit artifacts—a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report—so results can be trusted as upper bounds rather than as validated alpha.","feed_headline":"Analyst quality, not constructor, drives LLM portfolio gains","feed_subtitle":"OpenPM's availability gate and deterministic risk critic make the backtest auditable; equal-weight is the baseline to beat.","key_machinery":"The carrying mechanism is the tiered allocator running inside a point-in-time availability gate. Every record the agent sees carries an availability timestamp and enters the state at decision time $T$ only if $\\text{ts\\_available} \\le T$, so filings are stamped by EDGAR acceptance time, macro series by FRED/ALFRED vintage, and news by publish-plus-thirty-minutes. A natural-language mandate is compiled once into typed RiskConstraints, and a deterministic in-loop critic projects the constructor's proposed weights onto the feasible set, so feasibility is guaranteed by projection rather than by trusting the model. The constructor-isolation design freezes analyst outputs into captures and replays them across constructor models, which is what lets the paper attribute performance differences to the constructor alone. Each run emits a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report.","core_discovery":"On the paper's own terms, the discovery is that LLM construction skill is real but conditional and localized upstream. Holding the six-analyst evidence byte-identical, four of five constructor models beat same-pool equal-weighting on a strong analyst capture, with the strongest adding about five points of net return, while on a weaker capture none of them beat equal-weight. Constructor ordering is stable across pool sizes, and the strongest constructors mainly re-weight the equal-weight roster rather than replacing it. The paper reads this as evidence that analyst quality matters more than constructor choice, that same-pool equal-weighting is the baseline any constructor must clear, and that turnover—not spread—is the dominant cost at daily cadence. All returns are presented as single-window, no-market-impact upper bounds, with the caveat that a single trending name can drive much of the spread between constructors.","pith_inferences":["If external timestamps are wrong—say an acceptance time is back-dated or a model's cutoff is mis-stated—the gate can certify a run that still leaked; a natural extension is cross-checking stamps against independent sources.","The analyst-over-constructor ordering is demonstrated on one 44-day window with a handful of trending names; extending the replay design across regimes and dates would show whether it generalizes.","The same availability-gate and audit-artifact pattern transfers to other open-ended LLM agents, such as web browsing, retrieval, and tool use, where unavailable information can leak into state.","Per-name attribution shows a single name can dominate the constructor gap; a stress-test that blanks the top-attribution name would quantify how fragile each reported ordering is."],"forward_implications":["Any reported LLM portfolio return should be read as an upper bound unless it is accompanied by a contamination certificate and a cost-sensitivity curve; OpenPM makes both routine artifacts.","Compute should be spent first on the analyst and evidence tier; swapping analyst captures moved constructors more than any constructor choice did.","Same-pool equal-weighting is the baseline to beat, and it is hard: on weaker analyst evidence no constructor cleared it.","At daily cadence, turnover is the main cost driver, so a disciplined low-turnover constructor can beat SPY and cash net of cost while high-churn models fall below SPY.","Natural-language risk mandates should be compiled to typed constraints and enforced deterministically at execution time, not scored post hoc."],"supporting_citations":[{"why":"Establishes the need to measure LLM benchmark contamination per benchmark, which the contamination certificate operationalizes.","marker":"(Sainz et al., 2023)"},{"why":"Documents backtest overfitting and the inflation it causes, motivating the leakage controls and the no-alpha caveat.","marker":"(Bailey et al., 2014)"},{"why":"Supplies the point-in-time and leakage-prevention methodology that the availability gate instantiates.","marker":"(López de Prado, 2018)"},{"why":"StockBench is the contamination-free, cost-aware benchmark that OpenPM extends to S&P 500 scale and sub-daily cadence.","marker":"(Chen et al., 2025)"},{"why":"DeepFund shows the live post-cutoff failure mode that point-in-time gating is designed to prevent.","marker":"(Li et al., 2025a)"},{"why":"PortBench's finding that most model-profile pairs fail to beat equal-weight supports OpenPM's equal-weight baseline.","marker":"(Zhao et al., 2026)"},{"why":"Provides the Sharpe-ratio statistics that justify reading single-window results as orderings, not validated alpha.","marker":"(Lo, 2002)"}],"fun_headline_variants":["Analyst quality, not constructor choice, drives LLM portfolio gains","Auditable backtesting: analyst quality beats constructor model for LLM agents","LLM portfolio agents: analyst quality matters more than constructor choice","OpenPM audit: analyst quality, not model, is key for LLM trading gains","LLM portfolio returns hinge on analyst evidence, not constructor model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The leakage guarantee rests on external availability stamps being accurate (EDGAR acceptance times, FRED/ALFRED vintages, GDELT publish lags) and on each model's stated knowledge cutoff being honest; if a timestamp or cutoff is wrong, the contamination certificate can pass even though future information reached the agent.","fun_headline_variants_meta":{"raw":{"variants":["Analyst quality, not constructor choice, drives LLM portfolio gains","Auditable backtesting: analyst quality beats constructor model for LLM agents","LLM portfolio agents: analyst quality matters more than constructor choice","OpenPM audit: analyst quality, not model, is key for LLM trading gains","LLM portfolio returns hinge on analyst evidence, not constructor model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000886,"raw_usage":{"total_tokens":3831,"prompt_tokens":958,"completion_tokens":2873,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":2779}},"tokens_in":574,"tokens_out":2873,"duration_ms":20230,"temperature":1.0,"reasoning_tokens":2779,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:49:59.361245+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Plant a deliberate leak: at decision time $T$, feed the agent a news article whose body contains a fact first reported after $T$ but whose timestamp is $T$ minus one minute, and check whether the contamination certificate flips to FAIL; if it stays PASS, the gate is not actually binding. A second check is to ask each backbone about an event from late in its stated cutoff window and compare its answers to the certificate's cutoff audit.","supporting_citations":[],"review_version":1}