{"id":"77f8628e-fa42-4390-b281-d57b74fb40fa","arxiv_id":"2508.11152","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A role-based multi-agent LLM system is evaluated for stock selection and portfolio construction against established benchmarks.","lead":"This paper tests whether a team of specialized AI agents, built on large language models, can pick stocks and build investment portfolios. The authors compare the agents' stock choices against standard benchmarks at different risk levels.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM training-data contamination may invalidate the backtest; abstract does not show that the evaluation period postdates model knowledge cutoff.","rationale":"The reader's weakest assumption correctly identified the fairness and bias of the evaluation setup as unverifiable from the abstract alone. I agree with that. However, there is a more specific and load-bearing concern that the reader did not name: LLM training-data contamination, a form of lookahead bias unique to LLMs. The abstract does not provide any information about model cutoff dates or evaluation periods, leaving this risk open. This concern does not change the verdict: it remains UNVERDICTED because the full text is unavailable and the claim cannot be trusted without this information. My recommendation is to keep the reader's verdict as UNVERDICTED rather than moving to ACCEPT or REJECT, since the concern is about missing evidence rather than observed counter-evidence. I have provided a concrete test that could settle the contamination worry if the full text were available, and I acknowledge that if the authors used post-cutoff evaluation, the concern would dissolve.","tokens_in":541,"tokens_out":2321,"duration_ms":29823,"concrete_test":"Obtain the full text and identify the exact evaluation date range and the LLM versions used. First, verify that the entire evaluation window is after the model's knowledge cutoff (or that the authors quarantined a dataset downloaded post-cutoff). Second, as an empirical test, rerun the same agent protocol on a held-out period 12 months after the stated knowledge cutoff using point-in-time market data; if the outperformance over benchmarks shrinks or reverses, contamination is likely. Third, inspect agent outputs for stock mentions or rationale that reference future events (e.g., earnings surprises, price movements) that would not have been known at the selection date.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a role-based multi-agent LLM system outperforms established benchmarks. For this to hold, the backtest must be truly forward-looking. A distinctive and plausible threat specific to LLM agents is training-data contamination: LLMs are pretrained on web text that includes historical stock prices, news, and corporate actions up to a knowledge cutoff. If the evaluation window overlaps with or predates that cutoff, the agents may recall what actually happened to stocks in the test period, so their 'predictions' are not genuine. The abstract does not disclose the evaluation period, the specific LLM versions, or their knowledge cutoffs, making this risk unassessable. If contamination occurred, the reported outperformance could be an artifact of memorization rather than multi-agent reasoning. This is a load-bearing concern because it directly attacks the validity of the performance evidence, not merely a peripheral implementation detail. Additional concerns like benchmark cherry-picking and transaction costs exist, but the contamination mechanism is the most severe and is unique to LLM-based agents.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes 'AlphaAgents,' a role-based multi-agent LLM system for stock selection and equity portfolio construction. The abstract claims that a team of specialized agents can analyze equities and outperform established benchmarks under varying risk tolerance, and that the study provides insights into the practical efficacy of multi-agent frameworks. However, the abstract presents no quantitative results, no evaluation window, no benchmark definitions, no model versions, and no methodology. The full text is not available; this review is abstract-only.","tokens_in":765,"tokens_out":1661,"duration_ms":21796,"significance":"If the performance claims were substantiated, the work would be significant: it would provide evidence that multi-agent LLM systems can add value in quantitative finance, an area where LLM adoption is growing but rigorous validation is scarce. The multi-agent role specialization idea is interesting and plausible. However, the current abstract provides no verifiable evidence, no reproducibility artifacts, and no falsifiable predictions. The significance of the claimed result is thus entirely conditional on unpublished evaluation details. There is no machine-checked proof, code, or data to assess, so the paper's contribution cannot currently be evaluated.","major_comments":[{"comment":"The central claim—that the multi-agent system outperforms established benchmarks—is stated without any supporting data. There are no numerical results, no performance metrics, no benchmark list, and no evaluation period. As written, the claim is unverifiable. For a portfolio-construction paper, the abstract should at least report the main comparison (e.g., Sharpe ratio, alpha, or return) and the test window; otherwise the stated conclusion is unsupported.","section":"Abstract"},{"comment":"The evaluation setup is not described, so a load-bearing contamination risk remains. LLM agents may have been pretrained on data that includes the evaluation period, making 'predictions' reflect memorization rather than reasoning. The abstract does not disclose the LLM versions, their knowledge cutoffs, or whether the evaluation window postdates those cutoffs. This is not a peripheral detail: if contamination occurred, any reported outperformance is artifactual. The authors must disclose the evaluation period, model versions, and a contamination-mitigation strategy.","section":"Abstract"},{"comment":"The claim of practical efficacy for portfolio construction is incomplete without accounting for real-world frictions. The abstract mentions 'varying levels of risk tolerance' but gives no indication of whether the backtest includes transaction costs, liquidity constraints, or survivorship-bias controls. These factors are first-order for equity portfolio construction and can easily reverse an apparent outperformance. The absence of any such detail in the abstract makes the practical-effectiveness claim premature.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'varying levels of risk tolerance' is vague; it is unclear whether this refers to different utility functions, portfolio constraints, or benchmark comparisons. A concrete specification would help readers understand the scope.","section":"Abstract"},{"comment":"The abstract does not cite any prior work on LLM-based trading agents or multi-agent finance, making it difficult to place the claimed contribution in context. At least representative references should be added in the full text.","section":"Abstract"},{"comment":"The abstract says 'We present a comprehensive analysis' but no analysis details are shown. Minor wording such as 'preliminary findings' might be more appropriate until the evaluation is disclosed, though this is a presentation issue relative to the major concerns.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This is an abstract-only review; the full text was not available. The abstract makes strong performance claims without any methodological detail, so I cannot assess soundness. The training-data contamination concern is real and cannot be evaluated from the abstract. I recommend that the editor either obtain the full text or require the authors to submit a version with a detailed evaluation section. I would not accept or reject based on the abstract alone; the appropriate status is 'uncertain' pending full disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the abstract is all there is, and it's too thin to judge. The central claim—that role-based LLM agents beat benchmarks in stock picking—is plausible but not verifiable from what's shown. The one thing you should know is that the validity of that claim depends entirely on whether the backtest is truly out-of-sample for the LLMs. The abstract doesn't say, so the result is unverifiable.\n\nWhat's actually new: a multi-agent team of LLM roles (researcher, risk officer, etc.) selecting stocks. That's a reasonable extension of known multi-agent debate and memory methods. The paper's potential contribution is in the evaluation design, not the architecture. The abstract also mentions 'limitations' of the framework, which is a good sign of honest intent.\n\nWhere the soft spots are: they're everywhere, but in proportion to the information available. There are no data, no error bars, no benchmark names, no evaluation window, no LLM version or knowledge cutoff. The stress-test concern about training-data contamination is on point. For an LLM, the default hypothesis is that it has seen historical stock prices and news up to its cutoff. If the test period overlaps that cutoff, the agents can recall what happened, and 'outperformance' is just memorization. The abstract doesn't rule this out, and the problem is load-bearing—it zeroes out the whole result.\n\nAlso missing: transaction costs, risk adjustment, and a comparison to simple models like a momentum or value factor. Without those, beating a broad index is weak evidence. The phrase 'under varying risk tolerance' suggests they construct portfolios, but we don't know if the risk model is meaningful or just cosmetic.\n\nIs it a serious thinker? The abstract reads like a standard system paper, no internal contradictions. But there's too little substance to tell if the reasoning holds. I'd want to see the full text before giving credit.\n\nVerdict: this is a desk reject as submitted. The abstract gives no reason to believe the performance claim is real. If the full paper provides a proper out-of-sample evaluation with the cutoffs disclosed and compares against standard baselines, it becomes a viable application paper worth a referee. But as it stands, you'd be refereeing a marketing summary.","headline":"You can't evaluate this paper from the abstract, and the contamination worry is the one that matters.","tokens_in":1152,"tokens_out":3421,"would_cite":false,"duration_ms":39328,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Role-based LLM agent teams can pick stocks that beat established benchmarks at varying risk tolerance levels.","keywords":["multi-agent systems","large language models","equity portfolio construction","stock selection","benchmark comparison","risk tolerance","LLM agents","portfolio management"],"falsifier":"Run the identical multi-agent system on market data from a period after the LLM's training cutoff, using paper trading, and compare its portfolio returns against the same benchmarks; the central claim would be refuted if the outperformance disappears or turns negative out of sample.","tokens_in":512,"feed_emoji":"📈","tokens_out":3519,"duration_ms":40505,"temperature":0.7,"pith_summary":"The paper claims that a multi-agent system built from large language models can handle equity research and stock selection, constructing portfolios that outperform established benchmarks across different risk tolerance levels. The goal is to show that role-based LLM agents, each handling a specialized part of the investment process, are practically useful for portfolio construction. If the claim holds, asset managers and individual investors could use such agent teams as a systematic stock-picking tool.","feed_headline":"LLM agent teams beat benchmarks on stock picks","feed_subtitle":"Study shows role-based multi-agent AI can construct equity portfolios at varying risk tolerance levels.","key_machinery":"The central object is the role-based multi-agent system: a set of large language model agents, each with a specialized role in equity research and portfolio management, that collaborate to produce stock selections. The mechanism is the division of labor and collaboration among agents, where each agent contributes analysis or recommendations that are then combined into portfolio decisions. The system is evaluated against established benchmarks at several levels of risk tolerance.","core_discovery":"The central claim is that a team of specialized, role-based LLM agents can select stocks and build equity portfolios that outperform established benchmarks under varying levels of risk tolerance. The study assigns distinct roles to different agents, has them collaborate on equity analysis, and aggregates their outputs into portfolio decisions. The paper reports that this multi-agent system performed favorably against the benchmarks and discusses the practical advantages and implementation challenges of using multi-agent frameworks in equity analysis.","pith_inferences":["The reported edge may shrink once transaction costs, slippage, and market impact are included, since the abstract does not say whether the backtest accounts for them.","The specific LLM and prompt design likely drive much of the result; swapping models or roles would be a natural test of whether the multi-agent structure itself is the source of performance.","If the LLM's training data includes the backtest period, the stock selections could reflect memorized outcomes rather than genuine forecasting; testing on post-cutoff data would resolve this.","The approach could extend beyond equities to other asset classes, but fixed costs and liquidity differences would require separate empirical validation."],"forward_implications":["Multi-agent LLM systems become a credible, benchmark-aware method for constructing equity portfolios automatically.","The role-based division of labor among agents can be reused as a template for other financial analysis workflows.","Risk tolerance can be dialed into the same agent team without redesigning the system.","The benchmark-comparison framework offers a starting point for evaluating future LLM-based investment agents."],"supporting_citations":[],"fun_headline_variants":["Role-based LLM agent teams beat stock benchmarks","Multi-agent AI outperforms on equity picks","LLM agent teams construct winning portfolios","Specialized agents beat market in stock selection","AI agents collaborate to beat equity benchmarks"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The evaluation setup is a fair and unbiased test: the historical data, benchmark choices, and agent design were not selected to flatter the multi-agent system, and no information about future returns leaked into the agents' selections.","fun_headline_variants_meta":{"raw":{"variants":["Role-based LLM agent teams beat stock benchmarks","Multi-agent AI outperforms on equity picks","LLM agent teams construct winning portfolios","Specialized agents beat market in stock selection","AI agents collaborate to beat equity benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000128,"raw_usage":{"total_tokens":873,"prompt_tokens":584,"completion_tokens":289,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":328,"completion_tokens_details":{"reasoning_tokens":224}},"tokens_in":328,"tokens_out":289,"duration_ms":3387,"temperature":1.0,"reasoning_tokens":224,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:04:28.082473+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical multi-agent system on market data from a period after the LLM's training cutoff, using paper trading, and compare its portfolio returns against the same benchmarks; the central claim would be refuted if the outperformance disappears or turns negative out of sample.","supporting_citations":[],"review_version":1}