{"id":"5d8c77d3-5741-4918-b4a8-563356f8a62c","arxiv_id":"2606.03918","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Hedge-Bench is a benchmark of 102 professional hedge fund tasks where frontier AI agents score below 16%.","lead":"This paper introduces Hedge-Bench, a collection of 102 real financial analysis tasks drawn from hedge fund analyst work, complete with expert reasoning traces for objective grading. A smart generalist might read it to understand the current limits of AI on complex, open-ended professional reasoning rather than routine calculations.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest assumption matches the paper's central methodological claim; the full-text description supplies no additional technical detail that would falsify deterministic grading or task representativeness. Therefore the reported performance numbers can stand on the evidence given.","tokens_in":1610,"tokens_out":188,"duration_ms":15196,"concrete_test":"Execute the published evaluation harness on the released dataset and confirm that the frontier-model scores remain below 16% under the deterministic trace-matching procedure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper constructs 102 tasks from professional hedge-fund traces to enable deterministic grading against verified expert steps, directly targeting the noise and circularity problems of model-judged open-ended benchmarks. No internal inconsistency, hidden assumption in the grading procedure, or statistical weakness in the reported <16% scores is apparent from the construction described.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Hedge-Bench 1.0, a benchmark of 102 tasks drawn from actual professional hedge-fund analyst workflows, each paired with verified expert reasoning traces to enable deterministic grading. It argues that existing benchmarks suffer from noise and circularity due to model-based judging of open-ended outputs, and reports that frontier models and agents score below 16% on this new benchmark. The dataset and evaluation harness are released publicly.","tokens_in":1661,"tokens_out":533,"duration_ms":13271,"significance":"If the construction and grading procedure hold, the benchmark provides a concrete, falsifiable measure of the gap between current AI agents and expert-level financial reasoning on realistic, open-ended tasks. The grounding in professional traces and deterministic evaluation against expert steps is a methodological strength that could influence how future financial-reasoning benchmarks are designed.","major_comments":[{"comment":"§3 (Benchmark Construction): The claim that the 102 tasks are representative of expert Analyst work rests on selection from professional hedge-fund traces, but the manuscript provides no explicit inclusion/exclusion criteria, sampling procedure, or quantification of task diversity (e.g., by topic, time horizon, or information-source type). This directly affects the generalizability of the <16% result.","section":"§3"},{"comment":"§4 (Evaluation Procedure): The deterministic grading approach is presented as eliminating noise and circularity, yet no inter-rater reliability statistics for the expert traces, no rubric examples, and no analysis of grading edge cases or disagreement rates are reported. Without these, the central claim that scores are reliably below 16% cannot be fully evaluated.","section":"§4"},{"comment":"Results section: The headline performance numbers (<16%) are given without statistical error bars, confidence intervals, or breakdown by model/agent type and task category. This makes it impossible to assess whether the reported ceiling is robust or sensitive to small changes in the 102-task set.","section":"Results"}],"minor_comments":[{"comment":"The abstract and introduction repeatedly use “deterministic grading” without a concise definition or pointer to the exact grading algorithm in the evaluation harness.","section":"Abstract"},{"comment":"Table 1 (model scores) would benefit from an additional column showing the number of tasks each model/agent was evaluated on, to clarify coverage.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive comments. We address each major point below, indicating revisions where the manuscript can be strengthened by adding requested details.","responses":[{"response":"We agree that explicit documentation of the selection process is needed for assessing generalizability. The 102 tasks were drawn from a larger collection of verified professional hedge-fund analyst workflows, with selection focused on open-ended tasks requiring integration of multiple heterogeneous sources. In the revision we will add a subsection to §3 that states the inclusion/exclusion criteria, describes the sampling procedure from the trace pool, and reports quantitative diversity statistics by topic, time horizon, and information-source type.","revision_made":"yes","referee_comment":"[§3] §3 (Benchmark Construction): The claim that the 102 tasks are representative of expert Analyst work rests on selection from professional hedge-fund traces, but the manuscript provides no explicit inclusion/exclusion criteria, sampling procedure, or quantification of task diversity (e.g., by topic, time horizon, or information-source type). This directly affects the generalizability of the <16% result."},{"response":"The grading itself is deterministic because it verifies completion of specific steps listed in the expert traces rather than relying on open-ended judgment. Nevertheless, we accept that reliability information on the traces themselves would strengthen the claim. In the revised §4 we will report inter-rater reliability statistics for the expert traces, provide sample rubrics, and include an analysis of grading edge cases and observed disagreement rates.","revision_made":"yes","referee_comment":"[§4] §4 (Evaluation Procedure): The deterministic grading approach is presented as eliminating noise and circularity, yet no inter-rater reliability statistics for the expert traces, no rubric examples, and no analysis of grading edge cases or disagreement rates are reported. Without these, the central claim that scores are reliably below 16% cannot be fully evaluated."},{"response":"We will expand the Results section to include statistical error bars and confidence intervals around the aggregate scores. We will also add performance breakdowns by model/agent type and by task category so readers can evaluate robustness and sensitivity to the task set.","revision_made":"yes","referee_comment":"Results section: The headline performance numbers (<16%) are given without statistical error bars, confidence intervals, or breakdown by model/agent type and task category. This makes it impossible to assess whether the reported ceiling is robust or sensitive to small changes in the 102-task set."}],"tokens_in":1323,"tokens_out":543,"duration_ms":21747,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core of this paper is a benchmark built from 102 actual on-the-job tasks taken from professional hedge fund analysts, each paired with their explicit reasoning traces so grading can be done deterministically against those traces. That setup is the main thing worth noticing. It tries to fix the noise and circularity that come from having models judge other models on open-ended financial questions.\n\nThe approach is straightforward and the published dataset plus harness is a practical step. Showing that current frontier models and agents stay under 16% on these tasks gives a concrete signal that the problems are harder than the mechanical retrieval and calculation work that existing tests cover.\n\nThe abstract gives no information on how the tasks were picked from the larger pool of hedge fund work, whether multiple analysts reviewed the traces for consistency, or how the scores were aggregated with any measure of variance. Those omissions make it difficult to judge whether the 102 tasks are representative or whether the grading procedure introduces its own artifacts. The stress-test note says the construction looks internally consistent on the surface, and nothing in the abstract contradicts that, but the lack of those details is still the main limitation.\n\nThis is useful for groups working on agent evaluation in finance or other domains that need open-ended professional reasoning. The published resources lower the barrier to trying it out. It is worth sending to peer review so the task construction and grading protocol can be examined in full.","headline":"Hedge-Bench gives a new benchmark of 102 real hedge-fund tasks graded against expert traces, with frontier models below 16%, but the abstract supplies almost no information on task selection or grading reliability.","tokens_in":2103,"tokens_out":369,"would_cite":false,"duration_ms":10237,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Frontier AI agents score below 16 percent on a benchmark of 102 real hedge fund analyst tasks graded against expert reasoning traces.","keywords":["AI agents","financial reasoning","benchmark","hedge fund tasks","reasoning traces","deterministic evaluation","open-ended analysis"],"falsifier":"An independent replication in which the same agents are run through the published evaluation harness and achieve scores above 30 percent would contradict the reported performance levels.","tokens_in":2510,"feed_emoji":"","tokens_out":595,"duration_ms":21050,"temperature":0.7,"pith_summary":"The paper presents Hedge-Bench as a set of 102 tasks taken directly from professional hedge fund work, each paired with explicit analyst reasoning steps. This design supports grading by matching outputs to verified traces instead of relying on model judgments. Current frontier models and agents reach less than 16 percent success under this standard. The benchmark targets the open-ended synthesis and judgment that separate mechanical data tasks from expert financial analysis. Accurate measurement of this gap matters for understanding where AI still falls short in high-stakes decision domains.","feed_headline":"AI agents score below 16% on real hedge fund analyst tasks","feed_subtitle":"Benchmark of 102 tasks with expert reasoning traces enables grading without model judgments and shows current limits.","key_machinery":"Hedge-Bench 1.0, a benchmark of 102 tasks drawn from hedge fund analyst work and supplied with explicit expert reasoning traces that permit deterministic grading.","core_discovery":"Hedge-Bench 1.0 consists of 102 actual on-the-job tasks grounded in the explicit reasoning traces of professional hedge fund analysts working with relevant information sources. This approach enables deterministic grading against verified expert steps. Frontier models and agents score below 16 percent on the benchmark.","pith_inferences":["Success on this benchmark could serve as a template for creating similar traceable-task sets in legal or medical domains.","The performance gap suggests that simply increasing model scale may not close the difference without explicit mechanisms for step verification.","Agents trained or prompted to output intermediate traces matching the expert format might show measurable gains on the same tasks."],"forward_implications":["Current agents cannot yet replicate the traceable reasoning steps used in professional financial analysis.","Benchmarks that rely on model-based judging introduce circularity that deterministic trace matching avoids.","The released dataset and harness allow repeated, consistent testing of future models on the same tasks.","Progress on these tasks would require agents to handle synthesis and judgment rather than only retrieval and calculation."],"fun_headline_variants":["AI agents score below 16% on 102 hedge fund tasks","Hedge-Bench finds AI below 16% on 102 analyst tasks","Expert hedge fund tasks benchmark AI under 16% score","102 real tasks ground AI evaluation at below 16%"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 102 tasks drawn from hedge fund work and their accompanying expert traces are representative of open-ended analyst questions and support fully deterministic grading.","fun_headline_variants_meta":{"raw":{"variants":["AI agents score below 16% on 102 hedge fund tasks","Hedge-Bench finds AI below 16% on 102 analyst tasks","Expert hedge fund tasks benchmark AI under 16% score","102 real tasks ground AI evaluation at below 16%"]},"model":"grok-4.3","cost_usd":0.005329,"raw_usage":{"total_tokens":2521,"prompt_tokens":564,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":53287000,"prompt_tokens_details":{"text_tokens":564,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1887,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":564,"tokens_out":70,"duration_ms":15912,"temperature":1.0,"reasoning_tokens":1887,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T09:59:10.145222+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An independent replication in which the same agents are run through the published evaluation harness and achieve scores above 30 percent would contradict the reported performance levels.","supporting_citations":[],"review_version":1}