Pith. sign in

REVIEW 3 major objections 8 minor 48 references

Professional analysts disagree too much for single-reference grading of valuation models, and current AI agents still trail juniors on judgment.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-07-31 16:10 UTC pith:3VOTKBFQ

load-bearing objection Strong measurement paper: single-golden valuation grading is empirically broken, and the mechanical–judgment gap is real; the human bar is partly confounded by vendor provenance and time budget. the 3 major comments →

arxiv 2607.24889 v1 pith:3VOTKBFQ submitted 2026-07-27 cs.LG cs.AIcs.CE

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

classification cs.LG cs.AIcs.CE
keywords BenchmarkLLM AgentsFinancial ModelingLLM EvaluationValuation JudgmentObserved-Practice EnvelopeConstruct Validity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Financial valuation models mix hard accounting facts with soft judgment calls such as forecasts, discount rates, and target prices. Existing agent benchmarks still grade those soft parts against one expert’s numbers, which treats professional disagreement as error. The authors show that when one analyst’s workbook is scored against another’s for the same company, the median score is only 0.33 and almost none agree on implied price within 10%. They build GAUGE to grade agent-built models against an observed-practice envelope drawn from many analyst workbooks, plus mechanical checks and hard validity gates. On that score, seniors, juniors, and students separate cleanly, while the best of 24 agents beats students but loses to every senior and most juniors, and the whole fleet is far stronger at building the spreadsheet than at valuation judgment.

Core claim

Point-tolerance grading against a single golden analyst answer is the wrong construct for end-to-end valuation: same-company professionals already disagree so much that a single-reference rule mostly measures that disagreement. Graded instead against observed analyst practice, current agents can assemble structurally sound models but still fall short of junior-level valuation judgment.

What carries the argument

GAUGE’s three-layer observed-practice envelope (method-level sensitivity bands, industry distributions, and same-company dispersion) plus 56 facets, eight validity gates, and deterministic structural checks, aggregated into a failure-aware score φ₀ that zeros non-completions.

Load-bearing premise

The “defensible” bands come from the same vendor analyst corpus used to prove professional disagreement, so a high envelope score may only mean “looks like this vendor’s practice,” not independently correct judgment.

What would settle it

Re-run the same agent and human panels with envelopes and multi-analyst bands built from an independent broker network or market outside this vendor corpus; if the mechanical–judgment gap and agent-vs-junior ordering reverse or collapse, the central claim about judgment deficit under observed practice does not hold.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X LinkedIn Reddit HN

If this is right

  • Valuation and other non-unique professional artifacts should not be graded by proximity to one author’s point estimates.
  • Agent progress on finance should be reported separately for mechanical construction and judgment, not as one end-to-end accuracy number.
  • Hard structural gates and failure-aware scoring change leaderboards relative to additive rubric averages.
  • Known-groups human baselines (senior > junior > student) become a required validity check for open-ended occupational benchmarks.
  • Released methodology, gated data, a frozen 48-task core, and a withheld refresh pool enable longitudinal, contamination-aware leaderboards.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Any domain where experts legitimately disagree—legal opinions, medical plans, strategy memos—may need the same envelope-plus-gates pattern rather than single-reference rubrics.
  • Closing the 26-point mechanical–judgment gap likely needs training signals tied to practice distributions, not only more spreadsheet tool use.
  • If agents keep ranking companies by risk about as well as analysts agree with each other while still snapping WACC to textbook grids, the remaining deficit is assumption craft, not cross-sectional ordering.
  • Vendor-network human baselines and envelopes invite external replication before the score is treated as a universal professional bar.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 8 minor

Summary. The paper first audits the single-golden-answer assumption behind existing financial-agent benchmarks: grading one analyst-built workbook against a same-company peer under standard point tolerances yields a median score of 0.33 across 108 directed pairs (65 companies), with 92.6% below 0.70 and zero same-vintage implied-price agreement within 10%. It then introduces GAUGE, a benchmark scoring agent-built valuation models against a three-layer observed-practice envelope (per-workbook sensitivity ranges, GICS p10–p90 distributions, same-company cross-analyst dispersion) rather than a point reference, combined with 56 facets (29 deterministic, 23 LLM-judged, 4 direct-rule), eight validity gates with score ceilings, and a failure-aware score φ₀ that zeros non-completions. Validation includes a 55-participant known-groups study (seniors 88.3, juniors 66.0, students 43.2), company-grouped cross-fitting, a perturbation selectivity control, and judge vote-sampling and cross-family judge audits. Across 24 agents on a 48-task core (1,011 scored cells), the best agent scores φ₀ = 53.4, above the student mean but below every senior; all agents show a mechanical–judgment pass-rate gap (fleet median 26 points). The authors release methodology, harness, run ledgers, raw judge votes, the peer-audit JSONL, gated data splits, and a withheld refresh pool.

Significance. If the results hold, this is a useful contribution on two axes. The peer-workbook audit is a direct, well-quantified, and reusable demonstration that point-tolerance grading against one analyst reference is construct-invalid for valuation — and the underlying 632 pair/tolerance records are released as JSONL, so the central negative result is independently recomputable. The benchmark itself is engineered with unusual measurement discipline: deterministic gate ablation (11/276 orderings flip), judge vote-count ablation (τ=0.944 at k=5), a cross-family judge replication showing a family-neutral uniform shift with preserved ordering (τ=0.857), a perturbation control bounding band permissiveness vs. selectivity, bootstrap rank-stability for the failure-aware ranking, verbatim failure trajectories, and full instrument provenance including disclosed detector defects and fixes. The mechanical/judgment decomposition and the WACC grid-vs-rank-ordering diagnostic (agents match analysts' cross-company risk ordering, ρ≈0.36–0.42 vs. 0.38) localize the agent deficit in a falsifiable way. The main limitation is single-vendor provenance of both the calibration corpus and the human panel, which the

major comments (3)
  1. [§6.3, Table 5, Fig. 5] The known-groups panel and the envelope calibration share one vendor network (§6.3: participants 'drawn from the commercial vendor network that produced the corpus'; §5.2: E-industry/E-company bands calibrated on that corpus). The paper's own numbers sharpen the concern: seniors' mechanical subscore (92.5, Table 5) essentially equals the best agent's (93%), so the entire senior–agent gap lives in judged/envelope facets — precisely the component calibrated to the vendor's own practice distribution. The App. N.2 perturbation control cannot detect a house-style offset shared by calibration corpus and panel. Request: (i) an explicit decomposition of the human–agent gap by grader family (deterministic+gates vs. judged vs. envelope), which is computable from released artifacts; (ii) framing in §6.3/abstract that φ₀ measures conformity to this observed-practice distribution, with external repli
  2. [§5.2, §6.5, App. N.1–N.2] Two load-bearing envelope numbers are weaker than the headline suggests. (i) Strict held-out E-method price coverage is 53.8% (CI 38.5–68.6%) — nearly half of real peer prices fall outside the full-credit band, so E-method functions mostly as a partial-credit device. (ii) The p90 implied-price tail rests on 17 undirected pairs, and the multiplicative near band (δ≈1.0) is asymmetric: it rejects 38.5% of upward but 0% of downward ±2× counterfeits (Table 20). The authors read the near band as partial credit only, which is defensible, but then the 91.2% near-coverage figure does little evidentiary work and should not be presented alongside strict coverage as validation. Please state the asymmetry and the n=17 basis in the main text (§5.2 or §6.5), not only in App. N.
  3. [§5.4, App. K.1, App. N.4–N.5] App. N.4 shows the leaderboard leans most on the judged facets (τ drops to 0.683 without J; top-1 preservation 0.305), yet the only human-correctness evidence for the judge is the second-hand 460-case note (App. K.1) with no annotator qualifications, sampling frame, consensus construction, blinding, or item-level labels — and its weakest slice is valuation (82.9% exact, κ=0.75), the facet family carrying the paper's central claim. The cross-family replication (App. N.5) addresses bias but both judges could share systematic errors vs. experts. The authors' downgrade to 'descriptive' is appropriate, but for a venue publication the judgment-facet correctness claim needs either documentation of this audit or a small documented expert re-audit of valuation facets (the five lowest-agreement facets flagged in N.5 are a natural target).
minor comments (8)
  1. [§5.2] The §5.2 prose defining the layers is garbled: 'available for 54uses same-GICS p10 to p90 distributions ... covering 8665-company multi-coverage corpus' — numbers appear fused (54%? 866 books? 65 companies). Please rewrite the paragraph around Eq. (1).
  2. [§6.3, App. D.4] Human–agent comparisons use φ₀ for both populations, but humans were untimed (median 200–273 min, App. D.4) while agent timeouts score zero. For the headline claim this is largely benign (the best agents complete 100%, so their φ₀ = φ), but for mid-fleet agents it is not. A sentence noting that the human–agent ordering also holds on completed-attempt φ (Table 18's Φ column vs. Table 5) would close this.
  3. [§5.3, Table 18] Notation collision: Φ in Eq. (5) is the gated final score, but Table 18's Φ column is a completed-cells conditional quantity that differs from Table 4's φ₀ (e.g., Grok 46.5 vs. 40.7; Doubao 2.1-pro 40.6 vs. 10.2). Rename one, and state in the Table 18 caption that rows are not rank-ordered by it.
  4. [§4, Fig. 2, Fig. 9] The 'flat score' of §4/Fig. 2 is never formally defined (fraction of stated criteria passed? equally weighted?). One line defining it, and noting unstated criteria are N/A, would help. Also Fig. 2 is near-duplicate of Fig. 9a; consider merging.
  5. [§6.3] Experience groups are 'vendor-classified' (§6.3): the known-groups result partly validates the vendor's own seniority labels. A caveat sentence is warranted alongside the existing credential-verification disclaimer.
  6. [App. J, §7] 22 of 25 industry overlays were generated by claude-opus-4-8, the same model as the ladder judge (App. J). The disclosure is commendable; please also note it in §5.3 or Limitations, since overlay activation affects which judged facets enter A(x).
  7. [§6.1, App. N.3] One generation per agent–task cell (§6.1) with generation variance measured only via a sentinel set: given the exact top-five set is preserved in only 22.4% of bootstrap replicates (App. N.3) and the judge flip rate is 2.2%, a sentence quantifying expected rank noise near ties would calibrate reader interpretation of the leaderboard.
  8. [Table 15, App. T] G5 (look-ahead) is described in App. T as catching time-inverted model mechanics rather than information leakage, since packs are as-of-clamped by construction. Consider renaming or re-scoping the gate description in Table 15 to match what it actually detects.

Circularity Check

3 steps flagged

Main empirical claims are not circular; partial circularity sits in the validity apparatus (same-vendor envelope + known-groups, and corpus context improving envelope-scored facets).

specific steps
  1. fitted input called prediction [§5.2 Three-Layer Defensibility Envelope; App. N.1 cross-fit]
    "eE_f(x) the same band widened by the p90 cross-analyst disagreement for that assumption (Section 4). ... E-company ... sets the near-band widths in eE_f and tests the two proxy layers. ... all folds remain within the same 65-company source corpus. These results provide internal validation of sampled practice, not external replication or proof that every in-band choice is correct."

    Near-band widths and the peer-audit disagreement tails that justify them are estimated from the same multi-covered corpus used to motivate and calibrate the envelope. Held-out coverage checks are company-grouped but still resample that single source distribution, so ‘envelope admits observed analyst practice’ is partly guaranteed by construction of the bands from that practice rather than an external referent.

  2. other [§6.3 Human Baseline: Known-Groups Validity; §7 Limitations]
    "Participants were drawn from the commercial vendor network that produced the corpus and grouped by the vendor’s seniority classification... The corpus and human study also come from one vendor network, with unverified credentials and overrepresentation of US listings."

    Known-groups ordering is offered as primary construct-validity evidence that φ tracks valuation experience. Because judgment facets are scored against an observed-practice envelope built from that same vendor’s workbooks, senior outperformance partly measures match to the house distribution that defines the bands—not an independently anchored standard of defensibility. This is a shared-provenance validity loop, not a formal self-definition of a derived equation.

  3. fitted input called prediction [§6.6 Training Signal; Appendix O]
    "E-industry tables raise the valuation-judgment subscore by +4.0 φ (95% CI [+1.3, +6.7]) in a leakage-controlled 200-workbook split... In both studies, the judgment gains concentrate on envelope-scored facets—the quantities the corpus distributions directly inform."

    Supplying corpus-derived industry envelope tables as context and then measuring gains on facets graded by those same envelope distributions is a statistically forced, scorer-aligned lift. The paper correctly notes concentration on envelope-scored facets; that localization is exactly the by-construction channel.

full rationale

GAUGE is a benchmark paper, not a first-principles derivation. The load-bearing peer-audit result—one analyst workbook scored against another under fixed tolerances—is an independent empirical measurement and does not reduce to its inputs by construction. Agent leaderboards use a frozen facet/gate stack on held-out tasks and are likewise non-circular. What carries a mild circularity burden is the construct-validity loop around “defensible judgment”: E-method/E-industry bands and p90 near-widths are estimated from the same multi-covered vendor corpus that motivates abandoning single-golden grading, and company-grouped cross-fits remain inside that 65-company source sample (explicitly internal validation, not external replication). The 55-person known-groups panel is drawn from the same commercial vendor network that produced the corpus, so high senior scores partly test conformity to a house practice distribution the instrument encodes. Separately, the training-signal study shows E-industry tables raise the valuation-judgment subscore on facets the corpus distributions directly grade—an expected, partly by-construction lift the paper itself localizes. These issues weaken how far φ can be read as transferable professional judgment, but they do not make the peer-disagreement statistic or the mechanical–judgment agent gap tautological. Score 3: real but partial circularity in the referent/validity chain; central comparative claims retain independent content.

Axiom & Free-Parameter Ledger

6 free parameters · 6 axioms · 3 invented entities

GAUGE’s interpretive claim—that higher scores mean closer-to-professional, defensible valuation practice—rests on domain conventions of three-statement modeling, on treating vendor analyst workbooks and seniority labels as practice referents, and on design choices (φ map, gate ceilings, p90 near bands, frozen LLM judge) that are calibrated or stipulated rather than derived from first principles.

free parameters (6)
  • Base point-tolerance bands (rev ±5%, WACC ±50bp, price ±10%) and 1×–4× sweep = rev±5%; WACC±50bp; price±10%
    Define the single-golden peer-audit pass/fail surface; median 0.33 and 92.6%<0.70 are relative to these choices.
  • Near-band width = p90 cross-analyst disagreement; industry bands at p10–p90 (operating p90) = p90 near rule; E-industry p90
    Sets partial-credit region eE_f and E-industry coverage/selectivity operating point (App. N perturbation frontier).
  • Facet score map φ(0)=0, φ(1)=60, φ(2)=100 = 0/60/100
    Nonlinear Fail→Pass jump dominates aggregation and known-groups/agent numeric scores.
  • Eight gate ceilings κ_g (e.g. G1→40, G5→35) = G1:40, G2:45, G3:55, G4:50, G5:35, G6:50, G7:60, G8:60
    Hard caps on Φ; removing caps flips 11/276 model-pair orderings.
  • Judge vote count k=5 majority reduction = k=5
    Controls sampling variance for 23 judged facets; stability reported, not correctness.
  • Deterministic detector thresholds (BS 0.1%, cash 0.5%, formula density 90/95%, etc.) = as in Tables 10–15 / gates.json
    Binary mechanical facets and gate triggers depend on these cutoffs validated on schema-conformant fixtures.
axioms (6)
  • domain assumption Standard three-statement accounting identities and DCF/WACC (or bank DDM) mechanics are the correct structural targets for “model construction.”
    Underpins deterministic facets and gates G1–G8; industry overlays reinterpret for banks etc. (App. H–I).
  • domain assumption Independently produced vendor analyst workbooks constitute a valid sample of “observed professional practice” for envelope bands.
    Core referent for E-method/E-industry; vendor labels explicitly do not verify credentials (§3).
  • domain assumption Vendor seniority classes (senior/junior/student) are ordered proxies for valuation experience in the known-groups study.
    Ordering validity claim in §6.3; paper notes credentials not independently verified.
  • ad hoc to paper Values inside empirical analyst envelopes are more defensible than values outside, without claiming unique correctness.
    Stated measurement philosophy of GAUGE (§5.2, §7); internal cross-fit is not external replication.
  • ad hoc to paper A frozen LLM ladder judge with k-vote majority is an adequate operational grader for 23 qualitative facets.
    §5.3–5.4; stability and partial human-agreement stats; documentation gaps acknowledged.
  • ad hoc to paper Non-completions and invalid artifacts should score zero on a fixed task denominator (φ₀).
    Failure-aware leaderboard convention §6; changes ranking vs completed-only φ.
invented entities (3)
  • Three-layer defensibility envelope (E-method, E-industry, E-company calibration) independent evidence
    purpose: Replace single point estimates for judgment-bearing quantities with observed-practice bands and partial credit.
    Central instrument of the benchmark; validated mainly by internal coverage/selectivity and peer audit, not external market outcomes.
  • 56-facet taxonomy with C1/C2/C3 mechanical–judgment split and 8 validity gates independent evidence
    purpose: Decompose workbook quality and prevent structurally unusable models from high scores.
    Rubric-first artifact with calibration history (App. G–J); gates are design choices tied to deterministic detectors.
  • Failure-aware score φ₀ / gated aggregate Φ no independent evidence
    purpose: Single leaderboard number combining facets, N/A dropping, and gate ceilings.
    Defined by Eqs. (3)–(5); not an externally standardized psychometric scale.

reviewed 2026-07-31 · how reviews work

0 comments
Cite this review

Pith. "Pith review of GAUGE: Grading Agent-Built Financial Models Without a Golden Answer." pith.science (2026). https://pith.science/paper/3VOTKBFQ

@misc{pith2026260724889,
  author       = {Pith},
  title        = {Pith review of: GAUGE: Grading Agent-Built Financial Models Without a Golden Answer},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3VOTKBFQ}},
  note         = {Machine review of arXiv:2607.24889}
}
Share X LinkedIn Reddit HN
read the original abstract

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rates, and target prices often admit multiple reasonable answers. Existing benchmarks nevertheless tend to grade such outputs against a single expert reference. Using independently built analyst models for the same companies, we find that across 108 directed pairs covering 65 companies, the median single-reference score is 0.33, 92.6% score below 0.70, and no same-vintage pair agrees on implied price within 10%. Point-tolerance grading can therefore penalize disagreement already present among professionals. We introduce GAUGE, a benchmark for evaluating agent-built valuation models against observed analyst practice rather than a single point answer. GAUGE uses 1,001 vendor-classified analyst workbooks and a 196-task evaluation set, with a three-layer observed-practice envelope, 56 auditable facets, eight validity gates, and deterministic structural checks. We validate the benchmark with a 55-participant known-groups study, company-grouped cross-fitting, and judge-stability audits. On the failure-aware score $\phi_0$, senior analysts average 88.3, juniors 66.0, and finance students 43.2. Across 24 agents and 1,011 scored generations, the best agent scores 53.4, above the student mean but below every senior and most juniors. It passes 93% of mechanical facets and 78% of judgment facets, with a fleet-median gap of 26 points. Current agents are substantially stronger at model construction than valuation judgment. We release the methodology, a gated de-identified data tier, a controlled training split, a versioned 48-task evaluation core, and a withheld refresh pool.

Figures

Figures reproduced from arXiv: 2607.24889 by Beidi Luan, Cheng Hua, Daxin Jiang, Haibing Guan, Hui Cai, Jiacheng Lu, Jing Li, Lingjing Teng, Rui Sun, Sinuo Wang, Tao Song, Wentao Zhao, Yijia He, Zhengze Wu, Zuo Bai.

Figure 1
Figure 1. Figure 1: GAUGE end-to-end. The corpus contains 1,001 analyst-built workbooks. Agents receive an Excel-in / Model-out task and are graded with a three-layer observed-practice envelope, 56 facets, and validity gates. From left to right, GAUGE turns analyst artifacts into hidden tasks, defensibility-aware scores, and fleet-level capability diagnostics. The bottom strip gives each stage’s role; the 24-agent mechanical–… view at source ↗
Figure 2
Figure 2. Figure 2: Single-golden tolerance sweep. Among 108 directed same￾company pairs, the median score is 0.33; 92.6% fall below 0.70 at 1× tol￾erances, and one third remain below 0.70 at 4×. The two-panel audit and criterion rates are in Appendix A. 3 The Finance Artifacts Corpus The Finance Artifacts corpus contains 1,001 vendor-classified analyst-built valuation workbooks: 922 tickers, 25 GICS groups, and 583 tab archi… view at source ↗
Figure 3
Figure 3. Figure 3: Observed same-company disagreement. Median/p90 abso￾lute differences are 147/374 bp for WACC and 25/107% for implied price. Tail estimates use 61 and 17 undirected pairs, respectively. them at 1×–4×. Median scores are 0.33/0.50/0.67/0.80. At 1×, risk￾free rate passes 66%, WACC 13%, and implied price 24%; none of 14 same-vintage price pairs passes. Restricting to same-vintage pairs leaves the median at 0.33… view at source ↗
Figure 4
Figure 4. Figure 4: Mechanical and judgment pass rates. The plot shows 12 provider flagships (of 24 agents), in leaderboard order, with mechanical pass rate, judgment pass rate, and full-stack 𝜙0 (48-task denominator; capability failures scored 0). The fleet-wide median gap is 26 points. Qwen3.7-Max has mid-pack pass rates but 17 unfinished tasks, giving 𝜙0 = 24.9. a hidden stratified wave is drawn from the unused bank and wi… view at source ↗
Figure 5
Figure 5. Figure 5: Known-groups validity of the human baseline. Each dot is one participant’s mean 𝜙0 over three assigned tasks, with non-completions scored zero (𝑛=55: 25 students, 18 juniors, 12 seniors; groups are vendor￾classified). Horizontal bars mark group means (43.2 / 66.0 / 88.3). The dashed line is the best agent, Claude Fable 5 (𝜙0 = 53.4); 33 of 55 participants outscore it, including all 12 seniors and 15 of 18 … view at source ↗
Figure 6
Figure 6. Figure 6: Score decomposition per agent. The 48-task ceiling is parti￾tioned into retained 𝜙0, gate deduction, failed active facets, and capability failure. Gate deductions are at most 4.2 points; failed facets dominate the loss for 21/24 agents. 0 10 20 30 40 50 Score on the 48-task core (failures scored 0) Claude Fable 5 Claude Opus 4.8 GPT-5.6-sol Kimi k2.7-code GLM 5.2 Grok 4.5 GPT-5.6-terra Claude Sonnet 5 Deep… view at source ↗
Figure 7
Figure 7. Figure 7: Full stack and additive credit on identical facet outcomes. Removing gates flips 11/276 model-pair orderings; the additive-minus￾GAUGE difference has 𝑟 = 0.73 with gate rate (in Appendix N). the p90 near band covers 91.2% (82.6–97.2%). At p90, strict held-out E-industry value coverage is 75.4% for beta, 79.0% for tax, 80.6% for ERP, 82.4% for WACC, 84.5% for risk-free rate, and 90.2% for terminal growth. T… view at source ↗
Figure 8
Figure 8. Figure 8: Corpus-as-context. Paired deltas with 95% intervals. E-industry tables change valuation judgment by +4.0𝜙; the mechanical interval in￾cludes zero. Exemplar cards produce a non-significant +1.0 change. ment gains concentrate on envelope-scored facets—the quantities the corpus distributions directly inform (in Appendix O). 7 Accessibility, Ethics, and Limitations Access and ethics. Rubrics, envelope statisti… view at source ↗
Figure 9
Figure 9. Figure 9: Full peer-workbook audit. (a) Cumulative single-golden scores for 108 directed pairs as tolerance bands widen from 1× to 4×. (b) Criterion pass rates at 1× and 3×. The senior analyst will flex assumptions; hardcoded outputs are a silent bug. 4. The model must produce a final **implied share price** that is clearly labeled and easy to find. 5. Use real accounting conventions (GAAP-style). The Income Stateme… view at source ↗
Figure 10
Figure 10. Figure 10: Tool use across the fleet (leaderboard order; frozen w2_main final attempts). (a) Executed tool calls per completed task, split run_bash vs. write_file: a 17× spread in interaction granularity (GPT-5.6-sol 2.9 to Claude Sonnet 5 49.6) with no monotone relation to rank. (b) Share of calls whose fed-back result is an error (stub-floor for Python classes, exact for harness-synthesized classes); dashed line: … view at source ↗
Figure 11
Figure 11. Figure 11: Self-verification discipline vs. the balance-sheet gate, per agent over completed tasks. In-run recalculation (LibreOffice or a formula￾evaluation library, floor detection) associates with a lower G1 rate at the group level (dashed means), but is neither necessary (Qwen3 Coder, Grok 4.5) nor sufficient (Claude Sonnet 5): running the checks and acting on them are different capabilities. Relation to scored … view at source ↗
Figure 14
Figure 14. Figure 14: reports, per agent, the share of completed tasks triggering each validity gate — G1 balance-sheet identity, G2 cash-flow tie￾out, G3 segment reconciliation, G4 hardcoded forecast cells, G5 look-ahead leakage, G6 unresolvable source citation, G7 unresolved circular references, G8 EV→equity bridge — alongside the any-gate rate from [PITH_FULL_IMAGE:figures/full_fig_p031_14.png] view at source ↗
Figure 12
Figure 12. Figure 12: The full per-facet matrix: pass rate (%, score ≥ 1 among active cells) for every agent × every facet activated on the 48-task core. Rows are facets grouped by pillar (navy separators; right-edge tags P1 input comprehension, P2 model construction, P3 forecast & reasoning, P4 valuation & sensitivity, P5 communication & auditability; [D] deterministic, [J] judged, [R] envelope); columns are the 24 agents in … view at source ↗
Figure 13
Figure 13. Figure 13: Facet pass rate by workbook tier. Across 24 agents, me￾chanical and judgment pass rates vary little across the four vendor tiers; the gap is 24.6–25.8 points. Tier is a scale covariate, not a measure of verified analyst quality. 35 [PITH_FULL_IMAGE:figures/full_fig_p035_13.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

48 extracted references · 14 linked inside Pith

  1. [1]

    Anthropic. 2026. Claude Fable 5 & Claude Mythos 5 System Card. https://www. anthropic.com/claude-fable-5-mythos-5-system-card. Accessed 2026-07-23

  2. [2]

    Anthropic. 2026. Claude Opus 4.8 System Card. https://www.anthropic.com/ claude-opus-4-8-system-card. Accessed 2026-07-23

  3. [3]

    Anthropic. 2026. Claude Sonnet 5 System Card. https://www.anthropic.com/ claude-sonnet-5-system-card. Accessed 2026-07-23

  4. [4]

    Andrew M Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, et al. 2026. Measuring what matters: Construct validity in large language model benchmarks.Advances in Neural Information Processing Systems38 (2026)

  5. [5]

    DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Con- text Intelligence. arXiv:2606.19348 [cs.CL]

  6. [6]

    Efthimios G Demirakos, Norman C Strong, and Martin Walker. 2004. What valuation models do analysts use?Accounting horizons18, 4 (2004), 221–240

  7. [7]

    Google DeepMind. 2026. Gemini 3.1 Pro Model Card. https://deepmind.google/ models/model-cards/gemini-3-1-pro/. Accessed 2026-07-23

  8. [8]

    Google DeepMind. 2026. Gemini 3.5 Flash Model Card. https://deepmind.google/ models/model-cards/gemini-3-5-flash/. Accessed 2026-07-23

  9. [9]

    Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Steven Wu, and Alexandra Chouldechova. 2026. Validating llm-as-a-judge systems under rating indeterminacy.Advances in Neural Information Processing Systems38 (2026), 112282–112350

  10. [10]

    Valentin Hofmann, David Heineman, Ian Magnusson, Kyle Lo, Jesse Dodge, Maarten Sap, Pang Wei Koh, Chun Wang, Hannaneh Hajishirzi, and Noah A Smith. 2025. Fluid language model benchmarking.arXiv preprint arXiv:2509.11106 (2025)

  11. [11]

    Shahed Imam, Richard Barker, and Colin Clubb. 2008. The use of valuation models by UK investment analysts.European accounting review17, 3 (2008), 503–535

  12. [12]

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2024. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974(2024)

  13. [13]

    Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

    Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2024. SWE-bench: Can Language Models Resolve Real- World GitHub Issues? arXiv:2310.06770 [cs.CL] https://arxiv.org/abs/2310.06770

  14. [14]

    Kimi Team. 2025. Kimi K2: Open Agentic Intelligence. arXiv:2507.20534 [cs.LG]

  15. [15]

    Kimi Team. 2026. Kimi K3 Tech Blog: Open Frontier Intelligence. https://www. kimi.com/blog/kimi-k3. Accessed 2026-07-23

  16. [16]

    Michael Krumdick, Varshini Reddy, Shivani Chaudhary, William Day, Maarij Ahmed, Hayan Haqqi, Muhammad Ahsen Fahim, Hanzallah Amjad, Ahmad Orakzai, Aqsa Gul, et al. 2026. FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks.arXiv preprint arXiv:2604.05912(2026)

  17. [17]

    Srivatsa Kundurthy, Clara Na, Colton Moraine, Anoushka Mohta, Case Winter, George Fang, John Ling, Emma Strubell, and Zach Kirshner. 2026. BlueFin: Bench- marking LLM Agents on Financial Spreadsheets.arXiv preprint arXiv:2605.30907 (2026)

  18. [18]

    Elaine Lau, Markus Dücker, Ronak Chaudhary, Hui Wen Goh, Rosemary Wei, Vaibhav Kumar, Saed Qunbar, Guram Gogia, Yi Liu, Scott Millslagle, et al. 2026. BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows. InRLEval: Methods and Reinforcement Learning Environments for Evaluating AI Agents

  19. [19]

    Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu, Xiaokang Zhang, Xiaohan Zhang, Sijia Luo, Xi Wang, and Jie Tang. 2024. Spreadsheetbench: Towards challenging real world spreadsheet manipulation.Advances in Neural Information Processing Systems37 (2024), 94871–94908

  20. [20]

    MiniMax. 2026. MiniMax M3: Frontier Coding, 1M Context, Native Multimodality — All in One Model. https://www.minimax.io/blog/minimax-m3. Accessed 2026- 07-23

  21. [21]

    OpenAI. 2025. gpt-oss-120b & gpt-oss-20b Model Card. arXiv:2508.10925 [cs.CL]

  22. [22]

    OpenAI. 2026. GPT-5.6 System Card. https://deploymentsafety.openai.com/gpt- 5-6. Accessed 2026-07-23

  23. [23]

    Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024. tinyBenchmarks: evaluating LLMs with fewer examples. arXiv preprint arXiv:2402.14992(2024)

  24. [24]

    Qwen Team. 2025. Qwen3-Coder: Agentic Coding in the World. https://qwenlm. github.io/blog/qwen3-coder/. Accessed 2026-07-23

  25. [25]

    Qwen Team. 2026. Qwen3.7-Max Model Documentation, Alibaba Cloud Model Studio. https://www.alibabacloud.com/help/en/model-studio/models. Accessed 2026-07-23

  26. [26]

    Anka Reuel, Amelia Hardy, Chandler Smith, Max Lamparth, Malcolm Hardy, and Mykel J Kochenderfer. 2024. Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices.Advances in Neural Information Processing Systems37 (2024), 21763–21813

  27. [27]

    Bytedance Seed. 2026. Seed1. 8 model card: Towards generalized real-world agency.arXiv preprint arXiv:2603.20633(2026)

  28. [28]

    StepFun. 2026. Step 3.7 Flash: A High-Efficiency Flash Model for Real-World Agents. https://static.stepfun.com/blog/step-3.7-flash/. Accessed 2026-07-23

  29. [29]

    Tencent Hunyuan Team. 2026. Tencent Hunyuan Officially Releases Hy3, Ad- vancing Agent Capabilities and Deeper Product Integration. https://hunyuan. tencent.com/research/100064?langVersion=zh. Accessed 2026-07-23

  30. [30]

    A Wang, G Meinhardt, J Katz, JH Kim, PK Chaudhary, C Blagden, and E Xu. 2026. BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents.arXiv preprint arXiv:2606.03829(2026)

  31. [31]

    Sinuo Wang, WANG PIAOHONG, Tianrui Qin, Maojia Song, Qianben Chen, Qiexiang Wang, Gengze Zhou, Zeyu Zhang, He Zhu, Dingfeng Shi, et al. 2026. EVOLVING ROLLOUTS: Harnessing Historical Experience for Web Agent Evo- lution in Reinforcement Learning. InForty-third International Conference on Machine Learning

  32. [32]

    Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, et al. 2024. Livebench: A challenging, contamination-limited llm benchmark.arXiv preprint arXiv:2406.19314(2024)

  33. [33]

    xAI. 2026. Introducing Grok 4.5. https://x.ai/news/grok-4-5. Release announce- ment, July 2026

  34. [34]

    An Yang, Anfeng Li, Baosong Yang, et al . 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL]

  35. [35]

    Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, et al. 2026. Glm-5: from vibe coding to agentic engineering.arXiv preprint arXiv:2602.15763(2026)

  36. [36]

    Taojie Zhu, Wentao Zhao, Rui Sun, Beidi Luan, Jiacheng Lu, Sinuo Wang, Jing Li, Daxin Jiang, Yonghong He, and Zuo Bai. 2026. From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets. arXiv preprint arXiv:2605.28359(2026)

  37. [37]

    922 tickers, 25 GICS industry groups

    Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Shu Liu, Sasha Cui, Sayash Kapoor, Shayne Longpre, Kevin Meng, Rebecca Weiss, et al. 2026. Establishing best practices in building rigorous agentic benchmarks.Advances in Neural Information Processing Systems38 (2026). KDD ’27, August 2027, San Jose, CA, USA Appendix Appendix Contents The appendix i...

  38. [38]

    Create the ./output/ directory if it does not exist

    Save the workbook to ./output/{TICKER}_model.xlsx. Create the ./output/ directory if it does not exist

  39. [39]

    Produce a SINGLE workbook with all sheets in it

  40. [40]

    11 KDD ’27, August 2027, San Jose, CA, USA Appendix 0 0.25 0.50 0.75 1 Flat single-golden score of one analyst vs

    Every projected/forecasted number must be a live Excel **formula** that references inputs -- never a value computed in Python and written as a number. 11 KDD ’27, August 2027, San Jose, CA, USA Appendix 0 0.25 0.50 0.75 1 Flat single-golden score of one analyst vs. a peer 0 20 40 60 80 100 Cumulative % of 108 pairs 1× 2× 3× 4× 0.70 92.6% of pairs below 0....

  41. [41]

    The model must produce a final **implied share price** that is clearly labeled and easy to find

  42. [42]

    Use real accounting conventions (GAAP-style). The Income Statement, Balance Sheet, and Cash Flow Statement must tie together -- net income flows to retained earnings and to the cash flow statement, ending cash on CF equals cash on BS, etc

  43. [43]

    TODO" cells, no

    No placeholder text, no "TODO" cells, no "Excel Data Table feature" notes. The workbook must be fully functional when opened. Do not ask clarifying questions. Make reasonable analyst-grade assumptions. Document them in the workbook (cell comments or an Assumptions sheet) but do not block on them. Work efficiently. Spend your reasoning on the model itself,...

  44. [44]

    {model_xlsx} A single institutional-quality .xlsx with these sheets: Cover, Assumptions, Revenue_Build (segment drivers), Income_Statement, Balance_Sheet, Cash_Flow, Debt_Schedule, WACC, DCF (Valuation), Sensitivity, Checks. Hard requirements (machine-graded + judged): - 5 forecast years after the last actual; every forecast cell is a live FORMULA referen...

  45. [45]

    {memo_md} A short investment memo: thesis with 3-5 falsifiable, quantified claims; key risks mapped to model drivers; headline numbers (implied price, EPS) that MATCH the workbook

  46. [46]

    name","value

    {assum} JSON list of every key assumption: {"name","value","source"} where source is a resolvable pointer ("Sheet!Cell", or the input row it came from). Every number you cite in the memo must appear here. Build the workbook now. When finished, reply with ONLY the path you wrote. Do not narrate. D.3 Rendered Input Pack: Excerpt The pack is dumped sheet by ...

  47. [47]

    A5"] = "Cash & Equivalents

    spans 2022A–2029E): Goodwill & Intangibles 15,000, Other Non- Current Assets 5,000, Other Current Assets 500 — none of these values appears anywhere in the input pack — while the pack’s gen- uine FY2024 cash figure (3,127) is copied backwards into 2022A and 2023A as well. Each constant is styled blue_font, the analyst con- vention for a legitimate hardcod...

  48. [494]

    abstain, don’t fabricate

    — so the partial-coverage limitation of Section 7 is checkable, not just confessed. The analyst-vs-analyst audit (Section 4) releases its 632 per-pair, per-tolerance records (158 directed same-ticker pairs at four tolerance multipliers), so the paper’s central negative result is recomputable from JSONL. And every judged cell in the frozen run retains all ...

This paper was first reviewed by grok-4.5 on July 31, 2026.