Pith. sign in

REVIEW 3 major objections 4 minor 21 references

FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper introduces FinEvo-Bench, a longitudinal benchmark of 120 financial tasks that measures whether agents improve on later tasks from retained experience, and reports that all four self-evolving scaffolds outperform non-evolving…

desk verdict FinEvo-Bench's paired state-reset design is the right way to measure self-evolution, but the headline gains rest on a judge validated by one expert on one run, so treat the numbers as promising rather than established. read the letter →

arxiv 2608.06144 v1 pith:4IY2S5WT submitted 2026-08-06 cs.AI

classification cs.AI
keywords self-evolvingagentslongitudinalbenchmarkfinancialworkflowsretainedexperiencerubricevaluationcompliancepairedcontrolstaskstreams
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

FinEvo-Bench is designed to answer a question most agent benchmarks miss: does experience from one professional task actually help an agent on later, related tasks? The paper builds 120 real-case-grounded financial tasks grouped into 20 scenes that share professional procedures, then runs four self-evolving agent scaffolds through three independently shuffled task streams. Each evolving run is paired with a state-reset control that shares the same task order, so any score difference isolates the benefit of retained experience. Across all scaffolds, the evolving condition raises mean scores by 9.33–19.37 points and reduces compliance issues by 0.12–0.44 per task. If these gains hold, FinEvo-Bench provides a reusable way to measure self-evolution ability separately from absolute performance.

What carries the argument

The measurement apparatus is a paired longitudinal protocol: each evolving run processes a globally interleaved task stream with an execute–score-and-feedback–reflect-and-consolidate cycle, while a paired non-evolving control resets all agent state before every task and shares the same task order, backbone, decoding, and scoring procedure. Any difference between the two conditions is attributed to retained experience. Score quality and financial compliance are assessed by scene-level rubrics derived from professional procedures, applied by an independent automated rubric judge whose scores on one validation run agree closely with a financial expert (ICC(A,1) = 0.95, mean absolute difference 1.6 points). Experience utilization is tracked by logging when stored memory or skills enter the execution context.

What would settle it

Score an additional complete run of 120 deliverables, or a stratified sample covering each scaffold and both feedback conditions, with independent financial experts using the same scene rubrics; if the intraclass correlation falls substantially below the reported ICC(A,1) = 0.95 or the mean absolute difference rises well above 1.6 points, the measured self-evolution gains cannot be trusted as measures of professional quality.

Watch

Extended reading notes

Core claim

The central claim is that retained experience produces positive self-evolution gains in professional financial workflows. Using paired state-reset controls on three globally interleaved task streams, the paper reports that all four scaffolds improve over their non-evolving counterparts: relative to controls, mean scores rise by 9.33–19.37 points and mean compliance issues fall by 0.12–0.44 per task. The highest evolved score reaches 91.65 and the largest self-evolution gain is +19.37. Within scenes, paired score gains at later occurrence ranks (4–6) exceed those at earlier ranks (1–3) by 6.10–8.70 points, and compliance reductions also grow, supporting gradual experience accumulation. The paper also finds that rubric-based feedback outperforms reference-answer feedback for every scaffold, that skill-only retention beats memory-only and combined memory–skill retention in one scaffold, and that all five measured capability dimensions improve, with report quality showing the largest gains.

Load-bearing premise

The automated rubric judge must keep agreeing with financial experts on every run, scaffold, and feedback condition; it was checked on only one run of 120 deliverables, and if that agreement does not generalize, every central score and gain in the paper is questionable.

Editorial extensions

If this is right

  • Self-evolution ability can be measured separately from absolute capability: a scaffold can have a high final score, a large gain from experience, or low cost, and these do not coincide.
  • Because later same-scene tasks show larger paired gains than earlier ones, the benchmark supports the interpretation that experience accumulates and transfers across related professional cases rather than merely improving within a single episode.
  • Rubric feedback that identifies missing evidence, incomplete analysis, and compliance failures is more effective for later tasks than providing a complete reference answer, suggesting that diagnostic feedback is what drives cross-task improvement.
  • In one scaffold, skill-only evolution outperforms memory-only and combined memory–skill evolution, suggesting that reusable procedural knowledge carries more of the benefit in recurring workflows than raw memory traces.
  • All five quality dimensions improve, with report quality improving most, indicating that transferable structure and expression improve more readily than case-specific evidence use and analysis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The automated judge's agreement was validated on one run of 120 deliverables; a fuller validation on additional runs, scaffolds, and feedback conditions could reveal whether the reported gains are partly scoring noise, and that validation is a natural next check.
  • The activation gap (activated tasks score 8.31–9.06 points higher than non-activated tasks) suggests a testable design rule: improving the retrieval index for scaffolds with low activation rates may raise their self-evolution gains and reduce their cost–quality tradeoff.
  • The mixed cross-scene interference results hint that always-on memory may benefit from broad cross-scene accumulation while selective retrieval may perform better in homogeneous scene stores; this could be tested by varying memory-injection policies independently of the scaffold.
  • If the protocol generalizes beyond finance, the paired-control interleaved-stream design could become a template for measuring self-evolution in other professional domains where tasks share procedures but differ in facts and conclusions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. FinEvo-Bench is a longitudinal benchmark of 120 professional financial tasks organized into 20 scenes, with six related cases per scene sharing an expert-reviewed rubric. Four scaffolds (Claude Code, Codex, Letta, GenericAgent) run on the same Qwen3.7-Max backbone over three independently shuffled global streams, and each evolving run is paired with a state-reset control to isolate the effect of retained experience. An automated Claude Code judge scores all deliverables. The paper reports positive self-evolution gains for all scaffolds (score +9.33 to +19.37 points, compliance issues down 0.12 to 0.44 per task), larger gains at later within-scene ranks, skill-only superiority in one carrier analysis, rubric-feedback superiority over reference-answer feedback, and gains across five capability dimensions. The appendix documents scene construction and provides a full feedback-to-memory/skill consolidation trace.

Significance. If the measurement assumptions hold, FinEvo-Bench is a valuable and relatively unique instrument: it combines professional workflows, open-ended deliverables, multi-aspect scoring, and longitudinal interleaving, and its paired state-reset design is a principled way to separate retained-experience gains from task-order effects. The shared-backbone comparison across four scaffolds, three shuffled streams, and the detailed appendix trace are clear strengths. The main reservation is that the entire empirical contribution rests on an automated judge whose agreement with a human expert has been checked on only one run, and whose feedback is fed back into the evolving condition but not the control.

major comments (3)
  1. [§4.2] The human validation of the automated judge is load-bearing but narrow. The 120 deliverables come from one complete main-experiment run, and the manuscript does not state which scaffold, which shuffled stream, or which feedback condition produced them; it also validates only the 0–100 score (ICC(A,1)=0.95), not the count of triggered compliance issues that drives the compliance-reduction claims. Because the evolving condition receives judge feedback in Stage 3 and the reset control does not, a judge-specific scoring bias toward rubric-like output would inflate every reported gain in Tables 3–7 and Figures 3–4. Please validate the judge on multiple scaffolds and on both evolving and control outputs, add compliance-issue-level agreement, and state whether the validating financial expert was one of the two rubric authors.
  2. [§4.3 and Tables 3–7] All central estimates are means over K=3 independent shuffles with no variance, confidence interval, or significance test. With only three runs, the claims that all four scaffolds show positive gains and that late gains exceed early gains are not yet statistically supported. Report per-run paired differences and a dispersion estimate (per-run means, bootstrap CI, or a paired test) for the headline score gain, the compliance reduction, and the early-versus-late comparison.
  3. [§4.1 and Appendix B] The stated rubric isolation is contradicted by the described pipeline. §4.1 says the scaffold never receives the judge's full item-level record and receives only 'rubric-based feedback summarizing the problems,' but the Stage 2 prompt in Appendix B.3 instructs the judge to write 'the complete scoring result and feedback' to judge.md, and Stage 3 resumes the conversation with that file. The worked trace in Appendix C contains the complete section-level score and every deduction. If judge.md is the full scoring record, the evolving condition receives substantially more rubric information than the text claims, and the reported gains may partly reflect feedback-specific optimization rather than generalized professional learning. Please either restrict Stage 3 to summarized feedback or revise the isolation claim and discuss the consequence.
minor comments (4)
  1. [§2 and Table 1] The comparison table would be easier to read if each benchmark's exact evaluation metric (e.g., success rate, rubric score) were stated, because some rows mark a property with a checkmark but do not indicate how that property was operationalized.
  2. [§4.1] The phrase 'Claude Code' is used both for one of the evaluated scaffolds and for the independent scoring agent; please disambiguate in the text, for example by calling the latter 'the rubric judge' throughout.
  3. [§4.3 / Figure 3] The early-versus-late gain comparison is acknowledged to be confounded with global position, but the conclusion 'supporting gradual experience accumulation' is stated more strongly than the design allows; the sentence should explicitly attribute the trend to both same-scene and global experience, as the earlier caveat does.
  4. [§5] The limitations paragraph does not mention the judge-validation scope; adding a sentence noting that automated-scoring validity was checked on a single run would make the limitations consistent with §4.2.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central self-evolution gains are estimated by a paired evolving-versus-state-reset control with identical task order, backbone, decoding, and judge; the conclusion is not equivalent to its inputs.

full rationale

The paper's central claim—retained experience raises scores by 9.33–19.37 points and lowers compliance issues by 0.12–0.44 per task—is a between-condition difference computed from Equations (2)–(3), where evolving and non-evolving runs share task order, backbone, decoding configuration, and scoring procedure, and differ only in whether prior-task feedback is retained. No fitted parameter is renamed as a prediction; the judge is validated externally (ICC(A,1)=0.95 on 120 deliverables) and then applied identically to both conditions. The judge-derived feedback given in Stage 3 is not the full rubric, and later tasks in a scene are substantively distinct cases, so improvement requires generalizing procedures rather than copying answers; the gain is therefore not forced by construction. Self-citations that appear (e.g., the GenericAgent scaffold description and SEAGym/FinMCP-Bench references) are used to describe scaffolds or related work, not to justify the benchmark's central inference, and no uniqueness theorem or ansatz is imported from prior author work. The main caveats—automated-judge validation on a single run and one financial expert, and the acknowledged scope limits (one backbone, non-parametric evolution, one-carrier comparison, five-scene diagnostic)—are measurement-validity limitations rather than circular reductions. Accordingly, the derivation chain is self-contained: the empirical claim would fail if the control condition performed as well as the evolving condition, which is a genuinely falsifiable outcome.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The measurement rests on expert-curated artifacts (procedures, rubrics, anonymized cases) and an automated judge; the preprint does not release these artifacts, so the central numbers cannot be independently recomputed. The hand-chosen rubric weights and thresholds are the closest things to free parameters: they are not fitted to data, but they directly determine all reported gains.

free parameters (3)
  • Scene rubric point allocations = e.g., Financial Statement Analysis: 8/15/12/13/10/15/5/12/10 over 100
    Hand-set by two domain experts in Section 3.3 and Appendix A.3; these weights determine every score, gain, and compliance metric, so the central claims are sensitive to them.
  • Rubric thresholds and tolerances = e.g., over-one-year receivables share above 30%; 0.5 percentage-point rounding tolerance
    Defined in Appendix A.3 Tables 15 and 16; they determine full versus partial credit and compliance triggers, and changing them would shift reported gains.
  • Early/late rank split = within-scene ranks 1-3 vs ranks 4-6
    Chosen in Section 4.1 equations (5) and (6); the conclusion that late gains exceed early gains by 6.10-8.70 points depends on this fixed split.
assumptions (5)
  • domain assumption Institution-provided professional procedures P_s are valid and representative of real financial workflows.
    Invoked in Section 3.2 as the source of required operations, deliverables, and constraints; the benchmark's professional validity rests on this.
  • domain assumption Two-expert scene rubrics correctly operationalize task quality and financial compliance.
    Section 3.3 describes rubric drafting and review; the rubrics are the sole scoring instrument and are not independently auditable from the preprint.
  • domain assumption Anonymization of institution-provided cases preserves numerical logic, chronology, and decision conditions.
    Section 3.2 states identifiers are removed while preserving data relationships; no quantitative leakage check is shown.
  • domain assumption The Claude Code judge's scores generalize from the one validated run to all runs, scaffolds, and feedback conditions.
    Section 4.2 validates ICC(A,1)=0.95 on 120 deliverables from one run; all other reported numbers in Tables 3-7 use this judge without further validation.
  • domain assumption State-reset controls remove all prior-task experience without introducing other behavioral changes.
    Section 4.1 defines the paired non-evolving condition as resetting agent state before every task; the entire self-evolution gain estimate depends on this clean counterfactual.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows." pith.science (2026). https://pith.science/paper/4IY2S5WT

@misc{pith2026260806144,
  author       = {Pith},
  title        = {Pith review of: FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4IY2S5WT}},
  note         = {Machine review of arXiv:2608.06144}
}
read the original abstract

Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains. Institution-provided professional procedures define the required operations and constraints. Eligible institution-provided and publicly documented cases supply the task facts. Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance. We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams. Paired non-evolving controls estimate each scaffold's self-evolution gain from retained experience, while an independent Claude Code scoring agent backed by Claude Opus 4.6 evaluates all outputs. Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37). Across scaffolds, the evolving condition raises scores by 9.33-19.37 points and reduces compliance issues by 0.12-0.44 per task. Paired score gains at within-scene ranks 4-6 exceed those at ranks 1-3 by 6.10-8.70 points. In Claude Code, skill-only evolution produces higher task quality and fewer compliance issues than memory-only and combined memory-skill evolution. Across all four scaffolds, rubric feedback also yields higher scores and fewer compliance issues than reference-answer feedback. FinEvo-Bench measures both professional performance and self-evolution ability: how effectively an agent turns prior experience into later improvement.

Figures

Figures reproduced from arXiv: 2608.06144 by the authors.

Figure 1
Figure 1. Overview of FinEvo-Bench. Left: 20 business scenes across six financial domains, comprising 120 tasks. Right: an [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. FinEvo-Bench construction and cross-task self-evolution. (A) Dataset construction. (B) Experimental setup. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Self-evolution score gain by within-scene occur [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Percentage-point self-evolution gains across five [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Stage 1 prompt and workspace state. The two text components are sent in one message and do not appear as workspace [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Stage 2 prompt and workspace state. The fixed Claude Code judge produces one scoring and feedback file. [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Stage 3 prompt and workspace state. Feedback is returned to the exact conversation used for Stage 1. [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: The Task 04 memory update. The retained content abstracts from the current company and records reusable failure [PITH_FULL_IMAGE:figures/full_fig_p019_8.png]
Figure 9
Figure 9. Figure 9: Selected final content of the feedback memory after Task 04. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: Selected final profitability and asset-quality instructions in the updated skill. [PITH_FULL_IMAGE:figures/full_fig_p021_10.png]
Figure 11
Figure 11. Figure 11: Final chain-substitution instructions added to the updated skill. [PITH_FULL_IMAGE:figures/full_fig_p021_11.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

21 extracted references · 10 canonical work pages

  1. [1]

    SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities

    BenYoash, N.; Brief, M.; Ovadia, O.; Shenderovitz, G.; Mishaeli,M.;Lemberg,R.;andSheetrit,E.2025. SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities. InProceedings of the Fourth Workshop on Generation, Evaluation and Metrics. Bigeard, A.; Nashold, L.; Krishnan, R.; and Wu, S

  2. [3]

    Gross margin: 42.00% (2023), 40.02% (2022), and 40.00% (2021)

    Table 21 maps the archived artifacts to the three-stage protocol in Appendix B. Stage Archived artifact Content used in this example Result 1answer.txtThe complete financial-analysis deliverable generated from the task files 369-line Markdown report 2judge.txtSection scores, deductions, violation count, and textual feedback 88/100; 0 violations 3before_ME...

  3. [4]

    Islam, P.; Kannappan, A.; Kiela, D.; Qian, R.; Scherrer, N.; and Vidgen, B

    FinAgent- Bench:ABenchmarkDatasetforAgenticRetrievalinFinan- cial Question Answering.arXiv preprint arXiv:2508.14052. Islam, P.; Kannappan, A.; Kiela, D.; Qian, R.; Scherrer, N.; and Vidgen, B

  4. [7]

    BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

    BlueFin: Benchmarking LLM Agents on Financial Spreadsheets. arXiv preprint arXiv:2605.30907. Liang, J.; Han, J.; Li, W.; Wang, X.; Zhang, Z.; Jiang, Z.; Liao,Y.;Li,T.;Huang,Y.;Shen,H.;Wu,H.;Guo,F.;Wang, K.; Hong, Z.; Lu, Z.; Ma, L.; Jiang, S.; and Xiao, Y

  5. [8]

    Shinn, N.; Cassano, F.; Berman, E.; Gopinath, A.; Narasimhan, K.; and Yao, S

    GenericAgent:AToken-EfficientSelf-EvolvingLLMAgent via Contextual Information Density Maximization.arXiv preprint arXiv:2604.17091. Shinn, N.; Cassano, F.; Berman, E.; Gopinath, A.; Narasimhan, K.; and Yao, S

  6. [9]

    Wang, Y.; Zhang, Z.; Chi, M.; et al

    Voyager: An Open- EndedEmbodiedAgentwithLargeLanguageModels.arXiv preprint arXiv:2305.16291. Wang, Y.; Zhang, Z.; Chi, M.; et al

  7. [10]

    Wang, Z.; Wang, K.; Wang, Q.; et al

    EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Per- spective.arXiv preprint arXiv:2605.18421. Wang, Z.; Wang, K.; Wang, Q.; et al

  8. [11]

    Wei, T.; Sachdeva, N.; Coleman, B.; et al

    RAGEN: Un- derstanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.arXiv preprint arXiv:2504.20073. Wei, T.; Sachdeva, N.; Coleman, B.; et al

Show all 21 references
  1. [12]

    Xie, Q.; Han, W.; Chen, Z.; et al

    Evo- Memory: Benchmarking LLM Agent Test-time Learn- ing with Self-Evolving Memory.arXiv preprint arXiv:2511.20857. Xie, Q.; Han, W.; Chen, Z.; et al. 2024a. FinBen: A Holistic Financial Benchmark for Large Language Models. InAd- vances in Neural Information Processing Systems...

  2. [13]

    InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track

    PIXIU: A Comprehensive Benchmark, Instruction Dataset and Large Language Model for Finance. InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track. Xie, T.; Zhang, D.; Chen, J.; et al. 2024b. OSWorld: Bench- marking Multimodal Agents for Open-Ended ...

  3. [14]

    InInterna- tional Conference on Machine Learning

    Offline Training of Language Model Agents with Functions as Learnable Weights. InInterna- tional Conference on Machine Learning. Zhao,A.;Huang,D.;Xu,Q.;Lin,M.;Liu,Y.-J.;andHuang, G.2024. ExpeL:LLMAgentsAreExperientialLearners. In AAAI Conference on Artificial Intelligence. Zhe...

  4. [15]

    Zhou, S.; Xu, F

    LifelongAgentBench: EvaluatingLLMAgentsasLifelongLearners.arXivpreprint arXiv:2505.11942. Zhou, S.; Xu, F. F.; Zhu, H.; et al

  5. [16]

    InInternational Conference on Learning Representations

    WebArena: A Re- alistic Web Environment for Building Autonomous Agents. InInternational Conference on Learning Representations. Zhu,F.;Lei,W.;Huang,Y.;etal.2021.TAT-QA:AQuestion Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance. InProceedings of ACL-IJC...

  6. [17]

    FinMCP- Bench:BenchmarkingLLMAgentsforReal-WorldFinancial Tool Use under the Model Context Protocol.arXiv preprint arXiv:2603.24943. A Benchmark Construction: A Worked Financial-Statement-Analysis Scene This appendix documents the construction process from Section 3 for theFin...

  7. [18]

    C Complete Example of Feedback-Driven Experience Consolidation This appendix follows one Task 04 trace from the Financial Statement Analysis scene. The task asked the evaluated Codex agent to prepare a risk-oriented analysis of an anonymized biomedical company’s 2021–2023 cons...

  8. [20]

    The feedback identifies the problem and its evidence, but does not expose the complete rubric or a reference answer

    Table 25 reproduces all six feedback items. The feedback identifies the problem and its evidence, but does not expose the complete rubric or a reference answer. No. Rubric item Judge feedback Deduction 1 Core profitability metric The 2022 gross margin was reported as 40.02%; t...

  9. [21]

    Feedback target Evidence in the submitted report Required correction or extension Gross-margin arithmetic 2022 revenue was 115.05 and cost was 67.88, but the table reported 40.02%

    ThenumericalandcoverageimplicationsofthefeedbackareshowninTable26.Thesevaluesarecomputeddirectlyfromthe Stage 1 report and make explicit what the next execution should verify or add. Feedback target Evidence in the submitted report Required correction or extension Gross-margin...

  10. [2023]

    Jiang,S.;Ma,L.;Hong,Z.;etal.2026

    FinanceBench: A New Bench- mark for Financial Question Answering.arXiv preprint arXiv:2311.11944. Jiang,S.;Ma,L.;Hong,Z.;etal.2026. SEA-Eval:ABench- mark for Evaluating Self-Evolving Agents Beyond Episodic Assessment.arXiv preprint arXiv:2604.08988. Jimenez, C. E.; Yang, J.; W...

  11. [2024]

    Jin,S.;Li,S.;Zhang,S.;andYan,R.2026

    SWE-bench: Can Language ModelsResolveReal-WorldGitHubIssues? InInternational Conference on Learning Representations. Jin,S.;Li,S.;Zhang,S.;andYan,R.2026. FinRpt:Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation. InProceedings ...

  12. [2025]

    Cai,Y.;Hao,Y.;Zhou,J.;etal.2025.BuildingSelf-Evolving AgentsviaExperience-DrivenLifelongLearning:AFrame- work and Benchmark.arXiv preprint arXiv:2508.19005

    Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks.arXiv preprint arXiv:2508.00828. Cai,Y.;Hao,Y.;Zhou,J.;etal.2025.BuildingSelf-Evolving AgentsviaExperience-DrivenLifelongLearning:AFrame- work and Benchmark.arXiv preprint arXiv:2508.19005. Chen,...

  13. [2026]

    Choi, C.; Kwon, J.; Lopez-Lira, A.; et al

    Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engi- neering Tasks with Generative Optimization.arXiv preprint arXiv:2604.12290. Choi, C.; Kwon, J.; Lopez-Lira, A.; et al

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.