Pith. sign in

REVIEW 9 cited by

INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18174 v1 pith:FFWK224R submitted 2024-12-24 cs.CE cs.AIq-fin.CP

classification cs.CEcs.AIq-fin.CP
keywords financialdecision-makingagentagentstaskscomprehensiveinvestorbenchacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements have underscored the potential of large language model (LLM)-based agents in financial decision-making. Despite this progress, the field currently encounters two main challenges: (1) the lack of a comprehensive LLM agent framework adaptable to a variety of financial tasks, and (2) the absence of standardized benchmarks and consistent datasets for assessing agent performance. To tackle these issues, we introduce \textsc{InvestorBench}, the first benchmark specifically designed for evaluating LLM-based agents in diverse financial decision-making contexts. InvestorBench enhances the versatility of LLM-enabled agents by providing a comprehensive suite of tasks applicable to different financial products, including single equities like stocks, cryptocurrencies and exchange-traded funds (ETFs). Additionally, we assess the reasoning and decision-making capabilities of our agent framework using thirteen different LLMs as backbone models, across various market environments and tasks. Furthermore, we have curated a diverse collection of open-source, multi-modal datasets and developed a comprehensive suite of environments for financial decision-making. This establishes a highly accessible platform for evaluating financial agents' performance across various scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    cs.CL 2026-07 conditional novelty 7.0 of 10

    On a new 615-question business-case benchmark graded by AI against instructor rubrics, frontier LLMs score 87-88% partial credit but complete only about half the questions.

  2. CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    CLQT is a new closed-loop, cost-aware benchmark that diagnoses LLM trading agent capabilities through strategy-consistent metrics and hash-verifiable trails rather than outcome rankings.

  3. InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    InvestPhilBench is a new multi-layer benchmark for LLM procedural reasoning in investment philosophy, with BASP metrics showing composite scores saturate while gate reconstruction accuracy reveals procedural deficits.

  4. ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Across five 100-task agent streams, sequential experience improves normalized reward by 16.9% in 14 of 15 model-domain combinations, but explicit skill maintenance matches pure in-context learning (0.602 vs 0.605) and...

  5. Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CM-LRS is a workflow-output-layer reliability score for capital-markets LLM outputs; on five public-document workflows, frontier closed-source models score 4.09–4.31 and Llama 3.3 70B scores 3.15 under four LLM judges.

  6. NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management

    cs.AI 2026-07 conditional novelty 6.0 of 10

    NextFund unifies live multi-market data, multi-agent portfolio reasoning, and end-to-end decision traces with an interactive Trading Arena for fairer LLM agent benchmarking.

  7. StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets

    cs.CE 2025-07 conditional novelty 6.0 of 10

    StockSim provides a dual-mode simulated stock market, with order-level and candlestick-level execution, for evaluating LLM trading agents.

  8. CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 9,356-pair Chinese multimodal financial benchmark reveals that state-of-the-art multimodal LLMs, including GPT-4V, still score below 53% on objective and 39% on subjective financial chart tasks.

  9. FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A two-stage RL framework with length, image-selection, and adversarial rewards, trained on 89,378 ASP-built financial image-question pairs, improves multimodal reasoning over LMM-R1.

Pith tools