Pith. sign in

REVIEW 6 cited by

Financial Statement Analysis with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.17866 v3 pith:OK3EBBUO submitted 2024-07-25 q-fin.ST cs.AIcs.CLq-fin.GNq-fin.PM

classification q-fin.STcs.AIcs.CLq-fin.GNq-fin.PM
keywords financialanalysisanalystsmodelsearningsfindfuturehuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate whether large language models (LLMs) can successfully perform financial statement analysis in a way similar to a professional human analyst. We provide standardized and anonymous financial statements to GPT4 and instruct the model to analyze them to determine the direction of firms' future earnings. Even without narrative or industry-specific information, the LLM outperforms financial analysts in its ability to predict earnings changes directionally. The LLM exhibits a relative advantage over human analysts in situations when the analysts tend to struggle. Furthermore, we find that the prediction accuracy of the LLM is on par with a narrowly trained state-of-the-art ML model. LLM prediction does not stem from its training memory. Instead, we find that the LLM generates useful narrative insights about a company's future performance. Lastly, our trading strategies based on GPT's predictions yield a higher Sharpe ratio and alphas than strategies based on other models. Our results suggest that LLMs may take a central role in analysis and decision-making.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    InvestPhilBench is a new multi-layer benchmark for LLM procedural reasoning in investment philosophy, with BASP metrics showing composite scores saturate while gate reconstruction accuracy reveals procedural deficits.

  2. FinTexTS: Financial Text-Paired Time-Series Dataset via Semantic-Based and Multi-Level Pairing

    cs.AI 2026-03 conditional novelty 6.0 of 10

    A new dataset and pairing framework links stock prices to semantically relevant news at macro, sector, related-company, and target-company levels, improving stock forecast accuracy over keyword-based pairing.

  3. Mamba Drafters for Speculative Decoding

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Mamba-based drafters can match self-speculation throughput with lower memory and cross-model flexibility.

  4. Breaking Down Bias: On The Limits of Generalizable Pruning Strategies

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Pruning-based bias removal in Llama-3-8B reduces racial bias mainly in the context used to choose what to prune, and transfers poorly across contexts.

  5. LLMs Can Teach Themselves to Better Predict the Future

    cs.CL 2025-02 conditional novelty 6.0 of 10

    DPO fine-tuning on outcome-ranked self-play forecasts improves LLM Brier scores by 7 to 10 percent over base and randomized-label controls.

  6. Cognitive Agents Powered by Large Language Models for Agile Software Project Management

    cs.SE 2025-08 reject novelty 4.0 of 10

    LLM agents acting as Agile roles produced plausible project artifacts in simulation, but the claimed improvements over human teams are unsupported because no comparison or validated metrics are provided.

Pith tools