Pith. sign in

REVIEW 16 cited by

FinQA: A Dataset of Numerical Reasoning over Financial Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.00122 v3 pith:UZI26C55 submitted 2021-09-01 cs.CL

classification cs.CL
keywords financialdatasetreasoningnumericalcomplexdomainfinqadata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The sheer volume of financial statements makes it difficult for humans to access and analyze a business's financials. Robust numerical reasoning likewise faces unique challenges in this domain. In this work, we focus on answering deep questions over financial data, aiming to automate the analysis of a large corpus of financial documents. In contrast to existing tasks on general domain, the finance domain includes complex numerical reasoning and understanding of heterogeneous representations. To facilitate analytical progress, we propose a new large-scale dataset, FinQA, with Question-Answering pairs over Financial reports, written by financial experts. We also annotate the gold reasoning programs to ensure full explainability. We further introduce baselines and conduct comprehensive experiments in our dataset. The results demonstrate that popular, large, pre-trained models fall far short of expert humans in acquiring finance knowledge and in complex multi-step numerical reasoning on that knowledge. Our dataset -- the first of its kind -- should therefore enable significant, new community research into complex application domains. The dataset and code are publicly available\url{https://github.com/czyssrs/FinQA}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    cs.CL 2026-07 conditional novelty 7.0 of 10

    On a new 615-question business-case benchmark graded by AI against instructor rubrics, frontier LLMs score 87-88% partial credit but complete only about half the questions.

  2. On the Fitness Landscape in the $NK$ Model

    math.PR 2025-08 unverdicted novelty 7.0 of 10

    For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.

  3. Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CM-LRS is a workflow-output-layer reliability score for capital-markets LLM outputs; on five public-document workflows, frontier closed-source models score 4.09–4.31 and Llama 3.3 70B scores 3.15 under four LLM judges.

  4. FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain

    cs.CL 2025-07 conditional novelty 6.0 of 10

    FinGAIA is a 407-task Chinese financial agent benchmark where the best agent, ChatGPT DeepResearch, scores 48.9%, far below financial experts at 84.7%.

  5. CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A 9,356-pair Chinese multimodal financial benchmark reveals that state-of-the-art multimodal LLMs, including GPT-4V, still score below 53% on objective and 39% on subjective financial chart tasks.

  6. Human-Centric Evaluation for Foundation Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    In free-form research collaborations rated by humans, Grok 3 scores highest, followed by DeepSeek R1 and Gemini 2.5; OpenAI o3 mini trails.

  7. FinS-Pilot: A Benchmark for Online Financial RAG System

    cs.CL 2025-05 reject novelty 6.0 of 10

    FinS-Pilot is a small benchmark of real-world financial assistant queries with real-time API data and text corpus, used to compare Chinese LLMs on financial RAG tasks.

  8. When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents

    cs.CL 2026-02 conditional novelty 5.0 of 10

    A new French financial document benchmark shows vision-language models are strong at text/table extraction but brittle on charts and multi-turn dialogue, with accuracy converging near 50% in conversational settings.

  9. VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding

    cs.CE 2025-08 conditional novelty 5.0 of 10

    VisFinEval is a 15,848-question Chinese multimodal financial benchmark covering eight image types across three workflow depths, on which the best model still trails finance experts by more than 14 points.

  10. CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A fine-tuned Llama 3 model with a trained document critic and program-based reasoning beats GPT-4o and other baselines on a new carbon footprint QA benchmark.

  11. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.

  12. Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A plug-and-play external reward model with speculative rejection sampling cuts tree-search cost for LLM decision-making to about 1/10 while keeping or slightly improving accuracy on math, planning, and financial reaso...

  13. Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains

    cs.CL 2025-05 reject novelty 5.0 of 10

    METEORA uses DPO-tuned rationales to select and verify evidence chunks in RAG, and claims better recall, precision, evidence efficiency, and poisoning defense, though key evaluation details are missing.

  14. SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    SciCUEval provides a multi-domain, multi-modality benchmark for LLM scientific context understanding, and evaluation results show reasoning-augmented models outperform specialized scientific models, but the dataset is...

  15. Hierarchical Reranking for Scalable Financial RAG System

    cs.IR 2026-07 reject novelty 4.0 of 10

    A finance-specific RAG pipeline combining table-to-JSON conversion, two-stage reranking, and long-context split-fusion reports NDCG@20=0.7918 and second place in the ICAIF '24 FinanceRAG challenge.

  16. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools