REVIEW 16 cited by
FinQA: A Dataset of Numerical Reasoning over Financial Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The sheer volume of financial statements makes it difficult for humans to access and analyze a business's financials. Robust numerical reasoning likewise faces unique challenges in this domain. In this work, we focus on answering deep questions over financial data, aiming to automate the analysis of a large corpus of financial documents. In contrast to existing tasks on general domain, the finance domain includes complex numerical reasoning and understanding of heterogeneous representations. To facilitate analytical progress, we propose a new large-scale dataset, FinQA, with Question-Answering pairs over Financial reports, written by financial experts. We also annotate the gold reasoning programs to ensure full explainability. We further introduce baselines and conduct comprehensive experiments in our dataset. The results demonstrate that popular, large, pre-trained models fall far short of expert humans in acquiring finance knowledge and in complex multi-step numerical reasoning on that knowledge. Our dataset -- the first of its kind -- should therefore enable significant, new community research into complex application domains. The dataset and code are publicly available\url{https://github.com/czyssrs/FinQA}.
Forward citations
Cited by 16 Pith papers
-
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
On a new 615-question business-case benchmark graded by AI against instructor rubrics, frontier LLMs score 87-88% partial credit but complete only about half the questions.
-
On the Fitness Landscape in the $NK$ Model
For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.
-
Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable
CM-LRS is a workflow-output-layer reliability score for capital-markets LLM outputs; on five public-document workflows, frontier closed-source models score 4.09–4.31 and Llama 3.3 70B scores 3.15 under four LLM judges.
-
FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
FinGAIA is a 407-task Chinese financial agent benchmark where the best agent, ChatGPT DeepResearch, scores 48.9%, far below financial experts at 84.7%.
-
CFBenchmark-MM: Chinese Financial Assistant Benchmark for Multimodal Large Language Model
A 9,356-pair Chinese multimodal financial benchmark reveals that state-of-the-art multimodal LLMs, including GPT-4V, still score below 53% on objective and 39% on subjective financial chart tasks.
-
Human-Centric Evaluation for Foundation Models
In free-form research collaborations rated by humans, Grok 3 scores highest, followed by DeepSeek R1 and Gemini 2.5; OpenAI o3 mini trails.
-
FinS-Pilot: A Benchmark for Online Financial RAG System
FinS-Pilot is a small benchmark of real-world financial assistant queries with real-time API data and text corpus, used to compare Chinese LLMs on financial RAG tasks.
-
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
A new French financial document benchmark shows vision-language models are strong at text/table extraction but brittle on charts and multi-turn dialogue, with accuracy converging near 50% in conversational settings.
-
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
VisFinEval is a 15,848-question Chinese multimodal financial benchmark covering eight image types across three workflow depths, on which the best model still trails finance experts by more than 14 points.
-
CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
A fine-tuned Llama 3 model with a trained document critic and program-based reasoning beats GPT-4o and other baselines on a new carbon footprint QA benchmark.
-
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.
-
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
A plug-and-play external reward model with speculative rejection sampling cuts tree-search cost for LLM decision-making to about 1/10 while keeping or slightly improving accuracy on math, planning, and financial reaso...
-
Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
METEORA uses DPO-tuned rationales to select and verify evidence chunks in RAG, and claims better recall, precision, evidence efficiency, and poisoning defense, though key evaluation details are missing.
-
SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
SciCUEval provides a multi-domain, multi-modality benchmark for LLM scientific context understanding, and evaluation results show reasoning-augmented models outperform specialized scientific models, but the dataset is...
-
Hierarchical Reranking for Scalable Financial RAG System
A finance-specific RAG pipeline combining table-to-JSON conversion, two-stage reranking, and long-context split-fusion reports NDCG@20=0.7918 and second place in the ICAIF '24 FinanceRAG challenge.
-
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.
Discussion (0). Sign in to comment.