Pith. sign in

REVIEW 3 cited by

FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.15800 v3 pith:M5WNUHHN submitted 2025-04-22 cs.IR

classification cs.IR
keywords financialdomainfindermodelsinformationrealisticbenchmarkcontexts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the fast-paced financial domain, accurate and up-to-date information is critical to addressing ever-evolving market conditions. Retrieving this information correctly is essential in financial Question-Answering (QA), since many language models struggle with factual accuracy in this domain. We present FinDER, an expert-generated dataset tailored for Retrieval-Augmented Generation (RAG) in finance. Unlike existing QA datasets that provide predefined contexts and rely on relatively clear and straightforward queries, FinDER focuses on annotating search-relevant evidence by domain experts, offering 5,703 query-evidence-answer triplets derived from real-world financial inquiries. These queries frequently include abbreviations, acronyms, and concise expressions, capturing the brevity and ambiguity common in the realistic search behavior of professionals. By challenging models to retrieve relevant information from large corpora rather than relying on readily determined contexts, FinDER offers a more realistic benchmark for evaluating RAG systems. We further present a comprehensive evaluation of multiple state-of-the-art retrieval models and Large Language Models, showcasing challenges derived from a realistic benchmark to drive future research on truthful and precise RAG in the financial domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

    cs.AI 2026-08 conditional novelty 7.0 of 10

    A new SEC-filing QA benchmark shows that retrieval models lose 13 to 20.5 points of ranking accuracy when plausible but wrong disclosures are used as distractors.

  2. FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering

    cs.IR 2026-07 conditional novelty 6.0 of 10

    FinSAgent improves financial filing QA by conditioning sub-queries on a summary of the local corpus and gating semantic reranking with a learned validity signal, beating baseline systems on five benchmarks.

  3. Enhancing Document-Level Question Answering via Multi-Hop Retrieval-Augmented Generation with LLaMA 3

    cs.CL 2025-06 reject novelty 2.0 of 10

    A standard RAG pipeline with an undefined multi-hop module is reported to outperform baselines on financial QA datasets, without code or data.

Pith tools