Pith. sign in

REVIEW 2 cited by

SECQUE: A Benchmark for Evaluating Real-World Financial Analysis Capabilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.04596 v1 pith:55UUBMOG submitted 2025-04-06 cs.AI cs.CEcs.CL

classification cs.AIcs.CEcs.CL
keywords analysisfinancialsecquebenchmarkevaluatingmodelsperformanceacross
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce SECQUE, a comprehensive benchmark for evaluating large language models (LLMs) in financial analysis tasks. SECQUE comprises 565 expert-written questions covering SEC filings analysis across four key categories: comparison analysis, ratio calculation, risk assessment, and financial insight generation. To assess model performance, we develop SECQUE-Judge, an evaluation mechanism leveraging multiple LLM-based judges, which demonstrates strong alignment with human evaluations. Additionally, we provide an extensive analysis of various models' performance on our benchmark. By making SECQUE publicly available, we aim to facilitate further research and advancements in financial AI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering

    cs.IR 2026-07 conditional novelty 6.0 of 10

    FinSAgent improves financial filing QA by conditioning sub-queries on a summary of the local corpus and gating semantic reranking with a learned validity signal, beating baseline systems on five benchmarks.

  2. FinMaster: A Holistic Benchmark for Mastering Full-Pipeline Financial Workflows with LLMs

    cs.AI 2025-05 conditional novelty 6.0 of 10

    FinMaster introduces a simulator-driven benchmark with 183 financial tasks and finds LLM accuracy collapses from about 96% on basic literacy to below 40% on multi-step accounting, auditing, and consulting workflows.

Pith tools