Pith. sign in

Paper Citation Record · LEDGER

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 5 inbound Pith citation observations for arXiv:2501.18062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18062 v1

Coverage vector

measured 5 of 5 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:53:29.689406Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:16.726294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:56:25.216379Z

Reference resolution

5 of 5 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e72439a9-86fe-476f-aab3-5a609cab0709 · outbound

This paper cites FinQA: A Dataset of Numerical Reasoning over Financial Data.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models FinQA: A Dataset of Numerical Reasoning over Financial Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.668745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.668745Z digest=sha256:b70f3af468849128065cd47dc0e641bc1e27c6a4bfadf19162d3d0eb94a951b8

Observation 474df27b-e50d-4e34-bc2a-f91db76556c0 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models FinanceBench: A New Benchmark for Financial Question Answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.678903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.678903Z digest=sha256:904e51dd0ff8f15d8c5066e431d0dd4104d5db24db694817e895e4d0ee49b9f4

Observation 2a1c8f9d-32ed-492d-a6d9-ba31d633346c · outbound

This paper cites A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.689406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.689406Z digest=sha256:a2a5240f47494d0da115e247909fbfaf4dde12f237be3be77f438ef4610a1bf4

Observation 8a430616-9e27-4c79-a038-68fdfff5f118 · outbound

This paper cites LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.674023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.674023Z digest=sha256:e7f89cc1bbc83f1d29941cd0f82452b957f92095d55c5c8804a7faa121877edf

Observation 7a0ac1f2-210a-47c6-9806-17bf56a4be43 · outbound

This paper cites On Leakage of Code Generation Evaluation Datasets.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models On Leakage of Code Generation Evaluation Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.683526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.683526Z digest=sha256:8b860a2d82662ccde46b9d59dbd88575f1076037ef724116610751fe747db7ed

Pith citing papers

Observation b5cbb368-61f5-41b7-a401-2147d79c0faa · inbound

On Path to Multimodal Historical Reasoning: HistBench and HistAgent cites this paper.

On Path to Multimodal Historical Reasoning: HistBench and HistAgent FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:16.726294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:16.726294Z digest=sha256:0454c03eba661c7e88514e9a184de04c857324c754da2965c9a6bc08ae69a971

Observation 2fbc8455-1593-4cd7-8fff-6dce4cc607e4 · inbound

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents cites this paper.

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:51.605474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:59:59.383781Z digest=sha256:aa47aab091e3df1c8411385d2e48bfd7c0a94d0363912d9f858591df260e90c7

Observation 12905610-f501-471c-8f08-6dbd8eade104 · inbound

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications cites this paper.

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.234156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T05:52:40.026883Z digest=sha256:05a91ddd4eae35cc7080910ea18612cbfc07ad11c8bf84f708fb3118bcba88f5

Observation 13123e87-942a-4f52-a066-c3c8fae96dff · inbound

LATTICE: Evaluating Decision Support Utility of Crypto Agents cites this paper.

LATTICE: Evaluating Decision Support Utility of Crypto Agents FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:56:25.221197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-07T13:30:46.523784Z digest=sha256:37132dac381112c8607a3088c8e4f76d64e7565e61de3f5d30650b122ade8026

Observation 63f6b8a0-ee55-448e-b67a-77b42fa61a60 · inbound

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain cites this paper.

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:16:15.789703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-12T02:15:53.024591Z digest=sha256:9a09540ccf5e487ccb733ab85c206f608d58460710637b0c3a8cb36431c8c72d