Pith. sign in

Paper Citation Record · LEDGER

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.20491.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.20491 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:50:22.254315Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42904533-c7aa-4ee7-9909-53bf29b36b24 · outbound

This paper cites Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-02T11:54:15.471726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-02T11:50:20.579924Z digest=sha256:3d510b89ca48dc2a5ce99fa2ef93946e8ecf27ac45887c7184bd6c9d6313bf95

Observation c66a9a79-0f12-4358-a7dd-6ab01e355df9 · outbound

This paper cites Revised guidance on model risk management: Supervisory letter SR 26-2.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Revised guidance on model risk management: Supervisory letter SR 26-2

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:20.651970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:20.651970Z digest=sha256:67d240d95193b2498433f653f4fd877134ed0d8ae91609911b17f670b09045f2

Observation 56071d13-9f62-4da9-8248-af9d8e5c6bd6 · outbound

This paper cites FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-02T11:54:15.255546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-02T11:50:20.718612Z digest=sha256:13ea14b2dbed78cb0f7982d02a0ee455201fa5976f61c6ababffeaf2009bec29

Observation da51d50a-1d4f-46b1-9cc1-766d597f3eec · outbound

This paper cites Visibility into AI agents.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Visibility into AI agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:20.817091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:20.817091Z digest=sha256:817edbc9ef2ca3b01a43a02d12c236e3ba7b15e681355c0df4b46a2527bda510

Observation 9daec3ae-59d2-4c63-a3bf-57d214f3a59c · outbound

This paper cites TRAJECT-Bench : A trajectory-aware benchmark for evaluating agentic tool use.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making TRAJECT-Bench : A trajectory-aware benchmark for evaluating agentic tool use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.037586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.037586Z digest=sha256:2ec148c962867ae8081c86ae30bc6fd2d5f70f1f0e9c88e2ba368238f72ff720

Observation 8d78d997-5196-4dfa-932d-77e7a44461ad · outbound

This paper cites AI safety best practices for regulated environments.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making AI safety best practices for regulated environments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.144230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.144230Z digest=sha256:30dea951c426ecc35335d6da466d09fab7b1361a6259a35c202e6ba4d50e6223

Observation 7be5c78c-c390-4332-bb34-795808e262a4 · outbound

This paper cites A review of evaluation metrics for text similarity.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making A review of evaluation metrics for text similarity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.216412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.216412Z digest=sha256:26e9dc8ded2c807028bfacfc8f6e881b5264ea705739f0718d4962083254ff17

Observation 92138f80-1996-4ac2-a18e-59d2d09a959f · outbound

This paper cites ToolSandbox : A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making ToolSandbox : A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.298577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.298577Z digest=sha256:24b9d254eb3df2944bb5a566434aab1f5316a1d8d1b78e05c73fa0ab77e17b27

Observation 75723614-9fc7-4260-b0ba-3067f4de8771 · outbound

This paper cites ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.375965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.375965Z digest=sha256:5c855bd5b78b3db666f015efbc357bee0a447988e90b60a0384e5f719ab14ec1

Observation b0aabff8-f433-4208-923d-50d55c5a64c4 · outbound

This paper cites FinResearchBench : A logic tree based agent-as-a-judge evaluation framework for financial research agents.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making FinResearchBench : A logic tree based agent-as-a-judge evaluation framework for financial research agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.492854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.492854Z digest=sha256:b173402b22246c4186bf8dad2bd440a4cc67b2f7ea280c007c84c5b272896315

Observation bcd03b1b-b4ae-455b-a469-d68db2085fd6 · outbound

This paper cites Scalable runtime governance for agentic AI in financial services.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Scalable runtime governance for agentic AI in financial services

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.661266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.661266Z digest=sha256:c0d299cc9009e26532e15e9a086876c765c2904a1e08dc2d8dc934d6988d1f26

Observation d781a69a-bd23-4d97-b24f-630a92ca4bd9 · outbound

This paper cites Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.831933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.831933Z digest=sha256:07832362ddc11dec436ea4e7beaf82a290442d7d77877c5fec985c8faba33a61

Observation a8be4c58-d500-4765-9c72-9840ce4f7a08 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Self-consistency improves chain of thought reasoning in language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.988653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.988653Z digest=sha256:006108b662e88b98775204b2b601af61a4cdaa212b51878b7ef7bf368e77ce0d

Observation 8c3c2a93-36e7-4936-9ca1-6d6aa44ccd25 · outbound

This paper cites FinBen : A holistic financial benchmark for large language models.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making FinBen : A holistic financial benchmark for large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:22.099780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:22.099780Z digest=sha256:380f0b52d5cec62b913bfa8499c4af2f3a87faca121a2279ce010e02ac4779df

Observation afc9c533-1065-4080-bed3-7488956c3dc3 · outbound

This paper cites -bench : A benchmark for tool-agent-user interaction in real-world domains.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making -bench : A benchmark for tool-agent-user interaction in real-world domains

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:22.254315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:22.254315Z digest=sha256:2810ee272876a4c16671b5e5685606e888f93246a9ac737d6c9afba5c7886961

Pith citing papers

No inbound Pith citation observations are available.