Pith. sign in

Paper Citation Record · LEDGER

Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.16974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.16974 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:39.064313Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:56:27.114999Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation afce4c83-1dbc-448d-82ba-b27682fe53ec · inbound

EasyMath: A 0-shot Math Benchmark for SLMs cites this paper.

EasyMath: A 0-shot Math Benchmark for SLMs Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:39.064313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:33:39.064313Z digest=sha256:ab6b646626e7a2fde024383a85d170d8d1dafdec8b65abd8f05dc8272568cd19

Observation eda8e380-146e-4921-ac4f-6308acb1848a · inbound

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning cites this paper.

The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:58:33.525577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:58:33.219451Z digest=sha256:299d8dc63f462125b1bc115d37bb95c3a38287a53ca86105b7c32f4f84f189c6

Observation d42194f6-6da4-4a9b-bc1d-a72b0c8f9db0 · inbound

Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models cites this paper.

Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:47.834129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:47.834129Z digest=sha256:f64f5e234e917948ac10f42f38615038bc0de3bc88280691fa64b4efa5e35b32

Observation 13d0e133-3bb0-4eb9-8de0-dcf7bb612298 · inbound

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis cites this paper.

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:49.119437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:49.119437Z digest=sha256:3c53ecff1564c9071b14b29aaea3c1ce8ee75f6bbbed706b4186f4ce38a1ba1c

Observation 7525b13b-a51e-4a14-9485-0428f05dc043 · inbound

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis cites this paper.

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T06:19:09.680553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:19:09.680553Z digest=sha256:05e712e33ac0a50534f3e7317ba497f90a2844a115f5d591732096331bcb6e20

Observation 58b21271-5c75-4a4a-8ea2-59011ee69658 · inbound

Towards a Science of AI Agent Reliability cites this paper.

Towards a Science of AI Agent Reliability Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T22:32:02.023467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:32:02.023467Z digest=sha256:b31eed2d324d4480edd3fd36616d12bb292fd8d34197f60ede36373b92857852

Observation 5bf674ad-7849-4264-a9ca-469bcab8a4e1 · inbound

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models cites this paper.

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.402158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:31:46.940449Z digest=sha256:d344e1341b429051f91bea22dc42c3125ef1bd7cb05b546fd546fc505d8f14a0

Observation 6b05796b-0d86-42d7-8102-e322530cb9f5 · inbound

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis cites this paper.

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T20:50:53.567415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:50:53.567415Z digest=sha256:07341a223f4d8419b02e42cc36d5eaf50c205c5f8a896b01f814c2bd71d635d6

Observation 6d4a2621-690f-4808-ae15-0689c3f5d5c0 · inbound

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems cites this paper.

From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:45.811279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:27:30.692356Z digest=sha256:0649a8c5c9f1fb986212fb5521559eb9dab34888623845f5b7c9f9840201f49d

Observation 6a8b0efe-dbe4-48e5-94be-133175b7d75a · inbound

Shapley in Context: Explaining Financial Language with Domain Expertise cites this paper.

Shapley in Context: Explaining Financial Language with Domain Expertise Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 284

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.118015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T01:46:33.186116Z digest=sha256:4a856d04df4325ebdfd04aa49cf8cde5e80d7d17495891fc63f9286a53b33105

Observation d781a69a-bd23-4d97-b24f-630a92ca4bd9 · inbound

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making cites this paper.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T11:50:21.831933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:50:21.831933Z digest=sha256:07832362ddc11dec436ea4e7beaf82a290442d7d77877c5fec985c8faba33a61

Observation f0e011b8-26aa-419b-9208-9ca2b9a07db7 · inbound

Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement cites this paper.

Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:15:36.327966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:15:36.327966Z digest=sha256:b9882e3233e7d289b2427065a061ffa09077999df3166e5beb6a8aa6a3a3aa22