Pith. sign in

Paper Citation Record · LEDGER

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation

As of 18 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2505.17095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17095 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:22.701010Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact10
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 545e9168-9c1c-4fa1-a6d6-2160340b64ce · outbound

This paper cites online" 'onlinestring :=.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.545896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.545896Z digest=sha256:5e261a47867a04e1b87ac3f5456050293f8cd9f7e81bca7430c41e2c69a1f8fe

Observation 3e199913-0c7e-493e-b552-1f73826b8e57 · outbound

This paper cites write newline.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.676307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.676307Z digest=sha256:c7de02d38132b84d5af04d17ef52021deec47be4d32d7f590d4c76c5149eaa73

Observation 6f031bff-9e47-4538-8380-c782f87652aa · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Non-Determinism of "Deterministic" LLM Settings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.784104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.784104Z digest=sha256:207b05b10f20fce5c4b87b450b4f58869d81202197b7fee6020240795cc9bb1b

Observation cb749c51-d812-4432-8d9b-66baad453fd4 · outbound

This paper cites Rahmani, and Youlin Li.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Rahmani, and Youlin Li

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T15:28:26.006630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:18.916189Z digest=sha256:36d499721e29eaa1969c3a38d7497bed4f05e007aaed25a787b476a5bfc6bd83

Observation 35ad88a8-a0a5-4022-af59-67259a56e4ab · outbound

This paper cites Sebire, Saleh Khalil, Elham Asgari, Christopher Tan, Andrew Taylor, and Dominic Pimenta.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Sebire, Saleh Khalil, Elham Asgari, Christopher Tan, Andrew Taylor, and Dominic Pimenta

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.025921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.025921Z digest=sha256:99396754755dc599a639b370a94b9973b717e0b8f9a731f28e7a3488cd1cecb9

Observation f0354ff8-c2e3-4b61-9a13-cabf543b041e · outbound

This paper cites Prompt Stability Scoring for Text Annotation with Large Language Models.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Prompt Stability Scoring for Text Annotation with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.116860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.116860Z digest=sha256:c534ef147448b50a8719097eb30075071dc2ff4c3cbde968ad8b55f2be08f70c

Observation 03e9cc4a-01bd-4cd9-973f-6860447e380e · outbound

This paper cites Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.270196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.270196Z digest=sha256:e3ed490c5395222c4c5e7b26c4da54313d124de3d1f02d1424c773c3faf88020

Observation b5c38088-02c5-4497-a132-c5fa265f4781 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.789786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.350669Z digest=sha256:6d873c04948d3f53e1d40a8b68a9b7899c053d0b23d70c053ba3f51fa11bec27

Observation 4762bb75-b517-48c9-a52f-b6a177d21600 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.419917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.419917Z digest=sha256:af77b8c09535786d0473fc2f08ef266d6e7b7a4695caa557075804b77998170f

Observation 016cd5c8-f737-4163-8f59-2481fb8ccbc1 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.474288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.474288Z digest=sha256:f1abb510d4bfc0e3558385229005a4c2e800e893c8272e068146aa9456c19f24

Observation db66a195-658c-4d4d-ad70-a76d807b50de · outbound

This paper cites Nambudiri.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Nambudiri

Reference 11

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.516563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.553401Z digest=sha256:3dff8b05f096ffd2b4d04eb5bf80dca22b8056f0fe4fb70bab167a2388fe7a5d

Observation d218f442-d4aa-4f91-8526-093e25e86295 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.211566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.699109Z digest=sha256:48d41f979c27781d441aafa8e2212ebaa76ea133f58588bb4e0e5c95157a9d91

Observation 8d3c30a9-5ce3-4dfe-aea0-71832e21c7cd · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:24.911183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.835466Z digest=sha256:1a893041cefec15646c088ca84010c9cfe4d2ade26a07c2a9c32461925574776

Observation b4f155b0-a035-4adf-bc0b-acfbe676b719 · outbound

This paper cites Gold, and Vishnu Mohan.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Gold, and Vishnu Mohan

Reference 14

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:24.646083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.997095Z digest=sha256:dfcecad758cfaef5dd02b87d57cd70cabb136c379438d15ffcbd0adf150901eb

Observation 0b00c702-a108-430f-b331-d05736248e93 · outbound

This paper cites Akdogan, Jessica Atkins, Mohamed B.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Akdogan, Jessica Atkins, Mohamed B

Reference 15

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:28:26.928726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.118227Z digest=sha256:43e97351d345d9f036215314ae5cf9b9477018547281119cc66cef2ef79341d2

Observation 360da593-0879-41ae-b68d-d180032a3f19 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.709683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.257178Z digest=sha256:084f4c55048f8a2d2495dacbce47696fefd98811fe9f19019f41e3ec12f8eeb8

Observation 471f0add-3d1d-4d98-b9e8-5ada45aac9cc · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 17

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:28:26.715873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.356613Z digest=sha256:576c025bb6808be8e5341d226e64e8aa8dc55aaf92b99debc873f3a043b9d4c5

Observation 66b77b6c-3cb5-46a4-bb30-34e4f5e0ae11 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:28:26.543053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.457250Z digest=sha256:b7d753915dce42e2acb1ee16b37198e937e9cfcfa02b56b84f7ea6d525de58ee

Observation f0bda5ea-ae83-4587-9ad0-18d2cf1f5e63 · outbound

This paper cites McCoy, Faye Yu Ci Ng, Christopher M.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation McCoy, Faye Yu Ci Ng, Christopher M

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T15:28:24.346060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.573090Z digest=sha256:d26ed58f7f1661c164e29d9da63b2d60d394eb1fb2cdcb3ff8b393cffbd050a6

Observation 9c0539b9-67c4-4346-b6ec-8b8f6dc119f9 · outbound

This paper cites Pennathur.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Pennathur

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T15:28:24.065742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.719002Z digest=sha256:75d291eb2a1f1aa7d500d877ff3962e2c358d0ceb1ceb63ae02954c8b261d873

Observation 7076c52c-d4d9-4717-86ae-8babcc286d6f · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:20.840862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:20.840862Z digest=sha256:ade23bb2e3c52f10f2581ee34631534264e8c77148261c21c63cb2893b72ea06

Observation 3dd95755-3335-47c8-90c9-2a4a316c76e1 · outbound

This paper cites Quiroz, Liliana Laranjo, Ahmet Baki Kocaballi, Shlomo Berkovsky, Dana Rezazadegan, and Enrico Coiera.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Quiroz, Liliana Laranjo, Ahmet Baki Kocaballi, Shlomo Berkovsky, Dana Rezazadegan, and Enrico Coiera

Reference 22

Resolution
verified exact
doi, observed 2026-08-07T15:28:23.800870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.944217Z digest=sha256:80b9a39e82608fb953639a61372ef78e1921b8c09504f7b389abd53fa09d949f

Observation dcd1a8d0-97bd-4d35-a5ad-19a631f90eba · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.045394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.045394Z digest=sha256:abd3542847064b5551f0cae79b477ec6a7fd77a3f5b7977d837a20eee4cef73d

Observation 8bcd07dc-8050-4cc9-8e32-0f41c6a7c13a · outbound

This paper cites Evaluating Consistency and Reasoning Capabilities of Large Language Models.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Evaluating Consistency and Reasoning Capabilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.131043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.131043Z digest=sha256:a9bf660c588ba80570ab7a3794c510ad982cc2d4f1ee7136ef5dca374753fb23

Observation a08155d6-4545-4c81-90bb-908ddb7edd6c · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.571936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.251803Z digest=sha256:b4b06773d1642fa14c201d0e886fb438dd3e717e969281dd423457c4c302ac44

Observation a31d8ed8-6421-4745-91a0-bf2c0f81b4a1 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 26

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:23.505634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.397661Z digest=sha256:4e772362becb6831190d225d5cf42ef6615431f3fc8be3d4c4371e119f285daf

Observation 43ed2592-f723-4b4e-8681-1e7894b5f171 · outbound

This paper cites Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:26.272163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.525302Z digest=sha256:8fd7f31735c8ca385fbf0c4490a8050e47d03b90783dac760487e3b1dbed0216

Observation a6267c53-5da6-4ecf-a5a7-6a05736bf0a9 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:28:21.711043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.711043Z digest=sha256:38ade2ea10a6122cbbf7740b94d58c545a03a7b13e518649bff80825bb024254

Observation 52e6195b-9c5f-4d53-82b7-fd7a8ccc746e · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.819970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.819970Z digest=sha256:ac477544bf2abe336561e309a76c6bb4a074d5b7ee59f5c443f4e42bf60e49b8

Observation 524d9d08-9753-4281-bfed-e5bb647aad90 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 30

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:23.238463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.958266Z digest=sha256:5c3ffac93962398f3f22100d4b7f9f12a67677677e16ef2f9505ba06c207da4b

Observation 3359e3e9-8a9c-4779-8484-ce2658534173 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:22.130493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:22.130493Z digest=sha256:854c91d910b815d466f164501db268d68c9dd8611543fe7e58cd5c4893260398

Observation 224e5fb2-952e-4f34-9811-f7240bb54e4a · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.413515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.229663Z digest=sha256:a2e545c7a838ac596de2f1e2969f77d35e75381d75fa913d2c185a7dda16c2eb

Observation 2dd1728c-1d31-487b-b5fc-577f9164bcd2 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation BERTScore: Evaluating Text Generation with BERT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:22.370716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:22.370716Z digest=sha256:12de5259f9ecdf15532ca913f36bb4bddfaaeacf86616b9939f0df848d20c428

Observation 37bc608e-0626-4d53-ba53-32c804110b48 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 34

Resolution
verified exact
doi, observed 2026-08-07T15:28:22.995871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.540412Z digest=sha256:7635cb2394927aadbe170a8ecd3b7a9ee27069345d52af45e986d9027199b272

Observation 42334b87-d979-4a62-9709-4d43b0c01f1d · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.263988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.701010Z digest=sha256:726d8689005c5892e3b3fe8191a59c706daa2228ae9b9cf426aace778fdeacea

Pith citing papers

No inbound Pith citation observations are available.