Pith. sign in

Paper Citation Record · LEDGER

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation

As of 10 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2505.17095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17095 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:28:22.701010Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact10
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 545e9168-9c1c-4fa1-a6d6-2160340b64ce · outbound

This paper cites online" 'onlinestring :=.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.545896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.545896Z digest=sha256:e110fa176c382c5d1c26b44b5a51ef6fce56438b6e9455f9e11ad4a6d3842e38

Observation 3e199913-0c7e-493e-b552-1f73826b8e57 · outbound

This paper cites write newline.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.676307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.676307Z digest=sha256:7ca10ab97750b9743875b36279f224789acb16a6cbbac0f560d8feef1076d507

Observation 6f031bff-9e47-4538-8380-c782f87652aa · outbound

This paper cites Non-Determinism of "Deterministic" LLM Settings.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Non-Determinism of "Deterministic" LLM Settings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:18.784104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:18.784104Z digest=sha256:6f1606963491dd920941d63303734df679a5c2cdef08633a6a90cfd014aa459e

Observation cb749c51-d812-4432-8d9b-66baad453fd4 · outbound

This paper cites Rahmani, and Youlin Li.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Rahmani, and Youlin Li

Reference 4

Resolution
verified exact
doi, observed 2026-08-07T15:28:26.006630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:18.916189Z digest=sha256:021b91689c620edfea9ef1afa561aa5986374c31f606d38ced3acc96acc30017

Observation 35ad88a8-a0a5-4022-af59-67259a56e4ab · outbound

This paper cites Sebire, Saleh Khalil, Elham Asgari, Christopher Tan, Andrew Taylor, and Dominic Pimenta.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Sebire, Saleh Khalil, Elham Asgari, Christopher Tan, Andrew Taylor, and Dominic Pimenta

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.025921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.025921Z digest=sha256:f8fc76bb688ba5ce442eab26c1f3823a47c99a678deecb84df02f5c48b355e26

Observation f0354ff8-c2e3-4b61-9a13-cabf543b041e · outbound

This paper cites Prompt Stability Scoring for Text Annotation with Large Language Models.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Prompt Stability Scoring for Text Annotation with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.116860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.116860Z digest=sha256:f32eaa92eb90b10cc710f7ba7b9c46513797609bfed986bb584db703ef0d69a9

Observation 03e9cc4a-01bd-4cd9-973f-6860447e380e · outbound

This paper cites Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.270196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.270196Z digest=sha256:5142ddc58dc7dd3c1d174386ef134f1f2c3f2c9f5a1be4030cdeef16a9c95367

Observation b5c38088-02c5-4497-a132-c5fa265f4781 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.789786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.350669Z digest=sha256:25b90db1e3d763ae265a7f4981cd00f2641e445367cd4bfc09a9e7f63ca6aa31

Observation 4762bb75-b517-48c9-a52f-b6a177d21600 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.419917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.419917Z digest=sha256:d77e32a1531c03e50cfea35507f45a6b418d3b8846a7b01a1b0f15d058f19d45

Observation 016cd5c8-f737-4163-8f59-2481fb8ccbc1 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:19.474288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:19.474288Z digest=sha256:2bd0041acc592dfcc61728d4efa438a86d94d986677c6703909662d158788570

Observation db66a195-658c-4d4d-ad70-a76d807b50de · outbound

This paper cites Nambudiri.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Nambudiri

Reference 11

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.516563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.553401Z digest=sha256:6f580a3d0e10753dc95360c55218e638a53d6ad69aefaef060e889ce3f9735c6

Observation d218f442-d4aa-4f91-8526-093e25e86295 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T15:28:25.211566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.699109Z digest=sha256:c4dab8168862e60b6058302f6c3783f61c2b8ff17728019a5f6465c1cd4c432d

Observation 8d3c30a9-5ce3-4dfe-aea0-71832e21c7cd · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:24.911183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.835466Z digest=sha256:c2b986fe822feb2bc1caba43033a015c01b754491eaeb4aa012c413dd0438e24

Observation b4f155b0-a035-4adf-bc0b-acfbe676b719 · outbound

This paper cites Gold, and Vishnu Mohan.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Gold, and Vishnu Mohan

Reference 14

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:24.646083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:19.997095Z digest=sha256:8023b6df57d110a7e020feefe728b8e250ae0083d90fd8588c992dd2a8cbb74f

Observation 0b00c702-a108-430f-b331-d05736248e93 · outbound

This paper cites Akdogan, Jessica Atkins, Mohamed B.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Akdogan, Jessica Atkins, Mohamed B

Reference 15

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:28:26.928726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.118227Z digest=sha256:5daff8cada7b67aa0315f45c2df5ce28bf2b151a86302fc1fe3eece381efcb3d

Observation 360da593-0879-41ae-b68d-d180032a3f19 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.709683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.257178Z digest=sha256:9867e351e08e344b83622c394089688746e9e708d229f0124db2e9cab4663c1a

Observation 471f0add-3d1d-4d98-b9e8-5ada45aac9cc · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 17

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T15:28:26.715873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.356613Z digest=sha256:1332ca78d3132b515e8fcf1d519647e51e5da516adf7da65c4c9d1d0a2b71e16

Observation 66b77b6c-3cb5-46a4-bb30-34e4f5e0ae11 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-07T15:28:26.543053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.457250Z digest=sha256:f9a6975caab2c7410cf92d36b72c8623d405f7f7dafcb6e3df2fdda4f4abaffa

Observation f0bda5ea-ae83-4587-9ad0-18d2cf1f5e63 · outbound

This paper cites McCoy, Faye Yu Ci Ng, Christopher M.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation McCoy, Faye Yu Ci Ng, Christopher M

Reference 19

Resolution
verified exact
doi, observed 2026-08-07T15:28:24.346060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.573090Z digest=sha256:a3f6efcc1f6d76a79aa0828090333db67d8c36a87d6d8c677df59f448c5228f1

Observation 9c0539b9-67c4-4346-b6ec-8b8f6dc119f9 · outbound

This paper cites Pennathur.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Pennathur

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T15:28:24.065742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.719002Z digest=sha256:542840ab689113154457abfcec1c6da6dfebaf99bb7c6fb653f7900899a3c667

Observation 7076c52c-d4d9-4717-86ae-8babcc286d6f · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:20.840862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:20.840862Z digest=sha256:375dfece5bf1f2800c7ecb16d591a5ed6e4ac5eb7fa5f708a2785ad6983ebc04

Observation 3dd95755-3335-47c8-90c9-2a4a316c76e1 · outbound

This paper cites Quiroz, Liliana Laranjo, Ahmet Baki Kocaballi, Shlomo Berkovsky, Dana Rezazadegan, and Enrico Coiera.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Quiroz, Liliana Laranjo, Ahmet Baki Kocaballi, Shlomo Berkovsky, Dana Rezazadegan, and Enrico Coiera

Reference 22

Resolution
verified exact
doi, observed 2026-08-07T15:28:23.800870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:20.944217Z digest=sha256:062ecbf352d2d341e4c9654bc18eacea34a879bac7e35741f361bcd2c6c70577

Observation dcd1a8d0-97bd-4d35-a5ad-19a631f90eba · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.045394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.045394Z digest=sha256:c92840329933ba6330b4419d57e30f7f21dbe059c06cd5ad299735f9a756101d

Observation 8bcd07dc-8050-4cc9-8e32-0f41c6a7c13a · outbound

This paper cites Evaluating Consistency and Reasoning Capabilities of Large Language Models.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Evaluating Consistency and Reasoning Capabilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.131043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.131043Z digest=sha256:d90ff572f2879bee7ac5e16f3643b5172a5763204d3847218ba2c9bf437a7426

Observation a08155d6-4545-4c81-90bb-908ddb7edd6c · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.571936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.251803Z digest=sha256:4f1b9fd8a1dc7fde59b338186843cfd4dd20c47821f39ba3a8422dc2f526a4ae

Observation a31d8ed8-6421-4745-91a0-bf2c0f81b4a1 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 26

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:23.505634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.397661Z digest=sha256:d14679fd476b455d74e7c127e1eba3bf80222cd9d7028243d9918d5da0f85307

Observation 43ed2592-f723-4b4e-8681-1e7894b5f171 · outbound

This paper cites Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:28:26.272163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.525302Z digest=sha256:8562dfadd3a0f0a63e66e6b110582e8141d26fc3f53c7e6cbac7ef2a510ae22d

Observation a6267c53-5da6-4ecf-a5a7-6a05736bf0a9 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:28:21.711043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.711043Z digest=sha256:dd7ce807202f22502ad12a2cf96082824324254f3ed72a9a075d95523b1f7a69

Observation 52e6195b-9c5f-4d53-82b7-fd7a8ccc746e · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:21.819970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:21.819970Z digest=sha256:1996116faa38ab0147a79cb808c57dce69a5c54b1b9d103bcb585cd40e8732c3

Observation 524d9d08-9753-4281-bfed-e5bb647aad90 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 30

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T15:28:23.238463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:21.958266Z digest=sha256:6828a8b9fb6ea822b79f9a3e4030e2bf0f2b109a3078513cb120728115c8ca4d

Observation 3359e3e9-8a9c-4779-8484-ce2658534173 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:22.130493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:22.130493Z digest=sha256:30c3fcfecbdad5a4f4f03c2d85b61195ea98ad42cfa7a7058e5832353c70d98f

Observation 224e5fb2-952e-4f34-9811-f7240bb54e4a · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.413515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.229663Z digest=sha256:11f622446fa94fe7cafcfdad1947ef177d5115d1f3d7ef61c5715a02fcda04cd

Observation 2dd1728c-1d31-487b-b5fc-577f9164bcd2 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation BERTScore: Evaluating Text Generation with BERT

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:22.370716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:28:22.370716Z digest=sha256:7757c34fbc5d7984315709d677486ee3deafc573cbe7c50b5d1117eb7dbfe278

Observation 37bc608e-0626-4d53-ba53-32c804110b48 · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 34

Resolution
verified exact
doi, observed 2026-08-07T15:28:22.995871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.540412Z digest=sha256:6371a437350a4bd054808ec4bfa4d572159a5717a3eaeb99024c1252818c5be8

Observation 42334b87-d979-4a62-9709-4d43b0c01f1d · outbound

This paper cites an unresolved cited work.

Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:28:27.263988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T15:28:22.701010Z digest=sha256:16e3b46ffbfba6c1f27a2a1ad7cee7b1b03098e1f08e2420703bf93e875f57ff

Pith citing papers

No inbound Pith citation observations are available.