Pith. sign in

Paper Citation Record · LEDGER

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

As of 17 August 2026, this Paper Citation Record lists 100 of 151 outbound references and 2 inbound Pith citation observations for arXiv:2506.17163.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17163 v1

Coverage vector

measured 100 of 151 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:15:22.050331Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T19:02:46.991897Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 151 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9e4ffd32-3ebf-4699-982f-5fcfb0c4ded6 · outbound

This paper cites Evaluating large language models on medical evidence summarization.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating large language models on medical evidence summarization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.695152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.695152Z digest=sha256:f1746dadf4a6e7e759ce37fa2ad58efe4a6015cd4ba793cc12aa053665c1427c

Observation f4150240-e4d1-44a2-90c9-a862884d2a4d · outbound

This paper cites Adapted large language models can outperform medical experts in clinical text summarization.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Adapted large language models can outperform medical experts in clinical text summarization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.699276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.699276Z digest=sha256:1182a44376d75b28c4d65c2292f3075ee9cd30c3deed441bdc5e3cd0a58d944d

Observation bf3bde49-59c2-4ada-89d2-7e39b1f01e08 · outbound

This paper cites Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.703146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.703146Z digest=sha256:4167c455af7cab53cf70d4df2005a229ba789c97e1a67e6ad0e91990e3aa563c

Observation 2acbc252-7251-4b32-9fe8-e52d9b257a0b · outbound

This paper cites Llm-based agentic systems in medicine and healthcare.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Llm-based agentic systems in medicine and healthcare

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.706939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.706939Z digest=sha256:06162663518f920acb52fa92fb4cce476d5d64fc111946c5c4612da280df61f5

Observation 07767479-b3f5-4f23-b0fc-55911580a52b · outbound

This paper cites MedDM:LLM-executable clinical guidance tree for clinical decision-making.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making MedDM:LLM-executable clinical guidance tree for clinical decision-making

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.710633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.710633Z digest=sha256:f5c4c7d5fd8db5615952bf2f3ea6e1470c132ef811d14071fdcc89b01f3a1323

Observation f6e46a8b-8451-40f1-bb9c-4a5c6e56aae5 · outbound

This paper cites Large language models encode clinical knowledge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Large language models encode clinical knowledge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.714905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.714905Z digest=sha256:b2602219c0b4505072e66f50e474b5609ffb92b940d4520e1c0892ede1019665

Observation e9883c0e-2b68-4294-b0a1-ad3bc2e3d75a · outbound

This paper cites Toward expert-level medical question answering with large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Toward expert-level medical question answering with large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.718778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.718778Z digest=sha256:996cf80f4da49384eaca19fa2a46b35604edb250a857f9db857f159198eb069c

Observation 3032ad41-4e58-4e07-9730-0e3dda922e41 · outbound

This paper cites Luks and Zachary D.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Luks and Zachary D

Reference 8

Resolution
verified exact
raw_fallback, observed 2026-08-15T19:15:23.025016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:15:21.722269Z digest=sha256:4caebc8997333a552fb7b389e7e4469f8fa07d6e1db7d5fdb97b89fa77b1c81b

Observation 604dc004-ab51-4576-b071-c5aacd248574 · outbound

This paper cites Variability in language used on social media prior to hospital visits.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Variability in language used on social media prior to hospital visits

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.725660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.725660Z digest=sha256:d98f0901f9c4d00aa6ca64b2dddb999663dc286754197dfb6e4959889c99ffbb

Observation 3086486e-829f-4e22-a92f-4f862604e9be · outbound

This paper cites How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.728874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.728874Z digest=sha256:dc148ad281fedcbaf601c5285b0dd75815638441003c0e552101fd5640bff925

Observation f7f08073-c5a2-4ebf-810b-3166225e8530 · outbound

This paper cites Open medical llm leaderboard.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Open medical llm leaderboard

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.732634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.732634Z digest=sha256:ba9c1e205ef53c8bdf07843785eb98544b78cd7987af3e34ff2829c1198585f7

Observation 163ba107-58bb-4b94-9fc5-0faa9c41c254 · outbound

This paper cites A rapid review of gender, sex, and sexual orientation documentation in electronic health records.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A rapid review of gender, sex, and sexual orientation documentation in electronic health records

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.736183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.736183Z digest=sha256:1f8adf82e57658b0bafcc0e1c9b067b2edb5e1fcbaa82a3c1ffdf09b52e00e4e

Observation c1dc0c61-a6ca-4f36-9ccd-88d90607b332 · outbound

This paper cites Health Care Experiences of Patients with Nonbinary Gender Identities.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Health Care Experiences of Patients with Nonbinary Gender Identities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.739883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.739883Z digest=sha256:6dd1bc4f3e0f3ab711fb8185af7ce83d53890ef017ee8871eeb72ea4cc6e6a74

Observation b054846d-c701-4048-9a16-1075165d6f0e · outbound

This paper cites Hoffmann, Roger B.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Hoffmann, Roger B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.747154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.747154Z digest=sha256:acb28b03d2994f6581f1a6bb3f539cccfcb897b59ac66c6156c9d0447b05063a

Observation 1cc38915-e466-45da-89d9-9a1fa99bc006 · outbound

This paper cites an unresolved cited work.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Unresolved cited work

Reference 15

Resolution
verified exact
doi, observed 2026-08-15T19:15:22.291183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:15:21.750666Z digest=sha256:c6103c96a112373f32980599ba1bbe65bb6c50d36a1f10660d8223b9600391c5

Observation df9834ba-ed33-4765-be03-b5b8f5a6d95f · outbound

This paper cites Gender disparities in health care.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender disparities in health care

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.754161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.754161Z digest=sha256:bc1008df43b11a13a7d2771a78b564557b92da033dfe202617c43257ecbaf654

Observation b7a80d4a-d602-4a80-a0f8-ac13be4c3c0b · outbound

This paper cites Defining gender disparities in pain management.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Defining gender disparities in pain management

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.757503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.757503Z digest=sha256:17fd68fd3181b6d9ab2ae419cca44cb973503868e6038cec547837292a79e09d

Observation 80b854cd-ef13-4eb3-91cf-389d54d53754 · outbound

This paper cites Gender differences in outcomes of a multimodal pain management program.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender differences in outcomes of a multimodal pain management program

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.760789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.760789Z digest=sha256:aaf5f074e7de7d9c5126da3f15d8d2880c09fd25f2e6b2cf5e1d127fe2f5407e

Observation 9b8a5ea0-c979-47fc-b5b8-eeba7f4a7354 · outbound

This paper cites Health and healthcare disparities among us women and men at the intersection of sexual orientation and race/ethnicity: a nationally representative cross-sectional study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Health and healthcare disparities among us women and men at the intersection of sexual orientation and race/ethnicity: a nationally representative cross-sectional study

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.763778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.763778Z digest=sha256:63274818f9803b08effe6308eaefd914635d215745f00b8215fad0a6fd028ecc

Observation 3ca81320-71ba-4cbe-b4b8-4f2acbfcf452 · outbound

This paper cites Gender bias in transformers: A comprehensive review of detection and mitigation strategies.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias in transformers: A comprehensive review of detection and mitigation strategies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.767181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.767181Z digest=sha256:04075fef37b8fe09748b8020af926cd3d1fa90df9ab7d048b37861ea15357e31

Observation 3818626d-6e34-46af-80ab-5e088208544f · outbound

This paper cites Gender bias in natural language processing and computer vision: A comparative survey.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias in natural language processing and computer vision: A comparative survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.770938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.770938Z digest=sha256:d842089ffcd89106898a3bc0799145eb06c2f0967469ad7f8b2fcf29cd6b4ed4

Observation eae15c9e-3435-493d-b2cf-5b5970538e03 · outbound

This paper cites Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.774269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.774269Z digest=sha256:92b98993e689fdc3c49767451ee49835564f437111d9b2282cde9d38afe0391e

Observation 03496300-554e-48c5-8674-3c11f4013c81 · outbound

This paper cites Addressing gender-related performance disparities in neural rankers.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Addressing gender-related performance disparities in neural rankers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.778046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.778046Z digest=sha256:de1cf40d7ce31502a7012d9faf1b0fa137bbb3237d996f85f9ac090c7abc1258

Observation 5ad8a1c1-fa49-4ea2-89ad-58928594874e · outbound

This paper cites Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:15:22.897241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:15:21.781255Z digest=sha256:5ee3ee21fbfb3497e9618ed385bfb6ff198db3d99861ab6b07bc0ece4b938a1c

Observation ae8aaac9-2fa7-4bc4-b193-cb82ec9fc32a · outbound

This paper cites Bias in bios: A case study of semantic representation bias in a high-stakes setting.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bias in bios: A case study of semantic representation bias in a high-stakes setting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.784946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.784946Z digest=sha256:f5aece3c5fc929c106b19bcfc3f65a6c91987ccccc7738481cbb71317e0f3083

Observation b87c217d-1adc-4fbb-b025-c165642ae253 · outbound

This paper cites Assessing the potential of gpt-4 to perpetuate racial and gender biases in health care: a model evaluation study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Assessing the potential of gpt-4 to perpetuate racial and gender biases in health care: a model evaluation study

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:15:21.788803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.788803Z digest=sha256:6f0ec1a950a2301e60552baea33cda1cb22d4a5f1967324d1fa6f393fc12a335

Observation 8f18a0f6-7c6b-42fc-b0db-22bd40ee9df7 · outbound

This paper cites Bias patterns in the application of LLMs for clinical decision support: A comprehensive study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bias patterns in the application of LLMs for clinical decision support: A comprehensive study

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.792493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.792493Z digest=sha256:cfc79469d34a5f14e3367479fb771e1ea97a813a7d8becf2d7cef846a7fb869d

Observation a2e7c5b1-aa5d-46c0-b7c7-59e788268d01 · outbound

This paper cites An investigation into the impact of deep learning model choice on sex and race bias in cardiac mr segmentation.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making An investigation into the impact of deep learning model choice on sex and race bias in cardiac mr segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.796629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.796629Z digest=sha256:4ea77f905db9c250ddda7a8d3246b821a45e11ba64356efef846b607531065bf

Observation c5c3ee91-32a4-45ee-a99b-b1b5ed9739cf · outbound

This paper cites Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.799920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.799920Z digest=sha256:6610aade568b005c16870509168f65ab64b0e079cdb246a8fea325dfbf18dc03

Observation a6f76897-4016-44a2-a27b-cd88d1519c10 · outbound

This paper cites Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.803345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.803345Z digest=sha256:ad8fddaea4215b44794a885ae3c2d804ae9cff55341c007e4ba302f6d4be2ad2

Observation 06199c94-cb4c-4ba6-9fbb-ce320207cf17 · outbound

This paper cites Subbalakshmi.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Subbalakshmi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.806421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.806421Z digest=sha256:d61c2f6f38424c2f5b90c1453d3ad0cc86425b40a481dffdac802b57fea17a36

Observation 0028562f-ea4c-4883-9eba-f13b08177cfe · outbound

This paper cites Gender identification from e-mails.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender identification from e-mails

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.809840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.809840Z digest=sha256:0cca88980c8f2ab44b3ff1e0079facae7e4649417b5e806d755ac7433018c79f

Observation e852443a-e92c-4b30-8e51-55e8b839eb9a · outbound

This paper cites Gender, pseudonyms, and cmc: Masking identities and baring souls.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender, pseudonyms, and cmc: Masking identities and baring souls

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.814000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.814000Z digest=sha256:809695c6321d8f2868a932edf3920be2c2167024c888889686f7a97733e06418

Observation 045e3bf8-d25b-43cf-b475-a2694ddc4bca · outbound

This paper cites Bensing, W.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bensing, W

Reference 34

Resolution
verified exact
doi, observed 2026-08-15T19:15:22.274909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:15:21.817266Z digest=sha256:da4bf0eb30bb6e47886203f340701237e5d294bf7672ac45649021e7dd5b8aa3

Observation 0ba291eb-fc49-4ed3-8dc2-6a0a506eac05 · outbound

This paper cites Write it like you see it: De- tectable differences in clinical notes by race lead to differential model recommendations.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Write it like you see it: De- tectable differences in clinical notes by race lead to differential model recommendations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.821718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.821718Z digest=sha256:9625a93c8bf65b3e63aec783a80e09863c941d47a470a32fba5e5e27d336f12d

Observation 6849da2c-e308-4cdc-8854-13aed97681d3 · outbound

This paper cites Peek, and Elizabeth L.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Peek, and Elizabeth L

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.824800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.824800Z digest=sha256:9a8b09ec4f227e598e970cf659bc1ca7eaaf3b3977436b0a57ab88a4fd7c6809

Observation 03a31317-ae3e-4349-b2e1-c81e3dea94ec · outbound

This paper cites "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.828030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.828030Z digest=sha256:92dffce04bdcf773e8063b5b7c68b4863a7da965765481e13fc316bf53b4abd1

Observation 83c0a3fe-383c-4e48-adee-edae9b217784 · outbound

This paper cites How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.832911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.832911Z digest=sha256:3a1c0210c31598bb96a39e35ed484955ec15f057404ce3fdb26892775b198da8

Observation d213c9f1-f684-453a-b1cc-4664fba2219d · outbound

This paper cites Closing the gap between open source and commercial large language models for medical evidence summarization.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Closing the gap between open source and commercial large language models for medical evidence summarization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.836320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.836320Z digest=sha256:5135ea6a1ca7eafa429979a3decdc76c6a607ef69f9744a449bd8890c8377a6b

Observation 0106d6ba-7ded-4ebd-ace1-abd7f29e770a · outbound

This paper cites Conversational ai in health: Design considerations from a wizard-of-oz dermatology case study with users, clinicians and a medical llm.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Conversational ai in health: Design considerations from a wizard-of-oz dermatology case study with users, clinicians and a medical llm

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.839288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.839288Z digest=sha256:cd54500f026ddb290bb0cc4cd06b77ee5c64ec800dfb162f3095404159c2b30c

Observation 94f25ef1-0ee1-4586-87c7-d207a25eee21 · outbound

This paper cites Guidelines for rigorous evaluation of clinical llms for conversational reasoning.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Guidelines for rigorous evaluation of clinical llms for conversational reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.842895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.842895Z digest=sha256:59e29d4724b59def60ecd145c7fbab58cfaae8b786198de2d56690ad1614b8bc

Observation 5548ce9c-7ed3-40ab-9143-3f46c509df84 · outbound

This paper cites Effectiveness of a chatbot for eating disorders prevention: a randomized clinical trial.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Effectiveness of a chatbot for eating disorders prevention: a randomized clinical trial

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.846373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.846373Z digest=sha256:52f7115b277771e60a2ae5bfeb5ef83420acc365e4af73b1ca117bc8daf84bec

Observation d3a0d68d-6fa2-49d8-8688-276dc405dcc2 · outbound

This paper cites Performance of chatgpt on free-response, clinical reasoning exams.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Performance of chatgpt on free-response, clinical reasoning exams

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.849693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.849693Z digest=sha256:b18ef67a89996b5326c20778f4e8d8f7ecfda666bcfd4613fddacc9f3b52ea80

Observation 77facdb3-58ac-4d26-8c6b-ab36d281fba9 · outbound

This paper cites The next generation: chatbots in clinical psychology and psychotherapy to foster mental health–a scoping review.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The next generation: chatbots in clinical psychology and psychotherapy to foster mental health–a scoping review

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.852646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.852646Z digest=sha256:0ec55ceb1dc91c9e3034519ba565f7c950d47bcacb1cca6bda478d9a45b4863e

Observation 95c9aecb-921f-43b6-b917-99702cc6590e · outbound

This paper cites Large language model influence on diagnostic reasoning: a randomized clinical trial.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Large language model influence on diagnostic reasoning: a randomized clinical trial

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.856000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.856000Z digest=sha256:0352515f64f4c311131d9028ea1d9f1753e719a142441f8e08ef457fcbe067da

Observation aa7f93a6-87d2-4f5c-a4e9-0565fafd0eac · outbound

This paper cites Human-algorithmic interaction using a large language model-augmented artificial intelligence clinical decision support system.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Human-algorithmic interaction using a large language model-augmented artificial intelligence clinical decision support system

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.859010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.859010Z digest=sha256:a307f54241d0db8468ad09cb9df13ad8bbfa5bbfbcd7ec218b6974a9f4c46888

Observation 710a98b7-d38d-4063-b188-482e7675ed7b · outbound

This paper cites The impact of responding to patient messages with large language model assistance.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The impact of responding to patient messages with large language model assistance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.862248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.862248Z digest=sha256:235ba896ae0fa5d46c8f85e774a71771a608302265b0f9ba2f14c994d797a8c6

Observation 2fda06f4-ea88-4966-b4ff-ebb93d606c84 · outbound

This paper cites Com- paring physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Com- paring physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.865790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.865790Z digest=sha256:407c0589ee50fdbc3d20dd95bcb6e99def14ca27df1e6f42789e5e375943797f

Observation 6394c0de-e839-4e4b-abff-3253bd9aacad · outbound

This paper cites An evaluation framework for clinical use of large language models in patient interaction tasks.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making An evaluation framework for clinical use of large language models in patient interaction tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.868987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.868987Z digest=sha256:a9f4daeba985723ebe5a02e13ab6da1fe4cffd378e4a3d57fa24c7d07b77d390

Observation d492586e-0885-4e5f-8afb-d2f78a5bc812 · outbound

This paper cites The medium is the message: How non-clinical information shapes clinical decisions in llms.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The medium is the message: How non-clinical information shapes clinical decisions in llms

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.872398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.872398Z digest=sha256:5506d56331f556183aa42f651b09bfced746ff11ffd32264ee6a10ec35433870

Observation 938a4e14-eb07-4e1d-b368-9f2a02f7e45a · outbound

This paper cites Chatgpt: the next-gen tool for triaging? The American journal of emergency medicine, 69:215–217, 2023.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Chatgpt: the next-gen tool for triaging? The American journal of emergency medicine, 69:215–217, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.875617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.875617Z digest=sha256:b8b9807074cc67b704ca24c0d3e78964e2e770338e0dd44c88d966c091219c9a

Observation 9328e306-038c-457b-a345-05db4a4a59a5 · outbound

This paper cites The diagnostic and triage accuracy of the gpt-3 artificial intelligence model.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The diagnostic and triage accuracy of the gpt-3 artificial intelligence model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.879040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.879040Z digest=sha256:3ddd59254bb6e6a060447c50666c48e91b2380565ea3e738efb6cbafcbca83a2

Observation 12b16d0d-408b-4619-a8ce-fcfff8ea9885 · outbound

This paper cites Triage performance across large language models, chatgpt, and untrained doctors in emergency medicine: comparative study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Triage performance across large language models, chatgpt, and untrained doctors in emergency medicine: comparative study

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.882949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.882949Z digest=sha256:b690e8375ccea118fe4e13794718b6cbe8268d1ef059cd269275e8f24a12dbd7

Observation 68ccb041-be39-4937-9ebb-fa4b303f0fe6 · outbound

This paper cites Evaluating llm-based generative ai tools in emergency triage: A comparative study of chatgpt plus, copilot pro, and triage nurses.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating llm-based generative ai tools in emergency triage: A comparative study of chatgpt plus, copilot pro, and triage nurses

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.886354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.886354Z digest=sha256:3f6c491fce643cc859858bf16d3cc3c630630d6207cba28e7dfc97daa9758160

Observation e27deeb6-c602-4ad4-9298-7ce0c74f9acb · outbound

This paper cites Integration of customised llm for discharge summary generation in real-world clinical settings: a pilot study on russell gpt.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Integration of customised llm for discharge summary generation in real-world clinical settings: a pilot study on russell gpt

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.890590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.890590Z digest=sha256:bf4bc9fdf1ce5f0f9b638d46158174d3ad8f0caca9090115d188ba831c33a06a

Observation 126df0af-a510-4147-a76d-a8e726b53b44 · outbound

This paper cites A toolbox for surfacing health equity harms and biases in large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A toolbox for surfacing health equity harms and biases in large language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.894779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.894779Z digest=sha256:5f1dfe14f503111e9ecf36781523dfeebf97ca326215f80d0aad97b6e5adeb27

Observation 793693f4-dc90-4b61-af2a-c9f7fb331166 · outbound

This paper cites Can AI Relate: Testing Large Language Model Response for Mental Health Support.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Can AI Relate: Testing Large Language Model Response for Mental Health Support

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.898179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.898179Z digest=sha256:918924a75552e9aba452963aa4e85b4867ba995baeb8864f7da658ad00a202e1

Observation b9b7e40d-e11e-4c60-99bf-3baec94930de · outbound

This paper cites A systematic review of large language model (llm) evaluations in clinical medicine.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A systematic review of large language model (llm) evaluations in clinical medicine

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.901737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.901737Z digest=sha256:6b001ae994250c01f500f4b3d34fff1e49f4b916b8e3cf17d2ee5f9ba1c397c5

Observation 94c15ff3-939c-417a-ac27-d71eee752d28 · outbound

This paper cites Evaluating the clinical benefits of llms.Nature Medicine, 30(9):2409–2410, 2024.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating the clinical benefits of llms.Nature Medicine, 30(9):2409–2410, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.904754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.904754Z digest=sha256:a4876e679330919728347d6039f6ce72a60e0610491ae6b7f51d6bd76d07f7a4

Observation 99398c58-fd04-4ee2-ae43-53757bceb76d · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.908939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.908939Z digest=sha256:9feef182a94710e80146a02123c981374b903beb9a2b571820ffb62fba9c6b4c

Observation c245a714-1587-4c8f-8e5c-6cf2b428a2d1 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.912397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.912397Z digest=sha256:096a27bbf39f51f54a1300f5c3a276f66806305eabc0ac49e3604eaaae62f207

Observation 738c7030-4081-40ec-a74c-732def8691df · outbound

This paper cites DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models

Reference 62

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:15:21.916224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.916224Z digest=sha256:338670dfa28a857b75a5c1ac4379cbeb87514504c982f5694ecb862253ede401

Observation d9c327ef-6b38-42ba-8014-72870728a54a · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.919409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.919409Z digest=sha256:2e9b8a72cf7d0d11c533c41dd57d83c6e3845408ec0fca2821800d76aa5b6ab4

Observation 829fb1d9-8764-4bb3-946f-44676c6f096a · outbound

This paper cites Performance of large language models on medical oncology examination questions.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Performance of large language models on medical oncology examination questions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.922867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.922867Z digest=sha256:20853ce8bca2ae729d1f52c72af906b0f881b76ed461038301efabe8db5ffdbd

Observation 74f6f77b-b65e-4716-83ab-2d26855d1b0b · outbound

This paper cites Medical Large Language Model Benchmarks Should Prioritize Construct Validity.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Medical Large Language Model Benchmarks Should Prioritize Construct Validity

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.926256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.926256Z digest=sha256:f406946437917a2be773e0134df4d5f8352ff9e5897bd895b4dd88352ed4cff0

Observation fd456a41-7389-4100-bd54-ddfb12d613e5 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.930184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.930184Z digest=sha256:fc4f4d86c151597ba9bee8efc9455ab6e768278dbbfb0361bd7bec575a7b7348

Observation f32569ca-49f8-4a11-bbe0-c12970b98b9b · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.934638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.934638Z digest=sha256:1172b56d79e5bcb3cc84b63e0655a5317d8e0e0cabd0bab2f7262e7d7544b690

Observation b5873111-7447-490b-b0fd-4c507b035507 · outbound

This paper cites Automating evaluation of ai text generation in healthcare with a large language model (llm)-as- a-judge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Automating evaluation of ai text generation in healthcare with a large language model (llm)-as- a-judge

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.939396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.939396Z digest=sha256:86561c7f0373aee68fd76006520299e3d182912fd4834f2bf191890748650428

Observation aba09d39-5f6d-466b-944b-f6ba8c560a0f · outbound

This paper cites Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.942893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.942893Z digest=sha256:6d96ca4eb2c2402cd99c34917d81c01e2e2bf8b4dac1aa283c5e3f89a0c936c3

Observation 9950664d-9951-4067-b49c-804a0c7f3f1f · outbound

This paper cites A Survey on LLM-as-a-Judge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A Survey on LLM-as-a-Judge

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.946197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.946197Z digest=sha256:ee6c7e9ec71bbb3278e112d97eb1ad5a763ded0965833e7fe7e8653ff87eea4a

Observation 42676804-f7d6-42ac-bd6e-d2e393d58ebf · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Can Large Language Models Be an Alternative to Human Evaluations?

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.950376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.950376Z digest=sha256:2ee282f4b8076e7f7f7ed92e97e4c7515e94919da7a4a5a9ac13a698bc64862e

Observation 7d78a3f1-ba94-4a4e-965a-3d604a7b44f3 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.953810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.953810Z digest=sha256:184799b9752de7e233d263e555cede722bad75ab1ddbf01978ff01b9ca6de2a8

Observation b06258f5-6ac2-47b3-8c31-6a6c4f36e4ad · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.957520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.957520Z digest=sha256:49c3bbe887a20127cd44dd3ae6541364dbe558cdc1e1395438567b53f6a7900c

Observation 23e4f448-4fe3-499c-a296-62d0bb0a3edd · outbound

This paper cites As- sessment of pathology domain-specific knowledge of chatgpt and comparison to human performance.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making As- sessment of pathology domain-specific knowledge of chatgpt and comparison to human performance

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.961112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.961112Z digest=sha256:e1cd7ef17a9102dd11c207462338965533a28d59aa4b52dc0bd6fbb1b248a997

Observation 318db684-412b-4da4-8b6b-bd29aad69080 · outbound

This paper cites Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.964430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.964430Z digest=sha256:e5de581bfac8f4f061639a3b323e4af7e14f784f3bb145e95ab38057564b05dc

Observation e8d4c663-c964-4e59-b977-f379e7efeb7b · outbound

This paper cites Style Over Substance: Evaluation Biases for Large Language Models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Style Over Substance: Evaluation Biases for Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.967504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.967504Z digest=sha256:51bfbec74d88aa690be1e20e5956aca335ea26ca10828607432667aa14fb7b07

Observation ad2dbca1-09b4-414e-a234-3694e0db6415 · outbound

This paper cites Ehrnoteqa: An llm benchmark for real- world clinical practice using discharge summaries.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Ehrnoteqa: An llm benchmark for real- world clinical practice using discharge summaries

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.971049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.971049Z digest=sha256:743a14b786ac263f34f68b11433dfec7d490d814edc7e33c3f0e00bd868dd194

Observation 4ad5a4b3-0cb3-4981-a583-42f4ac9cab04 · outbound

This paper cites Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Zhuo Ran Cai, Roxana Daneshjou, and Pranav Rajpurkar.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Zhuo Ran Cai, Roxana Daneshjou, and Pranav Rajpurkar

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.974313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.974313Z digest=sha256:9c75e48946a04b97dfc9dd676132e97b71a2b3f87f19751fb203861236f69b1d

Observation a2fae7cf-5f85-49f8-b27f-5e8c2047f143 · outbound

This paper cites The Llama 3 Herd of Models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The Llama 3 Herd of Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.977984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.977984Z digest=sha256:bba013857793a6293b2f9431096f5d116cca2d9e6148897c879f791596621446

Observation 0d1bc52a-cb62-4ce3-b0e6-3c1024250754 · outbound

This paper cites Linguistic analy- sis of communication in therapist-assisted internet-delivered cognitive behavior therapy for generalized anxiety disorder.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Linguistic analy- sis of communication in therapist-assisted internet-delivered cognitive behavior therapy for generalized anxiety disorder

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.980919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.980919Z digest=sha256:f633f98251a00f3216f2bfb2f42ad3fd55eb2f8e77bdfbbc1f184c0df5067b90

Observation 34a263ac-a9f0-4e76-aab2-8184bb42c548 · outbound

This paper cites Toward linguistic recognition of generalized anxiety disorder.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Toward linguistic recognition of generalized anxiety disorder

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.984000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.984000Z digest=sha256:30fe8327f5e343be65e312f1144aae5258f0803276ce1b58d0f104c25846b475

Observation bab69140-78dd-404e-84cf-a0849dea1d7f · outbound

This paper cites Linguistic markers of anxiety and depression in somatic symptom and related disorders: Observational study of a digital intervention.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Linguistic markers of anxiety and depression in somatic symptom and related disorders: Observational study of a digital intervention

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.987812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.987812Z digest=sha256:80e96cbfa059918fab01d8205eb4e9d5ab41efb213035dc92859330a6f14298f

Observation 11dd3167-df4a-4bb8-b04d-32554d1dc644 · outbound

This paper cites Are patient linguistic tones associated with mental health and perceived clinician empathy? JBJS, 103(23):2181–2189, 2021.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Are patient linguistic tones associated with mental health and perceived clinician empathy? JBJS, 103(23):2181–2189, 2021

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.991152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.991152Z digest=sha256:ebc1b700fe9770f2bb9a53168958978528abe5c95dc4e5280e0bc52382f5d4ec

Observation 56185fba-c4b9-4e81-b80f-3eb7c82ac921 · outbound

This paper cites GPT-4 Technical Report.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making GPT-4 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.994385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.994385Z digest=sha256:c9646f1837c4f545f8b0f4a5755c87145707ee271d241b484fd16a38694d83b0

Observation 860ff5ca-935d-4850-898e-cd23879b55c0 · outbound

This paper cites Palmyra-med: Instruction-based fine-tuning of llms enhancing medical domain performance.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Palmyra-med: Instruction-based fine-tuning of llms enhancing medical domain performance

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:21.997880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:21.997880Z digest=sha256:574e46c39ef89a10ac5e7281bc378c298ff327bc1f76ef6b5d549b997fa4039d

Observation 1d2fba3d-82ac-4d12-a340-e95bd0f2af14 · outbound

This paper cites The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.001808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.001808Z digest=sha256:074be3319ce1f41e3d6e3fafb207931dc4b486f06d108518b5b6b2cea5427e74

Observation 3e6b9251-d1e5-4672-8a36-51befbafa48c · outbound

This paper cites Multiple significance tests: the bonferroni method.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Multiple significance tests: the bonferroni method

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.005542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.005542Z digest=sha256:b0dddc18256c582a149e6c9928606a10391ec6aa78b648f62186adb903ee0ae4

Observation d9b4bed2-ba0b-42aa-a6a3-bf50ac4e2945 · outbound

This paper cites Note on the sampling error of the difference between correlated proportions or percentages.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Note on the sampling error of the difference between correlated proportions or percentages

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.008516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.008516Z digest=sha256:4fc7a46bbab54558d7f09ffa7dc941d773ea442bfa619ac0f9362eb3c645344e

Observation b19016fc-cee7-4c16-b5a6-f9b4178cb3e8 · outbound

This paper cites Individual comparisons by ranking methods.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Individual comparisons by ranking methods

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.011752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.011752Z digest=sha256:4a91c1e40c4359690d7a024627c32379439757059d29005b122070fb238caaa6

Observation 68fada91-d8d6-4e98-9c56-1fdf01e888d6 · outbound

This paper cites The measurement of observer agreement for categorical data.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The measurement of observer agreement for categorical data

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.015339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.015339Z digest=sha256:382aee9264521466d00781931e04da56320b172c9cd2f85b2cb87c46cfd1c6c1

Observation 15692283-b172-4ac9-99ca-b5000a05cc08 · outbound

This paper cites On the use and interpretation of certain test criteria for purposes of statistical inference part i.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making On the use and interpretation of certain test criteria for purposes of statistical inference part i

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.018639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.018639Z digest=sha256:a9c9118f05c488762084e71274687c2638c6b9ae7c1d2815339b7baaf04a5c63

Observation dcf64b46-7631-4460-b53a-9f9cc53a0df0 · outbound

This paper cites On a test of whether one of two random variables is stochastically larger than the other.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making On a test of whether one of two random variables is stochastically larger than the other

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.021993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.021993Z digest=sha256:1de9b24a22ceced68e498bce4d23ce3211c01f3b52d81f722c1c4bebfc25f74c

Observation a72ee610-6d01-4221-8918-e0fe277004d1 · outbound

This paper cites Gender bias and stereotypes in large language models.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias and stereotypes in large language models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.024908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.024908Z digest=sha256:2a2acb2771c8df056ca8536437dced7723dd64057bcd847be675db7fe399774e

Observation 0a2be227-7012-4856-9130-dbed21282f16 · outbound

This paper cites Llm evaluators recognize and favor their own generations.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Llm evaluators recognize and favor their own generations

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.027941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.027941Z digest=sha256:cbb74f8bff9cc7aef9c0850901e4c68261b03b556bbb393137ac33f5eba3d17e

Observation b8147dac-cbeb-41b5-8556-710015d964b2 · outbound

This paper cites Bender and Batya Friedman.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bender and Batya Friedman

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.031212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.031212Z digest=sha256:460071bb0b5d2bc014a49ba9c72edff3de1f7d70529398681beab2a31bf0634d

Observation 906af124-436a-4233-a10b-574b49c57199 · outbound

This paper cites Basic demographics, health practices, and health status of us medical students.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Basic demographics, health practices, and health status of us medical students

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.034556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.034556Z digest=sha256:3664beb942faba651f7597f18cdba027fadaa6cb898438db740494315f9c4c2e

Observation df7aa351-d8ab-4712-bdc0-a1a192dcd0cd · outbound

This paper cites International medical graduates in the us physician workforce and graduate medical education: current and historical trends.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making International medical graduates in the us physician workforce and graduate medical education: current and historical trends

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.037896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.037896Z digest=sha256:54d4e9060aa8c1ff5b995307265f502da1eb5eccc0da7ce78f5a9060edebbda1

Observation ba0cffc4-169b-46d6-b428-a46af3745e46 · outbound

This paper cites Clinical reasoning education at us medical schools: results from a national survey of internal medicine clerkship directors.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Clinical reasoning education at us medical schools: results from a national survey of internal medicine clerkship directors

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.041997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.041997Z digest=sha256:aa68288cbc20528d5cf7aa14ad53fe1dc28b1e1899f27b99c4df9771160a9dee

Observation 27e35414-7bf1-44df-91fa-af90972b99f5 · outbound

This paper cites Teaching medical students the important connection between communication and clinical reasoning.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Teaching medical students the important connection between communication and clinical reasoning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.046929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.046929Z digest=sha256:2af5af24ea3d0a6768bf96ca4a459e415fb5db0c8a346a075d46f0b930b96126

Observation 335521db-c88c-40fc-a86f-ff80cedd88af · outbound

This paper cites Factors associated with medical student clinical reasoning and evidence based medicine practice.

The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Factors associated with medical student clinical reasoning and evidence based medicine practice

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-15T19:15:22.050331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:15:22.050331Z digest=sha256:1ae1c9e8d75022ec678a60619ed39f7902a5a32502a5f57e26c7c989a685e084

Pith citing papers

Observation fb758406-220f-4da9-b0e5-e28ca6c4b238 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T19:05:10.672097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:8e6d022e13b6d90dcb56fc0ee833f5033e5e1670e697aa275444304af6b2147b

Observation 17c60d7f-00ef-42b6-be75-24438d6c3d87 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.549200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:f64bbaba054f91267be152fc4d00bc17eeb95918540456f8181360ebcab485a4