Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:15:22.050331Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 100 of 151 outbound references and 2 inbound Pith citation observations for arXiv:2506.17163.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:15:22.050331Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-09T19:02:46.991897Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 151 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 9e4ffd32-3ebf-4699-982f-5fcfb0c4ded6 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating large language models on medical evidence summarization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4150240-e4d1-44a2-90c9-a862884d2a4d · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Adapted large language models can outperform medical experts in clinical text summarization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3bde49-59c2-4ada-89d2-7e39b1f01e08 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2acbc252-7251-4b32-9fe8-e52d9b257a0b · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Llm-based agentic systems in medicine and healthcare
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07767479-b3f5-4f23-b0fc-55911580a52b · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making MedDM:LLM-executable clinical guidance tree for clinical decision-making
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e46a8b-8451-40f1-bb9c-4a5c6e56aae5 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Large language models encode clinical knowledge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9883c0e-2b68-4294-b0a1-ad3bc2e3d75a · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Toward expert-level medical question answering with large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3032ad41-4e58-4e07-9730-0e3dda922e41 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Luks and Zachary D
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 604dc004-ab51-4576-b071-c5aacd248574 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Variability in language used on social media prior to hospital visits
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3086486e-829f-4e22-a92f-4f862604e9be · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making How does chatgpt perform on the united states medical licensing examination (usmle)? the implications of large language models for medical education and knowledge assessment
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f08073-c5a2-4ebf-810b-3166225e8530 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Open medical llm leaderboard
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163ba107-58bb-4b94-9fc5-0faa9c41c254 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A rapid review of gender, sex, and sexual orientation documentation in electronic health records
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1dc0c61-a6ca-4f36-9ccd-88d90607b332 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Health Care Experiences of Patients with Nonbinary Gender Identities
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b054846d-c701-4048-9a16-1075165d6f0e · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Hoffmann, Roger B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cc38915-e466-45da-89d9-9a1fa99bc006 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df9834ba-ed33-4765-be03-b5b8f5a6d95f · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender disparities in health care
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a80d4a-d602-4a80-a0f8-ac13be4c3c0b · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Defining gender disparities in pain management
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b854cd-ef13-4eb3-91cf-389d54d53754 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender differences in outcomes of a multimodal pain management program
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b8a5ea0-c979-47fc-b5b8-eeba7f4a7354 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Health and healthcare disparities among us women and men at the intersection of sexual orientation and race/ethnicity: a nationally representative cross-sectional study
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ca81320-71ba-4cbe-b4b8-4f2acbfcf452 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias in transformers: A comprehensive review of detection and mitigation strategies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3818626d-6e34-46af-80ab-5e088208544f · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias in natural language processing and computer vision: A comparative survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eae15c9e-3435-493d-b2cf-5b5970538e03 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03496300-554e-48c5-8674-3c11f4013c81 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Addressing gender-related performance disparities in neural rankers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad8a1c1-fa49-4ea2-89ad-58928594874e · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae8aaac9-2fa7-4bc4-b193-cb82ec9fc32a · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bias in bios: A case study of semantic representation bias in a high-stakes setting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b87c217d-1adc-4fbb-b025-c165642ae253 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Assessing the potential of gpt-4 to perpetuate racial and gender biases in health care: a model evaluation study
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f18a0f6-7c6b-42fc-b0db-22bd40ee9df7 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bias patterns in the application of LLMs for clinical decision support: A comprehensive study
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e7c5b1-aa5d-46c0-b7c7-59e788268d01 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making An investigation into the impact of deep learning model choice on sex and race bias in cardiac mr segmentation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5c3ee91-32a4-45ee-a99b-b1b5ed9739cf · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6f76897-4016-44a2-a27b-cd88d1519c10 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06199c94-cb4c-4ba6-9fbb-ce320207cf17 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Subbalakshmi
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0028562f-ea4c-4883-9eba-f13b08177cfe · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender identification from e-mails
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e852443a-e92c-4b30-8e51-55e8b839eb9a · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender, pseudonyms, and cmc: Masking identities and baring souls
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 045e3bf8-d25b-43cf-b475-a2694ddc4bca · outbound
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0ba291eb-fc49-4ed3-8dc2-6a0a506eac05 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Write it like you see it: De- tectable differences in clinical notes by race lead to differential model recommendations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6849da2c-e308-4cdc-8854-13aed97681d3 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Peek, and Elizabeth L
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03a31317-ae3e-4349-b2e1-c81e3dea94ec · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making "Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c0a3fe-383c-4e48-adee-edae9b217784 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making How Individual Traits and Language Styles Shape Preferences In Open-ended User-LLM Interaction: A Preliminary Study
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d213c9f1-f684-453a-b1cc-4664fba2219d · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Closing the gap between open source and commercial large language models for medical evidence summarization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0106d6ba-7ded-4ebd-ace1-abd7f29e770a · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Conversational ai in health: Design considerations from a wizard-of-oz dermatology case study with users, clinicians and a medical llm
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94f25ef1-0ee1-4586-87c7-d207a25eee21 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Guidelines for rigorous evaluation of clinical llms for conversational reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5548ce9c-7ed3-40ab-9143-3f46c509df84 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Effectiveness of a chatbot for eating disorders prevention: a randomized clinical trial
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a0d68d-6fa2-49d8-8688-276dc405dcc2 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Performance of chatgpt on free-response, clinical reasoning exams
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77facdb3-58ac-4d26-8c6b-ab36d281fba9 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The next generation: chatbots in clinical psychology and psychotherapy to foster mental health–a scoping review
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c9aecb-921f-43b6-b917-99702cc6590e · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Large language model influence on diagnostic reasoning: a randomized clinical trial
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa7f93a6-87d2-4f5c-a4e9-0565fafd0eac · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Human-algorithmic interaction using a large language model-augmented artificial intelligence clinical decision support system
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 710a98b7-d38d-4063-b188-482e7675ed7b · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The impact of responding to patient messages with large language model assistance
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fda06f4-ea88-4966-b4ff-ebb93d606c84 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Com- paring physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6394c0de-e839-4e4b-abff-3253bd9aacad · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making An evaluation framework for clinical use of large language models in patient interaction tasks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d492586e-0885-4e5f-8afb-d2f78a5bc812 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The medium is the message: How non-clinical information shapes clinical decisions in llms
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938a4e14-eb07-4e1d-b368-9f2a02f7e45a · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Chatgpt: the next-gen tool for triaging? The American journal of emergency medicine, 69:215–217, 2023
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9328e306-038c-457b-a345-05db4a4a59a5 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The diagnostic and triage accuracy of the gpt-3 artificial intelligence model
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b16d0d-408b-4619-a8ce-fcfff8ea9885 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Triage performance across large language models, chatgpt, and untrained doctors in emergency medicine: comparative study
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ccb041-be39-4937-9ebb-fa4b303f0fe6 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating llm-based generative ai tools in emergency triage: A comparative study of chatgpt plus, copilot pro, and triage nurses
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27deeb6-c602-4ad4-9298-7ce0c74f9acb · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Integration of customised llm for discharge summary generation in real-world clinical settings: a pilot study on russell gpt
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126df0af-a510-4147-a76d-a8e726b53b44 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A toolbox for surfacing health equity harms and biases in large language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793693f4-dc90-4b61-af2a-c9f7fb331166 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Can AI Relate: Testing Large Language Model Response for Mental Health Support
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b7e40d-e11e-4c60-99bf-3baec94930de · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A systematic review of large language model (llm) evaluations in clinical medicine
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c15ff3-939c-417a-ac27-d71eee752d28 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Evaluating the clinical benefits of llms.Nature Medicine, 30(9):2409–2410, 2024
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99398c58-fd04-4ee2-ae43-53757bceb76d · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c245a714-1587-4c8f-8e5c-6cf2b428a2d1 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738c7030-4081-40ec-a74c-732def8691df · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making DiversityMedQA: A benchmark for assessing demographic biases in medical diagnosis using large language models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c327ef-6b38-42ba-8014-72870728a54a · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 829fb1d9-8764-4bb3-946f-44676c6f096a · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Performance of large language models on medical oncology examination questions
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74f6f77b-b65e-4716-83ab-2d26855d1b0b · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Medical Large Language Model Benchmarks Should Prioritize Construct Validity
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd456a41-7389-4100-bd54-ddfb12d613e5 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f32569ca-49f8-4a11-bbe0-c12970b98b9b · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5873111-7447-490b-b0fd-4c507b035507 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Automating evaluation of ai text generation in healthcare with a large language model (llm)-as- a-judge
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba09d39-5f6d-466b-944b-f6ba8c560a0f · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Limitations of the llm-as-a-judge approach for evaluating llm outputs in expert knowledge tasks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9950664d-9951-4067-b49c-804a0c7f3f1f · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making A Survey on LLM-as-a-Judge
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42676804-f7d6-42ac-bd6e-d2e393d58ebf · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Can Large Language Models Be an Alternative to Human Evaluations?
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d78a3f1-ba94-4a4e-965a-3d604a7b44f3 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b06258f5-6ac2-47b3-8c31-6a6c4f36e4ad · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e4f448-4fe3-499c-a296-62d0bb0a3edd · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making As- sessment of pathology domain-specific knowledge of chatgpt and comparison to human performance
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318db684-412b-4da4-8b6b-bd29aad69080 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Quality of answers of generative large language models versus peer users for interpreting laboratory test results for lay patients: evaluation study
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8d4c663-c964-4e59-b977-f379e7efeb7b · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Style Over Substance: Evaluation Biases for Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad2dbca1-09b4-414e-a234-3694e0db6415 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Ehrnoteqa: An llm benchmark for real- world clinical practice using discharge summaries
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad5a4b3-0cb3-4981-a583-42f4ac9cab04 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Tran, Daniel I Schlessinger, Shannon Wongvibulsin, Zhuo Ran Cai, Roxana Daneshjou, and Pranav Rajpurkar
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fae7cf-5f85-49f8-b27f-5e8c2047f143 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The Llama 3 Herd of Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1bc52a-cb62-4ce3-b0e6-3c1024250754 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Linguistic analy- sis of communication in therapist-assisted internet-delivered cognitive behavior therapy for generalized anxiety disorder
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a263ac-a9f0-4e76-aab2-8184bb42c548 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Toward linguistic recognition of generalized anxiety disorder
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab69140-78dd-404e-84cf-a0849dea1d7f · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Linguistic markers of anxiety and depression in somatic symptom and related disorders: Observational study of a digital intervention
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11dd3167-df4a-4bb8-b04d-32554d1dc644 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Are patient linguistic tones associated with mental health and perceived clinician empathy? JBJS, 103(23):2181–2189, 2021
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56185fba-c4b9-4e81-b80f-3eb7c82ac921 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making GPT-4 Technical Report
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860ff5ca-935d-4850-898e-cd23879b55c0 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Palmyra-med: Instruction-based fine-tuning of llms enhancing medical domain performance
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2fba3d-82ac-4d12-a340-e95bd0f2af14 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e6b9251-d1e5-4672-8a36-51befbafa48c · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Multiple significance tests: the bonferroni method
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9b4bed2-ba0b-42aa-a6a3-bf50ac4e2945 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Note on the sampling error of the difference between correlated proportions or percentages
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19016fc-cee7-4c16-b5a6-f9b4178cb3e8 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Individual comparisons by ranking methods
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fada91-d8d6-4e98-9c56-1fdf01e888d6 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making The measurement of observer agreement for categorical data
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15692283-b172-4ac9-99ca-b5000a05cc08 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making On the use and interpretation of certain test criteria for purposes of statistical inference part i
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcf64b46-7631-4460-b53a-9f9cc53a0df0 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making On a test of whether one of two random variables is stochastically larger than the other
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a72ee610-6d01-4221-8918-e0fe277004d1 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Gender bias and stereotypes in large language models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a2be227-7012-4856-9130-dbed21282f16 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Llm evaluators recognize and favor their own generations
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8147dac-cbeb-41b5-8556-710015d964b2 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Bender and Batya Friedman
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 906af124-436a-4233-a10b-574b49c57199 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Basic demographics, health practices, and health status of us medical students
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7aa351-d8ab-4712-bdc0-a1a192dcd0cd · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making International medical graduates in the us physician workforce and graduate medical education: current and historical trends
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0cffc4-169b-46d6-b428-a46af3745e46 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Clinical reasoning education at us medical schools: results from a national survey of internal medicine clerkship directors
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27e35414-7bf1-44df-91fa-af90972b99f5 · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Teaching medical students the important connection between communication and clinical reasoning
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335521db-c88c-40fc-a86f-ff80cedd88af · outbound
The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making Factors associated with medical student clinical reasoning and evidence based medicine practice
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb758406-220f-4da9-b0e5-e28ca6c4b238 · inbound
Compared to What? Baselines and Metrics for Counterfactual Prompting The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 17c60d7f-00ef-42b6-be75-24438d6c3d87 · inbound
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering The MedPerturb Dataset: What Non-Content Perturbations Reveal About Human and Clinical LLM Decision Making
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.