Pith. sign in

Paper Citation Record · LEDGER

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

As of 3 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2604.20441.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.20441 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:23:11.924161Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:13:05.342181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2847d94c-c931-4c37-a0d7-30ecd0f1aec1 · outbound

This paper cites SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:20:05.782711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:fec22e9d9823d3632a54d4ea1596f5f61a6c5dd7ed93eb02b9e0c7eb552a2a0f

Observation f58c9f30-9fa4-49ea-aafa-a2f8016efb60 · outbound

This paper cites Available: https://arxiv.org/abs/2603.04448.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Available: https://arxiv.org/abs/2603.04448

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:24:46.851570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:410381770238719378b8d1f35f4f6874e3f877700237b7dc8ea814ad2842106f

Observation 387dfcc1-819c-40dd-a578-28cfb4369da3 · outbound

This paper cites Artificial hallucinations in ChatGPT: implications in scientific writ- ing.Cureus.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Artificial hallucinations in ChatGPT: implications in scientific writ- ing.Cureus

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.442383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:b58a400f423ae587d7abd706e2ec06024ff40181c2828bf9c80fc254a2f767a8

Observation bdad854f-bcd1-457c-845f-26ebfa7c95e0 · outbound

This paper cites Evaluating large language models and agents in healthcare: key challenges in clinical applications.Intelligent Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Evaluating large language models and agents in healthcare: key challenges in clinical applications.Intelligent Medicine

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.430105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:071fe9a0bbada9e0991ca1cf2401b7cf8878045f92451f1f061538e63cbb74c2

Observation 9408c621-496b-4684-bbb0-a26fdaac6da7 · outbound

This paper cites Survey of hallucination in natural language generation.ACM Computing Surveys.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Survey of hallucination in natural language generation.ACM Computing Surveys

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.426975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:d15d0e6339ecc92245b8cf32627d04b5f6f750cfa4b00ae27dee067411017ba4

Observation ad98ef44-2e97-4ca7-a20d-7c4655b1de5b · outbound

This paper cites Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.432830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:781e4819a7e85449e446a7dc6c169d73b2f820350fc9e060f85e3e4a764afbfd

Observation daf00cea-982c-4247-b2c7-5a2a64badf19 · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Capabilities of GPT-4 on Medical Challenge Problems

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:43:34.900774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:95964adb66a4aca4306a4cb054e46a97e8890386d19108a7d4e58194cf2dcad1

Observation 0c63dbd3-3003-4cd2-ab30-81e2ace8e169 · outbound

This paper cites Toward expert-level medical question answering with large language models.Nature Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Toward expert-level medical question answering with large language models.Nature Medicine

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.421182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:51be2378341f66429378095a444dc7c438d35458f59848a1979aab9729404607

Observation fd1967b1-fdc5-4b47-bb65-25ad8f361079 · outbound

This paper cites A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains.npj Digital Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains.npj Digital Medicine

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.418331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:5fd2a3b03c8b131f8500d9fc5f93f7a84c04f12b0fb6cc9192cea15325d10a7c

Observation 005e241e-5a16-42f2-bf25-2288a53cdda0 · outbound

This paper cites Large language model agents for biomedicine: a comprehensive review of methods, evaluations, challenges, and future directions.Information.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Large language model agents for biomedicine: a comprehensive review of methods, evaluations, challenges, and future directions.Information

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.424241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:5f41f07347c65a45c05992b2b2d3e104d024282792c7f4cd4b240437cf8ccf77

Observation 4d08f078-c778-4bac-a497-96a05758cf77 · outbound

This paper cites MedAgentBench: a virtual EHR environment to benchmark medical LLM agents.NEJM AI.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills MedAgentBench: a virtual EHR environment to benchmark medical LLM agents.NEJM AI

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.435941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:56c55c88b3c28ddfa731c8a13ad0449880248a2234522b7641d1e53025f7be33

Observation a22f7ee9-8407-44b0-9ea0-80cc2561f871 · outbound

This paper cites The clinicians’ guide to large language models: a general perspective with a focus on hallucinations.Interactive Journal of Medical Research.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills The clinicians’ guide to large language models: a general perspective with a focus on hallucinations.Interactive Journal of Medical Research

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.439189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:540c40e7f6ce67974af004b86eaa05791a9c19de300f6a1132a6171228902cd3

Observation 565ed65d-653a-4a3a-b728-8822b37d76af · outbound

This paper cites Human researchers are superior to large language models in writing a medical systematic review in a comparative multitask assessment.Scientific Reports.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Human researchers are superior to large language models in writing a medical systematic review in a comparative multitask assessment.Scientific Reports

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.408720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:fbd277a857a71e5a8e66ef01b6e67ce29a9e8121ef8de339af6460cb73a9bac0

Observation 53a6b705-4d3a-4897-b309-2cb9beaf5c3e · outbound

This paper cites Citation integrity in the age of AI: evaluating the risks of reference hallucination in maxillofacial literature.Journal of Cranio-Maxillofacial Surgery.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Citation integrity in the age of AI: evaluating the risks of reference hallucination in maxillofacial literature.Journal of Cranio-Maxillofacial Surgery

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.405836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:43ffbe2c26e127d4cd4927ed099a3a3f75a128bf9d66a5cc08828a0ba4f1813f

Observation 025bbcfd-554f-446a-bb02-500322b85120 · outbound

This paper cites 2025;12:e80371.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills 2025;12:e80371

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.445786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:19792d7553bd6e406bf5c10fd8f5b56e3635391c9d240df6c02874f81aa87a78

Observation 561e8192-e414-458d-9353-44f1a726abf8 · outbound

This paper cites Systems and software engineering — Systems and software Quality Re- quirements and Evaluation (SQuaRE) — System and software quality models.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Systems and software engineering — Systems and software Quality Re- quirements and Evaluation (SQuaRE) — System and software quality models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.411454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:d2a5fe312288bb4820c9ad421e202c600e718144b0a455c4e13d1be249b9c729

Observation 4da631c9-e38e-4451-b8c3-71e12503e2cc · outbound

This paper cites Data structures for statistical computing in Python.Proceedings of the 9th Python in Science Conference.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Data structures for statistical computing in Python.Proceedings of the 9th Python in Science Conference

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.414472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:b7f7597dc5ae42f8166e3e8d92cba3cbfd2778efef08f5a87e62f1a145e1ca22

Observation 63a37b86-0588-457a-aff9-f5ad6307b955 · outbound

This paper cites Pingouin: statistics in Python.Journal of Open Source Software.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Pingouin: statistics in Python.Journal of Open Source Software

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.462501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:dd55db5fc9e26e11fa5fd6bf94efb5f11588dddbf574e65a51331cdfb0e3b618

Observation a1fecdbd-0f00-43f0-b644-3fdd6af9e029 · outbound

This paper cites SciPy 1.0: fundamental algorithms for scientific computing in Python.Nature Methods.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills SciPy 1.0: fundamental algorithms for scientific computing in Python.Nature Methods

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.458971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:c08bd2009aa6fe7e1e509a94ba2bd7d24e46871db9a252728dad328db34187ac

Observation 5b40463b-23a6-4d72-b3a0-629a133577a9 · outbound

This paper cites Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.455243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:49e5c70a16cfd38e5f67c1a65550e11bab8bb7c23502a7d0c3f1cc2b8a0a1384

Observation b95c98ec-d528-42ba-bfa1-afe077dfc496 · outbound

This paper cites A guideline of selecting and reporting intraclass correlation coefficients for reliability research.Journal of Chiropractic Medicine.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills A guideline of selecting and reporting intraclass correlation coefficients for reliability research.Journal of Chiropractic Medicine

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.449377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:5223e6ef403042db7eb2419ede4aad941950727001c1c9a9581322fa0735026c

Observation 5408ba97-481c-4d44-9529-a39d5275c84b · outbound

This paper cites Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit.Psychological Bulletin.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit.Psychological Bulletin

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.452309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:8b5d70aa6c84091e4e4fc0db3158a8be21843215ca271f5eddf34e6428d6e3dc

Observation 11af07d3-5b69-4f4b-a94a-f244269f03fe · outbound

This paper cites Statistical methods for assessing agreement between two methods of clinical measurement.The Lancet.

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Statistical methods for assessing agreement between two methods of clinical measurement.The Lancet

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T11:25:38.466281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T00:23:11.924161Z digest=sha256:befac04aba0e6aff6a1f427d70407ff5f8bbf7e88be3f685c64ef47fb6b55c17

Pith citing papers

Observation 976039e1-885a-41ad-9f78-d3851b1bacd5 · inbound

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries cites this paper.

Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T14:13:05.342181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:13:05.342181Z digest=sha256:57ee1e7f5cb2f780b41b083b8943015437cdfaa4d5c81fcb08c1403ab074bde7