Pith. sign in

Paper Citation Record · LEDGER

Collaboration among Multiple Large Language Models for Medical Question Answering

As of 17 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2505.16648.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16648 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:08.628751Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c796a62-7f5a-42ac-8c29-6e6b8be5dbec · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Collaboration among Multiple Large Language Models for Medical Question Answering Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.621356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.621356Z digest=sha256:d2880d2f7a1bc5c1c353a5a497fc317250d32def5684d9ed4845b45b5bc42f0b

Observation 992ffb51-d662-4a81-ab2b-6ca34058509c · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Collaboration among Multiple Large Language Models for Medical Question Answering Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.672437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.672437Z digest=sha256:35bcd2d2e3e27cf72f926594627fe29df86910606b4389fb9390c9333682cb2d

Observation 0f4fa892-563d-481b-9ffc-1e0554f68c77 · outbound

This paper cites Meditron-70b: Scaling medical pretraining for large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Meditron-70b: Scaling medical pretraining for large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.394863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:05.831701Z digest=sha256:b1fa67e01f5d546f5a5063c7a3c11f0c2ec2850b646025a28bcbd3f01aa94980

Observation 0d17389c-b692-49f5-88cb-7b9f5c11f9de · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

Collaboration among Multiple Large Language Models for Medical Question Answering MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.922781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.922781Z digest=sha256:c5e624a968390f88769a6fbd158e9c6ef1d5febe7661214ce330b43ba9782853

Observation 34e413f6-317e-43a8-ba03-a49903e0d401 · outbound

This paper cites ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?.

Collaboration among Multiple Large Language Models for Medical Question Answering ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T15:02:09.035183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:05.989565Z digest=sha256:4616cf570704006eb6599bd2432c7b9671adde8a8384eb100d482ef7e7c780ad

Observation 6a53e10c-b5f0-41f5-bd94-620c917bf9bd · outbound

This paper cites Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.105502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.105502Z digest=sha256:590d8890c956a78d128a1c6769dbfb6bf57272cf88a9310c587df2a53eded197

Observation 7ca43031-3112-4a53-8b85-2c4bfbcbf81e · outbound

This paper cites How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,.

Collaboration among Multiple Large Language Models for Medical Question Answering How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.138444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.175193Z digest=sha256:028841c3c70fcdd91d06bae32f47315c16f116880546818bae781894a2bb3571

Observation ccf5b023-eedb-4ae1-9c7c-34151049bf1e · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

Collaboration among Multiple Large Language Models for Medical Question Answering Capabilities of GPT-4 on Medical Challenge Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.263833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.263833Z digest=sha256:441383f2b06cc1348936dc0c86900635773484fcb8211978eedfd3107e23d9cd

Observation 7ea434a5-7605-427c-a156-480504eb66f0 · outbound

This paper cites Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,.

Collaboration among Multiple Large Language Models for Medical Question Answering Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.876540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.367380Z digest=sha256:57446dfe5fba8c1427a10b2a3b63d4d2c764e5b22683060f1b4029a581fd17da

Observation 12c0d2f2-0c9d-4323-b9dc-397be47b9bcf · outbound

This paper cites Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?.

Collaboration among Multiple Large Language Models for Medical Question Answering Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.367099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.565073Z digest=sha256:eff64afba83a2bd56c984bbf2d4fee743dcf8f33a05fc61d966aa6065e9590d9

Observation 515bf289-3aec-4d6a-a9d2-ddba1ebcb442 · outbound

This paper cites Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,.

Collaboration among Multiple Large Language Models for Medical Question Answering Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.095048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.674469Z digest=sha256:164c8fe46d1707d9c65782f1c411b0a9eee04dd5e476a09b4694e5e4058f4fe7

Observation 611d8b7f-2510-46db-a5ed-78cab0fc3ae5 · outbound

This paper cites Reasoning with large language models for medical question answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering Reasoning with large language models for medical question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.424555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.797610Z digest=sha256:5ae162c5744815a59fbb8d12ec05a270c33d73b64ca033f1b5f897c132cfc3fc

Observation fff2947a-95ac-4d12-944c-032250e05372 · outbound

This paper cites LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,.

Collaboration among Multiple Large Language Models for Medical Question Answering LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.140056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.843950Z digest=sha256:dcb5b20a4e651b070b84747b858a1c47969db9fb21ec7c015039574f4f3f6890

Observation ac4e2b73-7c80-4877-9b78-f98ca099b073 · outbound

This paper cites Adaptive Chameleon or Stubborn Sloth:,.

Collaboration among Multiple Large Language Models for Medical Question Answering Adaptive Chameleon or Stubborn Sloth:,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.763553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.901674Z digest=sha256:9d2fd3f5e0f749f6a398839b2b02d77317020069c766f32c1313eb6e4a5ea2c1

Observation e06c1b25-ab5e-43bc-b541-fdd8c8d5340c · outbound

This paper cites Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.423552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:07.022628Z digest=sha256:e2826458e3381b7232838a9d55a240f2a3c3b21c099792f112dcc946ba011d84

Observation fa868ea9-770b-4a39-a717-4ac986eca2ab · outbound

This paper cites Exploring collaboration mechanisms for LLM agents: A social psychology view,.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring collaboration mechanisms for LLM agents: A social psychology view,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.087397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:07.118020Z digest=sha256:75be23b5f777c9d420440d4b95810ea2c45994caee294beb49c1f4be8a9a248f

Observation 68489507-5e2d-4677-83fc-7b14593f6f49 · outbound

This paper cites Counterfactual debating with preset stances for hallucination elimination of llms,.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual debating with preset stances for hallucination elimination of llms,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:07.219238Z digest=sha256:d96859dae3a0accd5f27dc34ab73867be686bdf3f3464060c7c9692b80166c0c

Observation ccb093c0-cc84-469f-bdda-bbdd8eea0de8 · outbound

This paper cites One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.388081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.388081Z digest=sha256:ef792007881f85b91629707dc08ad9851c2be5892c73506cc736cff6954d9a0c

Observation 20c29ee5-b7aa-46b5-9836-c08769d9753d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.471313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.471313Z digest=sha256:a7c6dcb96f8fd476c2f00024810261041b553a0c72d7bdf9134c0f71f24efd57

Observation 9191cfbb-6da6-4588-a62d-48da69099284 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Collaboration among Multiple Large Language Models for Medical Question Answering Large Language Models are Zero-Shot Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.588376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.588376Z digest=sha256:eaadf6fbe086f65f136fc8952b26d9b92556118f73c6df7c02021599d3cc6df3

Observation ecbd8b30-c832-4d9d-ad3e-668d19d1b922 · outbound

This paper cites Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,.

Collaboration among Multiple Large Language Models for Medical Question Answering Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,

Reference 21

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.921281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:07.670413Z digest=sha256:6ddc64d5d4093347ef27c268eb833b52a80304a97880e2c9c379b34a1f63a1f7

Observation b6e9e6ad-a41b-409a-8d67-e174dd393f44 · outbound

This paper cites Mixtral of Experts.

Collaboration among Multiple Large Language Models for Medical Question Answering Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.767358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.767358Z digest=sha256:450baf4aaf14eebf12697037acb16fc55b55b30ffeeefa43a48d94a33018e829

Observation b7d0c431-08c2-4769-8804-95373363d289 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.897800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.897800Z digest=sha256:176b108b53806ab2a4a44146bc05ab3a3d1fe88e81e6832881a98e2455f724bf

Observation 047e65b5-0d3b-449b-b67d-3212a72c01e7 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering QLoRA: Efficient Finetuning of Quantized LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.959967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.959967Z digest=sha256:151955d7fd18e8c9e5a77fe5373828ed066def8921926b0a4eacbd240fad93e5

Observation d4d06f39-afbc-4e11-947d-80b83b9585f3 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Collaboration among Multiple Large Language Models for Medical Question Answering Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.057788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.057788Z digest=sha256:745887f34ea0405f78e8ac9a4bc03973c0708bca6a0b27b04ba6fefd54dc8353

Observation 81563bc2-fe86-4676-b6f6-0f29cb59e984 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Collaboration among Multiple Large Language Models for Medical Question Answering Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.405653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:08.188443Z digest=sha256:0ab727c42470a18f8b1f2bfb1c4065e26fa3b14932b1a99a2dd9da6d2bca01ba

Observation 9afeee92-4d2a-42f6-8a53-c48a442c7e24 · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.244822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.244822Z digest=sha256:36356902487bdbba276cefedf9034e8e0fbb5c676f87e0e49db82e02257c45c5

Observation 0cdb15c5-516d-4ae5-9db1-7d8b4c6375d9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.351081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.351081Z digest=sha256:36053147f1667a52f6c3472478a58615ede5e47b9d1110faff6628416f96c7fe

Observation f9b95562-1d14-488c-9cf1-2d86e06e9841 · outbound

This paper cites Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.418344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.418344Z digest=sha256:68777046c52bd937d762cff21d1b837be6a8f16a2a2060b2c9eaefc64f02d771

Observation 079fbce1-db73-41f1-a41f-886d28c166c2 · outbound

This paper cites an unresolved cited work.

Collaboration among Multiple Large Language Models for Medical Question Answering Unresolved cited work

Reference 30

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.821707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:08.513674Z digest=sha256:5037abbff905e0e95037bba731e3b5bcc9d1d6b526145a00e7b9c7d21259551d

Observation e20efab2-1a1e-4050-935d-c6177babd4b7 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Collaboration among Multiple Large Language Models for Medical Question Answering Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.628751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.628751Z digest=sha256:ab033c41125cc3f402e4d2e4ae11253312c1a1f4632d15b568f6abac04524c40

Observation bae281e0-84d8-4ae1-ae98-baab67d6add2 · outbound

This paper cites Available: https://mededu.jmir.org/2023/1/e46599.

Collaboration among Multiple Large Language Models for Medical Question Answering Available: https://mededu.jmir.org/2023/1/e46599

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.653097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:02:06.462338Z digest=sha256:afb479cff7b3e5e9c4bd6909599136a0dca8539fb3e3d4767f19d9e7129aa5d2

Observation b297728b-67e7-4b33-8e62-14aaf4bb33f1 · outbound

This paper cites Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.316036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.316036Z digest=sha256:5b27daa8126a81c21d8d689b7330628913416a37c02b3cc95e47a7ca3f336f74

Pith citing papers

No inbound Pith citation observations are available.