Pith. sign in

Paper Citation Record · LEDGER

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2504.18080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18080 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:28:11.158639Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:32:16.930291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T06:11:23.712115Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73b2b576-4d87-421c-b95f-bf107ef1ed77 · outbound

This paper cites Superhuman performance of a large language model on the reasoning tasks of a physician.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Superhuman performance of a large language model on the reasoning tasks of a physician

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.006071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.006071Z digest=sha256:c94c2b337b57ed9d2fa1572381c78e3a2cbb8e81c278b73b04130a4b01128aea

Observation 98110f26-a7fc-47bf-a384-64c93f1bf09c · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.010868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.010868Z digest=sha256:8e82443d9a017d77afa4b4ecde2b1030fde95463810c946f49031be8b65bf4ae

Observation 73d05cbc-4d28-44e1-a33b-8e24486b64ac · outbound

This paper cites MEDITRON-70B: Scaling Medical Pretraining for Large Language Models.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.015505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.015505Z digest=sha256:28a69db2c5c4ded6e726632a100d7e55f368f130ea676aaebd1a6b6d868d2fa2

Observation 10d34416-f6c7-421b-b9d6-3a57d230530a · outbound

This paper cites Beyond fine-tuning: Unleashing the potential of continuous pretraining for clinical LLM s.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Beyond fine-tuning: Unleashing the potential of continuous pretraining for clinical LLM s

Reference 4

Resolution
verified exact
doi, observed 2026-08-16T10:28:11.447147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:28:11.019658Z digest=sha256:aeae3fb1928f01f8a905bf7efcd6ab85ebc67353f344d365904221a4204cce23

Observation 24ed6c3b-ce00-4dd1-99df-09b563b1a13d · outbound

This paper cites Towards a Personal Health Large Language Model.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Towards a Personal Health Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.023518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.023518Z digest=sha256:df9b3330d5125630bfb64beafb2ed9080cea0a1021cc53d7a92395a21620f8f4

Observation 6092b268-ad98-440c-9cee-d7759edc8355 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.027707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.027707Z digest=sha256:3cfffa6bd449bef5b0447d40b93a4302a38131b698176f356387b863deeea395

Observation 457ee354-005e-4c35-a054-5570fb6d34c9 · outbound

This paper cites The case for 4-bit precision: k-bit Inference Scaling Laws.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization The case for 4-bit precision: k-bit Inference Scaling Laws

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.032102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.032102Z digest=sha256:968c58aef6701c57672182b3717fd703cbd176b5339d46fedd0fef501a234f08

Observation 0c58a6b2-b601-4b17-87fb-29f5ff648528 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Qlora: Efficient finetuning of quantized llms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.035791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.035791Z digest=sha256:93d6c29ba831d2b588718dbe402ec1e018da97e41ceefd28ab8a1595f5ed4e09

Observation 917fb95d-8361-484a-a09e-1918df96bef9 · outbound

This paper cites Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.039320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.039320Z digest=sha256:721ac52d5d9569de70f20f9fce619845c5dcbca02c479adf02f3790870ffd78c

Observation 19c99389-363b-4da0-ad6c-b4fc0e67369e · outbound

This paper cites The Llama 3 Herd of Models.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.043167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.043167Z digest=sha256:cdd9f8a7fa20d6a2700be727814f42ba62702da098751665cc2c27f691f8aecf

Observation 81d060c2-dbc1-4d9e-8c77-a889a1d130e7 · outbound

This paper cites an unresolved cited work.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.047103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.047103Z digest=sha256:bf607707ece5e56db98af6b3dc7d978b6261348faf71c61e829b2ab826ef4d7e

Observation 3060ebd3-1cd8-4034-b2f2-352701584d4b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Measuring Massive Multitask Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.050549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.050549Z digest=sha256:906b79d702b51cb02b95ddfd3951a542fd1931976b840fea68e1d49c4cca0fe9

Observation 5c921b51-133c-41f9-a97d-a26f3cec6eab · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.054338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.054338Z digest=sha256:1b425aa8f93964e0ef3204a2e127eb005bead59ed84e0c9995c8907cb8c833ce

Observation 5c261eda-d21e-4720-adf2-46af41367546 · outbound

This paper cites JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.057854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.057854Z digest=sha256:0c8e8d87f105709cb8aab98e5260d94be12bf7f3d76de4f574ccdf2582d6cd7f

Observation 4eaf16e5-6cce-4550-b987-572c5adf0e95 · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.061879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.061879Z digest=sha256:fabd3e4a370010f50e58ed29f6eeb82bbf219b9b58ec6d9b88e52e0ffb9c8a41

Observation ca91b0bb-b531-42d0-bbad-1c23b7753a27 · outbound

This paper cites P ub M ed QA : A dataset for biomedical research question answering.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization P ub M ed QA : A dataset for biomedical research question answering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.065840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.065840Z digest=sha256:e7d02a56cd72dd0713e013eb794aed94f7f1d39b3880ad64b6d7c9b5f3cac5af

Observation cb7c69c1-495e-4a23-b7b0-3af71d1f3876 · outbound

This paper cites An evaluation framework for clinical use of large language models in patient interaction tasks.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization An evaluation framework for clinical use of large language models in patient interaction tasks

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:11.627980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:28:11.069902Z digest=sha256:99ab4c03d9fa091da926be0d88fc5d7d92e5c423eccaed910617ba427eca8425

Observation 48b23760-3ce1-4fc0-8875-80697cb08a23 · outbound

This paper cites MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.073411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.073411Z digest=sha256:206b5f6131e6f2e8a0c73be87622ef944d5942ea7c836237e68da2641387dfe6

Observation 6b748f13-a5e6-4a0c-ba0e-904bb31ace17 · outbound

This paper cites Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.077396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.077396Z digest=sha256:b6da432e35f5c89fe37fb6770f4d5c7ca2d88722ec876bc88ac8da367bf2b95d

Observation 73cb7e17-8456-4916-863a-c19e0428d06a · outbound

This paper cites Medical Hallucinations in Foundation Models and Their Impact on Healthcare.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Medical Hallucinations in Foundation Models and Their Impact on Healthcare

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.081154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.081154Z digest=sha256:734aabc9e14b3db02e15c2c2ab8e49bc164c486eb1b2d8aa99353b9f434a4500

Observation f15475b9-a3b8-44c3-a75f-052992fea374 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization ReFT: Reasoning with Reinforced Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.084518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.084518Z digest=sha256:b148b03031c53ea95c75f2e36ffefe57e2d3aefaaa652517fbb31bfc1501f9a5

Observation 98a9ab13-58b5-4882-b961-1ad3d3c2f96c · outbound

This paper cites Radlink: Linking clinical entities from radiology reports.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Radlink: Linking clinical entities from radiology reports

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:11.616832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:28:11.088357Z digest=sha256:0a229531a4b07801e4fa213bb2f0bb0679c8f18bb82ca0bd4fc354bea20eb19d

Observation b1b65886-20a3-42e7-9f14-eed8dcb70513 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.091968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.091968Z digest=sha256:373e594cec67bae08e904c02ad3161cc11f95692367c601b7cfbcfa1fdf87aea

Observation 5f8dc36a-7d12-4725-9937-29c4223836e0 · outbound

This paper cites Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.095769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.095769Z digest=sha256:6e7224ec7d67538cd2a096bbc879856bb3309e3a0dc522c8059a008bcdfadf7b

Observation 31daaaa7-af17-491b-a9cb-9bf932f12744 · outbound

This paper cites From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.099070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.099070Z digest=sha256:c226511b1b8b553fd421860fbb15f80db285f3f2edc14a742a7dad8324bbaf50

Observation 72b9065e-cbe2-47b3-a690-5043ba97eb17 · outbound

This paper cites openai/MMMLU , 2024.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization openai/MMMLU , 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:11.606432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:28:11.102529Z digest=sha256:37a7bdd6f9800affda0440b057c39c676cc374a6976309ec3529cbe21b7874c3

Observation 8105fb90-7260-4de8-bf8d-5d87853ef6f3 · outbound

This paper cites GPT-4o System Card.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization GPT-4o System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.107627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.107627Z digest=sha256:77a82d7e77583ac0058866b16c62ad12cb4a28f33d5de366d4ab2187ffafce9c

Observation 8e88db10-a75a-416c-959b-26fb15c1070b · outbound

This paper cites Training language models to follow instructions with human feedback.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.111525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.111525Z digest=sha256:8f3f09195545bfc51778243d951403f643f08246548235bc28ff602b42d85ddc

Observation 90ebc75e-2d21-4aa1-ab54-f6d02fbe2b3f · outbound

This paper cites MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.115101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.115101Z digest=sha256:be02b7069c2d26832e6879328f45d255ca14564563cf320f57942e0454a34403

Observation a12ed395-812b-4182-b83d-ddb8330022a0 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Iterative Reasoning Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.118774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.118774Z digest=sha256:f001588eb88fd22750e991ffb946d7be07bcf8debafad2e9e14158406658949b

Observation b1e73f53-7ab7-4ef8-9e3a-f2f32623bc35 · outbound

This paper cites PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:28:11.225598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:28:11.122594Z digest=sha256:529befedd7112d8ebb13d52a3d451c04597c059e36720ed45d09658e02673439

Observation 76b0664a-0ccf-48f5-9665-455016c42998 · outbound

This paper cites Qwen2.5 Technical Report.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.126407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.126407Z digest=sha256:63c91721a76436ff8a36555749508dcdbcf85ac25b05afd150ffd795fb5b63a6

Observation 77f88661-dd29-4745-9d94-a600beb808dc · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.129935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.129935Z digest=sha256:c9e1a46df2185721edc05debbecd7553377c57d2e45e1b2985850ccadc1570eb

Observation 1bd2fa28-c7df-427d-84d2-3e52dcf5aabf · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Capabilities of Gemini Models in Medicine

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.133508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.133508Z digest=sha256:2d72bd24cd020ef380bdd98e0de708b8ad9a7c2782c9aca279ae8c1c8db4ea8f

Observation 523c3ae4-07c3-431e-84ca-17c0ca95756d · outbound

This paper cites Toward expert-level medical question answering with large language models.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Toward expert-level medical question answering with large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.137385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.137385Z digest=sha256:a201ca734a779fc9f3564be1f2cc257ee2a47a28c707b45f697e4620dff3797e

Observation 2a38297a-6470-4172-8e5f-cea04e9e3511 · outbound

This paper cites Development and bilingual evaluation of Japanese medical large language model within reasonably low computational resources.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Development and bilingual evaluation of Japanese medical large language model within reasonably low computational resources

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.140703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.140703Z digest=sha256:96460347dd0f52cfa35f5810aaf23c1469d804b2f63f43e3a3a70bc8c69bc8a1

Observation 8f7c7534-552b-4cc9-b1e7-63a1b779ac28 · outbound

This paper cites 70B-parameter large language models in Japanese medical question-answering.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization 70B-parameter large language models in Japanese medical question-answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.144236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.144236Z digest=sha256:2766884ec6d17456b5cf21bf2a5bc019cb3a1b6e31d026a6ec0ce9ff4b4eaf37

Observation eb5034e7-d180-48f5-88a7-8867bba2d594 · outbound

This paper cites Towards conversational diagnostic artificial intelligence.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Towards conversational diagnostic artificial intelligence

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:28:11.582356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:28:11.148130Z digest=sha256:a1e18f3749117939cb917d1a53dd9409a063b9ff5600d55e975ac10e045b58c7

Observation 7becc1bc-dc26-40de-92e7-45a1ff2c6117 · outbound

This paper cites Adapted large language models can outperform medical experts in clinical text summarization.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization Adapted large language models can outperform medical experts in clinical text summarization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.151512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.151512Z digest=sha256:ac08335edf2f385071cf6e848f8c39afdb4b88bd870b65c4e9773913e8fce710

Observation c348976a-fe2c-4d1f-9b59-55dc43f7fabc · outbound

This paper cites A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.154915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.154915Z digest=sha256:9868232b2b3cd80e531e7fa711dc0252dc57cf1b1c81502cf0ae17c47670522b

Observation 5da5b1f8-702f-43bc-a4f8-3ac543734e9a · outbound

This paper cites HuatuoGPT, towards Taming Language Model to Be a Doctor.

Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization HuatuoGPT, towards Taming Language Model to Be a Doctor

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T10:28:11.158639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:28:11.158639Z digest=sha256:fc9f09ecb5b95751753a6d1f932c4d900116f71bbf07c40ad3c6a92218131844

Pith citing papers

Observation 6a795170-6248-4603-bd60-8e1386c46c69 · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:23.714985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:6737f5346d2e843a045aa6d937196f094b8bb0ce2dae6ab67bbaf9f4290decdc