Pith. sign in

Paper Citation Record · LEDGER

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 10 inbound Pith citation observations for arXiv:2505.17952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17952 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:40.334058Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:44:27.022437Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.448127Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddff6edb-6561-415f-a0be-7db1e94857ea · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.110924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.110924Z digest=sha256:00fb0ccf116a818a02580b576c8018551db1fe44f450e604c64b1d4fe22089d8

Observation 44390b50-5e53-4f62-a2df-4932ae45cb20 · outbound

This paper cites Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.205646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.205646Z digest=sha256:d6feffdd7bbf4bfe6f00741651ffdaba18f6eac7319d827787e2df18fcc41c23

Observation 49b20d09-7281-4d9a-ae42-599be00232fb · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.294119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.294119Z digest=sha256:7872c0c25d852d054daf3e4cc4e21531803212c3988de1d6ea80e657d0bbf0ca

Observation fc8a92c0-c574-4791-b7c8-9f36ba96a05b · outbound

This paper cites Deliberative alignment: Reasoning enables safer language models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Deliberative alignment: Reasoning enables safer language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.430945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:35.381361Z digest=sha256:7e72f330905882eb17b94899a0a8780ff16d0535d0d981508365ae5d97227556

Observation 270b00cb-7f12-4d70-b0c7-ff752bffd7fe · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Capabilities of Gemini Models in Medicine

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.475924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.475924Z digest=sha256:38f89c6177ff3ee1864d0deea497e25fa8ab2358f6ed75830429fe8b9581bfc9

Observation 9196e662-694a-46e4-b70f-c0b95d6fbabb · outbound

This paper cites CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.566459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.566459Z digest=sha256:ed0fa400e563376e3ce2aa16dc59d6fbda7c4cdaa352d6028c7918a3da03d620

Observation 8f56f633-7dac-414b-9fb7-4dba5eda8c96 · outbound

This paper cites Thinking and reasoning in medicine,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Thinking and reasoning in medicine,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.290576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:35.658180Z digest=sha256:6519fc4fcb3da4e695c7c0a944867f2145108ff77b2ea14a598a4e4380588aef

Observation dbb93beb-e864-4aac-a750-6ab2d4386646 · outbound

This paper cites Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Towards Next-Generation Medical Agent: How o1 is Reshaping Decision-Making in Medical Scenarios

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.780810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.780810Z digest=sha256:e8d786f9b1102a4409322052ffa0b4cae18956b3616305b55f89919873c2790d

Observation ccd0b47c-5ec5-4198-b2b7-dfcd303abc82 · outbound

This paper cites Openai o1-preview vs. chatgpt in healthcare: A new frontier in medical ai reasoning,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Openai o1-preview vs. chatgpt in healthcare: A new frontier in medical ai reasoning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.134395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:35.861039Z digest=sha256:16465ef1d272f01f30e2b8b2c3031706e60a99cb29fb3b4577aa5e1d483d1d85

Observation 9fb96894-1c47-4a7a-af57-6915fbf06936 · outbound

This paper cites A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:35.953046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:35.953046Z digest=sha256:a2b6dd66d30b63b48b2438739d20ecbf1522d8c785368b49dbafba79c1cd781a

Observation d1f660c4-a146-46db-9daa-c1e61bf84ea2 · outbound

This paper cites HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.033970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.033970Z digest=sha256:3376d7dbc8e1aef6208ee6643d791c7d43a1aaf7532ad9b71253825db2c24a7e

Observation dfb8cb23-fa89-4910-84ff-5fa8cc3b26fb · outbound

This paper cites Chain- of-thought prompting elicits reasoning in large language models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Chain- of-thought prompting elicits reasoning in large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:43.008370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:36.140824Z digest=sha256:e91529662a7ee6b52e4dc3ca603187c1280869af700e1445ea28cd332d8cdbaa

Observation e5a910c0-7061-4daf-85c7-5fe14a3228a8 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Scaling Instruction-Finetuned Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.223757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.223757Z digest=sha256:aac117b94490d2ee34f16dfae2e11e4a75fa6f1f0e9fb7aaed9ae435719cd4c8

Observation e7ea9409-fc35-482b-ae14-effe08d21422 · outbound

This paper cites Star: Bootstrapping reasoning with reasoning,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Star: Bootstrapping reasoning with reasoning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.838945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:36.292069Z digest=sha256:7aae59611db5b28939c2419a1e0bcb22ec948fad77df84c35698a7b9b0540bef

Observation d33fe216-0b70-4e32-a9cb-71d7dc42abfc · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.371201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.371201Z digest=sha256:48ef3f2fda43b929c3285af4479ac6cd85c9df6c7e90a429ce530fe3a6600411

Observation 8d8e1ea7-7979-4aa0-bee8-ec6cc2b88527 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.475828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.475828Z digest=sha256:8677ee8206529df57be5d8dae33456bad771c479cec1b6902467ddd91292dd00

Observation 957b55df-daf8-471d-a8ba-095d21816c1b · outbound

This paper cites Training language models to follow instructions with human feedback,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Training language models to follow instructions with human feedback,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.540075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.540075Z digest=sha256:5139649b65ed2a274c49e26028ae212b251b2bd0b716bd4da11a8d6e117cf483

Observation 8a6bfc6b-3108-4773-bdd1-0cbed13750a6 · outbound

This paper cites To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.635299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.635299Z digest=sha256:98e7b2a4f70a0e9acfa1c765a6ad41b6eaf181ff823f5fe33aa24d31e59ba2b6

Observation 1fbcc017-d600-47fc-b76d-df04a5fa541f · outbound

This paper cites Direct preference opti- mization: Your language model is secretly a reward model,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Direct preference opti- mization: Your language model is secretly a reward model,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.662548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:36.701546Z digest=sha256:ec99b7b6360d09ed546d96658738b0f20298861c45bed9d167aa0374d25dcbc0

Observation e386da9a-1b0f-43f7-94ec-e59668369e62 · outbound

This paper cites Openbiollm-70b: Advancing open-source biomedical llms with direct preference optimization,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Openbiollm-70b: Advancing open-source biomedical llms with direct preference optimization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.481327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:36.828546Z digest=sha256:26564bd6ac3aef052bf039680f2e32d4be22ce2c7cedd3f4642aae0992da181d

Observation 1754908f-7f39-47ad-b414-3c26bd069522 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:36.929289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:36.929289Z digest=sha256:ed38b64f8eefacd8a282b0570b7bb24bcdd526d569bf86a98392fef3b99b67bc

Observation c80f83ec-8b68-4e3d-a607-b2e950a5a54d · outbound

This paper cites ACECODER: Acing Coder RL via Automated Test-Case Synthesis.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL ACECODER: Acing Coder RL via Automated Test-Case Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.015968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.015968Z digest=sha256:205c60e90d3d1bbf63262de1fde07133fb95a163498ab8103bca9efc5c4fb108

Observation f99ccc4a-99b6-43ff-8516-31e3ba326ad1 · outbound

This paper cites MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.084607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.084607Z digest=sha256:aa821ec4477163fe18fa0d85a8c1e73e39c225f0e2473f440d5ca9a512dd2185

Observation 29a452ce-3036-4ade-a2c8-bbd105af3e82 · outbound

This paper cites HuatuoGPT, towards Taming Language Model to Be a Doctor.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HuatuoGPT, towards Taming Language Model to Be a Doctor

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.171031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.171031Z digest=sha256:ad524c603c0f4c9ffc4808e5c0c4c78843965f1588f500a22d71bb8128d38ef3

Observation e8570afb-a017-4443-a685-cdaceee4c015 · outbound

This paper cites BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.248353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.248353Z digest=sha256:e4dcdc4720205ae2cc2a7e4f651ee21b15c61c69801532e0516ff8f15b052395

Observation 73ddad87-78bb-4a09-a9e0-731211447055 · outbound

This paper cites Continual Pre-training of Language Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Continual Pre-training of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.342138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.342138Z digest=sha256:e5851d75d216f1737c415426ee10d57ebecf7464c6fcf0fd54ff4f273d451af4

Observation 1473c761-1476-43af-8f1e-91f3ff449fd6 · outbound

This paper cites Ultramedical: Building specialized generalists in biomedicine,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Ultramedical: Building specialized generalists in biomedicine,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.431561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.431561Z digest=sha256:2369f5c040e74dae73c858a90d62bd55cfeaddb3a888f9dfa5f9f6a494d91dff

Observation 4d5da96f-8c4b-4684-96e5-9e805d0d8759 · outbound

This paper cites m1: Unleash the potential of test-time scaling for medical reasoning with large language models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL m1: Unleash the potential of test-time scaling for medical reasoning with large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.649916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.649916Z digest=sha256:ed9a77cea8d41b273ef262aff7d0647244abfa646a8e3333828a987fd6868924

Observation d6ac4e88-471d-4433-beca-d35c9dca5d08 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.741316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.741316Z digest=sha256:a6574fb90ed9e455c65bdbde2ff1d8efada5ea9404d13f7434d526e67f328023

Observation b651a1eb-6835-48f3-8950-f88774b3528c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HybridFlow: A Flexible and Efficient RLHF Framework

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:37.853487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:37.853487Z digest=sha256:dcceca3be4060137dbd7a7d929297202b0d7e58a950cfe7668ddb00a38767311

Observation 954f89b7-6c7e-4fb7-885e-eca57af0df40 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.307188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:38.014800Z digest=sha256:9e326faae28dbf0f790a511cdb886e8c9e2dc3b58272eb9ad6ae590f7eb54c81

Observation ea87c37c-55e7-4a5b-a18e-4fc1419b4629 · outbound

This paper cites PubMedQA: A Dataset for Biomedical Research Question Answering.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL PubMedQA: A Dataset for Biomedical Research Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.111618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.111618Z digest=sha256:44d34d6f410d7e8d5269ac49ffcd5592f3c74d85fe3c107559794ada59bdad36

Observation 33dbd54b-e09c-4eb2-8c58-cbe4bddab2b2 · outbound

This paper cites Mmlu- pro: A more robust and challenging multi-task language understanding benchmark,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Mmlu- pro: A more robust and challenging multi-task language understanding benchmark,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:42.105838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:38.215005Z digest=sha256:6e8fc915f09ecc3049e8523a245a16b59a16f11b3ab1a9a2bc3c72f462b685d8

Observation df4ce292-511b-4cf3-b664-3bdd37b2a5fa · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Gpqa: A graduate-level google-proof q&a benchmark,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.322847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.322847Z digest=sha256:c36de739bbe9f20470fdebd31da19b823e2340104fa32f34292e64b6754ce1e4

Observation f4b35547-716f-4666-acfb-e821b41c857e · outbound

This paper cites MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.423359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.423359Z digest=sha256:c7b724b8783cf5df6184a26b0567b0542eb9ecf881fdef702d9503c2c72a3939

Observation 2fc5f574-142a-4694-b38a-c57e3ae2a1fa · outbound

This paper cites MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.579027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.579027Z digest=sha256:36f8c5d4e31b41b5ba1030b4e05dde89b1b00cd38d8f04785f429fec7c5f615a

Observation d514a37b-32bc-4cfd-adef-3c560f2ef65c · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:38.693288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:38.693288Z digest=sha256:79562073fb7b35179c700fea0c2b125b1fa470ca803da58e3f7add42ef9d9943

Observation 497ac864-91eb-43b3-bf04-b1d08468b539 · outbound

This paper cites Openbiollms: Advancing open-source large language models for healthcare and life sciences,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Openbiollms: Advancing open-source large language models for healthcare and life sciences,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.905072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:38.796996Z digest=sha256:ba1aa35f6f6d395bd35c5ead03851d3ce68621c9d6d702b3a49a502a05582fae

Observation a95fa484-fc52-4f60-b28f-9f93522260ab · outbound

This paper cites Towards building multilin- gual language model for medicine,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Towards building multilin- gual language model for medicine,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.710949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:38.908872Z digest=sha256:8ac95221fb396d4f8208e90e784b21484f30d293ec4afa1521c0738237645363

Observation fee30329-026f-44a0-8b51-306c9d1155a5 · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.026377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.026377Z digest=sha256:b6bded45dc5a3ff2c4514956b5dda3c2bea1ed6604057da2677302cd457dafc8

Observation b77641c6-aac8-485d-908b-4405227b7c09 · outbound

This paper cites What disease does this patient have? a large-scale open domain question answering dataset from medical exams,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL What disease does this patient have? a large-scale open domain question answering dataset from medical exams,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.162856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.162856Z digest=sha256:ae3a2467ff6c755b1adbd0a28ba01bf5f1dc47c117d0608a22e842644ec12eef

Observation d4c9f1c4-fb21-43f8-a5c9-812ee0279875 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.514776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:39.278178Z digest=sha256:7cb87d82bb5a96fe1fa8658c0655e8105340784da17c1ddc8bdaadf81a49cfd5

Observation e1c5258e-34e5-48bf-8a75-e9254d4a3c30 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL The Llama 3 Herd of Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.409211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.409211Z digest=sha256:d7f5cf697b5676114613c6e74a233942e2b66e166ea8c079e217d3ad3a7a482d

Observation a0c4a48a-0fd9-4268-b15c-a59d446e6bc3 · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.559067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.559067Z digest=sha256:9da023f5aa2139ba9a3ddacd2cd9738995b8bf70ca644edf88dde4df4d02bab2

Observation c9a988c5-1a8d-48cb-b5e2-854e6aa9078f · outbound

This paper cites GPT-4o System Card.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL GPT-4o System Card

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:39.677799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:39.677799Z digest=sha256:9b5d67c457aa7c91d8e7af2dd3e25b2b868ed8423af6d94d505ae6cf3495ba2e

Observation 9c02b31f-638b-4072-b27f-f4f295cdc9f6 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Qwq: Reflect deeply on the boundaries of the unknown,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.352436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:39.815897Z digest=sha256:88eb0c6e7ae0e5344c914ad2e1c1dd3073ba8250ae5cae797cb7f22dd22dc75f

Observation 4906fd2b-fb4d-478b-a2fc-d306805a407f · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL The claude 3 model family: Opus, sonnet, haiku,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:41.155809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:39.930681Z digest=sha256:468e8f5666abfaea1d3798db2415d1b0bd2df4be6346e284c7c50bc9cf097cb3

Observation 85eec960-6abc-4426-84d3-6e39bbd053cb · outbound

This paper cites DeepSeek-V3 Technical Report.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL DeepSeek-V3 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.068205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:40.068205Z digest=sha256:48ef5e2203902a6ba0a94f9dadbbd8888da9f9f2d8210bcdd098ca09b181a076

Observation f0042886-c4f9-4cf6-a277-04c03afbcfbe · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:40.218826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:40.218826Z digest=sha256:0b597fd4b30024f72b923aeff66a64361f9f14066ec4113631395d9fbd31e791

Observation 08e58909-ad10-49a8-9531-4114269efcd3 · outbound

This paper cites Qwen2.5: A party of foundation models,.

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL Qwen2.5: A party of foundation models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:40.962287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:41:40.334058Z digest=sha256:8e951019dbae9df9e77fe81e6719c152656db575d30364219130aad06d81d290

Pith citing papers

Observation cd855766-5f9c-42ba-8f38-91dd09040779 · inbound

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning cites this paper.

Med-R$^3$: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T10:44:27.022437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:44:27.022437Z digest=sha256:9324d953962f03a6f0f400efe96b5b603c0edc3621ebbc5352f7a58a7ebed09d

Observation 16937fd7-60aa-4f84-b71d-a777a78a7948 · inbound

Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning cites this paper.

Psyche-R1: Towards Reliable Psychological LLMs through Unified Empathy, Expertise, and Reasoning Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:18:43.366828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:18:43.366828Z digest=sha256:31ba5fb5560863f461781c34dfc62ad1aa4047b767655f66fc26438573050fd7

Observation 281d21d1-8ff1-4947-b757-d5b19d5a9f99 · inbound

Baichuan-M2: Scaling Medical Capability with Large Verifier System cites this paper.

Baichuan-M2: Scaling Medical Capability with Large Verifier System Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T11:50:13.102408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:50:13.102408Z digest=sha256:7cc1fb19238a64ef2e81d32f6528e389c117768c7aba65ed56f44d1b29b22847

Observation cb402813-c124-4aaa-ad62-52804fb69d88 · inbound

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification cites this paper.

AdaThink-Med: Optimizing Inference-Time Compute for Medical Reasoning via Uncertainty Quantification Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:51.561726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:51.561726Z digest=sha256:f6c50c699b67a342b6d389ee4e758962c71547887ca8951eb88c9f99eba86308

Observation 4744b8c1-bbb1-4855-a935-21941b67f116 · inbound

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training cites this paper.

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:22:43.249504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:22:43.249504Z digest=sha256:68e9a8287fad44feb77cfa8c37da38dc2b083b2e19783bedaa09cdda9bd5feda

Observation 36baa877-01c2-443a-8b05-923db99d484f · inbound

Medical Reasoning with Large Language Models: A Survey and MR-Bench cites this paper.

Medical Reasoning with Large Language Models: A Survey and MR-Bench Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:25:26.786876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T10:21:39.892271Z digest=sha256:866c468be27c3aa46dee993e495a3ee7d23bdd2767e6f0252b927d50fa95ce4d

Observation 5bcc3865-1977-42f6-9ce2-32d32e606e1e · inbound

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve cites this paper.

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:34.013256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T12:35:32.915685Z digest=sha256:b5e83279dc49ad72bc2689dd274f558a19f6b7e5414c68abddbcd44fb9bfa8fb

Observation 2cf0a86c-9e05-45f4-8702-dc3b3be5d501 · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:23.869943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:ce9f037fa4b4a8e88ac5cba5915de7a071565252471c6859ce8122b9ed9d1487

Observation d0d8dc84-b310-48dd-bfd9-68ccac497add · inbound

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA cites this paper.

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 186

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.449550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T10:21:12.782864Z digest=sha256:5c5c5fc5f4a96c553994e09e7fb6ba45e3f55598ec8c60675e41adddb59bf451

Observation 5b87cf8a-7fa2-4afb-9197-dc5a8f416f8b · inbound

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? cites this paper.

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Reference 155

Resolution
unresolved
no resolver link, observed 2026-07-30T16:01:42.470884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T16:01:42.470884Z digest=sha256:a357e323a87719f5ebb3934c50911aba4d20fce499e1eb6b75ef880ce00fb298