Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:36.509160Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 5 inbound Pith citation observations for arXiv:2505.19501.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:36.509160Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:16:36.029822Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
50 of 50 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation afea2159-1a30-42ab-b491-9dbae7c6f438 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning https://groups
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98928a95-ae8b-49ba-95c7-d9675e3662e6 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Publicly Available Clinical BERT Embeddings
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5f7560-c4fe-4f4e-b909-30935d26a96c · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning SciBERT: A Pretrained Language Model for Scientific Text
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95578303-0a4e-43f2-8890-0b52b5acb70c · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b22f4d-898f-4f46-9fba-5bcad3ffbc6f · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning MaxMin-RLHF: Alignment with Diverse Human Preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a7ae87-6a9a-46b3-ac06-a3b7d5c2023c · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75646ddd-da90-4e10-8468-4bf1706831ad · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A 5′ utr language model for decoding untranslated regions of mrna and function predictions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1918dbac-c09e-49cf-bc17-08613e799728 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ef9dc2-1148-494b-bc6c-906cf4dcb15e · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Temporal consistency for llm reasoning process error identification
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ff156e0-1d28-4814-a7ca-3c65c51b2479 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Embodied LLM Agents Learn to Cooperate in Organized Teams
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64f9b7dc-1b6b-405d-833b-2162ea59247f · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Math-perturb: Benchmarking llms’ math reasoning abilities against hard perturbations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24d7bf2b-47c5-4494-b0a8-a3bbcfec2ef8 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Crispr-gpt: An llm agent for automated design of gene-editing experiments
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac6f6e28-6b99-4c06-a28e-2d2507c3c7ea · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A Survey on Large Language Models for Code Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6eafb56-a2a7-4cfc-8e9c-54abbfa772b6 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f33a9524-a96c-4116-884a-000a202120a4 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cdd6523-8ddc-4899-b1e9-b1baf5e61344 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Bioasq-qa: A manually curated corpus for biomedical question answering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7fa524b-9938-4b67-87b6-074d9acbade0 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b06293d-b9af-4ed1-ae10-8ded4e7c4f18 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a98f1ac-515c-4c98-8df5-36f402dc2723 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77bbd0b2-17ea-4458-9a67-bf64eef549bc · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Biobert: a pre-trained biomedical language representation model for biomedical text mining
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a6037e-1978-4a79-95c8-48a164660ba6 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Mapping the Increasing Use of LLMs in Scientific Papers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33172ea8-d605-4bde-a5d1-e304650e5cb8 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Understanding R1-Zero-Like Training: A Critical Perspective
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8694ac4-f500-4491-96bc-cfff4e84cfcf · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Biogpt: generative pre-trained transformer for biomedical text generation and mining
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61c29d7e-f8ba-48c3-9696-7523461a9b8a · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9d5b13-ca7d-4bd3-86df-09cefabe942b · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Covid-qa: A question answering dataset for covid-19
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 212554fe-81aa-4ae6-826e-3be54a7976f7 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Training language models to follow instructions with human feedback
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec343161-45bc-43cb-ba87-66f2b60092bd · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d25e6f-a006-435a-be4b-15f8fd6de7a2 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A domain- specific next-generation large language model (llm) or chatgpt is required for biomedical engineering and research
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1153b7b9-16d3-4381-8ba0-ecd68bac0d78 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Humanity's Last Exam
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f1273f-6232-412d-8f21-579df929b853 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning SciFive: a text-to-text transformer model for biomedical literature
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398dcbeb-5310-4dbe-bf77-71a92d4be9d6 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f58c65-b194-47b9-bdd7-73f99773a326 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Gpqa: A graduate-level google-proof q&a benchmark
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1dbd3a9-b21d-4d71-ab50-a65683f5700f · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53a1d2b6-08e4-4829-8fe3-83f970070603 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0279c49-b79e-457c-9f38-f64049913834 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d81dc0-19e9-4631-953a-61b9f0b2d066 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11ccb532-811a-4fa2-a7c2-9f49545af144 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Galactica: A Large Language Model for Science
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e70ae3-52f4-48ff-9b2c-d3a0eb1a415b · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A call for built-in biosecurity safeguards for generative ai tools
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8f5076-31cb-4081-a8a1-dd64cc10cea3 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f645668-9c8f-4882-b23b-0fe5ee70045a · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Pairwise Proximal Policy Optimization: Harnessing Relative Feedback for LLM Alignment
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5447346-bd0c-432b-8df5-a6d32676b0b1 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DARWIN Series: Domain Specific Large Language Models for Natural Science
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc56ceb1-ca1e-45ab-a9e7-534e7fc3ea92 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc33c128-912a-46f2-b9d8-2183cc8c76cf · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092e32ff-81b6-4a4c-b27c-8426ce22c4ec · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54ca4a9-5f34-48a3-bb04-952e8cce4afc · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53ce49d-145b-4e6a-ac43-cc76e727053d · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e82881-5bed-4ba4-b91c-2b451aba0c5c · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning BioBART: Pretraining and Evaluation of A Biomedical Generative Language Model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c79a4eb-c102-4f33-bc36-b20ffc301b21 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A generalist vision–language foundation model for diverse biomedical tasks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f38840b1-f1c2-48c1-a399-4e3266f637bd · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Scientific large language models: A survey on biological & chemical domains
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3bbb883-9740-4886-bf2b-91166959f7c4 · outbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54ca4a9-5f34-48a3-bb04-952e8cce4afc · inbound
Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061e31b1-ba24-4b77-b94f-eef42f8dd929 · inbound
Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 836537d0-3232-4b45-9326-727cb0c321c3 · inbound
Evaluating Large Language Models in Scientific Discovery Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 35dc85c7-1183-4668-ac4d-9ddacfbfa3d3 · inbound
Heterogeneous Scientific Foundation Model Collaboration Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca5363fa-0429-4cc9-82e4-98d5859b14cb · inbound
How Post-Training Shapes Biological Reasoning Models Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.