Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:54.555335Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 8 inbound Pith citation observations for arXiv:2505.14625.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:54.555335Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:43:00.923422Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T02:16:26.183764Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 56aaa975-7b28-424b-8d22-68df808f6885 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa829132-5b02-4f00-813a-89c78d605902 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb3a90b-efdb-4823-a7f5-f85a8f59ef02 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Robotxr1: Enabling embodied robotic intelligence on large language models through closed-loop reinforcement learning, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 93031f66-d74d-4699-b9ff-c2290a75f510 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning xverify: Efficient answer verifier for reasoning model evaluations, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d5a7cf2f-c2a0-414e-b39f-ef90567747a0 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a93a0685-831e-4c90-9a01-795280ab935d · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9bafb76-c51d-45d4-a77c-563882fa8f9b · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Process Reinforcement through Implicit Rewards
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668db233-568d-4c5e-b3dc-0008c8606ebf · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b9eaf3c8-7c5f-47db-8983-e4f81e3053ce · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Raft: Reward ranked finetuning for generative foundation model alignment, 2023
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106e2374-3c76-46a7-a2d8-63ff106d590a · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04a842c-6fb0-46ca-8e4e-5fb1c67cd60f · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning The language model evaluation harness, 07 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc96dede-980f-400e-a3bb-297c8eed8dbe · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning A survey on llm-as-a-judge, 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c20689a6-cd8b-46be-93a5-e00c5012ee25 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6eebf7f-88ce-4155-a7f8-24bba27364bf · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de83242f-d66c-4d96-b460-91f78ea5548c · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Ultraeval: A lightweight platform for flexible and comprehensive evaluation for llms, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 78210343-cbd8-436f-822d-24ac91b47ace · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b835f02-d85b-4a76-97db-17fb968707ef · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Putting rl back in rlhf.https://huggingface
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9218d86e-64e4-4a9a-bb18-ab94f77c0901 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Math-Verify: A robust mathematical expression evaluation system.https: //github.com/huggingface/Math-Verify, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5eab1cc-2a3e-4122-98b2-6576af50da68 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OpenAI o1 System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30b3d3b-ea42-4ad3-85a2-3bff7c4554b5 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7f196a-3c6e-4d18-8406-31549108e161 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e549583-4066-48f7-ae5a-391be89490b8 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning From generation to judgment: Opportunities and challenges of llm-as-a-judge, 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 13bb5dbb-8698-4c18-b1f0-1141c77c40d6 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4aad93c-cd18-434c-a46d-a22d024f05c5 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Hashimoto
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 415006a0-c6c1-4439-971e-a4c733be84f6 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42cc20b7-78f1-4cc5-8994-39500db0b961 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18d1b64-2272-46c8-b00f-40b5de843138 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning General- reasoner: Advancing llm reasoning across all domains, 2025
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7f77a618-2304-42ac-a8eb-dff87f7c1c3d · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification, 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation caee832f-c2f3-47a8-800d-a9ac97acbc47 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning s1: Simple test-time scaling, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6203986-d348-4610-8853-308b98f5ad0d · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OpenAI Evals: A framework for evaluating llms.https://github.com/openai/ evals, 2025
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a4639574-c2f7-41c9-b344-6d1c7e8c1cec · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Manning, and Chelsea Finn
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3774910b-089f-4701-b8b2-f7d78db42b96 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Proximal policy optimization algorithms, 2017
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d82b86b-711b-4e51-8e57-d8c78ff1ee60 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e882606-ace1-4a33-81e1-dc723e8a65e6 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Hybridflow: A flexible and efficient rlhf framework
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acda1c18-7e8d-4349-b20a-cbfe59c44214 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a579662-71a9-4b86-abed-e584b3ee26b8 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988bea9c-818f-4c1a-9c72-eddf547bbe1f · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f4181c-2265-483b-8555-2bfc748af162 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb1a502-39cc-45e2-8769-3041624abbf4 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31866f05-745b-428c-85b5-d350a798a339 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Limo: Less is more for reasoning, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c960e078-d3cb-4cf8-a455-d4dc2fb840e4 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78eea0f-e438-4e9a-bad1-a1df752bc050 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70534c50-af02-4c81-ae12-951af34c4ea8 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65631411-b382-47d5-9dff-0ec989593bd8 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d0567c8-dec6-4884-aa07-53b613d8ee30 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning LLM as a judge
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9bed4c2a-33a8-443c-b0df-f3edfdac91f6 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a1be238-ee3e-4a3c-a577-be0e0032c2c9 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1413de55-c728-4406-9901-d4215721e8a8 · outbound
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3e48e768-8302-413e-96bd-f89841ee7373 · inbound
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 42f34f68-8adb-4ba9-baf8-69569e51f279 · inbound
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c04c54cf-8883-403e-b34a-b417301c2bce · inbound
High-Dimensional Statistics: Reflections on Progress and Open Problems TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 84e5ac04-f9fb-47ee-a91c-3ba5aa9b5650 · inbound
High-Dimensional Statistics: Reflections on Progress and Open Problems TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ae651858-8e70-472e-9370-c824e5cd6fd5 · inbound
Trust Region On-Policy Distillation TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 607992f7-4e8e-4307-abe4-9675df8e705b · inbound
Trading Human Curation for Synthetic Augmentation in RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4b0a04ab-7d48-4726-979c-99ad29cd69c5 · inbound
Trading Human Curation for Synthetic Augmentation in RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb0a385-ed47-474e-88e8-9680e753f793 · inbound
Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.