Pith. sign in

Paper Citation Record · LEDGER

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 8 inbound Pith citation observations for arXiv:2505.14625.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14625 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:54.555335Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:43:00.923422Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:26.183764Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56aaa975-7b28-424b-8d22-68df808f6885 · outbound

This paper cites Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.050333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.050333Z digest=sha256:b803df37efafcea6ce2560b960647e0e371fec3eb075239828793cb4160b788f

Observation aa829132-5b02-4f00-813a-89c78d605902 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.101516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.101516Z digest=sha256:383a42dafc2a093becf156578791a0354f20b0767182002edc658f7261c3736f

Observation 7eb3a90b-efdb-4823-a7f5-f85a8f59ef02 · outbound

This paper cites Robotxr1: Enabling embodied robotic intelligence on large language models through closed-loop reinforcement learning, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Robotxr1: Enabling embodied robotic intelligence on large language models through closed-loop reinforcement learning, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:59.677826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:49.161442Z digest=sha256:dd1abfbc382f364e095b9fcc2f8e7d7b916c172b4bb1201deecceeacc8dd846c

Observation 93031f66-d74d-4699-b9ff-c2290a75f510 · outbound

This paper cites xverify: Efficient answer verifier for reasoning model evaluations, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning xverify: Efficient answer verifier for reasoning model evaluations, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:59.384896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:49.231062Z digest=sha256:f2b687e3df03fde5bb29ea5f3a6b1699a6cf8f65b8496031c9f5aad8ff6dbc48

Observation d5a7cf2f-c2a0-414e-b39f-ef90567747a0 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.317584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.317584Z digest=sha256:6b5351adc2e4fc8f95c4c46b9aaef0bfb24ec287e13fe1c64c36cc8f3d7eb060

Observation a93a0685-831e-4c90-9a01-795280ab935d · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.383928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.383928Z digest=sha256:a9d272f61c418d6aff9a7ab9442c960a9867e659966a05818791e386444103ae

Observation a9bafb76-c51d-45d4-a77c-563882fa8f9b · outbound

This paper cites Process Reinforcement through Implicit Rewards.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Process Reinforcement through Implicit Rewards

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.435775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.435775Z digest=sha256:d4ce8384fb572a65ef11f41decb01a57cd005eb919dd6a1b89184eb7eefcde66

Observation 668db233-568d-4c5e-b3dc-0008c8606ebf · outbound

This paper cites an unresolved cited work.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:59.080504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:49.535231Z digest=sha256:82c8aa33ba9c6d6cb1fd6fe71bd8b43f9b2789511d08b6795ebb78b37aed4f84

Observation b9eaf3c8-7c5f-47db-8983-e4f81e3053ce · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment, 2023.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Raft: Reward ranked finetuning for generative foundation model alignment, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.637551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.637551Z digest=sha256:7170dc819a4bdd7b76b184771b9f674722e6954e52d93dfa52858d3782c5eaaa

Observation 106e2374-3c76-46a7-a2d8-63ff106d590a · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.712042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.712042Z digest=sha256:e1dda4cd5f594a291ca2e69092d67290b68c43744e6180ff544d5623bf7cc680

Observation e04a842c-6fb0-46ca-8e4e-5fb1c67cd60f · outbound

This paper cites The language model evaluation harness, 07 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning The language model evaluation harness, 07 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.802501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.802501Z digest=sha256:0b60fb2a5c34dee151f070a5931809ff4d5c812d911f96aeb550b2416cc63e00

Observation bc96dede-980f-400e-a3bb-297c8eed8dbe · outbound

This paper cites A survey on llm-as-a-judge, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning A survey on llm-as-a-judge, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:58.719702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:49.872918Z digest=sha256:70f1da7f297a6fe6ab58d86c9b3ec7ed61cbfd6a26845d3cc0e0d600f30e1300

Observation c20689a6-cd8b-46be-93a5-e00c5012ee25 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.951924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.951924Z digest=sha256:9acea27840bcb3a71d27e324cc1cedfc8fcd1588c41ee8ad0b99f670f255eb18

Observation a6eebf7f-88ce-4155-a7f8-24bba27364bf · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.028195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.028195Z digest=sha256:3e7b7644524030088491603756ca4c203205d166050e88e6c5bfb23aed149096

Observation de83242f-d66c-4d96-b460-91f78ea5548c · outbound

This paper cites Ultraeval: A lightweight platform for flexible and comprehensive evaluation for llms, 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Ultraeval: A lightweight platform for flexible and comprehensive evaluation for llms, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:58.400339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:50.106985Z digest=sha256:75ce6f4994f0d1a90656405021dd35ca3a9673c93eefa66be463f98eeed7cbdf

Observation 78210343-cbd8-436f-822d-24ac91b47ace · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.175071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.175071Z digest=sha256:945d5af72cb552ca54b741262678a8e64c9fdc65bf691019c418c4d5c541d2d1

Observation 4b835f02-d85b-4a76-97db-17fb968707ef · outbound

This paper cites Putting rl back in rlhf.https://huggingface.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Putting rl back in rlhf.https://huggingface

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:58.082900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:50.257487Z digest=sha256:6eaa8d3a6a2485fe0dd07b8b1127b1ef5e968817cd63dd7566e53020aa76fccc

Observation 9218d86e-64e4-4a9a-bb18-ab94f77c0901 · outbound

This paper cites Math-Verify: A robust mathematical expression evaluation system.https: //github.com/huggingface/Math-Verify, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Math-Verify: A robust mathematical expression evaluation system.https: //github.com/huggingface/Math-Verify, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:57.736623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:50.339116Z digest=sha256:f384304c4ccbcb128ac45422982e79e2d5f805d7371bd1b5a822089cdb9ddf41

Observation a5eab1cc-2a3e-4122-98b2-6576af50da68 · outbound

This paper cites OpenAI o1 System Card.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.434833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.434833Z digest=sha256:415b5f95406e2b1017110eeb1e8ccd500e704ad025d7b3bcd677a0fe168ed1c9

Observation c30b3d3b-ea42-4ad3-85a2-3bff7c4554b5 · outbound

This paper cites Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.521407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.521407Z digest=sha256:fdd0bd44f4f647c37aee822bc0fe299e433b5da2b2392d5d9eb2cdeacb3965e3

Observation 0d7f196a-3c6e-4d18-8406-31549108e161 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.636962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.636962Z digest=sha256:5106385255b1d50caed86843806feaf7eac23b8cd1292553fc7d933ef39ed87a

Observation 9e549583-4066-48f7-ae5a-391be89490b8 · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning From generation to judgment: Opportunities and challenges of llm-as-a-judge, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:57.431870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:50.743462Z digest=sha256:93fb464c4223bdf7b12338511c30bd324d14e58cc9f0bc3175b48d59e6c589e4

Observation 13bb5dbb-8698-4c18-b1f0-1141c77c40d6 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.844232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.844232Z digest=sha256:84b480fb91a8c35765975a14cafc8dfc9698a0ff7994eaffb63e54763989e2ac

Observation c4aad93c-cd18-434c-a46d-a22d024f05c5 · outbound

This paper cites Hashimoto.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Hashimoto

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:57.166159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:50.940181Z digest=sha256:c9c145f428e133b3f6dbb1680950618ab8cfb95afffa5b2179b93b951aa2135f

Observation 415006a0-c6c1-4439-971e-a4c733be84f6 · outbound

This paper cites Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.077787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.077787Z digest=sha256:f0e48cb2882452c8d439a5ac3b05da8b8b6409b95efaebe73478eb22c0572484

Observation 42cc20b7-78f1-4cc5-8994-39500db0b961 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.220864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.220864Z digest=sha256:da220fb95fa5ec87ab6b1448ee4b00c541d9cd36787bf5810c2e5329ec8cee06

Observation a18d1b64-2272-46c8-b00f-40b5de843138 · outbound

This paper cites General- reasoner: Advancing llm reasoning across all domains, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning General- reasoner: Advancing llm reasoning across all domains, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.875421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:51.390548Z digest=sha256:d4968f24c02d9c1977aadfbb9b10f46a6584f5c913468601835dd3f017bde26b

Observation 7f77a618-2304-42ac-a8eb-dff87f7c1c3d · outbound

This paper cites Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.563907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:51.532455Z digest=sha256:2693a4fcea7ba5499791067e7dcd09c1ba3a3807a0d635d3efc418d9e91d88f7

Observation caee832f-c2f3-47a8-800d-a9ac97acbc47 · outbound

This paper cites s1: Simple test-time scaling, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning s1: Simple test-time scaling, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.687694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.687694Z digest=sha256:61f3f9a61db7980031f6f3c9b215271525074658925b73b7b09926844135fa32

Observation a6203986-d348-4610-8853-308b98f5ad0d · outbound

This paper cites OpenAI Evals: A framework for evaluating llms.https://github.com/openai/ evals, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OpenAI Evals: A framework for evaluating llms.https://github.com/openai/ evals, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.331189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:51.831728Z digest=sha256:ac5131d45d744755a8670c8b2bff99200cd6777b31cf4c01e5f03ddc82bd3d2d

Observation a4639574-c2f7-41c9-b344-6d1c7e8c1cec · outbound

This paper cites Manning, and Chelsea Finn.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Manning, and Chelsea Finn

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.982792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.982792Z digest=sha256:eca0d7f0fa6b5b6ddac29f00b0452cfd18e7dbb1fd7f8a483ae65a09d7140ddb

Observation 3774910b-089f-4701-b8b2-f7d78db42b96 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Proximal policy optimization algorithms, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.128786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.128786Z digest=sha256:79cba5c2d42d322fc662e6966beef037f16ce89dfcf7ad7548d71bacd417765b

Observation 0d82b86b-711b-4e51-8e57-d8c78ff1ee60 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.262883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.262883Z digest=sha256:879f1f740e21a636e4b75335415d2c61a270d12ea1f17d3015d9d11b46ef854d

Observation 9e882606-ace1-4a33-81e1-dc723e8a65e6 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Hybridflow: A flexible and efficient rlhf framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.386742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.386742Z digest=sha256:a6738f84a5befcb61a1e05d4772b63a3b63092b8de1c9032a414c228253a21ec

Observation acda1c18-7e8d-4349-b20a-cbfe59c44214 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.514084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.514084Z digest=sha256:bd25336ea04a669e317fb2ed6f633a9e7f700935f699605b7d8373855fef88dd

Observation 3a579662-71a9-4b86-abed-e584b3ee26b8 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.625163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.625163Z digest=sha256:00950055b52229f77e65447589baa50d1aa7de0b66ad77f873b1a4c5aed9d7c4

Observation 988bea9c-818f-4c1a-9c72-eddf547bbe1f · outbound

This paper cites A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.715159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.715159Z digest=sha256:2ce7d06064cea6330150ae99515f5b6ab5f58d08ba7311549f4a3ee36b50ffa2

Observation b8f4181c-2265-483b-8555-2bfc748af162 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.916366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.916366Z digest=sha256:47065cfe7308694d4ce9ae99e05e6d6917f649ccc4056f435f210136ab1ed289

Observation 5bb1a502-39cc-45e2-8769-3041624abbf4 · outbound

This paper cites Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.097735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.097735Z digest=sha256:c6166ba2af68d9d7f900a87b66f8abcae96b58c3c73b581c043ad8a048483388

Observation 31866f05-745b-428c-85b5-d350a798a339 · outbound

This paper cites Limo: Less is more for reasoning, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Limo: Less is more for reasoning, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.220799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.220799Z digest=sha256:51aad0c8fe9e09929231b46470f9a331841b6ad2cb7bbf4ac0f89279ee22b449

Observation c960e078-d3cb-4cf8-a455-d4dc2fb840e4 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.453264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.453264Z digest=sha256:22aeab0696d249cd5d32480830895c1f5150aa40c9acfb2d302f84da5b7e9559

Observation f78eea0f-e438-4e9a-bad1-a1df752bc050 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.603007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.603007Z digest=sha256:c2d36c2b30269850de965b3213986c021057a4499de523c882762aee4c8b6a82

Observation 70534c50-af02-4c81-ae12-951af34c4ea8 · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.787633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.787633Z digest=sha256:80eeffc0fff9880adf963e427fde056619a597e274f5d23d6aa34a47042efeb0

Observation 65631411-b382-47d5-9dff-0ec989593bd8 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.974048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.974048Z digest=sha256:c338c45ede34cce17092da13946889726bba75ef99059c678f70e9f1c12554cc

Observation 6d0567c8-dec6-4884-aa07-53b613d8ee30 · outbound

This paper cites LLM as a judge.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning LLM as a judge

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.054865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:54.177987Z digest=sha256:f1039c946483a389ea9274d510d79aab658b934da0794ada335391491b340adf

Observation 9bed4c2a-33a8-443c-b0df-f3edfdac91f6 · outbound

This paper cites an unresolved cited work.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:55.783306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:54.297538Z digest=sha256:c7bf30286fc3810cf57cc098e261772b001dbc012b561b567aa6ea955880cecd

Observation 3a1be238-ee3e-4a3c-a577-be0e0032c2c9 · outbound

This paper cites an unresolved cited work.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:55.515728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:54.447166Z digest=sha256:a69b01d2ced80a8041120624441f4e689dc68c97258e4763bd00a4328ed5dffb

Observation 1413de55-c728-4406-9901-d4215721e8a8 · outbound

This paper cites reasoning.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning reasoning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:55.233979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:36:54.555335Z digest=sha256:ba4c4493b0729adb002b08d78b57646a75dd498a6540194fda70f219d27b4632

Pith citing papers

Observation 3e48e768-8302-413e-96bd-f89841ee7373 · inbound

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers cites this paper.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.721581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:751ed9c6eae3c5ee5a0e646b0bea1568d58c8185d5277adb85da7f31cbe66677

Observation 42f34f68-8adb-4ba9-baf8-69569e51f279 · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.117674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:3642248479da4849577cd9d780341fb8ef25ac8c015f371d64f0757767135543

Observation c04c54cf-8883-403e-b34a-b417301c2bce · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:06.803837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T15:35:08.202464Z digest=sha256:a7a7bbc76aaf92f76c4ab173dbdc2c8373a7e2f9a900eeb277f7a7cd0efae886

Observation 84e5ac04-f9fb-47ee-a91c-3ba5aa9b5650 · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:15:46.717605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T23:25:55.375541Z digest=sha256:7be2f85f3a0ac37a75bd338e711eafa1e94a3762e559ef185f208a183bb64684

Observation ae651858-8e70-472e-9370-c824e5cd6fd5 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 187

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.318017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:5e012bc0b9ba1ce227de5eb52e19b260edb28382f5c8a8e5209786789d8a1fc0

Observation 607992f7-4e8e-4307-abe4-9675df8e705b · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.185430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:6af11e522b0b898ab073e0fc5a1d97404ec7c8ca9c66bbc5fe11267eb713f971

Observation 4b0a04ab-7d48-4726-979c-99ad29cd69c5 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:3e4b3d1a901eff00ddc8e879203c169d2c610c0475f1bae93e45f4d34af611a7

Observation fcb0a385-ed47-474e-88e8-9680e753f793 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.923422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.923422Z digest=sha256:077788fb5b82f88c99813a74b363fb6f81b53cbe35995e87f48afdd93f2cd097