Pith. sign in

Paper Citation Record · LEDGER

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

As of 18 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 8 inbound Pith citation observations for arXiv:2505.14625.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14625 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:54.555335Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:43:00.923422Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:26.183764Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56aaa975-7b28-424b-8d22-68df808f6885 · outbound

This paper cites Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.050333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.050333Z digest=sha256:b803df37efafcea6ce2560b960647e0e371fec3eb075239828793cb4160b788f

Observation aa829132-5b02-4f00-813a-89c78d605902 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Online difficulty filtering for reasoning oriented reinforcement learning.arXiv preprint arXiv:2504.03380, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.101516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.101516Z digest=sha256:383a42dafc2a093becf156578791a0354f20b0767182002edc658f7261c3736f

Observation 7eb3a90b-efdb-4823-a7f5-f85a8f59ef02 · outbound

This paper cites Robotxr1: Enabling embodied robotic intelligence on large language models through closed-loop reinforcement learning, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Robotxr1: Enabling embodied robotic intelligence on large language models through closed-loop reinforcement learning, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:59.677826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:49.161442Z digest=sha256:04d8ea94ba34164cc09e7d08695229f2cb528ece7ac274906cdeee33cdfbd326

Observation 93031f66-d74d-4699-b9ff-c2290a75f510 · outbound

This paper cites xverify: Efficient answer verifier for reasoning model evaluations, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning xverify: Efficient answer verifier for reasoning model evaluations, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:59.384896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:49.231062Z digest=sha256:0aa2821cb05c8b7974ff501f53e4e54ffaf808f4b9f25c5b9b82dab48d39ea51

Observation d5a7cf2f-c2a0-414e-b39f-ef90567747a0 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.317584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.317584Z digest=sha256:6b5351adc2e4fc8f95c4c46b9aaef0bfb24ec287e13fe1c64c36cc8f3d7eb060

Observation a93a0685-831e-4c90-9a01-795280ab935d · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.383928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.383928Z digest=sha256:a9d272f61c418d6aff9a7ab9442c960a9867e659966a05818791e386444103ae

Observation a9bafb76-c51d-45d4-a77c-563882fa8f9b · outbound

This paper cites Process Reinforcement through Implicit Rewards.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Process Reinforcement through Implicit Rewards

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.435775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.435775Z digest=sha256:69511ed1124a8042b5ae11920212e439d500c9168e9d9bd03fa9bb5826760541

Observation 668db233-568d-4c5e-b3dc-0008c8606ebf · outbound

This paper cites an unresolved cited work.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:59.080504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:49.535231Z digest=sha256:756e3390f8c4e73aa3372a9d666315794e2f0732a5b8ffc9dfb3f24d4d2f31ea

Observation b9eaf3c8-7c5f-47db-8983-e4f81e3053ce · outbound

This paper cites Raft: Reward ranked finetuning for generative foundation model alignment, 2023.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Raft: Reward ranked finetuning for generative foundation model alignment, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.637551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.637551Z digest=sha256:7170dc819a4bdd7b76b184771b9f674722e6954e52d93dfa52858d3782c5eaaa

Observation 106e2374-3c76-46a7-a2d8-63ff106d590a · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.712042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.712042Z digest=sha256:e1dda4cd5f594a291ca2e69092d67290b68c43744e6180ff544d5623bf7cc680

Observation e04a842c-6fb0-46ca-8e4e-5fb1c67cd60f · outbound

This paper cites The language model evaluation harness, 07 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning The language model evaluation harness, 07 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.802501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.802501Z digest=sha256:0b60fb2a5c34dee151f070a5931809ff4d5c812d911f96aeb550b2416cc63e00

Observation bc96dede-980f-400e-a3bb-297c8eed8dbe · outbound

This paper cites A survey on llm-as-a-judge, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning A survey on llm-as-a-judge, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:58.719702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:49.872918Z digest=sha256:851bdd3ad6b97501ab3cc7add6b1de44742a369c25017474738fe1f8f8fbec69

Observation c20689a6-cd8b-46be-93a5-e00c5012ee25 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:49.951924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:49.951924Z digest=sha256:1535dd2b81c44c9430da5e83d1b4186bee2cff3448c0e0a4002c43656326d618

Observation a6eebf7f-88ce-4155-a7f8-24bba27364bf · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.028195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.028195Z digest=sha256:3e7b7644524030088491603756ca4c203205d166050e88e6c5bfb23aed149096

Observation de83242f-d66c-4d96-b460-91f78ea5548c · outbound

This paper cites Ultraeval: A lightweight platform for flexible and comprehensive evaluation for llms, 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Ultraeval: A lightweight platform for flexible and comprehensive evaluation for llms, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:58.400339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:50.106985Z digest=sha256:c442994cd7e1e990412de6f7c0ee71f6118e269032701ac1deda7f264cf25393

Observation 78210343-cbd8-436f-822d-24ac91b47ace · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.175071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.175071Z digest=sha256:945d5af72cb552ca54b741262678a8e64c9fdc65bf691019c418c4d5c541d2d1

Observation 4b835f02-d85b-4a76-97db-17fb968707ef · outbound

This paper cites Putting rl back in rlhf.https://huggingface.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Putting rl back in rlhf.https://huggingface

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:58.082900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:50.257487Z digest=sha256:fc9530b89eb50bd938b1f2422484adcdac55429d1eabdaf8e91920a58f9beb03

Observation 9218d86e-64e4-4a9a-bb18-ab94f77c0901 · outbound

This paper cites Math-Verify: A robust mathematical expression evaluation system.https: //github.com/huggingface/Math-Verify, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Math-Verify: A robust mathematical expression evaluation system.https: //github.com/huggingface/Math-Verify, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:57.736623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:50.339116Z digest=sha256:f0bbaff77c540e7dc5d0a201fcd08c6fb2b08cef0de743b03f4580b74a3e4c5f

Observation a5eab1cc-2a3e-4122-98b2-6576af50da68 · outbound

This paper cites OpenAI o1 System Card.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.434833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.434833Z digest=sha256:415b5f95406e2b1017110eeb1e8ccd500e704ad025d7b3bcd677a0fe168ed1c9

Observation c30b3d3b-ea42-4ad3-85a2-3bff7c4554b5 · outbound

This paper cites Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.521407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.521407Z digest=sha256:fdd0bd44f4f647c37aee822bc0fe299e433b5da2b2392d5d9eb2cdeacb3965e3

Observation 0d7f196a-3c6e-4d18-8406-31549108e161 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.636962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.636962Z digest=sha256:2be3842f1ddf18ce8f24137d5ae56d005b19016fb2d0721322fadb05e73939b8

Observation 9e549583-4066-48f7-ae5a-391be89490b8 · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning From generation to judgment: Opportunities and challenges of llm-as-a-judge, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:57.431870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:50.743462Z digest=sha256:48e3710a147f6d9babc4b859b225507d0968de51bfd4693e731a6bd195abb8d6

Observation 13bb5dbb-8698-4c18-b1f0-1141c77c40d6 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:50.844232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:50.844232Z digest=sha256:84b480fb91a8c35765975a14cafc8dfc9698a0ff7994eaffb63e54763989e2ac

Observation c4aad93c-cd18-434c-a46d-a22d024f05c5 · outbound

This paper cites Hashimoto.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Hashimoto

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:57.166159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:50.940181Z digest=sha256:187d5f789f384e2494ba497809030daf3c5027bb89e83daa233af5dae650f969

Observation 415006a0-c6c1-4439-971e-a4c733be84f6 · outbound

This paper cites Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Wildbench: Benchmarking llms with challenging tasks from real users in the wild, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.077787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.077787Z digest=sha256:f0e48cb2882452c8d439a5ac3b05da8b8b6409b95efaebe73478eb22c0572484

Observation 42cc20b7-78f1-4cc5-8994-39500db0b961 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl.Notion Blog, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.220864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.220864Z digest=sha256:da220fb95fa5ec87ab6b1448ee4b00c541d9cd36787bf5810c2e5329ec8cee06

Observation a18d1b64-2272-46c8-b00f-40b5de843138 · outbound

This paper cites General- reasoner: Advancing llm reasoning across all domains, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning General- reasoner: Advancing llm reasoning across all domains, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.875421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:51.390548Z digest=sha256:9ec3c28f39f40b5c276090929fb97726be16b7ffbb45455d6f013f3f5a358e73

Observation 7f77a618-2304-42ac-a8eb-dff87f7c1c3d · outbound

This paper cites Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Reinforcement learning with verifiable rewards: Grpo’s effective loss, dy- namics, and success amplification, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.563907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:51.532455Z digest=sha256:8e5cd1b3ef3fd891798943e8da06e1d25b02756e46f8ab06f0b492f4382a4670

Observation caee832f-c2f3-47a8-800d-a9ac97acbc47 · outbound

This paper cites s1: Simple test-time scaling, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning s1: Simple test-time scaling, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.687694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.687694Z digest=sha256:61f3f9a61db7980031f6f3c9b215271525074658925b73b7b09926844135fa32

Observation a6203986-d348-4610-8853-308b98f5ad0d · outbound

This paper cites OpenAI Evals: A framework for evaluating llms.https://github.com/openai/ evals, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning OpenAI Evals: A framework for evaluating llms.https://github.com/openai/ evals, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.331189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:51.831728Z digest=sha256:6a7d8c968f0973a8e70cd1d5137bada83b8056808dcb9e158f805554593e6dbb

Observation a4639574-c2f7-41c9-b344-6d1c7e8c1cec · outbound

This paper cites Manning, and Chelsea Finn.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Manning, and Chelsea Finn

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:51.982792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:51.982792Z digest=sha256:eca0d7f0fa6b5b6ddac29f00b0452cfd18e7dbb1fd7f8a483ae65a09d7140ddb

Observation 3774910b-089f-4701-b8b2-f7d78db42b96 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Proximal policy optimization algorithms, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.128786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.128786Z digest=sha256:79cba5c2d42d322fc662e6966beef037f16ce89dfcf7ad7548d71bacd417765b

Observation 0d82b86b-711b-4e51-8e57-d8c78ff1ee60 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.262883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.262883Z digest=sha256:879f1f740e21a636e4b75335415d2c61a270d12ea1f17d3015d9d11b46ef854d

Observation 9e882606-ace1-4a33-81e1-dc723e8a65e6 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Hybridflow: A flexible and efficient rlhf framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.386742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.386742Z digest=sha256:a6738f84a5befcb61a1e05d4772b63a3b63092b8de1c9032a414c228253a21ec

Observation acda1c18-7e8d-4349-b20a-cbfe59c44214 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.514084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.514084Z digest=sha256:bd25336ea04a669e317fb2ed6f633a9e7f700935f699605b7d8373855fef88dd

Observation 3a579662-71a9-4b86-abed-e584b3ee26b8 · outbound

This paper cites DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.625163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.625163Z digest=sha256:f5af6d42c4bffc75f7f0457dc9a42ef70f7e9a74978d4d57f856f14ae7cde39b

Observation 988bea9c-818f-4c1a-9c72-eddf547bbe1f · outbound

This paper cites A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.715159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.715159Z digest=sha256:2ce7d06064cea6330150ae99515f5b6ab5f58d08ba7311549f4a3ee36b50ffa2

Observation b8f4181c-2265-483b-8555-2bfc748af162 · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:52.916366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:52.916366Z digest=sha256:47065cfe7308694d4ce9ae99e05e6d6917f649ccc4056f435f210136ab1ed289

Observation 5bb1a502-39cc-45e2-8769-3041624abbf4 · outbound

This paper cites Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Qwen2.5-math technical report: Toward mathematical expert model via self-improvement, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.097735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.097735Z digest=sha256:c6166ba2af68d9d7f900a87b66f8abcae96b58c3c73b581c043ad8a048483388

Observation 31866f05-745b-428c-85b5-d350a798a339 · outbound

This paper cites Limo: Less is more for reasoning, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Limo: Less is more for reasoning, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.220799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.220799Z digest=sha256:51aad0c8fe9e09929231b46470f9a331841b6ad2cb7bbf4ac0f89279ee22b449

Observation c960e078-d3cb-4cf8-a455-d4dc2fb840e4 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.453264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.453264Z digest=sha256:22aeab0696d249cd5d32480830895c1f5150aa40c9acfb2d302f84da5b7e9559

Observation f78eea0f-e438-4e9a-bad1-a1df752bc050 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.603007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.603007Z digest=sha256:c2d36c2b30269850de965b3213986c021057a4499de523c882762aee4c8b6a82

Observation 70534c50-af02-4c81-ae12-951af34c4ea8 · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.787633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.787633Z digest=sha256:80eeffc0fff9880adf963e427fde056619a597e274f5d23d6aa34a47042efeb0

Observation 65631411-b382-47d5-9dff-0ec989593bd8 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:53.974048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:53.974048Z digest=sha256:c338c45ede34cce17092da13946889726bba75ef99059c678f70e9f1c12554cc

Observation 6d0567c8-dec6-4884-aa07-53b613d8ee30 · outbound

This paper cites LLM as a judge.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning LLM as a judge

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:56.054865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:54.177987Z digest=sha256:4e88fd9b55476d36cffd3bff957aa691e1b326ae410f39f3de7cea0f9c2d8699

Observation 9bed4c2a-33a8-443c-b0df-f3edfdac91f6 · outbound

This paper cites an unresolved cited work.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:55.783306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:54.297538Z digest=sha256:f7cd67fd219be7248f250b2823b4fdf86c65db8093d5aee1208d708bd08432ad

Observation 3a1be238-ee3e-4a3c-a577-be0e0032c2c9 · outbound

This paper cites an unresolved cited work.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:55.515728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:54.447166Z digest=sha256:e0a3a8a0e58bf0728cf810588dfdb3f054c8be506379b76763e0cdc3c7dccf30

Observation 1413de55-c728-4406-9901-d4215721e8a8 · outbound

This paper cites reasoning.

TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning reasoning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:55.233979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:36:54.555335Z digest=sha256:c4223a97e4f8f37eb489544d291f33efef5bbc362240e1f798533f576bec819b

Pith citing papers

Observation 3e48e768-8302-413e-96bd-f89841ee7373 · inbound

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers cites this paper.

Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:40:28.721581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T07:39:02.616294Z digest=sha256:0c44256d7ba9019c738da6511bf1923f98435419cc1fd92fbf2166448ada17c2

Observation 42f34f68-8adb-4ba9-baf8-69569e51f279 · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.117674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:5bfc4486eb521642dc3633085f8071b9c3f81ff887c984c9c40a6860bea4c607

Observation c04c54cf-8883-403e-b34a-b417301c2bce · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:06.803837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T15:35:08.202464Z digest=sha256:d73890e8b95065530cc0cea49ffbc0fec8bc841e2b04fc4882aa304e7b889f4f

Observation 84e5ac04-f9fb-47ee-a91c-3ba5aa9b5650 · inbound

High-Dimensional Statistics: Reflections on Progress and Open Problems cites this paper.

High-Dimensional Statistics: Reflections on Progress and Open Problems TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:15:46.717605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T23:25:55.375541Z digest=sha256:2cf6359646ce36bef0abd2497e995bea72bd0b6fcd05afbff8d21b9d265b03cc

Observation ae651858-8e70-472e-9370-c824e5cd6fd5 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 187

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:46:14.318017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:7820d41caac095978ba3751ac29244f75e98d57c276762a0ff8c9d435d6725a1

Observation 607992f7-4e8e-4307-abe4-9675df8e705b · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.185430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:019070b6aa296862ad092c033ef9f52e6b3e1efc832fb21dd749be0b14005c75

Observation 4b0a04ab-7d48-4726-979c-99ad29cd69c5 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:3e4b3d1a901eff00ddc8e879203c169d2c610c0475f1bae93e45f4d34af611a7

Observation fcb0a385-ed47-474e-88e8-9680e753f793 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T00:43:00.923422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:43:00.923422Z digest=sha256:6606436a0ef474afcae1a6c9f5df59332e713ccf24d4a0b48726848ff9d8f00e