Pith. sign in

Paper Citation Record · LEDGER

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

As of 13 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.03119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03119 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:57:30.540771Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecba873f-4762-45c8-81f9-16fbd33de8e5 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.386244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.386244Z digest=sha256:643aac9e986ce76ecee9c06b4ac3bc6d65e3d41c6e0e987931f22ca2cfb9ff03

Observation 07b09dfb-23e2-4526-b328-55a6c87417bd · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.390868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.390868Z digest=sha256:4f3a99e1d63e4079be4191cbfd1f8a6a47548caa5949c2f64ef6b7f857245b5d

Observation d87970c5-f31c-4c38-a7f7-36f9bb00f687 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.394592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.394592Z digest=sha256:935742c682dccf7fb9f44fdce0b3b8e0c78ad439bec328a78126c4b1688d2d3d

Observation 18005705-c99c-4694-8124-143e4fa47855 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.399182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.399182Z digest=sha256:42200ac540634b70744eb06094ab1e8c4008a832dc5dfc8fcbdb3dd5c9a175b8

Observation 97f01257-2ca2-47b9-901a-e1134505b677 · outbound

This paper cites The Llama 3 Herd of Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.403611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.403611Z digest=sha256:ba4a009c039aa76ea121680f6419e0b48b6d69717e256c2c47bfc8be49bd9cc5

Observation e80c2880-744f-4faf-8eb8-199be991df52 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.408060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.408060Z digest=sha256:ebe365b4c18534090fff1e5c09cebb2dee75eab0724d95a121c4f93b12e545c8

Observation 74b9c92d-b17b-4834-8392-bbf267c0b8a6 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.465507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.412971Z digest=sha256:e2a8fd5f3942f232ffa80d39a288631d2ad829487517314bd52e9314fbb2e169

Observation 484fb415-0e14-4aac-8b65-118275eaffee · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.416789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.416789Z digest=sha256:229889fefaa3cf0b029c1831d2bd09ab54db70f18effe992abac0a2e5e0bb531

Observation 31830e38-500f-4397-aac4-55b103e5e2ae · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.454145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.421426Z digest=sha256:3a31c3dd9aad5686d71ca10992fc183ecbd120e919de5785228b1f815e3dd36d

Observation 99c47dc0-d0af-4189-87b4-953c5e17baf9 · outbound

This paper cites OpenAI o1 System Card.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.425198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.425198Z digest=sha256:99674ca4450c354954f1735ae29cb502b68496d4bf8b9fa50e4ff192a16d4cb0

Observation 4d683406-7572-4bfa-ad39-ffa84c38e2fb · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.429284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.429284Z digest=sha256:428563784b64d2625de65c513f8a3e4e7c2a2fe04c1a8983eb7c02590abf1051

Observation 8a4d22af-d5c6-4369-b5b9-1450eff70389 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.433618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.433618Z digest=sha256:98cda64c52bcc12176b2745f5620907d694013ede236f8152eeda685ec24d8d5

Observation 856f1773-7b5e-4e2c-af7e-5ab9a838f22f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.437630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.437630Z digest=sha256:f7ceb166a1f348d0a1f544802ce987e8294dee5384601c6248ab049f53b2d3d9

Observation acccc0d8-e472-4e29-b736-14e9ece3952e · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.441747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.441747Z digest=sha256:93ff76c27cbba029a00813d7db79c345c1225538e405603d2c06485b67cc1f35

Observation 6a4a8b41-64b2-4fe0-8847-50fbe8d8f044 · outbound

This paper cites Let's Verify Step by Step.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.445813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.445813Z digest=sha256:a1112c7d9d84acbeb47fab242b4bf57dba51190d8f00cc7d3f0b53ec3a5aad83

Observation 639fe561-f601-4de2-9b25-cb60c4cb8fcf · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.443605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.450080Z digest=sha256:e9836de2f8a5f96d8e995d9bd58e374bb7ec317e223c898de04c09468fbbafa6

Observation d9b300f8-7a82-4088-a60b-468e9679e659 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.432998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.454210Z digest=sha256:8ed5aa0db21c57a5f2a78c357e12b5f9729af90ff7acaccbd710b2182ce930d5

Observation 031d6294-c5f6-4f4e-848e-3be310f62c2b · outbound

This paper cites Training language models to follow instructions with human feedback.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.458018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.458018Z digest=sha256:101ba1bcc9a039aed3678c279ea132b0ac02d4721109ab6c14cea03c92485424

Observation a1838133-bd51-4c1f-bf27-58ee53ce2e5c · outbound

This paper cites Language Model Self-improvement by Reinforcement Learning Contemplation.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.462207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.462207Z digest=sha256:691d2a7b159826499655e96e1de01881ee07397015192fe3b059b9c96c640701

Observation a5c27ee9-18a9-44e4-a8f8-2f1ba7aba2c1 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Maximizing Confidence Alone Improves Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.465960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.465960Z digest=sha256:4130bf45e42628d121ae40ce2e5831d0eb2fc58cf3d4b9e338567afeab815778

Observation e4444bbf-596d-4703-8f07-b06a47a936e5 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.469657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.469657Z digest=sha256:0131d378701c94ec851a9e569988685f3e5ab11f741157409611b8eb4238eced

Observation 2407f428-d323-4c52-b40f-01f50de7819f · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Manning, Stefano Ermon, and Chelsea Finn

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.473324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.473324Z digest=sha256:bf566d60c69f4040bb8e63a16619afe9d34b7a4cf6820932a0ba6465774a058e

Observation ba4562bc-c250-46e2-b7e3-d20bcea45c59 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-08T00:57:30.854553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.476769Z digest=sha256:8196c6d0213504202d52a80e6eaebb2c66dbd523fa2bc4394b4217d371bd4565

Observation 4ebe5c2a-8a41-4d26-96b6-4460b9107f1b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.480119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.480119Z digest=sha256:a4ef6cba2f3247aa19c4840c6a5f37264ba188af1ab67903f802db59c5d19b34

Observation 6e233841-6b2d-4913-abe1-03758e78e6cb · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.484139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.484139Z digest=sha256:4207c0be4335879c894922a855adbee47450e63d991edf90e250bad163634370

Observation 52e9b7ae-8a61-4a64-b77f-673325afb5da · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.487540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.487540Z digest=sha256:cc0f5973493af330aa31ae9737807fff98233b2869d9769c7c6ffb917cee27a0

Observation bd697cd9-bfbe-4342-a776-d9aa8ca12e60 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.491054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.491054Z digest=sha256:76f8ecf9bbd0476e271545e5fb5687123f9ba2404c011db76db4fc874e9e57a4

Observation bffc2168-0b82-4e2f-8f03-a67983067e19 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:57:31.414598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.494817Z digest=sha256:d8d83f8b832e8e2e788d90e8f9e3b3f656b3e00a7f8fdc5fa0ae6f50f8e32131

Observation c19251bb-e489-49b7-8ef3-3a65dfaf0ad7 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.498471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.498471Z digest=sha256:87aff85d917c80e998e66255b38b6fe4b41951b2825772d12614d212b6cf6965

Observation 0b698316-a173-408c-be4d-e7d77d98a12c · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.502123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.502123Z digest=sha256:b890f9f43507dc6584d6f1c580f7750eba09f7d3d9d1b05cad81907393f0f348

Observation a6584d1b-77ba-405b-814e-f56b6451ac95 · outbound

This paper cites Qwen3 Technical Report.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.506098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.506098Z digest=sha256:3ce98767cbb12548518fc904a5ef0a9ed0e0d18f860b55df5991efa694228c51

Observation f2b0c7ad-135f-4a0e-90bd-f0ec6b1178b7 · outbound

This paper cites Qwen2.5 Technical Report.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.510294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.510294Z digest=sha256:34443b6491f7f4502d4891454ffe4a5a46710184c5d41a628c1037b0856f9f47

Observation e31dbe6a-4007-4b9e-85c6-03f5a818230c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.514403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.514403Z digest=sha256:ecbd39e1be77369744dc97e7a70166af3c61adeb3abc2e5b3fef18faf7d4b6b8

Observation 56d7b9d6-bb2d-43c0-819b-d3f6c8771485 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.518105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.518105Z digest=sha256:f555fa09726d55aa56bcf3dd2dddbeeaa42367d6becd2520e3e42fa3c003f236

Observation 71176864-784a-4c8d-a116-44d3b33deb16 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.521694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.521694Z digest=sha256:9ef39c0140daf7e8da79597247a9d8bd8fe81085e8a88be5e1d69eecc400a3c1

Observation 5bdbee35-1a8f-45dd-833d-2f3053cbf2a8 · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 36

Resolution
verified exact
doi, observed 2026-08-08T00:57:30.666565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T00:57:30.525459Z digest=sha256:2718ec33d388282a87b5f48fd82b997142e8b10cc40f42ea2ead39d418dee178

Observation 91f079ec-09f0-4ead-865f-40038c090f19 · outbound

This paper cites Absolute Zero: Reinforced Self-play Reasoning with Zero Data.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Absolute Zero: Reinforced Self-play Reasoning with Zero Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.529006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.529006Z digest=sha256:9236a9b1a0d0d97d304057f5999ad4738b8640c70a3db6ea90ebcd5e753c9dd0

Observation a3e3bafa-9008-4732-aef9-bfd00e0bbdf1 · outbound

This paper cites Learning to Reason without External Rewards.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Learning to Reason without External Rewards

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.532594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.532594Z digest=sha256:e74429dc34c787d3a3a4c876b5325daf6fc0930493ef8fee749aea044cc60bcb

Observation 9b1374cf-3fea-4de6-aa10-19a0e09406a3 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Instruction-Following Evaluation for Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.537043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.537043Z digest=sha256:eb0f6f04e0f051cc1ed841a878f856d22ad37f0b40ed3a14abe598597113d10f

Observation df8cc502-09ba-410a-9a0b-c6e8240da32c · outbound

This paper cites an unresolved cited work.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.540771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.540771Z digest=sha256:fe84c9c7fd1215dacdc05dd23ecf1a9deb82ee20bf2770c0d7d3ab5e518e5b13

Pith citing papers

No inbound Pith citation observations are available.