Pith. sign in

Paper Citation Record · LEDGER

Spurious Rewards: Rethinking Training Signals in RLVR

As of 19 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 73 inbound Pith citation observations for arXiv:2506.10947.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10947 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T13:37:51.086217Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 73 of 73 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:52:15.381183Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved10
  • parse uncertain2
  • malformed identifier2
  • metadata mismatch3

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4a385ed9-c2c4-4bf4-9ed3-c95874899ef5 · outbound

This paper cites doi: 10.1038/s41586-025-09422-z.

Spurious Rewards: Rethinking Training Signals in RLVR doi: 10.1038/s41586-025-09422-z

Reference 1

Resolution
metadata mismatch
doi, observed 2026-05-16T13:37:51.109158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:3b4ebff4dec83126ac527d34073698a4629435d5806300a75a726aefbe031a33

Observation 2dea4de5-be44-48d8-b64c-93c47805eaa6 · outbound

This paper cites 2 OLMo 2 Furious.

Spurious Rewards: Rethinking Training Signals in RLVR 2 OLMo 2 Furious

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:37:51.117231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:d96c728790c6b92fd0a590f3e076e0733977408c15468c792430d2fff761a578

Observation 2a5e8549-3eef-4533-a2e1-d79e8df51518 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Spurious Rewards: Rethinking Training Signals in RLVR Maximizing Confidence Alone Improves Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.121149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:aceb8d2784413b65b63970179a4e12fbf88fc45742fa054ccc3119ef584f6bb6

Observation c9d64740-703c-4c47-b418-e60d75ab1a8f · outbound

This paper cites ISBN 979-8-89176-288-6.

Spurious Rewards: Rethinking Training Signals in RLVR ISBN 979-8-89176-288-6

Reference 4

Resolution
verified exact
doi, observed 2026-05-16T13:37:51.104488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:34373c8d723336c30ad33bc47423a5ba6e935bfa4f5ded5d1194b8edf3f9724d

Observation a1f7f2c2-4234-4f6a-ad6b-3b505e044238 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Spurious Rewards: Rethinking Training Signals in RLVR Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:57:51.136140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:5be24120d2e31549307823b5a6c4ad23cbf8478b2ceccbd771a230de27fd5e79

Observation 59b052cd-f9ff-4880-9dd4-617d1ce15ab0 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Spurious Rewards: Rethinking Training Signals in RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:37:51.113502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:f85533cc1d4ad0446154ce259bcd719fa4aa2aad5ea9a2aec168fe770e5d52a2

Observation bde87d3a-33ec-4a7e-95c3-0ba6794ae34c · outbound

This paper cites clipping bias.

Spurious Rewards: Rethinking Training Signals in RLVR clipping bias

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T13:37:51.151752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:cb3c57a58ca4d2729557d169a47731a0bd928b19e94b3b1babc3c67c732856d1

Observation 973f233d-76be-4905-85e2-892b00cd5d91 · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.153992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:1c71a8c8e47c507939b529f2a942af36fd0e347c0499fabe4ceaf36d73f6f727

Observation 600501ef-5d6f-4b5a-8405-621481040869 · outbound

This paper cites r={r},θ= {theta}.

Spurious Rewards: Rethinking Training Signals in RLVR r={r},θ= {theta}

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T13:37:51.156186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:6d7ced3561376721740a557266eed93256b87a8fb6991e2c71f437cc7f973243

Observation 4a01737a-1ad0-401c-ab0d-6290fa19e4f9 · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.158345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:7392ba56be31a37254807d4206de724473c470e8f48c9bca8f9e4091cd1622d1

Observation 41602f8a-22e4-4d04-b61a-fbd9a478714f · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.160343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:3f64d5a00dfb8be3896fe6d2169d6d9c05de726fc32afe2795d3d6e92b4a628a

Observation a629e801-ab4c-410a-b36b-3827bc20669b · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.127702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:5831e8ba5e96080065f44d59cff586bedd1ad1cdac06dea027a5e79aba38d8de

Observation 7076f53d-f1f8-474b-956c-aaddbd3892dc · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.130453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:30c773465b68d08d4a1f2102af17b4ddf7dcc3a2840a4c74139b4b4856a92ab7

Observation da9ee582-868e-45f1-a06d-751106e9c93a · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.132731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:db20a2b1cefe0fd905a2b3f0d40c070f2fef55a3bdcd1fdf3b650ed8adaf0e65

Observation d5deb67a-ce15-48fa-ac4b-50ebb92cbe7d · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.135158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:d42e983980c22a106e154d2e2c92d1d48ff47bb4d4ad09bb7c56587f879dec9a

Observation 6b536b42-0532-4243-8aa7-e0cc01c4a81a · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.137643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:c6b5a1fa5583e37b6c3a91bec6f0867657ad25ef1ed6864f7ee515911fa53b4e

Observation 057e586d-008b-4832-b47e-429b76f7e8f2 · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 20

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T13:37:51.140204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:4ac34e8cc40f540fbde9d1083083e845a919c2226bc09be591f715e3e8afa91d

Observation 7916eca4-cbc1-4fb4-945c-9b1933db954b · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 21

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T13:37:51.142465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:630ead56e743a53e1f7abf07d4913abfe194dcefd09f48b5a2bb940fb53af26a

Observation af755b84-41b2-4e3a-be8a-3ff8821da76b · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.144688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:17ea6d9902603f069eb8160ea2aff6cc152eea6a0de6233cfedb9d8fba676a03

Observation a16ecced-68fd-4e4e-8c10-7cb5b99b973e · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.146842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:ed3fa458f3c5c031bc8e0d3af46b2e57215336b5c4d5ea22e663b38f5c3a8ed9

Observation f6993bc0-5fda-48f1-99e0-57f59dacc218 · outbound

This paper cites Let’s convert10010 to base six using Python.

Spurious Rewards: Rethinking Training Signals in RLVR Let’s convert10010 to base six using Python

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:37:51.149197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:95d447765306f1bcbf3406ee54259843f55ea40fe5d7243197271ccc8a610c03

Pith citing papers

Observation 41dbb91b-0dd0-4c7b-b90f-4ef002dcd815 · inbound

PRL: Prompts from Reinforcement Learning cites this paper.

PRL: Prompts from Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T14:26:40.402303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T14:25:22.002977Z digest=sha256:d7a0686c94cbafb975ad13e77bc8b64f6aa7b9afb90897568871c39cbcacb8fa

Observation f1032a4a-c066-4bc0-bcab-e6fa66a95644 · inbound

Maximizing Confidence Alone Improves Reasoning cites this paper.

Maximizing Confidence Alone Improves Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:49.797135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:49.797135Z digest=sha256:2e27529efc3acfd45b556b3d8de73a32fdf168eeed5f7e72ba88f37abd49539d

Observation 4c0b8fcf-9b99-4446-8d7e-3af410306ce9 · inbound

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling cites this paper.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.462042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.462042Z digest=sha256:d10f42911272a89649a4f3948d060b0245a8ca969ee63dc6759467a2ff88fb9e

Observation 870979c2-4e34-441d-8bad-e60b4bcbc6c5 · inbound

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation cites this paper.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.881915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.881915Z digest=sha256:b55cf3c8df013e99254d899afa4e09cfb0e7775de934c16e9dc8b1e0a710ce47

Observation 4cef041f-37dd-4886-a7f6-26e174124b9c · inbound

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context cites this paper.

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:41.942088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:21:41.942088Z digest=sha256:75fb4fbe8eefe30649f9ad03c45fa7bf214eb9c223a2988c856cb2bb1073af06

Observation bf7edc7d-f230-4347-9a5c-06a11217689c · inbound

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess cites this paper.

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:15.096866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:15.096866Z digest=sha256:af9129fe01cd24c38158a42b102575ae6e509cba409557b79e429cbab5436c0d

Observation 349ccc8e-5008-48bf-8d04-fdc34e30e3e7 · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:c611d42327e6c63c27990316b6b8998a0791ae56d6b8f1b0b479679e41b95c8b

Observation 9ebc9ced-6798-441c-9ce9-e5b486c46400 · inbound

KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning? cites this paper.

KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning? Spurious Rewards: Rethinking Training Signals in RLVR

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:17:06.284535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:17:06.284535Z digest=sha256:e8a4399c0ec2a2573e6e625815d368b8d58c51c980a0f46be02773e92dc63259

Observation c615db8a-b98d-4bb0-aa2a-419d1f96996b · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Spurious Rewards: Rethinking Training Signals in RLVR

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.176219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.176219Z digest=sha256:eb749bc542c6d190523aa8ed99794649dfa146cc772be9e770fb6753afef2b01

Observation 24d1b1ce-3df5-47ed-81f1-c779200c85d5 · inbound

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding cites this paper.

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding Spurious Rewards: Rethinking Training Signals in RLVR

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:13.444701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:13.444701Z digest=sha256:7bdbf76d2d86f9e6f78aefd4c13abbbf29f024f941840388e460b2e9cc3c1392

Observation ae3c49ea-7b7a-47b2-9344-b4211a087cd4 · inbound

Post-Completion Learning for Language Models cites this paper.

Post-Completion Learning for Language Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:15.381183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:52:15.381183Z digest=sha256:935a4350229bbcbb0f728b82de91037431aca76d3d7c80a832a79d241fe1e44d

Observation e6e576f3-303d-4653-9fa9-de1668690b16 · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.501625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.501625Z digest=sha256:c5ec4bd133272de0920c03452f3f762e7d6af2c484f9f34b7c6e3812913acde2

Observation 2adda67e-b61d-4bfa-a97b-07329c450e82 · inbound

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning cites this paper.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.863783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.863783Z digest=sha256:24bb442e77b555b5ca8037fb3d996f6decfa4429dfb5b1202eaa589e9731723f

Observation 06d7b561-f235-48b0-80e6-e4599b64f9c2 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.167577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.167577Z digest=sha256:d99f56ffce36d502423f41e13fa30779914ca952e18ab7a78b0e108fcf20da25

Observation 936e5643-2132-4e42-b7ff-7ab06ed3bc6b · inbound

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism cites this paper.

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:02.795840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:02.795840Z digest=sha256:13760370a04ce11b884d06fd423ac1b329e9f327d84f58cebc5590b25f8f443b

Observation e84b70b4-5be6-47c1-8bf3-377ccab63756 · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.837840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:fb0cc872bc880484ae392ddaf320dd1d07894e471d649ed83337dd80a64c660c

Observation 70d53b55-92e9-418f-8b2e-2d0617b2bff0 · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Spurious Rewards: Rethinking Training Signals in RLVR

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.869115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.869115Z digest=sha256:b21ff601ffcadb6209f58c1961917d291472236b93bb2e46111586cfa0e2b6ca

Observation 37a43eb2-b4c8-43e4-a591-5430be5eb915 · inbound

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies cites this paper.

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Spurious Rewards: Rethinking Training Signals in RLVR

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:31:29.537032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:31:11.974395Z digest=sha256:ec05ed44542db91fdd05cb51658627d02d64c42f8ea0a3ea012ac16fe67f1c68

Observation d69aa6bd-aea1-46e4-914e-affbf4746d30 · inbound

A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining cites this paper.

A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:59:01.622788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:59:01.622788Z digest=sha256:715fbe2f85c9e8177f6054d0af889c933781074f1887399474e0d09e20a5877e

Observation c28a2971-a4d6-4140-b04f-4f1f68ed129d · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.230179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.230179Z digest=sha256:2e691a008e61da6e9182088ef34702d50a84614635e206f40de4f410c2a93a71

Observation cdfd8145-d278-4c87-bc4a-31f25d7eed51 · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:31.017232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:31.017232Z digest=sha256:c2fd1f481caf7326c2793477db2e26e3fe25e81c23d153bf679ff7dd88737036

Observation 74485d24-0a0e-45a5-88f1-77a2265e499a · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:12:23.642649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T05:11:29.205366Z digest=sha256:fc262d68dd3c6256c0ebd1db985f477938cc416d7c3248f84a498b4cc6c71432

Observation 335e52f9-56cd-4d6e-b4f0-1f8cd628784a · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:10:34.994167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T20:06:16.172916Z digest=sha256:f6860c14909dd305c8eba3e2598f0f05341481114f050de8536e5b7f70c28423

Observation 1d569e59-5604-43f3-ab7f-8d9ffc7f311f · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.162212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.162212Z digest=sha256:feb9a3a119f962ea0ba920f7a7d2218b8538e95afc806025abf73e15cf7e2826

Observation cd0c33e4-7ff5-4b20-b972-aed121ca57f9 · inbound

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards cites this paper.

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards Spurious Rewards: Rethinking Training Signals in RLVR

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:40:17.700706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T21:37:54.702010Z digest=sha256:a1635a5faec53435208f68c8d88c5405cca1f51c74ef5ef070d97959ecf7d075

Observation 1b0b2d8a-1d75-447c-9764-7a64e8d2eed6 · inbound

ThetaEvolve: Test-time Learning on Open Problems cites this paper.

ThetaEvolve: Test-time Learning on Open Problems Spurious Rewards: Rethinking Training Signals in RLVR

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T13:16:27.293449Z digest=sha256:b939718362fcd7739d616053e15f8ff477527cc8a7dea9c9b44fbb0a46db899b

Observation 5ec94834-9f43-4502-8714-d84f68b76924 · inbound

What Is Preference Optimization Doing, and Why? cites this paper.

What Is Preference Optimization Doing, and Why? Spurious Rewards: Rethinking Training Signals in RLVR

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:40:28.955681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T18:37:48.161545Z digest=sha256:a08df960fcb9dbf0c73bdbd3b82582504b9eb6be298e409c7d38b991c84ffe76

Observation b03b36f4-d86f-4452-be41-ff46f16d8cf7 · inbound

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs cites this paper.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Spurious Rewards: Rethinking Training Signals in RLVR

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.212310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.212310Z digest=sha256:726a647b1ec9cc160640aacc79b7d9294ab318d08852981341a28d0e33327f1f

Observation 2a3f5c8f-3e25-4b3a-b2a9-fe248c574759 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.283684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.283684Z digest=sha256:2e6ed777879dd9309994c177fa291902bf7e1ff50a7814eeb1cde9a956e101e1

Observation 0359888a-b996-489a-a866-91570c1013e8 · inbound

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics cites this paper.

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Spurious Rewards: Rethinking Training Signals in RLVR

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:55.977680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:55.977680Z digest=sha256:ad0cc7f2fe31c1956ce8250def913f864ffe17d6c02307b8d0d3b4099c3c6ba3

Observation 33539e0f-63bf-4dd7-99df-98f992a9cac4 · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Spurious Rewards: Rethinking Training Signals in RLVR

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T19:23:10.901441Z digest=sha256:ade6b2036ebeda470a037cccb4ce72e675cd1e2fa214a924717489fe74937ada

Observation 00440b3d-9da3-4d3c-9dec-6331025970ee · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Spurious Rewards: Rethinking Training Signals in RLVR

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T16:52:29.593103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:52:29.593103Z digest=sha256:54fc4b15902e4ff4bc1fa6cfb74be489187742401eee0d53a5e99988675bf037

Observation 888f70ca-4c29-4f1e-b9cb-33d0d227567c · inbound

Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning cites this paper.

Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:30:40.368494Z digest=sha256:1f2c980979f1e6ad0bfc75f3682dba8e14def606aa5c6d39ff563153510a6532

Observation 766f1fee-28d3-414c-892a-3872cab38956 · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards Spurious Rewards: Rethinking Training Signals in RLVR

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:323c84eb8f20a7428544d8d56371090a30f19162c9e88b93040aa3952a62e322

Observation fa7f8d2c-efbb-4c43-bc3a-a4de58daa56a · inbound

A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions cites this paper.

A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions Spurious Rewards: Rethinking Training Signals in RLVR

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T06:38:22.041842Z digest=sha256:83d14880319f9fd499668946519b72c2b236edf785a33d7fca6ca31a4c3b35b2

Observation c4a486c9-bda1-4833-9791-dc67fc342346 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Spurious Rewards: Rethinking Training Signals in RLVR

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:f950a239d2a4cbb11f78fce666052b4e7c42778d9777bc090ae20ef5b56c1d9b

Observation 8597c541-ade7-44b2-990d-16382fbae164 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Spurious Rewards: Rethinking Training Signals in RLVR

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:6346e63b2be6383ce37420602a200b0269a83fd52019e780b53782d370c016e1

Observation c955551c-4937-4047-b106-13062012e8e4 · inbound

SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning cites this paper.

SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T06:44:34.968571Z digest=sha256:8df5c7f6dbb49262f4e98bd69cfb72ef4860f30d8f54c2a1bba6eaf62a5bb5dd

Observation 9d6aea07-dd05-4d32-abd4-019fd8f4337e · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:a4af5f7c78d810b9f441c108e8324ef48eaa9c0281999468a7f5d0e5afbe4863

Observation f0af7455-69d4-43f4-a332-c2eef49e8601 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:aecc25cb1a411d32a2fdebc8163bb7dd666f366b6bf65e6367d3a717551a39a8

Observation c23d3da4-be60-49fc-9edb-5f45b380306f · inbound

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling cites this paper.

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T18:58:22.865080Z digest=sha256:4377c8c2a874106678c74cf193d0422fded060cda48b745ef4efa61547d29f83

Observation c802e97d-8c70-4acd-bf93-69ab6ad87b2a · inbound

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling cites this paper.

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T02:14:17.870184Z digest=sha256:5f16de324b7f1bc5bfcbac39524eeea8de74f9b8fa035ed82319d622e26c1e1b

Observation 3c315a2b-e9eb-49b3-ae7e-cf7e89165677 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:110681a507228b71fb921e2081783aa731a6e5371b0f4df8205091f9a2396b6c

Observation 3d0b2257-17ee-4361-83c5-0afc8b6c230a · inbound

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation cites this paper.

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T16:53:00.860162Z digest=sha256:f6693d8b877efa3445ebfc07a6f5c8ff1757db38f54e95535c5a4c405a4284f9

Observation 50827bcc-6bfd-4371-ba76-e13d04f92767 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:075518ce1bef1e80f36d5fd42148958e4481b2ce98b794fe67b4f8e87801da6c

Observation 1f867728-d6a6-491d-b6aa-2af9f8caa81e · inbound

Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models cites this paper.

Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T06:01:15.799220Z digest=sha256:a2da0561bbec4ba9fba5a41f62a9379c7b9c52870584c35c188331efffde91c7

Observation cfdb2cf9-250d-4247-80ff-9601d85bf7b4 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:4d7f851c271aeefb3f4de524dd45d0dea52e59abeb5938c2f95cec2f59fd119e

Observation 726a0883-aeaf-4dd3-a01c-6345737dee68 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:22.952772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:75707bc59922385e4f175a9aa95c21deaa78aa60a129f38f6b97693100576437

Observation dc8355c6-b564-4f6c-a35c-fce4c1af9507 · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:80bae85244bda81a615c4a9046ccbf4dc7ede42e09b16948a3ac4687d54564f7

Observation 2a608b4e-ddbc-43ba-8a18-56c3a284fbe7 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Spurious Rewards: Rethinking Training Signals in RLVR

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:5a5e9066055b173d679660750e8e8410f4d1dceed3c6a1fe9a59955d6fdc42df

Observation 982276db-f770-4500-a269-4fe318c9a47a · inbound

Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP cites this paper.

Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:33:03.662351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:e4e63210714cced51b9c4837f9ed2727fd9928f7efab8998280f73889fd860d3

Observation a9bcd287-bdb8-46ac-af5f-ae14cf350702 · inbound

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR cites this paper.

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Spurious Rewards: Rethinking Training Signals in RLVR

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:03:03.508869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T05:02:35.271960Z digest=sha256:f1e1237ab8789f2b73391a939e286b4393be6bfbc9011efbfe25d8a5ac2058b6

Observation 25f5bb25-b4c2-4e0b-9f3c-1341ed627632 · inbound

Label-Free Reinforcement Learning via Cross-Model Entropy cites this paper.

Label-Free Reinforcement Learning via Cross-Model Entropy Spurious Rewards: Rethinking Training Signals in RLVR

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T14:23:30.616183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T14:18:38.733276Z digest=sha256:1551b872972e88edcf4ec17871282f658d6c0abff5430c2f258c03ddd00333b6

Observation 43107580-6423-41fe-8bb9-d97f04f04b5d · inbound

Reasoning with Sampling: Cutting at Decision Points cites this paper.

Reasoning with Sampling: Cutting at Decision Points Spurious Rewards: Rethinking Training Signals in RLVR

Reference 3

Resolution
malformed identifier
local_arxiv, observed 2026-06-29T08:23:15.534726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T08:17:08.488565Z digest=sha256:88de7b87e4c44103960942ba8748443a30a5f97fd5088a594ef2f708799b7ec8

Observation fb845395-bada-49c1-aeb5-8b25ae86e72a · inbound

Consolidating Rewarded Perturbations for LLM Post-Training cites this paper.

Consolidating Rewarded Perturbations for LLM Post-Training Spurious Rewards: Rethinking Training Signals in RLVR

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:42:46.870525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:37:34.673402Z digest=sha256:3628e263fdf5b272491e75e22b0c75cf8512e74583d0c1a50332945224d45c6d

Observation 03f90375-e681-4bc1-915a-4f83010bb428 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.873180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:128bc0c47b36c88d8dff1648de6cc452f60e04bf37055a23a93189e78f2b5c48

Observation c9a30e50-d8dc-478f-a9ed-d53e814b1198 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.467969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:b4106e6501eb8f975040f9db6ec136e361f1643f6b444a09ee6586a1c27f0115

Observation e8950c1b-8edd-464d-bc8d-23c354c61ced · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:06:40.763205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:59314535aec77a719e2ca9af21480da5abc7bd6eb21192f611009f7f7fb07359

Observation e3aadc56-5a37-423e-a0c8-87ff3bbb679c · inbound

A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR cites this paper.

A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR Spurious Rewards: Rethinking Training Signals in RLVR

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:06:59.082200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:35:26.306648Z digest=sha256:30613dfdd5e6798208139e3c1492b3f8b4603e435cda4643fb30376e6e1339ec

Observation e497ec34-05e1-4055-b5dd-2c784799cdb0 · inbound

RREDCoT: Segment-Level Reward Redistribution for Reasoning Models cites this paper.

RREDCoT: Segment-Level Reward Redistribution for Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:57.646397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T02:12:16.029526Z digest=sha256:8b03da4ef6d18adacd1e17fc50c75e0ea9f16fdecc1cae694f7949e6e7fbd1ea

Observation 4ca2855a-fccd-46ad-a9e0-0db11187f3a5 · inbound

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO cites this paper.

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO Spurious Rewards: Rethinking Training Signals in RLVR

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T13:20:56.495398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:0d876f5c4383df6a6701b95973312b3540576b667b96301683756ebba46cd263

Observation 028afbad-e9eb-470a-b0e9-1f03e1a6cc86 · inbound

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO cites this paper.

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO Spurious Rewards: Rethinking Training Signals in RLVR

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T11:54:46.435397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:54:46.435397Z digest=sha256:f0574693ab530c7b279549a87d570af88ccf9f86a97bf2995bd65014f21dc436

Observation 701020a6-8764-4e8c-886e-c8bbc61bc510 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:38:55.978596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:034448b135480fe9d3d04d8af2afb3fb6d1afbbcf2a89ab8f6b6243cd4847876

Observation 45b40a9c-d615-4e3b-b516-ed6555117f7a · inbound

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents cites this paper.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Spurious Rewards: Rethinking Training Signals in RLVR

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:40:07.890098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T19:55:51.114244Z digest=sha256:94fa8e121a09e32a4d22a89c511a2e67db8fe606ab5eee32acae644bf3dd7a9f

Observation 520d274b-f92c-4b51-a10f-b89807016227 · inbound

RLVP: Penalize the Path, Reward the Outcome cites this paper.

RLVP: Penalize the Path, Reward the Outcome Spurious Rewards: Rethinking Training Signals in RLVR

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.379832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:c6691cf31db03b57bd57970cd2d1f2878df6daa1d829823d5c197d459d528733

Observation 07f73784-95e0-4cbd-8408-e75d5948032c · inbound

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning cites this paper.

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T09:26:08.755245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-09T09:20:08.212837Z digest=sha256:b65ee70e4f9e421e7c2059d0ba99af4b38b01ff650254ff043012399026910cf

Observation 353c0b54-a1fe-4d06-adee-9740cc3fc99d · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:9ed7875f3ccdbd7e3694333c8062f47f9692474c9f2ec7c48098288557e0317f

Observation 35fffbe0-93cb-4d5a-8829-f4f83000f389 · inbound

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning cites this paper.

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T17:39:27.292075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:39:27.292075Z digest=sha256:5f09387e8cf74c3532eef2eb8cb134e9c2410b1068fe9baecffe35387b09fae5

Observation fb1244d4-f004-4d78-b606-b1105598a081 · inbound

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR cites this paper.

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T07:36:47.336670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:36:47.336670Z digest=sha256:dd772a22ddeba607fbae3f539e88726f506e128d809ece9fe27c7b5527950c2f

Observation 3348cd36-1e17-4155-90d2-f7211246e2ce · inbound

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift cites this paper.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.835722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.835722Z digest=sha256:95e65cc01a5abf86caccd306aa1749a7fba2e8b82f12b6111615c87b24074896

Observation 66b24e57-486c-443f-a922-b8654f6369e8 · inbound

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models cites this paper.

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T02:03:05.632468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T02:03:05.632468Z digest=sha256:c271f9450940df73621435ad9b232b00937163b4f63b8cf875891128ef188e60

Observation d7ab1b69-45a5-4627-893c-40d3a1179c0b · inbound

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning cites this paper.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.867049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.867049Z digest=sha256:349636a12d2ada2e95708fc6ef493cd0c1ea243a00c1882ef1647ac929c74683

Observation 972f5e9b-f015-43b8-ad1d-a35d8b32f25f · inbound

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning cites this paper.

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:40.900230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:31:40.900230Z digest=sha256:190c766b30d2978f3e5fd8a8ccba6dc63fb664967779642ad5e487656a5b7e76