Pith. sign in

Paper Citation Record · LEDGER

Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2503.23829.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23829 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:44.467130Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4bcdf1b4-c21a-4413-bc33-70e03eb18358 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.467130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.467130Z digest=sha256:4bbd063b97cf6e8d13e36d0e21176203ead725772c38aeaea0108699935bbb39

Observation e14f3f10-e97e-4664-b51d-4925cd83e528 · inbound

Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards cites this paper.

Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:14.066640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:14.066640Z digest=sha256:f8ab85524b055a9d2ac26ea5f83125e80b7b12a160c9ba38dedf02cc554ef202

Observation 5e84209f-bbff-45ee-a5d2-77e813fae459 · inbound

General-Reasoner: Advancing LLM Reasoning Across All Domains cites this paper.

General-Reasoner: Advancing LLM Reasoning Across All Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:09.276522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:09.276522Z digest=sha256:c6c4307a31167cdc2697542661947bf1d0d4e8dade8045016563a51be7e47262

Observation 73c61492-c207-493b-b743-8d840017c1c7 · inbound

Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models cites this paper.

Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:17.653974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:17.653974Z digest=sha256:361213a8f2bc92d0e13c31a01fb4fdc9bab8adca93262abc13a54fc2a9419d4e

Observation facf22d4-454e-4613-b858-c10da0a74038 · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.067551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.067551Z digest=sha256:536a624ef7a60af29942043a04043d4173eae3de30983c1dc161ad114a4221b9

Observation cf464214-97c7-4498-990d-9e23b52bfaea · inbound

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning cites this paper.

OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:02.793750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:02.793750Z digest=sha256:00ba742985b02ed77484416c2679147d9986eaa612fa3eb0734df278c1f18687

Observation 645593a1-35a9-47e3-9965-95a705d71bbf · inbound

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning cites this paper.

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:12.121976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:12.121976Z digest=sha256:cbf53dd8130fbfb74d210102e8704b4bc038c3f853c69a17c5aaddd779e7ee42

Observation b96aebd1-5915-490a-b807-7b216794f814 · inbound

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs cites this paper.

Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:06:28.437391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:06:28.437391Z digest=sha256:75ae9d5bb6ab08e12e62e783c5e1bbae91222c510cf79bfa40ce5b97354b9588

Observation fa4c3211-c002-4827-8936-5f75068aeef6 · inbound

ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing cites this paper.

ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:35.320702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:35.320702Z digest=sha256:e9be530f1ff82b760ebce9b903e0ab4377ded4ca9fd1f7fd4932fa92950f822b

Observation 4f20c883-72af-4be5-b670-ffd717f77947 · inbound

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library cites this paper.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:33.035346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:33.035346Z digest=sha256:1f9ffdf9bb340f44a47ef1095e6a414994c72785093fb2340e59a72a1958371d

Observation 89f5da95-4b79-4e1e-ba65-f83acba760b6 · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:45.062113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:45.062113Z digest=sha256:48ff3edc53dd1f0e51718dc22775e3b010efa1f83bd39c84bf9e5feabcf32509

Observation 45d137f3-ca87-4e2f-b01a-e084297c6e34 · inbound

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models cites this paper.

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:39.063087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:39.063087Z digest=sha256:aebe44de7977246ce6c4d751886ad12322101d44779c120bd05e21cb49a954cb

Observation 7745fa08-d56a-4cc0-aa9c-f06996f5d17d · inbound

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks cites this paper.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:52:14.093197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:4c4b64007424c57f3e02735db698939890f6e2937fa71d0230d94f7438e6429d

Observation 65d37c96-ba03-4096-9c57-857ff75d8237 · inbound

Energy-Based Transformers are Scalable Learners and Thinkers cites this paper.

Energy-Based Transformers are Scalable Learners and Thinkers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:42:28.038550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:42:28.038550Z digest=sha256:0499e29f4e4be165da060d11f74eca295dcab7462025455230fc008391a362fd

Observation 03ff0fa0-99bd-4b98-b2c2-b7454de116c4 · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:38.987143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:38.987143Z digest=sha256:3256f303061456d8786e191bc1c7d6e928768c40df61d2189adeab939ce207ea

Observation 0b6983d6-f37e-4f3d-acf4-257c3df171fc · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.851371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:8f7cd736021345dc32b68ffcd0c937c2bdde541c135ef4f70dc7b1e3b8d889e9

Observation 0ae26018-5913-4345-bf2f-319ccf727136 · inbound

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models cites this paper.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.062832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.062832Z digest=sha256:957ddc851ce940ccd109eaffe77875dd707cfcf8154447424dced013d58c45f4

Observation 7b93d9e6-be98-401e-aee1-b81ff7a36c70 · inbound

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers cites this paper.

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:52.916321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:52.916321Z digest=sha256:b889f195e9167fc6f413dddadf69bf562293529cfb2bdb7f6979b2dcf7c51ea8

Observation 459fb54b-206d-4965-8a55-b0e54227814b · inbound

Reverse Browser: Vector-Image-to-Code Generator cites this paper.

Reverse Browser: Vector-Image-to-Code Generator Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T05:46:44.417762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:46:44.417762Z digest=sha256:2426bbea175b275be2148f58fdda25300ff3e30372584fbb03065ee0ca5cfb2b

Observation 500c22dd-a511-42f3-b0b0-879be3b0da6b · inbound

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs cites this paper.

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:52.481344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:52.481344Z digest=sha256:e8a6a98cc57461dd23ca8a33424016e5bc9d666d1e6f69847a3d343a3c8a0bf8

Observation 21974947-640f-41f2-956a-4a4e42720073 · inbound

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search cites this paper.

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:12:35.834905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T12:12:25.437344Z digest=sha256:b8286f3aaa442ff090a53225db3bb0201d22f89b1b588e840f7434e3c1bf6ce9

Observation 5678ad83-14b1-4b82-b552-9a4d16097830 · inbound

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts cites this paper.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:48.485814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:48.485814Z digest=sha256:a0a11a7a4db2941178763d37800eaccfa430e09ffc6503b85edb4377cee4383e

Observation f6d6e0de-6812-4916-84ef-98aeff15b891 · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.360338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:70a769ff670398b540969ed5a1aeec5a875ed3d0419d859f83ce18f93d934f7e

Observation 49959a33-24e4-4b0d-861a-15ade233c683 · inbound

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training cites this paper.

DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:48:50.980608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T01:46:21.744857Z digest=sha256:e07b5be0289c49e3b7bac3aefe14253f318009baf23db5bf946dbfd34cd8be86

Observation 553adc0d-e76a-4293-a47a-83b6d66855ed · inbound

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning cites this paper.

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:20:57.559827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T13:20:49.919833Z digest=sha256:1242f91ae37687500037a3bcc57b93172283d33d80b5e652580693e4580e5c95

Observation 53947a8e-5330-48d8-88c6-f47d8edd3061 · inbound

Specificity-aware reinforcement learning for fine-grained open-world classification cites this paper.

Specificity-aware reinforcement learning for fine-grained open-world classification Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:46:17.815819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T16:44:47.866126Z digest=sha256:f68f5f802f69129a061b1b6d6874a33fb253506c36c921aa5cbda6d8c9dfabfd

Observation 1570ade8-c9af-4dc6-bf9f-63b133f018d8 · inbound

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation cites this paper.

Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:51:40.644170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T21:51:10.972744Z digest=sha256:728df2ef8b634d1b2e704d7a06b7fe6f57f528c32306f37c36bfc48890316ac8

Observation 3b7189d9-5807-4d9e-aa2f-658142edfee5 · inbound

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks cites this paper.

LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T20:08:17.066898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T20:08:17.066898Z digest=sha256:f6e71959a614d2e3abc4ec5b4acf16fc01a0a63607b8119c53234299713af009

Observation a62ac229-1b51-47de-b77c-2b9ccaf8fd5b · inbound

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease cites this paper.

Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T14:23:49.346727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:23:49.346727Z digest=sha256:7a40ecc3b8f183f97f4aa7e622590aa116cf68b8539ab17bad64b31314c441cc

Observation cbb70e96-f794-4e69-a379-9ec69c09355e · inbound

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards cites this paper.

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.172294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:49:00.343580Z digest=sha256:189fbee5d597cb9b9b7492c0cb668be25450849b18f18de4d6d2d7d23a24a4c7

Observation 369ad831-a096-4da0-8abc-13cba5e99796 · inbound

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization cites this paper.

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:01:04.148354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T23:48:32.613988Z digest=sha256:ab1f4af9c9fd541c761645c2d870d55b3c1481064dfa7ee47334be1ebc0f42f3

Observation 5803ac88-fbb8-4ee0-bba0-c31b7b85402d · inbound

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring cites this paper.

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T19:35:39.078929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T19:33:35.690030Z digest=sha256:7a14918d6e2f72e25f0e8e11adb1772b6fc91520dbc1e2de02d3042767ac2e22

Observation b3786b5d-2427-4e79-8e10-970e17cfd6c1 · inbound

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring cites this paper.

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:15:52.434975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:56:43.707252Z digest=sha256:305f256da1085fe59c9702f40c14cfddaec1809085926de11431e3a3dcdd138a

Observation 1dd83977-cef3-48d3-8789-8ceb9697ab7a · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.250698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:1b27f46e59053217e3f5a82fdd9b5e3b7dd31b08dc0b91a0273a67d036839490

Observation 6e20c37c-f7c7-41eb-ab24-00109f9543f2 · inbound

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models cites this paper.

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:22.662442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T04:43:50.384980Z digest=sha256:0d60ad9983db8f6924c365c6b9f9dd199bffa97fdf5f1d82de681df427da0746

Observation c2ca204c-6241-40e0-afc0-1f11f80191cb · inbound

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents cites this paper.

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.525894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T05:35:45.084011Z digest=sha256:b2469f45bb9dcedfc1cbf6f07b4a32c4a8070c724748a35274ff4a2b8c44ab44

Observation 1f31b556-3816-4be0-bc71-8924d6b7bc72 · inbound

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination cites this paper.

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:47.109855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T23:08:05.597810Z digest=sha256:e3ebb8223b8abb7793a5bac49475fe5dc1186e8b877f4fb8228d68620780ffa2

Observation bfe80b35-bcae-4429-a91f-e2caf0a75a60 · inbound

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO cites this paper.

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:01.064488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T22:25:04.067339Z digest=sha256:61bc49631bbab44bed8b89428f4ac9d5bbc3d7c929c28c06904a41a01b06154b

Observation f305aa37-c9eb-46a2-9aa9-e80b02291d7d · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:33.947566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:a6a1cd5b9ad9711dd8db219f67d67c76d1fc4e8aa89186bdf956fe7a29246df4

Observation 0569465f-718a-4fb8-91cd-7afc15e533d5 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:06:40.787545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:9661fa6e6fe4e371480e25ec7b99e4b98278fc7f62c20ed7ae0406f4b61a8149

Observation 69cbd603-7104-40b4-853a-246a243078a3 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.019297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:1a36ce1f5411cedbc04bcda12f68e5cd533704bb76f1abfca8273cd89338e87e

Observation 55788ff5-3d72-4a1a-888c-f2014b89a2f6 · inbound

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views cites this paper.

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:39:45.338242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T08:38:46.044079Z digest=sha256:beec6d4c477b4d55402b9d9a44d8dadef311f11b8cbdcb41a19dd1e007e46f67

Observation 1bb171b7-5dd2-4e15-b5ba-78bbaeb0644f · inbound

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought cites this paper.

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:19:57.794350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T00:46:17.094339Z digest=sha256:d5840116a9334c41c756309d6f81c1f1f55962ae54c3296626c7925bcdf9a23f

Observation dcf008c4-a9b1-4f38-9e35-7bf8114e9a0c · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.146628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:3c5ff8cbc63d35e6aae7066c9743d0c1c4e7976998def87423ac878d2f2b3a80

Observation 61f6e4fe-1eae-489a-950e-6014c1940867 · inbound

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth cites this paper.

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.482991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T19:26:21.833145Z digest=sha256:78b55b2c9a4d6532149ef1de0328e95b30122d794776ee0125c9efbe87c45104

Observation 0c369fed-3427-4cc4-84b1-b2f2437d9bc2 · inbound

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments cites this paper.

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:28:18.294461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-03T13:21:06.380698Z digest=sha256:282f21255e8c0c998389acf1012c94d466ce8711ea1eee66fc237d642bc9f079

Observation 66b68307-6bdc-4878-ab54-31e38b2cac55 · inbound

MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models cites this paper.

MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:39:15.338019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:39:15.338019Z digest=sha256:cb4e89188a8121d53a3be29372f8978a8c9146dd5596eb1b50d684a55e6b3dd5

Observation 52e9b7ae-8a61-4a64-b77f-673325afb5da · inbound

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR cites this paper.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.487540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.487540Z digest=sha256:ebd6084055f906508af106255ec3d02c954cf5f5ec2a354219566f6d3833ccaf

Observation 2707c228-3d18-4c3a-96b0-dd1d011bc224 · inbound

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning cites this paper.

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T10:46:10.723383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:46:10.723383Z digest=sha256:b60ebfa23415e8232cbb8d43ccba8f119a6b1115a61a646a5565e9b30211864c

Observation 4c2611e7-fe7e-4729-8abe-f6948bf32ead · inbound

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning cites this paper.

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:33:18.256691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:33:18.256691Z digest=sha256:a158050f8c7b5ef9759e0a2751c49eed05373405abb813a5352f4e10dccaad88

Observation c91d0b00-6ba4-407d-9136-bf043f5db81e · inbound

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation cites this paper.

SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:03.935373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:29:03.935373Z digest=sha256:919feac8895997d2616f010d62cacc0864ed0dd4d3c7191d91a76fb6fdaffb01