Pith. sign in

Paper Citation Record · LEDGER

Lessons from Training Grounded LLMs with Verifiable Rewards

As of 15 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 3 inbound Pith citation observations for arXiv:2506.15522.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15522 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:59:49.564233Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:48:29.575024Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:47.204727Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 35225b90-78d7-4752-be6b-ec1e6cd2e9f6 · outbound

This paper cites I apologize, but I couldn't find an answer to your question in the search results.

Lessons from Training Grounded LLMs with Verifiable Rewards I apologize, but I couldn't find an answer to your question in the search results

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:59:50.547041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:59:49.564233Z digest=sha256:045bc0a563906151dbc867bb86b390a9631c50090bc862dff3813ce82d8474c2

Observation 0230525b-aa15-4555-b9e9-234d6bde6b96 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lessons from Training Grounded LLMs with Verifiable Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.123573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.123573Z digest=sha256:de9d1775a0508d2b00f57464b9b6d2661cf480b2e6771205eb4aa6d556cfddd7

Observation 45617367-ef31-4fb2-9d91-6a7148d34bca · outbound

This paper cites Training Language Models to Generate Text with Citations via Fine-grained Rewards.

Lessons from Training Grounded LLMs with Verifiable Rewards Training Language Models to Generate Text with Citations via Fine-grained Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.358493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.358493Z digest=sha256:8c2da215d4db4427c23c005aeca02f9ca5f5b7796ad0fc00ac20b1e1decdbd68

Observation 3daa139d-305f-433a-ab13-407f4c4408fc · outbound

This paper cites RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement.

Lessons from Training Grounded LLMs with Verifiable Rewards RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.438690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.438690Z digest=sha256:e93571b5f452337891cb22c25e4c9805aac2ee34c57605cc892527064562d23a

Observation 300387b0-b5d7-485e-a35a-6b66106dc482 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Lessons from Training Grounded LLMs with Verifiable Rewards Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.887449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.887449Z digest=sha256:0db6a9e03745af2cb89d3a933bcc7c721253ab446ae7dacd5cc09fb334a33b31

Observation a7343cb3-fb7d-42ca-b9e2-6bbd434ee48e · outbound

This paper cites Attribute First, then Generate: Locally-attributable Grounded Text Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Attribute First, then Generate: Locally-attributable Grounded Text Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.945371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.945371Z digest=sha256:db7b315ecc285c47c1eb0e86babe0293cd4336c4c044db727846acbf938ed6f4

Observation 38d883f4-49ef-47d8-9173-84b1fcecba0e · outbound

This paper cites Qwen3 Technical Report.

Lessons from Training Grounded LLMs with Verifiable Rewards Qwen3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.122177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.122177Z digest=sha256:017e128bd93a785e5c2376ce41003cc9558c7e13a953e81bc42e941a178f1bec

Observation 9a973654-1c0c-499d-bdfe-82e746fde4a2 · outbound

This paper cites Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.203531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.203531Z digest=sha256:cb62f0194b0673fe8189307f8cf56dfb386edc70e0ad77ae58496c5fb364ee03

Observation bc6b8e4e-2b09-48c3-8f94-c45a11e00523 · outbound

This paper cites RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation.

Lessons from Training Grounded LLMs with Verifiable Rewards RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.286759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.286759Z digest=sha256:f5f43b2ee8e362c3726444ccdf63de137bf88a5efce845b2dbb4b937abf3d6a9

Observation 18acf74a-00a1-4c3b-9ccc-2daf31e15d24 · outbound

This paper cites Effective Large Language Model Adaptation for Improved Grounding and Citation Generation.

Lessons from Training Grounded LLMs with Verifiable Rewards Effective Large Language Model Adaptation for Improved Grounding and Citation Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.340475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.340475Z digest=sha256:a4ce0e71b7c72fd2cb9bf04471291463c71a4cdb3f78c16c9366ea274e6069d7

Observation 125741e5-9e1a-47bc-8313-69b1c8f423f7 · outbound

This paper cites Making Retrieval-Augmented Language Models Robust to Irrelevant Context.

Lessons from Training Grounded LLMs with Verifiable Rewards Making Retrieval-Augmented Language Models Robust to Irrelevant Context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:49.441945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:49.441945Z digest=sha256:8d5eca39f59d84046cb7cc574d7392b42e0ee691c7f804a1535f5845fe02df09

Observation d956c89f-400c-4804-9a74-e5be51119b0c · outbound

This paper cites In Webber, B.; Cohn, T.; He, Y .; and Liu, Y ., eds.,Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781.

Lessons from Training Grounded LLMs with Verifiable Rewards In Webber, B.; Cohn, T.; He, Y .; and Liu, Y ., eds.,Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:59:50.767047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T23:59:48.580786Z digest=sha256:bf3f29623c58e44c9153f904e4e5bab3726e0946e7bf6f7c83fdfd38ea3079b7

Observation ee134af0-4c9d-4123-91d5-14d8f7f6ab1d · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Lessons from Training Grounded LLMs with Verifiable Rewards Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.657029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.657029Z digest=sha256:5d2de3670237b5440578c87a5f7dcfab574f7f86f645932758b76280969eef62

Observation 9c66798b-d9aa-4d47-8cd1-81052fcb18f4 · outbound

This paper cites Training language models to follow instructions with human feedback.

Lessons from Training Grounded LLMs with Verifiable Rewards Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.756238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.756238Z digest=sha256:87ee7d67ded9b62017161e4523cd1de4d11e05c8df8c6f848c237d3d3c1d6888

Observation 96fcc4ea-7249-48d0-9135-ca1ca25c81ce · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

Lessons from Training Grounded LLMs with Verifiable Rewards Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:47.915712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:47.915712Z digest=sha256:b6b23a2e98dfc344a6546ff03dc5ca7546e8377bf79ef0ae1d72f86bfcb0c1f1

Observation 9fe83f6a-6464-4523-bb30-8bd7adb02a88 · outbound

This paper cites The Llama 3 Herd of Models.

Lessons from Training Grounded LLMs with Verifiable Rewards The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:48.240829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:48.240829Z digest=sha256:99cb635c1440e501835bf5136952aee5a42c302e9c65b1a40bacae619867f68b

Observation fb49ac51-e732-439e-9968-a068c4997491 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

Lessons from Training Grounded LLMs with Verifiable Rewards ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:47.976127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:59:47.976127Z digest=sha256:152feca1f14c5a86bfa6e7691e1e5b7457e3216023a59b8fc4509883babb76bf

Pith citing papers

Observation 49264500-90d5-4536-82ab-ac3512717904 · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.297833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:306880b7a56df6c2d2f3e273751976fe5a00161e3d30467e67ae61fd4b5828db

Observation 3f8a0b7a-1887-4d68-8f7d-2f074cdb4939 · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:48:29.575024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:48:29.575024Z digest=sha256:a10837ae17a6aa3a1c3b16203925e8fdba7991a07783bc061f3826e5c80bc059

Observation dc08d080-03c2-41cf-88a2-737c511970a8 · inbound

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs cites this paper.

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs Lessons from Training Grounded LLMs with Verifiable Rewards

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:47.206413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T08:07:34.112333Z digest=sha256:8c11a2051ca67535c9c4881c4b275fe56d6c1f4714b70ec63f390bdf2a64471c