Pith. sign in

Paper Citation Record · LEDGER

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2507.20150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20150 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:56:41.084268Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:19:34.840954Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a7b6546-5f47-4b35-ac84-a2438a216ce7 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:40.998568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:40.998568Z digest=sha256:ba0ab99ce8e3faecba3f373fe3ef1cf1aa81930f4e972ff52d3e432398ed03e4

Observation b1e01e78-4e71-473c-94a6-2deb1c3f5060 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.008415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.008415Z digest=sha256:d2841ca2267875f630096ea7d4e99255f55c8b42b71fb194435d96f91b814936

Observation ef49e7f3-9dab-4f2f-8752-72992aa824ae · outbound

This paper cites Defense Against Reward Poisoning Attacks in Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Defense Against Reward Poisoning Attacks in Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.018945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.018945Z digest=sha256:59dc43affabef8f88b3292592a49bb057a2d3c7622d74bbff3f6b812c81e673a

Observation 08201b88-4bc2-466f-99b0-7405c37ae0b0 · outbound

This paper cites Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.028809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.028809Z digest=sha256:f54ce41e0ddb78e0b11369fd9f567e121ab0b82dc8dd103c19e74cc9961c5749

Observation 0fc187ff-f529-490d-ac32-1ed0bab4e3bd · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.033080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.033080Z digest=sha256:9d63f2b8d2772c70ab919d5f45ae8ac2e3e94a47726803877d8bc39011044e68

Observation 0ae26018-5913-4345-bf2f-319ccf727136 · outbound

This paper cites Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.062832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.062832Z digest=sha256:0a58d03ab0cc6b1c8db287928830b5f92f9db5dbb352508644052c9274a97984

Observation 00fe4cc6-6b70-4895-b35d-e92c2d86bd2c · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.067681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.067681Z digest=sha256:af7110dde29a2274e2afacf0bdfb2cc11b30d7b606e48e230cebb47c257e48e2

Observation 803f4fed-0a88-4288-a3d0-ba591b2e511d · outbound

This paper cites Qwen3 Technical Report.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.075895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.075895Z digest=sha256:2fb065a52218ac1ed10b891ef78021b4f90373dcbeb372c4c1ac8dc83581a257

Observation 5d513ce0-0a9e-4640-9569-6636b90a197a · outbound

This paper cites an unresolved cited work.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:56:41.491657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:56:41.079766Z digest=sha256:b8d6aeed33190e1ba3da4df77b354864191dc1e588f57f53a3b21ad5178081f8

Observation 15b88e43-d525-4a12-8c13-5251498e696f · outbound

This paper cites spurious reasoning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models spurious reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:56:41.477462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T17:56:41.084268Z digest=sha256:7a6c4c4aa9d4593453c52616f7c0a36e53a056b5b2a3546015d801917fd66bd9

Observation 78913d69-3cbb-4bef-9023-1c2d92327ffc · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 1963

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.023330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.023330Z digest=sha256:d39f95910d6ea781700752cf1fa7f0bb114ab65d3eacc2d97bad24485387e711

Observation d70c39aa-a11a-4a9d-816c-329a2cd21bb1 · outbound

This paper cites A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models A survey of efficient reasoning for large reasoning models: Language, multimodality, and beyond

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.050554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.050554Z digest=sha256:9d423d64cb073fe880f2596439529f803b8df57a216e881bfdaf2364ca73a7cc

Observation f50b2653-10c6-49bf-975f-081d70d6f876 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Proximal Policy Optimization Algorithms

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.054541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.054541Z digest=sha256:36d5911c7b990bfdc73cc95c64a0549188078cd550dc2651f99f8a53c5f06792

Observation 4eb82b6c-d6ae-483f-bd21-7cc8eca0f49b · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.058636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.058636Z digest=sha256:039f20ba014de887bfafb6207276ea3606c41b31b14666cb92a8d200309843c8

Observation 6183f9fc-4be1-4bdf-953a-7cd4438515d7 · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.041784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.041784Z digest=sha256:cafccdd0d4d7b94dcb4313f944b18918ed4e462d761d255bcd09877fdea716d5

Observation fd17ebcf-4b39-4eac-af62-47d1c12e5df4 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.071930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.071930Z digest=sha256:91f22ed11da524f406cb5fbc9c4fd9d749ab86232f4377a6e1910d04c007c270

Observation 4c93f222-a987-42fb-b239-849d117f537f · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.014266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.014266Z digest=sha256:aa28291b7038ad76bc5f750c423149184b57fd31dd9612f4bb363c39803a2790

Observation 0c63d43f-c85a-4e12-882d-bfe71776b201 · outbound

This paper cites MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.046033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.046033Z digest=sha256:145ed9b4511d3e47ac6bd14a8015864bccb1b5fb8efdc33ff54ae5e0cfbfd652

Observation 3826fd40-049f-4b4b-bdde-f91ab48284b9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.037293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.037293Z digest=sha256:fe31d886f2a405c453f1be57184247e2cb780d943c25d3f5a69a1ac9fd2605e7

Observation fc062070-efd6-41d5-ba8a-75fbb2cef70e · outbound

This paper cites Qwen2.5-VL Technical Report.

The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T17:56:41.003583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:56:41.003583Z digest=sha256:a9bc552ec943641aec572f81059a49fbad39d57063b25a97299d29da018af037

Pith citing papers

Observation b86ffbcb-f091-4df2-8a86-c7aa1ba77b3b · inbound

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation cites this paper.

Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:34.840954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:19:34.840954Z digest=sha256:0f8c2f849d1ad5eb3e39f3a07caf72ce771bdf67cceec5f6fc6a3c3babed65c7