Pith. sign in

Paper Citation Record · LEDGER

Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.19595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19595 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:02.541976Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.153282Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7adeef12-ffdc-4e0e-8f99-e1f659d5405b · inbound

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening cites this paper.

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:02.541976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:31:02.541976Z digest=sha256:6633399ebb7ea5697b36385f007912b1e34498d652c5386062b85eb52d401f9b

Observation a5a03000-ddb6-4db8-a58a-639361ea3c9c · inbound

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories cites this paper.

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:13.002002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:13.002002Z digest=sha256:36f29450899420fa29428db583bd4d03dc33fb2436f4caf8a99fa372cdaa3327

Observation dc56e706-31a6-4160-9c14-1dbfeae88be1 · inbound

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them cites this paper.

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.663454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:01.663454Z digest=sha256:deb7ec8c5702a95ea2bb6db4913d79e9790159f2e3cbdd2ed2d719ece34e1559

Observation 8e997e85-8055-440b-a0ec-1515799f464c · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.184900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.184900Z digest=sha256:9f243c5fad94f7ce28bdc7c7cebac5009a796e633b97e6f3009c02c02aa7cc7e

Observation 0d78bad3-7a1d-4923-8247-095ae001384d · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.565055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.565055Z digest=sha256:0e828bfec359f85fc48e298ca66843b35d980e4e86a2729562e87af46769af1f

Observation d88b1852-3a90-4dcb-bae8-f1db68cd280a · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.949595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:18ce067e8f45b79002f9e20d8a444d16620005b2d7f0e7128140be5903868c3f

Observation 2a0054f3-e5f2-4003-b4dd-a3e9efec7692 · inbound

Compute Aligned Training: Optimizing for Test Time Inference cites this paper.

Compute Aligned Training: Optimizing for Test Time Inference Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.736090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:12:53.868648Z digest=sha256:c2d31cabd90c8ed790869cca43003c8814afc1483ddb6ca5dc4dffbc6cdfe06e

Observation 05744ee3-eb1e-4fbe-a30f-504c9100dca0 · inbound

Compute Aligned Training: Optimizing for Test Time Inference cites this paper.

Compute Aligned Training: Optimizing for Test Time Inference Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:44:04.622393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:41:49.623026Z digest=sha256:a5b512295e8c9b3c43c987d073cddfe783006e03d2c78727cf2fa484ce3624c8

Observation be2fa0bb-81f2-437a-a626-b5ae1935083e · inbound

What should post-training optimize? A test-time scaling law perspective cites this paper.

What should post-training optimize? A test-time scaling law perspective Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:21:23.886212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T05:20:18.754351Z digest=sha256:ad8f7d4774fee728392dab159d9d71dfc771ab583260b4c4f3f08b0e03733108

Observation 4bdf7ce1-4a7e-44ec-a1fb-7202912bab31 · inbound

Finite-Time Regret Analysis of Retry-Aware Bandits cites this paper.

Finite-Time Regret Analysis of Retry-Aware Bandits Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:03:59.243572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T06:01:17.988127Z digest=sha256:313da6a671adffad686724b2e577211ce29c5a43fc11d4b8548a37bfd4e285b7

Observation 21aa595f-745b-4e8d-88c1-dd996626182f · inbound

REVES: REvision and VErification--Augmented Training for Test-Time Scaling cites this paper.

REVES: REvision and VErification--Augmented Training for Test-Time Scaling Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:15.155048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T21:14:15.337979Z digest=sha256:577c9eb48278b19677ed89a53794444f939f7ab7737f35c02503f8187a08a4d5

Observation 448581c1-4c99-47da-9b30-d6f176b30e11 · inbound

SPIRAL: Learning to Search and Aggregate cites this paper.

SPIRAL: Learning to Search and Aggregate Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.154935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:29:22.303745Z digest=sha256:3f3a1b7b62958b1074256d8459c8c0d3f37474cca6fbfcc106248541c95e4b73

Observation 8b568f87-3394-4b10-8ef4-3182ee70016c · inbound

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL cites this paper.

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.761591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T20:59:57.539909Z digest=sha256:e59601c5dcfb095c6de8cc88600df6b404bc40b172ab60b2691373f62af3ba13

Observation 77e2b571-845e-4398-8209-3bd9b587979a · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.754675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:8b5334d3f15e762ef02772ec2597062e35c72867208893b764a0a5151cba5bef