Pith. sign in

Paper Citation Record · LEDGER

Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.19595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19595 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:31:02.541976Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.153282Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7adeef12-ffdc-4e0e-8f99-e1f659d5405b · inbound

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening cites this paper.

Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:02.541976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:31:02.541976Z digest=sha256:1fbdf1e12ca27cf7052d4811a558d855f8282b8c6b70f55a7c017f0626b1470e

Observation a5a03000-ddb6-4db8-a58a-639361ea3c9c · inbound

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories cites this paper.

Large Reasoning Models are not thinking straight: on the unreliability of thinking trajectories Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:13.002002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:13.002002Z digest=sha256:475d985dc2353a4835062a1b91f4110f51e47114ddb95fd0da97ee05f64b401a

Observation dc56e706-31a6-4160-9c14-1dbfeae88be1 · inbound

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them cites this paper.

Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.663454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:52:01.663454Z digest=sha256:296d0c59269ce3375342dd358af7addc4bd06c80a282cfde4e5f4c8a7d5f0da4

Observation 8e997e85-8055-440b-a0ec-1515799f464c · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.184900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.184900Z digest=sha256:c717ec52168a27558c6c432b90b5106439adaaa562fc74a1c8f17a1877b1d1ff

Observation 0d78bad3-7a1d-4923-8247-095ae001384d · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.565055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.565055Z digest=sha256:517d1c4e4795a334a2528ec2898e2b66c93ff60c611f2fd4880c25ee453a8829

Observation d88b1852-3a90-4dcb-bae8-f1db68cd280a · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:19.949595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:d4e2182b191bad141179fe33529244501f5d8334358150daa76a44baa57dea9b

Observation 2a0054f3-e5f2-4003-b4dd-a3e9efec7692 · inbound

Compute Aligned Training: Optimizing for Test Time Inference cites this paper.

Compute Aligned Training: Optimizing for Test Time Inference Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:11.736090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:12:53.868648Z digest=sha256:c9b92ceddf3c6690b660b11a7a9a61cd896d9141f838ffb27e7eb582f8332002

Observation 05744ee3-eb1e-4fbe-a30f-504c9100dca0 · inbound

Compute Aligned Training: Optimizing for Test Time Inference cites this paper.

Compute Aligned Training: Optimizing for Test Time Inference Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:44:04.622393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:41:49.623026Z digest=sha256:33cfc866f807fe965ac4816c9d27a5a8241eddab3db767da4c35fb86d93bf308

Observation be2fa0bb-81f2-437a-a626-b5ae1935083e · inbound

What should post-training optimize? A test-time scaling law perspective cites this paper.

What should post-training optimize? A test-time scaling law perspective Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:21:23.886212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:20:18.754351Z digest=sha256:72eceaf27fcf33028bcc6e741e8f978c5d21ca8d23320ad197abe44c7b853fa0

Observation 4bdf7ce1-4a7e-44ec-a1fb-7202912bab31 · inbound

Finite-Time Regret Analysis of Retry-Aware Bandits cites this paper.

Finite-Time Regret Analysis of Retry-Aware Bandits Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:03:59.243572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T06:01:17.988127Z digest=sha256:2398500d15d2ea91c4679afe6c84dec7f6be92ca3c2dc13d30d1a83623a63896

Observation 21aa595f-745b-4e8d-88c1-dd996626182f · inbound

REVES: REvision and VErification--Augmented Training for Test-Time Scaling cites this paper.

REVES: REvision and VErification--Augmented Training for Test-Time Scaling Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:15.155048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T21:14:15.337979Z digest=sha256:390d342e60df942c48761491b7cb69a931832b88b7b8d5fbeb646379bcc57df0

Observation 448581c1-4c99-47da-9b30-d6f176b30e11 · inbound

SPIRAL: Learning to Search and Aggregate cites this paper.

SPIRAL: Learning to Search and Aggregate Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.154935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T08:29:22.303745Z digest=sha256:efaa7ee957b3108a59233387588a810ee2cf2fd25fddaad5a9a7c6c60b7ed1f6

Observation 8b568f87-3394-4b10-8ef4-3182ee70016c · inbound

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL cites this paper.

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.761591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-03T20:59:57.539909Z digest=sha256:c5f4828cb4ddd5346220fd15a10a080ebb9fbb5cfa72153ffc4498b07c907edc

Observation 77e2b571-845e-4398-8209-3bd9b587979a · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.754675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:d777b4ed11374e5d6f27541067af874c914a229728c95f9586af426bb91372b9