Pith. sign in

Paper Citation Record · LEDGER

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

As of 9 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2601.18175.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.18175 v2

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:14:02.691163Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09029daa-0b2c-4c2e-86cb-a3943bd44ca3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.610152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.610152Z digest=sha256:9e8dd3622b52409e56b859ea67d05bd87bf1326f06e0b1b46603129ec3b6d4ed

Observation 5c8376d4-1666-4d32-9355-6b519a11430a · outbound

This paper cites The first equality is a definition of Lπ0 (π+).

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success The first equality is a definition of Lπ0 (π+)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.691163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.691163Z digest=sha256:dcfa63d8293e5991f89d1d30e1a4ed8593ba9466c8489ad78c98ebc6044866bc

Observation 62f79463-8cd1-407f-86c7-4435e5607a42 · outbound

This paper cites Bhandari and D.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Bhandari and D

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.023064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.023064Z digest=sha256:1b9a92cfc3f5220f2d58d3ec78b4ab090e9f4e3da4d1f3f17e70c9207fcafa16

Observation d529df4a-4268-4725-bf0c-dea6b014cf42 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Proximal Policy Optimization Algorithms

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.464290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.464290Z digest=sha256:1cc84bc2265fe9e9750b31f3baf953faf10bcd7f90736dfdd68baa0aa09ad89a

Observation 310d9c01-bbdb-4460-8592-6e9a51c52323 · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Training Agents using Upside-Down Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.544642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.544642Z digest=sha256:abaf0132b2103db6ee9380624c5999ba70335ab543c1fa4e07257cbb909cf281

Observation bcb313a7-679a-48e0-a15c-31d27690d890 · outbound

This paper cites an unresolved cited work.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Unresolved cited work

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.093983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.093983Z digest=sha256:d273ddf4afa14c112ec51d2426a9aff133b43832e6e6574087e0232f4fc7f303

Observation b281e5ca-9f1b-4672-afe5-f200cffd2c3f · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.253123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.253123Z digest=sha256:8960e5404f849306861cd51ae32d09e6e41beecb1a6900de41817ae233a16099

Observation 221a88ec-2bdf-4443-ac4e-ba250853f5da · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.352724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.352724Z digest=sha256:6c6f2a62c4a7d6323178022810f7a6f47e3f7e1087489456f16d8abd05ef1f50

Observation 58bcb1bc-0d35-4f65-8783-f27f9934e8e3 · outbound

This paper cites The Llama 3 Herd of Models.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success The Llama 3 Herd of Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.173867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.173867Z digest=sha256:e87fec0f73fb3ce69a44373b061ca575c370d4c9e4921a78c61cc1abda528808

Pith citing papers

No inbound Pith citation observations are available.