Pith. sign in

Paper Citation Record · LEDGER

TCPO: Turn-Level Credit Policy Optimization

As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2608.01667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01667 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:23:14.694452Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b721a9c-415e-4ccb-a10e-af5f6890dec9 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

TCPO: Turn-Level Credit Policy Optimization ToRL: Scaling Tool-Integrated RL

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.645199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.645199Z digest=sha256:10ba5fad9c13bae57d6f8f6f00ea83fb02eeed13fe76587cb712d6feb732c74a

Observation 92693cd6-c1ba-428b-99a1-50faf8fae3ce · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

TCPO: Turn-Level Credit Policy Optimization HybridFlow: A Flexible and Efficient RLHF Framework

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.653931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.653931Z digest=sha256:88eccd0f6be448a8851d6a2c6fd7bd2b5545a1127278c57b534e7f7126006950

Observation a1b0240b-e3f3-48f3-9235-56130f21c21e · outbound

This paper cites Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning.

TCPO: Turn-Level Credit Policy Optimization Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.658329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.658329Z digest=sha256:022344df5300c0cb5ce2203954dc2bb98a4bf468222a2a9939fdb5a416ab39a8

Observation 344e4dc4-f6fc-4f6f-a7eb-7302855a467e · outbound

This paper cites TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents.

TCPO: Turn-Level Credit Policy Optimization TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:23:14.809554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T23:23:14.662916Z digest=sha256:6f4a6dac9db6967fe174f3b1d2f24c3a8e31b0de4f01441a87ef7e8128c99ef6

Observation b824a36a-6db2-479b-9853-46c1de6bd896 · outbound

This paper cites Exploiting tree structure for credit assignment in reinforcement learning with large language models.

TCPO: Turn-Level Credit Policy Optimization Exploiting tree structure for credit assignment in reinforcement learning with large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.927639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T23:23:14.667862Z digest=sha256:6cd622e2c24509e11aa281d1a32fb686b8ed1dac0a8d7344827c7026eb024b6f

Observation fdc7c8d9-5d4f-496c-a375-929f8d2adcdb · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

TCPO: Turn-Level Credit Policy Optimization Solving math word problems with process- and outcome-based feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.671864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.671864Z digest=sha256:8545fe321e854a140821ead9ba37824d2882ea2be6271e94d033d577fb5072f3

Observation a172f626-5a46-44ad-b41e-6b49a319ce06 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

TCPO: Turn-Level Credit Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.676492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.676492Z digest=sha256:acedce9f0d60709baf4fa4bc71d3f0ab16f6b7e7f7f7e8d1458abce21b9f9d43

Observation efce9824-9194-48da-8992-39b850606df8 · outbound

This paper cites Rein- forcing multi-turn reasoning in llm agents via turn-level credit assignment.

TCPO: Turn-Level Credit Policy Optimization Rein- forcing multi-turn reasoning in llm agents via turn-level credit assignment

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.912954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T23:23:14.681238Z digest=sha256:63dcf282e180d6a3658de8bcd276f91c4e3d40ef44f0c446b90a74b46c487e44

Observation 8ded1493-716c-4061-a89a-18136fec6c84 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

TCPO: Turn-Level Credit Policy Optimization Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.686054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.686054Z digest=sha256:52655ebe330b643ed0b49c93139504c15ea9b6b042e02307d2cd965950c16786

Observation 29aec9b2-098f-4024-ada1-5ec02895ecfb · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

TCPO: Turn-Level Credit Policy Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.690152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.690152Z digest=sha256:a9167d8cbe1866b65509c9728b4d103c02f4acbdc31a108b4cd5d0527b609a5a

Observation 2ba18ea4-ed78-43cd-bc54-deb51ad88f44 · outbound

This paper cites At 2po: Agentic turn-based policy optimization via tree search.arXiv preprint arXiv:2601.04767,.

TCPO: Turn-Level Credit Policy Optimization At 2po: Agentic turn-based policy optimization via tree search.arXiv preprint arXiv:2601.04767,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.694452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.694452Z digest=sha256:c25c541ea6794779af7c1ddb9be17c214e750ae91554d52e14e56931ab5c0ac0

Observation 74360757-7012-4b6a-9243-770addcc35bc · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

TCPO: Turn-Level Credit Policy Optimization Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.629933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.629933Z digest=sha256:d3ccb7856b9cbe172bb18b2952e33e8622eeeff625eee5cfb48c56404f6dd11a

Observation 3f450e9f-b910-4ae1-be33-eae85bac52ea · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

TCPO: Turn-Level Credit Policy Optimization An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.639942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.639942Z digest=sha256:1e45b3e3007adb12cf450311a877c1390be044ab1c9a71b22758177d5d2a7711

Observation ee6508c6-f791-4ad9-923b-8dccee513ef9 · outbound

This paper cites Let’s verify step by step.

TCPO: Turn-Level Credit Policy Optimization Let’s verify step by step

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:23:14.941960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-04T23:23:14.649732Z digest=sha256:5d6a90bf2841d11970911027d0ac38cfacb4d03c9e0a76871e1e7960bc68b4cb

Observation 9eac6ccf-3610-4ef0-9123-e2250fd1828d · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

TCPO: Turn-Level Credit Policy Optimization Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T23:23:14.635075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:23:14.635075Z digest=sha256:7e03b0b4a3bd5343e632b9c60e8238ef104b8cbd72bfacc6158d8302c8862868

Pith citing papers

No inbound Pith citation observations are available.