Pith. sign in

Paper Citation Record · LEDGER

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.08255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08255 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:20:31.002411Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d90394ae-8bf4-421a-81a6-97acf69a6280 · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.881986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.881986Z digest=sha256:9b6bce530c6b586684d24ab737986a537e8632777f7d0d2f73182a95fd63e9af

Observation 5dc7db82-d1bf-4bd7-a164-b663e62831fd · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Group-in-Group Policy Optimization for LLM Agent Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.898822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.898822Z digest=sha256:c8ea39aa864626653fc15331a0b281d118958d856ac9e831cf155de1626a2266

Observation fb549041-a873-461a-8fb4-e4a36e6be90a · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.903845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.903845Z digest=sha256:c005cc62bbcaee919fd4c5a04f549f4f95d5eef7c2ac55768dd51a91a48b6a22

Observation 5f552ab7-7564-4b60-9ea0-0f450dc67aa4 · outbound

This paper cites Shuo He, Lang Feng, Qi Wei, Xin Cheng, Lei Feng, and Bo An.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Shuo He, Lang Feng, Qi Wei, Xin Cheng, Lei Feng, and Bo An

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.915025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.915025Z digest=sha256:e7e78633169e13faf4fdce75ea45577dea1450bf03e32ae21f2fe0749fe3dd18

Observation f1bdb373-645c-4c2d-b478-0a7e8fc05de0 · outbound

This paper cites SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:20:31.418525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:20:30.920209Z digest=sha256:50ea152deb6ef62b8668818ec1b1225735903f1765615861804cea3d3524127c

Observation 20d99a4f-11f2-4944-b987-753ae45f1b6c · outbound

This paper cites Sparse Rewards Can Self-Train Dialogue Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Sparse Rewards Can Self-Train Dialogue Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.924964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.924964Z digest=sha256:449ff054e7d97a9b27530014c2a5f1864808f8f0168a6ea7a9f5202f826dfb56

Observation 951ad88b-d8f3-4151-8833-7524d42e3994 · outbound

This paper cites Let's Verify Step by Step.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Let's Verify Step by Step

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.935549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.935549Z digest=sha256:c03aaeade7636179217a2b99351b54364b0bee0a1860410118d44257d75b971e

Observation 0a1e9ed0-3451-4c16-857d-22bd399370dc · outbound

This paper cites A Survey of Temporal Credit Assignment in Deep Reinforcement Learning.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning A Survey of Temporal Credit Assignment in Deep Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.941220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.941220Z digest=sha256:3ea89287bb2e379ae5f6dd1f6942d9a95cff62cea3ad58b32df659acb07bd81f

Observation 9385ed41-af56-4b1d-8836-f1b0e782bdd0 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.951443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.951443Z digest=sha256:905130b19f9c9ea7d14dd2c4997da030d41b8023e1930cae7dc97d9c14d300e9

Observation ef0a00af-47bd-490e-adbb-aca16bc14dad · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.956539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.956539Z digest=sha256:425ac9317db1867e07ecfe11d54a2a196c398f3cd593a033b6aaedc0541f175a

Observation 284107c9-dd1b-4155-b113-ab125fd0a0fc · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.966366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.966366Z digest=sha256:b0d5ea8bf78da9a722a397f4a024a4a8bc77af557f55bddeb593b45bd495026d

Observation 15cddf84-14e0-4aa5-990a-35b31922c348 · outbound

This paper cites Chenglong Wang, Hang Zhou, Yimin Hu, Yi Huo, Bei Li, Tongran Liu, Tong Xiao, and Jingbo Zhu.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Chenglong Wang, Hang Zhou, Yimin Hu, Yi Huo, Bei Li, Tongran Liu, Tong Xiao, and Jingbo Zhu

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.977076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.977076Z digest=sha256:b214e44c085466ab555d96edf9bd2830595dfa5c11fce7d529566dade708083b

Observation 1e85fd38-0071-47a6-93d1-ab29797e7564 · outbound

This paper cites Qwen2.5 Technical Report.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Qwen2.5 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.982441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.982441Z digest=sha256:9d021a898643cc1f885100a79fd658cca0773174dc42d11830e5299c31989593

Observation ea9b2aac-c0ec-4774-ac53-d88079310eaf · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.987486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.987486Z digest=sha256:eac8653634c36f4a56e9f43722e863e7533d1d6c7779acaa054d5bd0bd3ec685

Observation af1ee5bb-ac15-47a5-a867-87075d555589 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.992560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.992560Z digest=sha256:a77d246c66ecfb013fe6f6cd5dd1092b1c995b2ef2068cf0fd7d2c2252863672

Observation 47e83f44-986d-4074-9e94-65de6da1013a · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.997565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.997565Z digest=sha256:fd36ee4c6cfbaa9329d789ced39b17b6e3f87449ae0b85d897b20a7e4fc44629

Observation 8165469f-0a38-4d57-b00c-5c60b6b73de1 · outbound

This paper cites SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:31.002411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:31.002411Z digest=sha256:08d87399bc97989e410324c706139bf6940e83a2664965cd55464f129d98abe5

Observation 40ac6398-d796-43b4-a6a3-86c19e1940d7 · outbound

This paper cites Hindsight credit assignment for long-horizon llm agents.ArXiv, abs/2603.08754,.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Hindsight credit assignment for long-horizon llm agents.ArXiv, abs/2603.08754,

Reference 1998

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.972193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.972193Z digest=sha256:8e1e4a956851fa6c3063253c6abe0cd160dc80ddb7cc36645651607550864880

Observation d5caeca4-f423-4630-a461-424602baec20 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.961438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.961438Z digest=sha256:5d893278a049c239825ad921da4b63b4ee8b787fd9ff752602ef1c7568acbb4d

Observation db603c01-77a3-440a-b6ab-8265efc89ac8 · outbound

This paper cites A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.909215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.909215Z digest=sha256:d7313ebd200384395c4058e2e006956c97473120f44130177d85e0176f59ab9c

Observation d6092662-97d3-483f-9352-65d6c435a8a9 · outbound

This paper cites Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.930351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.930351Z digest=sha256:b9c4f66549c6c84014e659b202ac875e0630c0d66eb2d87092a0ffa9aca66497

Observation e5ec0b2d-640a-441c-8c8d-848892cd796c · outbound

This paper cites Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.893043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.893043Z digest=sha256:77c6ca99cc2a7feb9d9c98eab7d4a2cd22b0ba37f0c6700bd125da57b5bd5d80

Observation bf9020c1-d71a-4fc9-9863-446c487e5f0a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:20:30.887706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:20:30.887706Z digest=sha256:93628ab852c7c5d31bd8aebe698c9aba8b82c0bb5448b0ddb57ec31169a7913e

Pith citing papers

No inbound Pith citation observations are available.