Pith. sign in

Paper Citation Record · LEDGER

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2607.13988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13988 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:10:51.670313Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:23:14.662916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T23:23:14.803053Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f21c86e2-fb41-4351-90c9-0a41ce7883ce · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.504547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.504547Z digest=sha256:d9a6f16b5ab888d448499edd98280c05115633291548728dbb6fc185e90f8907

Observation f73e6d76-0cd4-4ce0-ae26-c611a18fbd3d · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.514445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.514445Z digest=sha256:9d86bd13495711be7e3bdb8018b3380a95c107df3db81e324480a692394f4c60

Observation 9c6713de-62f7-4e67-a628-77c57601c653 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.519075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.519075Z digest=sha256:b638a9dc2c693d97b87d9a2397805dd89befec547cc6dc1178ab3138be4cad27

Observation 205c024e-c1ec-4be4-bfb0-36357f6de419 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.523255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.523255Z digest=sha256:4efbe5ce127690ea65b84463fc6b5c29b0d293fe355834386331c124194f03d5

Observation 7b01ea83-05d3-4bb3-a946-884540eef88f · outbound

This paper cites R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.528187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.528187Z digest=sha256:10aa3ad8259fc01fa23229e5cd4619e40390ffea950b6b71a6abb0e3ae85f908

Observation fe7c2143-8f14-4272-87fe-6dcc857aabba · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.532947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.532947Z digest=sha256:405c4872af22a7dd730341d16c8ae57b5a23de74682a14eab054be66168e223d

Observation 5d5f25e4-d7ea-466e-9319-ab5f1dd8c5d3 · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.537374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.537374Z digest=sha256:d82c4c384ae7cd12b2315543b99dd4cd67c4af5de2b9570724e77909af007937

Observation eca1964b-ea13-408a-ae1c-5b25746d26d5 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.541964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.541964Z digest=sha256:e0e1a83af2a917c68cdd25306d07f4ace4297990209a0c23989e8a35698860bc

Observation db9c136a-5340-4fc8-96ad-edeffac6f133 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.550887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.550887Z digest=sha256:4f7ef24ba56c208ff11e63f03998ea1cdf38c9b657a62f2ea250d69f82b3cb3a

Observation f1c36c4f-a132-4663-99ac-4466b198786c · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents WebCanvas: Benchmarking Web Agents in Online Environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.555037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.555037Z digest=sha256:2a29ca6bc6053ca63f61db89917ab2956faba7d385366b14a959b9b344cb884d

Observation ddf29d8b-df17-406d-8ace-795cdff9af3c · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.559083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.559083Z digest=sha256:f0a3024c044bf2ea9a88a4f8864d9fcd9d0686f150cbf23a71dd79a5a6945a3c

Observation 36014a80-2955-4406-be10-e95963b257ce · outbound

This paper cites Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.564249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.564249Z digest=sha256:0614dd6bb815baa34cf891b5606b5075bd1cad7bbaa594727d3d985261a8918b

Observation 6df697c2-8b11-4e9f-a747-90d051749517 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.568458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.568458Z digest=sha256:3042166d2cb8687497057252ee08cd4f82763196fb2f537c05b25a520954d2d0

Observation b6fecf0b-61e0-446b-868a-2c9019b23f4a · outbound

This paper cites Proximal Policy Optimization Algorithms.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.573315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.573315Z digest=sha256:4a3eed52924300c7d3a41e6b89a19a4859195fae650e67f45fa086ee7d80c0a3

Observation 0668243f-b7a8-418c-b483-cd30bab93933 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.577885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.577885Z digest=sha256:a9808bceee8ad967042c42cbf60f12eb9ab6dd1921a429f8204ddc53c55148b2

Observation 5210b9de-fd48-4c7b-af8f-257af5a6c803 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.581899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.581899Z digest=sha256:7418b310f987a95e26ea9f9b1e6791384dbd1e028fd6f7584e2ab68767103a46

Observation 8668f7ce-d362-405f-b796-e58bf922a197 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.586270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.586270Z digest=sha256:e500ace0d8e6b58527543e4f0501e920a8d607062fb4e6798b4ed76ed7e7ec1c

Observation db20d725-16fc-4756-8dff-c428a55c28e6 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.590861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.590861Z digest=sha256:cf01699c44b72c4a6a18fcc27392768f80a8c2dabcb604a2e82055ef51777597

Observation 9aa12024-a25e-4e43-98da-c46e8c6782fd · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.595241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.595241Z digest=sha256:21a0c23c50e1b3703873a8b7b0e96cdb034e5d94c70cc835ed2bec49bbf2fc9b

Observation 2dd274b2-41b4-4893-bfd7-fbfea6ab602e · outbound

This paper cites Tongyi DeepResearch Technical Report.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tongyi DeepResearch Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.602824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.602824Z digest=sha256:4d3903a496bde26a64df44d4ebf4cee131095d7cfaa0c032a10ed2187a0f9471

Observation 0626a104-8a73-416e-ad0f-7c545f9ba511 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Solving math word problems with process- and outcome-based feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.607019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.607019Z digest=sha256:09b568beec8837074f3d6b71256e5ab660801435e9907be994169cc07fffe15a

Observation ea196cc6-889b-4372-ac6e-50370063138b · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.616187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.616187Z digest=sha256:895a67072f51af85e9ef60c9f49933d19f29742547b2978b562f1b61cc760f84

Observation 06b7549d-5b6a-442e-a256-bac21fa3aea6 · outbound

This paper cites Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.620840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.620840Z digest=sha256:ab4daa37370343a7a032cf03ecc290aedc28501cdacf9ab9e7365b90186036a4

Observation de5ac6b2-5944-4264-9278-b9116f52d45d · outbound

This paper cites Qwen3 Technical Report.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.624897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.624897Z digest=sha256:fe7b1db09e353efb74ed212d3cb7d1ea02c5253cf3ba3ba8b921237a90dac26a

Observation b368e4a0-6da4-476e-ab01-3a1481441569 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.628833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.628833Z digest=sha256:1c636658c6a95d69ff8cff76b278b6923f8d1ed294b083c0d6d948ae7bf8ba9b

Observation e2d46ae3-7b17-4cff-a041-dbfd573a0dc6 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.632686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.632686Z digest=sha256:dcd4c35b6234b2fc85aa8c99988c41dd55aa3e853965fd51cf9cb573a76d882c

Observation a2d4769b-f6df-498a-a643-760a44d83a51 · outbound

This paper cites Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.636722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.636722Z digest=sha256:1eb34a53e652b23d87c44adff22df377419c96b29479fad611dd6e23906575f8

Observation 28bdb2ca-b582-4aa4-96d8-2fc7fe070402 · outbound

This paper cites Free Process Rewards without Process Labels.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Free Process Rewards without Process Labels

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.640636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.640636Z digest=sha256:5e857c94e787a44790f5826214c39ae1e87f2d2649ce47f275b15c4b13205056

Observation 84450d16-3158-487c-a59b-c3a2249d38b2 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.645371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.645371Z digest=sha256:2426da0aeed2e19d15e451a7cbfe65d725258902836f6f54d6f3397b79a6682b

Observation 70dade9b-c088-4898-b2ed-d25004e24f77 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents AppAgent: Multimodal Agents as Smartphone Users

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.649959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.649959Z digest=sha256:1e548cfe8feb51aea1ec78abea0a2500ccab76f71dc3490bee6e51d605b8e5aa

Observation 8a792367-ef91-4ae3-8638-051ccff89d0c · outbound

This paper cites Zhang, X.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Zhang, X

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.654620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.654620Z digest=sha256:cfac91964d1bc7e22efdf360cb8123fa4953863bce3ffb13015cd95c3e5b86ff

Observation 1111c82e-4abf-44cd-a07e-b9c6cf50c318 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.659361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.659361Z digest=sha256:59bd1dd461812ab39d561707b88ded4d5bf99f065a3cdf0bc557ca243970c49f

Observation a88927cf-ac3a-4722-93ab-cf50b951a53b · outbound

This paper cites at the Jakarta Expo Boxing Hall.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents at the Jakarta Expo Boxing Hall

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-08-02T03:10:51.663804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.663804Z digest=sha256:6dc05acb706dd8164ce2515d59800bd21e42672e98e6956d215cd619ad925489

Observation db19753c-bdb6-48ea-ab61-d39caa622bdd · outbound

This paper cites influxdata company name change from early 2010s to later name.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents influxdata company name change from early 2010s to later name

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.670313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.670313Z digest=sha256:6f9b22ca9884ddcbae431e61c3c5cb6a82e94cfeca96a232e8c538fc5fc7c4e4

Observation 25a870da-d1a6-4645-97e1-b2f99bbfada5 · outbound

This paper cites Tan, X.-W.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tan, X.-W

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.599182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.599182Z digest=sha256:7ab2c07c3d5d2adb12ab1e34c224bd21dcfee28e0d6d22ec963744483116c3f0

Observation 74c730fc-3178-4f5e-97f1-0176e0f8fc29 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.611387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.611387Z digest=sha256:ec24bcd9d46c269525bd9e464834dedf57b3fe0b1996bc4dd8876e8dafd0193d

Observation a6de5c47-ed61-469e-a188-460913426f03 · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.509228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.509228Z digest=sha256:871bfa365e45e0be7934e4adcdafed73a013d2e14d7fd582fe9a8403f453db77

Observation b7ec7ad2-4275-476e-8f7b-7e0c14e056df · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.493672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.493672Z digest=sha256:d1f92de7ca9316c2b0665752133c74a7f41c3cf634251d98a17c18a93219edf6

Observation 09568c4b-06bc-4695-b267-50ef40445613 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.499866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.499866Z digest=sha256:887c9b31816f4b5450f6662592c6c4b4cf8e0db79a3916f451664fa22dccdc91

Observation 61784b62-6e4f-4ba1-904f-038bf903a1b3 · outbound

This paper cites Let's Verify Step by Step.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Let's Verify Step by Step

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.546216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.546216Z digest=sha256:cc4061da4ce2a084f734fea427146e045447aec180800cce3a983f8ff163ac53

Pith citing papers

Observation 344e4dc4-f6fc-4f6f-a7eb-7302855a467e · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:23:14.809554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T23:23:14.662916Z digest=sha256:69288bfe071e69ea953ffb0d0abee4951f04e98b1cf2294d1afa22988e16a390