Pith. sign in

Paper Citation Record · LEDGER

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2607.13988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13988 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:10:51.670313Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:23:14.662916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T23:23:14.803053Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f21c86e2-fb41-4351-90c9-0a41ce7883ce · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.504547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.504547Z digest=sha256:a21eca1e32c590530542ffb4310dfb9be446860f66d70147da01b217e4aaba7e

Observation f73e6d76-0cd4-4ce0-ae26-c611a18fbd3d · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.514445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.514445Z digest=sha256:a045833e0a2c1de910bfac8551808e7117a73ce25bf22c646311dac5e2068306

Observation 9c6713de-62f7-4e67-a628-77c57601c653 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.519075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.519075Z digest=sha256:64d6dfdf20f7ae41aa51a95504bd5d51c321fe9a9187c7f4c8f2818418cc4364

Observation 205c024e-c1ec-4be4-bfb0-36357f6de419 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.523255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.523255Z digest=sha256:984bc2d81a7fcbee4e516f32fba7441291d4aee5f5ff59e0949bb43b06c112a8

Observation 7b01ea83-05d3-4bb3-a946-884540eef88f · outbound

This paper cites R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.528187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.528187Z digest=sha256:cc40943a7924d26197d964ad242d0bd3b8522b8c5b7cb676c82f3b1849b6064e

Observation fe7c2143-8f14-4272-87fe-6dcc857aabba · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.532947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.532947Z digest=sha256:79d993426ce3e84a4b0bdbf2f8eb8d628b8bb3f0552fb18303b74dcf085e7eab

Observation 5d5f25e4-d7ea-466e-9319-ab5f1dd8c5d3 · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.537374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.537374Z digest=sha256:4892ad88d7809f9553793c14f66af18f94e2253f7b2622826956ca8d1546a2b5

Observation eca1964b-ea13-408a-ae1c-5b25746d26d5 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.541964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.541964Z digest=sha256:e7058f6588ecae754eb16418eccde953878d938ca71520250bb36086b7d5ed9c

Observation db9c136a-5340-4fc8-96ad-edeffac6f133 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.550887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.550887Z digest=sha256:8ac85e22036cdbdc72c5255720259fc7af598ffbdc0e38b8bda9078421515ab0

Observation f1c36c4f-a132-4663-99ac-4466b198786c · outbound

This paper cites WebCanvas: Benchmarking Web Agents in Online Environments.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents WebCanvas: Benchmarking Web Agents in Online Environments

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.555037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.555037Z digest=sha256:3c68d69cb224364b63283ed3630bfd89a2eb23e51ceac005d8dd20e49fd72ba0

Observation ddf29d8b-df17-406d-8ace-795cdff9af3c · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.559083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.559083Z digest=sha256:0257cd13e5072d008c3ca14543be74bd4eaed0cfadb13b90e33f4096afa39bf2

Observation 36014a80-2955-4406-be10-e95963b257ce · outbound

This paper cites Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.564249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.564249Z digest=sha256:485c8fbefc43b4db741dc1edb71441b187728610a717e7a686452ff78683a128

Observation 6df697c2-8b11-4e9f-a747-90d051749517 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.568458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.568458Z digest=sha256:4d7e0f536721995a9b1d4d0c914c8b408637fcda484a93f0d7dc961f45fcf5ab

Observation b6fecf0b-61e0-446b-868a-2c9019b23f4a · outbound

This paper cites Proximal Policy Optimization Algorithms.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.573315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.573315Z digest=sha256:b06d9b864355a854672f1a44d0481c5b92863145def1a3cf2cbd8e5eb4a4df3e

Observation 0668243f-b7a8-418c-b483-cd30bab93933 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.577885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.577885Z digest=sha256:8696d4430feaf3bf1490975e069e7eb04dfbb7218511ba354f00d0328dca9ded

Observation 5210b9de-fd48-4c7b-af8f-257af5a6c803 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.581899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.581899Z digest=sha256:bc75747e45422ed5894e62c223ad88306083bc203f6192c972c0f460f81d148b

Observation 8668f7ce-d362-405f-b796-e58bf922a197 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.586270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.586270Z digest=sha256:2dfec47b0488c1b4d890b0e1e4c166cacede6fc63eb32542057ee7b8d0b18fc2

Observation db20d725-16fc-4756-8dff-c428a55c28e6 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.590861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.590861Z digest=sha256:3689d362fc9d43b7f618b779e97b45fdbd886373a0dfb40aea72343b002a570e

Observation 9aa12024-a25e-4e43-98da-c46e8c6782fd · outbound

This paper cites ZeroSearch: Incentivize the Search Capability of LLMs without Searching.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents ZeroSearch: Incentivize the Search Capability of LLMs without Searching

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.595241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.595241Z digest=sha256:5b64a8e9d4b7fb281f25da4ce3766fcd3b6fcc0ee65d62364f8253178998970a

Observation 2dd274b2-41b4-4893-bfd7-fbfea6ab602e · outbound

This paper cites Tongyi DeepResearch Technical Report.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tongyi DeepResearch Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.602824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.602824Z digest=sha256:e771ee5bea94757ddc0705976804ec86487e65651060dfa914a3203b12b5cdb1

Observation 0626a104-8a73-416e-ad0f-7c545f9ba511 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Solving math word problems with process- and outcome-based feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.607019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.607019Z digest=sha256:4f6d6bf9d5c4be9f25094580dcd7a05943797a5bbe6d45214cfa2da1b3114dc8

Observation ea196cc6-889b-4372-ac6e-50370063138b · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.616187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.616187Z digest=sha256:1e96e69c5d2e1f93d65db1c0c647abea683e94ae289774dd9c4c4a2e02bc7b37

Observation 06b7549d-5b6a-442e-a256-bac21fa3aea6 · outbound

This paper cites Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.620840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.620840Z digest=sha256:3132c15b3525f1144b10820b9dacca924d2c6a8731724f545f7cc7927006607a

Observation de5ac6b2-5944-4264-9278-b9116f52d45d · outbound

This paper cites Qwen3 Technical Report.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.624897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.624897Z digest=sha256:0c471890291dd4330755b1bb637022e71d83a0b99b663a0bf09949ecd228ccfc

Observation b368e4a0-6da4-476e-ab01-3a1481441569 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.628833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.628833Z digest=sha256:cc08c9d48e0fd7adf7e7b6e5b13d4438d0196814ceb8cc9202ff473129d4fd84

Observation e2d46ae3-7b17-4cff-a041-dbfd573a0dc6 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.632686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.632686Z digest=sha256:1af7e9156e958ee48bed7710628f596ff86103c3a5acbc55c71bab2dfeeb6cda

Observation a2d4769b-f6df-498a-a643-760a44d83a51 · outbound

This paper cites Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.636722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.636722Z digest=sha256:80dd0ce380fea676178f2c37a8902b7ddd1c48c11c89167642b4c132def4ad73

Observation 28bdb2ca-b582-4aa4-96d8-2fc7fe070402 · outbound

This paper cites Free Process Rewards without Process Labels.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Free Process Rewards without Process Labels

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.640636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.640636Z digest=sha256:0232ed166d5e04b83ccc223129f8c4d90d7d73a09786c0821049ce2c9810ea45

Observation 84450d16-3158-487c-a59b-c3a2249d38b2 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.645371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.645371Z digest=sha256:79adc9946051fa812c4cc66a10cc970ae27b97a46cb54717e66d3dd56e28bd63

Observation 70dade9b-c088-4898-b2ed-d25004e24f77 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents AppAgent: Multimodal Agents as Smartphone Users

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.649959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.649959Z digest=sha256:89cc6fd24a87933283086ffcde159e0645cacd6171b10de588bd810c90be85e1

Observation 8a792367-ef91-4ae3-8638-051ccff89d0c · outbound

This paper cites Zhang, X.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Zhang, X

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.654620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.654620Z digest=sha256:9d284708a417db53cb9001d3cb93656bfd6a6b2f536c59793c39fa19e4677e54

Observation 1111c82e-4abf-44cd-a07e-b9c6cf50c318 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.659361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.659361Z digest=sha256:adb3d041a6d0f4122f2f1181b97e79bbe4936a727acec6d9d1ffcb86e78ca42b

Observation a88927cf-ac3a-4722-93ab-cf50b951a53b · outbound

This paper cites at the Jakarta Expo Boxing Hall.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents at the Jakarta Expo Boxing Hall

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-08-02T03:10:51.663804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.663804Z digest=sha256:5a49e9c517d0c3c09539548c139d8d8700c1f871b9c73d4742d8361bae24a4d4

Observation db19753c-bdb6-48ea-ab61-d39caa622bdd · outbound

This paper cites influxdata company name change from early 2010s to later name.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents influxdata company name change from early 2010s to later name

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.670313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.670313Z digest=sha256:aa918356d8c3ce1e43cda42dd0770440fd950f0131cabbcb7e024c2a9bb72aec

Observation 25a870da-d1a6-4645-97e1-b2f99bbfada5 · outbound

This paper cites Tan, X.-W.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tan, X.-W

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.599182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.599182Z digest=sha256:4a303211cf855ff62eef806915da0dbfa59c5a648a29438b32e941a861f45ef2

Observation 74c730fc-3178-4f5e-97f1-0176e0f8fc29 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.611387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.611387Z digest=sha256:b76af0ff821f3d17f435e36e34437d3a6a17d89459dbaef2bc7f3dfd4284754f

Observation a6de5c47-ed61-469e-a188-460913426f03 · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.509228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.509228Z digest=sha256:d4d7684ff8e7b7c571cee0399879a2abcb58e2319de1208b7ad4fcf3ba8dda1f

Observation b7ec7ad2-4275-476e-8f7b-7e0c14e056df · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.493672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.493672Z digest=sha256:8304c3a72577412d19775b3f1e8da90c9a13b4c041ff22c5f3c6d03dea5b614a

Observation 09568c4b-06bc-4695-b267-50ef40445613 · outbound

This paper cites an unresolved cited work.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.499866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.499866Z digest=sha256:c7375f4c6e1b3cb5a29b7b579d7c7f6157a374e21a5b2e60ee2ba320b08cda8e

Observation 61784b62-6e4f-4ba1-904f-038bf903a1b3 · outbound

This paper cites Let's Verify Step by Step.

TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Let's Verify Step by Step

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T03:10:51.546216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:10:51.546216Z digest=sha256:97e4040ed48fcd699dc2ae7fb0a25152059d8da5dbf32b4d305eb5ac253b1ea0

Pith citing papers

Observation 344e4dc4-f6fc-4f6f-a7eb-7302855a467e · inbound

TCPO: Turn-Level Credit Policy Optimization cites this paper.

TCPO: Turn-Level Credit Policy Optimization TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:23:14.809554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:23:14.662916Z digest=sha256:149457a932e961255352b089420e8849c4a3d7965864a5a1c182dd41f88ebee2