Pith. sign in

Paper Citation Record · LEDGER

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2607.22724.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22724 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:55:00.149522Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 735f858a-003b-4482-bfb6-99bf523fe866 · outbound

This paper cites GPT-4 Technical Report.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:56.943572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:56.943572Z digest=sha256:4c929cdfacb70cbd18c958488887adbfb5c9ce4aab9f2f41703c56b2125c184d

Observation 153f6762-c155-4940-a48f-5abfbc6c7910 · outbound

This paper cites Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.024936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.024936Z digest=sha256:f255bd9037c3e79fcc605c2dbb6d5bcd42af51ede2ee76069b46c1bd2d24e292

Observation 76b916fa-ceae-4a1f-8c33-6bc7d40bfe1e · outbound

This paper cites XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.104603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.104603Z digest=sha256:4540291aa67b41f7c2e541e3d3979415e47ba5aaef93db4eec54eb6615ff4d96

Observation d1c681c6-96d6-4772-8643-d084061965fb · outbound

This paper cites Progra: Progress-aware reinforcement learning for multi-turn function calling.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Progra: Progress-aware reinforcement learning for multi-turn function calling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.183813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.183813Z digest=sha256:d8b35d1b6ebe753a1ffa4b5600c686bd59b95b589719ad79e623fd068edce828

Observation e9d9387e-8aa1-4762-8e92-2eabf8b75da1 · outbound

This paper cites Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.260816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.260816Z digest=sha256:4147ca6bf75f652d754f4a6ff6575296bae682112c24e1131927b7708b0451d4

Observation f39d3ef5-e163-4ba1-8e85-655eb5f7fe63 · outbound

This paper cites Proximity-based multi-turn optimization: Practical credit assignment for llm agent training.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Proximity-based multi-turn optimization: Practical credit assignment for llm agent training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.347969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.347969Z digest=sha256:1cdb714dfd4dcdc3b5b9eefd9374ef31e9bfab6ffac5cb1fe3e5ad142bc80ece

Observation 8b5ce984-ba5e-45cd-9070-d08ec0e81ae5 · outbound

This paper cites Towards efficient online tuning of VLM agents via counterfactual soft reinforcement learning.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Towards efficient online tuning of VLM agents via counterfactual soft reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.387214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.387214Z digest=sha256:e17be3bd2544aa543a7a5639cdda4c828d637cc01555d6afb9319d09012ef573

Observation 2aecaa62-a666-492c-9d25-c4c1980a138c · outbound

This paper cites Group-in-group policy optimization for llm agent training.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Group-in-group policy optimization for llm agent training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.408564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.408564Z digest=sha256:ca606e0f19feefb695ccdda1917cdcc08306904694d7ba8e6464f6eb3eaa0b3f

Observation 29d00fd5-e39d-4257-bf65-eec4519cf6a8 · outbound

This paper cites Multimodal web navigation with instruction-finetuned foundation models.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Multimodal web navigation with instruction-finetuned foundation models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.449494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.449494Z digest=sha256:fb1db0142da0f7f3e8f65adc16d994e04e972f4383a1c9b6490957b47c72da62

Observation 7b74b658-d362-43e5-8bd8-2e99cd0f9c94 · outbound

This paper cites Navigating the digital world as humans do: Universal visual grounding for GUI agents.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Navigating the digital world as humans do: Universal visual grounding for GUI agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.482108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.482108Z digest=sha256:7ad6f637e94261dbfd5cf3d07bc63ff3b318f4aa4dac981f70ac5303d834dfcf

Observation e70a9ec2-8237-4ecc-b7f4-4c2831afcfe5 · outbound

This paper cites Hierarchy-of-groups policy optimization for long-horizon agentic tasks.arXiv preprint arXiv:2602.22817, 2026.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Hierarchy-of-groups policy optimization for long-horizon agentic tasks.arXiv preprint arXiv:2602.22817, 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.550878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.550878Z digest=sha256:5286ce4f64f1cf3c26390f8a655f76dd5ad8e29e31cc0f381e08737c2a90df26

Observation 3066127f-4654-4345-9d09-bfd39d4d3d11 · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free! InICLR 2019 Workshop, 2019.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Buy 4 reinforce samples, get a baseline for free! InICLR 2019 Workshop, 2019

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.627348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.627348Z digest=sha256:a0d6e97b48194b7ed397d07fae8cde10a2049382aa287a4f404d796dc89c35a6

Observation 9bec2e20-02f0-4da2-b05f-5ebdf58e5339 · outbound

This paper cites No prompt left behind: Exploiting zero-variance prompts in llm reinforcement learning via entropy-guided advantage shaping.arXiv preprint arXiv:2509.21880, 2025.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks No prompt left behind: Exploiting zero-variance prompts in llm reinforcement learning via entropy-guided advantage shaping.arXiv preprint arXiv:2509.21880, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.697039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.697039Z digest=sha256:95c2e5e6bf1978ffd74433cfce2bbab5f7672cbfaf174c8ee39bae151e53e679

Observation ba852e80-66cd-4988-bddb-7dbc2b450f2d · outbound

This paper cites Salt: Step-level advantage assignment for long-horizon agents via trajectory graph.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Salt: Step-level advantage assignment for long-horizon agents via trajectory graph

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.763485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.763485Z digest=sha256:91b0ff9f2d22811fa1171bf43c1b192febbb111e1e0c5df23f681f016f146b5b

Observation 1bb46914-ba17-471c-8a21-b66cb2d554d0 · outbound

This paper cites Embodied agent interface: Benchmarking LLMs for embodied decision making.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Embodied agent interface: Benchmarking LLMs for embodied decision making

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:57.849466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:57.849466Z digest=sha256:170b5b10bde38e7e74848cc8c7507665211d81ed85a1c4ab1f203d8906b1ccfc

Observation b694f3a2-2a2b-40ea-ae60-66a4cd8aaa80 · outbound

This paper cites Agentic reinforcement learning with implicit step rewards.arXiv preprint arXiv:2509.19199, 2025.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Agentic reinforcement learning with implicit step rewards.arXiv preprint arXiv:2509.19199, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.026598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.026598Z digest=sha256:1a69d3977e5e5e261326b34dbb5572d26e2d34b1f3e7c6c197e07a40a9e673ad

Observation 6f170e07-8301-4bc9-a2e3-116956eb6366 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Understanding R1-Zero-Like Training: A Critical Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.149883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.149883Z digest=sha256:43994c4cc1028710606aace108e41b65831dc5a085ca81765bd6d9e0c69b48f0

Observation 9ef6eb7f-41df-451c-bf4e-00675be94ffe · outbound

This paper cites TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.217966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.217966Z digest=sha256:219ce2864e9cfc7026472812d106e75185b51c900350b72ee8a53c6aa848e9a0

Observation c0ed8ad7-bd09-4dbd-9a3b-bcb972e39b69 · outbound

This paper cites Ngrpo: Negative-enhanced group relative policy optimization.arXiv preprint arXiv:2509.18851, 2025.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Ngrpo: Negative-enhanced group relative policy optimization.arXiv preprint arXiv:2509.18851, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.321856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.321856Z digest=sha256:df75158807506b0eefd9b5a727501b42d9c178f1da04fffe695fed5d1eb8cdb1

Observation ab76c4dc-079a-47a8-ae20-a36978d09223 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ToolRL: Reward is All Tool Learning Needs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.370351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.370351Z digest=sha256:3d852a1598e6d419669f2759995f9e18ffdc71d7a42e60e9396411d4cf1d5c45

Observation e39b4dfb-3766-4988-bd8f-80d31b93802a · outbound

This paper cites RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.489272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.489272Z digest=sha256:f39d282e9cac9d0d0876a2cb7d4ea9860da954790b3e112ed0fc99bf76530ca8

Observation c6c5daa6-733f-4db9-b1b4-450f2e84d886 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551, 2023.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551, 2023

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.558728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.558728Z digest=sha256:5bf745d6388167aa80aa289841d0c66d467672f59c62a4130382c966f188b423

Observation fb84fd38-1f30-4ad0-b912-50e749c4a27f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.628083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.628083Z digest=sha256:7f5be05f2a79bc80c7b2dc87095a28fc08a2c6394630046374ab32c788b0bbcd

Observation fd73ad1f-60b7-4a24-9f29-5a0e76208184 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.733456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.733456Z digest=sha256:ba7418b495ed4256bd84dcd67146e3e8c54a3fc7977c544adbd0c449a59c657b

Observation a1464195-bee3-40ab-a24e-ce74a2e37d8e · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36, 2024.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.902231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.902231Z digest=sha256:5696082227b0c94bb0b2e91c6c81f0581a17c2cba82785ed180f3ea88dbe3309

Observation d56febe0-c7b2-4e16-a94f-834d91a6effd · outbound

This paper cites ALFWorld: Aligning text and embodied environments for interactive learning.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ALFWorld: Aligning text and embodied environments for interactive learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:58.960233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:58.960233Z digest=sha256:11a205b04d2056f045279cefc39731b14f358a06cee0465254dc7d790916533d

Observation 1add1d35-3663-46e3-8b31-dd1718c77724 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Gemini: A Family of Highly Capable Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.062582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.062582Z digest=sha256:093b5273b403e382bcd7c9ed34348a9b554fc674a74b10161f9311488ac87d87

Observation 12cecd86-58dc-47f0-a82d-cf3ee377e031 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.123594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.123594Z digest=sha256:b27a57aca4891d02a60a320be640cb9511125f6ce8b74af211adfd66ab928bfb

Observation 77f84b04-9d40-491f-9927-52328a607635 · outbound

This paper cites Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.187480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.187480Z digest=sha256:3eca722dd3c9ba4151a3f03b5248e13598aa1d289daab2c27d3b25dc44b7ad30

Observation 0c3f1976-d2f5-4aa9-8e5e-ea3e347b890e · outbound

This paper cites Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.arXiv preprint arXiv:2510.14967, 2025.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.arXiv preprint arXiv:2510.14967, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.299303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.299303Z digest=sha256:48ef35739cc605b16a06fac346983c4309eeaf0ae5519fe438c371920a8dcb45

Observation 69741546-e41a-4506-a1df-59b33b37d407 · outbound

This paper cites SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.427547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.427547Z digest=sha256:94ed4ab21f4c3c1c63a5712c90e58ce53d6afd04cadfc7ad8348dfc7c26013b2

Observation 8aa4bd78-cbee-4d74-a401-bcde75c7ef8e · outbound

This paper cites Mobile-Agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Mobile-Agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.487844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.487844Z digest=sha256:6de3f9d1210a8dc2531a6c6b5eae9ac1cdfb1e93e38dc3db461035d804199f42

Observation e6f8b6a3-ab00-49c3-9543-10dd64968821 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.568773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.568773Z digest=sha256:3315994e2750ef52330ed64cf34e54acb6e607b5fd6563bd82e3ac50ade99e39

Observation ede04870-efde-46f5-84f5-1c473f846b00 · outbound

This paper cites Agentprm: Process reward models for llm agents via step-wise promise and progress.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Agentprm: Process reward models for llm agents via step-wise promise and progress

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.658025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.658025Z digest=sha256:199f8de7e2d887e6bdd2b0b811e930a63148b32b733b25d1e53eb38227b524de

Observation 16f0f686-765a-471b-bf93-1715f5ea85ce · outbound

This paper cites Watch every step! llm agent learning via iterative step-level process refinement.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Watch every step! llm agent learning via iterative step-level process refinement

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.719592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.719592Z digest=sha256:4ed80e08b03d92216a9336ddaf68a2a42238b8dd0759b8ee10455ed5cba8ac00

Observation 051548c4-99fd-47c1-af84-d6ed190263df · outbound

This paper cites Qwen3 Technical Report.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Qwen3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.787781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.787781Z digest=sha256:7eb64952c88bd86e164da75c53c47cd52a79dc0174e09c58cd8c4ab07ef205b7

Observation 388bfc90-99b3-45cb-9700-2c87a312a1fe · outbound

This paper cites WebShop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks WebShop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.916475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.916475Z digest=sha256:71dcb048dc78e8425d02dd46b5bfceb43de345d9c702e7398e17746a15f65b1d

Observation 0c24956e-6f6d-4016-b592-9e6a32c08b93 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ReAct: Synergizing Reasoning and Acting in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:54:59.962766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:54:59.962766Z digest=sha256:0f0174e4c13947a642304ddeb7f47c8b47cbf805d2583ffee77efaaa25c8e024

Observation 0ce797b9-649e-4872-ab75-2f4d50c414a2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:55:00.025273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:55:00.025273Z digest=sha256:78e797f3d6c697543e792a4cfa961a61107e702c04a2e7a778fceebee061352c

Observation f7c6c403-0bc4-426d-9a67-637a11b4e003 · outbound

This paper cites Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:55:00.080830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:55:00.080830Z digest=sha256:1300cbac0e09726fee298e9fe1ff5bc51921e141bbd0d0d57a229cde98106eac

Observation 89303a57-3d4e-45c8-8f9d-684a626ada19 · outbound

This paper cites Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning.arXiv preprint arXiv:2510.19807, 2025.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning.arXiv preprint arXiv:2510.19807, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:55:00.117534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:55:00.117534Z digest=sha256:1d934308f9549fb27ebd5f75afaadeab4894c053e34bc725f022a0af99b1749a

Observation 7081ea75-5803-4eca-ad0f-75a2a6116586 · outbound

This paper cites put a cool tomato on the countertop.

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks put a cool tomato on the countertop

Reference 43

Resolution
malformed identifier
no resolver link, observed 2026-08-01T11:55:00.149522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:55:00.149522Z digest=sha256:e0b7b7ab2dcc076d328477eacfb3e346f9e57c0e71f37ab9e5d9cdd3d11dac1d

Pith citing papers

No inbound Pith citation observations are available.