Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:55:00.149522Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2607.22724.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:55:00.149522Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 735f858a-003b-4482-bfb6-99bf523fe866 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153f6762-c155-4940-a48f-5abfbc6c7910 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b916fa-ceae-4a1f-8c33-6bc7d40bfe1e · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c681c6-96d6-4772-8643-d084061965fb · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Progra: Progress-aware reinforcement learning for multi-turn function calling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9d9387e-8aa1-4762-8e92-2eabf8b75da1 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39d3ef5-e163-4ba1-8e85-655eb5f7fe63 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Proximity-based multi-turn optimization: Practical credit assignment for llm agent training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b5ce984-ba5e-45cd-9070-d08ec0e81ae5 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Towards efficient online tuning of VLM agents via counterfactual soft reinforcement learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aecaa62-a666-492c-9d25-c4c1980a138c · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Group-in-group policy optimization for llm agent training
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d00fd5-e39d-4257-bf65-eec4519cf6a8 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Multimodal web navigation with instruction-finetuned foundation models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b74b658-d362-43e5-8bd8-2e99cd0f9c94 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Navigating the digital world as humans do: Universal visual grounding for GUI agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e70a9ec2-8237-4ecc-b7f4-4c2831afcfe5 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Hierarchy-of-groups policy optimization for long-horizon agentic tasks.arXiv preprint arXiv:2602.22817, 2026
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3066127f-4654-4345-9d09-bfd39d4d3d11 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Buy 4 reinforce samples, get a baseline for free! InICLR 2019 Workshop, 2019
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bec2e20-02f0-4da2-b05f-5ebdf58e5339 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks No prompt left behind: Exploiting zero-variance prompts in llm reinforcement learning via entropy-guided advantage shaping.arXiv preprint arXiv:2509.21880, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba852e80-66cd-4988-bddb-7dbc2b450f2d · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Salt: Step-level advantage assignment for long-horizon agents via trajectory graph
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb46914-ba17-471c-8a21-b66cb2d554d0 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Embodied agent interface: Benchmarking LLMs for embodied decision making
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b694f3a2-2a2b-40ea-ae60-66a4cd8aaa80 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Agentic reinforcement learning with implicit step rewards.arXiv preprint arXiv:2509.19199, 2025
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f170e07-8301-4bc9-a2e3-116956eb6366 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Understanding R1-Zero-Like Training: A Critical Perspective
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ef6eb7f-41df-451c-bf4e-00675be94ffe · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ed8ad7-bd09-4dbd-9a3b-bcb972e39b69 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Ngrpo: Negative-enhanced group relative policy optimization.arXiv preprint arXiv:2509.18851, 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab76c4dc-079a-47a8-ae20-a36978d09223 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ToolRL: Reward is All Tool Learning Needs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e39b4dfb-3766-4988-bd8f-80d31b93802a · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c5daa6-733f-4db9-b1b4-450f2e84d886 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Toolformer: Language models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551, 2023
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb84fd38-1f30-4ad0-b912-50e749c4a27f · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Proximal Policy Optimization Algorithms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd73ad1f-60b7-4a24-9f29-5a0e76208184 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1464195-bee3-40ab-a24e-ce74a2e37d8e · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56febe0-c7b2-4e16-a94f-834d91a6effd · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ALFWorld: Aligning text and embodied environments for interactive learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1add1d35-3663-46e3-8b31-dd1718c77724 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Gemini: A Family of Highly Capable Multimodal Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12cecd86-58dc-47f0-a82d-cf3ee377e031 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f84b04-9d40-491f-9927-52328a607635 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3f1976-d2f5-4aa9-8e5e-ea3e347b890e · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Information gain-based policy optimization: A simple and effective approach for multi-turn llm agents.arXiv preprint arXiv:2510.14967, 2025
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69741546-e41a-4506-a1df-59b33b37d407 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa4bd78-cbee-4d74-a401-bcde75c7ef8e · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Mobile-Agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6f8b6a3-ab00-49c3-9543-10dd64968821 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede04870-efde-46f5-84f5-1c473f846b00 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Agentprm: Process reward models for llm agents via step-wise promise and progress
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f0f686-765a-471b-bf93-1715f5ea85ce · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Watch every step! llm agent learning via iterative step-level process refinement
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051548c4-99fd-47c1-af84-d6ed190263df · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Qwen3 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 388bfc90-99b3-45cb-9700-2c87a312a1fe · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks WebShop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c24956e-6f6d-4016-b592-9e6a32c08b93 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks ReAct: Synergizing Reasoning and Acting in Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ce797b9-649e-4872-ab75-2f4d50c414a2 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c6c403-0bc4-426d-9a67-637a11b4e003 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89303a57-3d4e-45c8-8f9d-684a626ada19 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks Scaf-grpo: Scaffolded group relative policy optimization for enhancing llm reasoning.arXiv preprint arXiv:2510.19807, 2025
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7081ea75-5803-4eca-ad0f-75a2a6116586 · outbound
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks put a cool tomato on the countertop
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.