Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T10:50:54.419477Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 0 inbound Pith citation observations for arXiv:2607.04963.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T10:50:54.419477Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 110 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 957f32bc-7e79-45ae-a7e2-d092602e8e1b · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2018 , publisher=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320b38b2-5c3d-4b08-8c12-b627df5136dc · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International conference on machine learning , pages=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05aee880-0362-4168-810e-49dea053f2fc · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Proximal Policy Optimization Algorithms
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1d901c-42d4-4b8b-ab74-aea5bebebc89 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training NeurIPS 2024 Workshop on Open-World Agents , year=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8064da12-92e1-404e-899d-c45d858f70e7 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f465362-e5d1-41b0-a1cf-338ac3980999 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b247f7c-3cfb-4205-9c25-e1ad118b43f2 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fccddd7b-c971-4912-81e3-57b1c10d13f0 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Dawn of
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59256957-4dc9-436f-8070-ee1b3757428e · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b263e96-8033-49f1-9f23-a1360f573ed5 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Gemini: A Family of Highly Capable Multimodal Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 047b9b9b-1387-4236-bdaf-474bf9d19fb9 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbee810a-7503-4044-b679-ae74aa84f9ac · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Transactions on Machine Learning Research , issn=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e019641e-235c-48d6-a6a1-98693a199282 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Twelfth International Conference on Learning Representations , year=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6818e8-cfa8-499d-b859-5054d651c4db · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Agent as Cerebrum, Controller as Cerebellum: Implementing an Embodied LMM-based Agent on Drones
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0fab0fd-ee0a-41e1-b57f-23b2304be306 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2023 , organization=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ea9c60d-52b9-4e98-80a8-d78fdecb332e · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training AppAgent: Multimodal Agents as Smartphone Users
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6790ad-0d11-4eb6-a205-0455cfbf5a0b · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3962adc-f4b6-4072-a5c7-0d08872bd172 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04fb3aca-3512-4525-8aee-e4fd3ed65f68 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14075823-08c4-46c8-be05-55673d61ee34 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea01bc3-da9c-4af1-bf71-9121e51c6aa0 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8ec340-b559-499c-98ba-7f17b61249c2 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f2f142-126c-4729-8d54-73ccda10b892 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66100cb3-994e-4bb5-a046-59ec672fded9 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f413e00-dcd8-4e95-a9dc-e4ef6bac9f52 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e2c3f41-3444-41c2-bcf5-724d9f21d4b6 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8976df23-1688-4147-bd56-ae470e7824a3 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International conference on machine learning , pages=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5528c1-323a-48f8-917f-0b08a64b44a4 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training ScreenAgent: A Vision Language Model-driven Computer Control Agent
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1045e8a1-e62a-4f88-b725-de8b7f2680ac · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Findings of the Association for Computational Linguistics ACL 2024 , pages=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd47b29e-9b45-40de-b877-0b02d1cd61da · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Reinforcing
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e310de8e-4725-4e8c-b7d1-5d3f49cf1eb5 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International Conference on Machine Learning , pages=
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56cecc91-1805-4e49-912f-9797b1ea6101 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training FireAct: Toward Language Agent Fine-tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acad7ed2-fc0c-4ac0-92b7-0d51960b3ed8 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Twelfth International Conference on Learning Representations , year=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bce8b4-78a5-4194-9de2-63698f423084 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16bd459-e178-41e8-8696-63e1d0919bbe · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8d8593-4603-4bd9-9b02-0ff3aa328784 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d16b72f0-e5b8-4d73-a4ee-57fbeee56165 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ba1777-7e6d-4cd1-9f46-426e23a72028 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Huang and Mustafa Safdari and Yutaka Matsuo and Douglas Eck and Aleksandra Faust , title=
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e558e9-5543-47fc-854b-a25724db6c26 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd33716-b5a4-48b4-bfbd-2a31c4b6ed40 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 123aa6e2-783e-4364-a3d5-597766887343 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bf2cb6-2040-41c0-b4e6-2463bd6c254f · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International conference on machine learning , pages=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9caaed7-851b-432f-a850-226988c2c7fd · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Large-Scale Study of Curiosity-Driven Learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb33799-c497-4663-980e-72d7e58c9d08 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2010 , publisher=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20e9de4-e927-41ab-9133-37dbc93b178a · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8af4c45-6483-4640-b418-f3b2c316ef6a · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d354d5-aa7b-41b6-83d0-cb24c42e6039 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2025 , url=
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35d7284-3b5c-4797-b1bc-9b2144c703da · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training , title =
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc8a6eaa-0aec-49cf-8452-22b4f1c2fea2 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e36be9a-30ee-4886-a157-05b101a813c7 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c382f814-2583-4ff0-af54-06778051013f · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2024 , url=
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb605e8d-26e0-4f5c-acda-177e31c5f407 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qwen2.5 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072b26d5-edc6-49f7-9e1b-ec34d12395aa · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Embodied Agent Interface: Benchmarking
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45bc7e29-f461-4d42-827f-26f8ad5f4261 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64faf99c-c04b-4c46-b25a-5e2cf8b57b42 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c042b441-17ec-4e08-8a90-594347e3badf · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training ICLR 2019 Workshop , year=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22545a3c-3b43-44ad-b52c-6d49c5e845ee · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Understanding R1-Zero-Like Training: A Critical Perspective
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6bc13d3-873c-42c5-985b-7c5b745d673e · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 745f386e-72a4-45ce-b849-a3bb3b9f73ea · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b644b619-26d8-48ea-91df-0d76deedb835 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b369be3-e4ab-4a7a-a8c3-8c31c7118a7a · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4728e245-64c2-4507-9861-c9d9a79d2649 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Fine-Tuning Language Models from Human Preferences
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78a91dff-a3f5-4096-8ef0-f7cea535354a · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4eb951-4ac6-4d41-b74f-82d96b3dc9bb · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22a6063-fe30-45f5-bc49-c388e89bce20 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2024 , organization=
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e344cb6-0fa2-48e5-9fac-b1695c121d57 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0600542-6d49-4567-809e-94e6b476ee83 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training ToolRL: Reward is All Tool Learning Needs
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f375978-c8a9-4d8a-9d97-90c889677d7b · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Kimi k1.5: Scaling Reinforcement Learning with
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a386376-582d-4a48-a9a9-1ae52656a5ee · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe314ac-1e82-4bc2-bda0-7bc89470683d · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Nature , volume=
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e562ecc5-f39a-4722-8bc8-e1cbaa3c576a · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Keep CALM and Explore: Language Models for Action Generation in Text-based Games
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80472bc-65df-429a-ad86-947049c3a6df · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Nature , volume=
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6bcefbf-1140-40fa-8182-2ebb5e28b9a4 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc881cec-8333-41b4-b89e-e740b974845d · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84942a6-1d88-4bdc-bd49-345545803f89 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95936b3f-a369-47d9-9bfa-caf7c672acf8 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qwen2.5-
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4228559b-a206-4066-8ed7-e518374af6a9 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be50895e-a126-4cad-b461-592417e6d5d6 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd621eac-c30c-4d50-9625-bd7ec5475a53 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Navigating the Digital World as Humans Do: Universal Visual Grounding for
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 657707dd-ad74-4f1c-8f47-f612d05f8ca1 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 016f0631-24b2-4e2f-b739-2550c3faa646 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738a5063-4efc-43b1-b186-08c188afc0df · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a00c42f-93ed-45b6-b4d7-d44b3414654d · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d853f87-8e48-45c0-8b45-32ab1b217429 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Transactions of the Association for Computational Linguistics , volume=
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c15aee-048a-4732-9285-1d60034f3533 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d989d7-942b-49b6-92db-389cc43b5e57 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2022 , publisher=
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08536e4e-425d-4d7a-b1b7-7dbcb0784b68 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Measuring and Narrowing the Compositionality Gap in Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4603ac8-d8d4-42c3-a630-992b2619fb6e · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede3a1b8-1754-4d9b-aea1-d9301445b575 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5f57ff-b147-4ba9-a093-d6f2e22f6878 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2500176-f6f2-4021-ad3c-4b837c51ddf4 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb9a0e8-37d0-435d-adb1-9727a6ce19ae · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Text Embeddings by Weakly-Supervised Contrastive Pre-training
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc178778-6115-419e-bc51-0a4c66471888 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Towards Efficient Online Tuning of
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32017a5b-3aed-4b52-83b2-f9939883db03 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qwen3 Technical Report
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e79b2e-d358-4394-ac5b-34614017d85e · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training , author=
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4ac855-29fe-4c3f-bd0b-6a6e657fdfe0 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39e0090-b9af-40e9-bb51-fdc755be7ea1 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14758683-266e-4648-9a30-c9efbd253752 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training arXiv preprint arXiv:2509.22576 , year=
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b155a972-255b-4d7d-9ff8-8174c9ea0e72 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f12623df-14a8-4c8c-baf3-cfd014d9d665 · outbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.