Pith. sign in

Paper Citation Record · LEDGER

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

As of 20 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 0 inbound Pith citation observations for arXiv:2607.04963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.04963 v1

Coverage vector

measured 100 of 110 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T10:50:54.419477Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 110 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 957f32bc-7e79-45ae-a7e2-d092602e8e1b · outbound

This paper cites 2018 , publisher=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2018 , publisher=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:eb338dcc4515e755d02db75dcdef85767546027d0c0675e92e7676bca2ba72f4

Observation 320b38b2-5c3d-4b08-8c12-b627df5136dc · outbound

This paper cites International conference on machine learning , pages=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International conference on machine learning , pages=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:67f960c21ed2b79a00e9a95749847fd6f4b3abd97847c4aee9d7a76d4648d0fe

Observation 05aee880-0362-4168-810e-49dea053f2fc · outbound

This paper cites Proximal Policy Optimization Algorithms.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Proximal Policy Optimization Algorithms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:271ed44a0f83f0d49474205c78342715249b56c63314dc12d4792e4dfbfaa4c2

Observation 4e1d901c-42d4-4b8b-ab74-aea5bebebc89 · outbound

This paper cites NeurIPS 2024 Workshop on Open-World Agents , year=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training NeurIPS 2024 Workshop on Open-World Agents , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:9321aa820268921c2e5be2cf5dd3bad4417ff3f72b9058004ff18c612c1d1916

Observation 8064da12-92e1-404e-899d-c45d858f70e7 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:7ef93144cc7ddf974824172f024560e328b0799f58bdeb74eeac716a9fe91e40

Observation 2f465362-e5d1-41b0-a1cf-338ac3980999 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a30dd74f62ed407c905700d1c155473e6b45ffa00a2f348504f2ea95bcc80fa5

Observation 1b247f7c-3cfb-4205-9c25-e1ad118b43f2 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:7bce2624bba2ae2ec37cf826a5f4fdf5fd52963bea6eb3a39a43f7d85195d009

Observation fccddd7b-c971-4912-81e3-57b1c10d13f0 · outbound

This paper cites The Dawn of.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Dawn of

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:ed6bbb635b28988e3b347abc8a3363b185e63c90eb764fd3e36276d66aa801f4

Observation 59256957-4dc9-436f-8070-ee1b3757428e · outbound

This paper cites DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:166627806ed750e4a419fe99ed3c25c500572aaabf565694ea6fd500a5a681b7

Observation 9b263e96-8033-49f1-9f23-a1360f573ed5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:5349355031867e58b86bcfa45b43def9f04fec9ee99dcce5c97a6db6a3bab3f3

Observation 047b9b9b-1387-4236-bdaf-474bf9d19fb9 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:cbec4f98c5cf8d54bbec8441d1f40410e3fec41a788427183d15a4e6adbe571c

Observation fbee810a-7503-4044-b679-ae74aa84f9ac · outbound

This paper cites Transactions on Machine Learning Research , issn=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Transactions on Machine Learning Research , issn=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:cc852a3cbc2c5742751e6291bdaeb185666326505948afda968643eab09180b3

Observation e019641e-235c-48d6-a6a1-98693a199282 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Twelfth International Conference on Learning Representations , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:caf011f6942d51c8175d3b678c40f7a6c45568b48e19a2cc07e1ea3d7647df67

Observation 1f6818e8-cfa8-499d-b859-5054d651c4db · outbound

This paper cites Agent as Cerebrum, Controller as Cerebellum: Implementing an Embodied LMM-based Agent on Drones.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Agent as Cerebrum, Controller as Cerebellum: Implementing an Embodied LMM-based Agent on Drones

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:0a09512ff737d94e8a1132b412c7749fc75768984d9e58ff253b95fd19d2fc87

Observation a0fab0fd-ee0a-41e1-b57f-23b2304be306 · outbound

This paper cites 2023 , organization=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2023 , organization=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:6fed7dca5d4cb8675d10a3464ce30afc6af125556c82d2a0e9ad81a44ad820ba

Observation 6ea9c60d-52b9-4e98-80a8-d78fdecb332e · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training AppAgent: Multimodal Agents as Smartphone Users

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:6611e6ddceda97e4ab88a6421f3e1bb96f5626895076f90ba10b02cf0adec5ab

Observation 0f6790ad-0d11-4eb6-a205-0455cfbf5a0b · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:d84b7e3388af3f84954132853756dabd00a5c0381298876332b8439cd4465341

Observation f3962adc-f4b6-4072-a5c7-0d08872bd172 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:b92f6c597aaa48f1a8576970477f0f025f760ee65d5a053d754bc81f7585ab0e

Observation 04fb3aca-3512-4525-8aee-e4fd3ed65f68 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:f73727b618eee4ce6a2fab04df664e431016cca35cdb4b4bb4bf2ee56ee6d846

Observation 14075823-08c4-46c8-be05-55673d61ee34 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:e134c5ba9e5f50edfb27d77d16358381ef983d60aab8b02e3c5bcccac96aea55

Observation 4ea01bc3-da9c-4af1-bf71-9121e51c6aa0 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:ca67057623c6fbb0c793895da0a93cb8e3dd31b78f5a7d1c1a18b1abf2580694

Observation 5c8ec340-b559-499c-98ba-7f17b61249c2 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:e13bfb8f1f818f303bf3414907fa66c69db053254f17add826cff3daa06fc37e

Observation 32f2f142-126c-4729-8d54-73ccda10b892 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:50f350f03d9ef38681254e9065eee46cbcad7ef00092cbf685ca854d7566eb12

Observation 66100cb3-994e-4bb5-a046-59ec672fded9 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 24

Resolution
parse uncertain
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:417c0dfbeffda996e285c44ba5678e909e4c4ced35e179ef858d6c67b84bc74b

Observation 2f413e00-dcd8-4e95-a9dc-e4ef6bac9f52 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:ad66f1e785b6bd9429091255664553a6246010449c69de0fc02f292c681c2ab1

Observation 9e2c3f41-3444-41c2-bcf5-724d9f21d4b6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:8a9b56cb44533477c97515a1ae8b47a9fa809da1889bf254b8923496bd15a629

Observation 8976df23-1688-4147-bd56-ae470e7824a3 · outbound

This paper cites International conference on machine learning , pages=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International conference on machine learning , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:d533e20e6fcc561d20a4a078db42e7106635b413da2bcd56fbc79687f60f8b0f

Observation 9d5528c1-323a-48f8-917f-0b08a64b44a4 · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:9efe1f940f926526e2d62dde100095d2ccf9832cd00e8f3004b6b39428f5c368

Observation 1045e8a1-e62a-4f88-b725-de8b7f2680ac · outbound

This paper cites Findings of the Association for Computational Linguistics ACL 2024 , pages=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Findings of the Association for Computational Linguistics ACL 2024 , pages=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:df1de0c0fe4edbaaecff17a808e3a0b969e5ff788aaa143df4c225afea0861e7

Observation fd47b29e-9b45-40de-b877-0b02d1cd61da · outbound

This paper cites Reinforcing.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Reinforcing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:ad9ed9ef0e4578e0542052ebf1d9b1bc76a11f5949581cf10235eb3b0cbedcbc

Observation e310de8e-4725-4e8c-b7d1-5d3f49cf1eb5 · outbound

This paper cites International Conference on Machine Learning , pages=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International Conference on Machine Learning , pages=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:4aa4f81672680e87c11cba41a9d8b08f5421f3b20e086cb873eeb50ba4f99170

Observation 56cecc91-1805-4e49-912f-9797b1ea6101 · outbound

This paper cites FireAct: Toward Language Agent Fine-tuning.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training FireAct: Toward Language Agent Fine-tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:1a9ab887c7c015c73a670300cabbfc56297c38567608a90f07a7e51ec34e1711

Observation acad7ed2-fc0c-4ac0-92b7-0d51960b3ed8 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Twelfth International Conference on Learning Representations , year=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:51efaa719da56e3284c7ddd31cb538875d4092ee9f14d02806bcd6d2981678eb

Observation 79bce8b4-78a5-4194-9de2-63698f423084 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:587e4a4952d340f607d35d1f3e62d6ae06dabbea5a3d013f98249d714a22c621

Observation c16bd459-e178-41e8-8696-63e1d0919bbe · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:7e914a309b80896ba23e9355c9d44d9ec442e963b221caeda4d1a873b3e7ecd6

Observation 4e8d8593-4603-4bd9-9b02-0ff3aa328784 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:5ac4b1bead976a49c802615d9b2b359224981aa7d1e3f2ae4db567f8720d230c

Observation d16b72f0-e5b8-4d73-a4ee-57fbeee56165 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a9b262c1578c6c75f3b43d04e9d175d9be01c542b4d6795395f107560543eb79

Observation 50ba1777-7e6d-4cd1-9f46-426e23a72028 · outbound

This paper cites Huang and Mustafa Safdari and Yutaka Matsuo and Douglas Eck and Aleksandra Faust , title=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Huang and Mustafa Safdari and Yutaka Matsuo and Douglas Eck and Aleksandra Faust , title=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:0b2900ddab6b60558f24ebe6d14a36c96bb51f11c530ef7b7c0b6f5536fec6cb

Observation 77e558e9-5543-47fc-854b-a25724db6c26 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:c3fc7615f7555aee5e3657b19610ae31a7fca6398cb6edb0fe12a0447e7fc2b2

Observation ffd33716-b5a4-48b4-bfbd-2a31c4b6ed40 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:090773f942e82da353cb44d11ae7d1af7e187ff756747ffc46fef36ca812d17e

Observation 123aa6e2-783e-4364-a3d5-597766887343 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:c8ed1d120a4459df3c47896a943b57e463457978fed692ccc1239ac40b3dc780

Observation 79bf2cb6-2040-41c0-b4e6-2463bd6c254f · outbound

This paper cites International conference on machine learning , pages=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training International conference on machine learning , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:2618dcf845bbf9c2caf3fe2a1db644b898d86a75b0ae0b5c21d88d47052c8dac

Observation a9caaed7-851b-432f-a850-226988c2c7fd · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Large-Scale Study of Curiosity-Driven Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:2467564101dac059cbca6e77ee1e226a98cdbb7d66e2c4b76568f41d0d1fac61

Observation 2eb33799-c497-4663-980e-72d7e58c9d08 · outbound

This paper cites 2010 , publisher=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2010 , publisher=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a7463e28c2dcfc63051e4bd22249f7831c6d541eb85901c7509669a818391c2f

Observation c20e9de4-e927-41ab-9133-37dbc93b178a · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:4f5c828ad13ecbcb613ab2fcbf1c74b7b910634d1928ed6f8882fde1cf72ea6a

Observation c8af4c45-6483-4640-b418-f3b2c316ef6a · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a2d412289136cbae83474133740bac645b522cc514ce8fb530c29a3a492f37e9

Observation 56d354d5-aa7b-41b6-83d0-cb24c42e6039 · outbound

This paper cites 2025 , url=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2025 , url=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:ec58d1f9e65cb1671476ce063398757462fbc4b05cc2ea414d9032583efe8cf4

Observation c35d7284-3b5c-4797-b1bc-9b2144c703da · outbound

This paper cites , title =.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training , title =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:d70249173eaf21229b4fd5599c67ef63a38eeda1f3eca12df8edf8db485375a5

Observation fc8a6eaa-0aec-49cf-8452-22b4f1c2fea2 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:d54f318de59786f5c1e330f9aee8ea47d47f5357fc708a899d777bb56bfc122e

Observation 3e36be9a-30ee-4886-a157-05b101a813c7 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:28b5f3596a71ca82912d53ff88ee3640c4cc48cac6e536a0ee5899688e26247e

Observation c382f814-2583-4ff0-af54-06778051013f · outbound

This paper cites 2024 , url=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2024 , url=

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:54724f9c3c21705cd6b61a945ed2a58c21972393015eb150c3f4d17aa76f035c

Observation bb605e8d-26e0-4f5c-acda-177e31c5f407 · outbound

This paper cites Qwen2.5 Technical Report.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qwen2.5 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:5330a740d95fb9980a40b18f29528773dee367a95bc5e10b26428e018895570e

Observation 072b26d5-edc6-49f7-9e1b-ec34d12395aa · outbound

This paper cites Embodied Agent Interface: Benchmarking.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Embodied Agent Interface: Benchmarking

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:5864772868b4c626d3e99f482cc89d14c3d6c3f74400c8a0896ddc29b970b3e0

Observation 45bc7e29-f461-4d42-827f-26f8ad5f4261 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:4e4884835542e13b57930e1d14da7605edd644f4f019c3eeb0f173b7899ca5af

Observation 64faf99c-c04b-4c46-b25a-5e2cf8b57b42 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:c1606dff2c2b745b9dd68c3335f393a245a2549b9893a8181d71cbea2a6d4e41

Observation c042b441-17ec-4e08-8a90-594347e3badf · outbound

This paper cites ICLR 2019 Workshop , year=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training ICLR 2019 Workshop , year=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:f05022afdc76865d9deca77080f0f5402c0603c17f0a82729f647c523dc56029

Observation 22545a3c-3b43-44ad-b52c-6d49c5e845ee · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Understanding R1-Zero-Like Training: A Critical Perspective

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:0552852e8c7d783aa665971562eec79ed88d646460884338857aeaaafb8e338f

Observation d6bc13d3-873c-42c5-985b-7c5b745d673e · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:5f3f7a444097a7e2c6f1304fd7a73263531ab99c058bfa69b1fbd5539d61d49c

Observation 745f386e-72a4-45ce-b849-a3bb3b9f73ea · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:71c3fc07e8198ab9dfbef8b8e5ab73ac6f6f56cf244e053cb415c86bb98c7bfa

Observation b644b619-26d8-48ea-91df-0d76deedb835 · outbound

This paper cites Reinforcement Learning for Long-Horizon Interactive LLM Agents.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Reinforcement Learning for Long-Horizon Interactive LLM Agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:25dbbc850dfc7a0324e5779707e1f0bb98ac8651326470cf5fb10b21b0489096

Observation 6b369be3-e4ab-4a7a-a8c3-8c31c7118a7a · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a9c706798726ca8ae3801c713bf5b5ebeeb8f5bcd488df18d7a70357e208f5d0

Observation 4728e245-64c2-4507-9861-c9d9a79d2649 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Fine-Tuning Language Models from Human Preferences

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:be4f3d75762820bf3393b366ebf696175478178b0faa7e0f0073828ace9dc6f1

Observation 78a91dff-a3f5-4096-8ef0-f7cea535354a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Advances in Neural Information Processing Systems , volume=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:b9b362cb701607a23080ee1caeadb3bd2942a5530494075088d79497aa5b0ef6

Observation 0a4eb951-4ac6-4d41-b74f-82d96b3dc9bb · outbound

This paper cites Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:2e3022675cc5b9071a71312b6863f10c6c121a42a7882a41829be383c08d096c

Observation c22a6063-fe30-45f5-bc49-c388e89bce20 · outbound

This paper cites 2024 , organization=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2024 , organization=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:31602465fdfd3906bf20df10fb9081a92143b36275156284c4932f17cdd709d5

Observation 0e344cb6-0fa2-48e5-9fac-b1695c121d57 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:3e2872209b959982efee125efcc85d19dc3ea1f06402969d4ff0e12719b1c3b9

Observation f0600542-6d49-4567-809e-94e6b476ee83 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training ToolRL: Reward is All Tool Learning Needs

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:f44b5191960f579201a134bf30d0ff940c6998d38a9dc9c134c2d24842fc8bc2

Observation 5f375978-c8a9-4d8a-9d97-90c889677d7b · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Kimi k1.5: Scaling Reinforcement Learning with

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:85bbde81b7ccff11b8f0c1a21b94b36508d56025af515a7f4548f3ce13326308

Observation 3a386376-582d-4a48-a9a9-1ae52656a5ee · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:1405caca2244a954253c7eaa075ee58184e714187da19f84a48b1e71775e9284

Observation ebe314ac-1e82-4bc2-bda0-7bc89470683d · outbound

This paper cites Nature , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Nature , volume=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:15c56096700fbec7c6d4bddf515c49c18e9a4513276728ce16762a5db4f5e447

Observation e562ecc5-f39a-4722-8bc8-e1cbaa3c576a · outbound

This paper cites Keep CALM and Explore: Language Models for Action Generation in Text-based Games.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Keep CALM and Explore: Language Models for Action Generation in Text-based Games

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:bf03ac0d918ccff28944049e1dfb912db944dadff66e32dd454131b0aba04eb8

Observation e80472bc-65df-429a-ad86-947049c3a6df · outbound

This paper cites Nature , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Nature , volume=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:246afc70fb26271e76de526ba8555afd2c7a38844833ce92ff2688b1efda9374

Observation f6bcefbf-1140-40fa-8182-2ebb5e28b9a4 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:4df987459fe67a89b78d15a80e413c8dbde23289a7dd856008d2f0138a7fac51

Observation fc881cec-8333-41b4-b89e-e740b974845d · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:40d5774ed67a0946d1dc405025d926b00def2205565186d7e9e19d14c0c0cc84

Observation b84942a6-1d88-4bdc-bd49-345545803f89 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:289306eeb755e84891429c9567f8a6048d2db1c4ba51c70d5fb67f93f47a483d

Observation 95936b3f-a369-47d9-9bfa-caf7c672acf8 · outbound

This paper cites Qwen2.5-.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qwen2.5-

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:891e758c197ff01ca9d92ba9cb8818c6feef456b0b4a18b15902e9f40ed5d22e

Observation 4228559b-a206-4066-8ed7-e518374af6a9 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:adf501696229d01f221c531d0a5247b9ed69a84083a7891341f211dc6091a857

Observation be50895e-a126-4cad-b461-592417e6d5d6 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:f5028d378013236031cff9d25075d4ab7206f9ee59b3c1fd9917e440b8b732a5

Observation cd621eac-c30c-4d50-9625-bd7ec5475a53 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Navigating the Digital World as Humans Do: Universal Visual Grounding for

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:783ca178ca2e2285962cde8b24b7d1df9e5661db6e9c462c2545c38be9802efa

Observation 657707dd-ad74-4f1c-8f47-f612d05f8ca1 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:5668b9db2bc943f9cd00dd0dc567707af6ef7856af816f5fce16c9fb5ab1d25d

Observation 016f0631-24b2-4e2f-b739-2550c3faa646 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:3f69aa358a180290763bb9869dab61b6554f004dff7eea7ee83df7fec3e92c3d

Observation 738a5063-4efc-43b1-b186-08c188afc0df · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:3b76bdc3aba5fa2217ee33dcdde296bc2bd8467db90fcee273d0bebe2d90ea88

Observation 8a00c42f-93ed-45b6-b4d7-d44b3414654d · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:33798a82719c3e904769828d7a5e842cf998aa5e2bbb1f08011765c51ca34449

Observation 9d853f87-8e48-45c0-8b45-32ab1b217429 · outbound

This paper cites Transactions of the Association for Computational Linguistics , volume=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Transactions of the Association for Computational Linguistics , volume=

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a8c07ac4c5e8f42f4d7ea3cb0b9173f4e31fc38d3ca5aeb6e7bd08b064020168

Observation e6c15aee-048a-4732-9285-1d60034f3533 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:7da80902c4dba42d97964aa5712ddc399312b50c2907d275e42536420f2bde35

Observation 24d989d7-942b-49b6-92db-389cc43b5e57 · outbound

This paper cites 2022 , publisher=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training 2022 , publisher=

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:933462c829363564c36920a6eb2cb54a711c1a54364caab9ec2d0f73032cc77d

Observation 08536e4e-425d-4d7a-b1b7-7dbcb0784b68 · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Measuring and Narrowing the Compositionality Gap in Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:c869d878e47c95a2653e2250fde1c61250b7bfe94dc3cdf44cc6a31cc8a3b87b

Observation a4603ac8-d8d4-42c3-a630-992b2619fb6e · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:25b98f15b2b6c304acbf83c421106821a3cc51182e3db881003fa091952e81dd

Observation ede3a1b8-1754-4d9b-aea1-d9301445b575 · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:073ee478d45133d52f0c27c26644c82ce048c00deb84fb60c84a8f2729122992

Observation 5d5f57ff-b147-4ba9-a093-d6f2e22f6878 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:ccb647c40439b105c6d2163b818786441e4b089506bc1716eac2d336cc3b3744

Observation f2500176-f6f2-4021-ad3c-4b837c51ddf4 · outbound

This paper cites an unresolved cited work.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:1e17d084c88a0707bda463f7a0350eb64527cbc651e513fea37245aba2d531f6

Observation adb9a0e8-37d0-435d-adb1-9727a6ce19ae · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:4686f4db46beaef3c556099c4638e20eed8e0fff2fcf651bdc33e30d4b7dcfc3

Observation dc178778-6115-419e-bc51-0a4c66471888 · outbound

This paper cites Towards Efficient Online Tuning of.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Towards Efficient Online Tuning of

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:a2a097b0c98835e537f40b50738f08d6576ccad92864786fa83f92e6867691a2

Observation 32017a5b-3aed-4b52-83b2-f9939883db03 · outbound

This paper cites Qwen3 Technical Report.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qwen3 Technical Report

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:614a9569d7d3b84074474e555a48264481088aebc18dba0f68c3176103a68ba4

Observation b0e79b2e-d358-4394-ac5b-34614017d85e · outbound

This paper cites , author=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training , author=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:53a1331ce13efb9e306cd5c7e219f3f9d5237214ef5ddae26f1814ddae34d20a

Observation 1b4ac855-29fe-4c3f-bd0b-6a6e657fdfe0 · outbound

This paper cites Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:e3de080a699890edbe359ee3ff22e6cbc552df6334ccb53c24e6abc892309f4d

Observation f39e0090-b9af-40e9-bb51-fdc755be7ea1 · outbound

This paper cites RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:66217cab31eeced8404d428ea12764fbc4fa38383bbe9bff1e3eafc0ab663d6f

Observation 14758683-266e-4648-9a30-c9efbd253752 · outbound

This paper cites arXiv preprint arXiv:2509.22576 , year=.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training arXiv preprint arXiv:2509.22576 , year=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:577bdbb1a54c36a6aa267ff0c03d472490e80755c271fa6ab3771014735da7f3

Observation b155a972-255b-4d7d-9ff8-8174c9ea0e72 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:be307f94c5deb96507d649317091c3d89fb0361cbe2366ef3d4a10cded64c025

Observation f12623df-14a8-4c8c-baf3-cfd014d9d665 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:12528745709a49e377107f10b65b9afc75be265a8799b4ef2f8b8ffb145f6a45

Pith citing papers

No inbound Pith citation observations are available.