Pith. sign in

Paper Citation Record · LEDGER

TAPO: Transition-Aware Policy Optimization for LLM Agents

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.27973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27973 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T21:44:39.699017Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b4aa3bc-6628-42d1-a393-165418b0f174 · outbound

This paper cites GPT-4 Technical Report.

TAPO: Transition-Aware Policy Optimization for LLM Agents GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.585532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.585532Z digest=sha256:d428281a0027078750c614cefa2ddb7b2b75928b7c89ed186a5014f4d19e3668

Observation 5831a163-a6c7-44db-8028-17f4af99b395 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.588812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.588812Z digest=sha256:81f6dd61a02a4dd4c004536b53ca7aee57520c40eb477d174aa3e01cea3644d6

Observation b8e22867-69a9-4e42-88da-edfa1627f947 · outbound

This paper cites Qwen2.5 Technical Report.

TAPO: Transition-Aware Policy Optimization for LLM Agents Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.592129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.592129Z digest=sha256:4b2cd7a1e3fbc87ae2a0b2e930363495bcb94038ee93b4c05fdeebb65b8c8f9b

Observation a9fa85a6-9ae5-46c7-a736-b8f3626439dc · outbound

This paper cites DeepSeek-V3 Technical Report.

TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeek-V3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.595092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.595092Z digest=sha256:6fd79c08198e6cefad493a52d500d0bda48cb78e9edc27fb5c628d351b4410dd

Observation 1f02e486-dab4-461f-8114-afd539c2ae60 · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.598107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.598107Z digest=sha256:528b7cff76a32c33d75df0f7600485c3d4e2b6d37c817f2903dba3d33461b715

Observation 634df6b6-7e15-4fa5-8211-c2c498cd8216 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

TAPO: Transition-Aware Policy Optimization for LLM Agents GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.601537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.601537Z digest=sha256:e5931928441cf92075010b832ba35d53d83b5bdb44fb87264c3183678eb2dd48

Observation 903a6313-e57f-489c-88f1-55f281659901 · outbound

This paper cites Embodied agent interface: Benchmarking llms for embodied decision making.Advances in Neural Information Processing Systems, 37:100428–100534, 2024.

TAPO: Transition-Aware Policy Optimization for LLM Agents Embodied agent interface: Benchmarking llms for embodied decision making.Advances in Neural Information Processing Systems, 37:100428–100534, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.604589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.604589Z digest=sha256:1ad64a052c92f3cdbf3644761aba562b57a480292e61342f35ed715a550d9d9d

Observation 6556022e-ee80-4c45-81db-d214d24cd033 · outbound

This paper cites MIT press Cambridge, 1998.

TAPO: Transition-Aware Policy Optimization for LLM Agents MIT press Cambridge, 1998

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.606756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.606756Z digest=sha256:d718aa2fe197829aa39fe47770e88d068f162628d2ff5b216f50b08d04ef70ea

Observation 385bcef2-03b7-4696-b0e5-0e3a6d3d3b84 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

TAPO: Transition-Aware Policy Optimization for LLM Agents Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.609034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.609034Z digest=sha256:14efeb17c548827f02a58f7604ef9d8100795671506c2dc6769bcfd87443d240

Observation 71c2bca8-b13a-487f-9e18-37234ffc5b1b · outbound

This paper cites OpenAI o1 System Card.

TAPO: Transition-Aware Policy Optimization for LLM Agents OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.611311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.611311Z digest=sha256:85729c97b4198f2b0e513d1174d81e45fe0fe8762c0253fdd2d134d807337c3c

Observation 411df13f-8af4-443c-8036-881b431581ac · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.614561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.614561Z digest=sha256:ad43535fa73ddf130dee6ea67a21af91f22bfeb26b7241dce3308f8a7195dc7f

Observation 888a4484-cdc8-4d06-85cc-3db0d89068bb · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.617527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.617527Z digest=sha256:1e586890c761fce0855afb9b46c93bb0a6e87a19394601a47ee63e246a4e289a

Observation c8735516-36a0-43e9-89f9-5203bb40c61a · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

TAPO: Transition-Aware Policy Optimization for LLM Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.620076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.620076Z digest=sha256:171bc0bfa6a67a453537b81624cfeadac84e3f3400a515adb526ae42b7f7c635

Observation 1c8ae0c1-cf39-4b32-98a2-c000c4af9b68 · outbound

This paper cites General agents need world models.

TAPO: Transition-Aware Policy Optimization for LLM Agents General agents need world models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.622957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.622957Z digest=sha256:ba2d5eb7694278e27dc23ad60d724807aa9f7c37c3d3dadc1806a52952277585

Observation c7e4a8bd-41a0-4f7a-87c1-794798b2141c · outbound

This paper cites Reinforcement Learning with Unsupervised Auxiliary Tasks.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reinforcement Learning with Unsupervised Auxiliary Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.625477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.625477Z digest=sha256:7618ad85ee0af9fefbf9492e7a608fa002579d20a91fb2e40b11dfbb3f9e4999

Observation e6db1d2c-16ab-4d3a-9085-5c7787abe26a · outbound

This paper cites Deep- mdp: Learning continuous latent space models for representation learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents Deep- mdp: Learning continuous latent space models for representation learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.628079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.628079Z digest=sha256:30525b40891f2632fae9554aa2e2b3e8fa15d8a8488093e099a84c06c7f30644

Observation 4d86f2f0-9407-47eb-9b2e-6975c1723600 · outbound

This paper cites Vlms- guided representation distillation for efficient vision-based reinforcement learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents Vlms- guided representation distillation for efficient vision-based reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.630204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.630204Z digest=sha256:7cc47616e45b43f32df0dac00258217990c108ddc33f47c3f2f380e9f8941909

Observation 5328c020-c40e-4471-a32f-1bef29fd3f08 · outbound

This paper cites Data-Efficient Reinforcement Learning with Self-Predictive Representations.

TAPO: Transition-Aware Policy Optimization for LLM Agents Data-Efficient Reinforcement Learning with Self-Predictive Representations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.632400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.632400Z digest=sha256:f1069cc8d186735c3f731c02875c98f9adacbfaeb8f9f4e2b66b013cbc798746

Observation 5d95d0b2-43f3-403f-98bf-6a6ec15cb62e · outbound

This paper cites Agent Learning via Early Experience.

TAPO: Transition-Aware Policy Optimization for LLM Agents Agent Learning via Early Experience

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.634912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.634912Z digest=sha256:5cd994284d2804395f1857c6b00ad7c74c7606d60df5b725756ed62590c9c537

Observation fb474bf4-dcbe-46be-bb53-1dfdf73bab9c · outbound

This paper cites Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reinforcement world model learning for llm-based agents.arXiv preprint arXiv:2602.05842, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.637487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.637487Z digest=sha256:2e1a834374cb9fda5ce4e6039549dd7149f5154dee2d1f950dbccde00244a171

Observation bdac9022-35f5-449c-8d0b-a75e148c70f1 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022.

TAPO: Transition-Aware Policy Optimization for LLM Agents Webshop: Towards scalable real-world web interaction with grounded language agents.Advances in Neural Information Processing Systems, 35:20744–20757, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.639844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.639844Z digest=sha256:176c6cae36c1013af710ccd52f44d4e2b14b5a4780b6158ab8dfbcd6d611e161

Observation 4c015d4a-6083-4f83-9e76-5284edfb809e · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.642114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.642114Z digest=sha256:208fceb4ffe85ea053b77fce750dc873041bd29eb3bd8339400554e38b2ac2f5

Observation aac405ef-6282-48e6-be71-c8d28ba29d35 · outbound

This paper cites CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges.

TAPO: Transition-Aware Policy Optimization for LLM Agents CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challenges

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.644614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.644614Z digest=sha256:a8da5116e0a9756c4b0c1e9a5c82045564d6a868c1def6a7d2cd1079717ae79e

Observation 65cc9ddf-b3c2-4104-9aed-8f694fcc11c9 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

TAPO: Transition-Aware Policy Optimization for LLM Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.647018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.647018Z digest=sha256:b8527964c89c6dc295dc1cdbd2b9877d586f17ca0bd869979379746c5e379711

Observation 2d413b60-1d57-46da-87c7-279b4f935e39 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.649499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.649499Z digest=sha256:1127663ffb164288bdc3b38c5a1303ededfae49f05728c967b92b8569921b601

Observation 6c0496fa-01bb-4c21-8b8d-35eb7d750d84 · outbound

This paper cites WebDancer: Towards Autonomous Information Seeking Agency.

TAPO: Transition-Aware Policy Optimization for LLM Agents WebDancer: Towards Autonomous Information Seeking Agency

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.651881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.651881Z digest=sha256:54cf667a7b724bb4126fc1377b919cb0023358a3f1877368b1d05819010ee517

Observation d3d1f9e4-8cea-4cb8-a25d-12a5c6f16966 · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

TAPO: Transition-Aware Policy Optimization for LLM Agents WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.654314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.654314Z digest=sha256:0d98e63bc2b1660898342538200a68659eec61b7115cd50959ece89e4d8b7803

Observation 32e21fec-cc70-4b66-83e2-2944416c30cd · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

TAPO: Transition-Aware Policy Optimization for LLM Agents Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.656683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.656683Z digest=sha256:cde0b5374e29d8ef0dec1e698090ad62c48ea3e189db4130f15ace64c8ebd408

Observation e1bbfc75-3b80-4683-8353-be3af30bad12 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

TAPO: Transition-Aware Policy Optimization for LLM Agents React: Synergizing reasoning and acting in language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.658791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.658791Z digest=sha256:d575aff810147d0a3d85d02da0d034efe1564bb95de360c4edff915cdc917dff

Observation 774269c8-0ef8-4ddf-8c5e-e370ef54b1d2 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reflexion: Language agents with verbal reinforcement learning.Advances in Neural Information Processing Systems, 36:8634–8652, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.661240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.661240Z digest=sha256:72b352e47772aa30c343435c073be1f418028700164bf125d239f4b32aeccf34

Observation 05423b04-eb26-4c2c-a480-7126191203d3 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

TAPO: Transition-Aware Policy Optimization for LLM Agents Fine-Tuning Language Models from Human Preferences

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.663478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.663478Z digest=sha256:a12be0102caaafecee76fffaceda4914ea8ffe720422a2277662707fce478683

Observation 189ff840-67d9-4843-8c0e-de4eae2f4613 · outbound

This paper cites Proximal Policy Optimization Algorithms.

TAPO: Transition-Aware Policy Optimization for LLM Agents Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.666212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.666212Z digest=sha256:8d7565f7d41182071d2abad3233f1823f850f88bf27a8de11273effc4c92c8f6

Observation b915ecb9-cb16-42a9-beb5-851cc140193c · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

TAPO: Transition-Aware Policy Optimization for LLM Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.668576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.668576Z digest=sha256:e1328cc2939c1a093333b50d3019e11b0fbeb39aaa2c79c1780ffc83df7e78ec

Observation 858ec9ad-6c65-497f-9cb6-c0e3ea8e723b · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

TAPO: Transition-Aware Policy Optimization for LLM Agents Understanding R1-Zero-Like Training: A Critical Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.670928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.670928Z digest=sha256:764d34bec2c07cd1018c0b8c2d9f541cc1c17aeb829caa36ce3d948d8cfbc1b5

Observation f06992f8-47c7-4e56-8a08-c7341f06a8c4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TAPO: Transition-Aware Policy Optimization for LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.673802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.673802Z digest=sha256:b8add49fe032f8bd801f183776e93ad2366260eeb75045f56be2b73e1cc4ea09

Observation 0ca8784e-c1de-493f-8d81-83c2274d1d71 · outbound

This paper cites Agentic Reinforced Policy Optimization.

TAPO: Transition-Aware Policy Optimization for LLM Agents Agentic Reinforced Policy Optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.676348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.676348Z digest=sha256:6588a7c163c93ee886f8c0b59f68a173490ee63e5a3ba1498c6e01acff6854ad

Observation d929e877-9682-493b-9c17-ef7bfa0dd1f8 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

TAPO: Transition-Aware Policy Optimization for LLM Agents Dream to Control: Learning Behaviors by Latent Imagination

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.678831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.678831Z digest=sha256:2c4c9789d61d845ef5dffa29fba9beed5228421bb2e0e92934b77af6d837c2c6

Observation 93875012-557a-42a8-90df-798d46b23199 · outbound

This paper cites Mastering Atari with Discrete World Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Mastering Atari with Discrete World Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.681305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.681305Z digest=sha256:a63f70c07ece5b63b6d903ffa7f25c3b41c01b0e721d7f7913b7810617272ef4

Observation 6c8da77b-8f99-4c4b-8fb3-d0898bc32cc7 · outbound

This paper cites Mastering Diverse Domains through World Models.

TAPO: Transition-Aware Policy Optimization for LLM Agents Mastering Diverse Domains through World Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.683897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.683897Z digest=sha256:1cba18d79ece2e0cd3a22c8b3cba492343e10382b9cf4666a9f4908afb446c0c

Observation 9d205fcf-7e3c-482f-8f60-5dbe3578bfcc · outbound

This paper cites Reason- ing with language model is planning with world model.

TAPO: Transition-Aware Policy Optimization for LLM Agents Reason- ing with language model is planning with world model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.686699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.686699Z digest=sha256:715945a68b69da2d0653f2413324c7e0d0c940106d5b00d762939f7ca7ebf881

Observation 5aa4ef01-dc0d-47cc-a9e7-253c083f8ffb · outbound

This paper cites Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation.

TAPO: Transition-Aware Policy Optimization for LLM Agents Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.688810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.688810Z digest=sha256:19ec9353107c1a25a77cad097ac451b7fdd7e12c3d59bc49083305b3e72ddebc

Observation f202d84d-b982-4228-9ab8-288546579a13 · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

TAPO: Transition-Aware Policy Optimization for LLM Agents Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.691367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.691367Z digest=sha256:53e1046e5cdc1de933d68ced11328c57e89fbfe40e2c3c1cc92109cd560612a7

Observation 53c918da-4385-4142-8181-dcf78676b305 · outbound

This paper cites WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model.

TAPO: Transition-Aware Policy Optimization for LLM Agents WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.693932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.693932Z digest=sha256:3fd181c527725a90472cf98fef034be853ae4722ca9fd1e9730a7faf0a385979

Observation 6198f19b-ce98-49eb-b616-8555e4fa7ad7 · outbound

This paper cites RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents.

TAPO: Transition-Aware Policy Optimization for LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.696164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.696164Z digest=sha256:05979a6bcef8d3255143ce6421f4b8c8f782527ebbba79a58de429b81d15d5c9

Observation ff20e295-7b85-4f67-b39b-369ce0b7a07f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TAPO: Transition-Aware Policy Optimization for LLM Agents Training Verifiers to Solve Math Word Problems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.699017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.699017Z digest=sha256:032618716476383b4971a4dbea88104b0021bcc3b8b85529d6437bac4305534a

Pith citing papers

No inbound Pith citation observations are available.