Pith. sign in

Paper Citation Record · LEDGER

Learning Agentic Policy from Action Guidance

As of 6 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2605.12004.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12004 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:02:49.206053Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:34:39.055539Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact56
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 1639d43e-7ea4-4e77-9b48-1f80e58be624 · outbound

This paper cites Claude Opus 4.6 model card.

Learning Agentic Policy from Action Guidance Claude Opus 4.6 model card

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.831018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:0db0bad4e23c8f581b10fa6eca56358812fffab809135f21efa82a38da7ca969

Observation 5691e81c-db21-469e-914f-a15cb41526d4 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Learning Agentic Policy from Action Guidance $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.802510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:63450e673374cb01ce745ace7dcbad695fc3af48a8f94a64cd5989b72e2449a2

Observation cf10f969-fe6e-4803-be0b-9cd962e0eda4 · outbound

This paper cites Fine- tuning web agents: It works, but it’s trickier than you think.

Learning Agentic Policy from Action Guidance Fine- tuning web agents: It works, but it’s trickier than you think

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.816125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:947c76dbab0f1dec3f855f7d44be73bdc5089cba7a49e64035ea2f66442389a1

Observation f2176436-10eb-4ea0-bbf8-da2b115a9287 · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

Learning Agentic Policy from Action Guidance SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:43:34.765363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f4d98e490e76e5829561df821bf89f77e883a25c4b0538aa275f5c6a67c8f5d4

Observation a2d792c1-5c8e-444f-8ae2-27ecb873b73a · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

Learning Agentic Policy from Action Guidance xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.799791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:68623f51e51bd43a1c4bc08b812fffefb783ad291fa6aa210c7a6e88803d948c

Observation 357d6141-dc4b-4d87-8a14-a127566deef6 · outbound

This paper cites Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning.

Learning Agentic Policy from Action Guidance Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:33.109892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:baba4f90e03bbedc5c8fc6449f0064f3b936075689161678474737df9405638a

Observation 987e40b8-e036-4fc2-85eb-758838d08a28 · outbound

This paper cites GPG: A simple and strong reinforcement learning baseline for model reasoning.

Learning Agentic Policy from Action Guidance GPG: A simple and strong reinforcement learning baseline for model reasoning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.825455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:55dba615d5a5eac6b587f5042d7697ce8bda34b710b876791122c1e72f255492

Observation fbe351f8-dc2b-4f91-8b4d-b0e3a4584bae · outbound

This paper cites Redsearcher: A scalable and cost-efficient framework for long-horizon search agents.

Learning Agentic Policy from Action Guidance Redsearcher: A scalable and cost-efficient framework for long-horizon search agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.587389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:90611f385121f5adf25b6d1fec021cf48c743780194b57794f0a2cef9841c35b

Observation b3e77c83-4a51-4118-9c6d-fc14d1fed60d · outbound

This paper cites Harder is better: Boosting mathematical reasoning via difficulty-aware GRPO and multi-aspect question reformulation.

Learning Agentic Policy from Action Guidance Harder is better: Boosting mathematical reasoning via difficulty-aware GRPO and multi-aspect question reformulation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.776771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:fa855ce22d9a0996378676c79cb1ea58ce491b98b1390bc5acb7bbe56026b1cc

Observation de9932e6-8f5c-4a75-86a9-c4c8d28d60d8 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114.

Learning Agentic Policy from Action Guidance Mind2web: Towards a generalist agent for the web.Advances in Neural Information Processing Systems, 36:28091–28114

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.818141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1433b66dba4789e5ba20ceadcda1fe411645916f069b0797fa93a8f652c12e74

Observation f5550734-41ba-41e7-80b5-8341adc77647 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

Learning Agentic Policy from Action Guidance OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.519833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:a8b303be4f1f323c4f895092148fc1c0c7e1a402c63d938bf8cb1ab91570aa22

Observation a0d8fdd6-9f20-4899-a0f7-7fd7a4f750e6 · outbound

This paper cites Wildclawbench.

Learning Agentic Policy from Action Guidance Wildclawbench

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.823623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:3fce1cd52a402f85d13780428dac69b14f4000e89b66fd09fede70b1e812b3a9

Observation 751d4fcb-e620-4ec5-abcc-dc6947058c96 · outbound

This paper cites Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning.

Learning Agentic Policy from Action Guidance Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.727198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d663151c9c7c5ce8dbd9169c3de3222edd84f4ccb7d4984c20d38680f74a9134

Observation 37eda508-fa87-42b7-be1d-f85b7e9c116e · outbound

This paper cites Agentic Reinforced Policy Optimization.

Learning Agentic Policy from Action Guidance Agentic Reinforced Policy Optimization

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:57:12.173630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:e2d0917338dd200d3638a5f714e8227c5cebf252dd0d7c6c7907fc2aed4e2f31

Observation bec250b6-7d34-4c30-bf01-d732a8d11b3a · outbound

This paper cites Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence.

Learning Agentic Policy from Action Guidance Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.621909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:9de811cdafa30024fb42977107c95ae8e2fee70c0670e40ccf9bcfe7d8e1bb18

Observation 5efedcd7-945a-4fa4-8eda-6e88d8374c97 · outbound

This paper cites Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks.

Learning Agentic Policy from Action Guidance Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:32:18.834213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:a2459c6e9579b0ad7a0db7438741e28125d31a8399881711ac0854bdf83b273a

Observation 4a61f7ee-90f2-460f-9b23-03f66a981999 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

Learning Agentic Policy from Action Guidance Group-in-Group Policy Optimization for LLM Agent Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.632913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:7053bf499ade799889e606a2306ec7ac17a373fe2a2d357f6b5bec8c6af84f51

Observation 859ffaa4-7137-4c9d-bc2d-2b543d12052e · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

Learning Agentic Policy from Action Guidance SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.645623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f8368944aeebbd4f0ceafe4fb94b024053268901b428f1a1f229f713dc8a1d43

Observation 1e817b8d-464e-40b4-b648-316d0bf06b9a · outbound

This paper cites Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl.

Learning Agentic Policy from Action Guidance Beyond ten turns: Unlocking long-horizon agentic search with large-scale asynchronous rl

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.649435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:a5d55a4e7679a805b9e60465c0793bd25b20e2dd865b5906ae8a3e5f4ee93abe

Observation 009f5b83-d135-430d-a75b-484f7a33c050 · outbound

This paper cites Actor-curator: Co-adaptive curriculum learning via policy-improvement bandits for rl post-training.arXiv preprint arXiv:2602.20532.

Learning Agentic Policy from Action Guidance Actor-curator: Co-adaptive curriculum learning via policy-improvement bandits for rl post-training.arXiv preprint arXiv:2602.20532

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.590870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b307705e326ba6efcd9df176b68634496b95414ea391d0cf5b96f7ef69ef6674

Observation 3af81f34-30cf-47d1-bbd8-d1e391518386 · outbound

This paper cites Deep q-learning from demonstrations.

Learning Agentic Policy from Action Guidance Deep q-learning from demonstrations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.778716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:6ed0aa61bdb451ffa9d8fd95f41417e7faee841292fd83feea5447b0e43a00f5

Observation 50fbd06d-8e5a-424c-bb98-d95aa186f427 · outbound

This paper cites Boosting mllm reasoning with text-debiased hint-grpo.

Learning Agentic Policy from Action Guidance Boosting mllm reasoning with text-debiased hint-grpo

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.798467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b219f47d81908d044810bff7bf224c0ccd379288a930596d191fee25477d9d22

Observation b82336af-dcd4-47ec-bc91-ffc20befc04b · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Learning Agentic Policy from Action Guidance Reinforcement Learning via Self-Distillation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.597502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:dbccfb6865acffd8dee5e05330729f00bfffa072cc6704d97c8f69b40cce951d

Observation 08144029-b394-4e13-8a6d-176e06855d6c · outbound

This paper cites Tree search for LLM agent reinforcement learning.

Learning Agentic Policy from Action Guidance Tree search for LLM agent reinforcement learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.780576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:20d61c03f13c43b22446c434ef208d45fdd3735d680a22cb547b5716b75bd120

Observation 0be9105e-bea4-4158-beac-d4f44700b3c5 · outbound

This paper cites Thinking with map: Reinforced parallel map-augmented agent for geolocalization.ACL.

Learning Agentic Policy from Action Guidance Thinking with map: Reinforced parallel map-augmented agent for geolocalization.ACL

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.812409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:493dd69d5ad18b43fe77d48e60720dfb76a81afb8fb580fd834e79317d73d1b2

Observation 8635f681-43ae-46dd-94c9-dfaef4cb5faa · outbound

This paper cites Vcrl: Variance-based curriculum reinforcement learning for large language models.

Learning Agentic Policy from Action Guidance Vcrl: Variance-based curriculum reinforcement learning for large language models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.676707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:0c6786d6f5168d6ee4be4de2887bde1a6d5808acdb94afdad3f188257c7d8037

Observation 80061224-480d-4b83-b19b-6846e08e04c9 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Learning Agentic Policy from Action Guidance SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.688917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:5771eea9f03e6a91e818e2fba60da9673c68665983303c806df95222242a2d56

Observation afda1e21-8088-47ae-9ae2-b7a6effdc9d4 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Learning Agentic Policy from Action Guidance Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.606250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:a9ac906309286c3140aebe412bcdd6724069ab42d1dcaf0635f7084f635da289

Observation 52a638ee-6fb7-4688-97ab-09df64d5fd0f · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Learning Agentic Policy from Action Guidance WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:37:09.773663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1033647dc7d07027cab7f4c0978778a0a78a068e6924beb4b65251b63b760818

Observation b970073b-9c08-4745-9f20-11aae7c9e7b9 · outbound

This paper cites Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting.

Learning Agentic Policy from Action Guidance Adacurl: Adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.814168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:2852c5618525b9874ee92b9a038eac0f629309f3cded0cd5c41e4f1cdee5ff7c

Observation 55358957-318a-4682-941b-47aa643d0dac · outbound

This paper cites WebThinker: Empowering Large Reasoning Models with Deep Research Capability.

Learning Agentic Policy from Action Guidance WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:14:25.573630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:21fbee428ef971e9b736bf786f6ec2c24f274a836458e19592a804eaee41473e

Observation 92c5b202-9bfe-4c5c-b64a-08617c072a60 · outbound

This paper cites Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model.

Learning Agentic Policy from Action Guidance Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.618672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:57ecb8d0237b0f19f3ebb74b8a9b1b104c87c13f619e2b03d6ebd9a62a7da739

Observation cfd20f2a-199e-4d14-ae46-5d9da119b7bb · outbound

This paper cites Guided exploration with proximal policy optimization using a single demonstration.

Learning Agentic Policy from Action Guidance Guided exploration with proximal policy optimization using a single demonstration

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.821791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:5b6e75a12ad3d2e1fbbd8666d1bb39ab95f5a7e18607aab85d0b6400855c104a

Observation 66b3e568-7ed0-4777-bb46-04d01342c9e7 · outbound

This paper cites Truthfulqa: Measuring how models mimic hu- man falsehoods.

Learning Agentic Policy from Action Guidance Truthfulqa: Measuring how models mimic hu- man falsehoods

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.820085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b9f8a3dbf99f882e4af38a96e070890a46f269787293c7c74011d86378529198

Observation 8caa542f-b769-4914-8565-be8a51ce2d1a · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning Agentic Policy from Action Guidance DeepSeek-V3 Technical Report

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.796733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1ca40c5dbb3f0fea389f2a3abc488acce1982079267de1b8b40a6bbcb9124448

Observation 2361523b-ecd7-42c7-a7aa-dac6d962aa89 · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

Learning Agentic Policy from Action Guidance Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.594178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:41d2d858fee013c73c93cff9edf96b38bf880cadafa8142730e664e724128718

Observation b737ab17-d520-460c-a169-a0d582078b9c · outbound

This paper cites Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions.

Learning Agentic Policy from Action Guidance Learning what reinforcement learning can't: Interleaved online fine-tuning for hardest questions

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.787952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:e776ade6da911803a14cf4b230feaf911bdd3855bc0a469fbc134975e5688ef3

Observation 826ac4a2-1c2d-4ade-9115-ee3388b75676 · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

Learning Agentic Policy from Action Guidance SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.776747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:bcac670fb6610b14e90a8507b003f2d9c7b435ae7c53579519b5187c8ba78ee3

Observation cf36beac-9f5d-45f3-9e96-e84907002074 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

Learning Agentic Policy from Action Guidance GAIA: a benchmark for General AI Assistants

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.603371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:1f6c2133c44a2566c0bcae03a474ee360a3f19a2aaf85e123bb64ded75395557

Observation 14e1b130-4e90-4896-aea3-d3393945bed6 · outbound

This paper cites Minimax m2.1 system card.

Learning Agentic Policy from Action Guidance Minimax m2.1 system card

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.829342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:243d731689b0fc7c59886cb81c304e2bace037484c6a19e295c4be812fcd51d9

Observation bdfa9d7a-d43d-4f0b-b3dc-39f5de834860 · outbound

This paper cites Over- coming exploration in reinforcement learning with demonstrations.

Learning Agentic Policy from Action Guidance Over- coming exploration in reinforcement learning with demonstrations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.806867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:c21b264b96f9c03b060ea1c16991681ab49b8fbb61163cced2d5edc9d0880920

Observation 2cf22fff-26b7-4055-817f-97f6907687bc · outbound

This paper cites Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models.

Learning Agentic Policy from Action Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:07:17.751601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:bade72e0cb4ec3b6e5ef854d85ed32b138f8cf0180f13ba6b89db2c623b5e20b

Observation a00052b6-468d-4553-a72a-21ff3ba94931 · outbound

This paper cites Gpt-5.4 thinking system card.

Learning Agentic Policy from Action Guidance Gpt-5.4 thinking system card

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.810613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:cf519cc0e572392cf8e7b851f6425782387ab93bb457e10f25ff5a553854a2a6

Observation b097e1fa-7b21-481a-acf6-bde3b98e240a · outbound

This paper cites Iterative reasoning preference optimization.Advances in Neural Information Processing Systems, 37:116617–116637.

Learning Agentic Policy from Action Guidance Iterative reasoning preference optimization.Advances in Neural Information Processing Systems, 37:116617–116637

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.802870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:eb364dc450707a2c69d26073927d613e12a76afc111ad3ae15a2c8f777bc8d2a

Observation 2fba3040-3165-4db1-8b1a-54dc1ca17d4e · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Learning Agentic Policy from Action Guidance UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.699644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f9695b69d27788b2c0ab09bfbe55e0b26aa8f8b7cc444e894ad2300ee83141be

Observation a57effd7-b998-4217-8826-62f5937f4fe7 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

Learning Agentic Policy from Action Guidance Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.755910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:71bca165c92e018b8d012529d171926683629581ac61e6e31642861289ef964e

Observation 9c44e73f-ff6c-4474-9312-9bd559a419c2 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Learning Agentic Policy from Action Guidance GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.773989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b9d5bdd03b00f1a49be58851470aeb2c57d971a39c238998b54a7db87d01dcbe

Observation d467159e-4082-404f-8d31-199b819351da · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning Agentic Policy from Action Guidance Proximal Policy Optimization Algorithms

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.600157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:9bb20358bb0a9cff775c259cb797819fdeff19e213ff1b270fbb5cfe58648c62

Observation 1633b736-430c-4d9c-862d-6bc615c3575a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning Agentic Policy from Action Guidance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.696982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:599e317725f3eac2e2e9966dd8b4f969581bc0289049e5f4a1b342ca02c2aa5f

Observation fc9c0f91-b266-4d95-b705-4562cdd3e576 · outbound

This paper cites Self-Distillation Enables Continual Learning.

Learning Agentic Policy from Action Guidance Self-Distillation Enables Continual Learning

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.685825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:338f878588a66cdf3fd9d98a2ef779df0bb1dccd65449bd544eac571ec42650b

Observation fa15cef6-ae82-438f-9a32-97dd9aa78c15 · outbound

This paper cites OpenAI GPT-5 System Card.

Learning Agentic Policy from Action Guidance OpenAI GPT-5 System Card

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.679643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:bb111d7f7b0624935e99951f1b5ae22b14a890f1705cb356a9dd7b35eca63783

Observation 95c47958-1b89-4fc5-ad6c-b886c217c52a · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Learning Agentic Policy from Action Guidance Kimi K2.5: Visual Agentic Intelligence

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.768488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:9a5abb88989274d4c57621cbe248abd721a5b56e4c0ce94efefc06582d6ee72e

Observation 485864f6-187f-404a-a5ab-3f6c1c9529ab · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February.

Learning Agentic Policy from Action Guidance Qwen3.5: Accelerating productivity with native multimodal agents, February

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.805037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:ddfe87eb2ae93259acb247e0f6586c963319a7efb4e1e976d53c8290c3f3cf84

Observation e2912aa8-4b67-4e68-b97c-f70d1e74a105 · outbound

This paper cites an unresolved cited work.

Learning Agentic Policy from Action Guidance Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-13T11:07:39.800591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:291055aa3cdcebbf9df0fea32ec513fb48662278039a83639836b53a47c8b409

Observation 68746824-b279-4f01-862c-ff4041c7e0d0 · outbound

This paper cites Tongyi DeepResearch Technical Report.

Learning Agentic Policy from Action Guidance Tongyi DeepResearch Technical Report

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:56:57.433092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d09ab1dd5aaf353f5a6948080dbaa3b591381e18451e4d898f50b15a0d4b08cf

Observation f4e2c5f8-93f7-4bc0-b219-a26eaf80f776 · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965.

Learning Agentic Policy from Action Guidance Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.774687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:0975782c3867ce2f560c5a3431867ffffc07808b31786188285ddb4f6d0492b1

Observation 8dc39b33-8e3c-4204-ba86-c86d6618adae · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Learning Agentic Policy from Action Guidance Deep Reinforcement Learning and the Deadly Triad

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.711777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:acabcd2b4368d4a923577558e0a3713617adbd8d3608e34af024613a5a2f0d93

Observation 87baeeb7-af4a-4cee-b02e-f8467ffb4008 · outbound

This paper cites Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards.

Learning Agentic Policy from Action Guidance Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.758867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:2163e400baec0a6b503cc2427194bd4bee8732659d1261a3cb62a3400d56eaa6

Observation 74e5c648-4a6d-4977-9805-bacabdda8f9f · outbound

This paper cites Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem.

Learning Agentic Policy from Action Guidance Let it flow: Agentic crafting on rock and roll, building the rome model within an open agentic learning ecosystem

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.761647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:9734e914b7625be9b15f4ab1cb1c225e49332996e92a5cb44ce00181eaf2734f

Observation db34b9fc-5d71-432d-993e-62773e8d8907 · outbound

This paper cites OpenClaw-RL: Train Any Agent Simply by Talking.

Learning Agentic Policy from Action Guidance OpenClaw-RL: Train Any Agent Simply by Talking

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.765019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:e78ece9fc2697d9e6e784c11750c9d8536322c4418ad6cdea99918aaff1f2494

Observation 7711b20f-5074-4cc0-b9e5-82086d47af10 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Learning Agentic Policy from Action Guidance RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:b04cb0b98f20f017753fa5ed8196245fe3d2d0a43660e650c016da690a957d96

Observation f9d25609-d38b-42b2-8332-d678f8004c54 · outbound

This paper cites Agentic Reasoning for Large Language Models.

Learning Agentic Policy from Action Guidance Agentic Reasoning for Large Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.657956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:0e67a75ce34b19ed44741352f68fa00d5a769b0d9b4774ce4378474153c76eec

Observation f4c1fd2f-5f2e-44e6-af5b-fb0dce07b088 · outbound

This paper cites WebWalker: Benchmarking LLMs in Web Traversal.

Learning Agentic Policy from Action Guidance WebWalker: Benchmarking LLMs in Web Traversal

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:07:17.791143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:a3461d4549bdc843452ba3f9834db0f652ffeba4319dbcd01cf34ab6296c70ab

Observation 82846ed5-0a12-479c-a3b3-b13b3b5eddd7 · outbound

This paper cites Learn hard problems during rl with reference guided fine-tuning.

Learning Agentic Policy from Action Guidance Learn hard problems during rl with reference guided fine-tuning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.702708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:59d2d739c48975530dc695bc1f5ff9ec315227b17d190468f70e22da5f14d9ae

Observation a51275f2-6cbc-4d04-9b51-8f2c2dfe705d · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

Learning Agentic Policy from Action Guidance Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.808900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:bd62a372c2b1d688b3e3dd7fa18dc7d3f3a3402ee8d5eaf00e4001ba07cf3228

Observation 1d69ddf5-36fb-4f9f-ae7f-c240932e4ed3 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Learning Agentic Policy from Action Guidance Learning to Reason under Off-Policy Guidance

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:17:03.075876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:47cc14047483ddd40c9f60a0ddef752ad6cd297f9491b640cbd16253500708ac

Observation 64b3bdc3-82c7-4ed2-8889-7e699ce8bea5 · outbound

This paper cites GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL.

Learning Agentic Policy from Action Guidance GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-27T02:04:34.683381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d9dc7f2876d42abf0fd400915f6daddbf51c54c0cf280ee0463e85843be39632

Observation 9f6a887b-9d47-4c88-ae0f-10733a8b4686 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Learning Agentic Policy from Action Guidance React: Synergizing reasoning and acting in language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T11:07:39.827410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:c52b7f4cae5e184e355834573eb997eecb814f92aa31990c540f974ee25c2ea1

Observation afad9ae9-cc3b-4580-a25f-ea603716bab3 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Learning Agentic Policy from Action Guidance $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.666389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d7a14b4497449771adc8d6256af51763ed1632bbb44a054eecf553c12e5dcf76

Observation 69d3835c-60f0-442b-9165-e06fad3beb15 · outbound

This paper cites Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms.

Learning Agentic Policy from Action Guidance Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.669790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:2092dce41bd36dcc899c9307ff6183589934c2e3531a5addf2983d2123b1735c

Observation 7f6c0cb2-ad96-4afd-8d81-644cb60a574a · outbound

This paper cites arXiv preprint arXiv:2603.21383 , year=.

Learning Agentic Policy from Action Guidance arXiv preprint arXiv:2603.21383 , year=

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.663444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:528f45a59db6f6f594d95ba1f1f532d95ac1b1d7b97831e38bda4623c6c91a8d

Observation 41b76d43-58f0-4bec-a003-97980b44661c · outbound

This paper cites MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework.

Learning Agentic Policy from Action Guidance MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.656657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:7385c30e1e9722200a30596e7db017006c0feadd4a5591606468f28c608f6076

Observation 1c777d36-c31c-4b72-bea6-5de65d9332bc · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learning Agentic Policy from Action Guidance DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.694521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:84acd8926ecdb4bf55a09f1ede7ebef10722401ecad63918cedbbb600a1f69c6

Observation 66a6d182-aa4a-4d38-b0c8-a96694d6572c · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Learning Agentic Policy from Action Guidance Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.659881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:f7063997fe32e359f51f982dc5ce04c16ea7191bbbdd268828524b30018c9dd6

Observation a850c3b3-a2dc-4ad3-8e8c-acf481ef8173 · outbound

This paper cites Agentevolver: Towards efficient self-evolving agent system.

Learning Agentic Policy from Action Guidance Agentevolver: Towards efficient self-evolving agent system

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.785140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:6f50e64ad9b959a66aad1e00a103934343906c35626599ea7f8dd444774a16a8

Observation ff787be5-ad14-4556-bcbd-6f07b151209e · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

Learning Agentic Policy from Action Guidance The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.608879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:8ba5bc987f1e0c5d367c52fd2953d8ce36324ca3f404f35cb1ad876c95bd0996

Observation 5eddc1e9-79b6-4ecb-998d-34d290444f99 · outbound

This paper cites arXiv preprint arXiv:2508.11408 , year=.

Learning Agentic Policy from Action Guidance arXiv preprint arXiv:2508.11408 , year=

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:07:17.705737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:9f5cd6bdff311b94d89eef51a2504506b17e3230ffef3d8f08b7b2096980784e

Observation d45a8540-4c29-4f0d-8030-aef3cc060ce8 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Learning Agentic Policy from Action Guidance Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.779421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:d0d5ac4e688e3cf667c293afd7f46e8367d2b156df2ef0b90d5df8a5332b7e54

Observation 4bd1baa4-5550-45eb-b3a4-96e166e18099 · outbound

This paper cites Prosperity before collapse: How far can off-policy rl reach with stale data on llms?.

Learning Agentic Policy from Action Guidance Prosperity before collapse: How far can off-policy rl reach with stale data on llms?

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.708884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:647a6f1a05f8cd7c58caa030af9282157792b6bc578ec4fffef003af0876e219

Observation 379ba4ab-3a65-469a-bd7e-d9cdcdfc50cf · outbound

This paper cites Code2world: A gui world model via renderable code generation.

Learning Agentic Policy from Action Guidance Code2world: A gui world model via renderable code generation

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:07:17.639387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:0ca8b31cbb2528ef20e5555665f8b9642aa86a84110acddde824e4cc4bc966bb

Observation 66043e16-8a57-43b6-9729-a26b0da6908c · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Learning Agentic Policy from Action Guidance Instruction-Following Evaluation for Large Language Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:07:17.642066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:08b36df43134095670d3efc53a8ccdce63609c1f05866a7fcace32c38de437ba

Observation 3bb80b65-9f3d-4e17-a552-07be46bb0b8d · outbound

This paper cites BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese.

Learning Agentic Policy from Action Guidance BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:04:50.009871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:e6fd392ac66cbf4baab51b3bbe2b80695b213b09eb8e20c8147f01a28a094a59

Pith citing papers

Observation 1e1e43b4-438e-4481-bbd8-f32e36544441 · inbound

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning cites this paper.

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning Learning Agentic Policy from Action Guidance

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-08-01T07:39:17.666069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-01T07:34:39.055539Z digest=sha256:f96703e27268da58f0c53b457c3c4d749b165f776b5f503d4e9117c50f44ee7e