Pith. sign in

Paper Citation Record · LEDGER

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

As of 22 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 0 inbound Pith citation observations for arXiv:2608.12764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12764 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:51:43.368701Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved32
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 548ae55b-55e2-41eb-954d-b76e5e591933 · outbound

This paper cites Glm-4.6.https://docs.z.ai/guides/llm/glm-4.6, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Glm-4.6.https://docs.z.ai/guides/llm/glm-4.6, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:44.050035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.161572Z digest=sha256:8a022155417123a9e6ffab08dd22c37d470b1cb97a027659d11a4508a220dc07

Observation c1e85063-e048-44c7-b5a2-a1c149ddbf0a · outbound

This paper cites Glm-5.1.https://docs.z.ai/guides/llm/glm-5.1, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Glm-5.1.https://docs.z.ai/guides/llm/glm-5.1, 2026

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:44.042037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.164991Z digest=sha256:b78a4ed2fb192750a43d11447cb1680b05c5360027f012660f6351022a9ffe80

Observation 9ed6c33b-fb1f-40be-93b9-0ba7e7fc2e4f · outbound

This paper cites Introducing Claude 4 — anthropic.com.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing Claude 4 — anthropic.com

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:44.033676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.167930Z digest=sha256:b454589aeb11c5db4a5590a6515ff4e14a12016ae3c7ce8f9b759c73ff352a2a

Observation dd43c5cd-cd67-4d67-8ec2-b4b3ee26c669 · outbound

This paper cites Introducing claude opus 4.7.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing claude opus 4.7

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:44.025747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.171598Z digest=sha256:a60ead988ea98b931e4e95dc2d62a0af9e53a64ea03c3da1653601a2f44109da

Observation 61ac298d-1de2-46b8-80eb-5bac227bbd23 · outbound

This paper cites Is gpt-oss good? a comprehensive evaluation of openai’s latest open source models, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Is gpt-oss good? a comprehensive evaluation of openai’s latest open source models, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:44.017417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.174610Z digest=sha256:0d12cee906a5f78f46b34f595f3d23fbdaf17ce1830cd9421302f3213a5a2dfa

Observation c25e1169-fde7-4057-85bb-fbf7c94c7f7d · outbound

This paper cites SFT memorizes, RL generalizes.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents SFT memorizes, RL generalizes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:44.008434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.177631Z digest=sha256:1c2bf8c236193d500612538a91c6cc152cb0e2b0831714e33a8315e714d4943b

Observation 418d6d98-af6d-414e-bdaa-a607bb5b60bd · outbound

This paper cites Gemini 3.1 pro: Best for complex tasks and bringing creative concepts to life.https://deepmind.google/models/gemini/pro/, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Gemini 3.1 pro: Best for complex tasks and bringing creative concepts to life.https://deepmind.google/models/gemini/pro/, 2026

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.999583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.180525Z digest=sha256:f47ecf152f4402b9d316142b23ec9dc96ecde8575f5df81f3b20f7c831cc5ada

Observation ffdab16d-2070-416f-ba54-76c838ebfb19 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.183156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.183156Z digest=sha256:20f043c5da60aff17ec58feaa07ea0bc0fba063741be1772db55d9740ca32a6e

Observation 6345e824-41d7-403e-8924-25c31cd738c2 · outbound

This paper cites Deepseek-v3.2: Pushing the frontier of open large language models, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Deepseek-v3.2: Pushing the frontier of open large language models, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.985774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.185839Z digest=sha256:f885a93cb3f139b1d0c4afde0b1fa23beb57325fb9ba4fef00410ec20b726ed4

Observation 83d43f61-deb0-4545-aa73-af2292a24da3 · outbound

This paper cites Openthoughts: Data recipes for reasoning models, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Openthoughts: Data recipes for reasoning models, 2025

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.976138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.189463Z digest=sha256:93966517ab2925fed6e40596435f20c04d5b1e556e8c0b2a856e0ba77777b6cb

Observation ff42beca-82c0-4741-b8d7-724ef4aea5cd · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.965115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.191958Z digest=sha256:b0daf1c2f4e5f126c69b0bc238240cbe78abf7ec9eb3a6cb175843cbb29e3311

Observation 2904fb17-06de-41a6-a14e-4fa7bd5933af · outbound

This paper cites Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.194617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.194617Z digest=sha256:bb046b6d89cf5e4880ac044eba2f2c499538b65bcddbe4bbea75cbbd876d315e

Observation cf34c1f6-2774-42bd-9aa3-68e0632b0f3c · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Reinforcement Learning via Self-Distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.197480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.197480Z digest=sha256:6b5a6c37beb8bec5f3f94f67840cb7da242c8076032e21cc03084094040bec59

Observation bbcfe5a8-7c4d-4729-bd5c-0faf5bb04904 · outbound

This paper cites An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.200350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.200350Z digest=sha256:bed6db84e60f1229a0d53f2172f14c52efb80832eb01bf34079254be92aa09b1

Observation 97fbe675-2bfe-49eb-9b8a-d2e9faf3e251 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.203767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.203767Z digest=sha256:f128f4032049df359901ecf870f8dfcd65417c5faf44a5adf0059391fc965e27

Observation 92f2cbd0-5407-4c46-816c-332d28b43ff2 · outbound

This paper cites Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation, 2024.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.947758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.206665Z digest=sha256:961361ba215978dab5bafc5634e294ac994ff8efb61529d29382cfd3178bc995

Observation f05af0a4-9eb9-41e7-bfb8-5d88a1596e67 · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unifying group-relative and self-distillation policy optimization via sample routing, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.209595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.209595Z digest=sha256:175e330dfc3b75629d2876ec7b110f63207094058f60aa8190a3baa422d7b61d

Observation e4e0c641-23b8-482d-9d52-34c178b8b2c5 · outbound

This paper cites The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents The tool decathlon: Benchmarking language agents for diverse, realistic, and long-horizon task execution, 2026

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.930895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.213232Z digest=sha256:95fea15617f2c1d1a071bdaf6f472b07f907ee13667eb0819c77b294b3e10b99

Observation a140c354-1a2d-4ab6-b8f4-a0801f3f26f2 · outbound

This paper cites Websailor-v2: Bridging the chasm to proprietary agents via synthetic data and scalable reinforcement learning, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Websailor-v2: Bridging the chasm to proprietary agents via synthetic data and scalable reinforcement learning, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.922352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.216094Z digest=sha256:128b2e7388ea6e7e2bc9e512458a54bea95e1461f40ae0db60fbaa26c798d325

Observation 1b758483-8e39-456d-af58-47c56d82d287 · outbound

This paper cites Websailor: Navigating super-human reasoning for web agent, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Websailor: Navigating super-human reasoning for web agent, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.913604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.218917Z digest=sha256:1cd1543fbe2a4acdda27c5be7c15f4d9177b2384f9625425818768ca7c6d3c3c

Observation 8a65a022-0282-4aa2-a37a-067dcfa28e0f · outbound

This paper cites Chain-of-agents: End-to-end agent foundation models via multi-agent distillation and agentic rl, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Chain-of-agents: End-to-end agent foundation models via multi-agent distillation and agentic rl, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.904459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.221863Z digest=sha256:dd2445b7079be085d34254b367e8ce08f675c9a853a2f46ab5b98945e4352e77

Observation d7383d2c-d8df-4543-ad74-e1e76aa73a4a · outbound

This paper cites Webthinker: Empowering large reasoning models with deep research capability, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Webthinker: Empowering large reasoning models with deep research capability, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.894115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.224555Z digest=sha256:df6e49e9406cb9c5c39f2b466c2babebea18c6da8b3d5acd809593211f3df4a9

Observation ee6b7ef5-bd7e-4eaa-b120-b617099f344a · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.227221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.227221Z digest=sha256:1fa5bb4a03c9eb81267fdd824eb460d3fd3471ae936e6fd38e19b8e36ceb1fed

Observation d4bf8ac0-7843-4c0e-87d0-d91b7edf8f46 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Code-r1: Reproducing r1 for code with reliable rewards

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.230442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.230442Z digest=sha256:a5da7dc205c3a7cd695b96044761050ffe022f9b52fae87dd1a77cb7a6b87975

Observation d2e772a2-db15-4381-a0e3-384c9ba65cdd · outbound

This paper cites Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.233345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.233345Z digest=sha256:4a647c48dedb5c55545dc5d4b133510b47e5edd30f897cc3a68a767ef7d63422

Observation ddeb2804-c677-44c9-a553-71645a4bf174 · outbound

This paper cites Webexplorer: Explore and evolve for training long-horizon web agents, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Webexplorer: Explore and evolve for training long-horizon web agents, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.871231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.236115Z digest=sha256:2185ef186507df3d460072cbefab881367d804c60782632d2e9e18ca46939e97

Observation efd09cfe-6597-4652-a0fc-febcdad64339 · outbound

This paper cites On-policy distillation.Thinking Machines Lab: Connec- tionism, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents On-policy distillation.Thinking Machines Lab: Connec- tionism, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.238860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.238860Z digest=sha256:7df5deea6a961b77a7118322875ea34370fc83a8e229d24ddb2a3bb0fbb5e76d

Observation 67d044b7-5e1d-4cfe-80b8-677198465c45 · outbound

This paper cites Deepdive: Advancing deep search agents with knowledge graphs and multi-turn rl, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Deepdive: Advancing deep search agents with knowledge graphs and multi-turn rl, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.856509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.242571Z digest=sha256:58f1242d30f81cd3e5abc432c898ca5c4d1c2078a3a6b113b7e4578cbda05819

Observation 14420623-eb4a-42d0-b418-dee0e3931b24 · outbound

This paper cites Gaia: a benchmark for general ai assistants, 2023.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Gaia: a benchmark for general ai assistants, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.246253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.246253Z digest=sha256:0067bbb652f769d395ac9e009ddb902180b94580c2447ba0f18ded38742bc330

Observation 0e369d84-6b97-4c31-b204-86d92bc8f893 · outbound

This paper cites Minimax m2.7: Early echoes of self-evolution, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Minimax m2.7: Early echoes of self-evolution, 2026

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.842820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.249288Z digest=sha256:6bb10fd22092ef7991c65d98154b4dd761af6fcbadaee767764b85b4391a4502

Observation b4d8224f-499a-4c67-a16d-4c557279332b · outbound

This paper cites Gpt-5-nano.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Gpt-5-nano

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.834397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.252043Z digest=sha256:d781912f0ef5cca2e240979bc5359ec01b99f8b704596d874887afb7b2f1d593

Observation 1b80dfb7-199c-4a0d-8cb6-d9cb74fe7f51 · outbound

This paper cites Introducing gpt-oss.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing gpt-oss

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.826063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.254662Z digest=sha256:f45d105224e5ec439e40401fb70f58732bfa55359d203a41151a130a6e056300

Observation c7a3c941-fbae-4ded-854e-614078dc060d · outbound

This paper cites Introducing gpt -5.5.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Introducing gpt -5.5

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.816936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.257321Z digest=sha256:1451749a0c7316f8560ee20e794c87dc81bd181c8ca27f5153d5bee93a40fd64

Observation 57bd4ec2-cd31-4f68-b75c-2c7bf272849f · outbound

This paper cites Qwen3.6-Plus: Towards real world agents, April 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Qwen3.6-Plus: Towards real world agents, April 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.260011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.260011Z digest=sha256:f051d92f1bd730aeb4e753329bcff26ceaf83365a179f47e7b592e6c7bdb8518

Observation 6a18700d-d4b5-4446-b0bf-2a88035e295b · outbound

This paper cites Crisp: Compressed reasoning via iterative self-policy distillation, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Crisp: Compressed reasoning via iterative self-policy distillation, 2026

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.804533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.262642Z digest=sha256:3240418358579ac0beb678b04b29cd35e08978d90149979f09c861a1814ba2fc

Observation 4c5b4d03-3257-41e3-9e70-c77e4a31c06a · outbound

This paper cites Generalization in generation: A closer look at exposure bias.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Generalization in generation: A closer look at exposure bias

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.795590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.265300Z digest=sha256:ebd0da310b720b07605707927e8acd3ee1fc4194861a9e46aa625ba84516a944

Observation 4cf2727c-0a05-4416-914f-0597dda75701 · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.268137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.268137Z digest=sha256:b0912075b26dc78070967d6bd6b2ae550122478d56a6267eb1db04e7fafb4a35

Observation f54d5f89-7ecc-4ad0-bb98-e67927e0f891 · outbound

This paper cites Self-distillation enables continual learning, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Self-distillation enables continual learning, 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.270828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.270828Z digest=sha256:da77e1e5514c7eafbc54270c4ff64b0b5138aca67da1c3e31ff4fd223ef22c9a

Observation b310cec0-6933-4c75-9c8a-faa11b68b926 · outbound

This paper cites A survey of on-policy distillation for large language models, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents A survey of on-policy distillation for large language models, 2026

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.776760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.273666Z digest=sha256:bb46f10827deaae6aa765364817dd13db43908960a162d31a701587b25e75c24

Observation 83067a0b-77ad-4284-bc08-cb3c031d2576 · outbound

This paper cites Webshaper: Agentically data synthesizing via information-seeking formalization, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Webshaper: Agentically data synthesizing via information-seeking formalization, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.768299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.276194Z digest=sha256:6e8a9a61353ff6a35f58f027c53f279744170ad03940051cfcd4b1c65cae9eb2

Observation 9558d3d7-cf9a-4498-a7cc-94240f65c2e3 · outbound

This paper cites Kimi k2.6: Advancing open-source coding.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Kimi k2.6: Advancing open-source coding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.760182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.278666Z digest=sha256:730836ef5429840bcf8353895e09ba522f92fb852acf2786fdb5f2390281b371

Observation 3c517d4d-a00d-4422-8f88-bc637c8986c4 · outbound

This paper cites MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.281315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.281315Z digest=sha256:275b15687ba7a6ad329d5053bc0b7009a2c901aa1368e8c11b2f0b0c8e46694b

Observation 8a772cdf-7424-4a82-accf-872f59ca70a2 · outbound

This paper cites Browsecomp: A simple yet challenging benchmark for browsing agents, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Browsecomp: A simple yet challenging benchmark for browsing agents, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.284154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.284154Z digest=sha256:11ee70c9507bf01da242f63a079a75ac9c8ab6ef6edd4ba88d1f18fdece23ffa

Observation 4788aa76-868c-4c01-bf0f-2843fee84807 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.286776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.286776Z digest=sha256:d7339e8b36ddebbeb015bf6ca7b3882c20eaf47d5fb5b57855b71391ccd29104

Observation 29ff97a0-2086-4e2b-9d1e-c04652128481 · outbound

This paper cites Smartsearch: Process reward-guided query refinement for search agents, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Smartsearch: Process reward-guided query refinement for search agents, 2026

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.742516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.289651Z digest=sha256:1513ebbea1d72e7685cc369210e6ae0f3370f887150a19c6bdafb21512337df5

Observation 191dff8d-fe0f-433c-9749-f9147ab6ad48 · outbound

This paper cites Mirage or method? how model-task alignment induces divergent rl conclusions, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Mirage or method? how model-task alignment induces divergent rl conclusions, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.734668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.292212Z digest=sha256:0d2b38ef2baf1145867a32c4204a33f0f05ea26fad9fc6762336a4bb9b05a78c

Observation 8005269b-060e-43ba-82d3-6734b9123600 · outbound

This paper cites Recode: Updating code api knowledge with reinforcement learning, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Recode: Updating code api knowledge with reinforcement learning, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.726062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.294878Z digest=sha256:f7dac236c7d9aaf6aa3a92ff28f39b06d65e50a5184f0a5a47d617f11be3372e

Observation c6938df4-51c4-4d03-94bd-24fb7aaad3fc · outbound

This paper cites On the generalization of sft: A reinforcement learning perspective with reward rectification, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents On the generalization of sft: A reinforcement learning perspective with reward rectification, 2026

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.718023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.297524Z digest=sha256:6df5e0ee6a68eda6d8e380cb4e13c56f80eb1af63341916da1e7d408cf1b8d2a

Observation 4f3c1a6e-fe1b-4e02-9fb5-1422126c4526 · outbound

This paper cites Grok 4.1 fast and agent tools api.https://x.ai/news/grok-4-1-fast, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Grok 4.1 fast and agent tools api.https://x.ai/news/grok-4-1-fast, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.709829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.300110Z digest=sha256:73122827f9074682c8c4ce401510bbfda8fb960ab4c8c85e1ad5915f946b65ea

Observation 8f6fd0bb-f01d-4ffc-8a13-8bae0310a77d · outbound

This paper cites Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.302629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.302629Z digest=sha256:ebab9934632c05e22b6e1529ecf4fff4800424ea066dc64a97a71ae0437be828

Observation f3271436-d47e-4138-85c1-e3730eee63c4 · outbound

This paper cites Principle process reward for search agents, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Principle process reward for search agents, 2026

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.696941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.305487Z digest=sha256:fafa2e9c4e89bca978ba3d19b1d27bf87ecebcd1e3e9dbbef5194eb1d9904e5a

Observation 12da9215-0b9a-420e-a5d5-b714f08c0a3f · outbound

This paper cites Tip: Token importance in on-policy distillation, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Tip: Token importance in on-policy distillation, 2026

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.688894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.308274Z digest=sha256:b76b5e82e554d77a46658d3d54625ae809ccd164104d4cd79fcb2a65025321e4

Observation 42f2809f-408f-4a3b-9531-c6ea05356dee · outbound

This paper cites Qwen3 technical report, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Qwen3 technical report, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.680631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.311087Z digest=sha256:55827f17d37d29a1b5d28f2c68f7fa1ecd9883796d225a36392b1ffe52c0f469

Observation 312d3551-dfb5-4e9a-9e0d-4382c12354ac · outbound

This paper cites Self-distilled rlvr, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Self-distilled rlvr, 2026

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.313743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.313743Z digest=sha256:20c5a6b9e6c2411f137d5a3451db56726186dc1c432995f6e8c5017563e7688f

Observation 35d6e048-4992-48ca-a0d3-0e0bc682f0b5 · outbound

This paper cites Cohen, Ruslan Salakhut- dinov, and Christopher D.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Cohen, Ruslan Salakhut- dinov, and Christopher D

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.316248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.316248Z digest=sha256:b0f75225e3834d0c64a4c9600673670a2ea04cdc604d909b4645d92636616a36

Observation d742c681-1544-463c-a09f-331a18f4816b · outbound

This paper cites React: Synergizing reasoning and acting in language models, 2023.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents React: Synergizing reasoning and acting in language models, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.318984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.318984Z digest=sha256:231edcdea905bccda16e9d1f44fc4f68a5babfc9eea219bc1a6431421e835744

Observation 840f85e9-87c7-4670-86a4-15b0d5d55b88 · outbound

This paper cites On-policy context distillation for language models, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents On-policy context distillation for language models, 2026

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.321940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.321940Z digest=sha256:c3c41d56323a5f72fad1ba0ae2526c61aa7e5d945cc933f3df4f33a58c427084

Observation e179a6bd-ed91-47fb-8320-b9a32f1fb84b · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.324449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.324449Z digest=sha256:52de9d6f32546ca9fb317a026ec05a7b9412602dc8deafeace55ba04058ce9e8

Observation eba60e96-0f9f-486a-90bd-0fb221ce7f8a · outbound

This paper cites Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Simplerl-zoo: Investigating and taming zero reinforcement learning for open base models in the wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.649857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.326932Z digest=sha256:ab90bf20042380c236dd3a7e34da78e6cb38dccf197ecd7baf72ededf69a48ba

Observation 1847a17f-81e8-46b7-a374-781a1ebd4558 · outbound

This paper cites The landscape of agentic reinforcement learning for LLMs: A survey.Transactions on Machine Learning Research,.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents The landscape of agentic reinforcement learning for LLMs: A survey.Transactions on Machine Learning Research,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.640876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.329537Z digest=sha256:276b706ab65c1007aa5c5a78a33ba370ba013553f2d9bc74523fa8c2ab146996

Observation 625fdb2f-a8e4-489a-8087-d9f3ab320c85 · outbound

This paper cites Tool-r1: Sample-efficient reinforcement learning for agentic tool use, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Tool-r1: Sample-efficient reinforcement learning for agentic tool use, 2025

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.627760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.335666Z digest=sha256:e69b0fdeb2bb87b01d7149a196f9f6ff02870b0473a4d7bad267ad86f6d34f4c

Observation 8eb1cad7-5bfc-4ba1-bb2d-71eb96edfe97 · outbound

This paper cites Criticsearch: Fine-grained credit assignment for search agents via a retrospective critic, 2025.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Criticsearch: Fine-grained credit assignment for search agents via a retrospective critic, 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.618413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.338362Z digest=sha256:9485f57e429252c8e00147d29d25a7c57246b65f88650f5b825e1727ce3066bd

Observation 5569005a-0110-47dc-accd-32fd96bfa726 · outbound

This paper cites Self-distilled reasoner: On-policy self-distillation for large language models, 2026.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Self-distilled reasoner: On-policy self-distillation for large language models, 2026

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:51:43.340966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.340966Z digest=sha256:be97d101a4636046cb82075d5b5e07dfe74dc6ae1f70a894dac392224980f028

Observation 375f757c-d9e1-4c0b-b105-7a844f775214 · outbound

This paper cites aha moments.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents aha moments

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.603372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.343631Z digest=sha256:b7862f54c0e65f91454eb14436bad3dd3a2264ac0b99f97ec38fd95fadff2474

Observation 895914a5-9a82-4a58-901a-763b5af889fa · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.595159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.346793Z digest=sha256:d1ae20c3175f2a744977d15a4f6612d6212085ab14429d98f889363589ce5120

Observation 1a19588e-4b42-4fa6-9b8d-2366147fa7e7 · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.587466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.349454Z digest=sha256:9a9a589517ec2a1f83b4ffead4aea52bad5dc2d6cb5374a3ebc172eb70ab1f65

Observation b0742e58-0293-4282-beb3-d51d6bcd15dd · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.578910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.352242Z digest=sha256:7062dd2d9fcf461c8ae2d59de02cd3e21d6ff257855b5d7d61cb2cd4e4719b8e

Observation 45cf3b16-9885-4a08-98ce-3cbc662b9405 · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.570822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.354989Z digest=sha256:e448d88ea8e88ce88c9eafc53b59ca710eb13325a6c30ea550b37aa572095968

Observation 9bbcc85a-0281-456a-9068-466bffddca48 · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.562305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.357939Z digest=sha256:9da03f1d45ee195e6fdf6f6a04156b2f628cb5c8816b7809cd8ce48a3d095216

Observation 6636605a-6564-4d6c-86f4-90b27da15767 · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 71

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T23:51:43.553528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.360589Z digest=sha256:cabc296eebe78dba91365fddc9dfcf80dab0f13e6d94abe276e19b3da40f6015

Observation 21f1c732-01b6-48bf-a80d-1f7f50ab5d17 · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.545276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.363312Z digest=sha256:054046dcf5b713fe6e12d9672fea5da79f0a9b5a7360c16b8821732e19869b11

Observation 11882a9e-120b-4e7f-865e-3da02dbbca80 · outbound

This paper cites who is the COOP leader in Amarillo.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents who is the COOP leader in Amarillo

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:51:43.536539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.365890Z digest=sha256:e00a4f5a1fa570243cd2ffc60343727f9f9d29559120e4d2716015193e31f59d

Observation 00f96cd1-8200-47cd-83c8-1ec6ff9fe91b · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:51:43.527742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T23:51:43.368701Z digest=sha256:5402aec171ba5853704da3a41c8c1063fecfe67a132c500437a8e9f82ca016cf

Observation 3668f22e-e1fc-404e-806f-09de3c9ffead · outbound

This paper cites an unresolved cited work.

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents Unresolved cited work

Reference 2026

Resolution
parse uncertain
no resolver link, observed 2026-08-15T23:51:43.332744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:51:43.332744Z digest=sha256:ad2559024494008990568d998594321839c81efbca2d53a8f243f4a264a93114

Pith citing papers

No inbound Pith citation observations are available.