Pith. sign in

Paper Citation Record · LEDGER

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 6 inbound Pith citation observations for arXiv:2507.01489.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01489 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:55:09.996272Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:55:55.496339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:53:28.611729Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2980a235-2774-4187-bc1b-f97f589fd2ef · outbound

This paper cites Brown, B.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Brown, B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:55:10.276010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T20:55:08.809952Z digest=sha256:1b5b9f113726ada67fb06f45b4113355ac3ef31b138045cfdbd56789152e1cbd

Observation 741dd78a-e701-4532-96fa-7b09cd18633c · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.044949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.044949Z digest=sha256:02f9aef934961366f78e4d45fbea356f80e593e9ace45e8094b1798503e7d898

Observation b56633a1-1bb9-4945-9047-497a5ffb5378 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.251681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.251681Z digest=sha256:50f7f660b6a7b3cea0117cedc0dde3b07582fdefab01c04b75e43babca1ef827

Observation cd883f29-6f39-45e4-a322-8bb6646e2d57 · outbound

This paper cites Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.340912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.340912Z digest=sha256:601364f9a3463549e7a309b55d2084ae6e1f9b976ba310b01e09a4cad7892c39

Observation 579394ad-e4ce-4fc6-8220-841b831c109b · outbound

This paper cites Measuring and Narrowing the Compositionality Gap in Language Models.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Measuring and Narrowing the Compositionality Gap in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.438805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.438805Z digest=sha256:4818162875f824fb8720e14756831549bd09e975fab213e7c6984abb78f85eb0

Observation e4058c2a-53ed-4ea3-a043-9ebed6af1788 · outbound

This paper cites Qwen2.5 Technical Report.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Qwen2.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.563729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.563729Z digest=sha256:780535b182a28b0a797cbc8555ff047fafe4fd539edd12d65b160c9d8a2af155

Observation 65766e1d-9f9c-413a-b2a4-8a57bdb89f94 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.825758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.825758Z digest=sha256:7a82129db611c2425162418357b781a66aac72c14f79b2c2165275d9121a5648

Observation 5a78b56a-609d-4642-bbdd-b527cbff4f49 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.879680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.879680Z digest=sha256:963f17d06d4790269cf44c915905ea27b69f6c5b3cc0d73c877ee123aa512916

Observation cc12c887-380f-4ca6-a429-67fadbedd70f · outbound

This paper cites OpenResearcher: Unleashing AI for Accelerated Scientific Research.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning OpenResearcher: Unleashing AI for Accelerated Scientific Research

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.948075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.948075Z digest=sha256:7a209014ef2bdaeb6d215d58750e288a517c32f78489434e283947a0a995975d

Observation 3779de78-c0be-4671-ac6a-69f58a4549d4 · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.996272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.996272Z digest=sha256:cb9e427d49a5eb7265175b5aa055c1b5b2121bf13ac2ccc25f14a57f99e2cba9

Observation bb878f04-f70a-467a-a495-7e8d242f156b · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.108495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.108495Z digest=sha256:4543bb8e7ae1c973ae02fb23510f478b1f4b65b1148c96783fbcbc890432a7c0

Observation e7748746-bc8d-4743-82b9-64a1f7610247 · outbound

This paper cites MuSiQue: Multihop Questions via Single-hop Question Composition.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning MuSiQue: Multihop Questions via Single-hop Question Composition

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.770719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.770719Z digest=sha256:2d85117ada62baea6ffb6d7d5dd6e793a2a233da489d60fdaf9e23807c4b2c11

Observation cc514193-ee30-4d29-9e02-81507033a662 · outbound

This paper cites GPT-4o System Card.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning GPT-4o System Card

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.165242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.165242Z digest=sha256:9b7ae7ef9fcd801741a95260a6f8da0a6623d7e1603222a2c728ea9f5449f685

Observation 44c22c11-5e34-44c1-a8b0-410e413b38c8 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:09.695260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:09.695260Z digest=sha256:7b89086efabe680c3e1b2d28e35abf571ba9db8b42746a5e29b9e68868beb996

Observation 4b64dbb8-5014-4415-9e30-908677aa8367 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:08.936392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:08.936392Z digest=sha256:74a93edecec5f84701cf0bf0038a6a59858742ecc6c33e29004ee33c54c54080

Pith citing papers

Observation 88b47baf-aff7-4e85-bc8e-def515eb4e0b · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:15.635577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:4d4fd0880d63ab0d46d7e4dec672dacf3e2ff381d70d091689113040d7f07535

Observation 56960e50-3558-4410-8bca-ed3bd3392f08 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.597990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:a2d14755533fae995e49a15a38402f62506566f75bb81800d5a8d43f306a937d

Observation 91881407-ef2d-4594-a886-d4bd1467fc30 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:48.707183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:a2fe5eda99898942efc1ba4e6607da05573107bf44656c512d6308bc08f47ff7

Observation 2d126fbc-da39-44de-b05f-8ef1dc014e1f · inbound

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents cites this paper.

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:24:40.531043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T06:21:56.971204Z digest=sha256:de943c7ef6d753c01836b562ae6f9ddb9d83ff3b841c988d92cf03ad17c52a68

Observation 3ae45dc3-c1ce-487e-9d3b-aa30ec874642 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.613407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:b598eb41097807d44dd275d6fde2fb8432f5f305dd5595808bdf01bff52719dd

Observation 4efbff9a-c058-4e2b-9631-ae4d821c1538 · inbound

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems cites this paper.

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:55.496339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:55:55.496339Z digest=sha256:33bcd08fd3cbfc60966e0dbea2b5e679f26803471946f796830b4b507aa556ec