Pith. sign in

Paper Citation Record · LEDGER

StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2403.07714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.07714 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:47:55.636246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:06:19.681075Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 51ac8584-6412-46fe-a3b9-26e357b0b439 · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.636246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.636246Z digest=sha256:8d79e43978ca3f36fa68715a518cc3a65885e0a15b3590e35edc15ed4b6dac14

Observation 05cca2ad-bdce-4466-ba58-36b46f12aca8 · inbound

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists cites this paper.

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.965163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:41.965163Z digest=sha256:45eb330762662ecb9d56c7b64ad0f88e378b934560c79d70d2ffa02167b2ffab

Observation 84c72a2d-a0d9-4de3-813c-6a0277e8d042 · inbound

Self-Challenging Language Model Agents cites this paper.

Self-Challenging Language Model Agents StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.061630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.061630Z digest=sha256:c349f616985455c3864d8d2657aa7c81d61a1806c197a31465049d1009a358e6

Observation c70c1609-b90e-40cb-a10c-cca927199e16 · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:49:28.125588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:89242690ce5be799b7fc59f75ee4c48765c351ff020dceaa6159ba12b2f5a816

Observation a3d7c793-3654-490f-927a-e3eda35011e3 · inbound

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use cites this paper.

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:51.595736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:51.595736Z digest=sha256:9cfce5f4a909a58c94bb1bc7dc749d2c007a4030158cf2aad34d53173430566a

Observation 20ee8583-fa32-4e82-8269-ef8aeed25810 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:34.979167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:34.979167Z digest=sha256:8d3d02c9517e04417846b42391eaf5921b73ccc843b71d990fedfe43fbc55954

Observation d34f950c-1b42-48e7-b072-2db7111ea030 · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:30:45.649037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:8badbc9c9c4916f7867315058beb5002588ea83465206793b11b14389de5b1f2

Observation dde28d33-9925-48d9-a6c4-240baa247fac · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.391787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:ca3450d99c72ea1d664030c7a02310e9298fddd89513bbdefe665f690f9a5ae0

Observation 17d9d2fd-8810-4899-9066-72b2c9535641 · inbound

Efficient Multi-round LLM Inference over Disaggregated Serving cites this paper.

Efficient Multi-round LLM Inference over Disaggregated Serving StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T23:16:44.239465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:16:44.239465Z digest=sha256:eb65acea7d751a85b891161203cb3cf9f76cbf4e0bcfb1638bdb3e75d1698b3d

Observation 4aed8ea9-6fbe-410c-a594-fb98dc41faad · inbound

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use cites this paper.

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T05:54:53.374955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:54:53.374955Z digest=sha256:c330978c62975a27f83a7a1b2b36f84f9f901b517824dcd10d1e3b899e277a74

Observation b38ef717-2894-41d9-86f2-9d505d4e0d36 · inbound

ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution cites this paper.

ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:10:29.005883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:05:34.931798Z digest=sha256:dc19c4c597305cc04704aedae8871719681ec241af6f8616db6b948cc0b371e0

Observation 982b5d76-d92d-4e71-8dba-8e655f1ffd0d · inbound

CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness cites this paper.

CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:14.073372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T19:47:49.262632Z digest=sha256:903a3362a1c007d4fba3c71f218f1cf0309110ececbd654dbf7530f0f8b5566f

Observation 62b5da10-b4da-4a32-ad9c-10dc4035c4cc · inbound

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems cites this paper.

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:13:16.054259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:13:03.046209Z digest=sha256:2faa100199d511e46ec670c8a9d458b3b4ec7aee3ff5f6a2a096eb669551db10

Observation dd6a516c-0861-4011-9fd6-3a69032c13e5 · inbound

CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation cites this paper.

CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:19.683156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:45:58.990464Z digest=sha256:8665579747e54bc8fa43f163e350e31bd92f68e2fb036b4c394b08253879c960

Observation 13b04911-4930-4a14-b996-e942cbefe43b · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:3f5fcb57556d3961fe6e8c9e66758ac232c755c267105d6cd23d1872805b16c6

Observation 4fa97a3e-0b4d-4c0a-a1c1-cddde0c8a8ac · inbound

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis cites this paper.

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:46:11.564751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:46:11.564751Z digest=sha256:de4e9538314473b8e3958929510e71418a01311f88a4d4883597eae5ea07a2cf

Observation 87accfaa-1c5d-4813-ba9d-c775bfe6412d · inbound

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents cites this paper.

Execution-First Synthetic Tool-Use Trace Generation for LLM Agents StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T12:17:54.179264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:17:54.179264Z digest=sha256:acfeea02b4b97be8b42b1d5c7b3aadae35c3c7f281344017b9d557bb1f0830bc