Pith. sign in

Paper Citation Record · LEDGER

ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2408.04682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04682 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:34.819687Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0ad1fd50-ca3d-45cd-9304-6f5c1a5f1f4f · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:35:51.215921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:f7d4d950f5d55303bb0dbe382c73236428153b56b2198bdc180d24cc0168d0d4

Observation 9602fd78-0b64-4afa-9e40-bb39925730e3 · inbound

Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions cites this paper.

Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:02:40.376710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T09:02:40.294491Z digest=sha256:e649d542202afd2f47342ca3f2e814bd6fb0fef74ab6a9c94e1b14e7de72bf3d

Observation dbd88b35-1cd6-4f4d-a67a-46f92a7aa064 · inbound

Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services cites this paper.

Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:34.819687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:34.819687Z digest=sha256:f2a43bb9de65563047c6f8e5e63b55e05024cd70e8a4cd91df7020a56df65c81

Observation 55ff60d1-758d-4466-8142-ea7b84b5ab8c · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:01.284289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:01.284289Z digest=sha256:f5fa5b42f3bb8d0349bd530ceaff37dfc506c0429770f94c7a0547b773e9ab19

Observation cd43e93d-4be5-4004-9d58-f5834c6c6f73 · inbound

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions cites this paper.

CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:38.571897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:36:38.571897Z digest=sha256:99db99d76d11f5ba99c656ef5d97a6eafc32dcd730d4e372026ef2f50c88070e

Observation f533e183-3092-407c-839d-ffecd6c4733d · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:52:17.432112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:59179c441f8e4528853a14f7d6cce29c4bbbb821d7bd373ba7f6f0aa2827563f

Observation 546c32d9-a530-4dc0-9c0d-1bde942dbe87 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.628526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.628526Z digest=sha256:58ac80406b309a61eb493ab2ecebccdde82e29bb44cd0c4bec31dec7f521ed48

Observation 90fb7440-f45f-4f79-b44e-d8638061cfa0 · inbound

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities cites this paper.

Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:54.497019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:54.497019Z digest=sha256:0bdc1faf78a8b5476b50f3c68bd9f982ff5591ba36afcc62e15a3f08944cfc12

Observation 4e5ac1d4-53b1-4a8f-881b-74a20ced8816 · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:47.975577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:47.975577Z digest=sha256:05ad08c6d221cd7fcbd01b400ae57a4ee0cd1901612e46e9233b6a746a93e58f

Observation 9c5b7e91-0ef3-44f7-ba03-612cb3886353 · inbound

Apple Intelligence Foundation Language Models: Tech Report 2025 cites this paper.

Apple Intelligence Foundation Language Models: Tech Report 2025 ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:26:58.622125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:26:58.622125Z digest=sha256:919e94d1015eae83733764cff639c341cf26d396d5ed49aa489123dca7f720f3

Observation 2065dfca-4776-4ec4-b6fd-34ad20d7434b · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:57.418255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:57.418255Z digest=sha256:13c2fd0da0e687cccf93f3d8659b4b10e5f96cb7c2a1da14a6ba322515e9e92e

Observation 17b85ac1-5fe6-4332-a284-bd791f769b3d · inbound

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution cites this paper.

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 4976

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:30.254612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:30.254612Z digest=sha256:e36bf10089198a5ca9c01f4e51d179708ab19a9200b3417727d39bf923278a19

Observation 8985ac9c-af6b-44df-94c6-5f026bf80eae · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.486170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.486170Z digest=sha256:98c6b13838621bea9b9e9e8799edeac48298616fb132c347569c3ae53226d673

Observation 384f4a7a-5d99-4c71-ad96-646df71eaa06 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.733725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:6b839f141c51515f9db6c7dca02bf7bda5266929cd2769c1391b6ec5dfbff4a2

Observation 32adb8cd-f031-49ee-82e3-dc87d140ae16 · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.408309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.408309Z digest=sha256:fe7e9bdd8ef6e7ebe5004411dbab72a3ddfd3e7c75fbddd6942970b68a1d957d

Observation 96b63d17-0dd0-4f62-96de-03c8dc107cf8 · inbound

PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback cites this paper.

PyTOD: Programmable Task-Oriented Dialogue with Execution Feedback ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:57:28.102393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:57:28.102393Z digest=sha256:c43e166bfc4c9b35ed7c4bb100b041d9f3988afcfbf61c7881ea011853e83f3b

Observation dcde5d88-8512-4bb9-bea9-bbb2d01f6287 · inbound

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench cites this paper.

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:00.685400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:00.685400Z digest=sha256:15917f16d5a0c7ce373221c6dd0110f27e5625a1e111c62a2e7a47e302667feb

Observation 2acc08f8-df26-4463-b47c-75c18ec09782 · inbound

COMPASS: Benchmarking Constrained Optimization in LLM Agents cites this paper.

COMPASS: Benchmarking Constrained Optimization in LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:46:07.880209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:45:11.334594Z digest=sha256:9289fe6a046e70e5cd2b01b285bc029a59700c3756393aee2f336da599f9a515

Observation cdcb1055-214e-49e5-9335-e7073b300fa1 · inbound

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory cites this paper.

Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T18:03:37.341452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:03:37.341452Z digest=sha256:08e37c98fccf08273cc8ecb47d48f7badd2e2b8c9160d3b4eb961808476d70a8

Observation cbc00d68-b687-49dd-93f7-3e8128ef1c4f · inbound

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents cites this paper.

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:31.274272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:31.274272Z digest=sha256:3bb94c3581e469ab62d565531a5687c46686685642199e9be96a93dc564d45bd

Observation 7671de91-1088-487d-ae66-8bcdfc4e3373 · inbound

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls cites this paper.

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:46:33.312636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:46:33.312636Z digest=sha256:ce2b838c479ca21f7fd01f86951b5ea7c710337d2581f0e002f7faeb13f0c606

Observation 1e99cad8-ad5f-4498-af27-6b33117f8164 · inbound

Mind the Sim2Real Gap in User Simulation for Agentic Tasks cites this paper.

Mind the Sim2Real Gap in User Simulation for Agentic Tasks ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:29.394088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:51:29.394088Z digest=sha256:fee3c58f789086eda48ba64855a60a129c7f6545e80d3cb3cd190bad9bf54513

Observation 303b2907-089e-4a5c-b292-3012610353e9 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.276299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:7db2f19d0302e4feaccab8e2b3bac9b8c8b7eb26afeee162609850a3dacc1eba

Observation f512d45f-3c96-4cc1-8cb9-a12791cb8bc4 · inbound

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills cites this paper.

Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T09:29:51.204678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:29:51.204678Z digest=sha256:397b4653e6589c0af952214a16c3716c95f9f712aa305266ac38d98f921df057

Observation b64fd55d-3560-4fae-b973-1d1382b4c779 · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:02.948205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:23:26.128263Z digest=sha256:2639bb7a9075c3f5098c876929ce29846adbb4ea964dc2b69df4976742bbdf30

Observation cb9d7391-7d34-42db-a282-a7f59a02b766 · inbound

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment cites this paper.

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T21:35:10.173253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:35:10.173253Z digest=sha256:5ffbb7cd7b330451a5d5664ed3932ba997afe20aa31b6ae993e1db41103f31b3

Observation a9d309d2-bd28-4d6d-b504-e79a0a1ae000 · inbound

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation cites this paper.

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:25.973584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T14:01:23.894966Z digest=sha256:3bc929fef2ee01e9487e5c249a0fd72afd3eea5012d08b38b2931d7195c4d873

Observation 78643d9a-d7a8-4f1b-9013-46192bceb7f3 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.975780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:138f6c54696692af92a381754be9e0112749b9c2510390bd744a08f329234994

Observation bf0365b3-03c4-4603-84ae-bfa4a8c932e4 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:43.057004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:bf1580d382e797cc94fb71a128ce2d13966a74835338849f38e639a88bf4b0ff

Observation 77ad442d-6a89-4507-8670-4154703a5604 · inbound

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions cites this paper.

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:13:44.721016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T17:10:02.936605Z digest=sha256:8b3e6247841efb94bc68f094a7ba63a58811a5f5fead334a9aa8041ed07737d6

Observation 5df495e5-804e-45c2-8469-186133156986 · inbound

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents cites this paper.

Designing for Doubt: The Case for Informed Abstention in Autonomous Agents ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:46:23.684356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T14:03:02.703499Z digest=sha256:8ff058cd6606c4ecb251794c57ab9d61f379d5aa469b7b9a3b1e27d294c80ed4

Observation 5fa44723-d6a5-4bca-a7d3-fb6e66293b35 · inbound

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose cites this paper.

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:59.725697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:37:50.570757Z digest=sha256:9ebf5812d31dbc68b7df77b94deaf41410a2ea786c4967efa6e7b952bd01d868

Observation 91ffc82a-47cf-4552-9eec-f94d765e5ebb · inbound

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows cites this paper.

Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:55.533094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T20:18:07.134598Z digest=sha256:4e82380e130bf68027b852c0813b2a3eba6747a926c692a6d6ef7e998e5ad889

Observation 29be73f0-dce0-478d-bf47-45b0d02832d4 · inbound

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use cites this paper.

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T07:36:53.348031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:36:53.348031Z digest=sha256:194692c8d58c44372bf75241eb04f153f11e4ae646c5ecf85af4689768cdbc79

Observation 63af01f3-1fd4-4326-a9cd-47ddd31efa8f · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:06.175548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:06.175548Z digest=sha256:92cea9c14f6bb14b6ec943849d2621f223f2ad7a43b6cdf35bb0d41e12511871

Observation 5eea9802-d3ed-463f-8f49-be91be9f9510 · inbound

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools cites this paper.

Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:39.396840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:39.396840Z digest=sha256:8918e7077e0531c69ae5d69c45f2a4f9b438fe035d113eefaf09672d8a93e22a