Pith. sign in

Paper Citation Record · LEDGER

MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2310.03128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.03128 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:11:11.989677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:18:59.715072Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bfdcb510-4117-4c31-bdb4-3897e0a75dd1 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.521093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:10546063aca151e34763342bd85a2339c8651e79c568ebd62dde00891ca3011d

Observation 15ec5b6d-4d56-46f8-a061-989c9f6bc2bc · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.181459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:e245df692847854e99a54a7977d7b558b459fc024e43f224f02aaf75921f5c8b

Observation 55495220-d2ad-4deb-8d59-2736b19284fe · inbound

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains cites this paper.

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:19:00.864649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T03:19:00.831153Z digest=sha256:b2b0705b9391f9c111872a0f69770dbf6c268621b7cae149d818d322399afa85

Observation 34424258-f287-4c29-95a8-71d578960cf5 · inbound

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond cites this paper.

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T20:11:11.989677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:11:11.989677Z digest=sha256:a68a1d64f626f587b0750f7899d5daa75f405df7de2089bfc42f09d68724030a

Observation c18c7bb9-1a8c-4176-9085-053352f1ff88 · inbound

Prompt Injection Attack to Tool Selection in LLM Agents cites this paper.

Prompt Injection Attack to Tool Selection in LLM Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:08:28.977781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T17:08:28.933831Z digest=sha256:c3aab07cf3e110c3b00791ac9c96fabc30402b1a6c7831d02cfe55dc26aab6a8

Observation 9e16af64-e263-45dc-b0d4-27ecf863a911 · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.470715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:94f1a8c78748ebd107a98e2d21b765482a419a83404d4450bdea9fa408c904b0

Observation b59b1997-5a74-45d7-8a59-a4d8437cf4cc · inbound

The Curious Language Model: Strategic Test-Time Information Acquisition cites this paper.

The Curious Language Model: Strategic Test-Time Information Acquisition MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:45.870083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:45.870083Z digest=sha256:0687f08b619d77501e2e40f31e7ee3520e228645ecc3e281ec5275b6052c63dc

Observation 20756517-6596-4bec-9837-d139ff5bdd0b · inbound

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues cites this paper.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.861022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.861022Z digest=sha256:2770018ac3f02c8bedbb39c7cbed7d55d346ac9b0810675d3e988c80e036fadc

Observation 6ac1df20-9361-4c43-82c9-f4898d7e2d0e · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:47.549550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:47.549550Z digest=sha256:5cdab61cfd4133913b56b4f4c3555d450778eeee8a8309775567d827d239761e

Observation 337c765f-ae53-452a-8937-39caf61300d2 · inbound

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models cites this paper.

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:43.754386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:43.754386Z digest=sha256:ab48aa6e6360601f95d3ca80370921ff6cab2b8b6d0a8a2c5b6bb3e4b6ed36fc

Observation 0b1ca4a9-c1c7-4390-8c31-1e095790d3e2 · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:57.024887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:57.024887Z digest=sha256:edb8837f6bf588b87d62097a16dab7f9a0c8021168796eafb47a324fe8754571

Observation 22550fcd-f160-4d2f-bbd9-269ad5bc33d1 · inbound

Towards Compute-Optimal Many-Shot In-Context Learning cites this paper.

Towards Compute-Optimal Many-Shot In-Context Learning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:38.656905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:38.656905Z digest=sha256:7f0e06d8c065ceeb3a174c2d9858d925f52c7e029d9a269debcc822b2777a916

Observation 51babeb4-227a-4516-95b0-2f094cd24194 · inbound

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis cites this paper.

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:40:51.463887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:37:11.945418Z digest=sha256:b899a04554c0b3d33b634acaddea1c68419d9976840c5bfb4f12c2142caf3f95

Observation 1576480d-e593-4139-8f1a-353f60d92f52 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.609937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.609937Z digest=sha256:9d9ae8d170b184a71cf07e9170a53327981385d7331bf89c635097f5325f7d35

Observation 67a0b2f9-4580-4f30-9e34-68085d9089bd · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.335428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.335428Z digest=sha256:ff3c4a7d6ba91b360560ee708c3d73bd4ead937e2209b2102d8f754afadc5256

Observation 915671c3-e5cc-4ec2-9679-5d17a1a45d50 · inbound

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever cites this paper.

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:16.475327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:46:16.475327Z digest=sha256:58721d61dc77e5c8d93b61b453a48b4d0fac73f69826a273dc70785db862be9a

Observation a03ba34e-3ea1-4c44-a50e-849e2df3cda7 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.792208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:02:55.529271Z digest=sha256:ed2584174882328bf9a486f83ee4e1c270465493e36643e15c8cb12603acc96f

Observation 0e9d9af4-cf15-48a8-824e-b99be0ce07c2 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:12.592641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:12.592641Z digest=sha256:44bd448fb31c259c93f753bc2320917499041eaff01e43029e9b4ef94387ab21

Observation 523d99e1-82e0-4b96-84b8-3501b8c110a9 · inbound

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception cites this paper.

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:50:52.539483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:46:03.228969Z digest=sha256:7c833a2dd2c4f96a37ce0c6d3328c270abc69c5d56e026cfe6fc52c2f07b4462

Observation 0700dca1-1a6d-46f8-8cd4-7a0dabd4525f · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 261

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:18:20.250588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:ee99743f27c61994fb8ce87dadeea8cb76edec399d3bfdecf14498139c558cf1

Observation ad97d443-4774-437f-b967-c37f48d6ce05 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:36.280577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:36.280577Z digest=sha256:f65645b9bb9555c784e1f1064c91539878b539734ae4e96d909d22cfcf8e71b4

Observation 59ec2ede-71a8-4033-86a9-01bfcebf3c9f · inbound

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls cites this paper.

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:46:33.268448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:46:33.268448Z digest=sha256:f9eb093ae0f97ef704edea15d8defe652840a5c2ecf0d3a02414a287fdaff8e7

Observation 4b853e52-1d2d-404e-83a7-6d6775680c58 · inbound

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents cites this paper.

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T19:59:59.030585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T19:59:59.030585Z digest=sha256:8198bfe5ef4afd1f2b18bcbc7958c534a34c67312bde0121d931b3b7554f188d

Observation 4f2bdf27-338b-487a-8c00-d32a047c8fd4 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:10.333342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:9e5f8808053f0771c97f890310559692b9fdc087ef4f237d7ab176df9eea8c9b

Observation 4a5476e6-df85-4c89-9155-3d38d16c87f8 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:33.837524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T19:04:53.639217Z digest=sha256:d6e4edb2c1f85ad913dbf16af0b84e9ea6f72a8c533114ba7db43ca4539fd760

Observation d00cbb03-51f7-491b-9c00-d4a541116888 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:12.424114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T00:25:46.507401Z digest=sha256:3f407f0bfe4c27d2ab6d86af8c54606cb73d8648c91d8d4bd6aaa6fc50901612

Observation f87b296a-6676-48f1-9d59-9f30edbfe7a2 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T05:24:44.710671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:24:44.710671Z digest=sha256:95c19ec65592eaf5575a45f2070c7ca4a4bd5cc844b2bdcf01c35e41d4dd08d8

Observation 7709fb7c-2ad0-4ef3-886e-0853e6383a01 · inbound

Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI cites this paper.

Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:37.672983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T19:06:36.952203Z digest=sha256:596ccf9da4389460c601a1ed648b95e38c4cc5a93ba9ed0cc53f49d531a5a57a

Observation 684b6e09-e7cb-4db7-9017-2bb70f976afd · inbound

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation cites this paper.

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:42.967662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:19:39.392516Z digest=sha256:1bb86831c8c65e58b9778b7acd703512404af2a658853538ceeca65a0e78dffb

Observation a86b0777-73cc-4fea-906b-4dfb1a7f260f · inbound

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning cites this paper.

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:51.785672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:29:47.384341Z digest=sha256:7b0eca6ca58849c15f4a16ea9e4e681259935dc9a7a3c91f7e7b2022c8fb6f90

Observation b053461d-5763-4d1d-8f70-7b549a050ee9 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:35:04.476298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:31:54.252816Z digest=sha256:30dd37ab52f1cc836d01d99c92efd3837e7c12f6033fbfe455367db3cb26e7f5

Observation 3e545d30-a3b9-44e7-b915-75cc50001943 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:53:43.552641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T20:52:20.975459Z digest=sha256:f82cd0243a021bcd34096bd5fc0ca129176f988d0098843e4fa3b233450ac46c

Observation d473db4c-3e82-4cd3-9767-bcb8063c5b44 · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:57.998107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T03:07:38.232966Z digest=sha256:c06eb7b6fa65e1c116ce0d944a91ff92c253f6f0dbbd758fa4f587e2990691c8

Observation 749d96c9-e863-41ee-9a88-6548c8b301dd · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 165

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:52:39.943407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T16:51:13.491389Z digest=sha256:cd09a59b00c7dad983a60ba8e8206256a661d776e04be722c833ac127b35fb5f

Observation b36e908b-d4c6-4335-a441-73f48c1b2f6c · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.401757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:8dea9fe14b2036a4d08c5aa1f0cd706d548582af20b4910382340057c3c92dfb

Observation b7301589-2913-4ce2-9e68-39c437d74e07 · inbound

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning cites this paper.

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.732455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T00:25:59.532059Z digest=sha256:c00e1367f09ed797733b317bf89ebb767a1b6c5b3f527fcc07fedf152ab9c9cc

Observation df94b234-2959-4db7-a73e-2469787000e9 · inbound

Capability Self-Assessment: Teaching LLMs to Know Their Limits cites this paper.

Capability Self-Assessment: Teaching LLMs to Know Their Limits MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.643921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:25:50.579196Z digest=sha256:9fd694cb6b01c16b569ef44a891cdecc108680304a9dee0b49cdc235b84a6712

Observation 0cbf862b-fdcf-46a9-adff-cdc6cc2f28a2 · inbound

NTILC: Neural Tool Invocation via Learned Compression cites this paper.

NTILC: Neural Tool Invocation via Learned Compression MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:04.773446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T00:05:36.600389Z digest=sha256:c2c6f50ebaa250a8d18771cb47a88ba0f13de49c2b9444f179fc0562be41be44

Observation 7f6a9ba2-b68d-4f41-8fe9-02eb8f016c6e · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.155289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:34:23.382705Z digest=sha256:a4399f2abc46fef54b06af7258b0a9ba13fc68d3abaf3b7892f3f1945ff79bd4

Observation c152ec48-6f8b-4634-82ed-6532c4cc7091 · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T14:59:04.152474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:59:04.152474Z digest=sha256:bb6bcf69888e93c93821763035abf455b059d868a60a82a0afdd3e2487e3bbbf

Observation b490de2a-c3a7-4ef9-8b45-7932e70decb2 · inbound

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose cites this paper.

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:18:59.716659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:37:50.570757Z digest=sha256:11ef470058ca4b2237d5adeb5231533a447d12e2f5edd7a0ff894663f8391863

Observation cf78261e-76a9-4938-8b66-59b2e68a263f · inbound

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing cites this paper.

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T01:35:09.657698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:35:09.657698Z digest=sha256:cdac782ecbf5ee4ece9300eb82095713651e6b8b7e25b3039a0d6978183e0e88

Observation f5bd2c1c-a64a-4a8f-bc81-35180af2a342 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:29.777251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:29.777251Z digest=sha256:c6368efcbcd92e973bcfb0694d13a2fb0c5e65d7811c74c9bb27418dc21bad96