Pith. sign in

Paper Citation Record · LEDGER

MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2310.03128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.03128 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:11:11.989677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:18:59.715072Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bfdcb510-4117-4c31-bdb4-3897e0a75dd1 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.521093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:0a1c7fb03d024a7f6a037c22af8a802a3db1e802b89d0316cf8e0f71c35d247d

Observation 15ec5b6d-4d56-46f8-a061-989c9f6bc2bc · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:43:11.181459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:fb805907e14f3c7176da83dc3dcf6ee7bfa47d25df6d19590cd9290b10380d39

Observation 55495220-d2ad-4deb-8d59-2736b19284fe · inbound

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains cites this paper.

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:19:00.864649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T03:19:00.831153Z digest=sha256:7059313765e8a90fbe65c58b3b59830586accd4300a2c9fa1c304f757e14db33

Observation 34424258-f287-4c29-95a8-71d578960cf5 · inbound

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond cites this paper.

Artificial Intelligence in Spectroscopy: Advancing Chemistry from Prediction to Generation and Beyond MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T20:11:11.989677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:11:11.989677Z digest=sha256:0fe0156dbb4434691e09427aa2b8a9ac299b8474c57575f6d14657d57d86bd64

Observation c18c7bb9-1a8c-4176-9085-053352f1ff88 · inbound

Prompt Injection Attack to Tool Selection in LLM Agents cites this paper.

Prompt Injection Attack to Tool Selection in LLM Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:08:28.977781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T17:08:28.933831Z digest=sha256:0c6f87c86a47354f37286bebb05408704e79a400dd8f9e4e7df63f85adb48225

Observation 9e16af64-e263-45dc-b0d4-27ecf863a911 · inbound

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment cites this paper.

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:52:17.470715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T07:52:17.174347Z digest=sha256:6cdb5c1c4fb9460aa00d4632e9e0f8d938d2eaf8987b31691a70858f9346e1d9

Observation b59b1997-5a74-45d7-8a59-a4d8437cf4cc · inbound

The Curious Language Model: Strategic Test-Time Information Acquisition cites this paper.

The Curious Language Model: Strategic Test-Time Information Acquisition MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:45.870083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:45.870083Z digest=sha256:cf78bbd55a6652100d769c6736f965f5fb28ad0119b7abf2cd9ea39f5e39a5fd

Observation 20756517-6596-4bec-9837-d139ff5bdd0b · inbound

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues cites this paper.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:38.861022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:38.861022Z digest=sha256:8e2df30eba01b4b1528f205ca6007eec6df9b0393ca6351a2d329bf1dc19b9bb

Observation 6ac1df20-9361-4c43-82c9-f4898d7e2d0e · inbound

Teaching a Language Model to Speak the Language of Tools cites this paper.

Teaching a Language Model to Speak the Language of Tools MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:47.549550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:47.549550Z digest=sha256:5cdab61cfd4133913b56b4f4c3555d450778eeee8a8309775567d827d239761e

Observation 337c765f-ae53-452a-8937-39caf61300d2 · inbound

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models cites this paper.

MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:22:43.754386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:22:43.754386Z digest=sha256:05a575aa6d19058ad763e70598754100171690697667240949241835a9ff7ccb

Observation 0b1ca4a9-c1c7-4390-8c31-1e095790d3e2 · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:57.024887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:57.024887Z digest=sha256:398751107bd77900cfdf814b070021e675698f867d240ea10d720ad830091193

Observation 22550fcd-f160-4d2f-bbd9-269ad5bc33d1 · inbound

Towards Compute-Optimal Many-Shot In-Context Learning cites this paper.

Towards Compute-Optimal Many-Shot In-Context Learning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:38.656905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:38.656905Z digest=sha256:8253e7f22bafce02e5317131f0ac193b2e9d78f3419f6ca822a42ddf80fed532

Observation 51babeb4-227a-4516-95b0-2f094cd24194 · inbound

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis cites this paper.

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:40:51.463887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:37:11.945418Z digest=sha256:d79d309f6a2912cb9f916f9af5ba3b5a24278159f8d2cf0896c11bce67ce6a92

Observation 1576480d-e593-4139-8f1a-353f60d92f52 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.609937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.609937Z digest=sha256:9d9ae8d170b184a71cf07e9170a53327981385d7331bf89c635097f5325f7d35

Observation 67a0b2f9-4580-4f30-9e34-68085d9089bd · inbound

UserBench: An Interactive Gym Environment for User-Centric Agents cites this paper.

UserBench: An Interactive Gym Environment for User-Centric Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:26.335428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:26.335428Z digest=sha256:ff3c4a7d6ba91b360560ee708c3d73bd4ead937e2209b2102d8f754afadc5256

Observation 915671c3-e5cc-4ec2-9679-5d17a1a45d50 · inbound

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever cites this paper.

Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T18:46:16.475327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:46:16.475327Z digest=sha256:58721d61dc77e5c8d93b61b453a48b4d0fac73f69826a273dc70785db862be9a

Observation a03ba34e-3ea1-4c44-a50e-849e2df3cda7 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.792208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T11:02:55.529271Z digest=sha256:39681518f172928b5daff27b46fbcadd30b34a7d70f89d67e049f32448d1f1ab

Observation 0e9d9af4-cf15-48a8-824e-b99be0ce07c2 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:12.592641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:12.592641Z digest=sha256:e7a67127fab4974f6e86cd1016b30540087ca27ec7fa194a7d181a2dd048de08

Observation 523d99e1-82e0-4b96-84b8-3501b8c110a9 · inbound

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception cites this paper.

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:50:52.539483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T03:46:03.228969Z digest=sha256:6784f11d28c9cbf7818e6b711c66e3716b9d92df9e7632b5c1124132a559600a

Observation 0700dca1-1a6d-46f8-8cd4-7a0dabd4525f · inbound

Memory in the Age of AI Agents cites this paper.

Memory in the Age of AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 261

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:18:20.250588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T18:18:19.911342Z digest=sha256:c6c96b93b4b63dcf81cda4aeef0d14a238ff6826ec32bb71dffaef436989bc3e

Observation ad97d443-4774-437f-b967-c37f48d6ce05 · inbound

Toward Efficient Agents: Memory, Tool learning, and Planning cites this paper.

Toward Efficient Agents: Memory, Tool learning, and Planning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T09:21:36.280577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:21:36.280577Z digest=sha256:f65645b9bb9555c784e1f1064c91539878b539734ae4e96d909d22cfcf8e71b4

Observation 59ec2ede-71a8-4033-86a9-01bfcebf3c9f · inbound

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls cites this paper.

Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:46:33.268448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:46:33.268448Z digest=sha256:f9eb093ae0f97ef704edea15d8defe652840a5c2ecf0d3a02414a287fdaff8e7

Observation 4b853e52-1d2d-404e-83a7-6d6775680c58 · inbound

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents cites this paper.

Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T19:59:59.030585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T19:59:59.030585Z digest=sha256:8198bfe5ef4afd1f2b18bcbc7958c534a34c67312bde0121d931b3b7554f188d

Observation 4f2bdf27-338b-487a-8c00-d32a047c8fd4 · inbound

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling cites this paper.

To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:10.333342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:32:57.054584Z digest=sha256:a997091c6ba2269021ccee7297b0e49aa290747582c59d93329170986d2291ad

Observation 4a5476e6-df85-4c89-9155-3d38d16c87f8 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:33.837524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T19:04:53.639217Z digest=sha256:2aac1a7d021f370599a7b243c106c7fd658b9901ec04dbfd05836835762c3205

Observation d00cbb03-51f7-491b-9c00-d4a541116888 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:12.424114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T00:25:46.507401Z digest=sha256:6a5e8a31c6860c0bd3db07aee001194a10634e8724de9e2e6586f67516c8b5fa

Observation f87b296a-6676-48f1-9d59-9f30edbfe7a2 · inbound

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval cites this paper.

FitText: Evolving Agent Tool Ecologies via Memetic Retrieval MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T05:24:44.710671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:24:44.710671Z digest=sha256:95c19ec65592eaf5575a45f2070c7ca4a4bd5cc844b2bdcf01c35e41d4dd08d8

Observation 7709fb7c-2ad0-4ef3-886e-0853e6383a01 · inbound

Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI cites this paper.

Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:37.672983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T19:06:36.952203Z digest=sha256:6b707af48b5597d9ef97cf44a6efd24c9b538b8ca3656a2dada963ae602055dd

Observation 684b6e09-e7cb-4db7-9017-2bb70f976afd · inbound

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation cites this paper.

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:42.967662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T16:19:39.392516Z digest=sha256:795695a5907c2e3f6c9527a01d4465dcf612e150d300b79a005d5269fb5329e0

Observation a86b0777-73cc-4fea-906b-4dfb1a7f260f · inbound

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning cites this paper.

Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:51.785672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:29:47.384341Z digest=sha256:fbdc1ee2f2ed58b8654de2fec9799454103790330c6c31ea66de8e186d7dea4c

Observation b053461d-5763-4d1d-8f70-7b549a050ee9 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:35:04.476298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T05:31:54.252816Z digest=sha256:28f519885ed976da504c131f6995a4385d84c47aba3bbf4474d43b46546c2474

Observation 3e545d30-a3b9-44e7-b915-75cc50001943 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:53:43.552641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T20:52:20.975459Z digest=sha256:398cb8c3e7d07ba823eabf9e9dda35ca867f34954d252fba5257d6556a09d006

Observation d473db4c-3e82-4cd3-9767-bcb8063c5b44 · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T03:08:57.998107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T03:07:38.232966Z digest=sha256:13a15ff26e32707f509f6db0941c223ab21a82bcba5c77660edd27c3f647cad1

Observation 749d96c9-e863-41ee-9a88-6548c8b301dd · inbound

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems cites this paper.

Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 165

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:52:39.943407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T16:51:13.491389Z digest=sha256:97edc968aee6c4f819c6e590c7a87324172a5db8843c9515f25b61d820c7a7bf

Observation b36e908b-d4c6-4335-a441-73f48c1b2f6c · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.401757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:67320309f09fc1c60fab14fed9f8b6d404e59aa1a2c437473b24b349cefd3fda

Observation b7301589-2913-4ce2-9e68-39c437d74e07 · inbound

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning cites this paper.

MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.732455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T00:25:59.532059Z digest=sha256:0fa3b8fee365c389548ec92c726effca9e6f4abd2a7b02e0bd2a204404137f0d

Observation df94b234-2959-4db7-a73e-2469787000e9 · inbound

Capability Self-Assessment: Teaching LLMs to Know Their Limits cites this paper.

Capability Self-Assessment: Teaching LLMs to Know Their Limits MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.643921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:25:50.579196Z digest=sha256:19ac7bf3f1eb8daff985759b88ac5879bfaa49942160a308fd23b134005b0839

Observation 0cbf862b-fdcf-46a9-adff-cdc6cc2f28a2 · inbound

NTILC: Neural Tool Invocation via Learned Compression cites this paper.

NTILC: Neural Tool Invocation via Learned Compression MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:04.773446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T00:05:36.600389Z digest=sha256:076cfd5e76703d094f90a6f5c5480ad6e51cf2a0558a233da6ae79860bd81527

Observation 7f6a9ba2-b68d-4f41-8fe9-02eb8f016c6e · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.155289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:34:23.382705Z digest=sha256:53e76bd5ea4b5de8355b62d841a6b09737dfc17b4db4e49f909e90ccd0931bb1

Observation c152ec48-6f8b-4634-82ed-6532c4cc7091 · inbound

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment cites this paper.

The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T14:59:04.152474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:59:04.152474Z digest=sha256:bb6bcf69888e93c93821763035abf455b059d868a60a82a0afdd3e2487e3bbbf

Observation b490de2a-c3a7-4ef9-8b45-7932e70decb2 · inbound

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose cites this paper.

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:18:59.716659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T00:37:50.570757Z digest=sha256:e187a3712a3b770b1b3e99128589f008ef1fcfc2fe22dfe4c7a64497b60dabd7

Observation cf78261e-76a9-4938-8b66-59b2e68a263f · inbound

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing cites this paper.

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T01:35:09.657698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:35:09.657698Z digest=sha256:35fb5841ae7eb1ce90b2519b6e6cb22a2ec9fe556ab925850a275b5aaeec07a9

Observation f5bd2c1c-a64a-4a8f-bc81-35180af2a342 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:29.777251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:29.777251Z digest=sha256:c6368efcbcd92e973bcfb0694d13a2fb0c5e65d7811c74c9bb27418dc21bad96