Pith. sign in

Paper Citation Record · LEDGER

ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2401.00741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.00741 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:01:44.628674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:07:04.799187Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3dcb81aa-b902-4f1b-b2e8-6311d4761d58 · inbound

Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training cites this paper.

Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-08T15:01:44.628674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:01:44.628674Z digest=sha256:818661f3f569c4a6e5d28750eb1c176d1ab8d6eaa9708f9b8375a00326a0c536

Observation 7c537a96-3284-4682-9589-d888e1d1cca0 · inbound

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving cites this paper.

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:18.664602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:18.664602Z digest=sha256:b67fefea1e259eb017d5109c3aa64be34150433637771e1f123c23fc4efd8218

Observation b304eab5-2ffb-4dca-acf6-3cbcc9d960da · inbound

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios cites this paper.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.285024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.285024Z digest=sha256:87dcd9463bba60ae3d1b6b553d8e76fa35539728252ca903257fedf5670fedb0

Observation aa42965a-feaf-4205-9eb2-6acc7788c794 · inbound

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues cites this paper.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.814433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.814433Z digest=sha256:4bcef452ab2cd9a5b748c3806933a2962ad644436fd72eb4f31b18f6b71150f2

Observation ee0df568-7c51-476b-acad-6e8f42a7fd51 · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:58.888932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:58.888932Z digest=sha256:44a7d357e92c767ee7027c89676b3ac98217971782b50e4dd25ccce465b28fc4

Observation 29acd129-a309-4c63-a61f-4df24546b5fd · inbound

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis cites this paper.

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:40:51.210432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:37:11.945418Z digest=sha256:5d37b169c12eddd87e68b3bce5cb3e5228358579f54b7a89092bc0610ac2b200

Observation ad188bc9-cd30-4daa-ac70-85e554974d2d · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:30:45.639649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:28:16.904091Z digest=sha256:3ce49b5180e9aafa4475a418c2747bbf553a9469714f1a65e1f499ee9f1cf8bc

Observation fe624dae-8cac-46b5-b1f4-f3b768992cae · inbound

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers cites this paper.

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:10:16.421156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T15:10:04.253250Z digest=sha256:227cc26eebb51219f3320cfff2e5c1b644fa3280824ed8dcd4b5bea4751176fc

Observation 3c091dde-fcd9-4222-8d54-a59ee4b03698 · inbound

NTILC: Neural Tool Invocation via Learned Compression cites this paper.

NTILC: Neural Tool Invocation via Learned Compression ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:04.800577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T00:05:36.600389Z digest=sha256:18f805b4997d20ce6f29cca2a5332ee55c831d330b80f3a8a9714deebf55b36f