Pith. sign in

Paper Citation Record · LEDGER

WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2407.12823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.12823 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:39:57.542284Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:26:22.106095Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1c547470-1b7c-4923-aecd-6c521ac251f5 · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:57.542284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:57.542284Z digest=sha256:9ea428ffe1d2be04c7afe64f9ce89e288bbd233b5956831c194701b3b0023a9b

Observation f3724a34-b6fe-4ef0-a8cf-04b09d6de3dc · inbound

AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org cites this paper.

AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T16:56:24.825344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:56:24.825344Z digest=sha256:266404872677cde343b43570e8bc4cb53fea80bed54f418da23241c0be60ac77

Observation 201dadd7-5a20-47cd-905d-9eac6ec4efab · inbound

The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge? cites this paper.

The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge? WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.109894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:25:58.875055Z digest=sha256:9e1aec194bcdb837e1df936f9fb4e78ef98fb7c16e089611e7a1ffbfce4beb09

Observation 3193dece-5d4d-4a79-bc06-453ba78448c9 · inbound

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning cites this paper.

TIDE-Bench: Task-Aware and Diagnostic Evaluation of Tool-Integrated Reasoning WTU-EVAL: A Whether-or-Not Tool Usage Evaluation Benchmark for Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:31:23.915767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T02:47:31.830410Z digest=sha256:683d8b427e0bac1a3e27144d4e5a7a1a757987fe3e2756d210168282980266b3