Pith. sign in

Paper Citation Record · LEDGER

NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2409.03797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.03797 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:04:46.846452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:36:29.696220Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 845a9e89-81e7-41c7-a4e2-7771fdf4b6f8 · inbound

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios cites this paper.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.849002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.849002Z digest=sha256:89a78ed108f9de5b36aafd526d2b61cc43bf084fd8b68d7e14876daa773eac41

Observation 57fe1622-e7d9-4369-90da-9e236bc5ade3 · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.420327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:c9a6b7facf405a5ef451aee88cb98bb5c58344ae2d20e7d828dafc115c703d95

Observation 89aaede8-7ad7-492b-ace9-1b6eeb8d658c · inbound

LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios cites this paper.

LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:04:46.846452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:04:46.846452Z digest=sha256:7fdb8703f3b0f97a9483a86ffd85c1aa03881f75326392c442650d24a685ab45

Observation 47a21541-db06-42f7-a64d-4d957cf2a0e7 · inbound

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench cites this paper.

How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $\tau$-bench NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:00.653584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:49:00.653584Z digest=sha256:14c1c93c4ec14b94f5401283e6ff9af44dccbc1c3496bf149e0df23265613107

Observation 91a285a0-ab89-412b-9e4c-467a7bfe332a · inbound

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling cites this paper.

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:30:59.514386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T06:30:39.858246Z digest=sha256:06d73e3788e35a71ab888a7bcf5424f1bbf6a6c230947ec3a27310c7a2d44b78

Observation 71622b61-2a6f-4e43-8e98-6b82048bf7f5 · inbound

RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation cites this paper.

RAG Strategies for Natural Language-Based SQL Query and REST API Call Generation NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T06:09:50.756340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:09:50.756340Z digest=sha256:89af9003383a3d433dfa54f7cdc9a710d9d7e9d88b0f2a0bb08da9aa93d96c30

Observation 4a652ed7-9a22-404c-aa7b-60c8d909eb08 · inbound

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments cites this paper.

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:29.697608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T09:54:00.111238Z digest=sha256:5316a7db1658c2c1207b42cda1e4ccf557eb2f4367d5c529aa8c95cc8bc702e8