Pith. sign in

Paper Citation Record · LEDGER

TaskBench: Benchmarking Large Language Models for Task Automation

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2311.18760.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18760 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:19:06.989495Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:18:59.738705Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7f4a8dc9-6794-45d8-a81b-113d55c31a5b · inbound

CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning cites this paper.

CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning TaskBench: Benchmarking Large Language Models for Task Automation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:06.989495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:06.989495Z digest=sha256:43dca2ad59baab0bdb74cfb6f3fe441f9ab9da7e73fb0ca6e14f03d27aa23cf6

Observation 22f25df4-17a1-4a54-bb3b-133f2b219cda · inbound

Action Engine: Automatic Workflow Generation in FaaS cites this paper.

Action Engine: Automatic Workflow Generation in FaaS TaskBench: Benchmarking Large Language Models for Task Automation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:10:22.011780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:10:22.011780Z digest=sha256:5313f97873bc1573a881bb76ee8b9550cc2a988f6143244508fe1e05e9b97a07

Observation 8dfca7e9-594d-40af-a297-4a0f3b50207c · inbound

AI PERSONA: Towards Life-long Personalization of LLMs cites this paper.

AI PERSONA: Towards Life-long Personalization of LLMs TaskBench: Benchmarking Large Language Models for Task Automation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:31:41.360494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:31:41.360494Z digest=sha256:90388c784834bea9958953f49a0414a88b370c20c7ecee0dd3e719896d77806b

Observation 9fd4e417-dddc-4af6-b6b0-b61d03924c2c · inbound

LLM4SR: A Survey on Large Language Models for Scientific Research cites this paper.

LLM4SR: A Survey on Large Language Models for Scientific Research TaskBench: Benchmarking Large Language Models for Task Automation

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-10T21:39:25.659017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:39:25.659017Z digest=sha256:4e8a5322c7a7aef1a12dd189c6ef0ae6340cdbd871b4e43d08511a94722b4b6c

Observation c3f77fc0-34fe-4860-a236-f3a330bb0d90 · inbound

AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement cites this paper.

AdaptBot: Combining LLM with Knowledge Graphs and Human Input for Generic-to-Specific Task Decomposition and Knowledge Refinement TaskBench: Benchmarking Large Language Models for Task Automation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T13:31:47.258759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:31:47.258759Z digest=sha256:ca692c8f16a171cce5475b05f7717432922296e7be78895ee83e4cd500357b50

Observation 06084d52-bd1f-4843-ba0b-22d00827e86e · inbound

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities cites this paper.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities TaskBench: Benchmarking Large Language Models for Task Automation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.265475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.265475Z digest=sha256:e36ab0099afd03d26fac4b6a4a373df9b985f8b2931aee39ce37e49c63775a9d

Observation 8a54ff64-1f1d-4520-a5f0-b99007a8a92c · inbound

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios cites this paper.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios TaskBench: Benchmarking Large Language Models for Task Automation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.797982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.797982Z digest=sha256:874fee911cb8988c9d2dd6e95b50759a6a15fa409f4fb3996c33bfad87615c4a

Observation ab839f1b-689b-47f7-a359-3fff8177c7b7 · inbound

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues cites this paper.

DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues TaskBench: Benchmarking Large Language Models for Task Automation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:40.257193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:02:40.257193Z digest=sha256:70e53e748a1e4718e8571d6a458987619933a13dc528fa8ca55f3c18882419f9

Observation 0ff91483-f9bd-427e-b52e-bc1f71953158 · inbound

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems cites this paper.

Butterfly Effects in Toolchains: A Comprehensive Analysis of Failed Parameter Filling in LLM Tool-Agent Systems TaskBench: Benchmarking Large Language Models for Task Automation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:58.036946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:39:58.036946Z digest=sha256:ed91aa3307db847013c732dd891a663297cec912f1a24fe8ace87f17bb9eb20c

Observation 8fd32775-a877-4f2f-8338-4d5d8e800dc2 · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey TaskBench: Benchmarking Large Language Models for Task Automation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.762114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.762114Z digest=sha256:0d05469363f752db7e5eaba9d441b2f8c74b6d678a13f874ea9f680582455f2a

Observation 5fd0df3b-8120-4a59-9dc2-31261a1047b8 · inbound

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints cites this paper.

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints TaskBench: Benchmarking Large Language Models for Task Automation

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-21T12:00:04.241730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T11:57:53.779767Z digest=sha256:979fa293d4e6d43ca7a374b78049098883737dd683ea557dce022e0bd5e7811d

Observation 0f88185b-6004-4a3b-9ec1-e81652714fd7 · inbound

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation cites this paper.

From Intent to Execution: Composing Agentic Workflows with Agent Recommendation TaskBench: Benchmarking Large Language Models for Task Automation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:42.825400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T16:19:39.392516Z digest=sha256:e6750e86b2810f62eab82ed020cf2393ff606fffdac612aceb2be11c4200c1d3

Observation e3d1ed72-7b21-47d1-989a-8ca379ac76df · inbound

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications cites this paper.

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications TaskBench: Benchmarking Large Language Models for Task Automation

Reference 131

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:20:57.226817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:47:39.926540Z digest=sha256:38a462df75d322ff4a6e8e1de5c22d35a1c3e4bd7faf631059391ba23b68528b

Observation c36dc98e-b038-476c-a0da-8d5223a49814 · inbound

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications cites this paper.

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications TaskBench: Benchmarking Large Language Models for Task Automation

Reference 133

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T23:19:14.564249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T23:15:44.550045Z digest=sha256:3667d7a82894793b7e4fe3d525c24bb2920efce84e4ac250e591e1e061247031

Observation f00eb9e9-5c6c-4fa8-840f-e52a854f5302 · inbound

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications cites this paper.

A Comprehensive Survey on Agent Skills: Taxonomy, Techniques, and Applications TaskBench: Benchmarking Large Language Models for Task Automation

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:25:07.428566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T23:23:42.883286Z digest=sha256:b4e8650100ee093ab9f875be834328ab1b9aabc3be0dc959d4e82b0bf80d08ce

Observation ba5974ba-b04f-4dc3-92ec-390f9993e4d4 · inbound

The Scaling Laws of Skills in LLM Agent Systems cites this paper.

The Scaling Laws of Skills in LLM Agent Systems TaskBench: Benchmarking Large Language Models for Task Automation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T18:13:37.585700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T18:10:08.737710Z digest=sha256:66823c084fb1e7f3cb0c12fc95e7aef24f1d9a88f4f9c26e57d34b19b62cef11

Observation 5ca47bc0-0474-4ff5-a8e9-43e394bf85e4 · inbound

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery cites this paper.

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery TaskBench: Benchmarking Large Language Models for Task Automation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:32:45.524880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T20:29:27.398862Z digest=sha256:ef3cdca288c71b82523dbc7bd7c8b904db988b54b4fe790fc4bbb5a4a5d1be7a

Observation 0c073145-80da-41e9-b334-c9fe81fe015e · inbound

Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams cites this paper.

Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams TaskBench: Benchmarking Large Language Models for Task Automation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:01.567575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T23:09:59.207650Z digest=sha256:da000b85589d8dadd60ed948365e3bcde6fb3351c02b8b05b5d15b103dc97a47

Observation 2f4cddb0-fea0-4544-bd36-1835165cd92e · inbound

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task cites this paper.

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task TaskBench: Benchmarking Large Language Models for Task Automation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:56.923529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T09:55:01.205573Z digest=sha256:80a5bed875b9308b7e18469838a2714787e2d61504409d64a04f9f00bdb1d9c2

Observation 416983a5-6e59-45f4-81fd-e3a9e03315dc · inbound

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose cites this paper.

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose TaskBench: Benchmarking Large Language Models for Task Automation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:18:59.740372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T00:37:50.570757Z digest=sha256:3958246064f4b19639909f6bd3961e88dc97f52b7cb66bea65cf2271537d7925

Observation 0d6b121e-3485-4c8d-903d-5a29cc237600 · inbound

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents cites this paper.

ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents TaskBench: Benchmarking Large Language Models for Task Automation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:35:41.902511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T05:23:21.162004Z digest=sha256:2de77bfd3cd2081ea2c5534d3d381220b442d5bfae2c78b76c9fc5116cb5843d

Observation 741ecbfb-e0ae-454d-9095-65ef8cbf5832 · inbound

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities cites this paper.

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities TaskBench: Benchmarking Large Language Models for Task Automation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T11:02:26.692134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:02:26.692134Z digest=sha256:39ecf6c2e8b5cd8f285595327417d86599c2d9e1180f2efe9af9d74fc90c151c