Pith. sign in

Paper Citation Record · LEDGER

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2506.13977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13977 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:44:06.111125Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:44:28.494145Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T09:55:41.045115Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 217bbc1a-257d-4b69-9794-05b4d991820e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.812223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:43:58.852016Z digest=sha256:bcc909b2fdef38a84be51ad567affccf56d736999a91ec33cbfd8a93252d2e0f

Observation 0c0cfbec-79ff-4061-9b7b-227bdb64af60 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.659430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:43:58.980243Z digest=sha256:8603d5b000122862d40a2d17cd5e9603702847f9882b6fd41f860d5d9d064d85

Observation 9f0377f4-e565-4569-89f6-490bbdc4321c · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.477070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:43:59.122421Z digest=sha256:cff28b058d4405ce1a5aeb63a89018652fffa802efe23e2d9987cb6e77645d09

Observation 7fecc16b-e1d9-4ccc-a44f-8c7668226c5e · outbound

This paper cites Learning From Mistakes Makes LLM Better Reasoner.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Learning From Mistakes Makes LLM Better Reasoner

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.291125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.291125Z digest=sha256:1627474f9654f32ef5f63a844d65a6a722e257a9cd8df4acadb086fc18e7c66e

Observation 3fadd0eb-838a-4d0c-8203-639c4e8f39d3 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.452589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.452589Z digest=sha256:c9dda0e803b72c28210f16f4284c13fad6663ae8a579afbdd6810469a0a165dc

Observation f8c47a9e-9490-41e8-bc0f-24251e672eb8 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.669916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.669916Z digest=sha256:262d11e4f20d3ccb01fb3918eb15bbb63a938bd7574aded2f47ce54528fe30d6

Observation 845a9e89-81e7-41c7-a4e2-7771fdf4b6f8 · outbound

This paper cites NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.849002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.849002Z digest=sha256:89a78ed108f9de5b36aafd526d2b61cc43bf084fd8b68d7e14876daa773eac41

Observation 1654c7bc-fb7d-4a53-a740-21d386f76742 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:00.072088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:00.072088Z digest=sha256:6b279980d6bc71ba83efbd1ba2f40b2f1ca14adc87117bf37d8cfed5288b4e42

Observation 689c701e-040c-47bc-b1a0-7613609aef61 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:00.225070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:00.225070Z digest=sha256:dd2afe84693879d47e3362b06105c04d8ea21ec37c84fbb97d5dd449873c3705

Observation 72fdc844-aeb9-4214-adb9-c7bfd31a1ef9 · outbound

This paper cites Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:00.419151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:00.419151Z digest=sha256:ed0ee518665f8cd7f6cacb328078c830eaf229328a74f029f74454dd184f2972

Observation e759c2be-f0f9-4544-99a3-e67b35b0900e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.292439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:00.639866Z digest=sha256:9424aa34ecf1d1eeb8859f825d6b34735e09515ab636c7ca3ff4edd48080202b

Observation ffd358cf-dc7e-4423-8913-b22c9e5d2a6e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.144833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:00.793887Z digest=sha256:595437d0e140a621309e1845115629888932a75077ca1a9dd2a369032176c862

Observation eeb95397-334c-442f-8fbb-74addd9d7918 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.033520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.033520Z digest=sha256:78d1ffe82d9e6d7a277353b8876bd4b9ad4d48b4c5db2402e8bb7e369da71f83

Observation c7931df2-f944-4b98-b455-0e7d77241325 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.243326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.243326Z digest=sha256:f382a1ac1912d84777cf8ed9612a6501ad802539d542594cb2dd21c9c8d9cd50

Observation 7e7d13f1-3ad6-41a8-9b1b-52245361afbc · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.019221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:01.421439Z digest=sha256:ff6c47a828abbdbfd019f7a7186bd60f2a2230c567b9851df614a48f7b491b1d

Observation d2444bff-9f8a-4a7e-a33c-c5eaed1e5a6d · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.856039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:01.644137Z digest=sha256:f772d646a02192524007572c28a74656bffaf452c69fbe3b1dce1fec3d7e9f6e

Observation 66880987-ac1e-4fc5-aa42-4bb68f4f11c0 · outbound

This paper cites GPT-4o System Card.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.792913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.792913Z digest=sha256:0b506aae7a034431f53659ea409fb853b8eb0c2b436e2d54750c2a121ce5b705

Observation bac4134f-cf69-46af-9fee-3fc533a43831 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios A Survey on Large Language Models for Code Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.993713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.993713Z digest=sha256:959384d330b2eec18e174aa11f22cb2a3a2853ab6af93d193764b91c54af349b

Observation c834917b-d2d1-4eb4-8348-e13f40af22a9 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.725540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.149884Z digest=sha256:e8f7daf210f2812717427ce42be580d6153a917d3217aa0c18c93a1dc83c92ab

Observation c7c21ec4-5f7a-43fd-b22e-34bb9c2be634 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.575434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.356201Z digest=sha256:567ed8d948b97b0a2f17490416cc11405e3a2417d464ae9e17e7c89a7eb375cf

Observation f3cafb9a-70c5-4a84-bdef-16b74ec4736e · outbound

This paper cites CriticEval: Evaluating Large Language Model as Critic.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios CriticEval: Evaluating Large Language Model as Critic

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:02.536287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:02.536287Z digest=sha256:9a87e8dac0663308966f72d615f02bcce41c5f1663c645b0caabcb2a63aaa7f9

Observation 9c78a4a2-24f3-47c4-9a50-4634594879d9 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.416356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.662842Z digest=sha256:c0df60a2e66ee9b127f944016aa168cd06658526edfa30dda38a9f993281ca9b

Observation 7ce7f19c-92ea-4f12-b406-aa7f1353fe1a · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.296384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.813366Z digest=sha256:c5de024aaf53e88f130dd02e305bf571f64db33269ee6ef86dc6004cd925717a

Observation acc7b5f5-65db-484e-bba3-d81f28142f86 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolACE: Winning the Points of LLM Function Calling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:02.936014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:02.936014Z digest=sha256:a55f741962168579d1288de222ee9c5de670e00f1211285a5f0b0193ad707740

Observation 02bc1dfe-46d6-417a-bbef-fef24ef561b9 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LLM Critics Help Catch LLM Bugs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.053326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.053326Z digest=sha256:b8f40b862093fb1d7fe0a123e580dd5b8d84f9b8c3651dc2b341ac2b25caf5ef

Observation 7883b501-a81c-4491-9c33-e6e412c7962d · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.157613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.157613Z digest=sha256:1ac2595b449e790e4da41b56a8c11d5bad3bd57450d70fe0c0a0cb10221b832f

Observation 8fbd3ae0-f63f-492b-a458-819abe786dac · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.134734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:03.278928Z digest=sha256:80e5389091d56819fc853df3dd4f9e163356876d66d512fc4d0702a312ddee84

Observation d1f1c625-1a6f-4184-94ef-08a192067740 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Gorilla: Large Language Model Connected with Massive APIs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.440975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.440975Z digest=sha256:72b284811c7b67f326ac54f9e21d00da1280e41b2f6a6b7f0b513b3b76638790

Observation 93279016-7dd9-445e-adf8-3b5823caf82c · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.514998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.514998Z digest=sha256:9d542baa4435f2130ff18bd53463a1753264b9e25ec34a76fbf101a4d0bd4180

Observation 172fa7a8-5323-4f0a-bd87-b9e72adeaf4c · outbound

This paper cites From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.629843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.629843Z digest=sha256:b85c828e29bc4285eadfba2f183c54501e08ba79ad68af5c0319bf595eafb929

Observation 8ffde937-60a1-4432-ac8b-4277c4bf7c0d · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.752979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.752979Z digest=sha256:134f5843ff33b7576332c504b024655c0071e9ddbd09d3f9e2ebf8411c271619

Observation 8a54ff64-1f1d-4520-a5f0-b99007a8a92c · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios TaskBench: Benchmarking Large Language Models for Task Automation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.797982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.797982Z digest=sha256:2c392ed1a46fc57f8404847f3d43bca9af09ab6b5252567935f45a56c8983f18

Observation 1c5587e5-b660-4644-a920-f09aac92215f · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.906567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.906567Z digest=sha256:9bc8a36c1e2a29c7d7410c4e768014be97341e55a04ca8caffe53b98327d9eff

Observation 38c48ce5-9139-484e-8c75-e579e826ac76 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.957618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:03.984334Z digest=sha256:ca97b9beb88836b5bfebdef2bdf3fb89e636eedbc65c7eb92289383e89ea7cc7

Observation 76a0d6b7-153b-412d-9ca2-33d62202f73e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.809782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.079080Z digest=sha256:b517490df03325062c286238d159f1d65eec0cbc7be1fc6dca54d7b21b9fad1c

Observation 0647b2b7-efd5-4d81-97fd-78b1d48a8f1e · outbound

This paper cites Qwen2 Technical Report.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.166038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.166038Z digest=sha256:dfaa313362d082708ab33d482e17104c9a24320bd416c8eafc102eb21047e740

Observation 47a6ee20-8ae0-4209-9e94-6ba84077b99c · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.305695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.305695Z digest=sha256:296f10abea9ae9c468de4bcb7bf2d8afdfcf5a1f206798b435722261bb0b6ac0

Observation 780b2f58-8832-4486-844d-86d6bf3a3b8b · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.423186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.423186Z digest=sha256:95e8b5d77336ddaef0dbdfd174c1197b70ea4fe80a2b69e7eb437c45c775996f

Observation 5ae17871-d694-4766-815c-5583c4ebcb13 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.649387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.555717Z digest=sha256:cd21b2a0e66c30159c95a3b49ebaae922fcd14b6d78833dbbca730f3985edb95

Observation e8efb481-ee2b-4016-ad62-e2f495fb91d5 · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.671307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.671307Z digest=sha256:ada037e6518358f108b3989ac8b05093e34ce2c65d207c7a5d625a2fd22e7a50

Observation 7fd377a5-55a9-47e0-8e70-c03d14bd779f · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.485880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.778585Z digest=sha256:264d0f686793cfb27d9379778a0f51cf1e1a3c3246806c25c577ea885e8af5d8

Observation 9a216b23-d2b7-4c7b-a3e1-e26f67b2ee8d · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios On the Tool Manipulation Capability of Open-source Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.866566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.866566Z digest=sha256:a457ad7c3c979952947dd89090e0c8c82201f3a086fe2a26f55f452e9a85d6f9

Observation 12a6544b-3b5d-4ac5-802f-efc3b7c78a1e · outbound

This paper cites Patil, Ion Stoica, and Joseph E.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Patil, Ion Stoica, and Joseph E

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:44:07.356573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.996971Z digest=sha256:9bfd545efc35854a729fb0b99e0ad839e162cce118ea9ee31e2918d5ca99f9cc

Observation be33c64d-3ebd-4ac8-af3b-46dcf181c08c · outbound

This paper cites Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:44:06.391470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.228437Z digest=sha256:509afcbafdfe1c1574c9aad732744b63d32bf29def6bec7741b3beb0c7049ca9

Observation b304eab5-2ffb-4dca-acf6-3cbcc9d960da · outbound

This paper cites ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.285024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.285024Z digest=sha256:0d6e49bba910761d9f6b9a19e304a17ad6a859f1dca4db9139f6dd3c9b00aa19

Observation f92ede64-81df-43ec-805d-2bc53fecf758 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.227197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.392905Z digest=sha256:737d2f476d3304eef8de1ea7ababd9d2c19092a09c0eab573f73fd684710cbe2

Observation ef958cb6-c84d-4d1c-9d3d-11da26766dbf · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.045908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.512316Z digest=sha256:7f60a3c9b900bc41202d9ee71000920ca25ed5068d0d152210fcba3a0eec8fbd

Observation c6049b07-f49e-4e96-9aed-d247befb0e8b · outbound

This paper cites EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.615100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.615100Z digest=sha256:2a91e8876f03f4f7bed318a0a1ea6620660042d1d4d4b50423386835008d859b

Observation 79c6ef08-dc2b-47df-8e58-ae41965ded89 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.694410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.694410Z digest=sha256:979b972d21955174d75cddc12a2ca4fa7ffbf6cf3cda816b3ed35bb1fcbfa178

Observation c75468e2-0c74-43f0-aaa8-0b4cb451c5ef · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:06.803432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.781953Z digest=sha256:1e0a9555d6c0ebb0f5a8738735a2a7e0f26671a894b361b51c82288014a39328

Observation abf54ff6-b669-4a22-bb08-2eb7a92e59db · outbound

This paper cites A Survey of Large Language Models.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios A Survey of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.874236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.874236Z digest=sha256:a3af7c90f65c3a7125129e0e7492f456cdae1e98ffffa7b5bf27fa027aff4441

Observation b177e4ee-7bfb-496a-b3b8-d879c4c037b6 · outbound

This paper cites online" 'onlinestring :=.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:06.009699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:06.009699Z digest=sha256:6dd3ecc9f7097d4fb4e28155cca8870857fe77b2a877c1122bafbc0b4320d4ed

Observation b76269b9-64a7-456d-af9e-e20f598c35dd · outbound

This paper cites write newline.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:06.111125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:06.111125Z digest=sha256:1a9a292c13ca04a8f4ea6518d169ab18231ba55317a0753e056fce66bbffadc4

Pith citing papers

Observation 568c73ec-2278-4e24-ac49-9a81921e4d2b · inbound

Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning cites this paper.

Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:49.792116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T19:24:46.023589Z digest=sha256:ed5d6af8a5e30b4abf9726a6533bd8250a76f4670b039d1215474b1e1591a6b8

Observation 4904fd2f-b77a-4838-b57f-5cab4c0c216a · inbound

Recursive Self-Evolving Agents via Held-Out Selection cites this paper.

Recursive Self-Evolving Agents via Held-Out Selection CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:14:37.523405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T11:13:15.923479Z digest=sha256:7ec9ece48b094f0af399a7056d83a41fc278172358f2ecbe2252be6ff9fa6fc7

Observation 1a7a4e19-1268-4d10-a232-8b4d46724ac3 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.046656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:3feeee5f693ed059d30f01572024eddf1485b69dcba45943e7f4c67560e08d91

Observation da4b10f7-2315-4881-8542-3bc14fb3835f · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.494145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.494145Z digest=sha256:35b7833458508bb104ffffc4dc4ae1a0d10580e2477ab2a626e0bc8f1cf1548f