Pith. sign in

Paper Citation Record · LEDGER

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

As of 17 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2506.13977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13977 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:44:06.111125Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:44:28.494145Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T09:55:41.045115Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 217bbc1a-257d-4b69-9794-05b4d991820e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.812223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:43:58.852016Z digest=sha256:ad5383b44b1d03a7a29c5c108a74984b72224687fcb616e7b4ba5977868f1c0b

Observation 0c0cfbec-79ff-4061-9b7b-227bdb64af60 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.659430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:43:58.980243Z digest=sha256:479c8ee1d4bc9e4514aeb7f450f206766bdffb0b748f35ec3de470eeb963129e

Observation 9f0377f4-e565-4569-89f6-490bbdc4321c · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.477070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:43:59.122421Z digest=sha256:43e2545e8bf33aeea75cee148a11ee9610a347eb2d0cdd02987812fe1696ca2e

Observation 7fecc16b-e1d9-4ccc-a44f-8c7668226c5e · outbound

This paper cites Learning From Mistakes Makes LLM Better Reasoner.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Learning From Mistakes Makes LLM Better Reasoner

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.291125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.291125Z digest=sha256:5003693d8dab07a567f15144558ef752119bab3820997530aec328a7549e23c8

Observation 3fadd0eb-838a-4d0c-8203-639c4e8f39d3 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.452589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.452589Z digest=sha256:c9dda0e803b72c28210f16f4284c13fad6663ae8a579afbdd6810469a0a165dc

Observation f8c47a9e-9490-41e8-bc0f-24251e672eb8 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.669916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.669916Z digest=sha256:262d11e4f20d3ccb01fb3918eb15bbb63a938bd7574aded2f47ce54528fe30d6

Observation 845a9e89-81e7-41c7-a4e2-7771fdf4b6f8 · outbound

This paper cites NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:59.849002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:59.849002Z digest=sha256:89a78ed108f9de5b36aafd526d2b61cc43bf084fd8b68d7e14876daa773eac41

Observation 1654c7bc-fb7d-4a53-a740-21d386f76742 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:00.072088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:00.072088Z digest=sha256:6b279980d6bc71ba83efbd1ba2f40b2f1ca14adc87117bf37d8cfed5288b4e42

Observation 689c701e-040c-47bc-b1a0-7613609aef61 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:00.225070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:00.225070Z digest=sha256:dd2afe84693879d47e3362b06105c04d8ea21ec37c84fbb97d5dd449873c3705

Observation 72fdc844-aeb9-4214-adb9-c7bfd31a1ef9 · outbound

This paper cites Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Advancing Tool-Augmented Large Language Models: Integrating Insights from Errors in Inference Trees

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:00.419151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:00.419151Z digest=sha256:d6423e0566e1f4068ce9b6f0d770b0a57aae3db9b07a6967a1b26d750dbc7537

Observation e759c2be-f0f9-4544-99a3-e67b35b0900e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.292439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:00.639866Z digest=sha256:e56e591dd5f20b021ea2b4b848e2166eb7febbf58d03ac251d4356816c8f16e7

Observation ffd358cf-dc7e-4423-8913-b22c9e5d2a6e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.144833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:00.793887Z digest=sha256:a08806251fffd577ac01485e77f09267e7a2e0c146f8993d4aca0f2d406fe886

Observation eeb95397-334c-442f-8fbb-74addd9d7918 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.033520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.033520Z digest=sha256:78d1ffe82d9e6d7a277353b8876bd4b9ad4d48b4c5db2402e8bb7e369da71f83

Observation c7931df2-f944-4b98-b455-0e7d77241325 · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.243326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.243326Z digest=sha256:f382a1ac1912d84777cf8ed9612a6501ad802539d542594cb2dd21c9c8d9cd50

Observation 7e7d13f1-3ad6-41a8-9b1b-52245361afbc · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:09.019221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:01.421439Z digest=sha256:806afa8b9684555254d5b3ebac1bf0ad56e7bd7c1641b1beb5cc23552e38c0df

Observation d2444bff-9f8a-4a7e-a33c-c5eaed1e5a6d · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.856039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:01.644137Z digest=sha256:4879aba1b528b655e49e5eff14952de00971f8a05fdc570358fc2ef399fe6caf

Observation 66880987-ac1e-4fc5-aa42-4bb68f4f11c0 · outbound

This paper cites GPT-4o System Card.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.792913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.792913Z digest=sha256:0b506aae7a034431f53659ea409fb853b8eb0c2b436e2d54750c2a121ce5b705

Observation bac4134f-cf69-46af-9fee-3fc533a43831 · outbound

This paper cites A Survey on Large Language Models for Code Generation.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios A Survey on Large Language Models for Code Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:01.993713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:01.993713Z digest=sha256:959384d330b2eec18e174aa11f22cb2a3a2853ab6af93d193764b91c54af349b

Observation c834917b-d2d1-4eb4-8348-e13f40af22a9 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.725540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.149884Z digest=sha256:7f2a7b4bea05c93de18e882510d268c71cc066b9cea2485170d1cd1b856c90d5

Observation c7c21ec4-5f7a-43fd-b22e-34bb9c2be634 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.575434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.356201Z digest=sha256:5332e03e94ace44859da46b1da0b82fda621dad90dce34f0658595adea61691a

Observation f3cafb9a-70c5-4a84-bdef-16b74ec4736e · outbound

This paper cites CriticEval: Evaluating Large Language Model as Critic.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios CriticEval: Evaluating Large Language Model as Critic

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:02.536287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:02.536287Z digest=sha256:b743388010203374077a468dd05b1638efdb667c8d4e533194455d3f497ac44e

Observation 9c78a4a2-24f3-47c4-9a50-4634594879d9 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.416356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.662842Z digest=sha256:66cf3e69cc14e417954ad092e41db6d4f3fe6683d017c19274fac62e0fc43e4f

Observation 7ce7f19c-92ea-4f12-b406-aa7f1353fe1a · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.296384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:02.813366Z digest=sha256:888688ad25086896d78a559c529f0d1cadcaf5c3ffdf9a38ad827e0284188797

Observation acc7b5f5-65db-484e-bba3-d81f28142f86 · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolACE: Winning the Points of LLM Function Calling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:02.936014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:02.936014Z digest=sha256:a55f741962168579d1288de222ee9c5de670e00f1211285a5f0b0193ad707740

Observation 02bc1dfe-46d6-417a-bbef-fef24ef561b9 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios LLM Critics Help Catch LLM Bugs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.053326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.053326Z digest=sha256:b8f40b862093fb1d7fe0a123e580dd5b8d84f9b8c3651dc2b341ac2b25caf5ef

Observation 7883b501-a81c-4491-9c33-e6e412c7962d · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.157613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.157613Z digest=sha256:1ac2595b449e790e4da41b56a8c11d5bad3bd57450d70fe0c0a0cb10221b832f

Observation 8fbd3ae0-f63f-492b-a458-819abe786dac · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:08.134734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:03.278928Z digest=sha256:b80dbb58ad1365e69d98c6924e0498d77f7f8ec95905bcb686419b7fd2f12127

Observation d1f1c625-1a6f-4184-94ef-08a192067740 · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Gorilla: Large Language Model Connected with Massive APIs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.440975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.440975Z digest=sha256:72b284811c7b67f326ac54f9e21d00da1280e41b2f6a6b7f0b513b3b76638790

Observation 93279016-7dd9-445e-adf8-3b5823caf82c · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.514998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.514998Z digest=sha256:9d542baa4435f2130ff18bd53463a1753264b9e25ec34a76fbf101a4d0bd4180

Observation 172fa7a8-5323-4f0a-bd87-b9e72adeaf4c · outbound

This paper cites From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios From Exploration to Mastery: Enabling LLMs to Master Tools via Self-Driven Interactions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.629843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.629843Z digest=sha256:b85c828e29bc4285eadfba2f183c54501e08ba79ad68af5c0319bf595eafb929

Observation 8ffde937-60a1-4432-ac8b-4277c4bf7c0d · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.752979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.752979Z digest=sha256:134f5843ff33b7576332c504b024655c0071e9ddbd09d3f9e2ebf8411c271619

Observation 8a54ff64-1f1d-4520-a5f0-b99007a8a92c · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios TaskBench: Benchmarking Large Language Models for Task Automation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.797982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.797982Z digest=sha256:2c392ed1a46fc57f8404847f3d43bca9af09ab6b5252567935f45a56c8983f18

Observation 1c5587e5-b660-4644-a920-f09aac92215f · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:03.906567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:03.906567Z digest=sha256:9bc8a36c1e2a29c7d7410c4e768014be97341e55a04ca8caffe53b98327d9eff

Observation 38c48ce5-9139-484e-8c75-e579e826ac76 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.957618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:03.984334Z digest=sha256:0be3519b238d8e5ed5158a60a00cf71b0b77a82180d6301437212028123af015

Observation 76a0d6b7-153b-412d-9ca2-33d62202f73e · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.809782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.079080Z digest=sha256:ba23d622fc8c4f6147bf036c0ca930b0f3709d39ca9b7c01014dda68405c266f

Observation 0647b2b7-efd5-4d81-97fd-78b1d48a8f1e · outbound

This paper cites Qwen2 Technical Report.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.166038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.166038Z digest=sha256:dfaa313362d082708ab33d482e17104c9a24320bd416c8eafc102eb21047e740

Observation 47a6ee20-8ae0-4209-9e94-6ba84077b99c · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.305695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.305695Z digest=sha256:296f10abea9ae9c468de4bcb7bf2d8afdfcf5a1f206798b435722261bb0b6ac0

Observation 780b2f58-8832-4486-844d-86d6bf3a3b8b · outbound

This paper cites Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.423186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.423186Z digest=sha256:bf3a402d2fd190c205c0fa89e36fbe418f7b0b97ae91ab5dac999292721c0859

Observation 5ae17871-d694-4766-815c-5583c4ebcb13 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.649387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.555717Z digest=sha256:33b4567449b9b7043d150de007f08b7fef89d5187caf50560ce793c9c7fb6114

Observation e8efb481-ee2b-4016-ad62-e2f495fb91d5 · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.671307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.671307Z digest=sha256:ada037e6518358f108b3989ac8b05093e34ce2c65d207c7a5d625a2fd22e7a50

Observation 7fd377a5-55a9-47e0-8e70-c03d14bd779f · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.485880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.778585Z digest=sha256:df0f7ae730ad3471905edd1a664f954add7f550fe01576f9a70a1a427db148ab

Observation 9a216b23-d2b7-4c7b-a3e1-e26f67b2ee8d · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios On the Tool Manipulation Capability of Open-source Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:04.866566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:04.866566Z digest=sha256:a457ad7c3c979952947dd89090e0c8c82201f3a086fe2a26f55f452e9a85d6f9

Observation 12a6544b-3b5d-4ac5-802f-efc3b7c78a1e · outbound

This paper cites Patil, Ion Stoica, and Joseph E.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Patil, Ion Stoica, and Joseph E

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:44:07.356573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:04.996971Z digest=sha256:b600082ff16fb14f4f4c96671c8ec646992c6741a9bb95c8ba15237da547eaac

Observation be33c64d-3ebd-4ac8-af3b-46dcf181c08c · outbound

This paper cites Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Varying Shades of Wrong: Aligning LLMs with Wrong Answers Only

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:44:06.391470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.228437Z digest=sha256:a2ab8cbd69a85b2e59305ef0df74f085b782f71615874f1b445d635c82ddbdbd

Observation b304eab5-2ffb-4dca-acf6-3cbcc9d960da · outbound

This paper cites ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.285024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.285024Z digest=sha256:0d6e49bba910761d9f6b9a19e304a17ad6a859f1dca4db9139f6dd3c9b00aa19

Observation f92ede64-81df-43ec-805d-2bc53fecf758 · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.227197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.392905Z digest=sha256:297fe817d8ee72a7d7421696bdace79a4a6d5d356867a297c2b0d21d85e6d626

Observation ef958cb6-c84d-4d1c-9d3d-11da26766dbf · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:07.045908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.512316Z digest=sha256:eaef08bcfee1b76f6197c3d614a322a2d6f6c3cf23d14f861033abf564078dad

Observation c6049b07-f49e-4e96-9aed-d247befb0e8b · outbound

This paper cites EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.615100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.615100Z digest=sha256:2a91e8876f03f4f7bed318a0a1ea6620660042d1d4d4b50423386835008d859b

Observation 79c6ef08-dc2b-47df-8e58-ae41965ded89 · outbound

This paper cites AgentTuning: Enabling Generalized Agent Abilities for LLMs.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios AgentTuning: Enabling Generalized Agent Abilities for LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.694410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.694410Z digest=sha256:24329cd1b2f70fe1ab822e161b8e13e7a1c338951ef6538aacd029cbc52da59a

Observation c75468e2-0c74-43f0-aaa8-0b4cb451c5ef · outbound

This paper cites an unresolved cited work.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:44:06.803432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:44:05.781953Z digest=sha256:3d54ff92b719fb656442335072197fa06bb4eac819a0c4dff672045f0908f115

Observation abf54ff6-b669-4a22-bb08-2eb7a92e59db · outbound

This paper cites A Survey of Large Language Models.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios A Survey of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:05.874236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:05.874236Z digest=sha256:a3af7c90f65c3a7125129e0e7492f456cdae1e98ffffa7b5bf27fa027aff4441

Observation b177e4ee-7bfb-496a-b3b8-d879c4c037b6 · outbound

This paper cites online" 'onlinestring :=.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:06.009699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:06.009699Z digest=sha256:6dd3ecc9f7097d4fb4e28155cca8870857fe77b2a877c1122bafbc0b4320d4ed

Observation b76269b9-64a7-456d-af9e-e20f598c35dd · outbound

This paper cites write newline.

CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:06.111125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:06.111125Z digest=sha256:1a9a292c13ca04a8f4ea6518d169ab18231ba55317a0753e056fce66bbffadc4

Pith citing papers

Observation 568c73ec-2278-4e24-ac49-9a81921e4d2b · inbound

Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning cites this paper.

Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:00:49.792116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:24:46.023589Z digest=sha256:3d40be00838cd3d157ee1dd74674892de98db5c8339115f47890f9a45ad51fd8

Observation 4904fd2f-b77a-4838-b57f-5cab4c0c216a · inbound

Recursive Self-Evolving Agents via Held-Out Selection cites this paper.

Recursive Self-Evolving Agents via Held-Out Selection CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:14:37.523405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T11:13:15.923479Z digest=sha256:a0413a8ffbb711a5a3f698312d68457eccfa1449a681c2e538d2b7b197423f9d

Observation 1a7a4e19-1268-4d10-a232-8b4d46724ac3 · inbound

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search cites this paper.

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:55:41.046656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-01T06:02:48.532478Z digest=sha256:87bb3965aea3d17be889454d7cfee20151ab01dd60b5fdb0f5c3edb1718bac50

Observation da4b10f7-2315-4881-8542-3bc14fb3835f · inbound

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent cites this paper.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T04:44:28.494145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:44:28.494145Z digest=sha256:35b7833458508bb104ffffc4dc4ae1a0d10580e2477ab2a626e0bc8f1cf1548f