Pith. sign in

Paper Citation Record · LEDGER

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

As of 9 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 4 inbound Pith citation observations for arXiv:2505.16113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16113 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:11:46.806710Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:46:14.378320Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:29:02.503467Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 025f2365-18a6-4afb-a3c6-37b30ae1d0e9 · outbound

This paper cites Benchmarking Bayesian Deep Learning on Diabetic Retinopathy Detection Tasks.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Benchmarking Bayesian Deep Learning on Diabetic Retinopathy Detection Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.024634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.024634Z digest=sha256:9500c53337447aec266f1396a1fe29ff68b49cb9a41345ba3b8471063c8084fc

Observation 3e90927a-26d4-4e50-88ac-cace3e22a086 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no ques- tions.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Boolq: Exploring the surprising difficulty of natural yes/no ques- tions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:48.530797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:45.065298Z digest=sha256:094bf2865ceb9d26372a004d0b5a46d81ab522f64bc144ec5d91d56b39f21603

Observation b5589958-8520-4e0c-87cb-fe19a46bdc04 · outbound

This paper cites Detecting hallucina- tions in large language models using semantic entropy.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Detecting hallucina- tions in large language models using semantic entropy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:48.349186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:45.141259Z digest=sha256:52eaaba5d3b3d65d14d761fea82da797e6860d3e7913db8089890754fe02a74a

Observation 4ce3ef00-44c0-4115-8dc4-6964fd2cbedf · outbound

This paper cites an unresolved cited work.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.244890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.244890Z digest=sha256:7b539e8acc06e46041d26b2e147dea79db203a032943da73383a98ecfb82fe1c

Observation 9b0a6620-f0d5-4f18-abaa-cbe46a657613 · outbound

This paper cites Dropout as a bayesian approximation: Representing model uncertainty in deep learning.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Dropout as a bayesian approximation: Representing model uncertainty in deep learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:48.038983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:45.455818Z digest=sha256:4fef6067b7f32112e90a6ac69d0630a0adb9cfe5ff15db1f7fff1802ec73a6b8

Observation ca60c6f2-bc36-483a-a178-859e90297ca2 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.611899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.611899Z digest=sha256:d7e4f8e3ffcd558aad88a00f840405711e2db388918bb407a03b5ab4b7b33415

Observation f774ab00-ed1c-437b-a2db-8fe632399df9 · outbound

This paper cites Unsolved Problems in ML Safety.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Unsolved Problems in ML Safety

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.769975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.769975Z digest=sha256:e1b20ebd88a7e39811a9f4a7965930688b766412b7e5f4c4146b10d44daa29a9

Observation c975e134-c24e-47c9-b900-1294177dc146 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:45.904860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:45.904860Z digest=sha256:9cc743b536fa4e0f5f2da118b8b2f4f0fafb7178b316d1739456503517e7f99e

Observation 5828cb78-9767-4de0-8c0d-3e3500a323a1 · outbound

This paper cites Large Language Models in Law: A Survey.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Large Language Models in Law: A Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.036999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.036999Z digest=sha256:0f7373e06571aec3687b48a4c7161de5538544b0de9da4cff94cfa017d05edc6

Observation 7b62fc94-35ec-4a67-a773-6b93691d7c96 · outbound

This paper cites Simple and scalable pre- dictive uncertainty estimation using deep en- sembles.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Simple and scalable pre- dictive uncertainty estimation using deep en- sembles

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.965482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:46.093367Z digest=sha256:b9751f4cc315d95efeef288a04703274d3ccbef22906d9ae70f82cac4379b196

Observation b76735cd-c92d-4c4a-acb3-aeb5b18021c5 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.764858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:46.141626Z digest=sha256:10813ecfe92fa341312503a8126bd84df41773407475a079ec2b759e6e9191ca

Observation ce0d25f6-6fa4-404f-ba4b-9f994604945a · outbound

This paper cites Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.190902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.190902Z digest=sha256:b5011707244d31a4aaccedf509dcfd04f907033ce3632df5db62c95998b8544d

Observation 924ee281-803d-45e2-a9b0-863ded1c79fa · outbound

This paper cites ToolACE: Winning the Points of LLM Function Calling.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools ToolACE: Winning the Points of LLM Function Calling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.258369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.258369Z digest=sha256:7839b48183edf2b3190a986fdc3673356cd853cf46ed3013146aca5f59f29ae2

Observation a481744f-c47d-4a40-8368-4cac467687bd · outbound

This paper cites Uncertainty Estimation in Autoregressive Structured Prediction.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Uncertainty Estimation in Autoregressive Structured Prediction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.329808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.329808Z digest=sha256:cee0532886333c629de20d208385b812ff9a8dc2fad2158619923e1c9af4f173

Observation 82700843-72b3-487d-b72a-572eccacb412 · outbound

This paper cites Revisiting, benchmarking and exploring api recommendation: How far are we?, 2021.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Revisiting, benchmarking and exploring api recommendation: How far are we?, 2021

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.667536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:46.380342Z digest=sha256:3f11b564ff04c9d6dbc92dfaa4a8b88a712544c2862b55f9a3efff4bea83e5ee

Observation 77667ac6-9d8f-4b70-8998-c3723011a061 · outbound

This paper cites Tool Learning with Foundation Models.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Tool Learning with Foundation Models

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-07T15:11:46.423549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.423549Z digest=sha256:4b5495921a8ea847e27f5347aaedab2d383bf86a2ac9e7795d80e30f1d52a47b

Observation 0b188d76-8f2e-4a63-b6bf-411e9f8cd342 · outbound

This paper cites Tool Learning with Large Language Models: A Survey.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Tool Learning with Large Language Models: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.515086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.515086Z digest=sha256:a284e56c7ccc0ff4346a5575de77456aafb5dda501dcfba435b1830516547275

Observation b61fa36e-31df-4c1e-921f-ee98d4dfafdc · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Toolformer: Language models can teach themselves to use tools

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.495778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:46.583746Z digest=sha256:15ed62bf0f2eec1c01aa8f266a9add55705fc68bba2d7bcf8202d9c39309b6f9

Observation 9e6f5e98-a010-40b5-9936-c4bcbb102d38 · outbound

This paper cites Using the adap learning algorithm to forecast the onset of diabetes mellitus.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Using the adap learning algorithm to forecast the onset of diabetes mellitus

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.379534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:46.615512Z digest=sha256:3753ee46db2fc9ee86d58504f10f1487e1d53f33e409d25be18ed18052989469

Observation dfe20a7a-299d-456e-abd6-4c2c6bfd6d07 · outbound

This paper cites Question generation as a competitive undergraduate course project.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Question generation as a competitive undergraduate course project

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.248402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:46.677222Z digest=sha256:5d9b54855f8803072213482753dcac6a42c769e6a534383a967af999ee628865

Observation 0a38c0fe-5290-47c4-9a3c-13e795e38be4 · outbound

This paper cites Large language models in medicine.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Large language models in medicine

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:46.763294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:46.763294Z digest=sha256:dc5ff65824caca1600f99fb2cf92a08fbdde593b4f0fe84610fac1e8bbd7a206

Observation a17d443b-48c9-454c-a8a3-2e5785c0ee5c · outbound

This paper cites Toolqa: A dataset for 12 llm question answering with external tools.

Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools Toolqa: A dataset for 12 llm question answering with external tools

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:11:47.080564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:11:46.806710Z digest=sha256:faf3167ce6faae01619fcad81691d6f311eb5348e58bfd9f69702f79e2bc30a0

Pith citing papers

Observation bd986cb2-9269-48f6-b838-cb2b1614f789 · inbound

Uncertainty Propagation in LLM-Based Systems cites this paper.

Uncertainty Propagation in LLM-Based Systems Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:16:23.687803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:10:07.638483Z digest=sha256:71436bc4bd959cce8aa1117e8547cbada719f46dce791bea0943e1954cd6e535

Observation 18f9caa4-06c0-4e9a-8547-b00f52a419ed · inbound

ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift cites this paper.

ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:29:02.505043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T22:16:17.385766Z digest=sha256:773e923451d46bff0a469f5b2ead60492d02b145366a2eb495e17eec43032e5a

Observation a9b73959-b485-4a26-b3ed-ff9adbc85a61 · inbound

Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference cites this paper.

Diagnosis-Driven Automatic Repair for Agentic Workflow via Symbolic Inference Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T06:24:52.086585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:24:52.086585Z digest=sha256:5b5ffe3198a025fc3b12106e88bfa0348f09b5ad46c82413ed175b01a8b53ab9

Observation 10f62dba-8f8e-4273-a3e8-f9a326ee90f8 · inbound

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers cites this paper.

Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:46:14.378320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:46:14.378320Z digest=sha256:adc506237b9462a18cfbbb4543e8bbdab7eea56ad97e337c954e331e337ed210