Pith. sign in

Paper Citation Record · LEDGER

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

As of 9 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.04899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04899 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:05:41.210502Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67c64fc7-7a01-438b-9011-c290c8f887c0 · outbound

This paper cites Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.879497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.879497Z digest=sha256:2824b31d11c96a7a8f2b188d7f9b0205eb3a5c031ba5388a26367dc1d0e9e4f3

Observation 7a7e6c93-a5b9-478e-b940-8ca88100279b · outbound

This paper cites Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.927074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.927074Z digest=sha256:51455eeb83f34496717b2da235298a286391aff9da4d3e8739a18c130c532419

Observation caba163f-e3e8-4a3a-9a87-3f11f97e151d · outbound

This paper cites Thinking Out Loud: Do Reasoning Models Know When They ' re Right?.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Thinking Out Loud: Do Reasoning Models Know When They ' re Right?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:37.956147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:37.956147Z digest=sha256:29e0f9907b819d926b89a78f233dc43438a7cfb1965e4b626c6553302e06580b

Observation fc2df6d9-8137-4ac7-b10d-8df1521ceb63 · outbound

This paper cites Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.004867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.004867Z digest=sha256:1784f3b5a4179ed6d938428ca8ca04b69ff205bfa5abf996fea8333b89754766

Observation f6c237d1-89ae-4f6a-b84f-a32fd0712050 · outbound

This paper cites C heck E val: A reliable LLM -as-a-Judge framework for evaluating text generation using checklists.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification C heck E val: A reliable LLM -as-a-Judge framework for evaluating text generation using checklists

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.057387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.057387Z digest=sha256:f0c2ce8a9f2f422ddca0940a47a9b78a4d08b4c1cd85b77e78e2d42860e75faf

Observation 8c082c08-0007-478a-97ce-a26110fa4f0a · outbound

This paper cites an unresolved cited work.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.102656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.102656Z digest=sha256:f2da1f05acce8b138deb97c798e5b2e88b474a0314979b5b06409531032c75a0

Observation 8424d035-3c56-4419-b41d-6805fe56d743 · outbound

This paper cites Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.149210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.149210Z digest=sha256:49421cf0b9d963d3dd2b8c85d6f2e26dd3a77fd3c02b47802088caf96b124594

Observation 43b691c4-2287-4b89-8772-193c3feeb604 · outbound

This paper cites Uncertainty Quantification for In-Context Learning of Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Uncertainty Quantification for In-Context Learning of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.184134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.184134Z digest=sha256:c6a7b45ff72964691bcfd8510f1ce7a9d8a457cce9a944fe3bc8441940ce9e72

Observation 0941cce5-dedf-45b8-8ad2-71979b193d44 · outbound

This paper cites A Survey of Confidence Estimation and Calibration in Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification A Survey of Confidence Estimation and Calibration in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.283292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.283292Z digest=sha256:9fd577204d24fadaeb3f0660b5f171d2b0dd9d40afe8e39ae1f7c68b0cb1197e

Observation d488a415-7af0-45ba-8ea0-c62dbd6cc024 · outbound

This paper cites B ayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification B ayesian Prompt Ensembles: Model Uncertainty Estimation for Black-Box Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.357788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.357788Z digest=sha256:426dfe8995ba47c0f7fa7bbafa7effa727c86a7a8e7b3188079cf69e614fdaf0

Observation 84f94c29-51eb-4379-8f71-c1d0d6811834 · outbound

This paper cites Calibrating the Confidence of Large Language Models by Eliciting Fidelity.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating the Confidence of Large Language Models by Eliciting Fidelity

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.423013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.423013Z digest=sha256:d00f91fee6f4438b207444c2b4138aac85648db0feb064a7bf33d6491645f684

Observation 09f6206a-bb22-4359-93d7-67e3bde20ca6 · outbound

This paper cites Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.523404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.523404Z digest=sha256:ab672a2038c5aee79b61c2962c2c448a00f8563e7e14d2d9d841cb3d4f98d5f6

Observation 22d28937-d08e-49bd-977c-ab61269cdb4e · outbound

This paper cites Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.634360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.634360Z digest=sha256:14c3cbe6f89c7c54f5d9e50897516e394c3ba9e2653467e231c9a504bd614730

Observation 9d454fba-fe3e-44e2-848e-c2ee5eb5ebcd · outbound

This paper cites Adaptation with Self-Evaluation to Improve Selective Prediction in LLM s.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Adaptation with Self-Evaluation to Improve Selective Prediction in LLM s

Reference 14

Resolution
verified exact
doi, observed 2026-08-06T14:05:41.484457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:38.744461Z digest=sha256:2de5a0f235afbb12e4101f489f5f94479eef980aa9d401f5f964013a26b1ece0

Observation e53c6dec-ae7c-441d-b73d-8f80b6d37099 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.842660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.842660Z digest=sha256:004d3e216266585ac6784c143aea96b3d6bb60d5dcdd87d4b78eeb04aacf812a

Observation c567c4c8-6bda-4b9c-940b-1bc8ea15e20a · outbound

This paper cites and Ng, Andrew and Potts, Christopher.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Ng, Andrew and Potts, Christopher

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:38.908813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:38.908813Z digest=sha256:cd3cad0d86e313cbd6c7cf401acec79eda764fb89a1f77e844e345dde16a9872

Observation f90ee083-453a-467e-a92d-8382fc2556a2 · outbound

This paper cites an unresolved cited work.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:05:44.292357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.020702Z digest=sha256:19b8639f5c70e85de18f13a7a0aec5989fd28cad017746b6982b25c1e9c31312

Observation 01d8053b-95cf-4bc5-8fe0-602a82ee3898 · outbound

This paper cites Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 36th International Conference on Neural Information Processing Systems , articleno =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.056971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.056971Z digest=sha256:8bdd010bd6ec43256949befe46245df58dabf6e35ce58d9d81f100abdcf59935

Observation d8db0919-2b11-47cf-9b1d-1aad27c1d093 · outbound

This paper cites 2025 , eprint=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2025 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:44.115273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.111242Z digest=sha256:efea354c0b10d720c1cd16ecc4dbfb52351063db52d75888430a32177255d67e

Observation 21498f79-42eb-4369-bd65-0d088b4bcf55 · outbound

This paper cites Transactions on Machine Learning Research , issn=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Transactions on Machine Learning Research , issn=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.163789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.163789Z digest=sha256:fd341f41cd44b659b44a66fc2f2e65bd2684a631d023c63c53968e3238ce327d

Observation d54ca5b8-118b-40d5-851d-12a2ad92a030 · outbound

This paper cites Calibrating Verbalized Probabilities for Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Calibrating Verbalized Probabilities for Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.202100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.202100Z digest=sha256:e1bf3393f5bed836228f90db26da99d81ae04c0a81f7009731841324ae85c22f

Observation 8483de6d-16eb-4114-8f3c-f00ff81c13e3 · outbound

This paper cites IEEE Transactions on Software Engineering , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification IEEE Transactions on Software Engineering , year =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.916227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.244339Z digest=sha256:b7459982e49e123b264d339676c01cb93a48cd304731ef5c6b0e1d2d334cd8a5

Observation 1ffeb2e8-126d-4899-9415-470b8ae2a15c · outbound

This paper cites The Eleventh International Conference on Learning Representations (ICLR 2023) , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification The Eleventh International Conference on Learning Representations (ICLR 2023) , year=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.730775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.260749Z digest=sha256:8eb1aabe06f3230c054b4eeb753272543a81742c9ed5ff28a8145c0e1bd2e7ed

Observation 0e39d061-dbc2-4868-b125-d5a1189a4618 · outbound

This paper cites 2022 , eprint=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification 2022 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.312295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.312295Z digest=sha256:992cb5b4505eae75daa53d7b90d5a947ca955f2c2ff5209fe20e4098c8aa517c

Observation c70988a7-752a-48df-8e71-efbbe0ed9fe3 · outbound

This paper cites Qwen3 Technical Report.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.330738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.330738Z digest=sha256:f3f9599008993350d5f067f58a8e21951f47c91119de7d43a4ab4fab0a7b44e7

Observation 37e4596c-face-48b6-bc6b-c340f95d6986 · outbound

This paper cites Journal of Machine Learning Research , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of Machine Learning Research , year =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.408396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.408396Z digest=sha256:fef29126f55c8e132edfe0af672c1de32247d80b47ad6c753fff231232158298

Observation e2e800fa-b6f8-40d2-b683-36cc8ff90bdb · outbound

This paper cites Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.546227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.484689Z digest=sha256:88e2b23b65efffcb1b46b66f9a10a6e96f71b93ce15de9f6ff186d65b891008e

Observation 7dc4f51d-b218-48ed-b037-d87c76c0ddbe · outbound

This paper cites BingoGuard:.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification BingoGuard:

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.390257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.551490Z digest=sha256:7808841e1973e0d66f28c566d31488bac6fb719ec262240a7add1c07c4998f8a

Observation b7e78a0a-5396-4e4a-97ee-a02988d68451 · outbound

This paper cites SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:05:42.357608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.618195Z digest=sha256:5f992d5e3ab80a1a50d920eb3e46e67b4e820a678cf8acb23e0ab8a397a06c78

Observation 71bfb71e-37b6-4834-ac8a-7c3b94e6a197 · outbound

This paper cites arXiv e-prints , pages=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv e-prints , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.243107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.720053Z digest=sha256:e3d18888d8caf0ee1e277243baac616705433450f7a1fb06ea0faee9495b76da

Observation 92c313d1-b75d-45ac-b315-01a4fc31fd31 · outbound

This paper cites Journal of classification , volume=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Journal of classification , volume=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:43.113941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.796050Z digest=sha256:41c271f8bba50bd63a34e251fbad3e9ab81182a54debbc9c21e17cad4290253e

Observation 7c47c8ed-00b2-41b4-a9cc-40c6a6c7ab16 · outbound

This paper cites Genome Biology , volume=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Genome Biology , volume=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.919907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:39.872590Z digest=sha256:adcb7ede70a0302340012d325aaf1651af714621f5c82892d3dd6d2249db2598

Observation 22dd0a2b-96f2-41d1-87cd-453080fe5b4d · outbound

This paper cites Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering , pages =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:39.938459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:39.938459Z digest=sha256:cc9d6f672d0efebc35ac11df45f2021c3408b63ad1a0f54a836d96af037061bb

Observation 8ecfb3a6-3148-4077-bbea-d14d1054e9b8 · outbound

This paper cites and Zhang, Hao and Gonzalez, Joseph E.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification and Zhang, Hao and Gonzalez, Joseph E

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.050280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.050280Z digest=sha256:d1a9bd8f9cfdb50be8d8e501ca88b33a280d427a2fa5c16503b957264a6bf14f

Observation 9f679a33-70be-4c48-81cc-a05595cbaf0a · outbound

This paper cites Proceedings of the third International Workshop on Machine Learning in Systems Biology , pages =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the third International Workshop on Machine Learning in Systems Biology , pages =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.099071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.099071Z digest=sha256:fb706cec02781e936c9e8787eb40c4612da9f62a71e5da0c700d8f58a2df1cdc

Observation 65764544-f0b3-448a-be91-daaf472f6bc6 · outbound

This paper cites Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining , year =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.739886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:40.173527Z digest=sha256:70a3665183b7e5879f65b66d848e9f74728850b55b22f536bbbce1d4f03139ec

Observation 02dcc27d-ea26-471e-9c6c-d554679344e9 · outbound

This paper cites I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification I Can't Believe It's Not Better: Failure Modes in the Age of Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.282309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.282309Z digest=sha256:17a609da857079a697f2ec5bc99fdbb0fdf3ff5982758776cf48c007b5b3373d

Observation db888785-34b8-4ad5-b510-cc87ad917766 · outbound

This paper cites arXiv preprint arXiv:2506.01734 , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2506.01734 , year=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.444848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.444848Z digest=sha256:d40848977cf3aba7760e4eb3ca2ee6bee4a433bdcf4064dd4294953d1921182f

Observation d99f4ffe-c164-4358-9fee-cf9e63bcdb3f · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS 2015) , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Advances in Neural Information Processing Systems (NeurIPS 2015) , year =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:05:42.607419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T14:05:40.534443Z digest=sha256:86f78633942da2b63e4bd776eb2be99bdbd2ffdd436206ddc39ec7e3efad4202

Observation 05d2a915-b4d7-4299-83f2-859488b06c60 · outbound

This paper cites Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.626084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.626084Z digest=sha256:a7f522e93ce44b97158148ede80d0ccce60c58056547af253479a292ef193156

Observation 6c5ccd92-3062-4f97-b229-241f5a76e48f · outbound

This paper cites , author=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification , author=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.784413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.784413Z digest=sha256:905d869fc5510534349388c4ae4b2477f8eb73c0a8298a23fcfb2c39ce38f7b9

Observation 104a98c7-f521-48ee-a043-2816272ac1bb · outbound

This paper cites arXiv preprint arXiv:2509.13813 , year=.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification arXiv preprint arXiv:2509.13813 , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:40.854578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:40.854578Z digest=sha256:093e71f4efc6486f91e35e9d850faaf67e0f7c79a297948258706be02f682257

Observation e46d0c1e-5bda-41ad-a4b4-31474d59b214 · outbound

This paper cites M onte C arlo Temperature: a robust sampling strategy for LLM ' s uncertainty quantification methods.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification M onte C arlo Temperature: a robust sampling strategy for LLM ' s uncertainty quantification methods

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.018426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.018426Z digest=sha256:1b150585360447a3308b7ce80f5feacae9a66b64f7279e94e5682ff1ac3db900

Observation 7c6fc369-868a-4849-aa51-723192ef9d2a · outbound

This paper cites Scikit-learn: Machine Learning in Python , year =.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Scikit-learn: Machine Learning in Python , year =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.097460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.097460Z digest=sha256:c3aa92e84d60d67454c52c667f451becfc80b118c183476ba4898c491317a511

Observation e9aff573-6c02-4c52-92a0-8b2142e1114a · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:05:41.210502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:05:41.210502Z digest=sha256:3fc2b30fde538e6775816bdfe8944801375fd8e9a98138d0005fc2c9a7be3071

Pith citing papers

No inbound Pith citation observations are available.