Pith. sign in

Paper Citation Record · LEDGER

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty

As of 4 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2604.17293.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.17293 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:42:36.727558Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T04:49:14.556768Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T04:55:58.814806Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact13
  • verified fuzzy37
  • unresolved0
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85e44662-b91e-4ba8-b48b-a28b511c749d · outbound

This paper cites The Twelfth International Conference on Learning Representations.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty The Twelfth International Conference on Learning Representations

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.233391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:9632302898046cae762e773b817e1751353a95f46d1f6ed8df729d0ebc92f045

Observation 71837a24-9a23-4b87-b6e2-88b6556a5cc3 · outbound

This paper cites Transactions of the Association for Computational Linguistics , pages =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Transactions of the Association for Computational Linguistics , pages =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.236442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:d7f100da11d4100566773b823350aeee716ab7dee92495f8e287ef4526b836d7

Observation 23f3c190-ce6f-495b-8268-bc4b6ad9b734 · outbound

This paper cites Do Large Language Models Know What They Don ' t Know?.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Do Large Language Models Know What They Don ' t Know?

Reference 3

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.750807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:621cb93560514a1580b57f310bfe393f9253c8ec4d6031401a3839256414303d

Observation 9c414fc0-6ebf-4a5b-9fe9-16019a6ce6e9 · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 4

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.752901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:bcc679e12b6fc784b9377189ce5a39fcd3bd5963fe1d2abe2573c0fcf56253a0

Observation 49e9ab11-7c1e-435f-90ec-d9b02f03e668 · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill? , url =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.222200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:36d6fb4e3a42d8870899db630b57ddc38cb70e65de9cb912c37a642cadc5c546

Observation 33c038ac-2933-4037-b207-7c7cbf8346ef · outbound

This paper cites Training verifiers to solve math word problems , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Training verifiers to solve math word problems , url =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.120762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:c3475ebd77422a037654dc7a52ec1c0852e7e00783145192f7f57156e8975701

Observation 2319da1b-874e-4832-ad6e-0f34e2ee59de · outbound

This paper cites Measuring mathematical problem solving with the math dataset , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Measuring mathematical problem solving with the math dataset , url =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.194925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:c9c85a0cd04a7555e2f0c60b6f7eecaf94328d6aae502e737b75d8aac47c7ac2

Observation b69d0bb4-3b19-4db5-9903-d5be98b25678 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Dapo: An open-source llm reinforcement learning system at scale , url =

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.206410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:e2528054582a20091787d9281003ff09307972376a6458475f4630830e86ae3a

Observation 7c5f3be1-b4cc-43c5-b77b-6edc45cc0b03 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Hybridflow: A flexible and efficient rlhf framework , url =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.188194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:5e34d63b561eb0934ccae3861f85eafaba357d5e8fc5ca9468615ef8eda4d7ed

Observation 22d20a86-8864-4c30-ab77-34ebf310d9e3 · outbound

This paper cites Metacognition: Answered and unanswered questions , volume =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Metacognition: Answered and unanswered questions , volume =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.185389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:0853f4c412b790989ca518fe25676b2f422547e903a544aa45b73a70879f2af4

Observation 23a6d4e6-37da-4fe8-800a-ed780d93486d · outbound

This paper cites Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Benchmarking Uncertainty Quantification Methods for Large Language Models with LM -Polygraph

Reference 11

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.756665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:7f401e4560ce7fc4ddc7ec624cceccc2d6320cc1d60e775db6a9346734394fcb

Observation 5396b990-1f2f-40d1-8e57-3ed427fd8196 · outbound

This paper cites Deliberative alignment: Reasoning enables safer language models , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Deliberative alignment: Reasoning enables safer language models , url =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.191424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:fa1a9d320b04c7cb7edca67946379fc110540da03e6f542cc9c553b3f282d2ae

Observation be5fcd18-f7c2-422a-a90d-c5d1ed4de737 · outbound

This paper cites Does Biomedical Training Lead to Better Medical Performance?.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Does Biomedical Training Lead to Better Medical Performance?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.209241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:87b601b7b3d22c3a475cb5a2a64c1b3c9576cafccfa9b7e39d194a611b5da70c

Observation ca6eea5d-877d-49ca-be8d-a06d510cee44 · outbound

This paper cites A Survey on Proactive Dialogue Systems: Problems, Methods, and Prospects , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty A Survey on Proactive Dialogue Systems: Problems, Methods, and Prospects , url =

Reference 14

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.754646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:ffc5248a27ba00fa5f00e6e1dd42081aaea8e24c91b6fb94e5791a872ab0b1dd

Observation 13d814fb-40ba-4f65-9371-3c465dd9bb5c · outbound

This paper cites AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions , url =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.215545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:522f4051f8e25e8d5be5e818e69bc65ada9cff7836ac8ca8d3182474d6237dfb

Observation b448555b-86b6-4d80-948c-a62eb834a9bf · outbound

This paper cites Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.198799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:87a46a56c35c659d4f09d57a2fa729c59f3deb609c528f8787b7231cbf96c57d

Observation bd54f8a1-2c3e-4a55-b976-afab52b33446 · outbound

This paper cites Don ' t Just Say `` I don ' t know''! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Don ' t Just Say `` I don ' t know''! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations

Reference 17

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.745047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:0cbd9d8111c6208dc79d9e532e94fe289f381f4a9c3ce6bbc2a62d55a57b3a49

Observation 5a750a5d-f4f2-4aed-a24d-dc633e63a640 · outbound

This paper cites The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability , url =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.175987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:2855631faccc45a6fd239fe71a500c83b042c9399a02371a70eee23113e3bcc3

Observation b4954075-0fa4-4195-8b74-2781713d46d9 · outbound

This paper cites Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning , url =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.179200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:59acabe9c78590dde235a02a61fd0bd187d5d717e0727312ed44a188d6f07dfa

Observation d8405eff-2bbc-42f3-a416-bd04eb246079 · outbound

This paper cites Search-r1: Training llms to reason and leverage search engines with reinforcement learning , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Search-r1: Training llms to reason and leverage search engines with reinforcement learning , url =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.173459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:23e82b5d4717ffab7584120455ad80190544da9e0c8f4a191d7059c4865cc787

Observation f85548d7-3ff2-47aa-b894-c1eb356f4af8 · outbound

This paper cites The Twelfth International Conference on Learning Representations.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty The Twelfth International Conference on Learning Representations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.165560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:581e5fd86e3077018ea31be4e0a8e8ca7c178f4cb3db515f32319c2e959b572d

Observation c96fccb1-66d7-4802-bcd0-97696597f03f · outbound

This paper cites Transparent and Robust RAG: Adaptive-Reward Reinforcement Learning for Decision Traceability , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Transparent and Robust RAG: Adaptive-Reward Reinforcement Learning for Decision Traceability , url =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.169263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:2fb933586bbff6f595e1ac14ded943ce405cff99c1e41334a861879d4ac20721

Observation 49730b9f-2186-4986-b120-066b89353d7f · outbound

This paper cites Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger

Reference 23

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.746910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:7ed4e0a7ed1e0ff21a3f2dc4848826fdd5527592103b7e01a27d50d9cb461500

Observation a6e2bdfe-6630-400f-ae44-e92f367c0a28 · outbound

This paper cites Edelman , bibsource =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Edelman , bibsource =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.182805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:fa1458f493619e0d805c173bae19980709e44a8703e90feb0e04489df2c9ccd2

Observation 60f2440a-ffbe-4df5-889d-776a23178a03 · outbound

This paper cites W i C ke D : A Simple Method to Make Multiple Choice Benchmarks More Challenging.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty W i C ke D : A Simple Method to Make Multiple Choice Benchmarks More Challenging

Reference 25

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.735441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:eb340ea44fe7c513eec6e9083aea0a1b72933544ede93d6d4f4d50507f766818

Observation a74d959a-a3f6-4784-85b4-0455d694b207 · outbound

This paper cites None of the Above, Less of the Right Parallel Patterns in Human and LLM Performance on Multi-Choice Questions Answering.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty None of the Above, Less of the Right Parallel Patterns in Human and LLM Performance on Multi-Choice Questions Answering

Reference 26

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.733511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:d20a3100befe4522314db3cc4dddcd53300b1c48f0b6e30e7646a94b2cdfb26d

Observation a6e4938c-637d-4b15-bf23-1ebb5a3c3f68 · outbound

This paper cites Asking clarification questions to handle ambiguity in open-domain qa.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Asking clarification questions to handle ambiguity in open-domain qa

Reference 27

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.748821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:37ac26ed4d17fd80c519c79b4da4bf1e2298c4bfc4f7767017316caeb4fff819

Observation 812ff6ba-319d-49d3-94b5-87f807d9ac35 · outbound

This paper cites CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty CLAMBER: A benchmark of identifying and clarifying ambiguous information needs in large language models

Reference 28

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.741092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:29f7fa60012a895c739f879809aa3fa3b9717a0dcc6ee5fb9ca27f0eadaa87ad

Observation a2b90ad5-642d-4103-97ba-b037972b253f · outbound

This paper cites Benchmarking Hallucination in Large Language Models Based on Unanswerable Math Word Problem , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Benchmarking Hallucination in Large Language Models Based on Unanswerable Math Word Problem , url =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.202158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:0c765d8aacb171ace80514f488342cd797036ea84cd1038338c5f007f9affdd2

Observation 63ed5d99-4c30-41ec-a831-9c1aa6a3ab55 · outbound

This paper cites ArXiv preprint , title =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty ArXiv preprint , title =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.239257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:d2417ab63352f4859cde420e015b3bbd8528951ecedf83af730f4370f1e9f456

Observation 295a598d-5844-449c-8b66-4bce6a1be88d · outbound

This paper cites an unresolved cited work.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Unresolved cited work

Reference 31

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.737178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:014ca06316665cd61e7c91fe4200d426cd365a940f49f5f40e11233422236c64

Observation c5122af2-06d6-41ec-acf3-0a71e202ff90 · outbound

This paper cites Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models

Reference 32

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.739052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:26f5b2f74593008f5eb7895da72ac8509681ed345982ec952e1a8de5510ec287

Observation 8ea51281-aa31-4d02-b321-dc8909ad625c · outbound

This paper cites Wong and Emine Yilmaz and Shuming Shi and Zhaopeng Tu , bibsource =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Wong and Emine Yilmaz and Shuming Shi and Zhaopeng Tu , bibsource =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.212296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:0036464f8fc9fb9d20700bba93c5b9ba8a31245ff4ef71e738e8e89c2ec0e2fa

Observation 83b270a9-b44c-4b6d-a7e5-1338bd9beb2c · outbound

This paper cites UR ^2 : Unify RAG and Reasoning through Reinforcement Learning , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty UR ^2 : Unify RAG and Reasoning through Reinforcement Learning , url =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.225199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:bf60cbef8ee89ebbba91560e1d79f64114d34a2070bbb7a183e414b128f62e0b

Observation 68edb052-361e-4e77-9f62-067c32e5888e · outbound

This paper cites Octothinker: Mid-training incentivizes reinforcement learning scaling.arXiv preprint arXiv:2506.20512, 2025b.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Octothinker: Mid-training incentivizes reinforcement learning scaling.arXiv preprint arXiv:2506.20512, 2025b

Reference 35

Resolution
metadata mismatch
doi, observed 2026-05-10T05:46:09.743032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:8d4744c9b58fa12e9882e60846842d95feb70456466c720c534a463f17836016

Observation ccbd8552-e1af-438a-aaf2-9b8c3d8056db · outbound

This paper cites A Survey of Confidence Estimation and Calibration in Large Language Models , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty A Survey of Confidence Estimation and Calibration in Large Language Models , url =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.149143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:244a826bb551b73120d6577f024fe5061c2ec8f0647c1f16b2d55c094914dd0c

Observation 1eb9e26c-009f-4832-a329-08af901a4f66 · outbound

This paper cites The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models

Reference 37

Resolution
verified exact
doi, observed 2026-05-10T05:46:09.731685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:ea0b2918b4f92809c9b41023e0af13b7fc6d0c1ce63ec2874893a7d5dcbb689e

Observation 15fe4bde-af2a-4418-bfba-1721456d8c4b · outbound

This paper cites Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution , url =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.151826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:a4b2c4fbf64663bd7dc6ecf179899bd77b351b9c6bed22a405a7b5ffbf9fdba0

Observation 3b3be7c2-dc74-4bb8-9556-e4d385af4da8 · outbound

This paper cites Grace: A generative approach to better confidence elicitation in large language models , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Grace: A generative approach to better confidence elicitation in large language models , url =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.155448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:4b5daedcb0f5303b32928bb9afb9de0b2347c98de1a03fc406dcbfa140768888

Observation d683f474-1ed4-4bad-9b63-a0e4fda7f145 · outbound

This paper cites Large Language Models Must Be Taught to Know What They Don't Know , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Large Language Models Must Be Taught to Know What They Don't Know , url =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.159836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:0adef0dc852f0bbec9625fc4a9f184bb7d928772c9d27f82462254b3f70cb038

Observation 93c4ccc0-d789-450a-bd7c-efb7c087aa36 · outbound

This paper cites Beyond binary rewards: Training lms to reason about their uncertainty , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Beyond binary rewards: Training lms to reason about their uncertainty , url =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.230589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:0c43901be88330bd2b45a0680ffa5af0b778659d58044124a948a595b0ee9c4f

Observation 619c0904-7de1-42cd-a857-8f1c989448cf · outbound

This paper cites Knowrl: Exploring knowledgeable reinforcement learning for factuality , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Knowrl: Exploring knowledgeable reinforcement learning for factuality , url =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.228295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:f29d1129a217a3d96cbf513c05daa44945a60d594f66f283e10bceccf329922b

Observation f7b2b212-5290-4817-89e5-429ed5460b0d · outbound

This paper cites KnowRL: Teaching Language Models to Know What They Know , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty KnowRL: Teaching Language Models to Know What They Know , url =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.138185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:e5a9afbbe3bca883cbb97cbf3210764c56a62ab29654404c6949487d79fde932

Observation ddeeffc7-24d3-4895-8bcf-6aae3391e7ca · outbound

This paper cites Qwen3 technical report , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Qwen3 technical report , url =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.141215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:047dfe820d60dfd461dce548c63ed01c29f812b8661c7b8ebb87388f5da0482c

Observation c4b6a4b2-04dd-4e85-aca7-3b8de40fd506 · outbound

This paper cites an unresolved cited work.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Unresolved cited work

Reference 45

Resolution
parse uncertain
raw_fallback, observed 2026-05-21T19:25:32.133395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:4d228bbbacd587f8cbc6782e594c98d7284d73a8e59613f830c0b0c074d17fc7

Observation 4a02d689-6749-4c0e-807a-9115a70636a5 · outbound

This paper cites GPT-4o System Card , year =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty GPT-4o System Card , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.131521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:a010b1d326a3cf3ade26c33d4e4009ccbc03007a4a0ed564bea6f75026b40db3

Observation bb7fd78b-caa6-4829-9394-6e1a451dcd6c · outbound

This paper cites GPT-5 System Card , year =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty GPT-5 System Card , year =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.135890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:84cad98051c9210ef644f43ab371eebd6620a7e008db4b31dd297098ee4ba6e5

Observation 4acb1771-73b2-4c5e-8890-cc036411b1ed · outbound

This paper cites Introducing GPT-OSS: Open Weights for Advanced Reasoning , year =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Introducing GPT-OSS: Open Weights for Advanced Reasoning , year =

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.144892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:4e3e1c633e395e0fba81968152c68c30fd3eba11ccc7cec49239a5546e5759a0

Observation daeb6169-2491-4e9c-84b2-186fd9d41663 · outbound

This paper cites The Claude 4 Model Family: Opus, Sonnet, and Haiku , year =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty The Claude 4 Model Family: Opus, Sonnet, and Haiku , year =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.163022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:697f83ca6d092045cfe20aea626c216e673ebb3f27ba111fb828d4d9fc3b4f32

Observation 7f3e52f7-392d-4ed0-99b5-63b7bebcbb50 · outbound

This paper cites an unresolved cited work.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Unresolved cited work

Reference 50

Resolution
parse uncertain
raw_fallback, observed 2026-05-21T19:25:32.129664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:91ab3b7a46170114fd4885f970e1bf677df5bf76779d58a52feaf8955c133606

Observation 0b0d5a55-2988-46b7-bb6e-7aab3567a462 · outbound

This paper cites Deepseekmath: Pushing the limits of mathematical reasoning in open language models , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Deepseekmath: Pushing the limits of mathematical reasoning in open language models , url =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.125306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:7d25402739eb78846328f98a3da80f552bbb7d373725d62fa7245020325dfd27

Observation e88358eb-36c5-46ca-9c47-95ce89ac8414 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena , url =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.127445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:d63b90834fe47c200dcd1142822a1ca8037012fa180079a7a16c6b556768e471

Observation df1ccd2b-3824-4e7d-a830-a199194013a0 · outbound

This paper cites Countering capability boundary collapse of llms in reinforcement learning with hybrid-policy optimization , url =.

Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty Countering capability boundary collapse of llms in reinforcement learning with hybrid-policy optimization , url =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T19:25:32.123021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-10T05:42:36.727558Z digest=sha256:c5a8fb49710fd8dc8a0940e3f5421cd398fdbce625f2cb819e28a06de63f3409

Pith citing papers

Observation aa905abd-b9dc-4bb0-800e-f217a1f6b396 · inbound

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models cites this paper.

Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:13:59.898066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:10:23.518054Z digest=sha256:e9fcba6543fa3c707a876f39fae16a2236fc2a9d35edc99f8423d4b311af1725

Observation fd47c197-873b-46e0-95a8-7adddc535d9d · inbound

Future Confidence Distillation in Large Language Models cites this paper.

Future Confidence Distillation in Large Language Models Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:55:58.816489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-09T04:49:14.556768Z digest=sha256:cd53559582ad9760ca9d26bda25b9856857ee4157ee517245ea418c30270cb48