Pith. sign in

Paper Citation Record · LEDGER

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

As of 15 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2604.19809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.19809 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:37:41.353329Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:35:47.679312Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b9cade3-da4a-4a5f-8a95-e1f6a1c904bc · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T08:40:32.974213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:9df81801eba0023f1d69717489713e951c004f469f2ba028c8c6966dbcfd3b33

Observation 3423b013-5f83-4a29-b54c-1058854f55ac · outbound

This paper cites MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models MedCoG: Maximizing LLM Inference Density in Medical Reasoning via Meta-Cognitive Regulation

Reference 2

Resolution
malformed identifier
arxiv_id, observed 2026-06-01T02:02:27.350423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:940ddcfb37b88eb66cbbf7157539b2d5f73223dc707a62c3d82f39c2974e9219

Observation d6892b09-76c7-454f-bcdc-e748c4d6b1cc · outbound

This paper cites Each question has a unique identifier, domain label, subcategory label (5 per domain, 40 total), difficulty rating, question text, 4 answer choices, and a verified correct answer.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Each question has a unique identifier, domain label, subcategory label (5 per domain, 40 total), difficulty rating, question text, 4 answer choices, and a verified correct answer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.383028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:725d42531473fbbe9920634b784574eb5698ebb6be9d6ef6d4d1f18a2c4e764a

Observation eb81bfdf-64d1-4d6f-9aa9-4204e6de1821 · outbound

This paper cites fixed” — “tailored.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models fixed” — “tailored

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.388045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:e536f89e6560407b5bd41b9140d198eb9dfa5246dd4af669d73325ba118c16fe

Observation f2874805-4f77-44d8-b850-80460e13eee6 · outbound

This paper cites strong”—“weak.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models strong”—“weak

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.385458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:24f8d1ad04437d43e2a36ecc3a6ba20ef32ff04b881b359cdfbe13135ef2dbf9

Observation ebb146e6-1bec-4afb-8cc2-c098974efa12 · outbound

This paper cites an unresolved cited work.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-18T23:22:53.427283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:5aea8d6c6ad7a6d8d2b47bfbc1d9a47199a17a553a82f221daa80e1d9240a364

Observation 2c46ba44-9e2a-4796-9642-fbd5210221c2 · outbound

This paper cites an unresolved cited work.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-18T23:22:53.396683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:aa3268672bc48ad643a61562791501d453fd0a8ba754c14e930ee5ec53c2b19e

Observation 216edc78-548e-4b08-826a-6f569014b74b · outbound

This paper cites All claims are empirical.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models All claims are empirical

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.407714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:d2660c106bbef6a41b8360fca948406c5e50b01612f11bf841d0a68a874d4578

Observation 9b8870bb-ef28-48ce-bcd4-3b7c9e756f0a · outbound

This paper cites Section 3.4 specifies infrastructure requirements (∼8,000 API calls, temperature=0).

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Section 3.4 specifies infrastructure requirements (∼8,000 API calls, temperature=0)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.416170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:a9d4c8d5316a82faf393e72914816157cce2c602178cce10b8fa7cafe53cc4ad

Observation 69c5cf0c-7fd7-4067-a623-77d0f182bb36 · outbound

This paper cites Dataset files will be accessible there, and the Croissant metadata file is intended to be included as data/croissant metadata.json.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Dataset files will be accessible there, and the Croissant metadata file is intended to be included as data/croissant metadata.json

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.419839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:fb0ceeaa80cce34b3ee4defb9be37fe0ce2feaa37c4b47c0d73a320e825d8648

Observation 38747487-bfd3-4b5a-a811-057232132f6f · outbound

This paper cites All hyperparameters (temperature=0, wager scale 1–10, scoring rules) are specified.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models All hyperparameters (temperature=0, wager scale 1–10, scoring rules) are specified

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.422420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:8d8a8558839a3e9010d1c5dbd26ef58c149e187b800c032619c90623446f292f

Observation ac4a0f93-8426-4f5a-bf2d-4d75d8d2eba0 · outbound

This paper cites Effect sizes (Cohen’sd) and bootstrap p-values are reported for the three escalation curve comparisons.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Effect sizes (Cohen’sd) and bootstrap p-values are reported for the three escalation curve comparisons

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.424824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:a7d9d6ec4ffbd6742a97a5b85db78d304b8e119cd71edc15e2b067aa96666d2c

Observation a75ade1b-5e89-4179-abd0-f65849b48119 · outbound

This paper cites Infrastructure: NVIDIA NIM (free tier), DeepSeek API, Google AI Studio.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Infrastructure: NVIDIA NIM (free tier), DeepSeek API, Google AI Studio

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.410330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:9b1f4887b9d9ce98ea47621b8bb31e73369a34bc8f933f58e52cb0e6fe3983b9

Observation 8d7b48ec-7dff-48b5-bcd7-a81b5dd671b2 · outbound

This paper cites Questions are factual across 8 cognitive domains.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Questions are factual across 8 cognitive domains

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.407118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:ba624b33853e1b2b59cd0af69b0d82390954e64b13e395be7f259deb95607e24

Observation 8530fa40-2926-4ca5-a151-edab2bfd1935 · outbound

This paper cites Section 6.2 discusses Goodhart risk (models gaming MIRROR scores) and ecological validity limitations.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Section 6.2 discusses Goodhart risk (models gaming MIRROR scores) and ecological validity limitations

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.403677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:dcf80b2596c10355eb9fce92b431839eb618fbe63d6c867aadffaf9ac313ec47

Observation d012cbf9-5560-45ad-a922-515a3e56720c · outbound

This paper cites Versioned releases support periodic updates.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Versioned releases support periodic updates

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.412894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:aa7f7bea6a233f1af908d39f87511c1866aae5b0f19b8b1cb69c56630289f5d1

Observation 61f93bfb-696a-4252-a5c0-9fbc143e9882 · outbound

This paper cites The repository is released under the MIT License, which is also reflected in the Croissant metadata.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models The repository is released under the MIT License, which is also reflected in the Croissant metadata

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.395646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:e6965fd2f5d8d336a2c1322976549cd520d90d83fc31f4f024007badb5b95d31

Observation e96fbb64-904e-4de1-b916-571c69d5ce8e · outbound

This paper cites an unresolved cited work.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-18T23:22:53.390472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:c436d79e49a712af710ef0d2585e4c43b49cb0c0d382d8f939ad020eee9767f6

Observation 567ebbc7-feec-40b3-8e77-6814aaab0321 · outbound

This paper cites The human audit (Appendix W) was conducted by the authors.

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models The human audit (Appendix W) was conducted by the authors

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T23:22:53.410811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:37:41.353329Z digest=sha256:02b02f8647eecaa18fdcf29aeb372110025fffdad19bdd03f70a203aaa7f9aa1

Pith citing papers

Observation 182ee173-6991-4f5f-8031-52d41ecddc7c · inbound

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory cites this paper.

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:47.679312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:47.679312Z digest=sha256:5bfa02e41371df873b0ad32467b08d00cfb1226a7c8eedf026856bf00d7367d8