Pith. sign in

Paper Citation Record · LEDGER

Process Supervision of Confidence Margin for Calibrated LLM Reasoning

As of 10 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2604.23333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.23333 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T08:19:09.437464Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T06:48:41.331963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact40
  • verified fuzzy43
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd1272f0-df01-4d45-bd9a-e081c8a22cd4 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning On-policy distillation of language models: Learning from self-generated mistakes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.604886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3c5edc8eee90b50e0c7baf9c30e108190bd4de62272bfb22bd38bd79756931f2

Observation 62a90c69-16d0-4d4a-a952-d2b07ca5febd · outbound

This paper cites Barber, R.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Barber, R

Reference 2

Resolution
metadata mismatch
doi, observed 2026-05-08T22:34:28.718729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:25a42af2db43cc74a8096727a6d681d029f55aa6c731f959bfd55f178d08b1ce

Observation 3b4e2a25-f8b2-4cb2-8cec-dd36660a7133 · outbound

This paper cites Conformal risk control.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conformal risk control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.632806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:20e6aaa4bf6f31af937c767697b4bc7534a48e14c273d9328a8613c5c6908cbc

Observation 225b676f-c243-467c-9aae-729aa5860759 · outbound

This paper cites Reconsidering LLM uncertainty estimation methods in the wild.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reconsidering LLM uncertainty estimation methods in the wild

Reference 4

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.669854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:72f20250c305c92dd7dac5f9c906b83932554d1e867b63cb027d021844f29755

Observation b1cecae0-2a6c-4066-9c2f-4771d76c1a43 · outbound

This paper cites Linguistic calibration of long-form generations.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Linguistic calibration of long-form generations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.599634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b70d8e44eb831fa5a76b263590134743abb81904449d53d142ef3a194d7f79a5

Observation c6414d67-98ad-45a9-b3ca-bde9d0fccd34 · outbound

This paper cites Rewarding doubt: A reinforcement learning approach to calibrated confidence expression of large language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Rewarding doubt: A reinforcement learning approach to calibrated confidence expression of large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.628017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:c2b8225587777793efea4fb32dca2e866dea7547dcd490bbe46cc7feb0aa2bc5

Observation fff23457-9f47-4199-a675-a01faf932cb1 · outbound

This paper cites Bereket and J.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Bereket and J

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.302769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3423cbcef9de0438a256b27c5acf42875696036dfea4c1554556210c1139fb1a

Observation 04ee208a-c0da-4018-878d-53383f949a62 · outbound

This paper cites Do Androids Know They’re Only Dreaming of Electric Sheep?.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Do Androids Know They’re Only Dreaming of Electric Sheep?

Reference 8

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.714791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:acd240539e1571f9d5789df0ce0b7fcfafddf18f645df167ee3536f4cf5f654c

Observation e45f9f31-1f6b-4c2a-9566-a3afb12a81c8 · outbound

This paper cites Uncertain Natural Language Inference.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Uncertain Natural Language Inference

Reference 9

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.677536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:c7edc3156a4ef6ddfafad71cdaa8d95db28bd38021f0384c7314e897ceaba972

Observation 194d0830-dfe6-4e88-8a46-9f14ad0fc725 · outbound

This paper cites A close look into the calibration of pre-trained language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A close look into the calibration of pre-trained language models

Reference 10

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.673639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b71bc10a8519c54c3005c787e993177752e184bcd5a84cd8b854ca68223f4629

Observation acad2e9a-1d91-407a-b682-c58cdc0cd67f · outbound

This paper cites Mind the confidence gap: Overconfidence, calibration, and distractor effects in large language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Mind the confidence gap: Overconfidence, calibration, and distractor effects in large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.625395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:659ce7568df8ebc1c4196fd36c208fc06f7f87fa1ea1cff9dbbfc25190a9b735

Observation ba73bd9c-b312-46b3-a77a-b3986201bb76 · outbound

This paper cites Evaluating language models as risk scores.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Evaluating language models as risk scores

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.597097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:00e986815f9412aa6ee80052728bcab46740234501e850cacf2ae9c118994e22

Observation 8118ff22-c021-447c-a758-fbe4a6bca34c · outbound

This paper cites arXiv preprint arXiv:2603.09309 (2026).

Process Supervision of Confidence Margin for Calibrated LLM Reasoning arXiv preprint arXiv:2603.09309 (2026)

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:41:12.062181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4087e086eeb42238f9f2816a4a7cd3244f7f2ca1ab31cf80b0fcfbd872e8d1fa

Observation a0df8c10-e56c-4d0c-9234-65c0e2dc5bd2 · outbound

This paper cites Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:22.282912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:710c3ce41543d3e3d36981fa5afb80dd4690cd39b75deedfc852168dc5de21ee

Observation 4274fba9-8265-405c-9bd1-0cd480105a2c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.070150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9ac24ffba39393cd9e2ec70a7f6ce625bce8415cc8e05db672d28dc143b9e92d

Observation d128a3c7-bd13-45a0-a906-8638ac224943 · outbound

This paper cites Calibration of pre-trained transformers.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibration of pre-trained transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.602127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:31b213a25ff8fcb7912b14c338a694f8784ef74362857b2f6ccee74b11d2b2d7

Observation 5c09aea8-7652-4c17-a01f-69ddf6ec0e34 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Detecting hallucinations in large language models using semantic entropy

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.607466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:a3e26adbc864c4d35c2074c44927689baed6b49c288b3aa294c394a0d9f5d7d3

Observation af936a36-0466-4beb-b767-25f7106813be · outbound

This paper cites Deep Think with Confidence.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Deep Think with Confidence

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:30:21.518780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:62d5e2ace3c7bcbecdd90f848d719788d9e20d625c73d6ba3dbc122510dceda8

Observation c9324866-5083-4271-b0cc-967a7d707468 · outbound

This paper cites Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.225348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:0e46285e214fb7072d17e6655b3266cd8f6a00779378d1d4203025fc64bf2f88

Observation 43982279-c7db-46c1-8238-4f257e4d80a2 · outbound

This paper cites On calibration of modern neural networks.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning On calibration of modern neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.630777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:884b437b49254fb117c2c0a34fdbb2300adc8b3aed5747cfcf9760ed2b6ec782

Observation d904375c-697d-4c78-bfac-a464c7c4f4d1 · outbound

This paper cites Language model cascades: Token-level uncertainty and beyond.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Language model cascades: Token-level uncertainty and beyond

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.589798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:55b556da6e91bfb545c99370d38df28ed426dbff40a22dc955ff08ffbc2256b4

Observation c4b1cfdc-ba02-415c-bf80-965992023719 · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 22

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.661171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:2e7b342455ca067734b39c3eaf3ae5fbda50404ca09d35c89261c08d123a7943

Observation f5ec1206-ee0a-46ff-bea3-076db50247b8 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Measuring mathematical problem solving with the MATH dataset

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.592482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:2b2d935f414e1ef44841c73f55c0fdc7f35080450d6f17716dd6df8405a0fdf0

Observation 23fb203e-c842-48a3-bbe3-dc713987d2b4 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Efficient Test-Time Scaling via Self-Calibration

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.347697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d22c972e161a661ae16b98810d782bb49ed5fd38cf1f29cfd91fbfc0e63dde35

Observation 97eb9b03-c697-44ec-bbf5-e44f25368210 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 25

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.706923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:d6de34edebd11ddb00c9b68559ebf0242dd9b0d5a5c578cf25937725fc28cb19

Observation 57394729-3a4e-46a8-96a9-9142b2c53f53 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reinforcement Learning via Self-Distillation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:29:18.795354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:481f3011eadb542eacb018efcdbd70af47afe54c4e2322a850373e7ab6ce1c83

Observation 209dc3a1-c0ce-493a-8d7e-bd530bf3520e · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.594691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:e3f03672d453699e954f08db0766a71f7d221f0edfa4cfc996b72e18e46a3092

Observation a4d3476e-ea41-4c22-b44f-09024919246c · outbound

This paper cites Calibrating zero-shot cross-lingual (un-) structured predictions.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrating zero-shot cross-lingual (un-) structured predictions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.609903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:eeaf0f71069985acbe361f0d20c8ab73f2fef6f76a91af67bcff7c2e2a82c672

Observation 272d0462-2136-49b8-8009-f159c4627da5 · outbound

This paper cites Addressing the Binning Problem in Calibration Assessment through Scalar Annotations.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Addressing the Binning Problem in Calibration Assessment through Scalar Annotations

Reference 29

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.665676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b985497a9e841b8eecc1b800c64c22a6a7515894ad1a2fb687b0ca5bab3d0038

Observation b26b2712-3bb0-462a-9470-2e43a2113618 · outbound

This paper cites Conformal linguistic calibration: Trading-off between factuality and specificity.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conformal linguistic calibration: Trading-off between factuality and specificity

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.623048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4655a74ba2beff75e44ebe303c5e916417d42040ebaf75749b9b90c2d0592d5f

Observation 17ab1cfd-b331-4c58-9cc3-cd6521422b36 · outbound

This paper cites Is that your final answer? test-time scaling improves selective question answering.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Is that your final answer? test-time scaling improves selective question answering

Reference 31

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.724549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:2cfff8d3579046a6e9f9f7e3a476a9711d4e62012348f49ed3cfb5fd5c687679

Observation ce38668a-50ac-4cdb-bc70-758d102a26fa · outbound

This paper cites Why Language Models Hallucinate.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Why Language Models Hallucinate

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:32:40.683342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:769a1ac10a84be8ab56b52ede27777dc3e7071511d606a30e4cf242f87ddf38e

Observation 8ac5acb3-0999-47cf-91ee-67809c2ad9e8 · outbound

This paper cites Scalable best-of-n selection for large language models via self-certainty.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Scalable best-of-n selection for large language models via self-certainty

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.614155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ae727f23397e8ac7d56bc969bdc73667f9ddd4a01314df2bc33bb1216c01eabc

Observation 16f1ae38-4509-4703-ba91-7747fa83ceb7 · outbound

This paper cites Large language models must be taught to know what they don't know.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Large language models must be taught to know what they don't know

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.616558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:99fe5d33023686b8a9e9085df9ca22ee76eb26d21fd5f4f9c60e18bfdaf3f86e

Observation bcee4fd1-bbbe-4aed-b8b1-c3d7b6ecc4da · outbound

This paper cites Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-27T03:05:05.764939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b4834f21da11ea0ab486ea24a1f3c2e8409960c2973a153f370db40db9d0ad19

Observation 97afc72d-bee6-4933-96a9-1af084c9ac6d · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:52:02.773759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:61ff575eeea419ff452f74bc91a0eae540754adc6a0fc5d40cc71ec3fc3ee925

Observation 1dee103a-586c-4888-8e3a-421aa1345968 · outbound

This paper cites Semantic entropy probes: Robust and cheap hallucination detection in LLM s.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Semantic entropy probes: Robust and cheap hallucination detection in LLM s

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.637437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:89c9f0f859071163469650499e0f69320de58944286267327ff61e35f37a503a

Observation 09986e8b-bda9-4e4b-a042-03608dbf0347 · outbound

This paper cites Think with moderation: Reasoning models and confidence calibration in the climate domain.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Think with moderation: Reasoning models and confidence calibration in the climate domain

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.576742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:5d04855f2ec1b0c839cd3dd77846ba5b9914e24d30085d1065001483dfeeefd7

Observation b1f51ffa-1a68-4259-bf99-b87b2264c5e8 · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.635129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:0f83a2f9bf31ed82bd5c53ebf4ed3ef1aa91e684767772de49fd172a95ca5505

Observation 77afc915-3c1a-4fc3-8d64-d3b92a4b9313 · outbound

This paper cites Taming overconfidence in LLM s: Reward calibration in RLHF.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Taming overconfidence in LLM s: Reward calibration in RLHF

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.579581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:900b1c7bd1e538111361237f0314053b8670c7c6f956a2417bed512db02319b6

Observation 082160ae-529d-4a88-befb-decc08150de8 · outbound

This paper cites Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:41:12.130965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:14d5d402409db9a411648a2e85913f190e50a1ecc958c1720d934535544d3a4d

Observation dd8714b0-c0d3-408b-babd-f7873d1e2c3b · outbound

This paper cites Conftuner: Training large language models to express their confidence verbally.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conftuner: Training large language models to express their confidence verbally

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.582151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:355900aebc10cc086d1fd625e1b9d4400c701a36026bee7d45dba1d45ba13d8f

Observation 2d407fa2-085d-4b4d-97f8-0d7402e37f9c · outbound

This paper cites Let's verify step by step.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Let's verify step by step

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.584677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:0450680a9636aac0afc47942ecec530579dd000da2ff51496d33e16b8bff58e8

Observation c87b0da9-7f0d-400a-9f90-cd8ba3d6a707 · outbound

This paper cites C 2gspg: Confidence- calibrated group sequence policy gradient towards self- aware reasoning.arXiv preprint arXiv:2509.23129.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning C 2gspg: Confidence- calibrated group sequence policy gradient towards self- aware reasoning.arXiv preprint arXiv:2509.23129

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.314348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:67f59d1ebeca396b5e16cfb6dd1e90a22305bf205b5035b1b6215e308a77e688

Observation 3058fae0-faa6-4bc3-b902-88b091090a92 · outbound

This paper cites C\ 2\ GSPG : Confidence-calibrated group sequence policy gradient towards self-aware reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning C\ 2\ GSPG : Confidence-calibrated group sequence policy gradient towards self-aware reasoning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.618713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:e60622e484fea3eb82ac83bea78d33d23d44af0e73cdafb4525fbf0ccdeee82c

Observation e3f73d66-006e-4162-b625-92efe674a175 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.620953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:c252c2cd7bc76b59e9d9b5f89859468f2e1a49e709fa9afd1ebac94205a00649

Observation 94d6c7db-a1ad-42e8-a479-bf9da999bdf4 · outbound

This paper cites Wong, Lidia S.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Wong, Lidia S

Reference 47

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.702470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:88f2033cba92abf4915a1de9c5c95d7ee0c532cd99151d72757e949c2df11d85

Observation f0c7e4a7-4d68-46ec-bdaa-445d282696b7 · outbound

This paper cites Uncertainty quantification and confidence calibration in large language models: A survey.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Uncertainty quantification and confidence calibration in large language models: A survey

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-08T22:34:28.695062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:fe923f39abd264d6868bf139f292ec865ac847919b7d523b71022f4285f95cf7

Observation c7a4a1e5-b2da-4a83-acb3-e32c4c20add6 · outbound

This paper cites Your pre-trained LLM is secretly an unsupervised confidence calibrator.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Your pre-trained LLM is secretly an unsupervised confidence calibrator

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.571118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ef46287a7cde5a0437e5fbf197604c062ab0042eecb6e89aa38595eb018b0cf6

Observation b11ff9b4-14cd-4bf4-af4a-9dc9f798cb6b · outbound

This paper cites Improve mathematical reasoning in language models with automated process supervision, 2025 b.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Improve mathematical reasoning in language models with automated process supervision, 2025 b

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.573802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:18946b1b2ed6b7cd5f08a3edb13db4b2e4bf0168fec79629aa28b8dba9e60f2d

Observation 9cce0c0d-49b6-4f68-acbd-e568ed75d00c · outbound

This paper cites Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.120451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9d06bc190f82aa5c07bcc1c2761f63c0d8c83ec72fb6fb01ca9ae6cb1d5a8423

Observation 56b1c3ce-564e-4dd5-8e15-9a1589a609e0 · outbound

This paper cites Reasoning about uncertainty: Do reasoning models know when they don’t know? In Findings of the Association for Computational Linguistics: EACL 2026, pp.\ 3408--3458.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reasoning about uncertainty: Do reasoning models know when they don’t know? In Findings of the Association for Computational Linguistics: EACL 2026, pp.\ 3408--3458

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.587369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:6ecc9a506c76b53559598a0810c7bd3810a7a48992be79a1508e0e813cb14420

Observation 042d080c-1e64-46f7-b973-e6b939b4385f · outbound

This paper cites Closing the confidence-faithfulness gap in large language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Closing the confidence-faithfulness gap in large language models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.139298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:30af92595f6aec662cc1983b600aa6ea2b93243cba95ae64575ad824184c1ad7

Observation 6e9acc8b-3d9e-49ec-9a54-1c0f3714c54c · outbound

This paper cites S1: Simple test-time scaling.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning S1: Simple test-time scaling

Reference 54

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.698725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9008aa9d596549a5c925802580d0756870f016be20796dd7a98aa636aeef2ee9

Observation afc0299a-d722-4046-8cff-7c92ce147dc4 · outbound

This paper cites When do LLM s need retrieval augmentation? mitigating LLM s' overconfidence helps retrieval augmentation.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning When do LLM s need retrieval augmentation? mitigating LLM s' overconfidence helps retrieval augmentation

Reference 55

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.721488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:cbf8eafc8c539f8cee56aae2efe207fba80bde0971d097a362b2a35722972da7

Observation 3f7fcfc4-3b30-4d0c-aa81-12789e3a1092 · outbound

This paper cites OpenAI o1 System Card.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning OpenAI o1 System Card

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.150846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:167ea3e6ab913f6b3fd9e1ef01384d98a69a23251358a68a2c515b10ad56ee62

Observation fc5695f9-3798-490d-b24a-2275385ab356 · outbound

This paper cites Obtaining well calibrated probabilities using bayesian binning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Obtaining well calibrated probabilities using bayesian binning

Reference 57

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.650353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:fbe5ab0dc80652ff7f1cfb40de0a220ea6db352a2a57d3fd4d6ba69034cc65d6

Observation 836030a4-e0c3-4e41-a2fd-dfa71c0c7075 · outbound

This paper cites Optimizing anytime reasoning via budget relative policy optimization.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Optimizing anytime reasoning via budget relative policy optimization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.568389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:1c9c39be69bc2a90b27c93082937296ea3e87ffe7bd0199a5cb234c09837286f

Observation f000ed63-aa98-4381-b60f-c7546516f89d · outbound

This paper cites Demystifying reasoning dynamics with mutual information: Thinking tokens are information peaks in LLM reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Demystifying reasoning dynamics with mutual information: Thinking tokens are information peaks in LLM reasoning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.612001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:be686063d0030c34bf33d0525202234d7bb0f27c388a179c9ed4eb724d7dd51a

Observation 63f45eef-7739-40cf-ae97-8b8b72d513f3 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.639483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3210ab41c13b7e7bd14ae09282086f35a865b404218675897aeccbb547079fe8

Observation dfd38fd5-50d0-4a01-b2cc-a2b5b3009daa · outbound

This paper cites Jaakkola, and Regina Barzilay.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Jaakkola, and Regina Barzilay

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.667772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:25a936f6042d06ce2220b740e62271ecf64a190b4a29581a636ff1794b5b3521

Observation 27357831-5c28-49cc-9e59-230ea2d037d3 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.669826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:57ebd6b50bc472843c481c29c31a7cfec968026c91df7769d2f15d782726d99c

Observation aa1f0fe6-6909-4013-956e-2eb5adae305f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Proximal Policy Optimization Algorithms

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.187773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:139fd2ba2c1e81d3619fc299b08b4974e1054d44a71519c455cc27bbf6bba742

Observation b2eedcdf-8de0-4752-a704-3c85b4f9af32 · outbound

This paper cites Rewarding progress: Scaling automated process verifiers for LLM reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Rewarding progress: Scaling automated process verifiers for LLM reasoning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.672070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f03e3c9f260bfe2b80fcea12a8f4eb6898eb318c7e4d831c7cc2a493a9a917d2

Observation 2607e0b4-fb40-4f1a-995e-5357094e4e07 · outbound

This paper cites A tutorial on conformal prediction.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A tutorial on conformal prediction

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.674355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3b825d14f6c61237fc5a6cce7816152939219d78b7cd0692aa314307557c55ae

Observation ce80752b-fd38-483a-8c1c-b25113cae32e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.211505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:260581cb1df07b9153da5b541bc2712b9bddd8d5e83119d1cd8b18f59c2ffd99

Observation d8352934-9bbe-4f7d-9cee-2fec2acacbbd · outbound

This paper cites Wornell, and Soumya Ghosh.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Wornell, and Soumya Ghosh

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.665602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:022bef38bc375dd8402e02ae055e90d94db11258f263ae445262f5a9a889a21a

Observation 705a36ed-57d3-4547-8498-a3a00013c4c4 · outbound

This paper cites Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.676674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4bb2d68ad704cc8d1ffb96f8d9fa9ac98a36a98e035b4c96dfa8b47ece152253

Observation 20160f7c-1ef9-4585-bc4a-6178ee6e8d61 · outbound

This paper cites Alison Noble.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Alison Noble

Reference 69

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.681327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ef42b42ff367755cc38a2a7f617042dcbf74c7d2ba7cbd8e3e3bde5bb9376579

Observation e499919b-990c-4cd5-a36b-7b3259358bef · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.330.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning doi: 10.18653/v1/2023.emnlp-main.330

Reference 70

Resolution
metadata mismatch
doi, observed 2026-05-08T22:34:28.689505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:041081b9ed2799aa06efc5a05a6f129361d325eb1b036a91b375e558e057a974

Observation 3d18eb14-21ed-43f4-a7a2-9535c6f8dc6f · outbound

This paper cites Calibrating Verbalized Probabilities for Large Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrating Verbalized Probabilities for Large Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.166451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b417f5ac18849de609b869d0dbb4e68d1177284af88c6945934f0ed9498a6bdd

Observation f29053d1-b153-4b8f-9b29-2e95abf12c62 · outbound

This paper cites Always tell me the odds: Fine-grained conditional probability estimation.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Always tell me the odds: Fine-grained conditional probability estimation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.678770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f1733e83700de3e3c180b9d90995af6a90fca3183a303eee87ad0a9ca59b1df3

Observation 70421b68-46a3-48ac-a7a3-2a368cf2f1de · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:34:16.071337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:847853c30b5c4e1602bd46c8da056caf996148228e25ae8050be3ccd93893236

Observation 00c0c9c8-38b9-4b76-bea2-d5827b685676 · outbound

This paper cites Calibrating verbalized confidence with self-generated distractors.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrating verbalized confidence with self-generated distractors

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.680939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:fa4a2d7e1c903d4b7490a04a36790daaa24c0dccafeaea56bb1a6591c60f2821

Observation 836ac16e-1846-40c7-85a1-89adc4d56b14 · outbound

This paper cites Conformal Thinking: Risk Control for Reasoning on a Compute Budget.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Conformal Thinking: Risk Control for Reasoning on a Compute Budget

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:43:11.628743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9cbfa5d9969549c9e029fb9eef050aa734e33d544069ca9f23f81b4f6695158c

Observation 9cc0b33f-ee76-43b3-a0f3-b0f029817e6a · outbound

This paper cites C on U : Conformal uncertainty in large language models with correctness coverage guarantees.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning C on U : Conformal uncertainty in large language models with correctness coverage guarantees

Reference 76

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.642387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9951feb9f66330dd4eb7ae9ab85c10624354272a98f0fa3cac28d439e5f70a5e

Observation 61e175ff-4b2e-480e-a8d2-7c638f20682a · outbound

This paper cites Thought calibration: Efficient and confident test-time scaling.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Thought calibration: Efficient and confident test-time scaling

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.663418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:a095e461ea68d3319c7b69cf67092e5c030ee83a164b2d74053ce0e6f495d284

Observation 785f3ed1-e96b-4b08-a9f1-d00e2cc542af · outbound

This paper cites Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Can LLM s express their uncertainty? an empirical evaluation of confidence elicitation in LLM s

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.658902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b5f23441a886541ba885bcbfd31b20b57aced36f9176f453205fedd1fb1e89a7

Observation 95aa9c4f-3655-48ee-9ff5-28c2c7d55180 · outbound

This paper cites Beyond correctness: Harmonizing process and outcome rewards through rl training.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Beyond correctness: Harmonizing process and outcome rewards through rl training

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.656468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:fbc60f581e97ed5b8a1fe3920c3f4b6537dab91a68ed7c4625f47cb5f252e557

Observation 166d5225-c877-4956-a144-ed5e3ffc80db · outbound

This paper cites LLM Probability Concentration: How Alignment Shrinks the Generative Horizon.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning LLM Probability Concentration: How Alignment Shrinks the Generative Horizon

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.261594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:6196e71a2eb6cbc3922f3e16e13133377c65606eacdc238b64635cfd7e78e23c

Observation 45eb69cf-7a3a-41b9-8ba4-87f4241e0d73 · outbound

This paper cites Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?

Reference 81

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.655110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:3b8e0deb1c239f2fac8ef07a194d137aa660e94fee00f3fb16446d474deb18d7

Observation 4bcc19fd-c171-4037-ac66-9ee1108858cb · outbound

This paper cites OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:12.177273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:36673ca65428b2e3f7a6bbd375265a34669064c90258a4e6cb80c4b2f281f8d8

Observation b7dca2aa-de39-4794-9998-c7cb861d1a73 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.159349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:62a2763a9aca5f9ca759bc99203d5d1ebfd14de5f4df70a2183f7e72d871ca63

Observation d8c2b988-e882-4782-9176-d703685d228b · outbound

This paper cites Reasoning models know when they re right: Probing hidden states for self-verification.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Reasoning models know when they re right: Probing hidden states for self-verification

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.660953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9e017c79be5c8213e9c9ef27755b3b0fc1a727d0d3ea72447e8ec7e3e1543465

Observation eae8f818-50af-405d-a294-635d846bbf80 · outbound

This paper cites GRPO - LEAD : A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning GRPO - LEAD : A difficulty-aware reinforcement learning approach for concise mathematical reasoning in language models

Reference 85

Resolution
verified exact
doi, observed 2026-05-08T22:34:28.685372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:386e46f1bedc42bc9a55e210c160811824c0818b83e8a53f16e73a82a4ebb3b6

Observation ac930bb2-b14d-4541-b5b8-aea21ce3fd93 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.648955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:24fe9d042eb7150aff6ff802217d541cacc088aaf842a44f0a7c57ed411b3469

Observation a154b8bd-8a7b-44ad-9d78-cee204d0317e · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:43:43.298530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:13aeec23fb7020ab51a46f61c005bfd2c9eda5beb3e9907d91aa3e00abca66f0

Observation 8ba7f292-78cd-4085-b09f-68b8f4eebc34 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:54:31.187112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:057a0a40b855e404d882f18ef3002519273293572c3547a97c7fd96382fd94d6

Observation 45abc3c0-fcdc-42c2-8ebc-8520eb949ac3 · outbound

This paper cites Calibrate before use: Improving few-shot performance of language models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Calibrate before use: Improving few-shot performance of language models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.643897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:124f5d6cb36e9498e987be6afe35c454bacb609135273e5284a3ff35cf9f80f9

Observation 5fdcce4e-9326-46b4-bab9-8b0875ca0526 · outbound

This paper cites A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.237411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:9f1da4e31e8bd85c1b6db273c28c4be00665c65364d49538348fe04fc76f54d8

Observation b1227ec7-72df-4e76-9641-87207889a73f · outbound

This paper cites write newline.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning write newline

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.646716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:ec80b94bb3287f6dce258a03284e84c87b2eec6341ac189cce3519d24158fd07

Observation bd5e107c-3794-471b-8f10-38d9c9321a18 · outbound

This paper cites @esa (Ref.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning @esa (Ref

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T17:07:40.651575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:fef5dc6c6303f5143faed33f5bc55a4cfdc24d3689e596f69ab1d3994ae79f72

Observation e5932a13-9554-468e-b6c1-c8f485b73438 · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.654239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:4db19e201951504c4d6493ac971fd739050ed0c28fd58ba582c22c1b493c0bc3

Observation a834bab9-d393-4e5d-bafc-03898712ac4d · outbound

This paper cites an unresolved cited work.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-05-26T17:07:40.641516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:f4da3c82f12f8abebfd280e19285622eddfb7932f629ecf85cd32c9579026e13

Pith citing papers

Observation bdca9a63-aad3-4285-995c-87118c8e3c97 · inbound

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute cites this paper.

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Process Supervision of Confidence Margin for Calibrated LLM Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T06:48:41.331963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:48:41.331963Z digest=sha256:f3d2f8df691e1de5b57535ff3c80af6e76a0c50cd8d7eece94fdb5de52924271