Pith. sign in

Paper Citation Record · LEDGER

Thought calibration: Efficient and confident test-time scaling

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 5 inbound Pith citation observations for arXiv:2505.18404.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18404 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:23.518356Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:10:15.553703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:55.947620Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0258f757-10db-4251-afb9-63c4cda29bdb · outbound

This paper cites The Llama 3 Herd of Models.

Thought calibration: Efficient and confident test-time scaling The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:22.674500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:22.674500Z digest=sha256:cf2413a27d19e64ec2627d666b5246afd6900e0bf0b771a7e474c74232cbfcfc

Observation 3bb155b5-1a51-4ea5-a88f-1d46bbf4f350 · outbound

This paper cites In The Twelfth International Conference on Learning Representa- tions.

Thought calibration: Efficient and confident test-time scaling In The Twelfth International Conference on Learning Representa- tions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:24.233758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:22.828122Z digest=sha256:660e197af56b82237300b161170ca1dd1409f03e855468d9b505b3cb6e2275a6

Observation 53af5fe8-d11c-409e-a8dc-794ceffe5efe · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Thought calibration: Efficient and confident test-time scaling Reasoning Models Can Be Effective Without Thinking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:22.927520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:22.927520Z digest=sha256:bbace094b8513660205197afa463c0cb34223d24f0aa3701dc5a298010b24f06

Observation 8df164f5-697f-43d6-b5ff-144f8a4bbc79 · outbound

This paper cites Calibrating Reasoning in Language Models with Internal Consistency.

Thought calibration: Efficient and confident test-time scaling Calibrating Reasoning in Language Models with Internal Consistency

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:23.112070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:23.112070Z digest=sha256:1678fdf5270ad24690b46a0138d5fff6e16834f2e5c6145bd493cc97abf4ec51

Observation d164c14a-5f06-4ae7-9b61-5609356657d7 · outbound

This paper cites Qwen2.5 Technical Report.

Thought calibration: Efficient and confident test-time scaling Qwen2.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:23.238646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:23.238646Z digest=sha256:442314a77ed19b1dfcfecf325880a2b7167d30eeb7e696a36c5267c4adfb2a99

Observation c4a2d088-2575-4444-8f3d-fda0fc9cc30f · outbound

This paper cites lmdeploy natively supports the saving of last layer representations, so it was used for almost all experiments.

Thought calibration: Efficient and confident test-time scaling lmdeploy natively supports the saving of last layer representations, so it was used for almost all experiments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:24.040421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:23.426563Z digest=sha256:b05ae30503ca26a3fa4a4abf6232aa78a7f3d030a52990eae12030c4f3fcc613

Observation 4f531f36-c6ba-4abe-8755-b4012a6a8f52 · outbound

This paper cites Due to computational constraints, we report the mean over a single run.

Thought calibration: Efficient and confident test-time scaling Due to computational constraints, we report the mean over a single run

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:23.859130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:23.518356Z digest=sha256:261b9bb597e826c39b9c41b2576a42c7fdb50d301f300af0664d43432cbeecc7

Observation f76a5ef6-a1d7-49d5-83ee-23c09141ce55 · outbound

This paper cites Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control.

Thought calibration: Efficient and confident test-time scaling Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:22.522797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:22.522797Z digest=sha256:c41c06e372ae0a07406e78158984e3fa2c68bb38d405bd3b7e34687c73eef88d

Observation d0b6a32a-9223-464a-bf6b-cb6ed13451a9 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Thought calibration: Efficient and confident test-time scaling Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:22.995288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:22.995288Z digest=sha256:35883bcc72f1455715ad2a8302dafa2159116f2ae9759f245245a701c71f64a9

Observation d9f1127d-cd47-4510-9e02-350d7df22f95 · outbound

This paper cites In International Conference on Machine Learning, pages 19274–19286.

Thought calibration: Efficient and confident test-time scaling In International Conference on Machine Learning, pages 19274–19286

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:24.430785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:22.737261Z digest=sha256:de62fd3f8bb179af7cf1396e9f01dfe74f9adec61f4d94c0e43f5d670779d5da

Observation 6cbbd8c0-6c99-4e97-a9e2-67089320a0a4 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Thought calibration: Efficient and confident test-time scaling Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:22.580355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:22.580355Z digest=sha256:3050cb6486c774e0cb4fb89e8151e71dad275072df1d46f3b1ad76de1441538d

Observation fde1f4ae-9b61-40ed-9d38-0bf9b2fe0258 · outbound

This paper cites in-progress.

Thought calibration: Efficient and confident test-time scaling in-progress

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:23.317044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:23.317044Z digest=sha256:3a4f6f7b22c92c2227e243fb6afdf0031e715f9fffd0051b8c6316d3e31f783d

Pith citing papers

Observation 0bb434ad-6b19-4129-a9c9-0794971002ac · inbound

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing cites this paper.

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing Thought calibration: Efficient and confident test-time scaling

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T19:10:15.553703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:10:15.553703Z digest=sha256:b0b852fa9556592aeb0466ad5a82e0df533b0ec6a638b9d45ea94ef5af226b23

Observation 5e3cfcf8-3ebb-458c-b49d-836e39674a95 · inbound

Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems cites this paper.

Reliable LLM-Based Edge-Cloud-Expert Cascades for Telecom Knowledge Systems Thought calibration: Efficient and confident test-time scaling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:48:32.150347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:47:59.128618Z digest=sha256:af61621303ba398771174d99c229bdd3cacadfe3332f3238043e9ed64641c263

Observation 6c06d2d7-4b16-4f51-88ef-f9fff249c9ff · inbound

When Should an AI Workflow Release? Always-Valid Inference for Black-Box Generate-Verify Systems cites this paper.

When Should an AI Workflow Release? Always-Valid Inference for Black-Box Generate-Verify Systems Thought calibration: Efficient and confident test-time scaling

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:02:50.281693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:00:55.658993Z digest=sha256:7418e102f4a97f83389648989d1a34fc66f54806f81fff2d60a01ab4992b6c68

Observation d9788a92-8813-4e2e-bb27-ca962f2faa01 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Thought calibration: Efficient and confident test-time scaling

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.949123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:5f0f5209571fa6b115278bfd8b31a074970d4db9e7b23eab1d62dd33ea7bb2e8

Observation 389dc7f3-29dc-46cf-aa59-77823004f54a · inbound

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents cites this paper.

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents Thought calibration: Efficient and confident test-time scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-30T18:50:46.081641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T18:50:46.081641Z digest=sha256:4255de600fe2cc9dd940479d21cdec432e7a819e52eacb766ce0925286c52c36