Pith. sign in

Paper Citation Record · LEDGER

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection

As of 21 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 1 inbound Pith citation observation for arXiv:2601.13735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.13735 v2

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:33:15.230764Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:38:32.269935Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-14T06:16:15.895014Z

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ac8f484-6880-4bfb-9e13-a17370e1c783 · outbound

This paper cites an unresolved cited work.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:14.880865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:14.880865Z digest=sha256:e3bc3c39f84941f422b1f8ae5b38be2d220e782fe29a0b943dcf9b6a07c39ac5

Observation 94caa034-92a6-4c85-b6fb-f76e8a92419b · outbound

This paper cites an unresolved cited work.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:14.984757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:14.984757Z digest=sha256:ae2717a5beed36b7720481721f37cef80e610e84cd014cc608bd4c5173cc347f

Observation ede44ebf-9efc-464d-94f7-81060dc46662 · outbound

This paper cites an unresolved cited work.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:15.039658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:15.039658Z digest=sha256:c2995f5a8cb5653410f346a06d055a3fc1ba8334c5b23fd75be15d98dc43e20d

Observation 63b49082-67c0-4478-98ad-06e20ee616cc · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:14.592921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:14.592921Z digest=sha256:8d1a9c588db7aafe2e1c2967014e85c73dc11e97ee580aaf755183d0fb629f97

Observation 61cb4f7a-2833-4d5b-b1f6-f3659f8cb6ed · outbound

This paper cites Qwen3 Technical Report.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Qwen3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:14.704610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:14.704610Z digest=sha256:1ecb25cbdfdc71fe577daa022a8702bdef25c93eeeeeedb880182003fd8d7f9e

Observation 30a2c15a-76bd-437c-8f33-00633af7694c · outbound

This paper cites an unresolved cited work.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:15.124755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:15.124755Z digest=sha256:8be3dbd57bce0266e619ce483805f1c9da4f119b6f7609d2a190e1f7130a947a

Observation b0258e07-98e9-4c8a-89f0-df89dd295869 · outbound

This paper cites an unresolved cited work.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:15.230764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:15.230764Z digest=sha256:39e4b8ca7eb326d1a6cdd3d1abc9827e92e4b542b403f3ef3c4847f1854c47e4

Observation be6a9301-f5cf-4ce7-a1b8-1999e448a798 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:14.360591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:14.360591Z digest=sha256:d55fc66f612189e3c1ffb4d8453ed7b4fe12579d2295adca0ab9f946731df444

Observation aab0740c-8623-4f61-8621-6b65062e2cc5 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:14.464964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:14.464964Z digest=sha256:bf344fa2d2455a6cb635c52230a5e3b2731c8ddfddfeaea025644778d04b6c04

Observation ba530d23-7f84-4ff1-ab3b-0c03622d7e35 · outbound

This paper cites Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yu- taka Matsuo, and Yusuke Iwasawa.

Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yu- taka Matsuo, and Yusuke Iwasawa

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T09:33:14.395571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:33:14.395571Z digest=sha256:ac00b5f65c4a43a92c3c8917fc7dd27e332ae890ffd462fa73c90a07332f8942

Pith citing papers

Observation 06450b15-ec85-4e10-bb5b-cfc1138584e7 · inbound

Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing cites this paper.

Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:38:33.518993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-14T04:38:32.269935Z digest=sha256:61fdabd82f9708392b3216416b3dad0a97a49ab8b5931c4491bbf7d57158162c