Pith. sign in

Paper Citation Record · LEDGER

Preventing Language Models From Hiding Their Reasoning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2310.18512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.18512 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:41:22.141140Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 06a3266a-2cee-4b34-8223-3dcf28092312 · inbound

Alignment faking in large language models cites this paper.

Alignment faking in large language models Preventing Language Models From Hiding Their Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T22:50:11.946185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T22:50:11.846863Z digest=sha256:a2cc18023cbeb744b0d370a2bda1a69f0fd2c5ac011afd030536471561ed07f0

Observation 42aceb86-a0bb-481f-b238-dfaa9b454102 · inbound

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking cites this paper.

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking Preventing Language Models From Hiding Their Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T16:41:22.141140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:41:22.141140Z digest=sha256:42aee33bd0cd5f16fd7980f54ca927583c2a90d6a96092f575cd58c0aea3204a

Observation 67be5f1f-89f9-49de-b9de-97969602ee1a · inbound

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning cites this paper.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.383501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.383501Z digest=sha256:9c8115a96dcde8e3143cc8cc08e737ace80e659cd7deec0c6433465d0f6f6179

Observation 31c74176-d014-4c06-b0de-dc539a7e1980 · inbound

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors cites this paper.

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors Preventing Language Models From Hiding Their Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:55.843136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:55.843136Z digest=sha256:d86fefcb922104c618d9d399c6ade6f82ef0edf1377c4513fda768ae70f38d23

Observation 1bf448f2-6dcd-4830-905a-704b7d8ece0f · inbound

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety cites this paper.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety Preventing Language Models From Hiding Their Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:19:44.761298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T14:19:44.695462Z digest=sha256:c392209af67a5a95f0f469be1c913d86b895aa79474bdc5d49b3ca810cbcbccd

Observation 1f21787f-b906-4a30-bddd-cd0a03fb9630 · inbound

Diagnosing Pathological Chain-of-Thought in Reasoning Models cites this paper.

Diagnosing Pathological Chain-of-Thought in Reasoning Models Preventing Language Models From Hiding Their Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:26:30.628642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:26:30.628642Z digest=sha256:50ac0b279c73c7a6d644d265b0bdf4bc30da9ef9de8ed92963d50a716cf7f85c

Observation 3d71b3f4-33e8-4484-be57-83039ffb90f6 · inbound

Detecting and Suppressing Reward Hacking with Gradient Fingerprints cites this paper.

Detecting and Suppressing Reward Hacking with Gradient Fingerprints Preventing Language Models From Hiding Their Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.572076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T08:18:37.665350Z digest=sha256:87d6671b1fbc43adf4436b4ed4c05b97b971efae1a9350828ab9266a01de3f4a

Observation bc293fa8-4e7f-4c12-b12d-7002267f0c87 · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Preventing Language Models From Hiding Their Reasoning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.393171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:fe47c6ce01b7d2379963406bda4144f00ec66190c85c9a981e5718b4ebf6925a

Observation 03f5e166-a971-4774-ad45-9f8e85f5c4a0 · inbound

Understanding and Mitigating Premature Confidence for Better LLM Reasoning cites this paper.

Understanding and Mitigating Premature Confidence for Better LLM Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:44.385114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-30T14:03:25.913615Z digest=sha256:1a21712abcf22f429f9de9c3fda1a16e62d1e7911f50e6ebd9724530a669ed69

Observation f752b928-caa1-4e30-a480-7f9f7a324065 · inbound

Conceptual Steganography cites this paper.

Conceptual Steganography Preventing Language Models From Hiding Their Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.621461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-29T18:46:39.761975Z digest=sha256:874e571dc0080104cc83ed33665debbc52603a006462464d500e70bfdd13bc3b

Observation 4ad341bd-0be5-4de4-80fe-cc441e544fac · inbound

A Note on the Strategic Confinement Problem cites this paper.

A Note on the Strategic Confinement Problem Preventing Language Models From Hiding Their Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.842944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T17:30:12.165074Z digest=sha256:777d66a163dc2f7040a9fbc850a65bce41aa4c49da1149d22e6d854f4819b2ca

Observation 122c54e7-21d8-478a-b498-880e09ae5697 · inbound

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs cites this paper.

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs Preventing Language Models From Hiding Their Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.423665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T13:23:28.061924Z digest=sha256:93ccefbeef22fc89183c4a4c846007f3b98c528906baded1657bc85305d57318

Observation e8a0d71d-82ff-46f5-ac74-7566df5fb004 · inbound

Comparing Linear Probes with Mahalanobis Cosine Similarity cites this paper.

Comparing Linear Probes with Mahalanobis Cosine Similarity Preventing Language Models From Hiding Their Reasoning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:18.923298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T20:40:03.169825Z digest=sha256:eea40648627baa28f8de299d1319df99e82f5c3ecee9cb88b51b4322b0917793

Observation 0b3af246-dbce-44ae-b803-3e0a0600999d · inbound

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs cites this paper.

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs Preventing Language Models From Hiding Their Reasoning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:50.102094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-30T00:38:21.949283Z digest=sha256:25885fa82cf10e8878363c0cbc89b8e34535ae85456ec29afd133a0c3f5ca970

Observation 3fd1c6e4-9836-428f-8115-7be3bbfda935 · inbound

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation cites this paper.

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Preventing Language Models From Hiding Their Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:50.791477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:29:50.791477Z digest=sha256:23bde9c6d8c533e67dae26c66b101bfb664e9072fa5245e4770c8413286380f0

Observation 9df42a77-34b4-467c-8d69-28d57c845de5 · inbound

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations cites this paper.

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations Preventing Language Models From Hiding Their Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:59.970123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:02:59.970123Z digest=sha256:d1c877ae8be5689d29cec159c28e21365f2f6b7709514f2a9e8f2036ad629b7a

Observation c7dbb06f-cfdd-481b-a3dc-6ebe56957df8 · inbound

Not All LLM Reasoning is Visible in the Chain-of-Thought cites this paper.

Not All LLM Reasoning is Visible in the Chain-of-Thought Preventing Language Models From Hiding Their Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T04:14:45.406064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:14:45.406064Z digest=sha256:bf4c495051b35e9d1bc6baf82d5905c0313987a0036b6c168bb818bd4f494d99