Pith. sign in

Paper Citation Record · LEDGER

Preventing Language Models From Hiding Their Reasoning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2310.18512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.18512 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:41:22.141140Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 06a3266a-2cee-4b34-8223-3dcf28092312 · inbound

Alignment faking in large language models cites this paper.

Alignment faking in large language models Preventing Language Models From Hiding Their Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T22:50:11.946185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T22:50:11.846863Z digest=sha256:4995cc6a46389e0b49658ad0430d25f2cb1e4b80593a8431f4bb6f26d6f2a277

Observation 42aceb86-a0bb-481f-b238-dfaa9b454102 · inbound

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking cites this paper.

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking Preventing Language Models From Hiding Their Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T16:41:22.141140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:41:22.141140Z digest=sha256:42aee33bd0cd5f16fd7980f54ca927583c2a90d6a96092f575cd58c0aea3204a

Observation 67be5f1f-89f9-49de-b9de-97969602ee1a · inbound

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning cites this paper.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.383501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.383501Z digest=sha256:9c8115a96dcde8e3143cc8cc08e737ace80e659cd7deec0c6433465d0f6f6179

Observation 31c74176-d014-4c06-b0de-dc539a7e1980 · inbound

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors cites this paper.

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors Preventing Language Models From Hiding Their Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:55.843136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:55.843136Z digest=sha256:d86fefcb922104c618d9d399c6ade6f82ef0edf1377c4513fda768ae70f38d23

Observation 1bf448f2-6dcd-4830-905a-704b7d8ece0f · inbound

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety cites this paper.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety Preventing Language Models From Hiding Their Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:19:44.761298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T14:19:44.695462Z digest=sha256:fdbbb7a9d2d57f9b1ebcc8f3f42202e9f38180c7ee633eb6824fb67df73cc9ac

Observation 1f21787f-b906-4a30-bddd-cd0a03fb9630 · inbound

Diagnosing Pathological Chain-of-Thought in Reasoning Models cites this paper.

Diagnosing Pathological Chain-of-Thought in Reasoning Models Preventing Language Models From Hiding Their Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:26:30.628642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:26:30.628642Z digest=sha256:50ac0b279c73c7a6d644d265b0bdf4bc30da9ef9de8ed92963d50a716cf7f85c

Observation 3d71b3f4-33e8-4484-be57-83039ffb90f6 · inbound

Detecting and Suppressing Reward Hacking with Gradient Fingerprints cites this paper.

Detecting and Suppressing Reward Hacking with Gradient Fingerprints Preventing Language Models From Hiding Their Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.572076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T08:18:37.665350Z digest=sha256:348f3b50208ccc8ba22734a8e5f8891b323dcd2abc133868f26c02d3d2a0c89b

Observation bc293fa8-4e7f-4c12-b12d-7002267f0c87 · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Preventing Language Models From Hiding Their Reasoning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.393171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:b60d24a4e052c74a04f069b8296aa2a307ee6a84472a82ed49a9c9ff080ad026

Observation 03f5e166-a971-4774-ad45-9f8e85f5c4a0 · inbound

Understanding and Mitigating Premature Confidence for Better LLM Reasoning cites this paper.

Understanding and Mitigating Premature Confidence for Better LLM Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:44.385114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T14:03:25.913615Z digest=sha256:205d10c25e6357d0f9f458e577b0cb70f837f1db98aca16024fb7330789efa7d

Observation f752b928-caa1-4e30-a480-7f9f7a324065 · inbound

Conceptual Steganography cites this paper.

Conceptual Steganography Preventing Language Models From Hiding Their Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.621461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T18:46:39.761975Z digest=sha256:254b0e862a7e75a17efd566df8172021009b9d3d8e2574f2aaa6a93e6e34515c

Observation 4ad341bd-0be5-4de4-80fe-cc441e544fac · inbound

A Note on the Strategic Confinement Problem cites this paper.

A Note on the Strategic Confinement Problem Preventing Language Models From Hiding Their Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.842944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T17:30:12.165074Z digest=sha256:1bc7160a50542a1455fb6a1715f3a5be5f5abadec0ddf3bdd2267e31c7008fb7

Observation 122c54e7-21d8-478a-b498-880e09ae5697 · inbound

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs cites this paper.

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs Preventing Language Models From Hiding Their Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.423665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T13:23:28.061924Z digest=sha256:8b1be1697d105cfa4100c5fd47294900bc56c9b152cb2d5a4f7f9aaa9d5b9046

Observation e8a0d71d-82ff-46f5-ac74-7566df5fb004 · inbound

Comparing Linear Probes with Mahalanobis Cosine Similarity cites this paper.

Comparing Linear Probes with Mahalanobis Cosine Similarity Preventing Language Models From Hiding Their Reasoning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:18.923298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T20:40:03.169825Z digest=sha256:1539799e992393650e7630de180df2a09795068c1ca95316275ea1d987cb2f08

Observation 0b3af246-dbce-44ae-b803-3e0a0600999d · inbound

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs cites this paper.

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs Preventing Language Models From Hiding Their Reasoning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:50.102094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T00:38:21.949283Z digest=sha256:5d035fbebf904c84c35b8808c2cb8a7e6737d800a6745deee82598721368728a

Observation 3fd1c6e4-9836-428f-8115-7be3bbfda935 · inbound

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation cites this paper.

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Preventing Language Models From Hiding Their Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:50.791477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:29:50.791477Z digest=sha256:23bde9c6d8c533e67dae26c66b101bfb664e9072fa5245e4770c8413286380f0

Observation 9df42a77-34b4-467c-8d69-28d57c845de5 · inbound

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations cites this paper.

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations Preventing Language Models From Hiding Their Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:59.970123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:02:59.970123Z digest=sha256:76c022f8208db7b7db6282f19931da4a37059dffbb0afd9d3a9f65da70a0678e

Observation c7dbb06f-cfdd-481b-a3dc-6ebe56957df8 · inbound

Not All LLM Reasoning is Visible in the Chain-of-Thought cites this paper.

Not All LLM Reasoning is Visible in the Chain-of-Thought Preventing Language Models From Hiding Their Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T04:14:45.406064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:14:45.406064Z digest=sha256:bf4c495051b35e9d1bc6baf82d5905c0313987a0036b6c168bb818bd4f494d99