Pith. sign in

Paper Citation Record · LEDGER

Auditing language models for hidden objectives

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2503.10965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.10965 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:40:06.682295Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4523999c-11ea-4675-8629-f91213f4ebe6 · inbound

Towards eliciting latent knowledge from LLMs with mechanistic interpretability cites this paper.

Towards eliciting latent knowledge from LLMs with mechanistic interpretability Auditing language models for hidden objectives

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:40:06.682295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:40:06.682295Z digest=sha256:bed1c03817e243967137ceeb19cef6eb67a501dedc4feb7c2ef1fe692332d075

Observation aa202458-0103-447c-b666-20ea3ed08adb · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Auditing language models for hidden objectives

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:51.942324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:51.942324Z digest=sha256:df1ffe844d8236103b268babf0405d86ea1aec8a53576cebc4021980173ff34f

Observation 3a3bdf26-6371-4dc1-9092-d0dff245fa82 · inbound

Fine-Grained Interpretation of Political Opinions in Large Language Models cites this paper.

Fine-Grained Interpretation of Political Opinions in Large Language Models Auditing language models for hidden objectives

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:14.758765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:14.758765Z digest=sha256:7090ead66d37292cb72929768a4afded68d89e69ba093f37422693194b71a84b

Observation 2f079495-60a1-40d6-a715-75f588fa91e0 · inbound

Because we have LLMs, we Can and Should Pursue Agentic Interpretability cites this paper.

Because we have LLMs, we Can and Should Pursue Agentic Interpretability Auditing language models for hidden objectives

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:21.160673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:21.160673Z digest=sha256:4241bebbd624d15749f7f42cf3deb1bc5e04c817386e0103382f06ec9eb18731

Observation 9b88a1e6-611f-41e5-b5dc-cb04a4693e75 · inbound

Emergent misalignment as prompt sensitivity: A research note cites this paper.

Emergent misalignment as prompt sensitivity: A research note Auditing language models for hidden objectives

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:53:58.827203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:53:58.827203Z digest=sha256:c47ed5ac3f1d9a9c87bd4c4e2d48868f78bfadacdcc5ec1fb9b9321bb6bbcca7

Observation af0beebe-97c8-4cdb-83f8-ffbd766a73ca · inbound

Simple Mechanistic Explanations for Out-Of-Context Reasoning cites this paper.

Simple Mechanistic Explanations for Out-Of-Context Reasoning Auditing language models for hidden objectives

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:43.680460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:43.680460Z digest=sha256:de7492e49782b73a6b7cd5824d134bc2f5c10ad144fdfec582268f8be64feba1

Observation 70e3acff-d50d-4d41-8bde-f83b37011a63 · inbound

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs cites this paper.

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Auditing language models for hidden objectives

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:57:08.878310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:57:08.878310Z digest=sha256:3be3e73f4a8365e94c2363cab3ba49caf999984ee87fb3c0aed3b7bf5fdb618a

Observation 17683638-5c63-49f4-8d3f-25b356055bbb · inbound

Mechanistic interpretability for steering vision-language-action models cites this paper.

Mechanistic interpretability for steering vision-language-action models Auditing language models for hidden objectives

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T13:50:17.777613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:50:17.777613Z digest=sha256:d770087bab035057c7596012f189c15b04ee5d3fe25b861b63fe8930cff32000

Observation ed4d2bdb-a0b3-4817-bb93-0bade8384761 · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI Auditing language models for hidden objectives

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:26.835978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:26.835978Z digest=sha256:5d2ef0c1c44f3d9b2293c292ffb3db1acd2c673f8235b47cbf80520bdf2a1221

Observation 1f8e93d7-8c25-406e-ad64-0b7e8daba457 · inbound

Internal Deployment in the AI Act cites this paper.

Internal Deployment in the AI Act Auditing language models for hidden objectives

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.534920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T17:51:47.841707Z digest=sha256:80e3bd83a3e15434dae9e3f2f7609e755dc0e9f8356f37cb0f34f55b469cccbd

Observation e2285006-db6a-4264-814f-b9b95cfdf594 · inbound

Pando: Do Interpretability Methods Work When Models Won't Explain Themselves? cites this paper.

Pando: Do Interpretability Methods Work When Models Won't Explain Themselves? Auditing language models for hidden objectives

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T10:26:02.378367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:29:07.939420Z digest=sha256:fa1796c27dd4f58f0f4687707eafb7998f38ac74ca3791e28b5017ede5b80dea

Observation ecabf288-40ff-4f20-8eee-4284d200fcb7 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Auditing language models for hidden objectives

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.385565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:9de6cc02998f499474984f1931b1520b8862dec0a91b230be956b071bd4bc554

Observation 7919d241-d8ef-400d-a631-0f3936192d04 · inbound

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives cites this paper.

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Auditing language models for hidden objectives

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:06:07.697769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-09T18:47:41.188989Z digest=sha256:7603c256b16b968506f6b2b00d428b0fef85f1ae719d132e4640dac7af42284f

Observation 5ce426cb-d4e5-4ef4-960a-c6d84b54944b · inbound

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives cites this paper.

Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Auditing language models for hidden objectives

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:55:31.599646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:45:18.365192Z digest=sha256:8f9c4493c7c7cdc1cf990c2d1f9abd0a5fc607cfcedcf9a2d0f089e9952d113a

Observation 56fdb4a2-ef28-4146-ba4a-8e4e09a98f61 · inbound

Narrow Secret Loyalty Dodges Black-Box Audits cites this paper.

Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:57.070015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T00:53:49.010929Z digest=sha256:ba5519e49301a0cd0df3b3074e19c6e84f603f3bd8b47b8ff959fbe1dd9fd440

Observation 12f71361-56e3-40e4-b345-1c4336b2f5b3 · inbound

Narrow Secret Loyalty Dodges Black-Box Audits cites this paper.

Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.905126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T06:07:42.567241Z digest=sha256:7d60cb1e211dfad291ae2079fb8088551a942802b938a05f7a9c3948926863c9

Observation d3c0b841-123a-4bde-a163-775c7ade017e · inbound

Narrow Secret Loyalty Dodges Black-Box Audits cites this paper.

Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:05:07.322088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T23:02:20.906168Z digest=sha256:951b4b717bc7e63906f43b0489c5eca1b35b42c06be8070d31ae52781b12b5da

Observation 41e33e97-cf37-43ac-b5a3-0e315219a3de · inbound

Positive Alignment: Artificial Intelligence for Human Flourishing cites this paper.

Positive Alignment: Artificial Intelligence for Human Flourishing Auditing language models for hidden objectives

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:47.974215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-15T05:56:56.902705Z digest=sha256:d3950e03ab200e4b23ea3e4071d07ddeb613e0b5c5ab7724197ccbe5d2c7840c

Observation f9e9e36c-bd9e-4583-b4c1-7b602369d35d · inbound

Deep Minds and Shallow Probes cites this paper.

Deep Minds and Shallow Probes Auditing language models for hidden objectives

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.365710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T02:19:42.346071Z digest=sha256:f48c55d15616ca7611a2f6c94b6c7ae82aa67bfe07ac96f519cfa13a6d911f5f

Observation 3aa6eb90-0c30-46f6-9c73-13884c2f24f5 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Auditing language models for hidden objectives

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:56.941027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:2df4d891ee02842e49290184221f61a1ed13eb364d6e39bc32ffb46b37d9f3d4

Observation 5e6652c7-54d7-4f01-a9f2-a4d661157ebc · inbound

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands cites this paper.

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands Auditing language models for hidden objectives

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:03.932936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T20:53:04.274840Z digest=sha256:4686e96439a067c4ba89a8657eaa9776c51a92834360591760afbb68e1766bd1

Observation 3550bd84-1b41-4078-bf8e-2a075683e031 · inbound

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs cites this paper.

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Auditing language models for hidden objectives

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:41:21.485724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-22T09:38:04.387777Z digest=sha256:5aaa4d3833fa599d1fe7466b3dafdbff3b7008b156ba1906d677742eef038dc6

Observation 819b93fe-de95-498a-9cb6-748d9e3e96e5 · inbound

PRISM: Recovering Instruction Sets from Language Model Activations cites this paper.

PRISM: Recovering Instruction Sets from Language Model Activations Auditing language models for hidden objectives

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:08.024529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T16:52:02.948457Z digest=sha256:9c262576cd8f434d0c7239ed6029a6493918627cfbfb93b8c0946d204dd6c1ea

Observation 6c3f657a-63bc-48ee-90d3-6563d1ff96a7 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Auditing language models for hidden objectives

Reference 250

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.478939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:a44c2adbf06f3281d2e4533119db6fd007505876932f7c1929cbb8b2a74ea97f

Observation 4eadd4bd-af71-4321-9a60-26aa20c3d2ec · inbound

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms cites this paper.

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms Auditing language models for hidden objectives

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.294364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T09:51:16.969884Z digest=sha256:5e95262eaa0c164646b9f72f66df7e5b6bb19444a91fed75065d27297f2c8654

Observation 7be5d1e9-cb3f-4c8b-a787-ba932cf5e064 · inbound

RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue cites this paper.

RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue Auditing language models for hidden objectives

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:33.682490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:34:39.457798Z digest=sha256:c0939710210c0e15b33e535926e12414fe0a09ee2186238428a51b52cf820b7a

Observation 093f88b0-768b-456a-a7c9-f556eb286e15 · inbound

Self-CTRL: Self-Consistency Training with Reinforcement Learning cites this paper.

Self-CTRL: Self-Consistency Training with Reinforcement Learning Auditing language models for hidden objectives

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:56.002103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T01:38:48.296421Z digest=sha256:d50a9cf68f32d999bc10d22e78cc55a740941561b771c37c8bc50330f124ef7d

Observation 1444b25b-1bd5-4a3d-b6ea-21c396c7ad16 · inbound

Channel Location Constrains the Auditability of Subliminal Learning cites this paper.

Channel Location Constrains the Auditability of Subliminal Learning Auditing language models for hidden objectives

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.415684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T11:52:03.948568Z digest=sha256:19dd6ed8fc1304f222681bae5fdd0fe1e400d6d35d0c530c7a4c147df1928ed1

Observation f52bdd47-9f53-4350-a462-8c2dd4ed46d3 · inbound

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology cites this paper.

The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology Auditing language models for hidden objectives

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:07:08.170512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-02T15:57:48.589980Z digest=sha256:fb2b6c8ea9145d617cb482244f589ad9fa96c85140d3d083c25e83370cdcd19e

Observation 3b2d4f8e-a49e-47ed-9b54-1dc75367f11d · inbound

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric cites this paper.

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric Auditing language models for hidden objectives

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-07-15T05:51:54.267975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T05:51:54.267975Z digest=sha256:db1c8519f13cee0b4959e9bca83d520038ee2c5e4ffc4257308c0ec9aeabfb64

Observation 4956763c-4437-48e9-986e-a835920f2897 · inbound

GDM AI Control Roadmap cites this paper.

GDM AI Control Roadmap Auditing language models for hidden objectives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T06:53:53.237703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:53:53.237703Z digest=sha256:f080548851e1361bfe20bfd4a3aa9894735e08dd91ee3df1f2db67ac206c471d

Observation 333f4bf7-a9ae-4ad4-a719-b8d30c0475c0 · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models Auditing language models for hidden objectives

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:30.250045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:30.250045Z digest=sha256:05ae47f650094e08c3e7aff13260fef20b49f6f72eddc47627dc56fb8befaf48

Observation 5048bffa-aef4-4ffd-ad6d-f7f24bdf424a · inbound

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation cites this paper.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.407247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.407247Z digest=sha256:ed44c4feaa6e7d238d1695a0d991136e785a889d252613a43ec8335b7ec6ebb0

Observation 33c2b679-5937-4d5c-b354-19eda83950bc · inbound

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation cites this paper.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.413345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.413345Z digest=sha256:690e18c2fce46468e3d25b46d72be373008d42305c4f7ee537cd29e0b57bb415

Observation c8933e60-ff59-4a3f-8efb-652aa548b415 · inbound

Reference Feature Atlases for Mechanistic Auditing of Language Models cites this paper.

Reference Feature Atlases for Mechanistic Auditing of Language Models Auditing language models for hidden objectives

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:16.824307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:35:16.824307Z digest=sha256:a10b331f11ff99e36007e160145139e7c2031b9b8b985f0499f3b5eee8f0cef8

Observation 2e720ff2-65da-40cc-bb98-22df2db2e570 · inbound

Not All LLM Reasoning is Visible in the Chain-of-Thought cites this paper.

Not All LLM Reasoning is Visible in the Chain-of-Thought Auditing language models for hidden objectives

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T04:14:44.927414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:14:44.927414Z digest=sha256:7999cdce65001fb798395656ece976f10f43f8a37b3c025f5da45dfffd1f01fa