Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2503.10965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:40:06.682295Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 4523999c-11ea-4675-8629-f91213f4ebe6 · inbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Auditing language models for hidden objectives
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa202458-0103-447c-b666-20ea3ed08adb · inbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Auditing language models for hidden objectives
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3bdf26-6371-4dc1-9092-d0dff245fa82 · inbound
Fine-Grained Interpretation of Political Opinions in Large Language Models Auditing language models for hidden objectives
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f079495-60a1-40d6-a715-75f588fa91e0 · inbound
Because we have LLMs, we Can and Should Pursue Agentic Interpretability Auditing language models for hidden objectives
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b88a1e6-611f-41e5-b5dc-cb04a4693e75 · inbound
Emergent misalignment as prompt sensitivity: A research note Auditing language models for hidden objectives
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af0beebe-97c8-4cdb-83f8-ffbd766a73ca · inbound
Simple Mechanistic Explanations for Out-Of-Context Reasoning Auditing language models for hidden objectives
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e3acff-d50d-4d41-8bde-f83b37011a63 · inbound
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17683638-5c63-49f4-8d3f-25b356055bbb · inbound
Mechanistic interpretability for steering vision-language-action models Auditing language models for hidden objectives
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4d2bdb-a0b3-4817-bb93-0bade8384761 · inbound
Participatory AI: A Scandinavian Approach to Human-Centered AI Auditing language models for hidden objectives
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f8e93d7-8c25-406e-ad64-0b7e8daba457 · inbound
Internal Deployment in the AI Act Auditing language models for hidden objectives
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e2285006-db6a-4264-814f-b9b95cfdf594 · inbound
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves? Auditing language models for hidden objectives
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ecabf288-40ff-4f20-8eee-4284d200fcb7 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Auditing language models for hidden objectives
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7919d241-d8ef-400d-a631-0f3936192d04 · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Auditing language models for hidden objectives
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5ce426cb-d4e5-4ef4-960a-c6d84b54944b · inbound
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives Auditing language models for hidden objectives
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 56fdb4a2-ef28-4146-ba4a-8e4e09a98f61 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 12f71361-56e3-40e4-b345-1c4336b2f5b3 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d3c0b841-123a-4bde-a163-775c7ade017e · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Auditing language models for hidden objectives
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 41e33e97-cf37-43ac-b5a3-0e315219a3de · inbound
Positive Alignment: Artificial Intelligence for Human Flourishing Auditing language models for hidden objectives
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f9e9e36c-bd9e-4583-b4c1-7b602369d35d · inbound
Deep Minds and Shallow Probes Auditing language models for hidden objectives
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3aa6eb90-0c30-46f6-9c73-13884c2f24f5 · inbound
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Auditing language models for hidden objectives
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5e6652c7-54d7-4f01-a9f2-a4d661157ebc · inbound
Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands Auditing language models for hidden objectives
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3550bd84-1b41-4078-bf8e-2a075683e031 · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Auditing language models for hidden objectives
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 819b93fe-de95-498a-9cb6-748d9e3e96e5 · inbound
PRISM: Recovering Instruction Sets from Language Model Activations Auditing language models for hidden objectives
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6c3f657a-63bc-48ee-90d3-6563d1ff96a7 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Auditing language models for hidden objectives
Reference 250
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4eadd4bd-af71-4321-9a60-26aa20c3d2ec · inbound
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms Auditing language models for hidden objectives
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7be5d1e9-cb3f-4c8b-a787-ba932cf5e064 · inbound
RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue Auditing language models for hidden objectives
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 093f88b0-768b-456a-a7c9-f556eb286e15 · inbound
Self-CTRL: Self-Consistency Training with Reinforcement Learning Auditing language models for hidden objectives
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1444b25b-1bd5-4a3d-b6ea-21c396c7ad16 · inbound
Channel Location Constrains the Auditability of Subliminal Learning Auditing language models for hidden objectives
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f52bdd47-9f53-4350-a462-8c2dd4ed46d3 · inbound
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology Auditing language models for hidden objectives
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3b2d4f8e-a49e-47ed-9b54-1dc75367f11d · inbound
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric Auditing language models for hidden objectives
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4956763c-4437-48e9-986e-a835920f2897 · inbound
GDM AI Control Roadmap Auditing language models for hidden objectives
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 333f4bf7-a9ae-4ad4-a719-b8d30c0475c0 · inbound
Verbalizable Representations Form a Global Workspace in Language Models Auditing language models for hidden objectives
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5048bffa-aef4-4ffd-ad6d-f7f24bdf424a · inbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c2b679-5937-4d5c-b354-19eda83950bc · inbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8933e60-ff59-4a3f-8efb-652aa548b415 · inbound
Reference Feature Atlases for Mechanistic Auditing of Language Models Auditing language models for hidden objectives
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e720ff2-65da-40cc-bb98-22df2db2e570 · inbound
Not All LLM Reasoning is Visible in the Chain-of-Thought Auditing language models for hidden objectives
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.