Pith. sign in

Paper Citation Record · LEDGER

Deliberative Alignment: Reasoning Enables Safer Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 90 inbound Pith citation observations for arXiv:2412.16339.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16339 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 90 of 90 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:09:15.269623Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03ec3792-32a7-4cce-bea2-5a5d78ca545d · inbound

Adversarial Reasoning at Jailbreaking Time cites this paper.

Adversarial Reasoning at Jailbreaking Time Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:49:08.728504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:49:08.728504Z digest=sha256:5ce9c8fce69ffb7d878c7f583e901e61ed818fdc247cc008092e8d1e42638287

Observation b145b221-3dee-488f-b0c7-06e3d27e9764 · inbound

BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation cites this paper.

BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T17:09:15.269623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:09:15.269623Z digest=sha256:de4eac03cba9f421ff245dd206d7b86dd1352c6d10f536ad4a32f7ec1e1e4ba9

Observation a8bba437-5849-475f-ab0d-338116b1d350 · inbound

Reason4Rec: Deliberative User Preference Alignment of Large Language Models for Recommendation cites this paper.

Reason4Rec: Deliberative User Preference Alignment of Large Language Models for Recommendation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T13:33:57.561753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:33:57.561753Z digest=sha256:58f20a26d133576eb72d05dd860e3fecf15ff88ec9985811c90e0e1b3abdd203

Observation db024240-ef8d-4c00-b3d5-d92dc678aa24 · inbound

Safety Reasoning with Guidelines cites this paper.

Safety Reasoning with Guidelines Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T23:50:35.737171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:50:35.737171Z digest=sha256:745cf2771fc602c40161cbfcc5246be07c77fd7a6b81b644d281fe49faf3d811

Observation 512edc13-4188-4aa5-84dc-413593bd86bc · inbound

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation cites this paper.

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:54.849494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:15:54.849494Z digest=sha256:52d2dc07b071c138ffc06df4acd23dde9b18cd08bfd27439c0842e7154125b9f

Observation 2022d310-a28d-4ce0-8134-f1ba8f434780 · inbound

MetaSC: Test-Time Safety Specification Optimization for Language Models cites this paper.

MetaSC: Test-Time Safety Specification Optimization for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.972604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.972604Z digest=sha256:e8cbb36b4749a3b8f7afabd23c366d941367a004974703081661d4a76bea2f10

Observation 23fc67e0-3355-4317-b1f5-14535894b972 · inbound

Compromising Honesty and Harmlessness in Language Models via Deception Attacks cites this paper.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.386584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.386584Z digest=sha256:2ff6930871a34db0655a41b2bf144b81812bb0624f7d07871935f56c1886dd6d

Observation 63c30fcf-e936-4b0b-bc26-23b50b978bfe · inbound

On the Promise for Assurance of Differentiable Neurosymbolic Reasoning Paradigms cites this paper.

On the Promise for Assurance of Differentiable Neurosymbolic Reasoning Paradigms Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:14:41.674716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:14:41.674716Z digest=sha256:3e64081f01f5f0b4af8daab134fb82d15b7a160e412b3b2157046e64e882f6ff

Observation 6ef9a4ad-f9b4-4eed-abe4-dde4c5d50bf1 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.165967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:268b32dee3e92508256d5bc87ccffcda13f1056ca6df35ee00cd52cd1a81b941

Observation 797b4cde-8920-4277-8e1e-9c7bd552206a · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.396795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:b0587592ccb1ea305aab5dafce452f17108e141760235af21e62794a01ed9e16

Observation 375fd8bc-f0e6-4c2d-84d8-c63872211c1b · inbound

Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance cites this paper.

Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T15:05:24.161001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:05:24.161001Z digest=sha256:511760e1b534bd076f8a72db77eb59824b1be01672a87fc91892c14205509493

Observation 2891e1a8-293d-4598-b105-b252aa409e06 · inbound

PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment cites this paper.

PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:00:26.038752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:00:26.038752Z digest=sha256:2a1140d219e159fc94ddb814a6e29705cf8a97d7ba43ad4004aea7c1ab19a09b

Observation b5cb668e-a047-436e-8f98-c11c3d112395 · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:24:12.967449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:718dcfc74004d9e1b7fe117aa45ac2185c02804ae92aa752f83c2c48149d2b35

Observation 736ec2ca-2f0d-4f7a-a40b-c03748cba596 · inbound

Adaptive Plan-Execute Framework for Smart Contract Security Auditing cites this paper.

Adaptive Plan-Execute Framework for Smart Contract Security Auditing Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:12.003241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:26:12.003241Z digest=sha256:626aad1287310502abb3bd6ad98c4a920587406dc2971430a4b7e8e89fd0e2ca

Observation a2b84897-36d2-4d2b-84ef-9eaac98ee83e · inbound

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation cites this paper.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.074711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.074711Z digest=sha256:24bf4e4753933e4e2d1fff5440d6416a19452b800695183be0519ff647415ac1

Observation 428190b1-bbfa-4d5d-bc45-1d9bbdbd8c59 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.474448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.474448Z digest=sha256:315b025336ac0ab8321b98880ade1cdff289787603c2f8a0fc64a63f3899bc36

Observation 7c674e3d-7334-4773-91d8-1318731fe9a6 · inbound

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models cites this paper.

Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:54.133916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:54.133916Z digest=sha256:bbe148e60e8c7814fa204c634eea04f606869a76281a22eb2353426ad372df72

Observation 45941abb-072d-4929-8b1a-afb08a0d0b05 · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.091365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.091365Z digest=sha256:f6ffc0f07e1afc393ffc3391acba6cadcc0d31f8da798bb4d77ca5ec97db9881

Observation a3316b4b-0177-4d54-a433-db6f8f99dade · inbound

Are Reasoning Models More Prone to Hallucination? cites this paper.

Are Reasoning Models More Prone to Hallucination? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:30.366747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:30.366747Z digest=sha256:0c8afc7defc1cde8f764e48aed5637b5ad5fbda7e68fd176ddac2e59eef68c58

Observation 73fccd74-5a99-4ea2-acf7-9c161fd76513 · inbound

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training cites this paper.

LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.860418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:25.860418Z digest=sha256:774cec920b8680625c14bdd8c79203561f6ce0642781787189f5661d5a7c78c4

Observation 87506f66-df2f-47ce-a3e4-7b48fdf1c126 · inbound

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It cites this paper.

The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:18.665783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:18.665783Z digest=sha256:2e2e50b8c5df1255fd3d1ddc30c8426da5395a0ee6c66af01c53870768e7f5b3

Observation 0764566c-4bae-4790-a32f-5248f95636f0 · inbound

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise cites this paper.

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:21.052638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:13:21.052638Z digest=sha256:74f351aeba711fa3c551cd7d7dadc231443a8320c70594f7eadf4e5dc7be02d9

Observation 9d95336f-2acc-4672-83ab-f461671f1728 · inbound

Lossless Token Sequence Compression via Meta-Tokens cites this paper.

Lossless Token Sequence Compression via Meta-Tokens Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:26.266557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:26.266557Z digest=sha256:090c87b0a2e747ed8e8e063412aaf70878086968d7f6e90c950ce2ae7653245e

Observation 0573b213-1033-4801-bd25-b42495bbc261 · inbound

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences cites this paper.

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:30.276451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:10:30.276451Z digest=sha256:f639b659386a6a75ff62a23a832fce7973cba858d48c25aa082f84e136e5a749

Observation 567b6b7a-433b-4046-8802-41b02ea292ce · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:19.235153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:19.235153Z digest=sha256:c120382de10ca2498ab0e13f2633268ce91747df411b733a33f042db948e65d7

Observation d690aabc-7d87-4f24-929f-34c3c953e1ee · inbound

SafeCoT: Improving VLM Safety with Minimal Reasoning cites this paper.

SafeCoT: Improving VLM Safety with Minimal Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:00.051825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:00.051825Z digest=sha256:e6797635613691a32639a67dc996930632b53d9c678fca9ed99c90139a115de9

Observation a952a564-86e3-4215-84b1-0e5002fc8b9b · inbound

InfoFlood: Jailbreaking Large Language Models with Information Overload cites this paper.

InfoFlood: Jailbreaking Large Language Models with Information Overload Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:02:28.553458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:02:28.553458Z digest=sha256:444f7e602b2b3b805164db0fd57a9f415a717a4608d1f9411d5e365556d4c618

Observation f2513c89-8d5f-42a1-a973-4811b3530d72 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:00.611082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:00.611082Z digest=sha256:dd0a338fc44399ef619f85f3bdfac0b4440906fc4f6a80e3f6ed544c1325c08c

Observation d78fe336-7b68-4781-a6de-0a99ff02065d · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:24.108617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:24.108617Z digest=sha256:7e411a76b3e0ff5809feb3badde8d165da9404e5efc67c6ae6244c23387a8de2

Observation e922bb64-c738-4675-aed0-39776862d27a · inbound

SAND: Boosting LLM Agents with Self-Taught Action Deliberation cites this paper.

SAND: Boosting LLM Agents with Self-Taught Action Deliberation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:19.985383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:47:19.985383Z digest=sha256:886e02b51baf300166b6699b9536707cd5521fbf438351f180db7d06825b8c50

Observation 209be614-bf2f-4e4a-8152-18d0bb36ce4b · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.986195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.986195Z digest=sha256:1b8e575099ecddbb6a8d8336e26e0f433755785bd1cb628544b22dcf37b2c0b1

Observation efeff014-0da3-4387-89e2-6fc3a28658e7 · inbound

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data cites this paper.

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:17.551777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:17.551777Z digest=sha256:d28b1e8f362bca26051796a575e8e6c512b49f3e3a3bfdf61847294ee52831e3

Observation 449e52f3-1076-4415-b5b0-64a3b80315c0 · inbound

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning cites this paper.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.386002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.386002Z digest=sha256:f41d99b24bd5c54f02d988d34a11c6e064bea0fca43fbb615d1a1a70c7aa220d

Observation 415de428-35f3-4565-9940-7ce12511a372 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 283

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.297537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.297537Z digest=sha256:960cce86b45b05902221e9b90b4594c004474361e6dabaebfb874f2f31a9741f

Observation 7b23d923-fb9a-45b2-9e6d-cf79ef6937c9 · inbound

Libra: Large Chinese-based Safeguard for AI Content cites this paper.

Libra: Large Chinese-based Safeguard for AI Content Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:38.663630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:17:38.663630Z digest=sha256:7a9ceec8bc8f92c21aa1c0fb3b68af0ce33418a12d40c6541dd4dcef69044c90

Observation a5a3d952-e6d6-49f3-a72c-8cf9204b488b · inbound

R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge cites this paper.

R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:18:24.466989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:18:24.466989Z digest=sha256:b92d2b8c18739d63b0c78ec4df76337ef92af133980d210ce4eda25dd5d73d96

Observation 078762c1-6d06-4376-85a0-197877d02b1d · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.538461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.538461Z digest=sha256:8ec62dfed2cf207c96bef54dbee93dbe7dfce3cb73fcbc8893e856d73d962d51

Observation ef1a2ba5-1e30-4db3-88ed-19856cf9274d · inbound

Towards terahertz nanomechanics cites this paper.

Towards terahertz nanomechanics Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:03:50.244026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:03:50.244026Z digest=sha256:5fe477f4ad23429b612b94e84bac6ce312aa86946bdca776a7d3f6cace29a32a

Observation 7ef559a2-84e8-49b2-81d8-62ea5359e245 · inbound

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants cites this paper.

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:33.591580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:04:33.591580Z digest=sha256:368bd1e3fdca31c93574f8f52e1c074e3ad663a7326e0b380fa1d35afc55da3c

Observation a1c86dd5-b54c-4c70-b9d3-856187e3ec0a · inbound

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI cites this paper.

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:23:31.962576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:23:31.962576Z digest=sha256:d03ac6500d60edf9f0e46dfb999de8867218404bba3d5c44e848883db721a476

Observation 11e7454a-7858-43d9-98a1-03d479d9b806 · inbound

gpt-oss-120b & gpt-oss-20b Model Card cites this paper.

gpt-oss-120b & gpt-oss-20b Model Card Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:22:54.679028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:22:54.633089Z digest=sha256:14603f33c64c4f0d936e6e87c6bc505cccf4c5acf1dc116a10faf4dbeff1bc54

Observation 80137bcd-470a-400a-b79a-4e705dcc8b32 · inbound

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement cites this paper.

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:20:14.188064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:20:14.188064Z digest=sha256:caa776882a43be744c10fdfe03c90a7a4616cc3a348f970b19df865f31d0bf9a

Observation 73b85185-a26e-4f6f-abc8-c828b1299b94 · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:54.701047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:54.701047Z digest=sha256:61b37abf0e174974696789647a0e1df718fbd19d4abebc03b2f1c57a068bc2ec

Observation 717369dd-b535-4414-a90e-5ed197e818ed · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.027988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.027988Z digest=sha256:510fdebb8b38b78d1bdd433133ce4454f36fd152e931488149953d6852e30b01

Observation f910b908-634e-4eb7-8108-e78f6441d5d2 · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:36:47.435169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T19:36:23.882344Z digest=sha256:2f09cdeb25d60491243a83f75191058f57a6680f21c10584f1cd5a8a64a04624

Observation 812cb2d2-f762-48b8-92ff-288ab0b62199 · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:15.617560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:15.617560Z digest=sha256:b154efe2f7db159d64247bdd6a18baae98be3a0da9824d0e1d151bb74b36dadb

Observation 7661fd46-165e-4792-9ccc-10ba8412cae6 · inbound

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents cites this paper.

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:49.835349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:49.835349Z digest=sha256:450fca99205b8e60f038dad3b45dcaff95698e6a29c0d4e58fd7eeb5b8113f04

Observation ce52139d-2018-4134-8251-1252b0e2ce92 · inbound

Reasoning Up the Instruction Ladder for Controllable Language Models cites this paper.

Reasoning Up the Instruction Ladder for Controllable Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:07:52.834565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:07:52.834565Z digest=sha256:349c9e5f5ffd7a7177ebf8c7c0ac655445d2ef9015538dcf081d4dfed43a6071

Observation 76760757-a24b-4d92-be01-fe907de2fc53 · inbound

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection cites this paper.

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:40:42.343205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:39:17.785715Z digest=sha256:a4653dcd82b3d2082aed85763fca9294ec42c9e317ad853192645243e53145e6

Observation b26cf182-2e61-46b3-a1f3-1629fa03bffc · inbound

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations cites this paper.

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T11:24:08.445919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:23:26.852673Z digest=sha256:1e5ac803d7516bb9e30583485130c8a416e2ddcadcf9aaf25be9d5d3ef18a3bd

Observation ed444e7a-71c3-4c47-a1fc-47ab4227d11f · inbound

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities cites this paper.

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:25:51.877085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:52:16.543718Z digest=sha256:33c20005329d31443c10718ff99022324071e46fd2d82c49816f977ad27dc55e

Observation 414ce9e4-f3f5-49b9-a2fe-8e74c4f7c7a2 · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:03.213083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:44:50.762366Z digest=sha256:b35259f5bd239a9ca2463e5d0b5e503ecaa4c39228fbe85870968d0f1c63bf8a

Observation 26d3d840-1830-414c-9e64-2f30f50431db · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:19:56.880861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:19:56.880861Z digest=sha256:f3bca9d61e7c19b061a0c3eb61d278c6b7a24eb2c8a800b3d4c48e794c9c2246

Observation f5fbc12d-ebfa-4525-a4d1-a21615b3988b · inbound

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language cites this paper.

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:22:37.099023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T08:19:50.909364Z digest=sha256:ad8b3e53e7cb1424525ad6c337cdd91646f4fd826cf8f280bd48b5200facd5b9

Observation 17fe05f8-fa6c-4f9a-957e-3e1ef5b721d5 · inbound

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories cites this paper.

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:10.098655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:48:44.687520Z digest=sha256:34e4f9f31fb0cce823e353711036f960af065075114a8bc3fb7628bbbeecbee1

Observation 27bc220b-5f2e-4b0e-8fcc-df1df3f8a7af · inbound

Reasoning Structure Matters for Safety Alignment of Reasoning Models cites this paper.

Reasoning Structure Matters for Safety Alignment of Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:04.339490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T02:36:57.093584Z digest=sha256:a01db373e803b8221e7accf44be7097187dc8f2bb2b49fef44a3ee5c14e8d2d1

Observation c8701d36-fcba-4409-85cc-d5bb4047210b · inbound

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems cites this paper.

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:18.981416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:26:04.152045Z digest=sha256:74f03a8080b19ef9940f54ddab427aec4b347345a0d467cbbeeb3a5a14d8fb40

Observation e64a7095-35c2-462c-b0ae-e9560148d988 · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:09.007290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:17:19.079380Z digest=sha256:b547a94174096607660f4570f91f649df0ed0f441a374b36c5f22c62d57debb6

Observation e132bfc6-f458-4968-bb02-2054c8aeab10 · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:57:31.670860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:c327e792b83caf6243ea829fc00932de3454ca23bfcd9969caaecce6ab545993

Observation 35d8e2ed-d827-4192-9da9-8eaedc0b0e00 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:08.576351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:6f0f44d2965dd938aaec377d3bf5aca1c759780d154eea339745a6ee3f6cb92c

Observation 479779e5-6147-404b-84aa-3e672a7f678a · inbound

Internalizing Safety Understanding in Large Reasoning Models via Verification cites this paper.

Internalizing Safety Understanding in Large Reasoning Models via Verification Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:51:14.318720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:50:59.283409Z digest=sha256:7ffe32789d63c6a662c0d0c3f819bb6edefceaa36993e42d87a02747b1bff6f7

Observation 272fc79d-4ef2-4ae5-ba07-222ec5920666 · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.921471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:2d1279a63b0fde6f1e91fb48f35de13ddadb3ffb246f0ffc96f2e11fb43ed0af

Observation 7f4d29f4-5a4d-4ba7-a7c7-0bbc2ddd70f1 · inbound

How Well Do Models Follow Their Constitutions? cites this paper.

How Well Do Models Follow Their Constitutions? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.874528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T15:33:43.569710Z digest=sha256:5bcac11571fece275292315031077ff9b4a8d3d5048124e42da3993887545bd3

Observation cdfbb531-120e-466a-b884-6941277f5414 · inbound

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection cites this paper.

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:35:50.610720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T00:29:50.433221Z digest=sha256:fbcb60d2eb794823a9e8151e4e2a14ff96ef4ca90411e6f5b512ed3c37f01c25

Observation ba830e08-2ec1-429e-ba58-0d5e4ad252e6 · inbound

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training cites this paper.

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.706723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T14:18:16.033063Z digest=sha256:45edc12d35594db0d7b764dc7db0e8e0c964c8abd821361420af70493f74ee8e

Observation 261b5b94-bf0f-4a3a-b4cb-96c188b31576 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.526096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:7eaed76616aaf73da29047cca904357f9c6001d138b3b11843606981735ed101

Observation 9967d366-91e4-4062-8624-8761f1e45cc6 · inbound

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation cites this paper.

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.124147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:35:21.388956Z digest=sha256:af47313d9297322e43b0e858c688bc8410109431388bb9a6339d0ed3fddd6bee

Observation 0a3bd05e-3e5b-42b0-866e-0586952688a2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 244

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.453603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:382752e6dc64842465cc8e63eed21c7a77420bd13116920bec56fc5fbd0b2570

Observation bd46bb0c-007c-4699-a9ec-b322b81061fc · inbound

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance cites this paper.

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:43.430699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:35:35.594085Z digest=sha256:51b4e45fd5fdc7e32147a1a57cbcfbd15cb1ebc45dbb1b333b6c04fe9bdd891c

Observation adb4d015-68ac-43f9-bbdb-4708714ccbe5 · inbound

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems cites this paper.

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:34.767260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:04:38.863201Z digest=sha256:ba94f24c07c7f9d750e77778db917b4c7f5715e91aaeeab9c0340015661ee551

Observation 043c0c62-8bf8-42ae-b1ee-ea97acc3187d · inbound

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems cites this paper.

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:34:36.560267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T10:30:03.047181Z digest=sha256:008b7be99a9a91185eb52f4e8e282f331e347b7db4f7a66cf915c84e7fae6af4

Observation 125b2c5f-13ce-4795-8bb9-9e81e6ad67a1 · inbound

Do Thinking Tokens Help with Safety? cites this paper.

Do Thinking Tokens Help with Safety? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:30:00.639214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T23:37:49.412578Z digest=sha256:f37ab94cfb21298ce2eb517e76818ee49aff23017884720ed433fd6a60acbf66

Observation 8e26d784-e1d1-4784-a47d-feedb03f575d · inbound

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models cites this paper.

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:06.722457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T21:09:19.727723Z digest=sha256:6dfe589736184d8ed3aae632ab75a7cbdc0ef0d656d940d6ff9e5ce94ac02a26

Observation 872dc00c-6b2b-4f47-8a70-0014c8a804a7 · inbound

Agent Safety Is Action Alignment cites this paper.

Agent Safety Is Action Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:35.001200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:50:45.759936Z digest=sha256:bc878f2b8d231ab5a8d7ee114dda9f38f559e94e03b1580956ffe5ec60918b38

Observation 1d0be5d4-c9cd-4fef-8161-6bc4d6f3d461 · inbound

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment cites this paper.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:06:58.846398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T12:57:00.805343Z digest=sha256:df10249bfdd386b2959a883680be90d2b6ae52a7c5791ad8bf3dd0666321d836

Observation 9b2a054d-3b7f-4ee8-bccf-a3d47cd92f47 · inbound

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment cites this paper.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:09290d815f593d5d18260d0058800475db6f3c9b3a6c8b9fdb916f9fd348ddc3

Observation c34d92c6-d355-4da5-a9f0-5f14463cc273 · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:54.934420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:9cf032c0ec2e7796f5cb3473af8ecae31286eaa06a937c5660dd439488a7f132

Observation b9561dc8-97ce-42f3-af4f-499781103678 · inbound

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models cites this paper.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:ba3d74041027e849cb9767043c682187ce8e05413191257106b28b3983f2f44f

Observation 5be75692-127e-4334-b317-c72f709d75c2 · inbound

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety cites this paper.

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-12T14:28:50.627444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T14:28:50.627444Z digest=sha256:de28252a468186278f73159a0f7a25db1102f92230df0baa9a5ca4702561e328

Observation 7c8e5ee8-1db0-4ee8-b32e-b1651172bed9 · inbound

Cost of Reasoning in non-English Languages: A Case Study on Japanese cites this paper.

Cost of Reasoning in non-English Languages: A Case Study on Japanese Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:12:44.286699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:12:44.286699Z digest=sha256:f6b0624674428ceb775c9c042d0c1e62258ea8e3c69277c82802a7cfd7e0c622

Observation d6aed98b-cb00-497b-991a-2288d69fe821 · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 230

Resolution
unresolved
no resolver link, observed 2026-07-15T08:35:47.870083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T08:35:47.870083Z digest=sha256:4573c7fe85847c4f0576cb5e62ebd6c8ef6a956a386151fe1b2d72c641f137df

Observation a01fcd39-8be5-4f88-84ba-1f363099a19e · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-02T06:46:40.391719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:46:40.391719Z digest=sha256:1afe1eb0f679c0cc74db4e27f1aa951f6a93e0495b73846a7af5e77dde3ab47c

Observation d406d612-cdeb-4712-a465-bb94445022c3 · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:23.274320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:23.274320Z digest=sha256:956f381063dbc511fffe00eaac3bb63dbfa0d6e528160947118a522cd577b769

Observation 4492457e-d1d8-4fe6-ac1c-0f87fa2c6a30 · inbound

A Geometric Perspective on Stabilizing Value Conflict Resolution cites this paper.

A Geometric Perspective on Stabilizing Value Conflict Resolution Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T16:35:30.952522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:35:30.952522Z digest=sha256:4e987d11efc1cc8f4da1fd3c3b18a65d6e7663c57f45b6dc4f1cf4ad6f4fbc85

Observation 79cea9df-6c32-4571-bcc4-7adcd1b3fb91 · inbound

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs cites this paper.

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T08:38:55.560417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:38:55.560417Z digest=sha256:aafb367b218cc59db6eec1fb72825fa566b570daa9eb27a6db9b055663e34442

Observation d4c5bf99-e663-4e71-8425-5902ed995ed8 · inbound

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models cites this paper.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.667721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.667721Z digest=sha256:480f10b405ed5f3a38f05ef1038f59db33ae2702076075d88eed1df47aec82e7

Observation 78a191ed-6886-4149-8e7f-11aeed5beb12 · inbound

Constitutional Midtraining: Content Presence Drives Alignment Gains cites this paper.

Constitutional Midtraining: Content Presence Drives Alignment Gains Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:35:00.225163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:35:00.225163Z digest=sha256:6bedbbcc3e2440e7cdfa42c1a186ba1016b790a85804eccf134daac82e79c30f

Observation db5790cb-4fa6-40dc-a93d-2b4355464f2f · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.100966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:26.100966Z digest=sha256:b0825af6bb8ca38fc5019a70235bb16c36fa4026a68b4123a673889db55b4b00

Observation 77fe7ada-8310-426f-8ddf-ff0fc104378d · inbound

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving cites this paper.

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:38:56.180832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:38:56.180832Z digest=sha256:168339d8e7d77614e33ee2cf04aba084e0d3c3d0c6a748edb2eb2a0d2b1e9efd

Observation 50f5ab3b-2cb7-4d9f-b6dd-cb0e3db16493 · inbound

AI Security Leaderboard: Methodology, Results and Minimal Standard cites this paper.

AI Security Leaderboard: Methodology, Results and Minimal Standard Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.841373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.841373Z digest=sha256:fee732f010eb8f8e8b1e679962854eb7382fbcb71b3c19ed30a11554893473c6