Pith. sign in

Paper Citation Record · LEDGER

Deliberative Alignment: Reasoning Enables Safer Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 71 inbound Pith citation observations for arXiv:2412.16339.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16339 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 71 of 71 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:14:26.266557Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6ef9a4ad-f9b4-4eed-abe4-dde4c5d50bf1 · inbound

MoBA: Mixture of Block Attention for Long-Context LLMs cites this paper.

MoBA: Mixture of Block Attention for Long-Context LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:15:46.165967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:15:46.085555Z digest=sha256:f9c9ad301dbcac6726cf41736402c7554c23f8a829bf0aab4359737503d0610f

Observation 797b4cde-8920-4277-8e1e-9c7bd552206a · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:27:21.396795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:b2c1c169e2e92333a3a249890ebb552b261a5e2576531dfbcf366fb485defef6

Observation b5cb668e-a047-436e-8f98-c11c3d112395 · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:24:12.967449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:7ff7ceaad6259ef4d1a50f28b7286e3b0a736103321c746fd38193baed0e6478

Observation 0764566c-4bae-4790-a32f-5248f95636f0 · inbound

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise cites this paper.

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:21.052638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:13:21.052638Z digest=sha256:c4924e658c9c6ac05b14d3d238258a3e8c4ba3b59a398ec725afa62ceb8e71ee

Observation 9d95336f-2acc-4672-83ab-f461671f1728 · inbound

Lossless Token Sequence Compression via Meta-Tokens cites this paper.

Lossless Token Sequence Compression via Meta-Tokens Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:26.266557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:26.266557Z digest=sha256:2df4180fb3a0b99103b977ab324b590703cdde96e1960623a32cd6a603b10dd6

Observation 0573b213-1033-4801-bd25-b42495bbc261 · inbound

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences cites this paper.

Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:30.276451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:10:30.276451Z digest=sha256:db4679526ec0f481312ea62e3e1be56b9a8154901059c7eabc8c37cb1b187425

Observation 567b6b7a-433b-4046-8802-41b02ea292ce · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:19.235153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:19.235153Z digest=sha256:8ffaaa48042e7536fdd9627539258515d5eab197e4cf15fd421b7a3251a5a4b7

Observation d690aabc-7d87-4f24-929f-34c3c953e1ee · inbound

SafeCoT: Improving VLM Safety with Minimal Reasoning cites this paper.

SafeCoT: Improving VLM Safety with Minimal Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:00.051825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:00.051825Z digest=sha256:344f8d84bb9e6881f43d2f1fa27b37d424387b99af6df0134b0ad158529bd8ed

Observation a952a564-86e3-4215-84b1-0e5002fc8b9b · inbound

InfoFlood: Jailbreaking Large Language Models with Information Overload cites this paper.

InfoFlood: Jailbreaking Large Language Models with Information Overload Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T01:02:28.553458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:02:28.553458Z digest=sha256:e008cbc240234e7e801781075f461d5b1f18dc3363017d2aa4159af376d783bd

Observation f2513c89-8d5f-42a1-a973-4811b3530d72 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:00.611082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:00.611082Z digest=sha256:24720616b1082fdc6e83e094ba7540b313db49bb53b479e50c807d80ec683c65

Observation d78fe336-7b68-4781-a6de-0a99ff02065d · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:24.108617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:24.108617Z digest=sha256:adc237ec8b103ac145c02f47838fd42cc8f61569753c99cc4a679741e4afa1ec

Observation e922bb64-c738-4675-aed0-39776862d27a · inbound

SAND: Boosting LLM Agents with Self-Taught Action Deliberation cites this paper.

SAND: Boosting LLM Agents with Self-Taught Action Deliberation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:47:19.985383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:47:19.985383Z digest=sha256:f7c0a664ea097b166ed49f48ed36e2fb2e9c77a8f6b5cd02f62c5ced5c5c7d1a

Observation 209be614-bf2f-4e4a-8152-18d0bb36ce4b · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:17.986195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:17.986195Z digest=sha256:8d8d6811d21ce39a11600499f6b7bd53c53fe41bd91a8c0c23f6d5400463a7ac

Observation efeff014-0da3-4387-89e2-6fc3a28658e7 · inbound

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data cites this paper.

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:17.551777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:17.551777Z digest=sha256:65da1622f678e89ba42970259f96d5d389337ad95ffb51c1cdc8f8b6eba8f910

Observation 449e52f3-1076-4415-b5b0-64a3b80315c0 · inbound

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning cites this paper.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.386002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.386002Z digest=sha256:1028c43dbbdfd95cfa1dd06ef2ad7d30ef5b79950e6419fd54accfc491d17c68

Observation 415de428-35f3-4565-9940-7ce12511a372 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 283

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.297537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.297537Z digest=sha256:26dc42d036ab9a06b022bd1998283c3677264cbea2c0ac138ebd337d43a728c6

Observation 7b23d923-fb9a-45b2-9e6d-cf79ef6937c9 · inbound

Libra: Large Chinese-based Safeguard for AI Content cites this paper.

Libra: Large Chinese-based Safeguard for AI Content Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:38.663630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:17:38.663630Z digest=sha256:01f0cd847b736446fe8fce1ebd15dec23c27a15a286b047a5c775b2c725d8b7a

Observation a5a3d952-e6d6-49f3-a72c-8cf9204b488b · inbound

R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge cites this paper.

R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:18:24.466989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:18:24.466989Z digest=sha256:d258435cea5524acdf98d9dbace9ca068fc4efd5e518ee2a5c16623e155c0eba

Observation 078762c1-6d06-4376-85a0-197877d02b1d · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.538461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.538461Z digest=sha256:ca606b211da6c4db0057b6fdab41cc00a1355db769720f7601d85d78d2d26e6b

Observation ef1a2ba5-1e30-4db3-88ed-19856cf9274d · inbound

Towards terahertz nanomechanics cites this paper.

Towards terahertz nanomechanics Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:03:50.244026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:03:50.244026Z digest=sha256:e609a4278eabe0faa1a322ce8c5381c15fe62369deca4b06f089010aad52ec7f

Observation 7ef559a2-84e8-49b2-81d8-62ea5359e245 · inbound

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants cites this paper.

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:33.591580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:04:33.591580Z digest=sha256:081b97ebfa6d7560f8ad9f2a7c8c1560137b098e026cdd1730b701a8e0a9914a

Observation a1c86dd5-b54c-4c70-b9d3-856187e3ec0a · inbound

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI cites this paper.

Whose Truth? Pluralistic Geo-Alignment for (Agentic) AI Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:23:31.962576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:23:31.962576Z digest=sha256:7b2b3b545a8413cc6544d6d23c89c37eb8b7e20ae684c4b51e8edb5e93a820a4

Observation 11e7454a-7858-43d9-98a1-03d479d9b806 · inbound

gpt-oss-120b & gpt-oss-20b Model Card cites this paper.

gpt-oss-120b & gpt-oss-20b Model Card Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:22:54.679028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:22:54.633089Z digest=sha256:b52e865989be4a0414e00ee2f735efe4fe5ffef537998b8f6469c0aa2bcbfcbb

Observation 80137bcd-470a-400a-b79a-4e705dcc8b32 · inbound

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement cites this paper.

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:20:14.188064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:20:14.188064Z digest=sha256:33f3102f4f472d66371c5088289f6f45f10eded4702d47c1f14a86ad714b7464

Observation 73b85185-a26e-4f6f-abc8-c828b1299b94 · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:54.701047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:54.701047Z digest=sha256:7ffa8e5abacafaac35ddc5e5d32c3410d52b5ea7e757a83fe639175c1461098f

Observation 717369dd-b535-4414-a90e-5ed197e818ed · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.027988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.027988Z digest=sha256:4495c56912a13385b20985084293a5594d87a508509499d851cd4da8d248d627

Observation f910b908-634e-4eb7-8108-e78f6441d5d2 · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:36:47.435169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:36:23.882344Z digest=sha256:03630241642c2d1d4a54b54928c9bf630a2da97c344ca3d91653bd638fbb739c

Observation 812cb2d2-f762-48b8-92ff-288ab0b62199 · inbound

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs cites this paper.

Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:15.617560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:15.617560Z digest=sha256:4424bd2892e1fe1db207fdcd5aa99613cdd8eb6e52ef4a148d58d8f7b3719cca

Observation 7661fd46-165e-4792-9ccc-10ba8412cae6 · inbound

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents cites this paper.

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:49.835349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:49.835349Z digest=sha256:39c5a51de02632b004da06c2a88a81cc267ced2203f648c3e03b543ab733b97f

Observation ce52139d-2018-4134-8251-1252b0e2ce92 · inbound

Reasoning Up the Instruction Ladder for Controllable Language Models cites this paper.

Reasoning Up the Instruction Ladder for Controllable Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:07:52.834565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:07:52.834565Z digest=sha256:d25b7e39f8c250625e70f437144a3a9eb074bf94ffbfe13a5d58895758ecb1b4

Observation 76760757-a24b-4d92-be01-fe907de2fc53 · inbound

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection cites this paper.

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:40:42.343205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:39:17.785715Z digest=sha256:50a72ed0596826903301d12b22b772c2260483f935d38b17a614412cb0353eef

Observation b26cf182-2e61-46b3-a1f3-1629fa03bffc · inbound

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations cites this paper.

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T11:24:08.445919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:23:26.852673Z digest=sha256:fc55071b5a58d26cb4d12950d2b18b17ae36400b146b8d016f13691c8f748447

Observation ed444e7a-71c3-4c47-a1fc-47ab4227d11f · inbound

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities cites this paper.

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:25:51.877085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:52:16.543718Z digest=sha256:b8b937cfc33b7f1f40758c7194e31195220c7a32f693b57917d97882fe2004cc

Observation 414ce9e4-f3f5-49b9-a2fe-8e74c4f7c7a2 · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:03.213083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:44:50.762366Z digest=sha256:731c8cb27ba4be4e83cc20a33c868f1e23aaaefb58a145b0940aeab287de7dfe

Observation 26d3d840-1830-414c-9e64-2f30f50431db · inbound

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation cites this paper.

Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:19:56.880861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:19:56.880861Z digest=sha256:3c9e594711d491985dc7fa0b89b5fb28afe40088d3f393194a470cb84e283dad

Observation f5fbc12d-ebfa-4525-a4d1-a21615b3988b · inbound

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language cites this paper.

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:22:37.099023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:19:50.909364Z digest=sha256:98ba4ada47fb491fa8ede045011b8d061bef2ce0943e424f698b68449febc715

Observation 17fe05f8-fa6c-4f9a-957e-3e1ef5b721d5 · inbound

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories cites this paper.

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:10.098655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:48:44.687520Z digest=sha256:5eee59b8576ceb9e0e97230b9f77adb1ae1ba028ab7d4bda254d4e31f5c382a6

Observation 27bc220b-5f2e-4b0e-8fcc-df1df3f8a7af · inbound

Reasoning Structure Matters for Safety Alignment of Reasoning Models cites this paper.

Reasoning Structure Matters for Safety Alignment of Reasoning Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:04.339490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T02:36:57.093584Z digest=sha256:333c3ad12c3b5c526f4cdbd9a635ae6a833b4eda8301c289a6ae9b18c84b88d3

Observation c8701d36-fcba-4409-85cc-d5bb4047210b · inbound

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems cites this paper.

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:18.981416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:26:04.152045Z digest=sha256:68382513793651caf236442e354c260754c065346ec83fce275c32670eb69cd7

Observation e64a7095-35c2-462c-b0ae-e9560148d988 · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:41:09.007290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:17:19.079380Z digest=sha256:44ad30a1059e3552bbf905c12f47204ac3d92fd745eaa1d4fbc967245eb17b15

Observation e132bfc6-f458-4968-bb02-2054c8aeab10 · inbound

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue cites this paper.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:57:31.670860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:aa39fb2fb6ac62d0a09239b9eaee5361e946716a47b1b9b18999abe964a44eed

Observation 35d8e2ed-d827-4192-9da9-8eaedc0b0e00 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:08.576351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:8975f04d2b22905e92e93f404ab7b28b6b6a24d15279b235806f55ba69e8557f

Observation 479779e5-6147-404b-84aa-3e672a7f678a · inbound

Internalizing Safety Understanding in Large Reasoning Models via Verification cites this paper.

Internalizing Safety Understanding in Large Reasoning Models via Verification Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:51:14.318720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:50:59.283409Z digest=sha256:55f75f42b444f1159950a124b483e2a7f60845dbe129c52be34ed07436beb3ac

Observation 272fc79d-4ef2-4ae5-ba07-222ec5920666 · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.921471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:b96c83da545904c6f6221054d2e814a7d3d60c224194ef1bb9841fc85d3fa953

Observation 7f4d29f4-5a4d-4ba7-a7c7-0bbc2ddd70f1 · inbound

How Well Do Models Follow Their Constitutions? cites this paper.

How Well Do Models Follow Their Constitutions? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.874528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T15:33:43.569710Z digest=sha256:28ac695366f34102219a5c08d83a3b4148ce46fdf16b4ed54a1cd2b4c78db345

Observation cdfbb531-120e-466a-b884-6941277f5414 · inbound

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection cites this paper.

Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:35:50.610720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T00:29:50.433221Z digest=sha256:db901b0706088f356fddcb6a97cfa697052b6c95c6cb7dcc9c890d217b61cb0a

Observation ba830e08-2ec1-429e-ba58-0d5e4ad252e6 · inbound

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training cites this paper.

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.706723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:18:16.033063Z digest=sha256:4bf50797e17bf6a69e007787c065a0bf2d8d88dce96d5702371ad9e42afb0139

Observation 261b5b94-bf0f-4a3a-b4cb-96c188b31576 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.526096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:b82d196381e369e5e3d8738da84893e1e5f5d19f9c8373e470f87925a7bb43cd

Observation 9967d366-91e4-4062-8624-8761f1e45cc6 · inbound

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation cites this paper.

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.124147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:35:21.388956Z digest=sha256:37afd6d61eb536b6553e67ff698c893a1b8174da55526e81d15a40c0f9c3d7a8

Observation 0a3bd05e-3e5b-42b0-866e-0586952688a2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 244

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.453603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:8095bed3b7213f3e03a28b31ce4bc8c3f2e83ece05387a0b85ea5c1fa7afc9b5

Observation bd46bb0c-007c-4699-a9ec-b322b81061fc · inbound

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance cites this paper.

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:43.430699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T04:35:35.594085Z digest=sha256:2392d68bc66d69cb5bf2f8bac338243dff2983b4167e79d477eebde1e6744bce

Observation adb4d015-68ac-43f9-bbdb-4708714ccbe5 · inbound

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems cites this paper.

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:34.767260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:04:38.863201Z digest=sha256:2fc9a0463d689550c5963cd17f19c136a28435c3793dd1f6eef16abbdbd31287

Observation 043c0c62-8bf8-42ae-b1ee-ea97acc3187d · inbound

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems cites this paper.

Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:34:36.560267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T10:30:03.047181Z digest=sha256:d18d71304eb5e9c3b1d9bcc79db847e219936055a9e09dfdbf26f18f9bf40356

Observation 125b2c5f-13ce-4795-8bb9-9e81e6ad67a1 · inbound

Do Thinking Tokens Help with Safety? cites this paper.

Do Thinking Tokens Help with Safety? Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:30:00.639214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T23:37:49.412578Z digest=sha256:3ca4db3d27df4b87384dc0868bec0caf38c67895b28fe1bec689136a8920bb0f

Observation 8e26d784-e1d1-4784-a47d-feedb03f575d · inbound

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models cites this paper.

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:06.722457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T21:09:19.727723Z digest=sha256:7c782fff7a33483a3e7875fd9c932871aca44d20a296fe383965fc9e8d6c7f31

Observation 872dc00c-6b2b-4f47-8a70-0014c8a804a7 · inbound

Agent Safety Is Action Alignment cites this paper.

Agent Safety Is Action Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:35.001200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:50:45.759936Z digest=sha256:fe8fe7cf79cd606ba061b46c65a487a572e091924e30c62198c57c0dc971b4fa

Observation 1d0be5d4-c9cd-4fef-8161-6bc4d6f3d461 · inbound

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment cites this paper.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:06:58.846398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T12:57:00.805343Z digest=sha256:64cd0426b0aa3bb3d7b3847e38750878efa383796b56951cf1dcda2ad61818a5

Observation 9b2a054d-3b7f-4ee8-bccf-a3d47cd92f47 · inbound

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment cites this paper.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5a02788bb1b4a9de7f60ee80500b4964dece5483cbdeea3189e617aa18b55fd4

Observation c34d92c6-d355-4da5-a9f0-5f14463cc273 · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:54.934420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:e3fc5f8085f44d3858fd14419145c422eb4af6fc5cd38acef1eccb848cfe5fae

Observation b9561dc8-97ce-42f3-af4f-499781103678 · inbound

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models cites this paper.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:6cc4c2853fc57b98331f05f2b68d73b759bfc24d9f2880ca78a5bfb7c5a1fd23

Observation 5be75692-127e-4334-b317-c72f709d75c2 · inbound

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety cites this paper.

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-12T14:28:50.627444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T14:28:50.627444Z digest=sha256:01b53704864238e7e8a0a4dabe96f88fd04943a571be77cfde29f3af2b28093b

Observation 7c8e5ee8-1db0-4ee8-b32e-b1651172bed9 · inbound

Cost of Reasoning in non-English Languages: A Case Study on Japanese cites this paper.

Cost of Reasoning in non-English Languages: A Case Study on Japanese Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:12:44.286699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:12:44.286699Z digest=sha256:8dd7edcf98ecae19c3d3b2df3a8eb17df119419d8305eef4fb6cc8aecce11098

Observation d6aed98b-cb00-497b-991a-2288d69fe821 · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 230

Resolution
unresolved
no resolver link, observed 2026-07-15T08:35:47.870083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T08:35:47.870083Z digest=sha256:c5120b6168bc0c50572a8693e2c02979475de821e97a0ac817c092509709cd27

Observation a01fcd39-8be5-4f88-84ba-1f363099a19e · inbound

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack cites this paper.

Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-02T06:46:40.391719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:46:40.391719Z digest=sha256:91324ad122352b6323df589c62d6d1e92ab467ed55a3018586f73c954622d5da

Observation d406d612-cdeb-4712-a465-bb94445022c3 · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:23.274320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:23.274320Z digest=sha256:c7be662778be94c2599d28e1ed441b71b55c841223baccdce5883e8220ab1276

Observation 4492457e-d1d8-4fe6-ac1c-0f87fa2c6a30 · inbound

A Geometric Perspective on Stabilizing Value Conflict Resolution cites this paper.

A Geometric Perspective on Stabilizing Value Conflict Resolution Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T16:35:30.952522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:35:30.952522Z digest=sha256:92917d133c1fa22dcd6866900ccb1e59d3a17e06146d5894f411e34e08cb44ff

Observation 79cea9df-6c32-4571-bcc4-7adcd1b3fb91 · inbound

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs cites this paper.

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T08:38:55.560417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:38:55.560417Z digest=sha256:b7b75b67b0cf34c965764288d5e7c0609f2056482629c1537623a03787bce7e5

Observation d4c5bf99-e663-4e71-8425-5902ed995ed8 · inbound

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models cites this paper.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.667721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.667721Z digest=sha256:5c95c249465687d7c26e1b319d42ca7e37410f6745fffd3f5e50e9f388f1e583

Observation 78a191ed-6886-4149-8e7f-11aeed5beb12 · inbound

Constitutional Midtraining: Content Presence Drives Alignment Gains cites this paper.

Constitutional Midtraining: Content Presence Drives Alignment Gains Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:35:00.225163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:35:00.225163Z digest=sha256:c51617a9dacb69782d29dc7f9cd8015ec905adab5ba729aed441dd4902584b48

Observation db5790cb-4fa6-40dc-a93d-2b4355464f2f · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:26.100966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:26.100966Z digest=sha256:6b1ccb6ff8503d190fcb165e49360143cf121d447e3dc3bea3f94409823a4aa0

Observation 77fe7ada-8310-426f-8ddf-ff0fc104378d · inbound

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving cites this paper.

Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T14:38:56.180832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:38:56.180832Z digest=sha256:58d507af41687af391b43c723c2a017de0fd13138272d337df71cc33c9189480