Pith. sign in

Paper Citation Record · LEDGER

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

As of 8 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 20 inbound Pith citation observations for arXiv:2508.09224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09224 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:35:34.616461Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:12:44.184834Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6fae85d6-77d2-416f-8c51-c429779c4d98 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:33.950466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:33.950466Z digest=sha256:47109b611b9a718340b92a4b4fe62713e636ede7b66bdb659b494002e60fe6c5

Observation f2be2c8c-8ac3-4957-8e84-5c49decdb846 · outbound

This paper cites [14]OpenAI.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training [14]OpenAI

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:35:34.923470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T21:35:34.366680Z digest=sha256:179cbdb8e23445beb85f885c64597bc84737e2720f52bf4c5369feaa5b2f7291

Observation 81b07747-e930-4ec7-94d4-ab893594fb53 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.494480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.494480Z digest=sha256:ac0758978230ee28cbb510224ea527f14ab8f9f48972a87ca961a5a0ca05ad5a

Observation 9c3d15c9-ca69-4187-af43-76ddc9f34c93 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.096588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.096588Z digest=sha256:da083b1dbc950078bdaeb1ef9383200c51f9c0f1e255b40a7384c899b2739cd4

Observation e05cc146-80ba-4697-b710-292789aa3a39 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.616461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.616461Z digest=sha256:44053c7aab1b5cf97e33ccc8024d028a9a692c64ef415b8c2ef2d79f42cabf80

Observation 9a63e92e-2b3c-4b20-afb6-6ea4a7b1b4db · outbound

This paper cites GPT-4o System Card.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.168865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.168865Z digest=sha256:1f8ca8ee0d4535915d97be9ba0a7d680c5b66c341eed2335b97713fbc49aa20b

Pith citing papers

Observation d5e3815c-4641-4358-ba47-d92df3d6a102 · inbound

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks cites this paper.

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:38.344182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T17:00:06.402954Z digest=sha256:06aad951a4180c901ae734dfeccd002495745c8dfa6dd5bc4d4f11ab88eeda05

Observation 5896c4be-d5ba-4f1d-b9d4-3591bf8b6ef7 · inbound

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models cites this paper.

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:42:36.853762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T12:41:48.620040Z digest=sha256:be6b269fddb174398025b065a553eab863bcf9346c9d9f2629ae320542cb7c2d

Observation 36ecda09-0937-4c16-ab76-00ee4d62cc54 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.616391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:118a4e3c3b517a6fe85c9f8bbbdd5e68bd118da798e56b2ff4d6fb61b02d1fa0

Observation 07612579-4de2-4a67-a42e-7fd7c0e3ed24 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:f4b1b8b696d6c912efcfbbffbc849eea980b1b1a2bfb41d77ea786045bf72041

Observation 66929448-40e1-4ddf-be66-fc7de901521d · inbound

Cat-DPO: Category-Adaptive Safety Alignment cites this paper.

Cat-DPO: Category-Adaptive Safety Alignment From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.750925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T05:33:32.642379Z digest=sha256:5aef1987e7605dd0653d656220e367e75988d172ff8f8670e7c3f2279940ee1c

Observation 41e0bb4f-39d5-48a6-83fc-c31191e12e91 · inbound

Using large language models for embodied planning introduces systematic safety risks cites this paper.

Using large language models for embodied planning introduces systematic safety risks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:20.104188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T04:45:46.483867Z digest=sha256:8d633e9e5835dd216a08aa09b210cdf2919b3f147206f407bf2ad3c2b4bb49da

Observation b0bec346-0217-440a-9c94-502e42612725 · inbound

Jailbreaking Frontier Foundation Models Through Intention Deception cites this paper.

Jailbreaking Frontier Foundation Models Through Intention Deception From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.010147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:17:51.039062Z digest=sha256:1113267beafcf85a7145995ab41e6e924bdd8faf18d2a45fa2bfcda590ad0241

Observation aec6227f-8076-41dc-94a9-9d0babd9efd4 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.563625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:d8849e2c202806c47b3b86c6a39f69859b72122fd3b21c49157ea25479403c39

Observation 27dddf4d-e254-4f60-953b-8ff3632f9d83 · inbound

Internalizing Safety Understanding in Large Reasoning Models via Verification cites this paper.

Internalizing Safety Understanding in Large Reasoning Models via Verification From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:51:14.238142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:50:59.283409Z digest=sha256:f5ee6f6288385a2a58bc78d02c53d8425701571f1ecb613bcc9bfa958048b005

Observation 5d489190-aa42-498a-9084-49fc23f6703c · inbound

Reducing Political Manipulation with Consistency Training cites this paper.

Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.233116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:32:19.312335Z digest=sha256:54264cbe55be6bda153eb774b2898226537920be2ea3db0dfc4158502b9e681f

Observation 27d5e496-392f-4511-a39f-bf7b49040453 · inbound

Reducing Political Manipulation with Consistency Training cites this paper.

Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.727042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T16:49:16.542582Z digest=sha256:4446b9a82a60634c6bfb13b22e401d2b2b46425c4412b7ed9ac72735f591ca37

Observation 1a845245-3a9c-46d1-9b9b-cdb20a841c83 · inbound

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories cites this paper.

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.785841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T22:33:26.600072Z digest=sha256:c9162feea61cbd773fe50f558576b36928e394ec9b0256146f1610cad3cda1e6

Observation d966b265-c9f3-4000-b32b-f8565878dee3 · inbound

Investigating and Alleviating Harm Amplification in LLM Interactions cites this paper.

Investigating and Alleviating Harm Amplification in LLM Interactions From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:24.040929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:31:52.027889Z digest=sha256:951ce958805ea3df945db8205e60778ccd6e54ff005439d1596cbaf763d7a035

Observation fed5df3e-33d3-415e-967a-4a6a64015832 · inbound

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability cites this paper.

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.803972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:00:30.904247Z digest=sha256:a2dc1f72686e913c4534216cb1f28c8e8e5c4a798def7e74abb756c316610442

Observation 28e34803-b869-4ca7-93d5-9fb1e5c85b01 · inbound

Understanding Censorship in Large Language Models: From Mechanisms to Governance cites this paper.

Understanding Censorship in Large Language Models: From Mechanisms to Governance From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.397555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:02:35.840261Z digest=sha256:ab2bde8395c57c299835540ab18a59c987864e8dbfb380c639bdd85a6c97d68c

Observation 8032e4dd-60b3-47f1-a891-572cd3fd65ad · inbound

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets cites this paper.

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.496890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T14:44:57.205766Z digest=sha256:2da57443b6773b7a580d9c37754a6b5e2925ebc5ea470a6bb17b6f5a8691fdea

Observation 13e43977-b44e-4383-bfa6-66195151621c · inbound

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models cites this paper.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:2c7a696c1a390266f93f2b742bf020db71a8234e45191308e9f9a778c7168ba5

Observation b3b09427-a1e8-4a06-a09e-bb0a28d19412 · inbound

GPT-Red: Automated Red Teaming via Self-Play at Scale cites this paper.

GPT-Red: Automated Red Teaming via Self-Play at Scale From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.184834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.184834Z digest=sha256:da3fe0c65c0e3c9dbe50b36de885e3a4f22ed7f263f7e7ccfc854deb6de9af48

Observation 3a138c4e-30e0-4b6a-8e2f-7bead74052f4 · inbound

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting cites this paper.

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:35:58.478085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:35:58.478085Z digest=sha256:0b6862fc4a389f942bc69571370b61372a9dea09b12c37610e6d5642c256ccb7

Observation 9486a6b5-410d-471b-afd1-06341419d6b5 · inbound

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs cites this paper.

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:22.224180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:25:22.224180Z digest=sha256:ade864e9bd09b61fdd2f272d664f02c554fb31c88c9db6b2f4a7805c6133900d