Pith. sign in

Paper Citation Record · LEDGER

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

As of 14 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 21 inbound Pith citation observations for arXiv:2508.09224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09224 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:35:34.616461Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:34:57.214509Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6fae85d6-77d2-416f-8c51-c429779c4d98 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:33.950466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:33.950466Z digest=sha256:7241d6facb76ae2d97668eb9f12f7209fad701018ea52dfc890d079201830beb

Observation f2be2c8c-8ac3-4957-8e84-5c49decdb846 · outbound

This paper cites [14]OpenAI.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training [14]OpenAI

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:35:34.923470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T21:35:34.366680Z digest=sha256:71a3f00f426f6fb3230e1cd9fcea5f627295e8434ddd93c8dbb82f200aa9d73c

Observation 81b07747-e930-4ec7-94d4-ab893594fb53 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.494480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.494480Z digest=sha256:84b5d4692529f3274e593f5840840bec0134372575b0c9c8f9491919d044404a

Observation 9c3d15c9-ca69-4187-af43-76ddc9f34c93 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.096588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.096588Z digest=sha256:73d7a229ddbd2d1f2e904ac9fb063a298fc70695cc5152d7f9f90043eae0ed74

Observation e05cc146-80ba-4697-b710-292789aa3a39 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.616461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.616461Z digest=sha256:54f79f895abfed30f1b383a89ce6ec97eb791c670804acf728738ce128e8bbc2

Observation 9a63e92e-2b3c-4b20-afb6-6ea4a7b1b4db · outbound

This paper cites GPT-4o System Card.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training GPT-4o System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.168865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.168865Z digest=sha256:edbe10ff503daaedf2db4a75f2f1bffaa04c5605a0848d7aa1cba9d85f953e92

Pith citing papers

Observation d5e3815c-4641-4358-ba47-d92df3d6a102 · inbound

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks cites this paper.

Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:38.344182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T17:00:06.402954Z digest=sha256:fc01bc1e7d4e1311e3d090cd27fd207ac5894bdb64f84fdad0143bbf2b0e9b79

Observation 5896c4be-d5ba-4f1d-b9d4-3591bf8b6ef7 · inbound

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models cites this paper.

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:42:36.853762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T12:41:48.620040Z digest=sha256:966b64e8411d6d4dc243bb47b6b52dda3ac862fc935280ed3dcd334e7959ebc7

Observation 36ecda09-0937-4c16-ab76-00ee4d62cc54 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:52.616391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T18:25:53.037936Z digest=sha256:2b7e23b9b227f7b158adb4b3e626484a83cf47885dfabbd688a837e125f88c00

Observation 07612579-4de2-4a67-a42e-7fd7c0e3ed24 · inbound

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures cites this paper.

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T00:19:33.861692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T00:19:33.861692Z digest=sha256:13ccf53abc5b171701924a9b5ce50fed55899253f33481010f29fc7ec3f50ec3

Observation 66929448-40e1-4ddf-be66-fc7de901521d · inbound

Cat-DPO: Category-Adaptive Safety Alignment cites this paper.

Cat-DPO: Category-Adaptive Safety Alignment From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.750925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T05:33:32.642379Z digest=sha256:921d6a0a5d72be8818975eb6e30426da59ea2cff3e7159df08bce0f61c6c711f

Observation 41e0bb4f-39d5-48a6-83fc-c31191e12e91 · inbound

Using large language models for embodied planning introduces systematic safety risks cites this paper.

Using large language models for embodied planning introduces systematic safety risks From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:40:20.104188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T04:45:46.483867Z digest=sha256:a5cc40e28d78667b80418e1323aba679d8a25058d85050446d04a6a83fc4dbe1

Observation b0bec346-0217-440a-9c94-502e42612725 · inbound

Jailbreaking Frontier Foundation Models Through Intention Deception cites this paper.

Jailbreaking Frontier Foundation Models Through Intention Deception From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:14.010147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T03:17:51.039062Z digest=sha256:7577e877006604a937e8459d1cb1efbc6cc86aa0acd9f57b6cd79f428faea051

Observation aec6227f-8076-41dc-94a9-9d0babd9efd4 · inbound

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering cites this paper.

Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.563625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T11:49:47.994456Z digest=sha256:5b59730e0e0027d94cf0a979283dc6a71de58c1f2392c528a35204f93e1faed9

Observation 27dddf4d-e254-4f60-953b-8ff3632f9d83 · inbound

Internalizing Safety Understanding in Large Reasoning Models via Verification cites this paper.

Internalizing Safety Understanding in Large Reasoning Models via Verification From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:51:14.238142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T01:50:59.283409Z digest=sha256:838e3488aae87ade8842ca9d9b511b6a70e8987d56de91de670f473c4ca42853

Observation 5d489190-aa42-498a-9084-49fc23f6703c · inbound

Reducing Political Manipulation with Consistency Training cites this paper.

Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:34:40.233116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T05:32:19.312335Z digest=sha256:db144ac01295551cc9502eabeea517458d0e26c8687c8699046e9691ef3cb864

Observation 27d5e496-392f-4511-a39f-bf7b49040453 · inbound

Reducing Political Manipulation with Consistency Training cites this paper.

Reducing Political Manipulation with Consistency Training From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.727042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T16:49:16.542582Z digest=sha256:f0c0e088761d3b3f190f8c15ad17091ca3a582844b46f92390f597903e5eca18

Observation 1a845245-3a9c-46d1-9b9b-cdb20a841c83 · inbound

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories cites this paper.

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.785841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T22:33:26.600072Z digest=sha256:62f4617c07734c8e97163928ca40859202169fdcf7002fccd75c7446a8de0b32

Observation d966b265-c9f3-4000-b32b-f8565878dee3 · inbound

Investigating and Alleviating Harm Amplification in LLM Interactions cites this paper.

Investigating and Alleviating Harm Amplification in LLM Interactions From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:24.040929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T14:31:52.027889Z digest=sha256:566b9e74118bf87509011cb79aa8628108c234242a2a7c8f8b645eb5f4cafe36

Observation fed5df3e-33d3-415e-967a-4a6a64015832 · inbound

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability cites this paper.

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.803972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T10:00:30.904247Z digest=sha256:b464ad44ddf2d62e686ae34d17e73120df69eb211b86cfc838d9a46076509ec8

Observation 28e34803-b869-4ca7-93d5-9fb1e5c85b01 · inbound

Understanding Censorship in Large Language Models: From Mechanisms to Governance cites this paper.

Understanding Censorship in Large Language Models: From Mechanisms to Governance From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.397555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T07:02:35.840261Z digest=sha256:769835feeca6940be5ac8a1f30e71af78b288b177e368695845375adbd81b36c

Observation 8032e4dd-60b3-47f1-a891-572cd3fd65ad · inbound

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets cites this paper.

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.496890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T14:44:57.205766Z digest=sha256:ca5a70ad70397cacf5e75e93babaaa26e2964c325e2e11d5a33aab7b2cc373b9

Observation 13e43977-b44e-4383-bfa6-66195151621c · inbound

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models cites this paper.

Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T06:11:38.403281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:11:38.403281Z digest=sha256:d11b70338e2970a976bc3e6490ba0388a891dd3c1126092f1741f5711268efc1

Observation b3b09427-a1e8-4a06-a09e-bb0a28d19412 · inbound

GPT-Red: Automated Red Teaming via Self-Play at Scale cites this paper.

GPT-Red: Automated Red Teaming via Self-Play at Scale From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.184834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.184834Z digest=sha256:dec0abbd044cf97f45b7e31f6e3d0741678b112051ebdbe55e97c2076ed23ee1

Observation 3a138c4e-30e0-4b6a-8e2f-7bead74052f4 · inbound

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting cites this paper.

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:35:58.478085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:35:58.478085Z digest=sha256:6fe1e1739c4f0b4ee7e746424c626b2b1eba69e13e1c2825923e6ffa7339606c

Observation 9486a6b5-410d-471b-afd1-06341419d6b5 · inbound

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs cites this paper.

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:22.224180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:25:22.224180Z digest=sha256:1302c56418d92d664d034701ad0ca945ab8281d3ab272a44bdf2bd955a26064f

Observation a2b2ff0b-4eeb-40cd-90bd-3899f92aa4c8 · inbound

Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents cites this paper.

Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:34:57.214509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:34:57.214509Z digest=sha256:4ba3a55d6d215802549a716cdbc6b622bade77bc5f0a505384a9976a7090f3f7