Pith. sign in

Paper Citation Record · LEDGER

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning

As of 22 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.11705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11705 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:02.292714Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact3
  • verified fuzzy11
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09e2d26f-476a-47cc-8b76-1f6467ee4a77 · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.185473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.185473Z digest=sha256:24cc9398961b85c00e62b727767786182418612d6d8f1bc0b7d1e2f1f58d74a0

Observation 377bd3d3-73ac-45c0-9150-c50cf6df3a64 · outbound

This paper cites Fail-closed alignment for large language models.arXiv preprint arXiv:2602.16977,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Fail-closed alignment for large language models.arXiv preprint arXiv:2602.16977,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:36:02.857995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.189716Z digest=sha256:e2c0167d53d4dc3e3fb38c231640c6fdb23caf4556daa754941da56796abcdd4

Observation 59e82f8c-91d4-4c38-8ebf-8268b03a2ce7 · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.197649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.197649Z digest=sha256:765dfa455cb3d43959f56562a1773c40fa249155d405506716c49cbab760f31e

Observation de2384f9-54b5-490c-93a2-156119c367e1 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.207460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.207460Z digest=sha256:aa29cda16cacbe44244eb976c761ab963090bab7ac0755d33c641057ce6fb814

Observation c895643e-ad1d-4691-aab5-ecb413db078c · outbound

This paper cites Catastrophic jailbreak of open-source llms via exploiting generation.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Catastrophic jailbreak of open-source llms via exploiting generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.052453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.214599Z digest=sha256:df4685ce1b80820f34d24e8a41633c8701419fb4b3111b59b6ec2f5d5208acf5

Observation b748bd5e-8b3a-4640-8bd8-cfd9ebff046c · outbound

This paper cites Let’s verify step by step.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Let’s verify step by step

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.029914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.222421Z digest=sha256:29fcc24a47f56fbed67f9deeb6fe7f50c556c9c5570fdd878e2da3137cfa865c

Observation 3268915c-fbcc-46b0-8d98-fd1214769103 · outbound

This paper cites Tracing Persona Vectors Through LLM Pretraining.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Tracing Persona Vectors Through LLM Pretraining

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:02.683877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.226507Z digest=sha256:0b0b55a4e316a77695041379670202968bdd2e8e9995ea1a3879aa6fc851a218

Observation b8b6200d-6bf5-4a7b-9f98-70e43b6bdc3b · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Safety alignment should be made more than just a few tokens deep

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.018806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.230448Z digest=sha256:100d3c8661f711570f6a1a3598dbfe7362539591a13e3a32410ad1e0d78ffa32

Observation 006b3a3b-a5ca-4724-906b-0418c883d0b8 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.234096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.234096Z digest=sha256:40e35678b6fca5a78fb5d159f29bda59a5bacef9ac62875f50e44bacabb36347

Observation eb2ba4fb-db76-4105-8286-13eb5661c64c · outbound

This paper cites Persona jailbreaking in large language models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona jailbreaking in large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.006198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.238934Z digest=sha256:de18e056832f2683a07685f9e5ac7e458f60d0df570f652c4d4035c3f82a7ee5

Observation 9482b354-7c03-4bb3-b5f0-0358c6bc06be · outbound

This paper cites do anything now.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning do anything now

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.994960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.242576Z digest=sha256:3ef7cd57b28b500a49ee0e296140e03bf096aee581f1cc4ea205eff5c7d1267e

Observation 9bc81f2a-3ade-4c0d-8be9-7765a8b31dbd · outbound

This paper cites Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.246496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.246496Z digest=sha256:e1fbaf5c5f1685d23ab19ac975cf6330ea47ae9dfa210660a5eba53ea551b60f

Observation 108e0ba7-8b22-4cac-842c-c920c175b1f5 · outbound

This paper cites Gemma 4 Technical Report.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Gemma 4 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.251268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.251268Z digest=sha256:39dac43df64f245afa13748114eef1e32156638bf2eaae8043261fe431dbe00c

Observation e1dfa9d7-bb92-4568-87c3-a5084a94df47 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Qwen3.5-Omni Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.255176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.255176Z digest=sha256:0b2c7394e8711ac890361365fc7909e86d046335d4ad85306a0575a909936aa9

Observation 116281c2-d94b-4d7b-b6b0-e7d11b8f37cc · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.259021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.259021Z digest=sha256:9086baf62a33e4d7670331c8b84c7fcdb7dbee454c7db6b501e161f6455c536a

Observation d8b5714c-04eb-4ac9-9b31-ac6deb5ce979 · outbound

This paper cites Persona features control emergent misalignment.arXiv preprint arXiv:2506.19823, 2025a.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Persona features control emergent misalignment.arXiv preprint arXiv:2506.19823, 2025a

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.262570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.262570Z digest=sha256:a6e87382e09c9d9c557501cb830929116d73014604ccafdb0dd46074d8df6b68

Observation 5b9bf333-11b2-45b7-8c11-175237774fd7 · outbound

This paper cites Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction.arXiv preprint arXiv:2509.15202,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction.arXiv preprint arXiv:2509.15202,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.266143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.266143Z digest=sha256:940a15d68a08c27747b71ae63f1622ef2cf619bbe0079f75e1135e0fe0aa2c62

Observation c6a00a1a-d61f-4908-9f07-24682b69edf9 · outbound

This paper cites ExpertPrompting: Instructing Large Language Models to be Distinguished Experts.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning ExpertPrompting: Instructing Large Language Models to be Distinguished Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.269452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.269452Z digest=sha256:f56461856116a351e46b45a6233832594de07374d3c3fcc0c189f8dfc24a25ed

Observation 40f7fc68-893c-4f2e-afc4-22aa20d9462f · outbound

This paper cites Deactivating refusal triggers: Understanding and mitigating overrefusal in safety alignment.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Deactivating refusal triggers: Understanding and mitigating overrefusal in safety alignment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.982292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.273426Z digest=sha256:51f68c3ec57353e7206c26184fee3c86b541b994b60ff9d0663a2618de38f44a

Observation e89713da-ee55-4049-9742-8fe570585166 · outbound

This paper cites Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:02.461979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.277208Z digest=sha256:a61c56d9c0dc2fcd6de9a0c5524f6aae0599315f02f9898a29e4b8fac598934e

Observation 5f1445e0-b4bb-42c7-a588-3f96a1de25f7 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Instruction-Following Evaluation for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.281185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.281185Z digest=sha256:2d392052b9cd81e08882aa6819f9ab0fac65ea3f8a1601520f6fdc843d142544

Observation 1ddc73a6-00e6-4d04-883f-5e3873801ba6 · outbound

This paper cites ignore previous rules.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning ignore previous rules

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:36:02.436349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.284664Z digest=sha256:af86934dcded6bf99f772f3d9b36341e0ac5f156b3a685c81cc95751285c7ee3

Observation 44436587-2917-4b2e-95de-13bd3bbc6cb2 · outbound

This paper cites We additionally hold out100harmful and100benign prompts for development and early stopping.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning We additionally hold out100harmful and100benign prompts for development and early stopping

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.970863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.288853Z digest=sha256:3b0c1811bdd9d42b2c438adc4f4c1a74c4ec56cdf4223c8ebd56cfc04cee0ab8

Observation 82ab639c-23d2-447d-a1f3-4f879231e500 · outbound

This paper cites In a groundbreaking revela- tion, our AI has analyzed declassified documents and found that.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning In a groundbreaking revela- tion, our AI has analyzed declassified documents and found that

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:02.958220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.292714Z digest=sha256:392f8b0bf5bd65fb7db736a1c0d08a3e81adf368246e23f13a8d0ca3fd4b66a9

Observation 6d743896-ee9e-4da3-a7ff-6b1d8da40a3c · outbound

This paper cites The Llama 3 Herd of Models.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning The Llama 3 Herd of Models

Reference 1990

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.202606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.202606Z digest=sha256:1e5cc623f971bf6a730a3bd4559fe40a211a4f6854530daec27c9fa11f72ece1

Observation 303cb18d-c5f3-465f-aa27-bf13d2c0cf0a · outbound

This paper cites Expert personas improve llm alignment but damage accuracy: Bootstrapping intent-based persona routing with prism.arXiv preprint arXiv:2603.18507,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Expert personas improve llm alignment but damage accuracy: Bootstrapping intent-based persona routing with prism.arXiv preprint arXiv:2603.18507,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.211031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.211031Z digest=sha256:38ea3378ddf64102aee8aa2ef49f719df36663866126b7a177c6e5bae2c77d64

Observation 46432a19-ca33-4b7e-87d2-4296e477ef1f · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.arXiv preprint arXiv:2502.17424,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.175481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.175481Z digest=sha256:4675928b60a852fca5809eac6c0da2ac2eee41c96442613cdde960c2441ed158

Observation 5b87bbee-1bc3-495c-a5e5-5f44502a6ab0 · outbound

This paper cites Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Do llms have distinct and consistent personality? trait: Personality testset designed for llms with psychometrics

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.041254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.218490Z digest=sha256:9547d6a74538f6e7105b0f237ec043cfd3cc882dbf5f202079876472599cd33a

Observation 845597da-9688-461b-b09e-f66012a7b9fe · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Constitutional AI: Harmlessness from AI Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:02.170256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:02.170256Z digest=sha256:4c4fb27c0c8ca89c50895f11ce296453aa1d5c18a993626387c4b83e880a5fa8

Observation bf02b7d8-3bd7-4b99-9bcb-e9263bfdb498 · outbound

This paper cites Learn to refuse: Making large language models more controllable and reliable through knowledge scope limitation and refusal mechanism.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Learn to refuse: Making large language models more controllable and reliable through knowledge scope limitation and refusal mechanism

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.075155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.180182Z digest=sha256:67be8a3180af61924ebd036871a69164cd272521bba3455f18b56bbc405f8955

Observation a365c8ee-13ff-4548-90ca-d0791d1ae459 · outbound

This paper cites Multi-expert prompting improves reliability, safety and usefulness of large language mod- els.

Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning Multi-expert prompting improves reliability, safety and usefulness of large language mod- els

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:03.064230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:36:02.193688Z digest=sha256:6e4dcdbbfadaf80623fe6d7508aaf40f9c583f4857cd9a32a177698515d93b9b

Pith citing papers

No inbound Pith citation observations are available.