Pith. sign in

Paper Citation Record · LEDGER

Lifelong Safety Alignment for Language Models

As of 14 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 4 inbound Pith citation observations for arXiv:2505.20259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20259 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:00:11.823481Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:50:36.178215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:19:41.953726Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a49d225-c6b9-4970-adc6-1f49221a6a08 · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

Lifelong Safety Alignment for Language Models Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:02.705923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:02.705923Z digest=sha256:437b12a56053150a372bd0ca927c2da0085b23177883d48ea20e9b70bee0556e

Observation 6caf749b-2035-4bf8-9153-46c27287f545 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Lifelong Safety Alignment for Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:02.783278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:02.783278Z digest=sha256:a8bd232fff3227231c8d904957af9fdc0beea8906e9ea8802bec826b6ae9c0f0

Observation adb2ad27-37bc-4f36-bc8a-97bae27c567e · outbound

This paper cites Many-shot jailbreaking.

Lifelong Safety Alignment for Language Models Many-shot jailbreaking

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.679381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:02.976816Z digest=sha256:e3f638b14ff22bfa95aa58b53b28503a3cbe7fade214a66d7f02b2264fe5b4b8

Observation f5b061f3-b0bc-4a0a-9446-ce20d2013e71 · outbound

This paper cites Program Synthesis with Large Language Models.

Lifelong Safety Alignment for Language Models Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.163692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.163692Z digest=sha256:902ceb8abfc6b83b66b1f900ebcdfa53f989bee58f67565b47486061d576b003

Observation 22bbc4bd-ca39-4bf6-94f1-13330c877a78 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Lifelong Safety Alignment for Language Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.293613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.293613Z digest=sha256:1adc1c203b988f830bdb5bff8f12f55c9b2dffa5ed2911e55141dcdb73892450

Observation baa57ae7-86c2-4a8e-90df-38e7964957d1 · outbound

This paper cites SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning.

Lifelong Safety Alignment for Language Models SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.410838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.410838Z digest=sha256:0a82761945a30a31d79691a964549389ec6ce309d6b1fb6214199d4699fb5db7

Observation 927c47d5-5857-48a1-a1ae-90fe22b8daa4 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Lifelong Safety Alignment for Language Models Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.557694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.557694Z digest=sha256:6437925b08d0ccd590a4ce426e5d27b8263a675eb128f9d0d9c10ac337842b3f

Observation 837e0876-3179-45f5-8460-4c80e46fd7fb · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Lifelong Safety Alignment for Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.743897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.743897Z digest=sha256:f4f47bcc0b83b60756327531f36ed3b69b2bfc0cb4eaa541665a9443ea53ef6a

Observation faedbfb1-52b4-4c8c-ad1d-e666659e3196 · outbound

This paper cites Self-playing adversarial language game enhances llm reasoning.

Lifelong Safety Alignment for Language Models Self-playing adversarial language game enhances llm reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:03.880607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:03.880607Z digest=sha256:32d107780c72c11c1495230cf3095432a5db799d2990c7941a1dab1b0756c1ca

Observation fb5e5162-8911-4c6f-9784-d0040b47ef3b · outbound

This paper cites On the Measure of Intelligence.

Lifelong Safety Alignment for Language Models On the Measure of Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.018701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.018701Z digest=sha256:463458764f0343a42e4a67df9491352f189ad360c72db3e3b273d4a86e15e461

Observation d278ce1f-7e5e-4f52-be5a-777b285777f3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lifelong Safety Alignment for Language Models Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.168371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.168371Z digest=sha256:ccabf7149f78f90025d20b76e788493b4447e7edea1251f9b68dd9d7639d5bf9

Observation 851ed2b7-d71f-467e-8ecf-5cf79da8a869 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Lifelong Safety Alignment for Language Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.302593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.302593Z digest=sha256:8eac9de6d8b32e847bb5b495ddabcc25ce4950a37ff48ec5c620424ff6d0eb69

Observation 33bb40bf-19b0-4bf2-8368-8d379e62a761 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Lifelong Safety Alignment for Language Models RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.449719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.449719Z digest=sha256:b3c802c0e51131dd5c64f9b411976e556c7930287a145ff4751f5a9aa8f9f61e

Observation 3aacc81d-2e4d-47b9-8d59-5dbf1d4b9286 · outbound

This paper cites Beam Search Strategies for Neural Machine Translation.

Lifelong Safety Alignment for Language Models Beam Search Strategies for Neural Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.618458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.618458Z digest=sha256:7cd052a032baba867ed9b7b66fe2115b953234c0f5332ae9353e851d7448e660

Observation 43aca306-c3ce-4dae-9f7b-7f25bff4b5b7 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Lifelong Safety Alignment for Language Models A framework for few-shot language model evaluation, 07 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.757264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.757264Z digest=sha256:72a58571e1652dc379f19c5f3619334d7060eb9a2a1599f8cd67f4c735392472

Observation 9298d76f-bc7f-4d4a-aa27-382ceecacf83 · outbound

This paper cites Attacking Large Language Models with Projected Gradient Descent.

Lifelong Safety Alignment for Language Models Attacking Large Language Models with Projected Gradient Descent

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:04.958799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:04.958799Z digest=sha256:99a27aea1a99c1ae21899052d5365d0ba1f75b8c650f547bd7cf1b86e83861c7

Observation 45941abb-072d-4929-8b1a-afb08a0d0b05 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Lifelong Safety Alignment for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.091365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.091365Z digest=sha256:38892558a7f2164af966bfab6a33d50debb7840f5d082cad03500d2dba7e8c5b

Observation bd54c7f8-2255-4319-9439-aa020ce9bed1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Lifelong Safety Alignment for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.215286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.215286Z digest=sha256:109e0e44df9c63a3bb207a105c1d13fd96f6e91b12d8c7571490f7d33f05b262

Observation 9dac24db-23b1-45a1-8b56-0a5a09ed3a42 · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Lifelong Safety Alignment for Language Models ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.301732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.301732Z digest=sha256:8a3febcb2da203d2b37dcaecb0c2928696aa731a0927de23a420254a6ebeaa6b

Observation 4ec77654-9fba-4408-bb0f-a49ac81aabd2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Lifelong Safety Alignment for Language Models Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.483583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.483583Z digest=sha256:9e3ee6809805eeb3ab4be36f372987d7e2cb5c5cf19029e5c96a693c40dcb416

Observation c04b777c-af51-47db-9894-a42ae8072038 · outbound

This paper cites GPT-4o System Card.

Lifelong Safety Alignment for Language Models GPT-4o System Card

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.610293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.610293Z digest=sha256:6429a870491bed9d5cfbb670431b2c519c5f7bfc383ca10c54c80b58b2831b0a

Observation 386af72b-a74b-4055-ba72-8f2a48cc9a0c · outbound

This paper cites OpenAI o1 System Card.

Lifelong Safety Alignment for Language Models OpenAI o1 System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.784733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.784733Z digest=sha256:581c13ddc63e384e9a285f15fe855013040d6c2ced192b12911fdcd098d4726c

Observation a158b6c6-fa85-4b43-9101-14ce2a31ae06 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Lifelong Safety Alignment for Language Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:05.888953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:05.888953Z digest=sha256:b66a8af18a93d873ae195eb379aa6bfab895e1db251d80f7b2670b3c08a2609d

Observation 295d1e45-2c41-4f87-bf24-de54338125a0 · outbound

This paper cites Improved techniques for optimization-based jailbreaking on large language models.

Lifelong Safety Alignment for Language Models Improved techniques for optimization-based jailbreaking on large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.453486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:06.038643Z digest=sha256:22103da4c4eca37fd3ede0785b5ce4ef5ae29d7ad3a71480c9ac4d4217f3aaad

Observation eb8e42e3-35b7-42d9-8c7d-616bbe7fc5ca · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

Lifelong Safety Alignment for Language Models Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.306843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:06.175140Z digest=sha256:a17cc270935dbd743a89e0c873e5ce54aece0743741ea6c8cc5deaf7dd950d19

Observation ae230cd9-f965-458c-b10e-c53bee44f1f4 · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Lifelong Safety Alignment for Language Models SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.273227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.273227Z digest=sha256:81bf12893233bf574503b9e5fee2ae4dd3a3a17feb95e992605db0f163472f0d

Observation ea4f5390-cf96-4439-b556-1028e05aec36 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Lifelong Safety Alignment for Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.392442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.392442Z digest=sha256:cda63add5b752afddf41cbe8433a6b5a2ef6446ecf1399f428e70d96a665539f

Observation 4c64cd0d-6ab5-43d8-af99-a72ae829e781 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Lifelong Safety Alignment for Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.455968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.455968Z digest=sha256:0d4b9e8591b59610d582f45e61f8f5fe68953ba6677457a9f693a3cccf70da34

Observation 8fa47a29-dac1-4fbc-9e73-b4346fcd2907 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Lifelong Safety Alignment for Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.556951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.556951Z digest=sha256:00d2896db24cd0c14de93aa57496f66cb1ab8ec6a5c840963a36900f79cdbc5b

Observation 1f3963e7-0887-4698-a8ef-f282c1479027 · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Lifelong Safety Alignment for Language Models AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.675923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.675923Z digest=sha256:de2a8532b319d2c485d2a95ed284d80a7931b639fd2d44a6212bb4d59b0e3291

Observation d0d40a5b-e08c-4b30-abcc-3a8f1b8ef07c · outbound

This paper cites Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models.

Lifelong Safety Alignment for Language Models Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.780943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.780943Z digest=sha256:5d212cdffd740d59adb9124807924de8396e82a7e6e8b33ea3af1e0459b90d94

Observation f626f28a-d3a3-47a7-b2d7-9f52a7c14ef3 · outbound

This paper cites The Llama 3 Herd of Models.

Lifelong Safety Alignment for Language Models The Llama 3 Herd of Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.866296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.866296Z digest=sha256:1b5eaf7afe73aa946a148d68aac680373807227cf9a5a1407fcf48b5b9678d59

Observation 55626e72-0c44-4da8-8dad-cfa52f80b2aa · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Lifelong Safety Alignment for Language Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:06.928931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:06.928931Z digest=sha256:d522ae7f71004d2d27cca194829ff3cac36351039baf70d63fcef125a059071d

Observation d7940072-c44b-47d1-b2a1-bcb868333d48 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Lifelong Safety Alignment for Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.052674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.052674Z digest=sha256:a6ee808d148b78b62ceeaf01054b57e2f6199d7e8c3595ff659a3c1a0431cfe4

Observation e9dce32f-a95d-4e33-b939-c6ac61df89eb · outbound

This paper cites Introducing ChatGPT, 2022.

Lifelong Safety Alignment for Language Models Introducing ChatGPT, 2022

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:15.138892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:07.143256Z digest=sha256:73abd09dd732b1d6c51427fc46be6bc97fb9dae4d2182d29ffffa5643c248657

Observation b9aaa55f-6c71-4f0e-81f0-4eb10ec73810 · outbound

This paper cites GPT-4 Technical Report.

Lifelong Safety Alignment for Language Models GPT-4 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.202081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.202081Z digest=sha256:ce16769fdeff591e8012d9ed7da4c96a21f68b3d4f34f386e98ef5233eaf6d85

Observation 65556b60-fb42-45dc-9a4d-5b4bc75d215f · outbound

This paper cites Red Teaming Language Models with Language Models.

Lifelong Safety Alignment for Language Models Red Teaming Language Models with Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.291288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.291288Z digest=sha256:87d0f7cd1a2bad8d80dfc364b631c82de4cbc69522cfc641884b84f415b897fa

Observation 849b3976-422c-4ee0-a798-c06e9eaeedad · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Lifelong Safety Alignment for Language Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.401279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.401279Z digest=sha256:3f5be18384930951a64e0f6ba57b2fbe4c8ebebbd9517275d689756c14bdea0d

Observation 0a37d858-049f-4c28-ad9d-d3eba83c208b · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

Lifelong Safety Alignment for Language Models CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.477311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.477311Z digest=sha256:3a847f1cbd35d7816db7641dbee00fc57517c5587e9bf71fb23a7103d49b140f

Observation 18bbad8a-be74-443d-8f49-58b755d27d26 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Lifelong Safety Alignment for Language Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.567600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.567600Z digest=sha256:82328f8bed4018e88f2a357fd2a033a51fea9086d96159a2877dc2099dc1f756

Observation 82d7f6fd-7a2e-4c7c-91f9-c738a20e6507 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Lifelong Safety Alignment for Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.655270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.655270Z digest=sha256:b553b8683d5e8d1242a43b0f0da233c3cbc5533e3586f3bdf724bc01e3b6fe78

Observation e4c1861b-9875-4628-86e9-e6c8780e54de · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

Lifelong Safety Alignment for Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.771929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.771929Z digest=sha256:71b8b525c01ce27a0b22926597f675fd8c802f7ec918f0656c1379e1fd4b4302

Observation 0812ed52-1180-45cd-be70-1446100ba121 · outbound

This paper cites do anything now.

Lifelong Safety Alignment for Language Models do anything now

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.878705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.878705Z digest=sha256:9e32446eb80d46f039c025af08affcb74eb3cc30d08158e7a201cc72fe2a9c37

Observation b523b869-fa67-42ca-86b2-755399ca5348 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Lifelong Safety Alignment for Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.956726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.956726Z digest=sha256:882dd942c6a8ab9179e16b0889ebea4df8cdd0f1b2486f2fe7776d9c7559646b

Observation 2522f766-b0ba-44aa-967a-5bebb3ebef10 · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Lifelong Safety Alignment for Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.034076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.034076Z digest=sha256:744eb9d7a44a13b887b1da949ba8fdbfb4c994ba335dbc5c21fd11ab697ca3ea

Observation cc0358ee-0918-4988-8548-a3f8f847270c · outbound

This paper cites Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles.

Lifelong Safety Alignment for Language Models Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.128869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.128869Z digest=sha256:27ef0b2f45dec15d1f035eb8d01a5a57898a7975e028f86d9b440d2bc8f071f0

Observation 4019a0f1-2670-49a5-87d9-04e42281de20 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Lifelong Safety Alignment for Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.229980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.229980Z digest=sha256:06d8363a9c1fc32df1dc8c98610d991118333300836523f5797528192d369cc4

Observation 724b2904-3844-42b1-a71b-b60c161a7bfa · outbound

This paper cites Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment.

Lifelong Safety Alignment for Language Models Stochastic Monkeys at Play: Random Augmentations Cheaply Break LLM Safety Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.309033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.309033Z digest=sha256:eb9f58c92857968f5f459f07cb6d1f0d773ab8213033a1b374144fc0ba4ef57f

Observation 3b319180-7c6c-4567-8f10-b5ce34044201 · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

Lifelong Safety Alignment for Language Models Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.447260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.447260Z digest=sha256:4bb0449f8866a831f0811a3c94a15465d0c26e49ca5373551b5b8763511f3729

Observation 2db408e4-ecf2-4635-ae6b-b74b8edb11a1 · outbound

This paper cites Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping.

Lifelong Safety Alignment for Language Models Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:00:12.238103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:08.558760Z digest=sha256:00b1f3c8d533bf9d31e53a3e091e5829a03d33709c0c3b5ea9f3ded91095efa8

Observation d9bbd787-d624-4161-8818-5c51ed8c466c · outbound

This paper cites Safety Reasoning with Guidelines.

Lifelong Safety Alignment for Language Models Safety Reasoning with Guidelines

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.671515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.671515Z digest=sha256:549e053539e81d81a6204811902520cf90423f3bd83f503f532d01bd7d067602

Observation bed8f58e-038b-49f4-a57c-08a4657d8522 · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.

Lifelong Safety Alignment for Language Models A comprehensive survey of continual learning: Theory, method and application

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.790652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.790652Z digest=sha256:ac571ac2ba12ab47cf6e0fea69f5412b4f89fd42614f7c8053e934d1f42f24b6

Observation d2b2cb03-e159-409a-88fa-04c041eb9997 · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

Lifelong Safety Alignment for Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.959298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:08.889926Z digest=sha256:e21e79dbef87f38d30fd15743000a28da5611ef78862458fed3b4cf0fd831a25

Observation cad982fe-1290-4b4a-915f-1dd9f427cea9 · outbound

This paper cites Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection.

Lifelong Safety Alignment for Language Models Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:08.944339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:08.944339Z digest=sha256:0a743952cf7ea6b37507df1bde8225d8ab88227b0bdd12fc5892ee94d3d8c5c9

Observation 52886168-f2d8-4c1a-8165-9f296ceb6053 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment.

Lifelong Safety Alignment for Language Models Self-Play Preference Optimization for Language Model Alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.095114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.095114Z digest=sha256:543e57ea0f71fd1b52ae1240bc91df8edc0f1ba39eedc117cfb708c5c39b2645

Observation ef8af138-15f7-4858-8ab0-fdd8e83ca8f2 · outbound

This paper cites Qwen2 Technical Report.

Lifelong Safety Alignment for Language Models Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.188727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.188727Z digest=sha256:87ae2a9a739ed5c63611adf90543de9bf52e6ac74a59c45711a1e22f079e51de

Observation 7bce9548-c0fb-4385-a5b5-0fbd53fc1461 · outbound

This paper cites Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play.

Lifelong Safety Alignment for Language Models Scalable Reinforcement Post-Training Beyond Static Human Prompts: Evolving Alignment via Asymmetric Self-Play

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.274466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.274466Z digest=sha256:8113c9bc25c2fe32631331bb28d5033a1c8cb2854c0862b42c3b8d55b9bf7595

Observation b333d74e-723b-4e2f-9f79-a3f8e37df4f0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Lifelong Safety Alignment for Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.367525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.367525Z digest=sha256:29d4c4e30f73d1c817d744fd72bb199165913e576f3d2a5a1f1247c0f0f61edc

Observation 00eaeda2-fadc-4cff-8b0c-0c3331b6c350 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Lifelong Safety Alignment for Language Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.433962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.433962Z digest=sha256:d5159d59dcc4e68fefb96467fe1ae35716d31f674c59a8547013ba57152c7c2f

Observation e5ba28bf-bc90-443c-b2b9-76d055c25e1f · outbound

This paper cites Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training.

Lifelong Safety Alignment for Language Models Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.569239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.569239Z digest=sha256:da0ea29ba6b92e15d77ab6e9751a6f4e5b383ee752133f8ae1f4ef49c8d3d583

Observation c333ee08-3f33-45eb-b0f9-1188e89bc963 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Lifelong Safety Alignment for Language Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.671636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.671636Z digest=sha256:441bad4f11d90664cd9292ee937b6ddbba5c48b9258ddadf76ebde7e43ee0fa2

Observation 0afd2886-d821-46ee-834c-dc51ce6275be · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Lifelong Safety Alignment for Language Models How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.797411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.797411Z digest=sha256:7a2c4c05f8590132fba6df35c9f1df80686fb1cf915684c6239f6c55f0edc4d9

Observation 389ca17b-67fe-4ee1-a84b-4395171c63ce · outbound

This paper cites STAIR: Improving Safety Alignment with Introspective Reasoning.

Lifelong Safety Alignment for Language Models STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:09.891654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:09.891654Z digest=sha256:7d3706abe2a9e92fcb7b9f864dedd93af3ae7ceceabdd4f606c1711f006a98d8

Observation 8057d003-cdd4-4e15-a27b-9562f1df6e40 · outbound

This paper cites Improved few-shot jailbreaking can circumvent aligned language models and their defenses.

Lifelong Safety Alignment for Language Models Improved few-shot jailbreaking can circumvent aligned language models and their defenses

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.738623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:10.035594Z digest=sha256:c6771aa9aa1f9fb3eb20c81d1d6b8cfecfbec163189eb3647fddd80aeae0ee70

Observation 7e6c0211-6d70-434c-9dfa-92cd99d84f0b · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Lifelong Safety Alignment for Language Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.147312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.147312Z digest=sha256:2640610fd9416d2cba7023b682522797d5f608ff2ce84f675deca287aa7d8a3c

Observation 098faab8-0d50-4915-a5cc-7e550b6fdf5e · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Lifelong Safety Alignment for Language Models Instruction-Following Evaluation for Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.249221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.249221Z digest=sha256:501f2635af33b0c01b6c422bc56944145ed04e584db97227abf8b9a1c9909248

Observation 0819da74-9a13-4941-8668-53d027ab7ba8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Lifelong Safety Alignment for Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.406991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.406991Z digest=sha256:fbf42fc45280158340b8e6e3aa8654c554235a7dd35c42ad41c685a6fa6ae4b8

Observation e5bd4d64-8957-49a5-af76-f3b812418876 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Lifelong Safety Alignment for Language Models Improving Alignment and Robustness with Circuit Breakers

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.571592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.571592Z digest=sha256:2aabdbb9e4d94375fadf8532415f340bdb5440b3300f12e6e1843eaabefb944f

Observation a8712d85-ecbd-45d8-8302-e6d7e2516aeb · outbound

This paper cites question.

Lifelong Safety Alignment for Language Models question

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.536150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:10.668254Z digest=sha256:f6f20ab6cfa4d6f82efe86895d540fb0108d12f571806b5580a49db216aca646

Observation 8c2e9e7d-283d-45f3-b655-4cfc11c2e46e · outbound

This paper cites This is the most preferred response when confronted with policy-violating instructions.

Lifelong Safety Alignment for Language Models This is the most preferred response when confronted with policy-violating instructions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.309703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:10.800784Z digest=sha256:4641a996621a277cf413d957dab6fb46a7071def7dff74e3094106bb85b9bc77

Observation 557685eb-0ec5-4e2c-a9b3-eb590da57da4 · outbound

This paper cites It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content.

Lifelong Safety Alignment for Language Models It either offers a warning, or provides generalized or theoretical information that stays impartial and devoid of any overtly objectionable content

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.888539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.888539Z digest=sha256:4153d18cae57530d0b110bf0416954f3658a3e2235572477a2a836d2d2688a21

Observation 00e3612b-5336-48b6-bc82-e9ee5537b7a0 · outbound

This paper cites It offers content related to the request but without embedding necessary precautions or disclaimers.

Lifelong Safety Alignment for Language Models It offers content related to the request but without embedding necessary precautions or disclaimers

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:10.967249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:10.967249Z digest=sha256:8d2fb6d1be23b3f55e38874014a21c250fcb62886503d1ea2c1e7b0721437e18

Observation 1da7b12d-24b3-4197-9137-27c4268732d4 · outbound

This paper cites However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5.

Lifelong Safety Alignment for Language Models However, the model’s response is not a direct fulfillment of the user’s intent and is not as severe as the most inappropriate response of Score 5

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:11.062426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:11.062426Z digest=sha256:69245271a89cf3515262e9150d49328d530ab89f02e8693c7498c426c7085493

Observation 2d2d1d37-669e-41c8-90fa-bbff023cecd9 · outbound

This paper cites It completely ignores the usage guidelines and provides information that clearly violates the Meta policy.

Lifelong Safety Alignment for Language Models It completely ignores the usage guidelines and provides information that clearly violates the Meta policy

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:14.070089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:11.144360Z digest=sha256:e2b1cd1c846fb56333e1effba7dcda6f00b29cbb2f9bbbd22c6367c259dd238c

Observation 7cfad83f-1844-4e05-bd39-545c05c87055 · outbound

This paper cites Frame the prompt as a discussion or analysis question to invite exploration.3.

Lifelong Safety Alignment for Language Models Frame the prompt as a discussion or analysis question to invite exploration.3

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.887723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:11.284481Z digest=sha256:cb9fa5439a17264eb85e2999a7f7bd09497331cb2c8c57468cbac0bcba8aea00

Observation 5a69f598-9492-4f2f-9fae-7ff0a3fae857 · outbound

This paper cites an unresolved cited work.

Lifelong Safety Alignment for Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:13.690433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:11.368910Z digest=sha256:7d63147738fa0c3bcac791abf4a73e98b2776e963ff78f32c840930fb74a8525

Observation 31c08b3f-f936-4835-a655-dd0b0814f96d · outbound

This paper cites 30th St, Los Angeles, CA 90007, United States.

Lifelong Safety Alignment for Language Models 30th St, Los Angeles, CA 90007, United States

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.472286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:11.477728Z digest=sha256:ad3352992c72052a0974aaa9f8f6aa57e7d9b7a7bbc317dd9513c7bb621ac2ca

Observation 40ae3064-2211-4898-a35e-6815b8005f54 · outbound

This paper cites 20th St, New York, NY 10011, United States.

Lifelong Safety Alignment for Language Models 20th St, New York, NY 10011, United States

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:13.291874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:11.616468Z digest=sha256:27bebb4233dc2d2bb7f682afff2e229c1aa6f6095ca498bd1aeb4329fa373a78

Observation 1c729329-57f8-4f55-9021-d2181740f2b0 · outbound

This paper cites an unresolved cited work.

Lifelong Safety Alignment for Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:00:13.095931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:11.708628Z digest=sha256:e3839d5dc4092f07253500bce67b2306c707e7e3665e6053d5211eb6c19c6105

Observation 2384860d-b0e0-4adb-a710-7757d15df4ec · outbound

This paper cites pythonchemicals =.

Lifelong Safety Alignment for Language Models pythonchemicals =

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:00:12.905731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:00:11.823481Z digest=sha256:a27e5a435506a572ac2197b16695814dcf813b8b1e24a2bd255bf96173b6bee9

Pith citing papers

Observation 9051076b-a6fd-4a00-90ea-312814908556 · inbound

Safety Reasoning with Guidelines cites this paper.

Safety Reasoning with Guidelines Lifelong Safety Alignment for Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T23:50:36.178215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:50:36.178215Z digest=sha256:00bc7efc7da33d467a17f0db1c0dc50705b08725beaa0befa063b21831225bce

Observation 8ff0fd93-172d-4e8f-80ff-a4cd1ad54c71 · inbound

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World cites this paper.

\texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World Lifelong Safety Alignment for Language Models

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:46.332617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:46.332617Z digest=sha256:d95a542af2ae9ae010969281c8e166528ffb5ed93a185186d373ebba0fbce40e

Observation fd4f8b5c-061e-4393-838b-cdfb51d8fe9d · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Lifelong Safety Alignment for Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:06:19.461615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:7cddad4bc9bd6c13c4261088d2d088acb1e9cb46a168ac56d5aa609e5d9a4ad5

Observation 19d35091-37dc-4c53-8a57-c613eefcd074 · inbound

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak cites this paper.

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak Lifelong Safety Alignment for Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:41.955346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-21T06:16:01.040236Z digest=sha256:2ced8cd3ff92fec60ceb2d958afdf4b8fe46eb9c1f36968ec1be5150ad3dccde