Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:18.375972Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 11 inbound Pith citation observations for arXiv:2505.15710.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:18.375972Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:28:42.361096Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T07:53:14.432882Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cc41ad3d-ea9b-4498-9424-1c28ce32430b · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Foundational challenges in assuring alignment and safety of large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8fcedc20-95b4-44fd-ae1d-6624ccdd3722 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Constitutional ai: Harmlessness from ai feedback, 2022
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bd145b7d-ecfe-4191-9a68-0afe8a21ed2a · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Safeinfer: Context adaptive decoding time safety alignment for large language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b4f2425-6ecd-4de2-b7e4-169ff89c6355 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a95981c-7205-4f03-a99f-6c0685003853 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Pappas, Florian Tramer, Hamed Hassani, and Eric Wong
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c260cec5-2926-48f5-b352-de53010be1ec · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Finding safety neurons in large language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b1e6a2-1d86-41f8-b6f0-79aa48b80e32 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Safe rlhf: Safe reinforcement learning from human feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bef79a3d-131a-480d-9ff0-5089a1ce3928 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Hierarchical neural story generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e82de10-2706-46b0-907e-02a3a12c1cd0 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Controlling linguistic style aspects in neural language generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d8d63c05-0df6-485c-b4aa-8b231811a896 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Evolving neural turing machines for reward-based learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ff5faf9-998c-4d13-9deb-9f3ec2b9e867 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Measuring mathematical problem solving with the math dataset
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf2f1b0f-e39a-49ba-bc7f-8963484a907f · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking The curious case of neural text degeneration
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c067e4f-932f-472e-9bdf-b610902fc184 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Learning to write with cooperative discriminators
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b7270254-4aab-4948-9855-19b980d6df2c · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18652234-1528-4a29-a5b9-14e1212899cd · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking AI Alignment: A Comprehensive Survey
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a82d71-3059-4ad4-98a7-e5da9457f925 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Mistral 7B
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f650702-cdbb-4804-8b73-96efef81a970 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Buckley, Jason Phang, Samuel R
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c8d59d88-fdb3-48c4-967d-22cefa7f3042 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Contrastive decoding: Open-ended text generation as optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac0a7142-37ff-468f-868a-eb0e608aa41a · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking LiPO: Listwise Preference Optimization through Learning-to-Rank
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c9ed786-fd0b-4fa6-b713-689f1cbdd234 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreaking chatgpt via prompt engineering: An empirical study, 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19a89539-8a0e-4f7b-b103-30961419c693 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 920b1562-7e12-4e7b-bbae-27fc4630439e · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Llm improvement for jailbreak defense: Analysis through the lens of over-refusal
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a4d2249-63e3-4d15-82be-54585b0cc9b8 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Learning to rank from relevance judgments distributions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3880deef-f863-478c-9248-f6abd9dd32cf · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7fcf81-b543-4765-b4ab-17960c77bb26 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Language models are unsupervised multitask learners
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9fe84c67-cc6c-4ead-b36a-217eec5e76ca · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dabb801-295a-4f10-89af-fd05c45770ad · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Does Representation Matter? Exploring Intermediate Layers in Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a59fc6a-d359-458e-9679-afa9fe96f782 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Reft: Reason- ing with reinforced fine-tuning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3f10672a-304a-4703-8ad3-1c04084eaa3f · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b13e146a-b522-4359-ae62-890b07fb123d · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Self-consistency improves chain of thought reasoning in language models, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 76d87c41-7550-4fcb-803c-95150fc61c8c · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Chain-of-Thought Reasoning Without Prompting
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01410d90-fb2e-45c3-b15d-3c1fdd2cf1dd · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Reinforcement Learning for LLM Post-Training: A Survey
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62436cd-9a11-4008-8d0f-b9fc08dc48da · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Jailbroken: How does llm safety training fail? In NeurIPS, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ff5d838-7e2d-436e-bfdd-7c848391488d · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Assessing the brittleness of safety alignment via pruning and low-rank modifications
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cba2e2e6-8940-425b-aecc-e698ad06b693 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb309155-1b6e-4320-9bc9-84c84954840b · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Sorry-bench: Systematically evaluating large language model safety refusal
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3624ce13-2207-4910-9955-e7b32ff13757 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Defending chatgpt against jailbreak attack via self-reminders
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afd28f97-5b04-4eb0-97ef-e529c0821519 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation be8921d9-42fc-469f-bbf1-5d22e894abd2 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking SafeDecoding: Defending against jailbreak attacks via safety-aware decoding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f00da5d2-9628-4f22-9a94-712d56489629 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Qwen2.5 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4897e823-b2ac-46b6-8d64-cfc0869bffb5 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking The ai alignment problem: why it is hard, and where to start
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6f8eb0eb-ac04-4593-97d1-5c4917021d8d · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Rest-mcts*: Llm self-training via process reward guided tree search, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 512519c3-7e96-443f-9916-7cb34fcf72c8 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c726b0-6815-4cf2-8850-8ae95c79ce73 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7335f8a0-5cf2-486a-a22c-278fdfb3b4be · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Alleviating Hallucinations of Large Language Models through Induced Hallucinations
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0399a2e8-9448-474a-bd85-77a28c06a640 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Identifying and tuning safety neurons in large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 135914eb-99a5-4ea8-9117-1f5972336177 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking On prompt-driven safeguarding for large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc32a28b-3f76-470c-9485-147fba3a9090 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6fb1934-e8d9-4420-86c5-a02c5fbc6c9d · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Representation Engineering: A Top-Down Approach to AI Transparency
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6565bc-a792-400e-bdc0-4283ea8e43e1 · outbound
Advancing LLM Safe Alignment with Safety Representation Ranking Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fbd259-3bde-4437-be77-bca5152543d5 · inbound
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c7dca5b9-3fc8-4cb6-8db9-aa8408d99216 · inbound
SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a53ea7aa-a633-4276-a61b-37f552736f54 · inbound
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4b926b3-8045-41df-b52e-6a8e2d99b298 · inbound
Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89529674-b1c6-4487-b39e-fb4ea1fe192f · inbound
Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699e43b1-d54f-4e87-8814-1f71419681ee · inbound
Enabling Performant and Flexible Model-Internal Observability for LLM Inference Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0052c1e3-00a1-45f5-815f-da6460ec14ae · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a07c4bf-19c0-4db0-8171-0a934fd3a2c8 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5804f189-3924-43bc-8d23-1def0dfe31b8 · inbound
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 26572c6e-29ee-4df8-8692-090764d1656b · inbound
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd001eb1-32a0-4b4e-8862-f544de839176 · inbound
One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs Advancing LLM Safe Alignment with Safety Representation Ranking
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.