Pith. sign in

Paper Citation Record · LEDGER

Advancing LLM Safe Alignment with Safety Representation Ranking

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 11 inbound Pith citation observations for arXiv:2505.15710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15710 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:18.375972Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:28:42.361096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T07:53:14.432882Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc41ad3d-ea9b-4498-9424-1c28ce32430b · outbound

This paper cites Foundational challenges in assuring alignment and safety of large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Foundational challenges in assuring alignment and safety of large language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.847508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.209470Z digest=sha256:ed9ea25396a4ce758aedaa1f5c230cc159df3181768506a3c7686f044b0fbff6

Observation 8fcedc20-95b4-44fd-ae1d-6624ccdd3722 · outbound

This paper cites Constitutional ai: Harmlessness from ai feedback, 2022.

Advancing LLM Safe Alignment with Safety Representation Ranking Constitutional ai: Harmlessness from ai feedback, 2022

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.840918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.220989Z digest=sha256:f7e432d84f14e310ab2c46b5920c502f229ecf34cf9eff23b39fc513d51773e3

Observation bd145b7d-ecfe-4191-9a68-0afe8a21ed2a · outbound

This paper cites Safeinfer: Context adaptive decoding time safety alignment for large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Safeinfer: Context adaptive decoding time safety alignment for large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.834310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.226207Z digest=sha256:beb1b0fc92bf9b2ef1f0a6a9d130a981c4ff015e601b2608014606e861959001

Observation 9b4f2425-6ecd-4de2-b7e4-169ff89c6355 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Advancing LLM Safe Alignment with Safety Representation Ranking Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.235655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.235655Z digest=sha256:35ce8c36cfc750d393523d3b687c9d48795c8d61762a48cb10970116d64c359e

Observation 1a95981c-7205-4f03-a99f-6c0685003853 · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Advancing LLM Safe Alignment with Safety Representation Ranking Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.827732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.238736Z digest=sha256:db5eaabf06fb5563c349af1876e3d2ed2f6cb28fb188747b31ea1d398188f857

Observation c260cec5-2926-48f5-b352-de53010be1ec · outbound

This paper cites Finding safety neurons in large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Finding safety neurons in large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.241298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.241298Z digest=sha256:c197ea890a2e90764c53a85b9bf0dd7c785d95b3c8da93dde0037d1ca47a7287

Observation c2b1e6a2-1d86-41f8-b6f0-79aa48b80e32 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Advancing LLM Safe Alignment with Safety Representation Ranking Safe rlhf: Safe reinforcement learning from human feedback

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.820841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.243835Z digest=sha256:1c1bd8e0a85c6050e9fc66abef56d9388b6c28b57a2f0ff080b883a373ce1f5a

Observation bef79a3d-131a-480d-9ff0-5089a1ce3928 · outbound

This paper cites Hierarchical neural story generation.

Advancing LLM Safe Alignment with Safety Representation Ranking Hierarchical neural story generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.813974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.246767Z digest=sha256:aaf1b0d34cdd3e1ebb74cf085adfcbe0e240c93b5eb65dd9f3955dedcc225754

Observation 5e82de10-2706-46b0-907e-02a3a12c1cd0 · outbound

This paper cites Controlling linguistic style aspects in neural language generation.

Advancing LLM Safe Alignment with Safety Representation Ranking Controlling linguistic style aspects in neural language generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.807351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.249470Z digest=sha256:00abf023ce7a7628603ddcccf73461d5f51bb263f2ac54dc620b490d49042e74

Observation d8d63c05-0df6-485c-b4aa-8b231811a896 · outbound

This paper cites Evolving neural turing machines for reward-based learning.

Advancing LLM Safe Alignment with Safety Representation Ranking Evolving neural turing machines for reward-based learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.800140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.252005Z digest=sha256:c229648028a8341588e53d94566317efbe168fbf3de230d9d30f24a10094bcb9

Observation 4ff5faf9-998c-4d13-9deb-9f3ec2b9e867 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Advancing LLM Safe Alignment with Safety Representation Ranking Measuring mathematical problem solving with the math dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.793287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.254954Z digest=sha256:e72284c6c6a20335dc238c8a77a7e8db220cac37bb58398a0badc40e98182c22

Observation cf2f1b0f-e39a-49ba-bc7f-8963484a907f · outbound

This paper cites The curious case of neural text degeneration.

Advancing LLM Safe Alignment with Safety Representation Ranking The curious case of neural text degeneration

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.786568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.257253Z digest=sha256:684a69ee416916f550803fac1babbe92ffa7991511e789d5a282871fb201584d

Observation 3c067e4f-932f-472e-9bdf-b610902fc184 · outbound

This paper cites Learning to write with cooperative discriminators.

Advancing LLM Safe Alignment with Safety Representation Ranking Learning to write with cooperative discriminators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.779084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.259704Z digest=sha256:5a51234d70761496635d09d62839e91b76a8266f8f0a9d02bfbfaea2cad6703a

Observation b7270254-4aab-4948-9855-19b980d6df2c · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Advancing LLM Safe Alignment with Safety Representation Ranking Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.261712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.261712Z digest=sha256:50238554b7472a97a057c639a30a3673b5a1965ab61fef5f2b9b0f38de4827b2

Observation 18652234-1528-4a29-a5b9-14e1212899cd · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Advancing LLM Safe Alignment with Safety Representation Ranking AI Alignment: A Comprehensive Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.265028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.265028Z digest=sha256:47ac07b505f426931f8a78938691f6bcc375733f5d84b420001d4dc91a3e4a0c

Observation 29a82d71-3059-4ad4-98a7-e5da9457f925 · outbound

This paper cites Mistral 7B.

Advancing LLM Safe Alignment with Safety Representation Ranking Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.267495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.267495Z digest=sha256:19642472b9638ea77408adabb1ca38f77d07b198c336d61af96ece33967da684

Observation 5f650702-cdbb-4804-8b73-96efef81a970 · outbound

This paper cites Buckley, Jason Phang, Samuel R.

Advancing LLM Safe Alignment with Safety Representation Ranking Buckley, Jason Phang, Samuel R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.771832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.270168Z digest=sha256:d5020dffc4a5c1ace378ca5a3cd66aaddc0bffaf0a04616654eb67760a54676f

Observation c8d59d88-fdb3-48c4-967d-22cefa7f3042 · outbound

This paper cites Contrastive decoding: Open-ended text generation as optimization.

Advancing LLM Safe Alignment with Safety Representation Ranking Contrastive decoding: Open-ended text generation as optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.764757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.273152Z digest=sha256:8b49f7f4fd13834b94be09e719e0b409fca728644db4c64ebe31e39353000b67

Observation ac0a7142-37ff-468f-868a-eb0e608aa41a · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Advancing LLM Safe Alignment with Safety Representation Ranking LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.276005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.276005Z digest=sha256:9d82c57246aed697ec4afb76fb118faca7c5c71a24b3add151d137b2c287630e

Observation 4c9ed786-fd0b-4fa6-b713-689f1cbdd234 · outbound

This paper cites Jailbreaking chatgpt via prompt engineering: An empirical study, 2023.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreaking chatgpt via prompt engineering: An empirical study, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.756285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.279632Z digest=sha256:5a02dcf460811d54101c6182e92f499a7bd2f7262b9420d38f1a8ff21d1b3f34

Observation 19a89539-8a0e-4f7b-b103-30961419c693 · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

Advancing LLM Safe Alignment with Safety Representation Ranking Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.748183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.283437Z digest=sha256:4746a1416c1460253246deaa38d9cb8fb0fe4bf519c21129abceacaf1a76f531

Observation 920b1562-7e12-4e7b-bbae-27fc4630439e · outbound

This paper cites Llm improvement for jailbreak defense: Analysis through the lens of over-refusal.

Advancing LLM Safe Alignment with Safety Representation Ranking Llm improvement for jailbreak defense: Analysis through the lens of over-refusal

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.739091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.286927Z digest=sha256:046bfb40eb64c37fa870bb32eaaec066d512897809998072bfdac9fb9dd35691

Observation 1a4d2249-63e3-4d15-82be-54585b0cc9b8 · outbound

This paper cites Learning to rank from relevance judgments distributions.

Advancing LLM Safe Alignment with Safety Representation Ranking Learning to rank from relevance judgments distributions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.731853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.289693Z digest=sha256:c2ce122bb9757fcce8ef6e19bb88153e0f34775819d7417af5dfd03c8d61b7d5

Observation 3880deef-f863-478c-9248-f6abd9dd32cf · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Advancing LLM Safe Alignment with Safety Representation Ranking Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.292131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.292131Z digest=sha256:b9e7fae7c7ccfc771daf07ec9faa9f3263ef2d920d24d3a9e989c8ae2ac7619a

Observation 1e7fcf81-b543-4765-b4ab-17960c77bb26 · outbound

This paper cites Language models are unsupervised multitask learners.

Advancing LLM Safe Alignment with Safety Representation Ranking Language models are unsupervised multitask learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.724627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.294903Z digest=sha256:e64327c5d17bce0c2244b2d237675e3d9adbd3d529f8e89fbe5ad34821ddde40

Observation 9fe84c67-cc6c-4ead-b36a-217eec5e76ca · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Advancing LLM Safe Alignment with Safety Representation Ranking Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.298675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.298675Z digest=sha256:3ca74646f58ed5a667ed18da71f5ca3e6c0114d98936df5ef954b7b50e4dbc46

Observation 7dabb801-295a-4f10-89af-fd05c45770ad · outbound

This paper cites Does Representation Matter? Exploring Intermediate Layers in Large Language Models.

Advancing LLM Safe Alignment with Safety Representation Ranking Does Representation Matter? Exploring Intermediate Layers in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.301774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.301774Z digest=sha256:7a2a04ed7f0cc4b0484b889f397558f45d90370aa12616a8cf07346fa1f231c7

Observation 5a59fc6a-d359-458e-9679-afa9fe96f782 · outbound

This paper cites Reft: Reason- ing with reinforced fine-tuning.

Advancing LLM Safe Alignment with Safety Representation Ranking Reft: Reason- ing with reinforced fine-tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.717271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.305406Z digest=sha256:86896d38e6eb9bed7c9617f7b37ea5cdfe5002721261db570a729b71182e6386

Observation 3f10672a-304a-4703-8ad3-1c04084eaa3f · outbound

This paper cites Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts, 2024.

Advancing LLM Safe Alignment with Safety Representation Ranking Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.710402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.308143Z digest=sha256:edec218a2e99f997c5a01f77d213754084572ab6ac3aa6895209718c724a9dcb

Observation b13e146a-b522-4359-ae62-890b07fb123d · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models, 2023.

Advancing LLM Safe Alignment with Safety Representation Ranking Self-consistency improves chain of thought reasoning in language models, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.703409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.312113Z digest=sha256:5c6b738bd32f35ccfcbb1a1931eeea6dcc7ae61c3a5247bd903d36387d7813b4

Observation 76d87c41-7550-4fcb-803c-95150fc61c8c · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

Advancing LLM Safe Alignment with Safety Representation Ranking Chain-of-Thought Reasoning Without Prompting

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.315364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.315364Z digest=sha256:9b9e3f40791cfb437962b1a9b44b88e5f02b26ec7ef8838f21b0b16c65b4bcdb

Observation 01410d90-fb2e-45c3-b15d-3c1fdd2cf1dd · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Advancing LLM Safe Alignment with Safety Representation Ranking Reinforcement Learning for LLM Post-Training: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.318170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.318170Z digest=sha256:e485801332c08b7952a954eec115fdc09a140f835e7076d0810866ffce400253

Observation a62436cd-9a11-4008-8d0f-b9fc08dc48da · outbound

This paper cites Jailbroken: How does llm safety training fail? In NeurIPS, 2023.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbroken: How does llm safety training fail? In NeurIPS, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.696723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.322142Z digest=sha256:f633f2d96f3934e1a9ee2c22c6b9be1d3a0e75ea1df14d22c4b0da9d9b0fa71b

Observation 4ff5d838-7e2d-436e-bfdd-7c848391488d · outbound

This paper cites Assessing the brittleness of safety alignment via pruning and low-rank modifications.

Advancing LLM Safe Alignment with Safety Representation Ranking Assessing the brittleness of safety alignment via pruning and low-rank modifications

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.689585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.324741Z digest=sha256:37c3800f1e99a3a870ce143e0b672c571d2848156892ec84e77790b66b865bb6

Observation cba2e2e6-8940-425b-aecc-e698ad06b693 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.326956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.326956Z digest=sha256:50369b1210d7474575c4359cffcaf1bc7a155bc621b61a79026e45c076b2e6ae

Observation bb309155-1b6e-4320-9bc9-84c84954840b · outbound

This paper cites Sorry-bench: Systematically evaluating large language model safety refusal.

Advancing LLM Safe Alignment with Safety Representation Ranking Sorry-bench: Systematically evaluating large language model safety refusal

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.680690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.330486Z digest=sha256:c1774b63a8daaa6ec2d045ee9daefa73c45e1c96f7ebdd9dbb4393865e9c07be

Observation 3624ce13-2207-4910-9955-e7b32ff13757 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Advancing LLM Safe Alignment with Safety Representation Ranking Defending chatgpt against jailbreak attack via self-reminders

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.673601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.333273Z digest=sha256:2180e7573a3bff25384b4c00a3736fe2ca27ecf826afa3993edfa78b051bd2f9

Observation afd28f97-5b04-4eb0-97ef-e529c0821519 · outbound

This paper cites Safedecoding: Defending against jailbreak attacks via safety-aware decoding.

Advancing LLM Safe Alignment with Safety Representation Ranking Safedecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.665257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.335543Z digest=sha256:d15e493a7465e22bac724db8285ecef39d3dfad1200c6417d43bf481706f619e

Observation be8921d9-42fc-469f-bbf1-5d22e894abd2 · outbound

This paper cites SafeDecoding: Defending against jailbreak attacks via safety-aware decoding.

Advancing LLM Safe Alignment with Safety Representation Ranking SafeDecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.657477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.337971Z digest=sha256:ea7f41e4939a234f976cbe1f8aa2ac52ca0cba8ae54c9277614776fb0a9fbb85

Observation f00da5d2-9628-4f22-9a94-712d56489629 · outbound

This paper cites Qwen2.5 Technical Report.

Advancing LLM Safe Alignment with Safety Representation Ranking Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.340533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.340533Z digest=sha256:3a0d635c4c90f048289cb9acac47835502f4e48f0bce4695d94964514b164555

Observation 4897e823-b2ac-46b6-8d64-cfc0869bffb5 · outbound

This paper cites The ai alignment problem: why it is hard, and where to start.

Advancing LLM Safe Alignment with Safety Representation Ranking The ai alignment problem: why it is hard, and where to start

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.648771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.342864Z digest=sha256:04b739d21e9177f3f2595e2c62286da53bf80d8c4d31328a69ef35fd89b02240

Observation 6f8eb0eb-ac04-4593-97d1-5c4917021d8d · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search, 2024.

Advancing LLM Safe Alignment with Safety Representation Ranking Rest-mcts*: Llm self-training via process reward guided tree search, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.640308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.345365Z digest=sha256:5cbd5d92a0e3aee53038954de81192339e16400d0978e55db4c05ed34c1e8782

Observation 512519c3-7e96-443f-9916-7cb34fcf72c8 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

Advancing LLM Safe Alignment with Safety Representation Ranking Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.352301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.352301Z digest=sha256:ac693f689f9f23694dfc582be3b5417b8605ae1e8ad253c3376a222a8edb11d8

Observation 11c726b0-6815-4cf2-8850-8ae95c79ce73 · outbound

This paper cites Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models.

Advancing LLM Safe Alignment with Safety Representation Ranking Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.355760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.355760Z digest=sha256:ee5edd5af1909180ad836e24005f8e068ff179314a7b96f6d2742835b6732ef2

Observation 7335f8a0-5cf2-486a-a22c-278fdfb3b4be · outbound

This paper cites Alleviating Hallucinations of Large Language Models through Induced Hallucinations.

Advancing LLM Safe Alignment with Safety Representation Ranking Alleviating Hallucinations of Large Language Models through Induced Hallucinations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.359366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.359366Z digest=sha256:54d5c842cce9db62a6f533bdb719add68c2cf4752253389b9077b8d51ef41891

Observation 0399a2e8-9448-474a-bd85-77a28c06a640 · outbound

This paper cites Identifying and tuning safety neurons in large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Identifying and tuning safety neurons in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.630684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.363165Z digest=sha256:f36fbbb6aa7d638094daedbe25e517770f6a18c2023cd1e0fedecce4a583180c

Observation 135914eb-99a5-4ea8-9117-1f5972336177 · outbound

This paper cites On prompt-driven safeguarding for large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking On prompt-driven safeguarding for large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.622074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.366199Z digest=sha256:22673e24bffffb63969bf26a6fd063b2c1962e9eb61c7ad2bf47d2ea7b678f4f

Observation fc32a28b-3f76-470c-9485-147fba3a9090 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Advancing LLM Safe Alignment with Safety Representation Ranking Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.614468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:15:18.369535Z digest=sha256:d9a1367d16b2ea700340cfd5e0f659db519ff4f2a8958a8ff8428fa4a5710570

Observation f6fb1934-e8d9-4420-86c5-a02c5fbc6c9d · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Advancing LLM Safe Alignment with Safety Representation Ranking Representation Engineering: A Top-Down Approach to AI Transparency

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.372422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.372422Z digest=sha256:195b70b707ec30c7a5d5d11210b8921951f32fe36bf3601e64752ef5fe0b6911

Observation 2f6565bc-a792-400e-bdc0-4283ea8e43e1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Advancing LLM Safe Alignment with Safety Representation Ranking Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.375972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.375972Z digest=sha256:4cef88475c1e10831678d4ccd05f2271a9a422e2f9faf978271d0cba13cbc30f

Pith citing papers

Observation 68fbd259-3bde-4437-be77-bca5152543d5 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.979737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:0eb7ead59bc835782fab1524f811d329712a25735fe06bfd2bcf99a9807c8791

Observation c7dca5b9-3fc8-4cb6-8db9-aa8408d99216 · inbound

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems cites this paper.

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:42.361096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:28:42.361096Z digest=sha256:0619c5206c45a0852c5730b3a82f7f5199cfc675e8f26d6f70200a9841bb948a

Observation a53ea7aa-a633-4276-a61b-37f552736f54 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.595720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:aee11f0c3146e50d1a87d52f042288051e3faf2f7d2e6876f96e5b870f5ad67e

Observation a4b926b3-8045-41df-b52e-6a8e2d99b298 · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:57.877835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:12:34.909229Z digest=sha256:39f8591147501d64697a055c116addc2e7eb55b7d9b78ff36a051eab6d59152e

Observation 89529674-b1c6-4487-b39e-fb4ea1fe192f · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T16:34:03.109901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:34:03.109901Z digest=sha256:8964ea1eb15208fdd3403cb22d6d46a8908bf9e664c273abcd9e64149a4f9833

Observation 699e43b1-d54f-4e87-8814-1f71419681ee · inbound

Enabling Performant and Flexible Model-Internal Observability for LLM Inference cites this paper.

Enabling Performant and Flexible Model-Internal Observability for LLM Inference Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.099483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:17:13.823853Z digest=sha256:d3d1ac47632c3e82b7b1ba311bd3f39eb5ab37f6c63d7fdcfb363d1221676baf

Observation 0052c1e3-00a1-45f5-815f-da6460ec14ae · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.438225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:4c87aeac8ee25de9b19a8387ce2675a2ca38b0a9b0408ab808e02a23865551e4

Observation 5a07c4bf-19c0-4db0-8171-0a934fd3a2c8 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.890972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:46bfc1d744d5e36e578846cd3c137f8413eac9eb1e6498209d2c92bfc8b38285

Observation 5804f189-3924-43bc-8d23-1def0dfe31b8 · inbound

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle cites this paper.

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:58:20.492100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:55:29.094138Z digest=sha256:67cbcb8242e865ff86fae3d3986e96303f0e63c4a1ce39d3b11379c5e48e93a1

Observation 26572c6e-29ee-4df8-8692-090764d1656b · inbound

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures cites this paper.

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:14.434543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:33:30.281064Z digest=sha256:48bbb62dce569fc17410cc580c6decf6a7ebfcd2eda63373cbe137a51a5f330b

Observation cd001eb1-32a0-4b4e-8862-f544de839176 · inbound

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs cites this paper.

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T23:07:40.652133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:07:40.652133Z digest=sha256:ea388c77d20f44dbcd7017ad320192892f837cafe3bab32b801c08b4769d9ed5