Pith. sign in

Paper Citation Record · LEDGER

Advancing LLM Safe Alignment with Safety Representation Ranking

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 11 inbound Pith citation observations for arXiv:2505.15710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15710 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:18.375972Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:28:42.361096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T07:53:14.432882Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc41ad3d-ea9b-4498-9424-1c28ce32430b · outbound

This paper cites Foundational challenges in assuring alignment and safety of large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Foundational challenges in assuring alignment and safety of large language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.847508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.209470Z digest=sha256:734ff41992972a00f1406b3056c7ec9558d9202478b9203876d090e718e8a019

Observation 8fcedc20-95b4-44fd-ae1d-6624ccdd3722 · outbound

This paper cites Constitutional ai: Harmlessness from ai feedback, 2022.

Advancing LLM Safe Alignment with Safety Representation Ranking Constitutional ai: Harmlessness from ai feedback, 2022

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.840918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.220989Z digest=sha256:6b4263951e2b8d0d8a058158bffd4cfb9ce3707cea733195342472351f542900

Observation bd145b7d-ecfe-4191-9a68-0afe8a21ed2a · outbound

This paper cites Safeinfer: Context adaptive decoding time safety alignment for large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Safeinfer: Context adaptive decoding time safety alignment for large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.834310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.226207Z digest=sha256:57d16551e2f82474bf41eed589ecfc84c570e9b4a8598a8fb60452c931c760ab

Observation 9b4f2425-6ecd-4de2-b7e4-169ff89c6355 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Advancing LLM Safe Alignment with Safety Representation Ranking Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.235655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.235655Z digest=sha256:c924343bc893b545f71207b7bb4c85ea29806499aa0de4d053ebe46e544acf52

Observation 1a95981c-7205-4f03-a99f-6c0685003853 · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Advancing LLM Safe Alignment with Safety Representation Ranking Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.827732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.238736Z digest=sha256:8ed97155cb872db3e63fe55e2a01196c7ee806709131a890101b3c1d3dd4767d

Observation c260cec5-2926-48f5-b352-de53010be1ec · outbound

This paper cites Finding safety neurons in large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Finding safety neurons in large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.241298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.241298Z digest=sha256:99934f1ec96d4e687f4290a133468d18676de2b2af7c854a3cd5aef53c2c5426

Observation c2b1e6a2-1d86-41f8-b6f0-79aa48b80e32 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Advancing LLM Safe Alignment with Safety Representation Ranking Safe rlhf: Safe reinforcement learning from human feedback

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.820841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.243835Z digest=sha256:926ddbc2855bfacb2590cac376c5d010d392789a22ea57efb316c9f5dace080a

Observation bef79a3d-131a-480d-9ff0-5089a1ce3928 · outbound

This paper cites Hierarchical neural story generation.

Advancing LLM Safe Alignment with Safety Representation Ranking Hierarchical neural story generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.813974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.246767Z digest=sha256:440d5e7ad48068e8e7e6e752447275f5fae4ff94fb44e10c7a791dc3b76e9751

Observation 5e82de10-2706-46b0-907e-02a3a12c1cd0 · outbound

This paper cites Controlling linguistic style aspects in neural language generation.

Advancing LLM Safe Alignment with Safety Representation Ranking Controlling linguistic style aspects in neural language generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.807351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.249470Z digest=sha256:7e92ea5e18d4f67de47b7158104023eddeca64ac6e1451047ff767c50c70aaa7

Observation d8d63c05-0df6-485c-b4aa-8b231811a896 · outbound

This paper cites Evolving neural turing machines for reward-based learning.

Advancing LLM Safe Alignment with Safety Representation Ranking Evolving neural turing machines for reward-based learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.800140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.252005Z digest=sha256:ea90f9e6f53300a1107a0dfeb09a680fbf81058645a49ce459071779a6554bd4

Observation 4ff5faf9-998c-4d13-9deb-9f3ec2b9e867 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Advancing LLM Safe Alignment with Safety Representation Ranking Measuring mathematical problem solving with the math dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.793287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.254954Z digest=sha256:897af56f218351841323d2398340c296c102b9804a300f49cf3240460626b123

Observation cf2f1b0f-e39a-49ba-bc7f-8963484a907f · outbound

This paper cites The curious case of neural text degeneration.

Advancing LLM Safe Alignment with Safety Representation Ranking The curious case of neural text degeneration

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.786568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.257253Z digest=sha256:3b8303380f4d0e7311f516e92e1d0e4dc287646c99d1c0716899f5e5dac3a995

Observation 3c067e4f-932f-472e-9bdf-b610902fc184 · outbound

This paper cites Learning to write with cooperative discriminators.

Advancing LLM Safe Alignment with Safety Representation Ranking Learning to write with cooperative discriminators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.779084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.259704Z digest=sha256:8b59b38aa576c775ef8016e66c89119e2c1a64b5fd22f358c4c81e268088e91a

Observation b7270254-4aab-4948-9855-19b980d6df2c · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Advancing LLM Safe Alignment with Safety Representation Ranking Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.261712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.261712Z digest=sha256:9c5f97baa18e29784063421599b3c2dd6da9d89a815a2cca8340082ca9a53021

Observation 18652234-1528-4a29-a5b9-14e1212899cd · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Advancing LLM Safe Alignment with Safety Representation Ranking AI Alignment: A Comprehensive Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.265028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.265028Z digest=sha256:f6326ad052f81208bd0ecd58ab856f9f65d2c531f2438f8fb76f95c1b28c46ce

Observation 29a82d71-3059-4ad4-98a7-e5da9457f925 · outbound

This paper cites Mistral 7B.

Advancing LLM Safe Alignment with Safety Representation Ranking Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.267495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.267495Z digest=sha256:bf6bfeb16619e00144f53bdb580ef63f75e6d3000856687afd5aa55c8d1d7d2d

Observation 5f650702-cdbb-4804-8b73-96efef81a970 · outbound

This paper cites Buckley, Jason Phang, Samuel R.

Advancing LLM Safe Alignment with Safety Representation Ranking Buckley, Jason Phang, Samuel R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.771832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.270168Z digest=sha256:bcd28874847d56534f8858a1b73f2ca81e4323783f711017840f1842583d3a13

Observation c8d59d88-fdb3-48c4-967d-22cefa7f3042 · outbound

This paper cites Contrastive decoding: Open-ended text generation as optimization.

Advancing LLM Safe Alignment with Safety Representation Ranking Contrastive decoding: Open-ended text generation as optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.764757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.273152Z digest=sha256:7b0d77386969c1cddf3c32ec68778559aa605166d944bc32b93c8f4d8248843d

Observation ac0a7142-37ff-468f-868a-eb0e608aa41a · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Advancing LLM Safe Alignment with Safety Representation Ranking LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.276005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.276005Z digest=sha256:5207a7a08ddc23a8aa7af042267e03e300ab6ede92f0f8854fe82483c01854d6

Observation 4c9ed786-fd0b-4fa6-b713-689f1cbdd234 · outbound

This paper cites Jailbreaking chatgpt via prompt engineering: An empirical study, 2023.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreaking chatgpt via prompt engineering: An empirical study, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.756285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.279632Z digest=sha256:5e3dce17b972a5750ec5fecbef32790341d09fbd19b4d6a9515b38f8c0a08ec2

Observation 19a89539-8a0e-4f7b-b103-30961419c693 · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

Advancing LLM Safe Alignment with Safety Representation Ranking Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.748183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.283437Z digest=sha256:919c6edccf4c893b0a0c300a197e509cc728c712d53e5748ca19a6d039ac5410

Observation 920b1562-7e12-4e7b-bbae-27fc4630439e · outbound

This paper cites Llm improvement for jailbreak defense: Analysis through the lens of over-refusal.

Advancing LLM Safe Alignment with Safety Representation Ranking Llm improvement for jailbreak defense: Analysis through the lens of over-refusal

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.739091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.286927Z digest=sha256:d76a976061de3d34fb863d485278f2345081db40f13ed5cd6bc615fc59a856f1

Observation 1a4d2249-63e3-4d15-82be-54585b0cc9b8 · outbound

This paper cites Learning to rank from relevance judgments distributions.

Advancing LLM Safe Alignment with Safety Representation Ranking Learning to rank from relevance judgments distributions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.731853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.289693Z digest=sha256:fc4c2751373bdf561ce45779d70558aafab0cd3585996cd4f0726362ba10a3e4

Observation 3880deef-f863-478c-9248-f6abd9dd32cf · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Advancing LLM Safe Alignment with Safety Representation Ranking Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.292131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.292131Z digest=sha256:e1c569d99c8debfbcd99a341b377a463f75ef8db09197b20cb68bdc3aaa78de4

Observation 1e7fcf81-b543-4765-b4ab-17960c77bb26 · outbound

This paper cites Language models are unsupervised multitask learners.

Advancing LLM Safe Alignment with Safety Representation Ranking Language models are unsupervised multitask learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.724627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.294903Z digest=sha256:917968b13fd85759591a1073bb87b2e572ca100e28bd97a67ef8fd16c54c0c64

Observation 9fe84c67-cc6c-4ead-b36a-217eec5e76ca · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Advancing LLM Safe Alignment with Safety Representation Ranking Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.298675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.298675Z digest=sha256:d668d13274f7c01d9dcde5f142fc9b43bba7358909b11ac76f63a7486c842bd7

Observation 7dabb801-295a-4f10-89af-fd05c45770ad · outbound

This paper cites Does Representation Matter? Exploring Intermediate Layers in Large Language Models.

Advancing LLM Safe Alignment with Safety Representation Ranking Does Representation Matter? Exploring Intermediate Layers in Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.301774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.301774Z digest=sha256:0e26801c8964569bbc1103315989d4dd77bc057dab080af8ace50948f1fce81e

Observation 5a59fc6a-d359-458e-9679-afa9fe96f782 · outbound

This paper cites Reft: Reason- ing with reinforced fine-tuning.

Advancing LLM Safe Alignment with Safety Representation Ranking Reft: Reason- ing with reinforced fine-tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.717271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.305406Z digest=sha256:e141e2582a1fe2c30d65fd714c6e742f9c7df38b049b69568965cda242fb5163

Observation 3f10672a-304a-4703-8ad3-1c04084eaa3f · outbound

This paper cites Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts, 2024.

Advancing LLM Safe Alignment with Safety Representation Ranking Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.710402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.308143Z digest=sha256:09a8b01abc70a9f9359bd03df2db8ea648a13accff535c50b85c1f81ec273449

Observation b13e146a-b522-4359-ae62-890b07fb123d · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models, 2023.

Advancing LLM Safe Alignment with Safety Representation Ranking Self-consistency improves chain of thought reasoning in language models, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.703409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.312113Z digest=sha256:fdddd59ec1963d7896d15158112fd7a63544627e6fc479570dc29d52c5849aab

Observation 76d87c41-7550-4fcb-803c-95150fc61c8c · outbound

This paper cites Chain-of-Thought Reasoning Without Prompting.

Advancing LLM Safe Alignment with Safety Representation Ranking Chain-of-Thought Reasoning Without Prompting

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.315364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.315364Z digest=sha256:ee3f000bcc4372edaab887d9f75b50b415ea8021952651d130fc76661df52bf7

Observation 01410d90-fb2e-45c3-b15d-3c1fdd2cf1dd · outbound

This paper cites Reinforcement Learning for LLM Post-Training: A Survey.

Advancing LLM Safe Alignment with Safety Representation Ranking Reinforcement Learning for LLM Post-Training: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.318170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.318170Z digest=sha256:99f295c5ab97cb730199ab838e81be6975f2a95c5add373a42e72906e9cdfbe8

Observation a62436cd-9a11-4008-8d0f-b9fc08dc48da · outbound

This paper cites Jailbroken: How does llm safety training fail? In NeurIPS, 2023.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbroken: How does llm safety training fail? In NeurIPS, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.696723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.322142Z digest=sha256:363d659d48fb4b97aea89b22227627fb7d2942bce0bfc3f5ac70ea96835f0fc1

Observation 4ff5d838-7e2d-436e-bfdd-7c848391488d · outbound

This paper cites Assessing the brittleness of safety alignment via pruning and low-rank modifications.

Advancing LLM Safe Alignment with Safety Representation Ranking Assessing the brittleness of safety alignment via pruning and low-rank modifications

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.689585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.324741Z digest=sha256:68c659b2151f8af65729e77b629d0d0a22fef349674ffc5baa29726f2506a062

Observation cba2e2e6-8940-425b-aecc-e698ad06b693 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.326956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.326956Z digest=sha256:4bf3b663657782febaf463f3d1afa936705c2f449a38d1b116f97995aa9b7e54

Observation bb309155-1b6e-4320-9bc9-84c84954840b · outbound

This paper cites Sorry-bench: Systematically evaluating large language model safety refusal.

Advancing LLM Safe Alignment with Safety Representation Ranking Sorry-bench: Systematically evaluating large language model safety refusal

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.680690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.330486Z digest=sha256:9959a7bdc38c33cf9bc772bb8e84aaeb47e15f50132e4914bcd19ebd6c21bbb0

Observation 3624ce13-2207-4910-9955-e7b32ff13757 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Advancing LLM Safe Alignment with Safety Representation Ranking Defending chatgpt against jailbreak attack via self-reminders

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.673601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.333273Z digest=sha256:66221ed707942b33f531221980bd9296c6b7e595b3f97bb2b04273111f4d47d6

Observation afd28f97-5b04-4eb0-97ef-e529c0821519 · outbound

This paper cites Safedecoding: Defending against jailbreak attacks via safety-aware decoding.

Advancing LLM Safe Alignment with Safety Representation Ranking Safedecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.665257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.335543Z digest=sha256:1147739f09bbc6feff5edf8cb7cd540063d7add98d3360ce21a905221e73ccfe

Observation be8921d9-42fc-469f-bbf1-5d22e894abd2 · outbound

This paper cites SafeDecoding: Defending against jailbreak attacks via safety-aware decoding.

Advancing LLM Safe Alignment with Safety Representation Ranking SafeDecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.657477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.337971Z digest=sha256:0d01e887c6cfb8e570a281413b125735a0e6b78184fbc6a4905ec19397c1513e

Observation f00da5d2-9628-4f22-9a94-712d56489629 · outbound

This paper cites Qwen2.5 Technical Report.

Advancing LLM Safe Alignment with Safety Representation Ranking Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.340533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.340533Z digest=sha256:0436d248ffb8677d04696fbf5af9c044eb26d397042fc5bd2990f924935672fa

Observation 4897e823-b2ac-46b6-8d64-cfc0869bffb5 · outbound

This paper cites The ai alignment problem: why it is hard, and where to start.

Advancing LLM Safe Alignment with Safety Representation Ranking The ai alignment problem: why it is hard, and where to start

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.648771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.342864Z digest=sha256:8bd2daef872073e3d366abc4216ee979527103c05c75a88234801cbf2e4a3c45

Observation 6f8eb0eb-ac04-4593-97d1-5c4917021d8d · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search, 2024.

Advancing LLM Safe Alignment with Safety Representation Ranking Rest-mcts*: Llm self-training via process reward guided tree search, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.640308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.345365Z digest=sha256:87bc93cc9cfea964c138c4d1e9eccd15dc7454ef9bd40b12daab075c7ef1cb80

Observation 512519c3-7e96-443f-9916-7cb34fcf72c8 · outbound

This paper cites Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment.

Advancing LLM Safe Alignment with Safety Representation Ranking Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.352301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.352301Z digest=sha256:297e04a2edc992ae5fd8e052ea69b8e1a73783174bfbd6b30732bdf8b76200b3

Observation 11c726b0-6815-4cf2-8850-8ae95c79ce73 · outbound

This paper cites Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models.

Advancing LLM Safe Alignment with Safety Representation Ranking Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.355760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.355760Z digest=sha256:caa4ea2ac06539a2f4c5a7174b6baf36cc2f79bd9f194ff4bb2ed0634bbf246b

Observation 7335f8a0-5cf2-486a-a22c-278fdfb3b4be · outbound

This paper cites Alleviating Hallucinations of Large Language Models through Induced Hallucinations.

Advancing LLM Safe Alignment with Safety Representation Ranking Alleviating Hallucinations of Large Language Models through Induced Hallucinations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.359366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.359366Z digest=sha256:c64cd4094d6895296ccb1e7f2d171bf17f5a21ab76fdec6d965ef7482cd721fb

Observation 0399a2e8-9448-474a-bd85-77a28c06a640 · outbound

This paper cites Identifying and tuning safety neurons in large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking Identifying and tuning safety neurons in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.630684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.363165Z digest=sha256:8658ff441589bd4fb479719d02c800531b3068918ad18527dc088a7b50f319cc

Observation 135914eb-99a5-4ea8-9117-1f5972336177 · outbound

This paper cites On prompt-driven safeguarding for large language models.

Advancing LLM Safe Alignment with Safety Representation Ranking On prompt-driven safeguarding for large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.622074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.366199Z digest=sha256:3400dbded50e0ce57364b1aab2d670db5f1d7950ed32b0656c36703f99cd0dae

Observation fc32a28b-3f76-470c-9485-147fba3a9090 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Advancing LLM Safe Alignment with Safety Representation Ranking Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:15:18.614468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:15:18.369535Z digest=sha256:aefb9f6f1296f05d0fedc09e1521bb2480076a67c15d61c5ccbc03db00f48e70

Observation f6fb1934-e8d9-4420-86c5-a02c5fbc6c9d · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Advancing LLM Safe Alignment with Safety Representation Ranking Representation Engineering: A Top-Down Approach to AI Transparency

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.372422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.372422Z digest=sha256:c70b6ce84c55d0c49cc64913e0250d10c322202cafee3f81d472d0b71ecca7b0

Observation 2f6565bc-a792-400e-bdc0-4283ea8e43e1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Advancing LLM Safe Alignment with Safety Representation Ranking Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.375972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.375972Z digest=sha256:67b7f0708b37a98792e7a1f6334864966031c59960c532a782cf335b735cdf30

Pith citing papers

Observation 68fbd259-3bde-4437-be77-bca5152543d5 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.979737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:61819b08a86a3d9540e5e220a47d7f0b3c6971256abfbaf3cd375e018b67b6f2

Observation c7dca5b9-3fc8-4cb6-8db9-aa8408d99216 · inbound

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems cites this paper.

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:42.361096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:28:42.361096Z digest=sha256:09c58b41c3ad9b542471c69e7a90f6a0e4fc6ba02b0662d695da12f3f84c990d

Observation a53ea7aa-a633-4276-a61b-37f552736f54 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.595720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:3b2c8c109f38bdb5c0e7a568344900ba893c25a130c2999ef0b99a04717898eb

Observation a4b926b3-8045-41df-b52e-6a8e2d99b298 · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:57.877835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:12:34.909229Z digest=sha256:731c07a7bb5e913303a5eb9f36714b538da10f79ce8f5847e74dbf56e2ac1062

Observation 89529674-b1c6-4487-b39e-fb4ea1fe192f · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T16:34:03.109901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:34:03.109901Z digest=sha256:60109ce0c8cc9e06b307d5d115c8312e34ebeab6838d49a0a35cf07af058076c

Observation 699e43b1-d54f-4e87-8814-1f71419681ee · inbound

Enabling Performant and Flexible Model-Internal Observability for LLM Inference cites this paper.

Enabling Performant and Flexible Model-Internal Observability for LLM Inference Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:17:28.099483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T07:17:13.823853Z digest=sha256:e670c50772ca28f942e12fe1f58c3e684deae8f1706fd2be613cd40e8c00022e

Observation 0052c1e3-00a1-45f5-815f-da6460ec14ae · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.438225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:7a120b7e5be8493d0ce853959c3447a7dc5467ec640ec81db4ab66ab3a592d1b

Observation 5a07c4bf-19c0-4db0-8171-0a934fd3a2c8 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.890972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:8e9fdd8e7bd7c22f6c401fdd103dcff65b3f37ffdb06df4b9fcad1c078fb657c

Observation 5804f189-3924-43bc-8d23-1def0dfe31b8 · inbound

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle cites this paper.

A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:58:20.492100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T13:55:29.094138Z digest=sha256:d2525ae8d264c883101de8d0d11378f38a3c2103cbf82a6fc3b757d1cf19d147

Observation 26572c6e-29ee-4df8-8692-090764d1656b · inbound

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures cites this paper.

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:14.434543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T07:33:30.281064Z digest=sha256:8139bea542fe692a135a3f82465357627400f85b9cad01ca6bca15e1acddde6b

Observation cd001eb1-32a0-4b4e-8862-f544de839176 · inbound

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs cites this paper.

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs Advancing LLM Safe Alignment with Safety Representation Ranking

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T23:07:40.652133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:07:40.652133Z digest=sha256:cbabaeb42bc06669fa8ec71585a644ec7a023e1c18112ac409a353d00afd9119