Pith. sign in

Paper Citation Record · LEDGER

On Almost Surely Safe Alignment of Large Language Models at Inference-Time

As of 10 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 1 inbound Pith citation observation for arXiv:2502.01208.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01208 v3

Coverage vector

measured 100 of 125 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:18:40.969178Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:07.249475Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:43:07.622958Z

Reference resolution

100 of 125 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved94
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ac5fa90-7d16-4348-a5ad-6f20f83efd4c · outbound

This paper cites An empirical survey on long document summarization: Datasets, models, and metrics.ACM computing surveys, 55(8):1–35, 2022.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time An empirical survey on long document summarization: Datasets, models, and metrics.ACM computing surveys, 55(8):1–35, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.470256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.470256Z digest=sha256:5e21f44f1669518e3609027559cf7148faf41d0f415f43998ef36b71d43d1ae2

Observation 1181761a-d634-4615-af8e-de209e60ba4f · outbound

This paper cites Learning to summarize with human feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning to summarize with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.476598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.476598Z digest=sha256:cc3ba331d0ed72f4cb5adfdaf7906e3570f12a55d5754f1990a28dcad37c2885

Observation b1e88c8a-2a73-44df-9c95-7a9713802ee6 · outbound

This paper cites Pal: Program-aided language models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Pal: Program-aided language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.481857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.481857Z digest=sha256:d4ffa2768f5390c94da6cd662dbc9f898dfc02b4103b868b13dc5f56eaafddfa

Observation f6d39b8d-bfc0-4123-823f-368013e915a5 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.487014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.487014Z digest=sha256:ae82e3f0dc2615eedd7eece6bc59f984129aca280d4a6ec986bccc26e16bff18

Observation 1e67b72c-ab4c-4c90-9378-ac0cbdc4abfc · outbound

This paper cites ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.492246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.492246Z digest=sha256:1bd3687a1620b0e9131c47efce8157de2d05c317843ef3699efc9f075aaab338

Observation af3910d5-6cb5-4ca6-8fb1-c8032dc86753 · outbound

This paper cites A sur- vey on integration of large language models with intelligent robots.Intelligent Service Robotics, 17(5):1091–1107, August 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time A sur- vey on integration of large language models with intelligent robots.Intelligent Service Robotics, 17(5):1091–1107, August 2024

Reference 6

Resolution
verified exact
doi, observed 2026-08-09T16:18:41.128413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T16:18:40.497065Z digest=sha256:ed0dc24fbddac5a1fe5b193bf6af3b749c007c1d7f24bbf7f48001c9c6687f20

Observation 2025b726-2696-4717-be04-c3c9820a338c · outbound

This paper cites Toxicity in ChatGPT: Analyzing Persona-assigned Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Toxicity in ChatGPT: Analyzing Persona-assigned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.502533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.502533Z digest=sha256:ecf15e6520a5cc383eab011915a337a54c32677ea184cf4ed496620981ed394e

Observation ef519bac-1552-4b43-ac33-199be15ae185 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.507237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.507237Z digest=sha256:db02de018e9b8e348d59300555c58abe32eb78cd74d86ed4616c733846890fb8

Observation 64a6923b-99a6-47c1-ab70-77ddc7620e35 · outbound

This paper cites Ethical and social risks of harm from Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Ethical and social risks of harm from Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.512328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.512328Z digest=sha256:b76fcf585ed84369d8bf1fcaeac278eb631a7bd04d1a3fd6b95465404d421528

Observation db4b25dc-92a5-41ff-b2e2-dd61d1ec7116 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.517461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.517461Z digest=sha256:956a88c1db18987a8dc7d58c5d073a1fd4aa9e78cc8ad9b6bc5901e92ff436e1

Observation f199c6c5-b8ec-4f5b-b0e0-a29eb46ccddb · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.522572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.522572Z digest=sha256:e21cce90963d38e375e2ea1c17230413b22e165fdddccd8afb9fcc14af0445f8

Observation 648cc058-61a0-4053-8b17-978cb4b54180 · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.527252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.527252Z digest=sha256:d0f999b05302d1f40012c4e6f87191092ba88a8531809225f563d4e051f1b307

Observation dba0b472-7b72-4142-8613-083d21d44aca · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time WebGPT: Browser-assisted question-answering with human feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.531827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.531827Z digest=sha256:7c1618596fda4b7c6e25c4f32d923c167d68a9095153fafe276303b7683eb9e1

Observation 227a9f82-3d23-44eb-8d64-70512565a5d7 · outbound

This paper cites Controlled Decoding from Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Controlled Decoding from Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.536823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.536823Z digest=sha256:26c630998a49dab8eb9cbdf9efa2933fdedfecf1fc11facbd7bd4afc44c0e211

Observation 294b9ea5-5545-462b-97b3-54e10ec76f02 · outbound

This paper cites Discrete-time markov control processes with discounted unbounded costs: optimality criteria.Kybernetika, 28(3):191–212, 1992.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Discrete-time markov control processes with discounted unbounded costs: optimality criteria.Kybernetika, 28(3):191–212, 1992

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.542544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.542544Z digest=sha256:518cae9715becc5637947fe3122c9936cfcb139b6655178895cb77ace8f2b7bc

Observation ee21386e-f652-4cd8-9501-9486c02f365e · outbound

This paper cites Sauté rl: Almost surely safe reinforcement learning using state augmentation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Sauté rl: Almost surely safe reinforcement learning using state augmentation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.547563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.547563Z digest=sha256:b35d99d109cc2b75e8ff2262e07a772d73d2dd2d64469991877bc2d41392e072

Observation b7c85b2a-13af-4795-abf5-a1817f9f6c71 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fine-Tuning Language Models from Human Preferences

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.552457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.552457Z digest=sha256:769142d1adc1a39ab8212acd618ca18a6bb3a8e0de0c13f860050ffac9bc0954

Observation 3fd0f1f1-d309-4023-bd13-2b80c4327f83 · outbound

This paper cites Proximal Policy Optimization Algorithms.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.557404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.557404Z digest=sha256:cbcf9f4e0f9d5b44b95e402bafb19e29141632e8f8e5c7eb7f400fdcff790c61

Observation dbf5f19f-ae5a-405f-8053-bb6ad02842ff · outbound

This paper cites Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.562457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.562457Z digest=sha256:d27c7154f59735199902a2e3462aad4408b027e95bfe4027f4c7cc0a8dfa48cd

Observation 649dace3-01ac-4ac4-8cca-b3a0e6853cb3 · outbound

This paper cites Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.567531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.567531Z digest=sha256:f70c67b3fadf332a7b0d9a8ca1c2a0ded2bdb2461239841eca17edb3dd6c79f2

Observation 90d13abe-0606-46a8-be83-baf1426b7be6 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.572720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.572720Z digest=sha256:1369a0e071d055dc70b1ac8942c6cae88101e0f508e183ee09229155cf031f31

Observation 7699aa7b-11f8-4f6e-8770-38fe3745ab72 · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.577727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.577727Z digest=sha256:5e73e28f4a817ecbf4861a9c5d7812eb5a8ae2b731ed541d8d0db55144962ae1

Observation a6eccc06-70ba-46d1-a3e0-5158e1fbea3c · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.583638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.583638Z digest=sha256:b41dab1a4fe38c44b6dce9322da68a9d54edabfd8316567a5eb13f4a81ffd316

Observation 2cd7047a-8241-4f99-aed5-d87a2d392949 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.588740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.588740Z digest=sha256:318431f0dd93401b4c59a734cabaee06bc753fb6abd1016a50965805c5e1de7f

Observation 6976cd11-331a-4878-8740-128ed3baadfe · outbound

This paper cites Preference ranking optimization for human alignment.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Preference ranking optimization for human alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.593541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.593541Z digest=sha256:8a57a424f548efbc29172cde642fb7e92ca1651f8607ab42b0d473b7da5fb66e

Observation 9213a5ed-efd0-4009-b833-fb1631a64bd1 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time KTO: Model Alignment as Prospect Theoretic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.598165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.598165Z digest=sha256:bc79d53b8725aaaa6484041da7cb67e2558c4343e5c434e8e21361ee389aabed

Observation e5a525b3-1930-4ccc-8751-ff5f63ae4758 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.603344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.603344Z digest=sha256:06376a352fe29079886243510059782b1ee03e815684c4425dfb1bb2878ab127

Observation 6d24c066-c825-45cd-ab93-ea24dfd0afef · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.608158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.608158Z digest=sha256:2c3ba7771ccfe8dd9943486ba626de1fbbed8765bef8a2012b627b145ba6ed7c

Observation 96e435dd-c57c-48ea-9d6a-557397562cc5 · outbound

This paper cites Machine Unlearning in Large Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Machine Unlearning in Large Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:42.415020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T16:18:40.612953Z digest=sha256:583a47c0284eceb2073d955f95b26ed9a8bf0aab4194fd2801fe94958922c9d3

Observation d5328590-8792-4f8a-98d4-bf55f78f9471 · outbound

This paper cites Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.617617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.617617Z digest=sha256:49ede791928f261aed92ced78ab2c29701dc9d8a6cb08a55beea30f542fcd84b

Observation 572d086f-55e2-4102-8495-32dee9c94a4e · outbound

This paper cites Model Merging and Safety Alignment: One Bad Model Spoils the Bunch.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Model Merging and Safety Alignment: One Bad Model Spoils the Bunch

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.622233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.622233Z digest=sha256:274f8eaa6156b824dd436cbb8de2e724a2040dc309697984a84f65843b0d8105

Observation b31caba9-2112-4e1f-85d9-8ba70e4cf7e2 · outbound

This paper cites Trustagent: Towards safe and trustworthy llm-based agents through agent constitution.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Trustagent: Towards safe and trustworthy llm-based agents through agent constitution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.626509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.626509Z digest=sha256:0ae3f81080bf5b76f3ed0a3648c746c534f54b87d04fc8e287dbb4da37a30afc

Observation 396404a1-a9e6-456f-957b-347827e1b2d8 · outbound

This paper cites Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.631160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.631160Z digest=sha256:0bca0f19b0635a3041b2c457304909639a96867a56152aa7af69ebbc629aafb5

Observation 4bcea3a2-2472-4028-ba55-23a4c040a8a4 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.636116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.636116Z digest=sha256:7a2cd7be84a6f27604a7e41165b0f67c0b95feb308f685a6d5a2d7c56c88a1af

Observation bc805eee-7db5-4c63-9bc9-dcad2a38642b · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.641145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.641145Z digest=sha256:2fa8211b2547d4c0d09ce3f865f5925bf162962a82fc473c3e9bc6a8e8a8cb4f

Observation 37ee5772-0a77-4982-9ca0-9adb1b4979c5 · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.645718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.645718Z digest=sha256:0ff13bdc748cd3589856d2ea01f0803913f4932d425196b2860b444f515a1ee1

Observation 43d54e6f-706a-475d-833c-5df61b19afa8 · outbound

This paper cites SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.650485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.650485Z digest=sha256:08199dce455fae6e835b2da6aabc0c49d5921c66550d9f125f83e6794db22659

Observation 421ee288-8eca-40b0-bcba-50bee8ed51d5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.655310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.655310Z digest=sha256:fea32899670cee4df17c86c57d7cfb36557d63fb5d67c11169ff510272b918fb

Observation 1da9e745-d102-4f7f-a36f-f3ae579b8f93 · outbound

This paper cites Fast Best-of-N Decoding via Speculative Rejection.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fast Best-of-N Decoding via Speculative Rejection

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.660327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.660327Z digest=sha256:6efac55bb70758407f2545e52f462970053306251d791df3c72fb826e088df64

Observation e5f58d55-9049-4992-b048-707d9eb6478e · outbound

This paper cites Fudge: Controlled text generation with future discriminators.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fudge: Controlled text generation with future discriminators

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.665059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.665059Z digest=sha256:87163d8bfcedb3476dbff2b01f0979185e86d021c252ae8bc3c2e4f455fc57c5

Observation 10373a1d-8d46-4e0d-af28-550656ed2934 · outbound

This paper cites Cold decoding: Energy- based constrained text generation with langevin dynamics.Advances in Neural Information Processing Systems, 35:9538–9551, 2022.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Cold decoding: Energy- based constrained text generation with langevin dynamics.Advances in Neural Information Processing Systems, 35:9538–9551, 2022

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.669670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.669670Z digest=sha256:64635270ebf55664646f17fb68252cfda3786ae4e541c8bba85659db9e477a07

Observation 533b95a6-6aa5-4fff-a450-6c09a6dc7d10 · outbound

This paper cites Aligning Large Language Models with Representation Editing: A Control Perspective.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Aligning Large Language Models with Representation Editing: A Control Perspective

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.674151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.674151Z digest=sha256:553cea0d1abc799404a31a410ca59de6e417d5f7140c411bc55be6be259bbc28

Observation 124f1ebe-35b4-4a99-8e30-73cfad6e3a45 · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ARGS: Alignment as Reward-Guided Search

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.679002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.679002Z digest=sha256:bccbb0bd6ca8fc8420f6864efd0aec639da2bfeb5f4c71b59d6ddacc8e3ac993

Observation 208bf629-eee1-486a-83d3-f438ff773a75 · outbound

This paper cites Decoding-Time Language Model Alignment with Multiple Objectives.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Decoding-Time Language Model Alignment with Multiple Objectives

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.683972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.683972Z digest=sha256:ce3baafd474804c73f327fa18bccad7707f18e0628d3f484a7304450d30230b1

Observation 766529f0-d5ea-44f3-9022-4364b48cfb8e · outbound

This paper cites Deal: Decoding-time alignment for large language models.arXiv preprint arXiv:2402.06147, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Deal: Decoding-time alignment for large language models.arXiv preprint arXiv:2402.06147, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.688825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.688825Z digest=sha256:0c11bae6aa5ad29264bc0086183fbd5fe7a9d836c17a0650c7122368ae17b377

Observation 8714e22e-567f-4363-9f27-0c577114942b · outbound

This paper cites Value Augmented Sampling for Language Model Alignment and Personalization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Value Augmented Sampling for Language Model Alignment and Personalization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.693492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.693492Z digest=sha256:8443df0fc127e6f7a9852c25a8f1d277189987877e7433ada1616deb84a08a18

Observation 123dee81-6968-497b-8966-f3b1c026d33a · outbound

This paper cites ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.699323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.699323Z digest=sha256:ee4df6858c9f5167fbf1181c4fe53af187ec934f4553a07c52958620958445cf

Observation 5830fdfc-6502-443e-a20d-d3c788e9e189 · outbound

This paper cites Adversarial contrastive decoding: Boosting safety alignment of large language models via opposite prompt optimization.arXiv preprint arXiv:2406.16743, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Adversarial contrastive decoding: Boosting safety alignment of large language models via opposite prompt optimization.arXiv preprint arXiv:2406.16743, 2024

Reference 48

Resolution
verified exact
raw_fallback, observed 2026-08-09T16:18:42.059810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T16:18:40.704343Z digest=sha256:a458323dbec1f07b9e062b5f5eb72cd5c166a04fc0afb86907271a5d89ec2776

Observation e17d599a-d879-4a45-9d9f-34a81585f77f · outbound

This paper cites Parameter-Efficient Detoxification with Contrastive Decoding.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Parameter-Efficient Detoxification with Contrastive Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.709137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.709137Z digest=sha256:a92995055afae7b766c2644fb6b459c3288a73187475ad80fa98e521c5ea8a23

Observation 4c174fe1-5db0-4db6-bd00-cded87ecad8d · outbound

This paper cites Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.714261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.714261Z digest=sha256:346d3ee1d1cf276d56afb4a0f4cdc5b152ee3b0b9104198a3edf41b048a56263

Observation 92d87873-ae93-4ffb-9f2c-8da28b49f108 · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.719150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.719150Z digest=sha256:12dc916ced14fede2e550ab1e3576adfd6e729434379b11030b3aaf767aa938b

Observation 9faece8a-cb64-4cee-8fae-f19b7dd27043 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.723901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.723901Z digest=sha256:31fd34a3d3add9a65d8856393d87d8344550d7571643eb99e1a65061ccf9c25b

Observation f656bf9e-92b7-42a5-bae9-d40de150813e · outbound

This paper cites Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:41.914435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T16:18:40.728371Z digest=sha256:b52e39cfc9d80d723753a812b307dfc6d3493f42a62a532622a73686e9d7754e

Observation 1f8db701-1072-47cc-8803-4c7c0730eb18 · outbound

This paper cites Mixture of attentions for speculative decoding, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Mixture of attentions for speculative decoding, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.732968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.732968Z digest=sha256:894822fc251fea5ac3ec71c539eea0db1591244c6dc788c7e356d3f43da4e290

Observation 651c2a9e-9828-4a4a-bdb5-73a55ba1f21c · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.737239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.737239Z digest=sha256:979b3b8a13523394a2a9e6f4d655b9c06fc12ee06804980fac11d4ebc3f3c5f9

Observation ed3494c6-c7ef-4308-aeb9-2c7b04a6a608 · outbound

This paper cites Large language models for robotics: A survey.arXiv preprint arXiv:2311.07226, 2023.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Large language models for robotics: A survey.arXiv preprint arXiv:2311.07226, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.741644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.741644Z digest=sha256:7c52493191ed155eeec638f79c1d5ac36fcc7cdbc9132bbe1a5297a547355734

Observation 7145db43-1cca-4f4d-ab05-92a39914a809 · outbound

This paper cites Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.746544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.746544Z digest=sha256:1a8ddda5b94c66c78ba768cd682f6ca460d28b2cfcac1125a941c30299a47281

Observation b1b9a395-8e6f-4b0b-a60b-482ac08faa59 · outbound

This paper cites Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.751431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.751431Z digest=sha256:e0b9acd92d620415dd68dcf9c989b9219e745ab8288518e9ac681badd61c500e

Observation da1743d3-e76f-4297-bd1c-4beeb9758733 · outbound

This paper cites Learn from Failure: Fine-Tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic Proving.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learn from Failure: Fine-Tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic Proving

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.756147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.756147Z digest=sha256:d350dd45cea3fa75c5fc32316f9010c64d205a6535cc7e12118a1112b2916f0b

Observation 9e5ba516-0a06-4f77-9154-6f82a0b8c9c8 · outbound

This paper cites Learning to Learn Faster from Human Feedback with Language Model Predictive Control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning to Learn Faster from Human Feedback with Language Model Predictive Control

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.761602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.761602Z digest=sha256:023b9d286c5d55df7901612c884fc806fdf90e61f530cef12a0307476e3ae3e4

Observation 5705fb0c-2d13-4b52-80af-c767cd16424e · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.766715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.766715Z digest=sha256:343b40695b9339241a923218f44462479453d6f8bfeb86fa8b6ffd07c66b6f8c

Observation 42aaf597-5211-412e-b112-53e20c5aebc7 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.771963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.771963Z digest=sha256:e43534426d41aaf2469cc55a0133a13dc0f04a8c2fffb2501e54fd03e76284ca

Observation ac704221-d03e-4dad-b832-cde745c3475e · outbound

This paper cites Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.777676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.777676Z digest=sha256:8a1063e32e0146565bac25c8cbace1cad13ef3570787973be5b504fd4a411544

Observation bef3432f-ac89-4a0e-9fd2-fcfa2702a5d4 · outbound

This paper cites Simulation-guided beam search for neural combinatorial optimization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Simulation-guided beam search for neural combinatorial optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.782889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.782889Z digest=sha256:a0a2ca7f5bdd14f1cade525ddc4ad82f0c01f77bc57f75dfd3fa28d37cd57d0f

Observation e3d415a0-4764-41cb-812e-69dcc65e9d12 · outbound

This paper cites Constrained policy optimization.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Constrained policy optimization

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.788403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.788403Z digest=sha256:eae455a3563497a5d5eafc90aa33232564ab95e861733c159902c10a8b2f8042

Observation de7ee1c1-9ed3-4adb-86e6-a016d3d0146e · outbound

This paper cites A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.793196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.793196Z digest=sha256:e8e8bf6e7e0bb9104477fcf1f1d3384715b823d9fc523346a55a9d938f873129

Observation e5619f61-9495-4266-a530-23b98cdca7b0 · outbound

This paper cites Alpaca: A strong, replicable instruction- following model.Stanford Center for Research on Foundation Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Alpaca: A strong, replicable instruction- following model.Stanford Center for Research on Foundation Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.798466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.798466Z digest=sha256:26a6466a2af8799e851311983a5403852b7f9e5e6aefd7d4dbc24589a544cf0c

Observation b7de5c36-f360-4a9f-b87f-ef62f7b04c7d · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Gonzalez, Ion Stoica, and Eric P

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.803395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.803395Z digest=sha256:885801fb3713a9bfbdccefc45270463ec59adf2866b0ca1b23e8ad214ed53351

Observation 7d07096d-3017-4fed-b163-1c123b3a5bbc · outbound

This paper cites The Llama 3 Herd of Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time The Llama 3 Herd of Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.808370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.808370Z digest=sha256:0bb424e2770a52e9fdafd03dc6c7bf69d1380d11ec3d4b5b5c406e2d37db1a84

Observation e12b652f-af52-4be7-b387-036e5cfe24d3 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.813673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.813673Z digest=sha256:c32a6797c6959b9a1f433c30f95d8330e3f8672ca999cfdb5922ddeaad652d7a

Observation 5d4c7c56-409a-4a5e-9726-ca7b20976493 · outbound

This paper cites Quantile Regression for Distributional Reward Models in RLHF.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Quantile Regression for Distributional Reward Models in RLHF

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.819124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.819124Z digest=sha256:c6a5c998b47c2ae1e9b3c024e40ba6da5f9adaa5b22f36b61a65652bd0518083

Observation 0869e124-9f78-413a-82c8-6378656a10ec · outbound

This paper cites CRC Press, 1999.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time CRC Press, 1999

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.824171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.824171Z digest=sha256:a399363edc918175515887b852707ab58854c6a4fb1c83884451d9529fdf2d7a

Observation 5b182e98-1916-42fa-bce7-c4ad533480d6 · outbound

This paper cites Safe exploration in finite markov decision processes with gaussian processes.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration in finite markov decision processes with gaussian processes

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.829059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.829059Z digest=sha256:63e63f6da212a70e203fb1cf4a01acd0b9e4b6183561c28dccdda037a4b1da05

Observation 1bf01d93-0d87-42ce-a967-2632b8e3c341 · outbound

This paper cites Learning-based model predictive control for safe exploration.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning-based model predictive control for safe exploration

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.834159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.834159Z digest=sha256:7ad15d5114dd6b3779ba3bed6a1e51465ac7895b7cb1aa108b3536e8abe78555

Observation 6c6c3dfd-6153-4bda-ab02-ec7d776a2e50 · outbound

This paper cites Safe exploration in continuous action spaces.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration in continuous action spaces

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.839073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.839073Z digest=sha256:83ddcaf684a137c44d7a434b1d4f54c36acc3c40299757bd8ba44118836bdd73

Observation 5634e76d-a329-41bf-bd30-0d3f27d24f45 · outbound

This paper cites Safe exploration and optimization of constrained mdps using gaussian processes.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration and optimization of constrained mdps using gaussian processes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.844031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.844031Z digest=sha256:d5b61bc7078bf8056f3cc2a2dc523ce847c9eb58d4be3983bab7b65b0dbda715

Observation c1d9e1c2-9de4-4dca-b490-7b4ede335b4e · outbound

This paper cites Conservative safety critics for exploration.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Conservative safety critics for exploration

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.849005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.849005Z digest=sha256:78f133b00397814f5de83688f92230d18cde12687e95ee933a9704cc3d3241c2

Observation 0d2d19ca-97bd-4922-9920-8a2271f945fb · outbound

This paper cites Lyapunov-based safe policy optimization for continuous control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Lyapunov-based safe policy optimization for continuous control

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.853910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.853910Z digest=sha256:a4fe6e4b7f77a3151ee970242a60a37e49b018a4011d45df3bcb8b81c2d2b3dc

Observation 64b53dcf-dbfb-4ebb-85cc-673393c5aa59 · outbound

This paper cites Lyapunov-based Safe Policy Optimization for Continuous Control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Lyapunov-based Safe Policy Optimization for Continuous Control

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.859342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.859342Z digest=sha256:4925770caf6a65dd6ab525fd5203d6ba21ed90aab81d777b71b1dbc31b2dad35

Observation 95d7b4cf-1ca3-4cdc-ab9a-8240afb01643 · outbound

This paper cites Safe model- based reinforcement learning with stability guarantees.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe model- based reinforcement learning with stability guarantees

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.864852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.864852Z digest=sha256:b2f4ac5d8d16aba50bb06125811120dde43fae328058a7e4417d2cf5dee5ea0a

Observation ac53c036-895a-4b39-9feb-c7b3807c3c44 · outbound

This paper cites Barrier-certified adaptive reinforcement learning with applications to brushbot navigation.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Barrier-certified adaptive reinforcement learning with applications to brushbot navigation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.869206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.869206Z digest=sha256:f206a1af290a86b2acee09d191480b832f541152a57c9c810e01c594aa095b28

Observation 7d4be5b3-756c-43d3-8f69-6bf505e44aab · outbound

This paper cites End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.873885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.873885Z digest=sha256:0e8c1c7715f4f910d7f295310624dfd576d28c146efea5874db5b0b89d21decd

Observation 4ae5cb98-1c8d-4ab4-b423-a7668f7417e0 · outbound

This paper cites Reachability-based safe learning with gaussian processes.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Reachability-based safe learning with gaussian processes

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.878073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.878073Z digest=sha256:bf19506146818c101144009155e15535477b509ad0aed5481d154e9e09114514

Observation ef3f6ca5-4277-4b1d-8799-272cf280e262 · outbound

This paper cites Safeguarding resource- constrained cyber-physical systems with adaptive control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safeguarding resource- constrained cyber-physical systems with adaptive control

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.882265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.882265Z digest=sha256:027196386b8dc4f11a8129bb7aa1bc4f46b32db5a96f513822b0188e3541e546

Observation 6696e5ab-2224-4843-b2df-141817be418d · outbound

This paper cites Bridging model-based safety and model-free reinforcement learning through system identification and safety-critical control.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Bridging model-based safety and model-free reinforcement learning through system identification and safety-critical control

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.886651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.886651Z digest=sha256:d1599996764a86923fd31e4cd1eeef1ce20d1cad2e1104a08fb6707af26d6826

Observation 8ab2b39d-5b6a-4d8f-9c07-73f538ae03a1 · outbound

This paper cites Benchmarking safe exploration in deep reinforcement learning.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Benchmarking safe exploration in deep reinforcement learning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.890979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.890979Z digest=sha256:760af1c9bc1e8109fd96e916ebdb180b8d516373314d4f295ed63982fc1c5785

Observation 8ef6e941-b931-4824-9cb7-770e74557421 · outbound

This paper cites Responsive safety in reinforcement learning by monitoring risk and adapting policies.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Responsive safety in reinforcement learning by monitoring risk and adapting policies

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.895451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.895451Z digest=sha256:bc055fe372254637a06d61f5ed800aadd4e716d10441c19ee458f4b27a034bf3

Observation 50a8bd5b-4ba4-4050-bcad-7497e9c4a49a · outbound

This paper cites Relative value learning for constrained reinforce- ment learning.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Relative value learning for constrained reinforce- ment learning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.899713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.899713Z digest=sha256:dd257c11fa76a539c06da14000cfd527dec64570aef09d456ccb5fcfa1b67ab4

Observation 1aaaa581-e9b4-4bb2-a00b-f002d65c6940 · outbound

This paper cites Natural policy gradient for safe reinforcement learning with c-mdps.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Natural policy gradient for safe reinforcement learning with c-mdps

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.904474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.904474Z digest=sha256:67003bd045f6f6fe47a3fcc62f26a1aff96262430a779b3445d5916e5d82b30c

Observation 5f4be3f3-1c0b-4456-b84c-0429c3c7a294 · outbound

This paper cites Group Robust Preference Optimization in Reward-free RLHF.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Group Robust Preference Optimization in Reward-free RLHF

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.909170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.909170Z digest=sha256:c6d8fce18c3b0e2eb8f30d017f56a884582d6a81b981f98a35435cae65b5beb1

Observation ced521c4-6a4b-4a5f-a0ec-90a99883d58a · outbound

This paper cites Mission Impossible: A Statistical Perspective on Jailbreaking LLMs.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Mission Impossible: A Statistical Perspective on Jailbreaking LLMs

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.914123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.914123Z digest=sha256:fd23d971d8642a6e311be06f202f852931ea9b21f29fa7ba78d5201d2e929073

Observation f4af410e-f2c1-4553-8007-fda5028b0cf7 · outbound

This paper cites Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.919238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.919238Z digest=sha256:cff4b27e407bd3e502ceedee6c44000314913d59ea5614ade2925f6c2f8409d8

Observation cba5d427-4d84-4a84-a1c6-7d81a20c60d7 · outbound

This paper cites Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.934070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.934070Z digest=sha256:5afcce3073fb224715cbb6c9ba3e88025b613338841ada7d316b5d280e150f40

Observation ca5bb15d-04cd-456c-8255-defec97a947e · outbound

This paper cites On prompt-driven safeguarding for large language models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time On prompt-driven safeguarding for large language models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.939000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.939000Z digest=sha256:1646896f84e68132cf8717512716758ac62ed2c00f18ee8ecc6bd220bbf1085e

Observation 834e3454-ed9c-4937-b478-b541885fde2f · outbound

This paper cites Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:41.475394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T16:18:40.944028Z digest=sha256:cc74e4d76747cf8708eb6557dc4c1d9ca1450f09e3e4810203be944bed4b70c3

Observation 9bf19622-0f6e-4d67-9f43-8ab2e04fed6c · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.949223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.949223Z digest=sha256:572f5c1671a783b246b1b376845d7bc0fb884c5635ff2d1b2d47004f7bf01e9b

Observation a7ec959a-c7c2-4bb0-bdc7-2514815d82f6 · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.954512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.954512Z digest=sha256:47cd41959759c38b4c86d5b48606bef23007866681b51ead84b00095918bc77b

Observation bdbb141f-9b63-42cf-95ec-a002ad5818ce · outbound

This paper cites FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.959416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.959416Z digest=sha256:881e173b26dc6c33ee41367309efc43139dbb3ca60524e4aa3705f12f7b6dca2

Observation e6d9fe43-9f48-4960-bd25-9b2c985a3985 · outbound

This paper cites Chain-of-detection enables robust and efficient jailbreak defense.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Chain-of-detection enables robust and efficient jailbreak defense

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.964386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.964386Z digest=sha256:19e2882874ddd7194a747f64ba317c2ee7f0dafc5166011303997b278c934548

Observation 7f5930cc-e48e-4a7f-99cc-629663c6796f · outbound

This paper cites Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-08-09T16:18:41.400163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T16:18:40.969178Z digest=sha256:9cb85b725c840b8888e8aacb6dedcc489cee25387be3116d716245d028be7233

Pith citing papers

Observation f47fd653-d640-4acc-b755-e93a1f270232 · inbound

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values cites this paper.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Almost Surely Safe Alignment of Large Language Models at Inference-Time

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:43:07.630186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:43:07.249475Z digest=sha256:2211a02f02dec402dc34096a706d3275bb1f649dda53ab6fc312eb075200078a