Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T16:18:40.969178Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 1 inbound Pith citation observation for arXiv:2502.01208.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T16:18:40.969178Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:07.249475Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T05:43:07.622958Z
100 of 125 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8ac5fa90-7d16-4348-a5ad-6f20f83efd4c · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time An empirical survey on long document summarization: Datasets, models, and metrics.ACM computing surveys, 55(8):1–35, 2022
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1181761a-d634-4615-af8e-de209e60ba4f · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning to summarize with human feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e88c8a-2a73-44df-9c95-7a9713802ee6 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Pal: Program-aided language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6d39b8d-bfc0-4123-823f-368013e915a5 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e67b72c-ab4c-4c90-9378-ac0cbdc4abfc · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time ROS-LLM: A ROS framework for embodied AI with task feedback and structured reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3910d5-6cb5-4ca6-8fb1-c8032dc86753 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time A sur- vey on integration of large language models with intelligent robots.Intelligent Service Robotics, 17(5):1091–1107, August 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2025b726-2696-4717-be04-c3c9820a338c · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef519bac-1552-4b43-ac33-199be15ae185 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64a6923b-99a6-47c1-ab70-77ddc7620e35 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Ethical and social risks of harm from Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4b25dc-92a5-41ff-b2e2-dd61d1ec7116 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f199c6c5-b8ec-4f5b-b0e0-a29eb46ccddb · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648cc058-61a0-4053-8b17-978cb4b54180 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dba0b472-7b72-4142-8613-083d21d44aca · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time WebGPT: Browser-assisted question-answering with human feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 227a9f82-3d23-44eb-8d64-70512565a5d7 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Controlled Decoding from Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294b9ea5-5545-462b-97b3-54e10ec76f02 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Discrete-time markov control processes with discounted unbounded costs: optimality criteria.Kybernetika, 28(3):191–212, 1992
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee21386e-f652-4cd8-9501-9486c02f365e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Sauté rl: Almost surely safe reinforcement learning using state augmentation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7c85b2a-13af-4795-abf5-a1817f9f6c71 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fine-Tuning Language Models from Human Preferences
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd0f1f1-d309-4023-bd13-2b80c4327f83 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Proximal Policy Optimization Algorithms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf5f19f-ae5a-405f-8053-bb6ad02842ff · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 649dace3-01ac-4ac4-8cca-b3a0e6853cb3 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90d13abe-0606-46a8-be83-baf1426b7be6 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7699aa7b-11f8-4f6e-8770-38fe3745ab72 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6eccc06-70ba-46d1-a3e0-5158e1fbea3c · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cd7047a-8241-4f99-aed5-d87a2d392949 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6976cd11-331a-4878-8740-128ed3baadfe · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Preference ranking optimization for human alignment
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9213a5ed-efd0-4009-b833-fb1631a64bd1 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time KTO: Model Alignment as Prospect Theoretic Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a525b3-1930-4ccc-8751-ff5f63ae4758 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d24c066-c825-45cd-ab93-ea24dfd0afef · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e435dd-c57c-48ea-9d6a-557397562cc5 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Machine Unlearning in Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5328590-8792-4f8a-98d4-bf55f78f9471 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572d086f-55e2-4102-8495-32dee9c94a4e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b31caba9-2112-4e1f-85d9-8ba70e4cf7e2 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Trustagent: Towards safe and trustworthy llm-based agents through agent constitution
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 396404a1-a9e6-456f-957b-347827e1b2d8 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bcea3a2-2472-4028-ba55-23a4c040a8a4 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc805eee-7db5-4c63-9bc9-dcad2a38642b · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ee5772-0a77-4982-9ca0-9adb1b4979c5 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d54e6f-706a-475d-833c-5df61b19afa8 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 421ee288-8eca-40b0-bcba-50bee8ed51d5 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1da9e745-d102-4f7f-a36f-f3ae579b8f93 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fast Best-of-N Decoding via Speculative Rejection
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5f58d55-9049-4992-b048-707d9eb6478e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Fudge: Controlled text generation with future discriminators
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10373a1d-8d46-4e0d-af28-550656ed2934 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Cold decoding: Energy- based constrained text generation with langevin dynamics.Advances in Neural Information Processing Systems, 35:9538–9551, 2022
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 533b95a6-6aa5-4fff-a450-6c09a6dc7d10 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Aligning Large Language Models with Representation Editing: A Control Perspective
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 124f1ebe-35b4-4a99-8e30-73cfad6e3a45 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time ARGS: Alignment as Reward-Guided Search
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208bf629-eee1-486a-83d3-f438ff773a75 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Decoding-Time Language Model Alignment with Multiple Objectives
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 766529f0-d5ea-44f3-9022-4364b48cfb8e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Deal: Decoding-time alignment for large language models.arXiv preprint arXiv:2402.06147, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8714e22e-567f-4363-9f27-0c577114942b · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Value Augmented Sampling for Language Model Alignment and Personalization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 123dee81-6968-497b-8966-f3b1c026d33a · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5830fdfc-6502-443e-a20d-d3c788e9e189 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Adversarial contrastive decoding: Boosting safety alignment of large language models via opposite prompt optimization.arXiv preprint arXiv:2406.16743, 2024
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e17d599a-d879-4a45-9d9f-34a81585f77f · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Parameter-Efficient Detoxification with Contrastive Decoding
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c174fe1-5db0-4db6-bd00-cded87ecad8d · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d87873-ae93-4ffb-9f2c-8da28b49f108 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9faece8a-cb64-4cee-8fae-f19b7dd27043 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f656bf9e-92b7-42a5-bae9-d40de150813e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Probing the Safety Response Boundary of Large Language Models via Unsafe Decoding Path Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1f8db701-1072-47cc-8803-4c7c0730eb18 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Mixture of attentions for speculative decoding, 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651c2a9e-9828-4a4a-bdb5-73a55ba1f21c · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed3494c6-c7ef-4308-aeb9-2c7b04a6a608 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Large language models for robotics: A survey.arXiv preprint arXiv:2311.07226, 2023
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7145db43-1cca-4f4d-ab05-92a39914a809 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Leandojo: Theorem proving with retrieval-augmented language models.Advances in Neural Information Processing Systems, 36, 2024
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b9a395-8e6f-4b0b-a60b-482ac08faa59 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Solving olympiad geometry without human demonstrations.Nature, 625(7995):476–482, 2024
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1743d3-e76f-4297-bd1c-4beeb9758733 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learn from Failure: Fine-Tuning LLMs with Trial-and-Error Data for Intuitionistic Propositional Logic Proving
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e5ba516-0a06-4f77-9154-6f82a0b8c9c8 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning to Learn Faster from Human Feedback with Language Model Predictive Control
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5705fb0c-2d13-4b52-80af-c767cd16424e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42aaf597-5211-412e-b112-53e20c5aebc7 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac704221-d03e-4dad-b832-cde745c3475e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bef3432f-ac89-4a0e-9fd2-fcfa2702a5d4 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Simulation-guided beam search for neural combinatorial optimization
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d415a0-4764-41cb-812e-69dcc65e9d12 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Constrained policy optimization
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7ee1c1-9ed3-4adb-86e6-a016d3d0146e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time A Survey on LLM Test-Time Compute via Search: Tasks, LLM Profiling, Search Algorithms, and Relevant Frameworks
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5619f61-9495-4266-a530-23b98cdca7b0 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Alpaca: A strong, replicable instruction- following model.Stanford Center for Research on Foundation Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7de5c36-f360-4a9f-b87f-ef62f7b04c7d · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Gonzalez, Ion Stoica, and Eric P
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d07096d-3017-4fed-b163-1c123b3a5bbc · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time The Llama 3 Herd of Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e12b652f-af52-4be7-b387-036e5cfe24d3 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d4c7c56-409a-4a5e-9726-ca7b20976493 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Quantile Regression for Distributional Reward Models in RLHF
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0869e124-9f78-413a-82c8-6378656a10ec · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time CRC Press, 1999
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b182e98-1916-42fa-bce7-c4ad533480d6 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration in finite markov decision processes with gaussian processes
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf01d93-0d87-42ce-a967-2632b8e3c341 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Learning-based model predictive control for safe exploration
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6c3dfd-6153-4bda-ab02-ec7d776a2e50 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration in continuous action spaces
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5634e76d-a329-41bf-bd30-0d3f27d24f45 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe exploration and optimization of constrained mdps using gaussian processes
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d9e1c2-9de4-4dca-b490-7b4ede335b4e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Conservative safety critics for exploration
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2d19ca-97bd-4922-9920-8a2271f945fb · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Lyapunov-based safe policy optimization for continuous control
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b53dcf-dbfb-4ebb-85cc-673393c5aa59 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Lyapunov-based Safe Policy Optimization for Continuous Control
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95d7b4cf-1ca3-4cdc-ab9a-8240afb01643 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safe model- based reinforcement learning with stability guarantees
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac53c036-895a-4b39-9feb-c7b3807c3c44 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Barrier-certified adaptive reinforcement learning with applications to brushbot navigation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4be5b3-756c-43d3-8f69-6bf505e44aab · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae5cb98-1c8d-4ab4-b423-a7668f7417e0 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Reachability-based safe learning with gaussian processes
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef3f6ca5-4277-4b1d-8799-272cf280e262 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safeguarding resource- constrained cyber-physical systems with adaptive control
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6696e5ab-2224-4843-b2df-141817be418d · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Bridging model-based safety and model-free reinforcement learning through system identification and safety-critical control
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab2b39d-5b6a-4d8f-9c07-73f538ae03a1 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Benchmarking safe exploration in deep reinforcement learning
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef6e941-b931-4824-9cb7-770e74557421 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Responsive safety in reinforcement learning by monitoring risk and adapting policies
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50a8bd5b-4ba4-4050-bcad-7497e9c4a49a · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Relative value learning for constrained reinforce- ment learning
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1aaaa581-e9b4-4bb2-a00b-f002d65c6940 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Natural policy gradient for safe reinforcement learning with c-mdps
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f4be3f3-1c0b-4456-b84c-0429c3c7a294 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Group Robust Preference Optimization in Reward-free RLHF
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced521c4-6a4b-4a5f-a0ec-90a99883d58a · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4af410e-f2c1-4553-8007-fda5028b0cf7 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Improving llm safety alignment with dual-objective optimization.arXiv preprint arXiv:2503.03710, 2025
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba5d427-4d84-4a84-a1c6-7d81a20c60d7 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca5bb15d-04cd-456c-8255-defec97a947e · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time On prompt-driven safeguarding for large language models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834e3454-ed9c-4937-b478-b541885fde2f · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Dynamic Guided and Domain Applicable Safeguards for Enhanced Security in Large Language Models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9bf19622-0f6e-4d67-9f43-8ab2e04fed6c · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ec959a-c7c2-4bb0-bdc7-2514815d82f6 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdbb141f-9b63-42cf-95ec-a002ad5818ce · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d9fe43-9f48-4960-bd25-9b2c985a3985 · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Chain-of-detection enables robust and efficient jailbreak defense
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f5930cc-e48e-4a7f-99cc-629663c6796f · outbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f47fd653-d640-4acc-b755-e93a1f270232 · inbound
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values On Almost Surely Safe Alignment of Large Language Models at Inference-Time
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.