Pith. sign in

Paper Citation Record · LEDGER

Learning Safety Constraints for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2505.24445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24445 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:16.628557Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:57:51.928730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved20
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 211c39db-1a9e-44e8-952e-e3a33be7cf9e · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Learning Safety Constraints for Large Language Models MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:14.985484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:14.985484Z digest=sha256:70d5517de5bb4728c4b2f809f150803aacdd41eaa7bfec1fc3a9c4100d59643a

Observation 70a87e9a-0d49-4527-b423-512a812a02e6 · outbound

This paper cites SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking.

Learning Safety Constraints for Large Language Models SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.354687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.354687Z digest=sha256:cdba18c36e6bea80f18987ef0a5e0bb978fc132d81dee349857a0fc299e9a101

Observation 7363c4c8-40c6-482c-b8e0-7d7bc2524288 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Learning Safety Constraints for Large Language Models Safe Exploration in Continuous Action Spaces

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.456333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.456333Z digest=sha256:bbf4b01d57f7cd825700cbf9ea194edbc67235c16dd4b79cce12156d95dd87a0

Observation 08849323-40da-4714-a7fe-c4502c5a2c33 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Learning Safety Constraints for Large Language Models RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.620283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.620283Z digest=sha256:8b870e516c14d7a1d751f24e84f2b7f7a28a68961028edaec6a35a25ad3ea992

Observation 519a09b1-99c1-4f8d-bf1e-1b428aa01b35 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Learning Safety Constraints for Large Language Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.758569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.758569Z digest=sha256:060414dde616f39e107c1ace259af4bc33a08d05adca291d2d058b9dec34d8f7

Observation d78c2c74-df8f-49c9-87d0-4dd07e0a2639 · outbound

This paper cites Backdoor Attacks for In-Context Learning with Language Models.

Learning Safety Constraints for Large Language Models Backdoor Attacks for In-Context Learning with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.821630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.821630Z digest=sha256:790b04ea5c8fde59348a1490e71836251fe329e0358f3489aaadffdfbae84060

Observation 56cec1a0-9d3c-46d4-b90a-bb3b4cf916c4 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Learning Safety Constraints for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.939144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.939144Z digest=sha256:cbde7acb441897973d0c9d6034c684cef40a0bdf1b5d6946c4102e7a204ce3de

Observation 7f1d5073-9f8a-45ce-be38-8a0bd2e8880e · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

Learning Safety Constraints for Large Language Models Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.012101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.012101Z digest=sha256:35951f65e2aa0806bd600fd50aa58621feb41fccb5603b182ce42257e74fea67

Observation 41797f56-e347-427e-879c-d3aa95e037af · outbound

This paper cites Mpax: Mathematical pro- gramming in jax.

Learning Safety Constraints for Large Language Models Mpax: Mathematical pro- gramming in jax

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.069748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.069748Z digest=sha256:ebc2c5bce4fb8d71280039d4cfe44cc4af0e6b77f5bb779fdf7fd70d4063f7dc

Observation 5df8c693-4082-40b4-82ae-d1d98aff73cb · outbound

This paper cites Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization.

Learning Safety Constraints for Large Language Models Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.143593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.143593Z digest=sha256:258c5776080cf45e2c3089ca5f6d7b5b71c9ee64f496cdb755a5e1cb7825d3ae

Observation 2481352e-529d-4429-8213-38d98983fd1c · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

Learning Safety Constraints for Large Language Models Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.217942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.217942Z digest=sha256:01fcd3774893addc2eaa3a7430b212e8e5165af586e85a627abfa0a250792a28

Observation b459100a-c042-471c-be37-2187899fe70f · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Learning Safety Constraints for Large Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.287165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.287165Z digest=sha256:ae055c0ac808e577855ee2f1c404a0c79782d8e70696de82fa6a05919e5c028b

Observation 74b81ec1-c15b-4fb3-a0bb-fb078d10d9bd · outbound

This paper cites Token-level Direct Preference Optimization.

Learning Safety Constraints for Large Language Models Token-level Direct Preference Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.341645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.341645Z digest=sha256:d29215e07d4916bd2479cf1934787c65ba1d17dd1b3ed59a4c84bbfdc8a54ea2

Observation d0b25abb-0b16-4501-8622-7fd70047f051 · outbound

This paper cites Panacea: Pareto Alignment via Preference Adaptation for LLMs.

Learning Safety Constraints for Large Language Models Panacea: Pareto Alignment via Preference Adaptation for LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.406578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.406578Z digest=sha256:0600b8dc9333bdf00e36c866581aa8cb8d9ecc7c9c54fce8172614c5bd8323fa

Observation cc3e2173-fffc-4d37-aa3e-989ffb38cc47 · outbound

This paper cites Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization.

Learning Safety Constraints for Large Language Models Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:17.936203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:16.462865Z digest=sha256:4a55ffc05c968adbf39093487cb2de811b6c308e1e2f30c89d27d1ee3d66daee

Observation c06c74f7-cc17-4d49-8621-04f06d5d164c · outbound

This paper cites an unresolved cited work.

Learning Safety Constraints for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:30:17.725886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:16.506016Z digest=sha256:67d899d4bcda15bcbac00872d6ee9fa779450adfcbd21f383752c9873031f39c

Observation 947a66c1-85a2-4efe-b390-8d096005b368 · outbound

This paper cites {human question}\n{model answer}.

Learning Safety Constraints for Large Language Models {human question}\n{model answer}

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:30:17.579256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:16.558669Z digest=sha256:f4596f0d182b98277d1ec85391a932b389f000c6119c2b4e6f2b2762b3f93527

Observation e2f0e159-4e83-4861-9165-53506c0bd2b9 · outbound

This paper cites Results show mean ± standard deviation.

Learning Safety Constraints for Large Language Models Results show mean ± standard deviation

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:30:17.301765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:16.628557Z digest=sha256:6393b488eaf0d1c425b6883b51ccc3701320ed00fb2910f9182334ca6312fdca

Observation 12ed9f93-7ac7-444e-845a-2394f397f838 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Learning Safety Constraints for Large Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:14.786704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:14.786704Z digest=sha256:9c281473fe31016e5d43459abed8730d279c4af546c35c84794ff59402e21c7a

Observation 3fc8dcb9-694d-40f7-bfa6-435116fa0bca · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Learning Safety Constraints for Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.888021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.888021Z digest=sha256:299327cea3099bada3305dd1f3753b454fdf2a3ab747303970ca9fa68cb63f99

Observation 5691e74e-0f86-4b94-8d91-7297b889fcc0 · outbound

This paper cites Interpreting Neural Networks through the Polytope Lens.

Learning Safety Constraints for Large Language Models Interpreting Neural Networks through the Polytope Lens

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.138252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.138252Z digest=sha256:e63bbda13a95bf2fb527eb7a91280dd8322b9c99cfaac82534f20e6b7145c6f8

Observation 62406f3f-8d92-4d13-9b40-ac547bef0427 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

Learning Safety Constraints for Large Language Models AI Control: Improving Safety Despite Intentional Subversion

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.698528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.698528Z digest=sha256:646ffd16679b09dbddade8524d13bb54e12ba0e7112114d2573e49207ebd8068

Observation 4c4f088e-3a56-4d72-9850-0fc309a42dff · outbound

This paper cites Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts.

Learning Safety Constraints for Large Language Models Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:17.032993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:15.545665Z digest=sha256:72c8aa63e9ed31a3d3adf600cbce37f162b055808411619b505059774551ffbf

Observation 28a20280-8c32-43e8-b0aa-224937fb62cc · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Learning Safety Constraints for Large Language Models Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.271464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.271464Z digest=sha256:15e434bdef788eb1aa466812761163e16c3c81ef833fb763f96e7ae76b4cf9a2

Observation 05e6391c-18c6-4105-a685-75cfa16f0747 · outbound

This paper cites and Bartlett, P.

Learning Safety Constraints for Large Language Models and Bartlett, P

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:18.150799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:30:14.835783Z digest=sha256:5e60a641f9acb3f2dcb04b95b390a9d1b6a9c9269efb715d4f23731d6dd1a09a

Pith citing papers

Observation e2bd960a-5dbb-4c95-9548-c457ff5dc172 · inbound

When control meets large language models: From words to dynamics cites this paper.

When control meets large language models: From words to dynamics Learning Safety Constraints for Large Language Models

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:54:13.126908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:52:44.632671Z digest=sha256:7a23d3f0939b0a16845a4e859776a30fd90fb4c36a4e47ba127626d5d3192d6e

Observation 9036905c-b097-4870-b700-ba09c4020dc2 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Learning Safety Constraints for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.216730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:96dde680f0920e7a3a74bd9c4fc987eb02353b5ab8ce2752e8fe7039bb8b8ea0

Observation c9ba4192-0e7a-4ee5-ab2c-d234dde2bced · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Learning Safety Constraints for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:01:20.847532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:26fde45a9bdaabeeea4463222ce6765434877202da1edaf341f5adcdf595bf98

Observation 511e8842-49c9-4e3a-a11b-475e12eb343f · inbound

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance cites this paper.

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Learning Safety Constraints for Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:43.438995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:35:35.594085Z digest=sha256:6ade963b8caa85dcca6baf38018700ba3c613a65e16fc8771e8d098b396de38f

Observation 9d400a3f-0fe4-42db-8206-794618102f72 · inbound

Geometry-Guided Constraint Learning for LLM Safety Classification cites this paper.

Geometry-Guided Constraint Learning for LLM Safety Classification Learning Safety Constraints for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.928730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.928730Z digest=sha256:ff040ee0bd6d39bacaaa594d6de2b843f34fb09c00fcf48521525a0645710b9d