Pith. sign in

Paper Citation Record · LEDGER

Learning Safety Constraints for Large Language Models

As of 11 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 5 inbound Pith citation observations for arXiv:2505.24445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24445 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:30:16.628557Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:57:51.928730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved20
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 211c39db-1a9e-44e8-952e-e3a33be7cf9e · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Learning Safety Constraints for Large Language Models MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:14.985484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:14.985484Z digest=sha256:974088edbe98fd32b551bd584edd050e0333e606c85361fb9852b5a2f7b241f1

Observation 70a87e9a-0d49-4527-b423-512a812a02e6 · outbound

This paper cites SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking.

Learning Safety Constraints for Large Language Models SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.354687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.354687Z digest=sha256:5d840b03923691aecfb5772d70ee7cb0768bb7114b486070ed1e4470591a1261

Observation 7363c4c8-40c6-482c-b8e0-7d7bc2524288 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Learning Safety Constraints for Large Language Models Safe Exploration in Continuous Action Spaces

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.456333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.456333Z digest=sha256:06004a8f27530f81d9734d38b0655238e5e4d56bcd30e749f8b17a0917a8ffab

Observation 08849323-40da-4714-a7fe-c4502c5a2c33 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Learning Safety Constraints for Large Language Models RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.620283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.620283Z digest=sha256:1490abddf03d46439acbb75876294d4cedc515225ab64f1fbb4661439f1d7d07

Observation 519a09b1-99c1-4f8d-bf1e-1b428aa01b35 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Learning Safety Constraints for Large Language Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.758569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.758569Z digest=sha256:742f733a360f80dee5f5dcf7f0567f5936f98720dac882315b9364872948cd54

Observation d78c2c74-df8f-49c9-87d0-4dd07e0a2639 · outbound

This paper cites Backdoor Attacks for In-Context Learning with Language Models.

Learning Safety Constraints for Large Language Models Backdoor Attacks for In-Context Learning with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.821630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.821630Z digest=sha256:765b8694feb7f3fbc8ab69e817e46ea74897347abd7a79dfd56cfde84c98b55a

Observation 56cec1a0-9d3c-46d4-b90a-bb3b4cf916c4 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Learning Safety Constraints for Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.939144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.939144Z digest=sha256:249e0fd345b8f3382f5f0ecb4cc0614f9e9d2e63693b1f1c9c3dfac54ae721a7

Observation 7f1d5073-9f8a-45ce-be38-8a0bd2e8880e · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

Learning Safety Constraints for Large Language Models Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.012101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.012101Z digest=sha256:529a8f4711a6dfe13acc50c4046afe2045386aec2c2a80b6edf2558345c49d70

Observation 41797f56-e347-427e-879c-d3aa95e037af · outbound

This paper cites Mpax: Mathematical pro- gramming in jax.

Learning Safety Constraints for Large Language Models Mpax: Mathematical pro- gramming in jax

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.069748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.069748Z digest=sha256:96d14788aab69ff986909a89f673c8af1cfdc3bfa6d3cd03db0b6a92e2e0db5e

Observation 5df8c693-4082-40b4-82ae-d1d98aff73cb · outbound

This paper cites Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization.

Learning Safety Constraints for Large Language Models Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.143593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.143593Z digest=sha256:3a270f9e042ab69c849df10ed924517825bae35acbf68a32c7a248f8b69a69e3

Observation 2481352e-529d-4429-8213-38d98983fd1c · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

Learning Safety Constraints for Large Language Models Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.217942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.217942Z digest=sha256:9e6e0bee96f8d7929f222ddb3390730382cb77697316f7d30dc605a858786eae

Observation b459100a-c042-471c-be37-2187899fe70f · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Learning Safety Constraints for Large Language Models Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.287165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.287165Z digest=sha256:77b72ba23cc83c01494722ed5d155604bc73fe47a0f2d5502892c757481dc846

Observation 74b81ec1-c15b-4fb3-a0bb-fb078d10d9bd · outbound

This paper cites Token-level Direct Preference Optimization.

Learning Safety Constraints for Large Language Models Token-level Direct Preference Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.341645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.341645Z digest=sha256:a27fe0d433d30e6f8c5d19aa72d4b01d7e17aaf6876593af84c60bb31d91beed

Observation d0b25abb-0b16-4501-8622-7fd70047f051 · outbound

This paper cites Panacea: Pareto Alignment via Preference Adaptation for LLMs.

Learning Safety Constraints for Large Language Models Panacea: Pareto Alignment via Preference Adaptation for LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:16.406578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:16.406578Z digest=sha256:da93a88f5e9ca85da85c06c291803246df84b02204d661e12412ed9bd3447a95

Observation cc3e2173-fffc-4d37-aa3e-989ffb38cc47 · outbound

This paper cites Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization.

Learning Safety Constraints for Large Language Models Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:17.936203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:30:16.462865Z digest=sha256:3f63cc0a9fa61b5b7e61a815e6b809c91f2d34c13a10052dda4e36c32233a431

Observation c06c74f7-cc17-4d49-8621-04f06d5d164c · outbound

This paper cites an unresolved cited work.

Learning Safety Constraints for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:30:17.725886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:30:16.506016Z digest=sha256:2958d6429b27dc9b150c6417349352e92b4e6fccc7a349c49d8a9c151fb6d7b5

Observation 947a66c1-85a2-4efe-b390-8d096005b368 · outbound

This paper cites {human question}\n{model answer}.

Learning Safety Constraints for Large Language Models {human question}\n{model answer}

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:30:17.579256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:30:16.558669Z digest=sha256:90014f05cdb41344ba84feb920ae1018f5ebc6a8c7461d40eba0c12215eaeafb

Observation e2f0e159-4e83-4861-9165-53506c0bd2b9 · outbound

This paper cites Results show mean ± standard deviation.

Learning Safety Constraints for Large Language Models Results show mean ± standard deviation

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:30:17.301765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:30:16.628557Z digest=sha256:1b62948b8ec9f03a5ec570043edd743a666ea6fa8a630b75e0dedaa13defa051

Observation 12ed9f93-7ac7-444e-845a-2394f397f838 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Learning Safety Constraints for Large Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:14.786704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:14.786704Z digest=sha256:aaf4a06b4ff49fbc939b9c530c7d5979bd4739ce5835084e2cf1d4b2c7ffbe3c

Observation 3fc8dcb9-694d-40f7-bfa6-435116fa0bca · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Learning Safety Constraints for Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.888021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.888021Z digest=sha256:a04a577c003d80e3306352a44491550ac687ddf9ae10e84e3cb6a7b45b982b82

Observation 5691e74e-0f86-4b94-8d91-7297b889fcc0 · outbound

This paper cites Interpreting Neural Networks through the Polytope Lens.

Learning Safety Constraints for Large Language Models Interpreting Neural Networks through the Polytope Lens

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.138252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.138252Z digest=sha256:f2426f658ae9ab3184d846bf270af9b9d37fd642f433f035b6c5a4173dfecfae

Observation 62406f3f-8d92-4d13-9b40-ac547bef0427 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

Learning Safety Constraints for Large Language Models AI Control: Improving Safety Despite Intentional Subversion

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.698528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.698528Z digest=sha256:32915a9abce6fca7af71af59a3f20702e02f43fc9fe5f5092e6dfc17502b4656

Observation 4c4f088e-3a56-4d72-9850-0fc309a42dff · outbound

This paper cites Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts.

Learning Safety Constraints for Large Language Models Unlocking Decoding-time Controllability: Gradient-Free Multi-Objective Alignment with Contrastive Prompts

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:30:17.032993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:30:15.545665Z digest=sha256:68f02376f0e73d759ebd0a0c0921e75408774ec516a10ef3f6c6f3beb66b1152

Observation 28a20280-8c32-43e8-b0aa-224937fb62cc · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Learning Safety Constraints for Large Language Models Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.271464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.271464Z digest=sha256:c1c533a5db78e6cfe6cf9ce22cffd1fcedd9a4328d0a186a9416946b4679ecbe

Observation 05e6391c-18c6-4105-a685-75cfa16f0747 · outbound

This paper cites and Bartlett, P.

Learning Safety Constraints for Large Language Models and Bartlett, P

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:30:18.150799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:30:14.835783Z digest=sha256:cf4f57c0bf49b1a25f2758e9c6140c2c1663940aaacd2642917b05c3880450c0

Pith citing papers

Observation e2bd960a-5dbb-4c95-9548-c457ff5dc172 · inbound

When control meets large language models: From words to dynamics cites this paper.

When control meets large language models: From words to dynamics Learning Safety Constraints for Large Language Models

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:54:13.126908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T14:52:44.632671Z digest=sha256:41fb55d9cce84028f3519bbd093f21aefa7e59bdc8e046e0ebf96131682467d7

Observation 9036905c-b097-4870-b700-ba09c4020dc2 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Learning Safety Constraints for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.216730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:07bbcc6e0bf18ef048bc0622a56af4acd13fc18736ac68ca454a3009ad450c65

Observation c9ba4192-0e7a-4ee5-ab2c-d234dde2bced · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics Learning Safety Constraints for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:01:20.847532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:222a8102658997a67a61447c522e5e00189a18a1918e2ce9bd20c87700035bfb

Observation 511e8842-49c9-4e3a-a11b-475e12eb343f · inbound

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance cites this paper.

Greedy Coordinate Diffusion: Effective and Semantically Coherent Adversarial Attacks via Diffusion Guidance Learning Safety Constraints for Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:43.438995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T04:35:35.594085Z digest=sha256:36966f5a75770cdde0146dabd2515426e126f7a798ed734fd16f4a395a255225

Observation 9d400a3f-0fe4-42db-8206-794618102f72 · inbound

Geometry-Guided Constraint Learning for LLM Safety Classification cites this paper.

Geometry-Guided Constraint Learning for LLM Safety Classification Learning Safety Constraints for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T11:57:51.928730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:57:51.928730Z digest=sha256:f44a0c0975fc69299eed28fd7e9b995f1ee477ff73d96481ea9db40c1f9bc4a9