Pith. sign in

Paper Citation Record · LEDGER

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs

As of 11 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 3 inbound Pith citation observations for arXiv:2501.02018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02018 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:35:27.013254Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:18:41.044579Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T12:54:07.656435Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b11812df-a75e-499f-a9e1-35055655b0f2 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.875258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.875258Z digest=sha256:3e1246841a09bef1a338e896b8e1d57b95e1c80d67d7756fcb84ab24e6fabbeb

Observation 559dfc18-6046-4501-81c5-960609ef393e · outbound

This paper cites Nudging: Inference-time Alignment of LLMs via Guided Decoding.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Nudging: Inference-time Alignment of LLMs via Guided Decoding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.886057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.886057Z digest=sha256:395e8ee1b47c4e48b69876a974d6667742b460322a906553b9f486eddc990ca8

Observation 3bc3727f-5260-4f4d-a829-963bcee6f22a · outbound

This paper cites LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.891500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.891500Z digest=sha256:7032922774c2666ea9c4083faae5b7629fa35fb8bffff0977dcc564b6d780cbd

Observation a1b025e5-7b99-4884-840d-ea2f15c73eb4 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.896977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.896977Z digest=sha256:eaa8d252e525b363f5c607ed2bec7c0c85caff29ca6bba408badd4ce5bda021a

Observation ba8e6823-c1d7-4ac5-a38f-12ad31215c0d · outbound

This paper cites Malla: Demystifying Real-world Large Language Model Integrated Malicious Services.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Malla: Demystifying Real-world Large Language Model Integrated Malicious Services

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.919917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.919917Z digest=sha256:9eee5fe4c4f1d8de6dada7bed30df0d41f4b5d2b785d0d0affe94f1ecf5e883a

Observation 4835ab57-1cb7-41c6-a964-c80b41c647e7 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.925688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.925688Z digest=sha256:077022c11563b65ae82922d7e5ac133bf8690e1fdb69ddc08450de553bcf0b58

Observation 8fceeec4-0ecb-4777-8147-795effd35506 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.930644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.930644Z digest=sha256:f1ecf2aaba4d8d0f836436dc572a243945e2128c613906299ba270f2d28ba476

Observation 12a1b381-7c3f-483d-b292-82a9dfcb8018 · outbound

This paper cites arXiv preprint arXiv:2408.15625.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs arXiv preprint arXiv:2408.15625

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.936571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.936571Z digest=sha256:cad3d0542eb6df198e2a77008aa93282b6a581390df8f13e5060ac21e6cfbe58

Observation 6119d039-e7b3-4081-8fb2-678139218fa9 · outbound

This paper cites Red Teaming Language Models with Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Red Teaming Language Models with Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.942938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.942938Z digest=sha256:afeb3f0d20d098c11e37bf2c0ad948f7fcf6e8bd397c64a134ef51c401002bc4

Observation 1b6ebc3f-f1d0-4aca-94ae-fcb85c3cac89 · outbound

This paper cites Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.948582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.948582Z digest=sha256:4a0499574cf392a51960a3746c5457cbab01b3a45235a0be21d31831881e6609

Observation d1ccc78b-cced-467b-ae26-ea4e64fe8d4f · outbound

This paper cites Controllable Natural Language Generation with Contrastive Prefixes.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Controllable Natural Language Generation with Contrastive Prefixes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.953888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.953888Z digest=sha256:e334e2933591a02cf008224d787492c5bc499dec5223548bad179f82e6c45e19

Observation 7a326ec1-0c5e-4501-a40c-99b3d892d5d8 · outbound

This paper cites In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:35:27.575761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:35:26.958828Z digest=sha256:26a00f60ba2dd9aaa6b829b6073e261d38c5c6957e447ea2d14fc7edb37adf97

Observation 26473cd4-d6af-4c70-815c-bc37de93f660 · outbound

This paper cites Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.963890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.963890Z digest=sha256:a8ab9a8984653a33ec7ada5557b882690df9af64ae1c3caa9a57c39afb4200ee

Observation 37970b0b-1c64-4e4f-bf80-8bc2e3f569fa · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.969237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.969237Z digest=sha256:5ffc34a6b2872a3297425a049bf4c338ecfc92ecdd33a2791a77a76d13f5e2a6

Observation c2c08826-62ac-4dc9-a540-48f58aaabbc5 · outbound

This paper cites Self-Guard: Empower the LLM to Safeguard Itself.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Self-Guard: Empower the LLM to Safeguard Itself

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.979803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.979803Z digest=sha256:5a60a9f4b0f1ee94788e4e32284b561e519aaaa4ff5353c4fc780c5d25e59e58

Observation d6ef1e05-0fd2-4d45-bd4d-a8d6458c5f61 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.984854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.984854Z digest=sha256:2ddb50f620515d6d104a4fb709f04d0919783883745b63e65a69c833370a64a9

Observation 5ef4fb56-bdc6-405e-9fff-bfa24ccbd7bb · outbound

This paper cites In Findings of the Association for Computational Lin- guistics ACL 2024, pages 7432–7449.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs In Findings of the Association for Computational Lin- guistics ACL 2024, pages 7432–7449

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:35:27.558205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:35:26.990851Z digest=sha256:3eead90464a97df62a26c563daaa4d3f53bbc2d811df8986e212d675cd4694ff

Observation a76bef2b-afc8-45d7-b375-199041cb38a0 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Low-Resource Languages Jailbreak GPT-4

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:27.002467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:27.002467Z digest=sha256:d52ead054ffb5a0532f6717c68d0dfa6e74530c8c663dbd87d934838cf769339

Observation 85e48ec5-0dfa-49f5-84f6-0a6fea4bd398 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Instruction-Following Evaluation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:27.007735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:27.007735Z digest=sha256:d7c13ce405655187902e6f3308f0bdacf6baec66949f09a66a092c4651aead9d

Observation b8c638d9-7e46-47fb-a89c-5652165497e9 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:27.013254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:27.013254Z digest=sha256:98824ed0dc56466b64143351594e3d529af31829eb5a19c928b5231a71e53b2e

Observation ff7fe4c1-f2c9-40e9-a15f-d1894817d866 · outbound

This paper cites A Framework for Real-time Safeguarding the Text Generation of Large Language Model.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs A Framework for Real-time Safeguarding the Text Generation of Large Language Model

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.880697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.880697Z digest=sha256:7409977d6611258e277593289ad05e87af5f079c508266249837eda6d6f9c502

Observation 2b35992b-33b9-42ff-928b-382d8ce35a2e · outbound

This paper cites Machine Teaching: A New Paradigm for Building Machine Learning Systems.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Machine Teaching: A New Paradigm for Building Machine Learning Systems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.974603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.974603Z digest=sha256:76cd1385d4c615eee253a52c74baeb738828476ee836638f65df33b13860c713

Observation 5b45da61-d38f-4f7a-8825-1bdee77aea47 · outbound

This paper cites Situating Sentence Embedders with Nearest Neighbor Overlap.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Situating Sentence Embedders with Nearest Neighbor Overlap

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.914579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.914579Z digest=sha256:09a4d67d2ca529f0e2297dc43f61b020638b2da91dba05d1d714bffbfe0797f4

Observation 25bf91ad-1078-4808-8f63-e48ad7594a15 · outbound

This paper cites GeDi: Generative Discriminator Guided Sequence Generation.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs GeDi: Generative Discriminator Guided Sequence Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.908543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.908543Z digest=sha256:9ef6374fc1fea8c50b2c89ccdc7201cd2c94dfebb3253e4c993541a347a64f73

Observation 94a17b93-d682-49d3-a001-16291b145f60 · outbound

This paper cites FUDGE: Controlled Text Generation With Future Discriminators.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs FUDGE: Controlled Text Generation With Future Discriminators

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.997182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.997182Z digest=sha256:c7d81717f2b8d51b1ca343e6f6b6b0c56a7ce3f66505697eece387c3e99aa837

Observation d264333e-f9dc-43d0-8b24-3d181803a16e · outbound

This paper cites Critic-Guided Decoding for Controlled Text Generation.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Critic-Guided Decoding for Controlled Text Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.903529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.903529Z digest=sha256:9eeb5d279cabeb2c9a00f33141e1a291b3a4832720abea388cac6c6349ac6e86

Observation 44ce380b-daee-45f6-aa11-b23be1c1c8f6 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Detecting Language Model Attacks with Perplexity

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.870090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.870090Z digest=sha256:2436848a5e9af234eb3e0ff3f86db31f4b4a29a5a6e46813269c5cbab07a8406

Observation bf947123-d390-486c-a39a-a8ba06f22841 · outbound

This paper cites In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:35:27.591525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:35:26.864276Z digest=sha256:f48d9ff31d757ed70461dc896dea12e8bf9b9a8feca74c641f6dfaf6f0047d96

Pith citing papers

Observation 88536e69-d799-494d-9d17-6eca02ab0ea3 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:41.044579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:41.044579Z digest=sha256:af3ebc858ec1fc9ff4d7486a64a3f87f35fe50b82c7d991712e35e762ec1ca04

Observation b9e927c0-0ce1-4119-870b-44d6bb06c864 · inbound

The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems cites this paper.

The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:18.014443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:18.014443Z digest=sha256:f9fd76a936bb1c39b2ca9272e9ff43b3177e7a5b7b7b5fb64b5a1451fd7fe473

Observation 105f2060-0c23-46f7-adee-753beda65c02 · inbound

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention cites this paper.

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:54:07.664223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:54:06.525256Z digest=sha256:a05700e3a4c4364ffe904ec3493a85e71fefcc1a52826d30e125df7a219fc043