Pith. sign in

Paper Citation Record · LEDGER

Strategic Deflection: Defending LLMs from Logit Manipulation

As of 15 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2507.22160.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22160 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:05:44.625951Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:24:24.267269Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 21742424-0228-49b6-b944-22f39c585070 · outbound

This paper cites Training language models to follow instructions with human feedback.

Strategic Deflection: Defending LLMs from Logit Manipulation Training language models to follow instructions with human feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.381035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.425087Z digest=sha256:447d5f21f1bc41adff09c5deae43a36da8ffdc8b67b8e73212b855bfc882b81e

Observation 16a914b2-ff48-422c-91e7-ace163ab7bcf · outbound

This paper cites Deep reinforcement learning from human preferences.

Strategic Deflection: Defending LLMs from Logit Manipulation Deep reinforcement learning from human preferences

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.360768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.433864Z digest=sha256:ae4929c1064cf81b42f8d653df739a96f2eb88dd9e6f308aed80859ec53fed93

Observation dfa9d83a-feca-46da-a357-460ec701c679 · outbound

This paper cites Mistral 7B.

Strategic Deflection: Defending LLMs from Logit Manipulation Mistral 7B

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.442565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.442565Z digest=sha256:27322cc1067cb093ada8fa6b3cb4a25cb7c37ff6c9ec3f70ff62320bbc3010f8

Observation f09ff4ca-dc03-46de-b451-fd2cba7fc169 · outbound

This paper cites The llama 3 herd of models.

Strategic Deflection: Defending LLMs from Logit Manipulation The llama 3 herd of models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.454732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.454732Z digest=sha256:574b7833cb17fc2319f34fb808f0d46b0ac5d07aaeb1ca4264cc2cec307c547a

Observation 8b56acc9-7aa6-4136-ba2a-9940a912e427 · outbound

This paper cites On large language models’ resilience to coercive interrogation.

Strategic Deflection: Defending LLMs from Logit Manipulation On large language models’ resilience to coercive interrogation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.320699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.462785Z digest=sha256:e7ab2a26f6052b0d04f1971a8d800044cbb1c1422472894cf028d555c7b0d10a

Observation 6ab8a4e0-999b-4b93-a1ca-820f02528e21 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Strategic Deflection: Defending LLMs from Logit Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.471626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.471626Z digest=sha256:8b664667704127f5bd084867c0abe36af3bd669aabb2a4ab66d4e0663f0a3c38

Observation 5d7d4768-5c38-442e-900d-7453e798e9ea · outbound

This paper cites Jailbreak open-sourced large language models via enforced decoding.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreak open-sourced large language models via enforced decoding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.297838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.477735Z digest=sha256:5df7964245211a405597672dde8982e723ae54c8ff59c887e75c01ba2b603637

Observation 99440afc-c6b0-4c6b-9eb2-c10f3cdb3c0e · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Strategic Deflection: Defending LLMs from Logit Manipulation Safety alignment should be made more than just a few tokens deep

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.275737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.486012Z digest=sha256:904ab0a22278966ac58bbeb9c6901edf4cea8af5a954dc7f4b14a6edb9efb59d

Observation 180b2f18-aaf7-4ad0-86ef-3d9682ca6f8c · outbound

This paper cites Jailbroken: How does llm safety training fail? In Advances in Neural Information Processing Systems , volume 36, pages 80079–80110, 2023.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbroken: How does llm safety training fail? In Advances in Neural Information Processing Systems , volume 36, pages 80079–80110, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.251296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.491524Z digest=sha256:54463add87f6ace085a045b55eeea98777683611fddb532de8a91f1d3c327000

Observation e838db4e-e370-4e8c-b125-56c08475e087 · outbound

This paper cites do anything now.

Strategic Deflection: Defending LLMs from Logit Manipulation do anything now

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.216533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.498317Z digest=sha256:b5b416bbda1595cdc0fee8763c15b0d5ada215061c378c947ce558c0b830f1be

Observation 4fd0b2e3-1ba8-43d1-9d0e-8a75fcc6e530 · outbound

This paper cites A Cross-Language Investigation into Jailbreak Attacks in Large Language Models.

Strategic Deflection: Defending LLMs from Logit Manipulation A Cross-Language Investigation into Jailbreak Attacks in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.503676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.503676Z digest=sha256:938ab6d165b14f9e031cadbd57b48b40b9cef7d312f3ed333f6e6c0b31f78bf5

Observation 6b12d06e-437a-4d7b-9924-6203a60c4777 · outbound

This paper cites Protecting your llms with information bottleneck.

Strategic Deflection: Defending LLMs from Logit Manipulation Protecting your llms with information bottleneck

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.192046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.509524Z digest=sha256:383941d447ce38783083040c15ad780f999909e42121f81ff33115fd8af66b07

Observation 959e6f3d-fa51-4b4f-ab98-39129bcc81b5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Strategic Deflection: Defending LLMs from Logit Manipulation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.514125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.514125Z digest=sha256:5b3fec1c2cc68fdc8d98877844fb59025222bc7bfc102e9bf07c89326c2ab10a

Observation 7f96471a-2d22-4819-acaf-8b43d1aebc20 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreaking black box large language models in twenty queries

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.169706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.520627Z digest=sha256:07dfb1dde52dbd652e86c7f5a483b7582b1554de2812d70f75d0d24994a537ea

Observation ec7747ca-aa95-40b9-829b-cb04fb8f2627 · outbound

This paper cites Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models.

Strategic Deflection: Defending LLMs from Logit Manipulation Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.145466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.526173Z digest=sha256:9b22fe386a81955f277e1385f1b2fd67fa940a08cbdfad196f29a6bfeaf77ec7

Observation df81673c-0842-46ba-859d-37721d4b7f7d · outbound

This paper cites Catastrophic jailbreak of open-source LLMs via exploiting generation.

Strategic Deflection: Defending LLMs from Logit Manipulation Catastrophic jailbreak of open-source LLMs via exploiting generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.120803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.531779Z digest=sha256:06f4077edc86bd87a505f8ecce5726ab9eb2e1c4bd5190f2c6f28d6ec7e614a6

Observation 3c873539-fcdf-4f0e-ac3e-142987dd9777 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Strategic Deflection: Defending LLMs from Logit Manipulation Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.536146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.536146Z digest=sha256:cc30af75e27dfe1bcd51bbb94e243467ee91126edae39daa6fbb1e3ec9d99ed3

Observation 15f793b9-b949-4dce-9e30-e3e82aa31e24 · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

Strategic Deflection: Defending LLMs from Logit Manipulation Certifying LLM Safety against Adversarial Prompting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.541606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.541606Z digest=sha256:51752adbc7aa11e01c827827dc718e4285acc6fcbca951ba9721382f6092bf2d

Observation fb9a6b82-8580-41da-95fa-7ed979b2a608 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Strategic Deflection: Defending LLMs from Logit Manipulation Detecting Language Model Attacks with Perplexity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.546775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.546775Z digest=sha256:86e30c6c74a887ea4e773c570641f8c7fce35428b8eb2ad4fcb8121d561ccdeb

Observation 7f9a2905-5fa5-454b-8b4f-e8392af0d7b5 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Strategic Deflection: Defending LLMs from Logit Manipulation Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.552474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.552474Z digest=sha256:4fe4b1a1220e7caa1a1288d80a76f2c70a775e4f23e35f512186cb33c46df01a

Observation 5f17508b-2095-457f-8ce1-88baba66d487 · outbound

This paper cites Robust safety classifier against jailbreaking attacks: Adversarial prompt shield.

Strategic Deflection: Defending LLMs from Logit Manipulation Robust safety classifier against jailbreaking attacks: Adversarial prompt shield

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.097945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.558156Z digest=sha256:f0badfb66e835d9bc865daac2b08b30d776828c2c4dd847e8431f9b18b5ceb76

Observation cf1f2e26-6ed2-4ed9-ae91-db14b03b78ea · outbound

This paper cites Jailbreaking leading safety-aligned LLMs with simple adaptive attacks.

Strategic Deflection: Defending LLMs from Logit Manipulation Jailbreaking leading safety-aligned LLMs with simple adaptive attacks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.078524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.562678Z digest=sha256:66bb7821b9f8470870d35c572ab801b737bc53949c8075c2d0023f7377e41908

Observation 96d42cbc-56fc-4e49-9716-d763f85de590 · outbound

This paper cites Contrastive preference optimization: pushing the boundaries of llm performance in machine translation.

Strategic Deflection: Defending LLMs from Logit Manipulation Contrastive preference optimization: pushing the boundaries of llm performance in machine translation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.058799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.567565Z digest=sha256:f4bc785fefc5dfce1b8bef8aa06a8418fd37a61fe2a23a27f2759a08111c0b59

Observation 5a10df76-0587-463d-ae29-e55144495ca9 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Strategic Deflection: Defending LLMs from Logit Manipulation Direct preference optimization: Your language model is secretly a reward model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.038739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.573509Z digest=sha256:07896875146683c0f355e483786c1d195356263737bead172dec157a107de73c

Observation 30688163-6060-4951-9142-cd1519a81b63 · outbound

This paper cites GPT-4o System Card.

Strategic Deflection: Defending LLMs from Logit Manipulation GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.578742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.578742Z digest=sha256:a85312aa57948d215f55b964d690b5c77d711cce35668146f3d2f153f7b84859

Observation 0ccbce3f-2f9f-4247-abac-6a0be005e2f2 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Strategic Deflection: Defending LLMs from Logit Manipulation LoRA: Low-rank adaptation of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:45.019654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.583833Z digest=sha256:622ffd560c78d591e5a814c809b68385d703c87b9538a258463fc9536c655f3d

Observation bc351a75-c6fb-427e-acbc-49c9ac3f7208 · outbound

This paper cites Trl: Transformer reinforcement learning.

Strategic Deflection: Defending LLMs from Logit Manipulation Trl: Transformer reinforcement learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.589072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.589072Z digest=sha256:fe9c4dca8a8ada2c1b12b9ed6150cd7cd2c0436d75faf3aaee84426efbf48c3c

Observation 873c020d-26d4-407f-8a25-49f083c58d78 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Strategic Deflection: Defending LLMs from Logit Manipulation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.597243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.597243Z digest=sha256:0ca0f0d7679b7ecd0387b5652c5f4590e19a1cd8ccf2b697fb913a3df7194655

Observation 0ae32380-0592-472e-9f90-fdee11c65965 · outbound

This paper cites tinybench- marks: evaluating llms with fewer examples.

Strategic Deflection: Defending LLMs from Logit Manipulation tinybench- marks: evaluating llms with fewer examples

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.987019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.603212Z digest=sha256:6f578ba5dc1269854d976eb1d57c1fb2c89a3aa53305c8d59ee5bf7097e36ba3

Observation 7681c4da-35cc-4c17-88e6-1bca9d12be53 · outbound

This paper cites Measuring massive multitask language understanding.

Strategic Deflection: Defending LLMs from Logit Manipulation Measuring massive multitask language understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.966908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.608545Z digest=sha256:16f5d967317548d90c9f088821c1cba1abf3ce1c994c18fd6b4ad26fed8e5aa9

Observation f5ba09d9-18eb-47c9-8748-12d938f62721 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019.

Strategic Deflection: Defending LLMs from Logit Manipulation Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.948768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.614731Z digest=sha256:7807f6a1b663fef87a3730ea497cebabb0994642899d38c97d6df71ae7973cdb

Observation e1449f35-7554-4725-8fee-7f70f47afc52 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Strategic Deflection: Defending LLMs from Logit Manipulation Truthfulqa: Measuring how models mimic human falsehoods

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:05:44.928307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:05:44.620968Z digest=sha256:87ffe66b93bd0b0989ca25c389dbf207cd5010db32185a514600b41ad7ee9507

Observation a8c10808-bfa8-4ae3-b582-4d13098932b0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Strategic Deflection: Defending LLMs from Logit Manipulation Training Verifiers to Solve Math Word Problems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.625951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.625951Z digest=sha256:2041979f9c77a3e5a2151f10275c243462e71af0dd71b01dc0cbead67af677e4

Pith citing papers

Observation 6555fb52-5a81-4d7c-a31d-47c9eba147db · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Strategic Deflection: Defending LLMs from Logit Manipulation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-03T14:29:07.776313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-03T14:24:24.267269Z digest=sha256:efa7504166d0e66ee49059a9a4fe72cc9c776628d3d4395097e910e19632d9bc