Pith. sign in

Paper Citation Record · LEDGER

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

As of 14 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.09624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09624 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:57:41.323535Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11f76a94-d370-4d21-be28-6eadc3d8b2c6 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Detecting Language Model Attacks with Perplexity

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.177641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.177641Z digest=sha256:83632418bdd44f75d25da611b26f1b881482e5a1a533f7d0e95018bcbe17c211

Observation f89b53b5-dca4-4b80-b5ba-2301e2351e20 · outbound

This paper cites A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.182681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.182681Z digest=sha256:70c8ae431574330a1307da34df9f707f5e146ab2c8714306819637da73dd79c1

Observation 3b83a876-54e8-415d-9f83-f43a69bc846b · outbound

This paper cites Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.925626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.187216Z digest=sha256:aa8518181c593e6d6b35f00e68a2a48fe83b93cf3131c9fd5568c141cb5a1598

Observation a04cc204-349b-43ec-b09d-ddc3bea168d4 · outbound

This paper cites Refusal in language models is mediated by a single direction.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Refusal in language models is mediated by a single direction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.914862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.191487Z digest=sha256:3e641564bc2708442e1d38624df630c36fe461fb7fb40c23ba4ea333d2f8d6c6

Observation fcf626c2-ee70-44d3-a918-c569af0fc9f5 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.195732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.195732Z digest=sha256:f596a295a19d324fab28c4099e1a065746e2ae7be6f2e9d295a4abb7e2c78eda

Observation 0d2f74b0-54f9-4ce0-88f5-f5513b2b2cb5 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.904148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.200166Z digest=sha256:842ad7ac58addb679d3d7f8c71616bf281962e7928e87aa6ec1326e47ddf27cb

Observation a4970e7e-fd06-4e0d-bd95-81a509f36856 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.204173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.204173Z digest=sha256:a80ec3f779f78755271cdd40f9e292801f867c2caba309082ea59d9b6a58e1ef

Observation 51222e43-59ca-49bb-a4c2-74c796d1d1f5 · outbound

This paper cites When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.208214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.208214Z digest=sha256:f81bdd2fe02473d2dcfeb609a01e1da67dc2f609ffcacea0eefaa8ff6588308f

Observation 3e8eb405-9d1f-4f3a-9839-316a080594e0 · outbound

This paper cites The Llama 3 Herd of Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.212053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.212053Z digest=sha256:4f7679333a978c957860513e4b1dce5d51719e00f6f3c9a9b076becf119bdb71

Observation 9802d500-019d-42fb-aa2b-d47c3f598744 · outbound

This paper cites Designing and interpret- ing probes with control tasks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Designing and interpret- ing probes with control tasks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.892835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.215699Z digest=sha256:58c752b87b4b88cb7ee18db758cfa2a2455c2128cfc62cc77cb37bf307ecda7f

Observation 043c4b59-721a-44df-a834-1fed402e9e18 · outbound

This paper cites Best-of-N Jailbreaking.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Best-of-N Jailbreaking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.219248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.219248Z digest=sha256:47bed29315dea54d69c2806c0046f7dd4e1418e8ab3b8b718e8b1433d95cd6bd

Observation 4b555bad-6552-4f44-8bb1-b2a7197f9a9f · outbound

This paper cites Attention Tracker: Detecting Prompt Injection Attacks in LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Attention Tracker: Detecting Prompt Injection Attacks in LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.223332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.223332Z digest=sha256:c4c64836c22d7efb6e3cb6ab447a65a8ac37206c7276a9d76ac706e0f7a80379

Observation 287afe03-ecfc-4207-9b98-fa687975c303 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.227334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.227334Z digest=sha256:214703cacd3dc108d8a88633e80d5eb13a7d24f106941196f2320260f986ec72

Observation 7f6fd031-542e-4f9a-8a8d-d847d0dc8f38 · outbound

This paper cites BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.881537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.231145Z digest=sha256:99d5e3440bbdd32a26d015cbefa79558228f1e02d04bb08553ea085d83cec451

Observation aea13e5c-18fb-4152-8bf2-93a40c74b98b · outbound

This paper cites WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.870448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.234667Z digest=sha256:11614b1e28b20b27c31811a77d3a599561126b848f7b2bc720849fc27a4cfb16

Observation 160cb7af-1975-46cf-9fd4-df20e894231f · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.238272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.238272Z digest=sha256:a05a9a4d3f6f008a00f62cb3a1972bd4f28bb4341cd0fd65c10a6c5be03c6315

Observation 576ff26a-77ce-40e3-85ba-3127b401354d · outbound

This paper cites What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.242047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.242047Z digest=sha256:dc26a05d028d1d18ccaa21f0175d5670ae6fd87d58bb698fc757d4ef32ab6ad0

Observation b88a2ff3-48d9-437b-8f6e-4c6d4157d603 · outbound

This paper cites A simple unified framework for detecting out- of-distribution samples and adversarial attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A simple unified framework for detecting out- of-distribution samples and adversarial attacks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.858768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.245714Z digest=sha256:a3c3d2a30eaa694cd3f63fd67334de3281d656d618c100af193e6ce5f12a4544

Observation 83c8478d-1212-4d6f-8f7b-1d48cc5fbd13 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The power of scale for parameter-efficient prompt tuning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.847377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.249320Z digest=sha256:1a0e5a8635eb2798cf6aa72155c06e9ab9d64e5496ebf2b432570feaef1f81b2

Observation 6dcb5a50-2ca9-45d2-831e-0245c8884244 · outbound

This paper cites Prefix-tuning: Optimiz- ing continuous prompts for generation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Prefix-tuning: Optimiz- ing continuous prompts for generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.835771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.253003Z digest=sha256:93a069171db8cd290fc03718db68b8c4bfd0a61e12bd34e1ea5500eaa23eaaa0

Observation 2a46e70c-dcd8-49cf-8743-d965ef4f9f7e · outbound

This paper cites Towards under- standing jailbreak attacks in LLMs: A representation space analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Towards under- standing jailbreak attacks in LLMs: A representation space analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.823289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.256420Z digest=sha256:ce5305af5b8d710017086948680f5306b68b72cca5289509e4e11fb1be83bf27

Observation 377f6f05-7647-4b7b-9270-e986d404555d · outbound

This paper cites AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.812145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.259805Z digest=sha256:a958cc0446d160fbfe8be8d85bc48cab1b19c31b1dc1fbe91c8d5ec35f866ac3

Observation 7dc85007-6352-4923-b4ec-b9c434b1d8e3 · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.800603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.263412Z digest=sha256:4809c7d722c7614876a1d6d68c2484370300e7bce64fc0dbc29395386b36015f

Observation be4a1f39-a700-40e1-a3a4-3b7bf060fc36 · outbound

This paper cites The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.788973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.267711Z digest=sha256:1c86fc66cc8a19f477b9d6b58e3434cb00f360bb673d84c51d64a36b04e0198a

Observation 55173b51-9ee4-4f7c-93a6-1d736bb659f6 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.777446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.271441Z digest=sha256:8828991612374f77e1e653507d08342d850d204517125f9f0109a747e3cae683

Observation 0234a541-4364-4790-a6b3-1e78e02536ea · outbound

This paper cites Tree of attacks: Jailbreaking black-box LLMs automatically.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Tree of attacks: Jailbreaking black-box LLMs automatically

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.275245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.275245Z digest=sha256:92f8473177dd16eab601b560f218c89244a2c40666b5f454ef43e42d019bd59d

Observation 30f0bf98-5fd6-4687-b53c-c1224faa306a · outbound

This paper cites JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.278918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.278918Z digest=sha256:f874f0e30ade8cd2e674276e257639075d34ec352036e44a1a56074278c20203

Observation d5b8e7e8-33f7-49e8-98d9-ac29d0d6d92d · outbound

This paper cites Rebuff: A self-hardening prompt injection detector.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Rebuff: A self-hardening prompt injection detector

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.758950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.283315Z digest=sha256:a55a08762d80e161be9368ebfdee8abf1669b2b7060fa77fa0c6d47c5cc1aa20

Observation 003fb30d-91a6-4429-8abd-2ca3371ff7c0 · outbound

This paper cites I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.286981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.286981Z digest=sha256:a33e1541647d14a50c949ef3a7dcef07ae6202443b7554e721fc05bd0cabbdcf

Observation 3f717bca-7e3f-46ac-8393-3b7154d6af70 · outbound

This paper cites Do Anything Now.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Do Anything Now

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.747306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.290741Z digest=sha256:592cc4f76cdbe13a02359a67268a99622de193a79354ab10b57668df193555c4

Observation 99c6b68e-5387-4f80-8cf3-a42747243495 · outbound

This paper cites AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.430736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.294843Z digest=sha256:fe89308147720d9283cf478298c6fd8c20c32832e2288c6f4ea34292542a8e45

Observation 230cbbd7-198a-457f-943e-17a00cd94097 · outbound

This paper cites A StrongREJECT for empty jailbreaks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A StrongREJECT for empty jailbreaks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.735206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.298913Z digest=sha256:fdf8fd51b9f8824763222aa71da816bb1f6313918861ce67814094f53ee9cb25

Observation 63302013-a04c-44d4-b23e-efff0de17d93 · outbound

This paper cites AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.302502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.302502Z digest=sha256:4fe146530da27745dfb6fba315315645e74879956161e4fde7b4af28d904d4b5

Observation 877cd1fd-863c-4801-9533-8dec62aaf5c0 · outbound

This paper cites GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:57:41.723347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.306733Z digest=sha256:3dcf53d0ffac8d8e235b66dadb71f216fccedda73911323278cbec1277fe6fdd

Observation 968afd6f-ffa7-4c97-bb32-657975f181b4 · outbound

This paper cites Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:57:41.401067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T13:57:41.310740Z digest=sha256:900683af343a52dbb046c22565d133cf4f748809e3fd9adf7d80cc65a771191b

Observation 7f4db920-1fcb-4cce-bbcf-8f0984e851a9 · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.315199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.315199Z digest=sha256:34603fe293f9504ba0386e66fb977195ebc56829e63eb2a5a5e5d74ebed2c6af

Observation 77cabf51-e992-43cd-ab3e-84a6343143ba · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.319506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.319506Z digest=sha256:b664a0068662168da13cfba980433a2675b7d4cab1f0ee82197b0b059a2bb392

Observation 4c6f2e7b-e024-4d6f-b550-23e3816fc8a9 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T13:57:41.323535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.323535Z digest=sha256:4f1db3c0e369bf68e96d659d239547020eacc6e227372f381d335e259b5757bc

Pith citing papers

No inbound Pith citation observations are available.