Pith. sign in

Paper Citation Record · LEDGER

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.02665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02665 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:48:49.845907Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73d5e304-eaf1-4e55-adbf-64fc2f51e490 · outbound

This paper cites Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:47.691899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:47.691899Z digest=sha256:68b456d9d2a15fe4e7c452e5c9b12c0814f889916e39efbc92942b07c8e42689

Observation 60ade032-b860-4a8c-bffc-634a86ebd9d8 · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:48:53.100173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:47.816131Z digest=sha256:99d6aa2d5d4583c97685655d64f1f5468a14685542e39c6a1c22dd5d0e6a8a37

Observation 58f0fcc2-0fb7-4eac-9743-37fcf82efc48 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:47.873645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:47.873645Z digest=sha256:faf149256ed460cbdf7b58da1a127489ff6163b98bf8f5a7da4dcc7b5d178a05

Observation 5dc7053d-6fff-4f28-a9c3-00cb61985a9a · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.868473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:47.964952Z digest=sha256:15e3cad1050283e671b8bf8fea24113bbb9234fa76c5403ffa768e60932a5216

Observation fe14dc20-04de-4e94-9e16-600890e0727c · outbound

This paper cites Best-of-N Jailbreaking.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Best-of-N Jailbreaking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:48.102460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:48.102460Z digest=sha256:30648bb1f837fde2d49139e0ffe3a244a654ffe742eeb175a230036809531420

Observation 6a825f54-2d85-4f97-bced-821b5569187d · outbound

This paper cites Jacobs and Hanna Wallach.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Jacobs and Hanna Wallach

Reference 6

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T00:48:50.695454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.228849Z digest=sha256:f2330d7c1efcc448abae699f032f8680f63c080707676a9a8f1cc9ac629a11b7

Observation 784b503d-3cbf-4283-86cc-58e181a40db7 · outbound

This paper cites A Cross-Language Investigation into Jailbreak Attacks in Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity A Cross-Language Investigation into Jailbreak Attacks in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:48.346153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:48.346153Z digest=sha256:60add0ef7ecf29b27fd4bb1fb0e380534e30cf54481a1f9e4be7710cc27e7d5f

Observation c1131312-66ab-401b-9424-f9a97a5eef6b · outbound

This paper cites Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:48:50.397734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.474079Z digest=sha256:786069fad20617b35d59735505d870286b0001f385ca8158843232a3f7848581

Observation 46e4f696-0677-40f2-bc22-4e2c92e88b0d · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.681403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.573825Z digest=sha256:eea87bfd618e0d84f977a8b293119ac197ebc496cf67169b54dfa8b03bd1b855

Observation bf59ab63-6167-4b78-957b-e628b03c71b0 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.434370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.689413Z digest=sha256:14dd9375ce709526dbecbb3f6eb6a4cfa43e48633d0a4c8c52e901bc19237a31

Observation ba4de435-e0d6-4f05-b7c2-f8a12de7a683 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.149128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.797993Z digest=sha256:16848d748fb580d833e5cc87a8ed9e80ba097f3fe2a91982b02c2ecf44df376e

Observation c39f7f01-617c-4c84-ad77-4574af70a24c · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.827995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.913829Z digest=sha256:70978c080c23e30f101cb93bd14ea78376637293dab3ba432c9b53e882da9991

Observation ddfb7a64-a005-4b07-ae67-0b2cca66b80b · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.587110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.069998Z digest=sha256:5ba9bd22b55e688cf98871cb78b4f0fd69380b8e1fa8cfee77521d35b719e5a7

Observation 5c0b3f9e-5023-42cf-bfe0-6d65e5718c24 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.223511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.223511Z digest=sha256:9e05eb374577c41f1451d9aaa77438a7b7cbb1cabf843540134b903856a88998

Observation cebaace7-c0f9-4320-b9e0-41c4c5bacbfc · outbound

This paper cites When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.304875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.304875Z digest=sha256:c14a56d2c56b4c511d18589a68d3fb7ec3ee9728cbfc32c58b8afd67b164c8fb

Observation 71d8efd2-22c5-40c3-b32e-9750cd307da2 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Low-Resource Languages Jailbreak GPT-4

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.383630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.383630Z digest=sha256:7a2b507b5dbc8dd8854a1c440844e14d98ce247b66c414d77e8e4ad30171ae00

Observation 90c79157-601a-4a9c-b628-c8fb11fc3899 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.352287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.488227Z digest=sha256:017abdb5c86764d49e499d72b7e91ee38e1632919004d2bf3aefdb8044cbe0e2

Observation 5eb38f5c-d248-4d90-a46d-cb6fd5a2b453 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.605652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.605652Z digest=sha256:99606cdf7ff2098e290b0ff32297f5be266c68abe2d1b5d58d3ee966e8988d6a

Observation 8f71313a-8b27-43e3-8cfa-672f2f284fc4 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.039113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.690266Z digest=sha256:c1bc9e0466664d9f73b8706de5ee3bfed2026189edb82bf2c7e54ed1d913e98c

Observation e3c2cb12-e5d4-4d55-8a36-ee36640ae800 · outbound

This paper cites Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:48:50.091776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.750627Z digest=sha256:f28dc795fe3ce346211a11c6bac8f3b06d85b9081e6eb3f85e1527a261c3395c

Observation 75d7b4ee-09fa-432f-8c66-4c1a3dd6297d · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.845907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.845907Z digest=sha256:85f88a036eb3eb023f89fa96af9ba562057ab4f303a1b37c81d1190865bb29f5

Pith citing papers

No inbound Pith citation observations are available.