Pith. sign in

Paper Citation Record · LEDGER

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.02665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02665 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:48:49.845907Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73d5e304-eaf1-4e55-adbf-64fc2f51e490 · outbound

This paper cites Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:47.691899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:47.691899Z digest=sha256:f6cf94963fc01bab18f5b18db0360290b6d0560657e503a338da765a391b27c7

Observation 60ade032-b860-4a8c-bffc-634a86ebd9d8 · outbound

This paper cites Pappas, Florian Tramer, Hamed Hassani, and Eric Wong.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Pappas, Florian Tramer, Hamed Hassani, and Eric Wong

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T00:48:53.100173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:47.816131Z digest=sha256:7d72ead91d2fbf2e41a440cff3aec4a9255076bc374e258a7082eedf32852bfe

Observation 58f0fcc2-0fb7-4eac-9743-37fcf82efc48 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:47.873645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:47.873645Z digest=sha256:34ecaff778089e3e2ba07395560263ac46893c5599938567c0e7906a0a892400

Observation 5dc7053d-6fff-4f28-a9c3-00cb61985a9a · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.868473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:47.964952Z digest=sha256:e9429fe25135d78e6bce37da5f34f68058204320a03eec1b84200fbc3f7bcd48

Observation fe14dc20-04de-4e94-9e16-600890e0727c · outbound

This paper cites Best-of-N Jailbreaking.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Best-of-N Jailbreaking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:48.102460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:48.102460Z digest=sha256:aa09dcf96874fd67907020148a8079c5c94abfdf70e7a7977567e9d9e8b005e5

Observation 6a825f54-2d85-4f97-bced-821b5569187d · outbound

This paper cites Jacobs and Hanna Wallach.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Jacobs and Hanna Wallach

Reference 6

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-05T00:48:50.695454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.228849Z digest=sha256:e0f7151c72aa5fc881023af97658ee5fe10e9d69f42106b27bf07ece45a14921

Observation 784b503d-3cbf-4283-86cc-58e181a40db7 · outbound

This paper cites A Cross-Language Investigation into Jailbreak Attacks in Large Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity A Cross-Language Investigation into Jailbreak Attacks in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:48.346153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:48.346153Z digest=sha256:b6e7f1a481360b24e3a41b6db389972ba7c29a0fe9b9c74f71cfcaeeea3d9f12

Observation c1131312-66ab-401b-9424-f9a97a5eef6b · outbound

This paper cites Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:48:50.397734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.474079Z digest=sha256:43cce0cd198fe38eaad62ac566e1b88169dfc1367c91177d3f63e99771c4f9e6

Observation 46e4f696-0677-40f2-bc22-4e2c92e88b0d · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.681403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.573825Z digest=sha256:b9f1e59e685944549bb59e79c0522333aa137e233ba1b8be0882cc4bb883846c

Observation bf59ab63-6167-4b78-957b-e628b03c71b0 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.434370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.689413Z digest=sha256:46664295eaf3a371563f32706ec8557381b6441e12f3a4acfc6255ea6d92a199

Observation ba4de435-e0d6-4f05-b7c2-f8a12de7a683 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:52.149128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.797993Z digest=sha256:c4bc9b0912900cd6734606900659f98ad82d37967b898eb828a09ba365275751

Observation c39f7f01-617c-4c84-ad77-4574af70a24c · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.827995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:48.913829Z digest=sha256:ac8ee4fcf79658d52fa44ba2161a46934313c1a1da9435353815866193cbb49e

Observation ddfb7a64-a005-4b07-ae67-0b2cca66b80b · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.587110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.069998Z digest=sha256:af4679f8461110e68f8e36a262b22b2a3aa2511793d9c1ebbbb8c946cdc3b0ec

Observation 5c0b3f9e-5023-42cf-bfe0-6d65e5718c24 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.223511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.223511Z digest=sha256:3949659d10894a2fd0d9ef9b4d527de1a48f622dbad4bc27eff03c25b7fbe68f

Observation cebaace7-c0f9-4320-b9e0-41c4c5bacbfc · outbound

This paper cites When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.304875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.304875Z digest=sha256:cc9017dd4441182c20dd5fe6bb17645be937157c810e65e67e7ee63f5f8abc3c

Observation 71d8efd2-22c5-40c3-b32e-9750cd307da2 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Low-Resource Languages Jailbreak GPT-4

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.383630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.383630Z digest=sha256:302f18305c6acfcc5eb8bbc2126f29eec16a8f1dbf4e133412a9c04e442838f2

Observation 90c79157-601a-4a9c-b628-c8fb11fc3899 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.352287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.488227Z digest=sha256:deb9c612efff33aefdb6d234035052dc838897bbdfc3f82c9e7a04ee223ddb38

Observation 5eb38f5c-d248-4d90-a46d-cb6fd5a2b453 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.605652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.605652Z digest=sha256:f242d0707c1d7d31c58d98c2fb1c40f9964dec2c8476d395b58041bc7f32d315

Observation 8f71313a-8b27-43e3-8cfa-672f2f284fc4 · outbound

This paper cites an unresolved cited work.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T00:48:51.039113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.690266Z digest=sha256:ed58a99fcaf93318e28811a67d42922848902e4e20549bc5772fcfb871e9cbd0

Observation e3c2cb12-e5d4-4d55-8a36-ee36640ae800 · outbound

This paper cites Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-05T00:48:50.091776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T00:48:49.750627Z digest=sha256:d0c54d2656511857ee6b1a1fb8735b1ef6ba4e0e3b087a9ee1b3b5e7b60be5c2

Observation 75d7b4ee-09fa-432f-8c66-4c1a3dd6297d · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.845907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.845907Z digest=sha256:9ce75f80ec1a73947052e37f9fee7dbf1f9b3ec4ca6cc450be7dc932fbae5600

Pith citing papers

No inbound Pith citation observations are available.