Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T00:48:49.845907Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.02665.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T00:48:49.845907Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 73d5e304-eaf1-4e55-adbf-64fc2f51e490 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60ade032-b860-4a8c-bffc-634a86ebd9d8 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Pappas, Florian Tramer, Hamed Hassani, and Eric Wong
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 58f0fcc2-0fb7-4eac-9743-37fcf82efc48 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dc7053d-6fff-4f28-a9c3-00cb61985a9a · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fe14dc20-04de-4e94-9e16-600890e0727c · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Best-of-N Jailbreaking
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a825f54-2d85-4f97-bced-821b5569187d · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Jacobs and Hanna Wallach
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 784b503d-3cbf-4283-86cc-58e181a40db7 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1131312-66ab-401b-9424-f9a97a5eef6b · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 46e4f696-0677-40f2-bc22-4e2c92e88b0d · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bf59ab63-6167-4b78-957b-e628b03c71b0 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ba4de435-e0d6-4f05-b7c2-f8a12de7a683 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c39f7f01-617c-4c84-ad77-4574af70a24c · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ddfb7a64-a005-4b07-ae67-0b2cca66b80b · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5c0b3f9e-5023-42cf-bfe0-6d65e5718c24 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cebaace7-c0f9-4320-b9e0-41c4c5bacbfc · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71d8efd2-22c5-40c3-b32e-9750cd307da2 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Low-Resource Languages Jailbreak GPT-4
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c79157-601a-4a9c-b628-c8fb11fc3899 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5eb38f5c-d248-4d90-a46d-cb6fd5a2b453 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f71313a-8b27-43e3-8cfa-672f2f284fc4 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e3c2cb12-e5d4-4d55-8a36-ee36640ae800 · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 75d7b4ee-09fa-432f-8c66-4c1a3dd6297d · outbound
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.