Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T16:06:06.687929Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.24392.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T16:06:06.687929Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b310d6db-cb48-4ee2-bb07-75581854612d · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0f0e8e9-48cd-486a-8129-9dd083cabf0f · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b509488a-218b-4b99-83aa-3e9b2ac83105 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41706c2-33b6-48e5-86df-aece5f6d7ede · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Detecting Language Model Attacks with Perplexity
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeeedf14-617a-4aed-8c71-cef91edcac68 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 253a77cf-dfb4-4256-b0df-ebfab08a46fc · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75a6e83e-7179-488d-8857-cf1a5ded795f · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15352712-e2f2-4e54-9f8d-cb9f642e26d9 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f055f78-ed1d-47af-9e52-cff5c7013f2b · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a8a4ce-aa82-44e0-91ed-98d7ce5e593c · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f69015bc-e1bd-4c3f-a985-fdb2956ef072 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs The robots are coming: Exploring the implications of openai codex on introductory program- ming
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5872d1-564f-43b3-995c-f00c0877dac7 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dab7cd4-8336-406d-9b0c-f54cbe7c67a6 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c447e52-6f8a-43f2-878c-c4f64c0f9dea · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0accd163-a524-4631-a378-50070b71ae7e · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1887b75-0793-4ee5-b58c-1ef94547d7d0 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Salad-bench: A hierarchical and comprehensive safety benchmark for large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 195d7399-377a-4024-b7bc-786eb3d8fdc8 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Counting function estimates for coherent frames and Riesz sequences
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1486b3ba-5640-47a6-b238-8ae4f0b015bf · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Llm self defense: By self examination, llms know they are being tricked,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b33c1953-7071-425f-b150-2142b1e69023 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0789fb59-fc4e-4340-8c58-463e14c888d5 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98057fa9-a562-4330-af10-eeafa3d03830 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Trustllm: Trustworthiness in large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65dc2ad8-1853-470a-88f6-8d4714f294c0 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs XSTest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9639ddd-ddfc-4704-bd7c-72e117b06068 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Gemma 2: Improving Open Language Models at a Practical Size
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8cb5949-4aea-4280-9c85-af3750d94922 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs LLM4Decompile: Decompiling Binary Code with Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b46bff-5621-4fd2-83c0-15d8f16652c3 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Defending LLMs against jailbreaking attacks via backtranslation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d998395c-8f7a-48c4-b26f-c5c87662bce9 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs The art of defending: A systematic evaluation and analysis of LLM defense strategies on safety and over-defensiveness
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a179e9a-bb3b-445a-b683-f7a1f9e23786 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967b658a-0064-47d9-b9c8-f210621a010e · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 015aee69-d2e8-49ad-ba73-34c93b04abc1 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Defending chatgpt against jailbreak attack via self-reminders.Nature Machine Intelligence, 5(12):1486–1496, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df3be41-3ad5-47e0-9b08-a214089f302f · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd3b31c-d92f-45fb-89ce-84553e2d4239 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs SafeDecoding: Defending against jailbreak attacks via safety-aware decoding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c08e564-bb40-4baa-ada9-8baa55f1d190 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Contrastive preference optimization: Pushing the bound- aries of LLM performance in machine translation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a06c57-1f00-44ac-ad24-2746479fd2ae · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Qwen2 Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c66ab8a0-923c-4170-922f-8fe067ecd526 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs A comprehensive study of jailbreak attack versus defense for large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2b5ed2-554b-4b60-913f-271f13795738 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Intention Analysis Makes LLMs A Good Jailbreak Defender
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1add4be-80e4-4697-adb3-8a270c806bb6 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Jailbreak open-sourced large language models via enforced decoding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43f900a8-bad3-4961-a189-78e99f0bcfd0 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Instruction-Following Evaluation for Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5465ff8f-bbda-494c-8e48-909797b2a1fe · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Defending large language models against jailbreaking attacks through goal prioritization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e375fec4-ae95-4a0f-9514-9a211127c156 · outbound
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs The Llama 3 Herd of Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.