Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:57:59.499913Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.23015.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T03:57:59.499913Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 796cccb9-c854-4cd8-bcf7-b97beb564428 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks {TwinBreak}: Jailbreaking {LLM}security alignments based on twin prompts,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0542ac60-b3ad-4695-9223-02353c38621f · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Neurostrike: Neuron-level attacks on aligned llms,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ed6338-5e31-4e91-b44b-98cfaf8cb0c8 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Safeneuron: Neuron-level safety alignment for large language models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f10810c-19e5-4cbf-be1b-fcaf1ee41741 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Training language models to follow instructions with human feedback,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 763ab5f8-1725-4eb3-9057-b09882cd4d27 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed44964b-fd73-48ec-9382-ab4d85dfd91a · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Direct preference optimization: Your language model is secretly a reward model,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2aad0b3-f51e-4ade-82c0-3d31167e7837 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d757a4c6-6f55-4f11-a82f-ecdbd1b29937 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e74c8c1-5d5b-4f6b-9a27-599da978a466 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Autodan: Generating stealthy jailbreak prompts on aligned large language models,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001c2a94-53fd-4dac-8ff8-bae64a4ebc47 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Fine-tuning aligned language models compromises safety, even when users do not intend to!
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aae0de6-c913-4bdc-bca5-5a1b869515cd · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Shortgpt: Layers in large language models are more redundant than you expect,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712b6511-5db3-410d-8381-a8908ea86194 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Blockpruner: Fine-grained pruning for large language models,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d755635-e9c2-414d-94b3-92b174d6c9ea · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Finding safety neurons in large language models,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ba6b8d-cf8f-4b51-bde4-6bf2cda23d72 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e420db-5126-429d-bbdc-2a73b18457bd · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d68b254-17c6-462e-8a8a-36d6ab56ad30 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Fine- grained safety neurons with training-free continual projection to reduce llm fine tuning risks,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab34325-0e0d-46ee-8b72-53f9ccbb073a · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Refusal in Language Models Is Mediated by a Single Direction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cc20ce3-f571-4e38-baf4-da20d7705a6c · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d682ec-adb4-400a-8506-dad91974dee9 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Homotopic language reorganization in the right hemisphere after early left hemisphere injury,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f09300d-11b9-4e25-b073-12ccf9f44107 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Speak- ing with a single cerebral hemisphere: fmri language organization after hemispherectomy in childhood,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2a5483-bada-415e-b776-1842bf376991 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c18103-c801-421f-907d-cd234a279709 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks The Llama 3 Herd of Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75ede269-a074-48c7-ac30-923e19625928 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c869549f-6e94-440a-98ca-4b640c0689ff · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Gemma: Open Models Based on Gemini Research and Technology
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9530001-e977-441a-8566-efbdeee74f12 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Phi-4 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd8d583-5c33-46ba-bdc9-e3042292e405 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks A strongreject for empty jailbreaks,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc742037-5b64-4fba-8eab-13828deb9e35 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Training Verifiers to Solve Math Word Problems
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92abd44-59b5-4ed1-8ab5-9f98cb132053 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97fa6519-e9d4-45ee-be84-5a35048870f5 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Jailbreakbench: An open robustness benchmark for jailbreaking large language models,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29940718-1331-43fe-b547-cc1e9e3f0ff3 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efebfdfd-f71f-4ec9-9b91-c0c93cb989e4 · outbound
Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Truthfulqa: Measuring how models mimic human falsehoods,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.