Pith. sign in

Paper Citation Record · LEDGER

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks

As of 9 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.23015.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23015 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:57:59.499913Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 796cccb9-c854-4cd8-bcf7-b97beb564428 · outbound

This paper cites {TwinBreak}: Jailbreaking {LLM}security alignments based on twin prompts,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks {TwinBreak}: Jailbreaking {LLM}security alignments based on twin prompts,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:55.174461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:55.174461Z digest=sha256:7c0669403a9176fb35584756f4f2064422948b808b0855d025178f656a7a87bf

Observation 0542ac60-b3ad-4695-9223-02353c38621f · outbound

This paper cites Neurostrike: Neuron-level attacks on aligned llms,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Neurostrike: Neuron-level attacks on aligned llms,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:55.328738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:55.328738Z digest=sha256:ae9852dda3098bd1728e609fe6cdcbbc5c3d975585762775132fb217ad730dcc

Observation 83ed6338-5e31-4e91-b44b-98cfaf8cb0c8 · outbound

This paper cites Safeneuron: Neuron-level safety alignment for large language models,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Safeneuron: Neuron-level safety alignment for large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:55.463492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:55.463492Z digest=sha256:aadf1f5dee5858dd6038ec4eeb46143cefb0126f1fd91a3c26392d1293d130b4

Observation 0f10810c-19e5-4cbf-be1b-fcaf1ee41741 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Training language models to follow instructions with human feedback,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:55.725939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:55.725939Z digest=sha256:3121c8ce561c7b2f0e4f7d71ae1ea955ac138211d38cf96503c7fc9a1749dce4

Observation 763ab5f8-1725-4eb3-9057-b09882cd4d27 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:55.883947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:55.883947Z digest=sha256:c525f9e6c71d9255aa2ae6202b19548af010264768caf923e2ce7fd6121c2dc5

Observation ed44964b-fd73-48ec-9382-ab4d85dfd91a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Direct preference optimization: Your language model is secretly a reward model,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.021659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.021659Z digest=sha256:872644958e57b7aec7a68d43c786316a565ccb3ca6315142cbb91da31ef5ba29

Observation a2aad0b3-f51e-4ade-82c0-3d31167e7837 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.099899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.099899Z digest=sha256:a002d5d30d446e0e18eda9d055b341f30c886fa163897fef8f229e10939cdcb5

Observation d757a4c6-6f55-4f11-a82f-ecdbd1b29937 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.216801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.216801Z digest=sha256:007de681f90fb0c187453fc7fce20591de90643c5957f2b763065f991fd08f56

Observation 1e74c8c1-5d5b-4f6b-9a27-599da978a466 · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Autodan: Generating stealthy jailbreak prompts on aligned large language models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.314487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.314487Z digest=sha256:6e52184977ce55d16daa095973a89af302cba3480ed1d064aa38afa348e1e644

Observation 001c2a94-53fd-4dac-8ff8-bae64a4ebc47 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to!.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Fine-tuning aligned language models compromises safety, even when users do not intend to!

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.376108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.376108Z digest=sha256:4c957cb8b65160acc963228b2f7fee8c8c570113ee73963ffc736926b555f617

Observation 2aae0de6-c913-4bdc-bca5-5a1b869515cd · outbound

This paper cites Shortgpt: Layers in large language models are more redundant than you expect,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Shortgpt: Layers in large language models are more redundant than you expect,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.418135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.418135Z digest=sha256:75e0d0417984ac3fd1edf3b98e5757f3043f7269fee979702905d024510a8af9

Observation 712b6511-5db3-410d-8381-a8908ea86194 · outbound

This paper cites Blockpruner: Fine-grained pruning for large language models,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Blockpruner: Fine-grained pruning for large language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.556177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.556177Z digest=sha256:aeedf703123c233c1dace446d9d99c1af7202642955c95092749cef9942f77f7

Observation 0d755635-e9c2-414d-94b3-92b174d6c9ea · outbound

This paper cites Finding safety neurons in large language models,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Finding safety neurons in large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.718260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.718260Z digest=sha256:5729d4d17bc796c564e5f58e952028edbca2b738abb0115bb9ed6521cb42dd75

Observation a7ba6b8d-cf8f-4b51-bde4-6bf2cda23d72 · outbound

This paper cites NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:56.952079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:56.952079Z digest=sha256:30a5c22e5ae005a0046b2df28afb4e4ffa8b3e08853ef41b848d8623a2909353

Observation c1e420db-5126-429d-bbdc-2a73b18457bd · outbound

This paper cites NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.052252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.052252Z digest=sha256:1be3b199877d584aad78198ad1391f46d568b01bd7773164f04cb5d389f70650

Observation 8d68b254-17c6-462e-8a8a-36d6ab56ad30 · outbound

This paper cites Fine- grained safety neurons with training-free continual projection to reduce llm fine tuning risks,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Fine- grained safety neurons with training-free continual projection to reduce llm fine tuning risks,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.136280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.136280Z digest=sha256:5ec15a959995c61b0a1d30ab3fa435d31ed800caebb0b1f16cfecf6f25c5f0ef

Observation 4ab34325-0e0d-46ee-8b72-53f9ccbb073a · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Refusal in Language Models Is Mediated by a Single Direction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.180833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.180833Z digest=sha256:4efd26c236f8381d9d71a7a5f94bd3daf5399fa84ffbebfd6a60220c29360385

Observation 1cc20ce3-f571-4e38-baf4-da20d7705a6c · outbound

This paper cites Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Beyond surface alignment: Rebuilding llms safety mechanism via probabilistically ablating refusal direction,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.344949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.344949Z digest=sha256:9180cf389068e5417fb45cc5caea011d96c036bc64593efdb37a840959ab42a3

Observation 49d682ec-adb4-400a-8506-dad91974dee9 · outbound

This paper cites Homotopic language reorganization in the right hemisphere after early left hemisphere injury,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Homotopic language reorganization in the right hemisphere after early left hemisphere injury,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.470050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.470050Z digest=sha256:4026194eeb5ebe5bbf4135cbddac8071729eed220ca7b3ebbbc985395c3527fe

Observation 4f09300d-11b9-4e25-b073-12ccf9f44107 · outbound

This paper cites Speak- ing with a single cerebral hemisphere: fmri language organization after hemispherectomy in childhood,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Speak- ing with a single cerebral hemisphere: fmri language organization after hemispherectomy in childhood,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.605873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.605873Z digest=sha256:0fd0f53b1e7193d04fdde9eae1ab053ca2fc6215b7d450ca757aede7a1706475

Observation 1e2a5483-bada-415e-b776-1842bf376991 · outbound

This paper cites Qwen2.5 Technical Report.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.763420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.763420Z digest=sha256:8a8464d786afbc4c03a837281ca5bfbeafb5ef4d1eabf372486f189c43ecea3d

Observation f8c18103-c801-421f-907d-cd234a279709 · outbound

This paper cites The Llama 3 Herd of Models.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:57.999000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:57.999000Z digest=sha256:423c813e7e13c0f17f9ddc86419a77dee80e557dde5c6476336234a1dd41f877

Observation 75ede269-a074-48c7-ac30-923e19625928 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:58.210781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:58.210781Z digest=sha256:5936f2baba03a79c632709d4989ccd7961a0099e4d557a56a88a05d759a10846

Observation c869549f-6e94-440a-98ca-4b640c0689ff · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Gemma: Open Models Based on Gemini Research and Technology

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:58.396900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:58.396900Z digest=sha256:0e85f1820800d85db8dfeba18df623afc45af7ea4fff6549b86376c45948d6cd

Observation a9530001-e977-441a-8566-efbdeee74f12 · outbound

This paper cites Phi-4 Technical Report.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Phi-4 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:58.568949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:58.568949Z digest=sha256:cadd8ec2cb67b9ee39ebeaf2c58eafc0f05829b3bf1131189c673a4734e3f81f

Observation fdd8d583-5c33-46ba-bdc9-e3042292e405 · outbound

This paper cites A strongreject for empty jailbreaks,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks A strongreject for empty jailbreaks,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:58.664584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:58.664584Z digest=sha256:a2030ace42c8c690e2d78833b619144edea0014782bd5b117b8cded0f47b9737

Observation cc742037-5b64-4fba-8eab-13828deb9e35 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Training Verifiers to Solve Math Word Problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:58.789019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:58.789019Z digest=sha256:37f9909841ae1503d225fee89da577fab2d6cdfbb0b3703eb9204c04cea91c22

Observation d92abd44-59b5-4ed1-8ab5-9f98cb132053 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:58.902932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:58.902932Z digest=sha256:94be8cafa0fd82d633c9727966ca0b30dcdc3b439ab2f2c8f9d094f31a278051

Observation 97fa6519-e9d4-45ee-be84-5a35048870f5 · outbound

This paper cites Jailbreakbench: An open robustness benchmark for jailbreaking large language models,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Jailbreakbench: An open robustness benchmark for jailbreaking large language models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:59.143074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:59.143074Z digest=sha256:bd8c2be8f6d7cad08289fae94ffa3e3b1fd579e6caceb847c58cfca96fdfc0ee

Observation 29940718-1331-43fe-b547-cc1e9e3f0ff3 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:59.326603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:59.326603Z digest=sha256:e07b39747a00b7dba3e3092342dce7d4db6d544e334897d72d8c1e05f6585ece

Observation efebfdfd-f71f-4ec9-9b91-c0c93cb989e4 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods,.

Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks Truthfulqa: Measuring how models mimic human falsehoods,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:59.499913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:59.499913Z digest=sha256:5206ad8de29986d832de62574f376485aaa0b3ce3017f46ca00d71ffd17b9520

Pith citing papers

No inbound Pith citation observations are available.