Pith. sign in

Paper Citation Record · LEDGER

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security

As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2509.06807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06807 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:09:41.587858Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact3
  • verified fuzzy8
  • unresolved58
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9dc5398e-f585-4da2-9d83-c20d7a04e93f · outbound

This paper cites Mogu: A framework for enhancing safety of llms while preserving their usability,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Mogu: A framework for enhancing safety of llms while preserving their usability,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.077038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.412785Z digest=sha256:ec9093951ac6deb242ac1459de1e706f9d0c6b7f6894db48ab3e43b0f258db94

Observation 4bb66a28-3120-4825-b8da-f5254d3cd532 · outbound

This paper cites Gpt-4 technical report,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Gpt-4 technical report,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.418878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.418878Z digest=sha256:b6d1ae617cf2daed21ea66d09176065060fe9e51dfd56eb2bd45bb40c724293c

Observation a1558a58-a354-443b-b40a-f503ba477256 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.421041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.421041Z digest=sha256:f6a642b6896eeee9e6f242abee382dd6ea7fe33cac574969f5af6ac46d940eb9

Observation edc29ab5-a9a5-4367-a3c0-04ee523c00e2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.423390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.423390Z digest=sha256:47b9412e989eaa3f1904f583ea63073fb38952e0f306d92b76209ce5b1d8cd4f

Observation 7c0d83d0-d2b6-471f-a4a5-fae1eb0f28cd · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.426641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.426641Z digest=sha256:5b730ea895f5a8b9a993dab1062f23ff090c98f58d6b3958063fb30543c45eea

Observation dc764439-8a44-4e88-95b7-ec1d378ff20f · outbound

This paper cites FLIRT: Feedback Loop In-context Red Teaming.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security FLIRT: Feedback Loop In-context Red Teaming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.429359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.429359Z digest=sha256:ec4bbef19b6c8fce2b4bbb5de7a7ce355d795dcbaa36880f93e06631c1568827

Observation 403519d2-2a8b-4bc2-8ae3-e9693caa9657 · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.432374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.432374Z digest=sha256:f3a2695754775551cbb4528969acd9ef3ecdaef60bf9ec379cb9fa2516a6d92f

Observation ead34801-dcd9-434d-b7cf-10580f06e3e0 · outbound

This paper cites A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.435223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.435223Z digest=sha256:fd60f268891c8f87d589c59a554417e51095b6a4e8902b3d03d893eaab5e2f7f

Observation 12cedbfb-1706-4921-8499-8150a0757aaa · outbound

This paper cites Lima: Less is more for alignment,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lima: Less is more for alignment,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.063715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.437938Z digest=sha256:9d5c59fc170688efe5cd5db5633fce781a1720f38265d845df3b365c52dfb24d

Observation 02db291a-5a5d-4e42-af66-a501dd18c588 · outbound

This paper cites Training language models to follow instructions with human feedback,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training language models to follow instructions with human feedback,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.440287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.440287Z digest=sha256:106de3efd8467566421bb208d7b10d4e00c12808202188716b6a7c38b7d27ea6

Observation fb86b84a-ac4c-4939-9349-a3745d597338 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.442984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.442984Z digest=sha256:a264e79c0d46eacc3c67c77050d1b320e6853a9448d20902e9a6ba629f538e70

Observation 262bc832-9a17-474f-8b41-3324f2f1f943 · outbound

This paper cites A comprehensive study of jailbreak attack versus defense for large language models,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A comprehensive study of jailbreak attack versus defense for large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.050203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.445835Z digest=sha256:63613dbc631058af460d4b25749209b8f8a65f9a592764b0b1d04022f99e83e9

Observation 8f4adb80-c946-4fdf-a9cc-2de908f44f53 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.448210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.448210Z digest=sha256:9bca58dd6242cd5ffc062524164a093a2551ed2c478cc47a3db9c25e22e00ecc

Observation ce4ffc35-5ad8-47fc-81f3-d207b73311b6 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.450991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.450991Z digest=sha256:e2ddce7b68e5f48f4af44a40b4955541dafe3ba3b3c2bc1be19ddef7b1efb581

Observation 6c917c07-f06e-4e09-9d7a-21b822e1d425 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.453737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.453737Z digest=sha256:8c2c0698722d9758cc6e036b3035aba165605d148574596fe1dc66b3684f2aa7

Observation af156b49-2e8e-48a1-b665-2b82b8b86dd3 · outbound

This paper cites Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.456061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.456061Z digest=sha256:44fde91b994054e3802a736b1d7c6e87a840176df3789667555d65d964642003

Observation 3d039875-38b8-4456-ac5f-702912bac0dd · outbound

This paper cites Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.896274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.458494Z digest=sha256:1c52ebcd0685623648186382e59714f661e267069d0bd687a28aa99813a7da38

Observation 12afa9f2-8495-4ea5-a717-ffd7c39d6a48 · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.460875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.460875Z digest=sha256:3ff44310fc0c503f0556c811aa1666c752c52c4d7c500146c95ed01474af8ecc

Observation 0a575274-5666-4764-9913-0df741e67ef7 · outbound

This paper cites Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.463474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.463474Z digest=sha256:432b5a7e015d7d4aaa89afdb33920b64d37c3db3602c03143e3a4ad795e17e68

Observation 98df20ed-9ab8-4277-9bf6-30125ef8485f · outbound

This paper cites The Llama 3 Herd of Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.465759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.465759Z digest=sha256:96e60d36b82fcfc3577056462e96356d1f9086e57ae3747a0870baa0f3bb153b

Observation f161f21a-9365-4361-ab7b-725b7b14fb53 · outbound

This paper cites A holistic approach to undesired content detection in the real world,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A holistic approach to undesired content detection in the real world,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.040961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.468117Z digest=sha256:b088c6fdf8a661e374dec8bd6b08b85600a348e26c3630d322730fd57da141d2

Observation 881ed3e0-8875-48cf-a485-a295d796bf27 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.470662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.470662Z digest=sha256:d51206a613f2f8a5cd90bc5954481aed39006649bbcc9bcd86a6fae185010e10

Observation 1eebe795-ea34-4255-a580-d7a0b4b0d66b · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Certifying LLM Safety against Adversarial Prompting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.473339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.473339Z digest=sha256:76a6484d530cc61d5e6e870865d4ddb7c424c9e81b0d018ee9ce636c58e7fc69

Observation 8b5417fd-5bef-4bc5-8652-0547f8590e12 · outbound

This paper cites Lightweight safety guardrails using fine-tuned bert embeddings,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lightweight safety guardrails using fine-tuned bert embeddings,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.475614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.475614Z digest=sha256:f89b04127b85ec53e48abe1e01ad30bffbe23916cb40b7dd5f3560ebf38c8baa

Observation 0986a88f-38f8-4ea1-9def-a33b45ba276e · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.478401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.478401Z digest=sha256:6574c158d5c447e996e4af67725f1ff574d463bb08dfb1f95b9180a2832bb053

Observation e3c0882b-84ae-4945-a51d-c351a28d9757 · outbound

This paper cites Navigating the OverKill in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Navigating the OverKill in Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.840571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.480864Z digest=sha256:caf269582f29e782b7b5e7635d812b2f3ca81fde74b0f5014aab99ea523ae0f3

Observation 10a3c887-315a-40de-a43f-b04723e1dbe2 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.483523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.483523Z digest=sha256:ac17d31e3fab0457b2a927c261d3e70701508b0e88387c77339e7d431c9eac13

Observation 8cfb6400-5c9e-4148-bcae-beb5aafdf18e · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.485922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.485922Z digest=sha256:1e410a484464c025db84e92f4ac5c03e2ef597cfd3820608f39df9e361f564e6

Observation 3b83ca1f-1e85-426b-b604-e817fe22386e · outbound

This paper cites Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.488370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.488370Z digest=sha256:6377274908a27a0f07670088d4f45a2e5cacea9aba8a044181afd8c898d30441

Observation ef3557ae-8039-45ce-8b44-83a2979f2965 · outbound

This paper cites Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.491116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.491116Z digest=sha256:2f0a25a5c67ecf3b7c9e936c3e962a0aef469456d5da8334aa59b4edd7b3683d

Observation d0abc586-af39-432a-8053-36a2ccc3571d · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Survey on Mixture of Experts in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.493535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.493535Z digest=sha256:683add10549d1e56ba2b61fd7303640b0eb59a9ccca6d154ea2689390920470f

Observation 4632282d-a94f-4fc2-b241-76f34871da81 · outbound

This paper cites Uni-moe: Scaling unified multimodal llms with mixture of experts,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Uni-moe: Scaling unified multimodal llms with mixture of experts,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.495912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.495912Z digest=sha256:06f5603a20b846c74c1e873f2844d2337e95e40093f8dacb3f272dd3a0cc4712

Observation 8371582d-84ec-4ec6-b30c-437d0b14014e · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.498475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.498475Z digest=sha256:e1e5bc2b2963c742505c78c8026b85e27844bedddc05909db464f3e6e5bd1ebb

Observation a6e098f9-c354-479b-b8f2-9964ef0d3085 · outbound

This paper cites How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.500858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.500858Z digest=sha256:89044c5aa8461aaf55ee9740b21f70adcfbcc1c113662c8024fcae0d64ac1d22

Observation d9becc34-a79c-413b-b080-229bd19f3777 · outbound

This paper cites Red Teaming Language Models with Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models with Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.503307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.503307Z digest=sha256:a2b58d0997d0bb4ade7a4ba1f69d24d365308928f7ebc3ddc33b54d3721b5f10

Observation 426a32ea-c602-491d-9fc6-da5b8952a489 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.505782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.505782Z digest=sha256:36292e8538ff279e9bf90da52eb8adc8ffdec4b3aff52a338f744f599475a411

Observation 5f56caf1-de30-425d-a0ec-e7d7fabde360 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.508015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.508015Z digest=sha256:a955ec0c83d76266ef6a2dc8f13c0ce94b9341a0ff8c89dc49033fe310437a61

Observation 9fe51abf-e933-4e20-802a-28412b67252e · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.510729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.510729Z digest=sha256:7c09d7a86f2ba95053e7cecb230e7e987329b4f281e63e614933702d446be4d5

Observation 64bd5f15-fa7e-4f0d-b2ed-e26b626adbd3 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.513177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.513177Z digest=sha256:ef1e4ee0e9e35b7307bc5630a54a52448a0dc5ae26270f167367794872c566fb

Observation c6ba4a17-9d7e-4a19-8b37-598ca4bf46c7 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.515431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.515431Z digest=sha256:7d60e787392d49b1d60f4b8bf0c2a408c3044f52e8affd6c4f87c9de43e81526

Observation fef8e249-e825-45c4-9063-a5adb4df700e · outbound

This paper cites Jailbroken: How does llm safety training fail?.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbroken: How does llm safety training fail?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.517713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.517713Z digest=sha256:dd08de9f1c3f37cb0854f84d0f15681878c1a7d962d23dbf71cbeb25057da4a5

Observation fe644a2a-601f-4f18-b136-82dca3400e09 · outbound

This paper cites Automatically auditing large language models via discrete optimization,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Automatically auditing large language models via discrete optimization,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.023516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.520081Z digest=sha256:ad56a46fcbbc2d7b50b07a62f347b6f3ac44681f0fa69babe1d2339f373895e3

Observation a6efadde-989c-4147-8d82-5b3e79debe8b · outbound

This paper cites Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.522170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.522170Z digest=sha256:e84fc2e60a953eae327bc19b53bea2400825a712a3138b3521222a4b3ad0fb20

Observation efbff674-42ba-4754-b406-ece032bb95c6 · outbound

This paper cites Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.524212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.524212Z digest=sha256:98e0d973bde9221e6ed22ecc31999d1b546cd10acc7b913729902cedeb00e255

Observation 8277e1ce-0848-4875-b947-5de95d990291 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.526476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.526476Z digest=sha256:24fd5ee0757881bdf90aceaeeeeaaad4603aee8328f5da91f54e0dee7b491246

Observation 7fed6c88-2376-428a-bba4-3f70b0047335 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.528615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.528615Z digest=sha256:7d2538dbf4faca5c8e0f44cc20fe3e7e6beadb8f4cb29596ba26cad4bcae6237

Observation 60840032-8432-4455-8d31-8136cb96982e · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.530996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.530996Z digest=sha256:30b711e6c794b3bb8359b42c8d5e1981e50d93bf0493b7c3f23d9ebd58edffda

Observation 0746e8d6-ca2d-490b-bf08-d3c599108cad · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.533142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.533142Z digest=sha256:6de0499a436b425433c1a3542c101f715d57d0fb8378215e0faf13a22e4feb49

Observation 1a343dbb-5ea4-4ee0-8f64-0f9446a2b4ad · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminder,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Defending chatgpt against jailbreak attack via self-reminder,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.535213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.535213Z digest=sha256:1b4f205cf7ecdcddf863e6d9ef9e945ea46ce56f14e5173c9b1e391d5f3b72f0

Observation 6b551ba0-8dcc-4009-9481-9a872dae0065 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.537604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.537604Z digest=sha256:0685e4288e609a53dc7bbf07abf9d23e71bb72d5098817b8459c85e1fc428197

Observation 41b7dcb4-1a29-42f4-825f-7e6bcf4fb978 · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.540879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.540879Z digest=sha256:0a15485cbc86df9a6c4ee3e04a59817d8ae976d00b86c8289e59b6c35564387d

Observation 27d14d86-01b1-421c-a413-6a07a1aa0db2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA: Low-Rank Adaptation of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.543359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.543359Z digest=sha256:127fc7d181586be665cc67bb02fb641fb19fd315faf3b793474bd4c21a9b601a

Observation 16c5c855-997c-46f1-a7ea-85d38f328d70 · outbound

This paper cites Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.545861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.545861Z digest=sha256:49b85780aca97f7cff141eeddc39d6aed480355ce55e961148fce486bb544552

Observation 10d452d0-ad48-457a-9e2f-2a8516dd405b · outbound

This paper cites A Closer Look into Mixture-of-Experts in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Closer Look into Mixture-of-Experts in Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.548125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.548125Z digest=sha256:6990efce1b3006d5112342ab4f2e3ed0f41665132084cecc09e9eb99935b35ac

Observation 5381ff8e-c8a9-4738-95c8-afc06896dcd8 · outbound

This paper cites Adaptive Attention Span in Transformers.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Adaptive Attention Span in Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.550995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.550995Z digest=sha256:f0583899ce19dec1fdb4f1e7082dc3cf1db30eb2bc34bfcab90d09f0041e0697

Observation 3e54ab47-115d-4f55-8f27-43baeb65ec61 · outbound

This paper cites Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.679537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.554098Z digest=sha256:3fa773bb77ea750a296b3849bad38be5290857fdf97b63f487c15749feed6fad

Observation 236bfd5a-4043-4a69-9b94-c05979f94c3b · outbound

This paper cites The hidden risks of large reasoning models: A safety assessment of r1,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The hidden risks of large reasoning models: A safety assessment of r1,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.556460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.556460Z digest=sha256:7729176b2398556fb154e9c9c1a16370c22c5da4c99a0005ecf4539ec22ad268

Observation 3817d6ae-1a90-4122-ad7d-1712b98f17b0 · outbound

This paper cites Falcon-40B: an open large language model with state-of-the-art performance,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Falcon-40B: an open large language model with state-of-the-art performance,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.004374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.558613Z digest=sha256:7cd4c4bb53fe53c7b6731750f0e1243f8c09be915379ff81ea33247e7ca10c73

Observation 1e150241-98f4-4055-aadf-150db08c3e80 · outbound

This paper cites Qwen2 technical report,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2 technical report,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:41.995121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.560941Z digest=sha256:f3d91c4b6a812efc985d7230db456bafb2610d527189ff3ce674a94c829cfb4d

Observation 72b6ea39-32e3-4be0-9da5-bd150d90686a · outbound

This paper cites Qwen2.5: A party of foundation models,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2.5: A party of foundation models,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.563287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.563287Z digest=sha256:8c1d9d90b266b4933049d7749abe98b2f8e94042493af451b45ef486fd7b1199

Observation f203041f-f3e2-43da-91ea-3b60b6da5816 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.565548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.565548Z digest=sha256:f647371b8616af9071f99eab30037e77481d50871eb036ec5c0113624b607563

Observation ccb3ab6d-60ab-4875-8b4e-ea684f463a66 · outbound

This paper cites DeepSeek-V3 Technical Report.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security DeepSeek-V3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.568088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.568088Z digest=sha256:b82b19c3dbd553f558a3b4858d7794e92d9085b1cc54e02ffb88c086a9e539a6

Observation f663f21d-6fa0-4780-abeb-cc0ad5f7768c · outbound

This paper cites The unlocking spell on base llms: Rethinking alignment via in-context learning,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The unlocking spell on base llms: Rethinking alignment via in-context learning,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:41.982843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-04T23:09:41.570710Z digest=sha256:95a26b0209254bcda1377dc3cf979b1d4654620696c4ed45dbf6e4e2bb82cbd0

Observation f4736afa-cf44-4f03-8fef-46af2cc792f0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training Verifiers to Solve Math Word Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.573265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.573265Z digest=sha256:1949a8f8977213019a4080ba9f9d16da9e55f332d448108e7461b84b418793aa

Observation 89dce95f-0d2c-4792-9e61-17f46f8792bb · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.575944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.575944Z digest=sha256:cf80600361a89a825b7d3fd428a1895536c0932dc41ebe599aa729bc4d3ee813

Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.578345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.578345Z digest=sha256:7103a758e7fec668860b5abfb97379e52849200997cab3b2cc6a6e8a3ca7e73a

Observation 4af7982f-a890-4203-9d1f-94fefb2fb084 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.580806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.580806Z digest=sha256:a41dcd34406aa76da754739252d19565d0668126beef99db37831873350a28bb

Observation fc050fd4-0ae3-4005-a987-602c34e7b2ed · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language models are super mario: Absorbing abilities from homologous models as a free lunch,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.583242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.583242Z digest=sha256:4bb1e95c822dd79cebb2f1ea5459d937f884783a7692b812fbf12a9b538bb6c2

Observation 9694cb0d-1a26-4cb1-8a28-613fb8689937 · outbound

This paper cites The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.585521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.585521Z digest=sha256:d5edb494d4d11102393b131e358ca3e583876f897562ada34a81365b58fa27c9

Observation b591270b-0b7d-4011-845c-880f0d93f679 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-08-04T23:09:41.587858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.587858Z digest=sha256:ddb51ce2efa4cf012db19ad5e9ff59c7b3a696ef68611264ccc95738b07fe28d

Pith citing papers

No inbound Pith citation observations are available.