Pith. sign in

Paper Citation Record · LEDGER

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security

As of 14 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2509.06807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06807 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:09:41.587858Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact3
  • verified fuzzy8
  • unresolved58
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9dc5398e-f585-4da2-9d83-c20d7a04e93f · outbound

This paper cites Mogu: A framework for enhancing safety of llms while preserving their usability,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Mogu: A framework for enhancing safety of llms while preserving their usability,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.077038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.412785Z digest=sha256:8db2ae1a77af754551a81484e6d73b150c835841eba2512d013ae0e6e625eb60

Observation 4bb66a28-3120-4825-b8da-f5254d3cd532 · outbound

This paper cites Gpt-4 technical report,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Gpt-4 technical report,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.418878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.418878Z digest=sha256:87ace5b7aa6388a7d3d72c9b5ddeda60cf2fbd37f15413ee8bc8bd4e98d323fc

Observation a1558a58-a354-443b-b40a-f503ba477256 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.421041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.421041Z digest=sha256:83a3eb2f3766239dbc64c5e21442c446e322fd06917372c7ac419ec6778a4f45

Observation edc29ab5-a9a5-4367-a3c0-04ee523c00e2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.423390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.423390Z digest=sha256:7b568fa9167ae982878ce8760e661c9410b633a1c371d98f837880bd34faf6ee

Observation 7c0d83d0-d2b6-471f-a4a5-fae1eb0f28cd · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.426641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.426641Z digest=sha256:0f2cff840e884484b6e5fed0da341b7e4a63ae1142ba65e375a8e3febffd0ec2

Observation dc764439-8a44-4e88-95b7-ec1d378ff20f · outbound

This paper cites FLIRT: Feedback Loop In-context Red Teaming.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security FLIRT: Feedback Loop In-context Red Teaming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.429359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.429359Z digest=sha256:e23a86c7a06e99ca8621d9d5c4c456c7a0e4bad4f3f0840c95a51daf31fbf227

Observation 403519d2-2a8b-4bc2-8ae3-e9693caa9657 · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.432374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.432374Z digest=sha256:aa438331feb20a69faaf84d62360a5ae2bfd22e4b9e4e4659bfcdf1edf6451d7

Observation ead34801-dcd9-434d-b7cf-10580f06e3e0 · outbound

This paper cites A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.435223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.435223Z digest=sha256:ba4a09523dd6716b15d31ab154f6137729353169f43748a7fed32fd01a59e586

Observation 12cedbfb-1706-4921-8499-8150a0757aaa · outbound

This paper cites Lima: Less is more for alignment,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lima: Less is more for alignment,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.063715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.437938Z digest=sha256:714f09ae0a7a03fb10c8cbfcfd178a22e53a7b7e6d7fb054a8603301ededb78f

Observation 02db291a-5a5d-4e42-af66-a501dd18c588 · outbound

This paper cites Training language models to follow instructions with human feedback,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training language models to follow instructions with human feedback,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.440287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.440287Z digest=sha256:1c60b2931c4716c6a8f5249ed7f7d1e57d0b87a60fbc32232865f887586ec625

Observation fb86b84a-ac4c-4939-9349-a3745d597338 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.442984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.442984Z digest=sha256:3231bb8ab29f62057368c7f4da0cdec786a34a980da0fec3d3d6db1b42d6156c

Observation 262bc832-9a17-474f-8b41-3324f2f1f943 · outbound

This paper cites A comprehensive study of jailbreak attack versus defense for large language models,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A comprehensive study of jailbreak attack versus defense for large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.050203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.445835Z digest=sha256:48bf1efeff6ba0ed33b5d97d6db8569a240fa2bdf4669443583c5d59f3e63265

Observation 8f4adb80-c946-4fdf-a9cc-2de908f44f53 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.448210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.448210Z digest=sha256:229d47bd466a962f98f6489175118912272c46e09e10ef123c9ae9a5e59048b6

Observation ce4ffc35-5ad8-47fc-81f3-d207b73311b6 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.450991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.450991Z digest=sha256:18ff3f2cf975f6211c90eb7f5d0bf580ba826cbefdb614c96c84f3313049da2e

Observation 6c917c07-f06e-4e09-9d7a-21b822e1d425 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.453737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.453737Z digest=sha256:0f228e49742893458eb0bd630bac57b3cdd1d2f2063744296b81afa6ba480485

Observation af156b49-2e8e-48a1-b665-2b82b8b86dd3 · outbound

This paper cites Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.456061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.456061Z digest=sha256:f29484ceee2e03b82206bb346264de20c825c15c65857f9d91218529d33ea4ef

Observation 3d039875-38b8-4456-ac5f-702912bac0dd · outbound

This paper cites Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.896274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.458494Z digest=sha256:6771c69c294c6449a1acacfc21ee22db4c667489347116717a857288c6fcb9b4

Observation 12afa9f2-8495-4ea5-a717-ffd7c39d6a48 · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.460875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.460875Z digest=sha256:54fa031b387c489d0599a7579ce674b40c6916b91602d52971647c5814f0782a

Observation 0a575274-5666-4764-9913-0df741e67ef7 · outbound

This paper cites Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.463474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.463474Z digest=sha256:37f4c4f68b4af68af849a017148e5079fe10f788053e6101c0e7e8f7ad836ada

Observation 98df20ed-9ab8-4277-9bf6-30125ef8485f · outbound

This paper cites The Llama 3 Herd of Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.465759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.465759Z digest=sha256:d47ced01bfed7a8b056d8160434c423f01357b56fcc61822d0d63de558394c08

Observation f161f21a-9365-4361-ab7b-725b7b14fb53 · outbound

This paper cites A holistic approach to undesired content detection in the real world,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A holistic approach to undesired content detection in the real world,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.040961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.468117Z digest=sha256:c48812543df6987bf2d9f2505f19ce9d2dd01b7da3f600ff2a72aa878a1b2aba

Observation 881ed3e0-8875-48cf-a485-a295d796bf27 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.470662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.470662Z digest=sha256:58750a023c8c544358c742b4ce7c20917ea8162e08a279748eb8412ab3ebe48e

Observation 1eebe795-ea34-4255-a580-d7a0b4b0d66b · outbound

This paper cites Certifying LLM Safety against Adversarial Prompting.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Certifying LLM Safety against Adversarial Prompting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.473339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.473339Z digest=sha256:3dcc1b1a04109bb37120d5d6a25dd544bef023a2edc17862eb59048987d5c3d8

Observation 8b5417fd-5bef-4bc5-8652-0547f8590e12 · outbound

This paper cites Lightweight safety guardrails using fine-tuned bert embeddings,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lightweight safety guardrails using fine-tuned bert embeddings,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.475614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.475614Z digest=sha256:e662aa5db417078fe9c89a252cf5fe02470efd125db1a4d92506bbc7079efa59

Observation 0986a88f-38f8-4ea1-9def-a33b45ba276e · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.478401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.478401Z digest=sha256:ea2b7e0206ad6bb12145359f199a329f4583b348c8aadb1b357f75dcac2b8e74

Observation e3c0882b-84ae-4945-a51d-c351a28d9757 · outbound

This paper cites Navigating the OverKill in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Navigating the OverKill in Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.840571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.480864Z digest=sha256:d67b5b36167798765516e7180712a93f012c3622c1af3dc3d7db7b3a357c0404

Observation 10a3c887-315a-40de-a43f-b04723e1dbe2 · outbound

This paper cites Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.483523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.483523Z digest=sha256:ddff68eb3d74fac7ebed52076569eb0fead63d5bb00492c0c103203830484482

Observation 8cfb6400-5c9e-4148-bcae-beb5aafdf18e · outbound

This paper cites Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.485922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.485922Z digest=sha256:062b463f99eb906583e56424be8f659f732d347889e58860fb6bf1c679afd15d

Observation 3b83ca1f-1e85-426b-b604-e817fe22386e · outbound

This paper cites Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.488370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.488370Z digest=sha256:4b7fa042da2e1743b89ebe1fee06baf0253ead4c636bceeff099701f7dd8b625

Observation ef3557ae-8039-45ce-8b44-83a2979f2965 · outbound

This paper cites Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.491116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.491116Z digest=sha256:290376bbc27f94850f316c535f55a23b22ef84a0ca9fb3a19834228751ba4f56

Observation d0abc586-af39-432a-8053-36a2ccc3571d · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Survey on Mixture of Experts in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.493535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.493535Z digest=sha256:ce3808806a4b7acdc330335c81baa1c4edcd06fcae3fe94c6123e798f0c012d7

Observation 4632282d-a94f-4fc2-b241-76f34871da81 · outbound

This paper cites Uni-moe: Scaling unified multimodal llms with mixture of experts,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Uni-moe: Scaling unified multimodal llms with mixture of experts,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.495912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.495912Z digest=sha256:a2d5f7be759b090e9e4cec03e3a0c0e23f081186f4bada5e11d182b2a7a74c17

Observation 8371582d-84ec-4ec6-b30c-437d0b14014e · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.498475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.498475Z digest=sha256:5fb214adc0eb090ea5d87feeb50b8addc521b9bac6b2492aa2afe735c9650894

Observation a6e098f9-c354-479b-b8f2-9964ef0d3085 · outbound

This paper cites How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.500858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.500858Z digest=sha256:d7d2e4715cf9f83518d8777d375c99475ffab96370a17d4903c78d1b17adefb9

Observation d9becc34-a79c-413b-b080-229bd19f3777 · outbound

This paper cites Red Teaming Language Models with Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models with Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.503307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.503307Z digest=sha256:72d7edf82e4cedfde6d2934eaa6d7a8c40c71be3f6776f325674416cec4bea46

Observation 426a32ea-c602-491d-9fc6-da5b8952a489 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.505782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.505782Z digest=sha256:15a35d151a483b1a0195dc8b7ec4a165a5cdd0de72790c334774e887082cee46

Observation 5f56caf1-de30-425d-a0ec-e7d7fabde360 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.508015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.508015Z digest=sha256:5ae6e144dd334d8f368d3138f5d8b7afbc53bb647bca7e9c2f5ed4ffa321c0c8

Observation 9fe51abf-e933-4e20-802a-28412b67252e · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.510729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.510729Z digest=sha256:f1dfc70c71140cf1ee863e318ed3150202043736aa2e8cab2b900f33971f1014

Observation 64bd5f15-fa7e-4f0d-b2ed-e26b626adbd3 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.513177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.513177Z digest=sha256:7e27dd9fef9d99b5872e069c8a53eb447c0883fb8555d83157f169f091800f20

Observation c6ba4a17-9d7e-4a19-8b37-598ca4bf46c7 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.515431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.515431Z digest=sha256:32e7e71b8a69b2dcaa2309b042845ade83250789ae98e803919eb53a863a7a9a

Observation fef8e249-e825-45c4-9063-a5adb4df700e · outbound

This paper cites Jailbroken: How does llm safety training fail?.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbroken: How does llm safety training fail?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.517713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.517713Z digest=sha256:658ecfe5c77ed34d9fb4a0e88205897e8101be7f26abff0236039f7cd7525518

Observation fe644a2a-601f-4f18-b136-82dca3400e09 · outbound

This paper cites Automatically auditing large language models via discrete optimization,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Automatically auditing large language models via discrete optimization,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.023516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.520081Z digest=sha256:1524a80226e4a8f0af68c858dcb5650cf4fa8c7dc59121bd446a2d23ada022ff

Observation a6efadde-989c-4147-8d82-5b3e79debe8b · outbound

This paper cites Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.522170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.522170Z digest=sha256:0b33231e36c7865356f286569e7c6c3c2fae576889c2314b6de6da66d92f5c83

Observation efbff674-42ba-4754-b406-ece032bb95c6 · outbound

This paper cites Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.524212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.524212Z digest=sha256:c3b496e484802fbb1cf52d877faf9075267bbe22b1259697b93391810c04cf51

Observation 8277e1ce-0848-4875-b947-5de95d990291 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.526476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.526476Z digest=sha256:cae8a482f594eeb40ff405046d0e7d4a55c3daf0ec3d85d55258cdfe3fea057b

Observation 7fed6c88-2376-428a-bba4-3f70b0047335 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.528615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.528615Z digest=sha256:51b8cb2e563f2fdf8effe9b0ed7e17ec330b07042326c910a86850d8a66272bf

Observation 60840032-8432-4455-8d31-8136cb96982e · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.530996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.530996Z digest=sha256:6a76ae6a5f88d329d0cf2fc7c50c06a976130f3f274405c2bd8b2e0b3c4778ed

Observation 0746e8d6-ca2d-490b-bf08-d3c599108cad · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.533142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.533142Z digest=sha256:d07ce9b0b86574c96fdb8a9637f55c56e4b7e0d2f6c32ba719db340fa02f95e5

Observation 1a343dbb-5ea4-4ee0-8f64-0f9446a2b4ad · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminder,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Defending chatgpt against jailbreak attack via self-reminder,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.535213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.535213Z digest=sha256:2f06e3a1f1f28f603c47faa9c293f93b6d7cde7e34688fb8a021dd52a58307ce

Observation 6b551ba0-8dcc-4009-9481-9a872dae0065 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.537604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.537604Z digest=sha256:62f15443261d3f39bd294c4e3193c38fddd9a83876354e3a35cc2a03d2f9850d

Observation 41b7dcb4-1a29-42f4-825f-7e6bcf4fb978 · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.540879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.540879Z digest=sha256:577381f0b0a154b4668b13f8ddb349e1426dc1c3caea4f2237541da10ea5c3ec

Observation 27d14d86-01b1-421c-a413-6a07a1aa0db2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA: Low-Rank Adaptation of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.543359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.543359Z digest=sha256:68dbfbfbcef73e012cf2dd0701c57527e9bfcf60dc5c7c80ca23e319fae9fb50

Observation 16c5c855-997c-46f1-a7ea-85d38f328d70 · outbound

This paper cites Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.545861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.545861Z digest=sha256:4788b14e6dabdba41b0e70bcd69da649a501dbbf0405323e586d8096a778516e

Observation 10d452d0-ad48-457a-9e2f-2a8516dd405b · outbound

This paper cites A Closer Look into Mixture-of-Experts in Large Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Closer Look into Mixture-of-Experts in Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.548125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.548125Z digest=sha256:c16c56c53922e9e7add4e4dfab35b52999a33aaccf35f3ba4311513ffa0de869

Observation 5381ff8e-c8a9-4738-95c8-afc06896dcd8 · outbound

This paper cites Adaptive Attention Span in Transformers.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Adaptive Attention Span in Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.550995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.550995Z digest=sha256:a0f4ac28bd86b90e92ca619a499ab626c52f5a8d1642cbba7094ca470723960e

Observation 3e54ab47-115d-4f55-8f27-43baeb65ec61 · outbound

This paper cites Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-04T23:09:41.679537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.554098Z digest=sha256:f93cd103a57020d17a10b2f5bd394a28016bd8c0906ccd247de94cd6bbca7fa9

Observation 236bfd5a-4043-4a69-9b94-c05979f94c3b · outbound

This paper cites The hidden risks of large reasoning models: A safety assessment of r1,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The hidden risks of large reasoning models: A safety assessment of r1,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.556460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.556460Z digest=sha256:e8835c52789d1d3e0d59c29912570cac682177ae836c2bc972a6f59bda892a3e

Observation 3817d6ae-1a90-4122-ad7d-1712b98f17b0 · outbound

This paper cites Falcon-40B: an open large language model with state-of-the-art performance,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Falcon-40B: an open large language model with state-of-the-art performance,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:42.004374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.558613Z digest=sha256:4087a365f706d1be5e4cc55067d326660f224f161c7558f262a1a6fcc958d9f5

Observation 1e150241-98f4-4055-aadf-150db08c3e80 · outbound

This paper cites Qwen2 technical report,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2 technical report,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:41.995121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.560941Z digest=sha256:d098e8b9e17613e65af303e03ed96b97be9818cdf85296d2646a6e7004ed84c7

Observation 72b6ea39-32e3-4be0-9da5-bd150d90686a · outbound

This paper cites Qwen2.5: A party of foundation models,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2.5: A party of foundation models,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.563287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.563287Z digest=sha256:fa73515f786fc22a9271c3de00a04511b231c8c837b4d852f2f78b0434d7c0cc

Observation f203041f-f3e2-43da-91ea-3b60b6da5816 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.565548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.565548Z digest=sha256:666bc0786bdf6cb8b7d6ac876a9a39ffe0f715bf26ddfcfb2d7fecd0b15e1ab5

Observation ccb3ab6d-60ab-4875-8b4e-ea684f463a66 · outbound

This paper cites DeepSeek-V3 Technical Report.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security DeepSeek-V3 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.568088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.568088Z digest=sha256:6d5ee8e624916fa0bc82c6e2ed2c2df7aacb1b4b4295b7ccbf38ca4e4e1a52b6

Observation f663f21d-6fa0-4780-abeb-cc0ad5f7768c · outbound

This paper cites The unlocking spell on base llms: Rethinking alignment via in-context learning,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The unlocking spell on base llms: Rethinking alignment via in-context learning,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:09:41.982843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-04T23:09:41.570710Z digest=sha256:9c6acad3e3733d1fbd18ddeaa3fd7b75aa1339e4a4b395106dc09322cd46be84

Observation f4736afa-cf44-4f03-8fef-46af2cc792f0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training Verifiers to Solve Math Word Problems

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.573265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.573265Z digest=sha256:755f6b566cf4a80a1f8979ef4a84c11a7b7010567256b6b8ee9e0dba6e7c16c9

Observation 89dce95f-0d2c-4792-9e61-17f46f8792bb · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.575944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.575944Z digest=sha256:7a0d7e79674bab6dc4b6d6fd76128ba9a4ddce933358285f51436637595b39b9

Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.578345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.578345Z digest=sha256:fc30e224d2d7bc074fe068b108cbeecea3ac9513682f9cdb4f163b6da8bf53c7

Observation 4af7982f-a890-4203-9d1f-94fefb2fb084 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.580806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.580806Z digest=sha256:b21273f3fc04783107d0e71ddb7905de82d6eee3eb1809c3f1830192eb0f82c3

Observation fc050fd4-0ae3-4005-a987-602c34e7b2ed · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch,.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language models are super mario: Absorbing abilities from homologous models as a free lunch,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.583242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.583242Z digest=sha256:a4d8dc7fbbc81318d93df94239adb381ee6140cce1359f5cf396e39314454b45

Observation 9694cb0d-1a26-4cb1-8a28-613fb8689937 · outbound

This paper cites The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.585521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.585521Z digest=sha256:b286098f0db9464df5d8dde97789df58b0eb32353c448edc7f6364d88f8c6d0e

Observation b591270b-0b7d-4011-845c-880f0d93f679 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-08-04T23:09:41.587858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.587858Z digest=sha256:2dc53b0edc42d31076a1979d3866d99b1ec380120e864d8a7629034a41fc208c

Pith citing papers

No inbound Pith citation observations are available.