Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:09:41.587858Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2509.06807.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:09:41.587858Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9dc5398e-f585-4da2-9d83-c20d7a04e93f · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Mogu: A framework for enhancing safety of llms while preserving their usability,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4bb66a28-3120-4825-b8da-f5254d3cd532 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Gpt-4 technical report,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1558a58-a354-443b-b40a-f503ba477256 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc29ab5-a9a5-4367-a3c0-04ee523c00e2 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c0d83d0-d2b6-471f-a4a5-fae1eb0f28cd · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc764439-8a44-4e88-95b7-ec1d378ff20f · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security FLIRT: Feedback Loop In-context Red Teaming
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 403519d2-2a8b-4bc2-8ae3-e9693caa9657 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead34801-dcd9-434d-b7cf-10580f06e3e0 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12cedbfb-1706-4921-8499-8150a0757aaa · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lima: Less is more for alignment,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02db291a-5a5d-4e42-af66-a501dd18c588 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training language models to follow instructions with human feedback,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb86b84a-ac4c-4939-9349-a3745d597338 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262bc832-9a17-474f-8b41-3324f2f1f943 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A comprehensive study of jailbreak attack versus defense for large language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f4adb80-c946-4fdf-a9cc-2de908f44f53 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce4ffc35-5ad8-47fc-81f3-d207b73311b6 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c917c07-f06e-4e09-9d7a-21b822e1d425 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af156b49-2e8e-48a1-b665-2b82b8b86dd3 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d039875-38b8-4456-ac5f-702912bac0dd · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12afa9f2-8495-4ea5-a717-ffd7c39d6a48 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a575274-5666-4764-9913-0df741e67ef7 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98df20ed-9ab8-4277-9bf6-30125ef8485f · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The Llama 3 Herd of Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f161f21a-9365-4361-ab7b-725b7b14fb53 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A holistic approach to undesired content detection in the real world,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 881ed3e0-8875-48cf-a485-a295d796bf27 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eebe795-ea34-4255-a580-d7a0b4b0d66b · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Certifying LLM Safety against Adversarial Prompting
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b5417fd-5bef-4bc5-8652-0547f8590e12 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Lightweight safety guardrails using fine-tuned bert embeddings,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0986a88f-38f8-4ea1-9def-a33b45ba276e · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c0882b-84ae-4945-a51d-c351a28d9757 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Navigating the OverKill in Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10a3c887-315a-40de-a43f-b04723e1dbe2 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cfb6400-5c9e-4148-bcae-beb5aafdf18e · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b83ca1f-1e85-426b-b604-e817fe22386e · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef3557ae-8039-45ce-8b44-83a2979f2965 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0abc586-af39-432a-8053-36a2ccc3571d · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Survey on Mixture of Experts in Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4632282d-a94f-4fc2-b241-76f34871da81 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Uni-moe: Scaling unified multimodal llms with mixture of experts,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8371582d-84ec-4ec6-b30c-437d0b14014e · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e098f9-c354-479b-b8f2-9964ef0d3085 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9becc34-a79c-413b-b080-229bd19f3777 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models with Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426a32ea-c602-491d-9fc6-da5b8952a489 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f56caf1-de30-425d-a0ec-e7d7fabde360 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fe51abf-e933-4e20-802a-28412b67252e · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64bd5f15-fa7e-4f0d-b2ed-e26b626adbd3 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ba4a17-9d7e-4a19-8b37-598ca4bf46c7 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fef8e249-e825-45c4-9063-a5adb4df700e · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbroken: How does llm safety training fail?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe644a2a-601f-4f18-b136-82dca3400e09 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Automatically auditing large language models via discrete optimization,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6efadde-989c-4147-8d82-5b3e79debe8b · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efbff674-42ba-4754-b406-ece032bb95c6 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8277e1ce-0848-4875-b947-5de95d990291 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fed6c88-2376-428a-bba4-3f70b0047335 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60840032-8432-4455-8d31-8136cb96982e · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Removing RLHF Protections in GPT-4 via Fine-Tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0746e8d6-ca2d-490b-bf08-d3c599108cad · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A survey on large language model (llm) security and privacy: The good, the bad, and the ugly,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a343dbb-5ea4-4ee0-8f64-0f9446a2b4ad · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Defending chatgpt against jailbreak attack via self-reminder,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b551ba0-8dcc-4009-9481-9a872dae0065 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41b7dcb4-1a29-42f4-825f-7e6bcf4fb978 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d14d86-01b1-421c-a413-6a07a1aa0db2 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security LoRA: Low-Rank Adaptation of Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c5c855-997c-46f1-a7ea-85d38f328d70 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d452d0-ad48-457a-9e2f-2a8516dd405b · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Closer Look into Mixture-of-Experts in Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5381ff8e-c8a9-4738-95c8-afc06896dcd8 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Adaptive Attention Span in Transformers
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e54ab47-115d-4f55-8f27-43baeb65ec61 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Is On-Device AI Broken and Exploitable? Assessing the Trust and Ethics in Small Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 236bfd5a-4043-4a69-9b94-c05979f94c3b · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The hidden risks of large reasoning models: A safety assessment of r1,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3817d6ae-1a90-4122-ad7d-1712b98f17b0 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Falcon-40B: an open large language model with state-of-the-art performance,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e150241-98f4-4055-aadf-150db08c3e80 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2 technical report,
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72b6ea39-32e3-4be0-9da5-bd150d90686a · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Qwen2.5: A party of foundation models,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f203041f-f3e2-43da-91ea-3b60b6da5816 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb3ab6d-60ab-4875-8b4e-ea684f463a66 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security DeepSeek-V3 Technical Report
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f663f21d-6fa0-4780-abeb-cc0ad5f7768c · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The unlocking spell on base llms: Rethinking alignment via in-context learning,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4736afa-cf44-4f03-8fef-46af2cc792f0 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Training Verifiers to Solve Math Word Problems
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89dce95f-0d2c-4792-9e61-17f46f8792bb · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af7982f-a890-4203-9d1f-94fefb2fb084 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc050fd4-0ae3-4005-a987-602c34e7b2ed · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Language models are super mario: Absorbing abilities from homologous models as a free lunch,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9694cb0d-1a26-4cb1-8a28-613fb8689937 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b591270b-0b7d-4011-845c-880f0d93f679 · outbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.