Pith. sign in

REVIEW 2 cited by

Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13833 v2 pith:2CRB4BTX submitted 2024-07-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsphi-3safetylanguagealigningbreak-fixcyclepost-training
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models are deployed in an increasing number of domains, it is critical to ensure that they are aligned with human preferences and safety considerations. In this report, we present our methodology for safety aligning the Phi-3 series of language models. We utilized a "break-fix" cycle, performing multiple rounds of dataset curation, safety post-training, benchmarking, red teaming, and vulnerability identification to cover a variety of harm areas in both single and multi-turn scenarios. Our results indicate that this approach iteratively improved the performance of the Phi-3 models across a wide range of responsible AI benchmarks. Finally, we include additional red teaming strategies and evaluations that were used to test the safety behavior of Phi-3.5-mini and Phi-3.5-MoE, which were optimized for multilingual capabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems

    cs.DC 2026-01 conditional novelty 5.0 of 10

    On simulated GPU-NDP-DIMM hardware, scheduling MoE experts with tensor parallelism, load balancing, and prefill-driven pre-fetching cuts end-to-end latency by 2.41x on average versus MoNDE.

  2. Securing LLMs in the Wild: Privacy and Security Challenges at the Edge

    cs.CR 2026-07 conditional novelty 3.0 of 10

    Edge LLM security is framed as a Security-Efficiency Paradox, with a three-wall constraint model, a composite SOES score, and a small FP/INT4 benchmark of six models.

Pith tools