Pith. sign in

REVIEW 6 cited by

NEFTune: Noisy Embeddings Improve Instruction Finetuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.05914 v2 pith:YOQACDHT submitted 2023-10-09 cs.CL cs.LG

classification cs.CLcs.LG
keywords neftunefinetuningimprovementembeddingsinstructionmodelsnoisytraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We show that language model finetuning can be improved, sometimes dramatically, with a simple augmentation. NEFTune adds noise to the embedding vectors during training. Standard finetuning of LLaMA-2-7B using Alpaca achieves 29.79% on AlpacaEval, which rises to 64.69% using noisy embeddings. NEFTune also improves over strong baselines on modern instruction datasets. Models trained with Evol-Instruct see a 10% improvement, with ShareGPT an 8% improvement, and with OpenPlatypus an 8% improvement. Even powerful models further refined with RLHF such as LLaMA-2-Chat benefit from additional training with NEFTune.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 13 citations worldwide. Full citation record

  1. Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study

    cs.CL 2025-09 conditional novelty 5.0 of 10

    A 102B Korean-English model, expanded from Llama 3 70B with LlamaPro and Masked Structure Growth and trained on 194B tokens, scores 64.74 on KMMLU and 83.34 on KorMedMCQA, roughly matching GPT-4 on Korean benchmarks.

  2. More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A generalized pseudo-label-robust loss plus dynamic CLIP pseudo-labeling achieves state-of-the-art mAP on VOC, COCO, NUS-WIDE, and CUB in the single-positive multi-label setting.

  3. RankLLM: A Python Package for Reranking with LLMs

    cs.IR 2025-05 accept novelty 5.0 of 10

    RankLLM is an open-source Python package that modularly supports pointwise, pairwise, and listwise LLM rerankers, with integrated retrieval, evaluation, training, and response analysis, and reproduces results from Ran...

  4. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

  5. Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

    cs.AI 2026-02 reject novelty 4.0 of 10

    Mixing an LLM's logits with an earlier weak checkpoint during fine-tuning yields math and code accuracy gains beyond standard SFT saturation.

  6. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

    cs.AI 2026-06 unverdicted novelty 2.0 of 10

    A survey-style reference book mapping the full agentic-AI stack from transformer internals to production deployment, with no new research result.

Pith tools