REVIEW 5 cited by
Instruction Pre-Training: Language Models are Supervised Multitask Learners
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Unsupervised multitask pre-training has been the critical method behind the recent success of language models (LMs). However, supervised multitask learning still holds significant promise, as scaling it in the post-training stage trends towards better generalization. In this paper, we explore supervised multitask pre-training by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train LMs. The instruction-response pairs are generated by an efficient instruction synthesizer built on open-source models. In our experiments, we synthesize 200M instruction-response pairs covering 40+ task categories to verify the effectiveness of Instruction Pre-Training. In pre-training from scratch, Instruction Pre-Training not only consistently enhances pre-trained base models but also benefits more from further instruction tuning. In continual pre-training, Instruction Pre-Training enables Llama3-8B to be comparable to or even outperform Llama3-70B. Our model, code, and data are available at https://github.com/microsoft/LMOps.
Forward citations
Cited by 5 Pith papers
-
How Low Can We Go? Minimum Spectroscopic Requirements For Supernova Subtype Classification
ABC-SN classifies ten supernova subtypes with no performance loss down to R_λ=50 and SNR=5, and only minimal loss at R_λ=25.
-
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
Long chain-of-thought and RL training on math problems improves general reasoning benchmarks, while short chain-of-thought math fine-tuning often degrades performance.
-
Localizing Persona Representations in LLMs
Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.
-
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
Running multiple short annealing runs at different token scales can reveal per-source utility scaling curves that change data-source rankings compared with single point estimates.
-
A Survey of LLM $\times$ DATA
A comprehensive survey of the bidirectional links between LLMs and data management, organized as DATA4LLM and LLM4DATA with a new 'IaaS' data-quality framework.
Discussion (0). Sign in to comment.