Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative Data

Bin Sun; Boyuan Pan; Heda Wang; Kan Li; Peiwen Yuan; Shaoxiong Feng; Xinglin Wang; Yiwei Li

arxiv: 2312.12832 · v1 · pith:7SEIB2QMnew · submitted 2023-12-20 · 💻 cs.CL · cs.AI

Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative Data

Yiwei Li , Peiwen Yuan , Shaoxiong Feng , Boyuan Pan , Bin Sun , Xinglin Wang , Heda Wang , Kan Li This is my paper

classification 💻 cs.CL cs.AI

keywords reasoningdatallmsnegativecomplexdistillingframeworkknowledge

0 comments

read the original abstract

Large Language Models (LLMs) have performed well on various reasoning tasks, but their inaccessibility and numerous parameters hinder wide application in practice. One promising way is distilling the reasoning ability from LLMs to small models by the generated chain-of-thought reasoning paths. In some cases, however, LLMs may produce incorrect reasoning chains, especially when facing complex mathematical problems. Previous studies only transfer knowledge from positive samples and drop the synthesized data with wrong answers. In this work, we illustrate the merit of negative data and propose a model specialization framework to distill LLMs with negative samples besides positive ones. The framework consists of three progressive steps, covering from training to inference stages, to absorb knowledge from negative data. We conduct extensive experiments across arithmetic reasoning tasks to demonstrate the role of negative data in distillation from LLM.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
cs.AI 2026-07 unverdicted novelty 5.0

C3RL is a new RL algorithm combining correctness, calibration, and reference accuracy rewards to improve LLM confidence calibration, enabling CAS to outperform majority voting with up to 12.33x lower inference cost.