Pith. sign in

REVIEW 4 cited by

Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20535 v2 pith:AJJMHSJT submitted 2024-05-30 cs.AI cs.CL

classification cs.AIcs.CL
keywords codingdatareasoningacrossfamiliesdifferentimpactmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction Fine-Tuning (IFT) significantly enhances the zero-shot capabilities of pretrained Large Language Models (LLMs). While coding data is known to boost LLM reasoning abilities during pretraining, its role in activating internal reasoning capacities during IFT remains understudied. This paper investigates a key question: How does coding data impact LLMs' reasoning capacities during IFT stage? To explore this, we thoroughly examine the impact of coding data across different coding data proportions, model families, sizes, and reasoning domains, from various perspectives. Specifically, we create three IFT datasets with increasing coding data proportions, fine-tune six LLM backbones across different families and scales on these datasets, evaluate the tuned models' performance across twelve tasks in three reasoning domains, and analyze the outcomes from three broad-to-granular perspectives: overall, domain-level, and task-specific. Our holistic analysis provides valuable insights into each perspective. First, coding data tuning enhances the overall reasoning capabilities of LLMs across different model families and scales. Moreover, while the impact of coding data varies by domain, it shows consistent trends within each domain across different model families and scales. Additionally, coding data generally provides comparable task-specific benefits across model families, with optimal proportions in IFT datasets being task-dependent.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study

    cs.CR 2026-04 conditional novelty 7.0 of 10

    Large-scale audit of SkillsMP agent skills finds 520 skills with 1,708 credential-leak issues, dominated by debug logging into the LLM context and hard-to-remediate forks.

  2. Evaluating Language Models as Synthetic Data Generators

    cs.CL 2024-12 conditional novelty 6.0 of 10

    AgoraBench shows that an LM's ability to solve problems does not predict its ability to generate useful synthetic training data.

  3. CoinMath: Harnessing the Power of Coding Instruction for Math LLMs

    cs.CL 2024-12 conditional novelty 5.0 of 10

    CoinMath improves math LLM accuracy by training on GPT-4o-generated code rationales with concise comments, descriptive naming, and hardcoded solutions.

  4. Thinking with Knowledge Graphs: Enhancing LLM Reasoning Through Structured Data

    cs.CL 2024-12 conditional novelty 3.0 of 10

    Representing knowledge graph triples as Python code improved LLM multi-hop reasoning accuracy over text and JSON in this study, though the effect is modest and possibly due to explicit inference steps in the code.

Pith tools