Pith. sign in

REVIEW 2 cited by

DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09136 v1 pith:KPF5VL3D submitted 2024-02-14 cs.CL cs.AI

classification cs.CLcs.AI
keywords codeinstructiondiverseabilitygenerationllmsperformancetuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Code Large Language Models (Code LLMs) have demonstrated outstanding performance in code-related tasks. Several instruction tuning approaches have been proposed to boost the code generation performance of pre-trained Code LLMs. In this paper, we introduce a diverse instruction model (DolphCoder) with self-evaluating for code generation. It learns diverse instruction targets and combines a code evaluation objective to enhance its code generation ability. Our model achieves superior performance on the HumanEval and MBPP benchmarks, demonstrating new insights for future code instruction tuning work. Our key findings are: (1) Augmenting more diverse responses with distinct reasoning paths increases the code capability of LLMs. (2) Improving one's ability to evaluate the correctness of code solutions also enhances their ability to create it.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Success is in the Details: Evaluate and Enhance Details Sensitivity of Code LLMs through Counterfactuals

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Code LLMs drop by more than 10% in accuracy when problem details are subtly changed, and fine-tuning on such counterfactual variants boosts performance on standard benchmarks.

  2. Mixture of Reasonings: Teach Large Language Models to Reason with Adaptive Strategies

    cs.CL 2025-07 reject novelty 4.0 of 10

    MoR fine-tunes Qwen2.5 on GPT-4o-selected reasoning templates, claiming up to 13.5% accuracy gains, but the reported gains are not robustly supported.

Pith tools