Pith. sign in

REVIEW 1 cited by

Preference-grounded Token-level Guidance for Language Model Fine-tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.00398 v3 pith:YZXRZNOB submitted 2023-06-01 cs.CL

classification cs.CL
keywords guidancelearningtraininggenerationlanguagepreferencelearnedlevel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Aligning language models (LMs) with preferences is an important problem in natural language generation. A key challenge is that preferences are typically provided at the sequence level while LM training and generation both occur at the token level. There is, therefore, a granularity mismatch between the preference and the LM training losses, which may complicate the learning problem. In this paper, we address this issue by developing an alternate training process, where we iterate between grounding the sequence-level preference into token-level training guidance, and improving the LM with the learned guidance. For guidance learning, we design a framework that extends the pairwise-preference learning in imitation learning to both variable-length LM generation and the utilization of the preference among multiple generations. For LM training, based on the amount of supervised data, we present two minimalist learning objectives that utilize the learned guidance. In experiments, our method performs competitively on two distinct representative LM tasks -- discrete-prompt generation and text summarization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A shuffle-based token classifier plus group-level loss reweighting improves supervised fine-tuning of LLM agents on tool-use benchmarks.

Pith tools