Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Correcting Large Language Model Behavior via Influence Function

cs.LG · 2024-12-21 · conditional · novelty 6.0

LANCET uses influence functions to find training examples that drive an LLM's harmful outputs and then fine-tunes the model with a pairwise ranking loss to suppress those outputs without human-labeled corrections.

citing papers explorer

Showing 1 of 1 citing paper.

  • Correcting Large Language Model Behavior via Influence Function cs.LG · 2024-12-21 · conditional · none · ref 1

    LANCET uses influence functions to find training examples that drive an LLM's harmful outputs and then fine-tunes the model with a pairwise ranking loss to suppress those outputs without human-labeled corrections.