Pith. sign in

REVIEW 3 cited by

Inference-Time Language Model Alignment via Integrated Value Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.17819 v1 pith:WFAW2UUC submitted 2024-09-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords largemodelsvalueguidancelanguagetextttfunctionsalignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Large language models are typically fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. In this work, we introduce $\textit{Integrated Value Guidance}$ (IVG), a method that uses implicit and explicit value functions to guide language model decoding at token and chunk-level respectively, efficiently aligning large language models purely at inference time. This approach circumvents the complexities of direct fine-tuning and outperforms traditional methods. Empirically, we demonstrate the versatility of IVG across various tasks. In controlled sentiment generation and summarization tasks, our method significantly improves the alignment of large models using inference-time guidance from $\texttt{gpt2}$-based value functions. Moreover, in a more challenging instruction-following benchmark AlpacaEval 2.0, we show that both specifically tuned and off-the-shelf value functions greatly improve the length-controlled win rates of large models against $\texttt{gpt-4-turbo}$ (e.g., $19.51\% \rightarrow 26.51\%$ for $\texttt{Mistral-7B-Instruct-v0.2}$ and $25.58\% \rightarrow 33.75\%$ for $\texttt{Mixtral-8x7B-Instruct-v0.1}$ with Tulu guidance).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mitigating Object Hallucination via Robust Local Perception Search

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free decoding method that uses an MLLM's own local object descriptions as a reward prior, combined with CLIP similarity, to cut object hallucination, especially under adversarial image noise.

  2. Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback

    cs.CL 2025-01 conditional novelty 5.0 of 10

    TPO iteratively refines LLM responses at inference time using reward-model scores converted into textual critiques, letting an unaligned 70B model beat its aligned counterpart on several benchmarks.

  3. Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

    cs.CY 2024-12 conditional novelty 4.0 of 10

    The paper proposes the AI-45 degree law, a Causal Ladder framework, and five trustworthiness levels as a roadmap toward trustworthy AGI.

Pith tools