Pith. sign in

REVIEW 2 cited by

Regression-aware Inference with LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.04182 v3 pith:TCJ6EJPL submitted 2024-03-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords inferenceregressionscoringllmsmetricsmodelsacrossalternate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown strong results on a range of applications, including regression and scoring tasks. Typically, one obtains outputs from an LLM via autoregressive sampling from the model's output distribution. We show that this inference strategy can be sub-optimal for common regression and scoring evaluation metrics. As a remedy, we build on prior work on Minimum Bayes Risk decoding, and propose alternate inference strategies that estimate the Bayes-optimal solution for regression and scoring metrics in closed-form from sampled responses. We show that our proposal significantly improves over baselines across datasets and models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking the Capability of Fine-Tuned Language Models for Automated Vulnerability Repair

    cs.SE 2025-12 conditional novelty 6.0 of 10

    Fine-tuned AVR models memorize overlapping training data, so reported repair rates fall from ~20% to ~5% on non-overlapping splits, and match-based metrics misjudge true fixes.

  2. Tackling prediction tasks in relational databases with LLMs

    cs.LG 2024-11 conditional novelty 5.0 of 10

    Pre-trained LLMs, fed serialized relational rows with related examples, achieve competitive AUROC/MAE on RelBench without fine-tuning, but the headline comparison is weakened by pretraining contamination on Formula 1 tasks.

Pith tools