REVIEW 2 cited by
Regression-aware Inference with LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have shown strong results on a range of applications, including regression and scoring tasks. Typically, one obtains outputs from an LLM via autoregressive sampling from the model's output distribution. We show that this inference strategy can be sub-optimal for common regression and scoring evaluation metrics. As a remedy, we build on prior work on Minimum Bayes Risk decoding, and propose alternate inference strategies that estimate the Bayes-optimal solution for regression and scoring metrics in closed-form from sampled responses. We show that our proposal significantly improves over baselines across datasets and models.
Forward citations
Cited by 2 Pith papers
-
Rethinking the Capability of Fine-Tuned Language Models for Automated Vulnerability Repair
Fine-tuned AVR models memorize overlapping training data, so reported repair rates fall from ~20% to ~5% on non-overlapping splits, and match-based metrics misjudge true fixes.
-
Tackling prediction tasks in relational databases with LLMs
Pre-trained LLMs, fed serialized relational rows with related examples, achieve competitive AUROC/MAE on RelBench without fine-tuning, but the headline comparison is weakened by pretraining contamination on Formula 1 tasks.
Discussion (0). Continue with ORCID to comment.