REVIEW 16 cited by
BLEURT: Learning Robust Metrics for Text Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text generation has made significant advances in the last few years. Yet, evaluation metrics have lagged behind, as the most popular choices (e.g., BLEU and ROUGE) may correlate poorly with human judgments. We propose BLEURT, a learned evaluation metric based on BERT that can model human judgments with a few thousand possibly biased training examples. A key aspect of our approach is a novel pre-training scheme that uses millions of synthetic examples to help the model generalize. BLEURT provides state-of-the-art results on the last three years of the WMT Metrics shared task and the WebNLG Competition dataset. In contrast to a vanilla BERT-based approach, it yields superior results even when the training data is scarce and out-of-distribution.
Forward citations
Cited by 16 Pith papers
-
LLM-Based User Personas for Recommendations at Scale
A framework for real-time LLM-based user interest personas in large-scale video recommendations, using distillation, async inference, and video clustering to balance interests with novel topics and improve viewer valu...
-
ArgCMV: An Argument Summarization Benchmark for the LLM-era
ArgCMV is a new LLM-curated benchmark of about 12,000 arguments from r/ChangeMyView, and current key point extraction methods transfer poorly to it.
-
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
An end-to-end simultaneous speech-to-speech translation model with voice cloning, trained with a two-stage reinforcement learning reward scheme, reports high accuracy and low latency on the authors' RealSI benchmark.
-
Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation
A dual visual encoder with contrastive visual-text pretraining achieves the best reported BLEU-4 score among gloss-free sign language translation methods on Phoenix-2014T.
-
Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework
A new error-attribution dataset and fine-tuned judge model that outputs score, error category, and feedback for LLM responses.
-
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.
-
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data
Fine-tuning GPT-3.5 and Llama 2 on r/Anxiety posts improves readability but raises toxicity and bias while reducing empathy and reflection.
-
LaQual: An Automated Framework for LLM App Quality Evaluation
LaQual automates LLM app-store quality evaluation through scenario classification, static indicator filtering, and LLM-generated dynamic metrics, with Spearman correlations of about 0.6 against human ratings.
-
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
Kvasir-VQA-x1 expands Kvasir-VQA with 159,549 LLM-generated question-answer pairs stratified into three complexity levels, plus a robustness track using weakly augmented images.
-
Federated In-Context Learning: Iterative Refinement for Improved Answer Quality
Fed-ICL iteratively refines QA answers via federated in-context learning with only label transmission, showing convergence on a linear attention model and gains on MMLU and TruthfulQA.
-
RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation
RIVAL iteratively re-trains a reward model adversarially against the current translator and adds a BLEU-predicting head, improving in-domain WMT and subtitle translation over SFT baselines.
-
Teaching with Lies: Curriculum DPO on Synthetic Negatives for Hallucination Detection
Using hallucinated benchmark answers as rejected DPO pairs, ordered by an external fact-checker's grounding score, improves hallucination detection in 1B-3B Llama models.
-
How Small Transformation Expose the Weakness of Semantic Similarity Measures
A diagnostic benchmark of text and code transformations finds embedding similarity metrics often conflate opposition with equivalence; LLM judges discriminate better, and Euclidean distance improves code embeddings.
-
Preserving Privacy, Increasing Accessibility, and Reducing Cost: An On-Device Artificial Intelligence Model for Medical Transcription and Note Generation
Fine-tuning a 1B Llama model on synthetic endocrinology data improves structured medical note generation and substantially reduces LLM-judged hallucinations and omissions in a browser-based, on-device deployment.
-
When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications
Adding generic prompt rules to task-specific LLM prompts is not monotonic: in 15-20 case local suites, Llama 3 and Qwen 2.5 sometimes pass fewer extraction and RAG checks, so prompt changes should be tested per task.
-
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
A proposed Persian biomedical LLM, BioPars, is evaluated on medical QA datasets and reported to beat GPT-4 on a self-built Persian QA benchmark, but the training setup is not described.
Discussion (0). Sign in to comment.