REVIEW 8 cited by
PromptBench: A Unified Library for Evaluation of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The evaluation of large language models (LLMs) is crucial to assess their performance and mitigate potential security risks. In this paper, we introduce PromptBench, a unified library to evaluate LLMs. It consists of several key components that are easily used and extended by researchers: prompt construction, prompt engineering, dataset and model loading, adversarial prompt attack, dynamic evaluation protocols, and analysis tools. PromptBench is designed to be an open, general, and flexible codebase for research purposes that can facilitate original study in creating new benchmarks, deploying downstream applications, and designing new evaluation protocols. The code is available at: https://github.com/microsoft/promptbench and will be continuously supported.
Forward citations
Cited by 8 Pith papers
-
MolecularCanvas: LLM-assisted Small-Molecule Drug Discovery via Structure-Guided Constraints
MolecularCanvas, an interactive system using structure-level annotations and constraints to steer an LLM's molecular edit plans, received higher user ratings and expert-assessed molecule quality than a baseline workfl...
-
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
In a 430k-evaluation study, plain baseline prompting beats most elaborate prompting techniques on non-reasoning LLMs across MCQA benchmarks, with only small role-framing variants gaining about 3 percentage points.
-
Characterizing Fitness Landscape Structures in Prompt Engineering
Prompt fitness autocorrelation appears smooth under systematic enumeration but rugged with an intermediate-distance peak under novelty-driven sampling, yet the two analyses cover non-overlapping distance ranges.
-
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
A fused exponential-multiplication hardware unit using logarithmic quantization and exponent adjustment reduces FlashAttention accelerator area by about 29% and power by about 18% without visible accuracy loss on GLUE.
-
FLASH-D: FlashAttention with Hidden Softmax Division
FlashAttention can be rewritten exactly so each softmax weight is a sigmoid of a neighboring score difference plus a log-weight term, removing max subtraction and simplifying hardware.
-
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
GuardVal combines role-playing jailbreak generation with an Adam-inspired optimizer and an Overall Safety Value metric, but the method is underspecified and not validated with released code or data.
-
The Science of Evaluating Foundation Models
A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.
-
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.
Discussion (0). Continue with ORCID to comment.