Pith. sign in

REVIEW 8 cited by

PromptBench: A Unified Library for Evaluation of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07910 v3 pith:54726C2Y submitted 2023-12-13 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords evaluationpromptbenchpromptlanguagelargelibraryllmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The evaluation of large language models (LLMs) is crucial to assess their performance and mitigate potential security risks. In this paper, we introduce PromptBench, a unified library to evaluate LLMs. It consists of several key components that are easily used and extended by researchers: prompt construction, prompt engineering, dataset and model loading, adversarial prompt attack, dynamic evaluation protocols, and analysis tools. PromptBench is designed to be an open, general, and flexible codebase for research purposes that can facilitate original study in creating new benchmarks, deploying downstream applications, and designing new evaluation protocols. The code is available at: https://github.com/microsoft/promptbench and will be continuously supported.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 13 citations worldwide. Full citation record

  1. MolecularCanvas: LLM-assisted Small-Molecule Drug Discovery via Structure-Guided Constraints

    cs.HC 2026-08 conditional novelty 6.0 of 10

    MolecularCanvas, an interactive system using structure-level annotations and constraints to steer an LLM's molecular edit plans, received higher user ratings and expert-assessed molecule quality than a baseline workfl...

  2. Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

    cs.CL 2026-05 conditional novelty 6.0 of 10

    In a 430k-evaluation study, plain baseline prompting beats most elaborate prompting techniques on non-reasoning LLMs across MCQA benchmarks, with only small role-framing variants gaining about 3 percentage points.

  3. Characterizing Fitness Landscape Structures in Prompt Engineering

    cs.AI 2025-09 reject novelty 6.0 of 10

    Prompt fitness autocorrelation appears smooth under systematic enumeration but rugged with an intermediate-distance peak under novelty-driven sampling, yet the two analyses cover non-overlapping distance ranges.

  4. Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators

    cs.AR 2025-05 conditional novelty 6.0 of 10

    A fused exponential-multiplication hardware unit using logarithmic quantization and exponent adjustment reduces FlashAttention accelerator area by about 29% and power by about 18% without visible accuracy loss on GLUE.

  5. FLASH-D: FlashAttention with Hidden Softmax Division

    cs.LG 2025-05 conditional novelty 5.0 of 10

    FlashAttention can be rewritten exactly so each softmax weight is a sigmoid of a neighboring score difference plus a log-weight term, removing max subtraction and simplifying hardware.

  6. GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing

    cs.LG 2025-07 reject novelty 4.0 of 10

    GuardVal combines role-playing jailbreak generation with an Adam-inspired optimizer and an Overall Safety Value metric, but the method is underspecified and not validated with released code or data.

  7. The Science of Evaluating Foundation Models

    cs.CL 2025-02 conditional novelty 3.0 of 10

    A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.

  8. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools