Pith. sign in

REVIEW 2 cited by

Evaluating the Prompt Steerability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.12405 v2 pith:OGKI3YRR submitted 2024-11-19 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords steerabilitymodelbenchmarkmodelsableacrossbaselinedegree
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model is capable of reflecting various personas. To this end, we propose a benchmark for evaluating the steerability of model personas as a function of prompting. Our design is based on a formal definition of prompt steerability, which analyzes the degree to which a model's joint behavioral distribution can be shifted from its baseline. By defining steerability indices and inspecting how these indices change as a function of steering effort, we can estimate the steerability of a model across various persona dimensions and directions. Our benchmark reveals that the steerability of many current models is limited -- due to both a skew in their baseline behavior and an asymmetry in their steerability across many persona dimensions. We release an implementation of our benchmark at https://github.com/IBM/prompt-steering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An Auditable Agent Platform For Automated Molecular Optimisation

    cs.LG 2025-08 conditional novelty 5.0 of 10

    A hierarchical multi-agent LLM platform with recorded provenance improved average predicted binding affinity for AKT1 by 31%, while single-agent runs favored drug-likeness.

  2. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.

Pith tools