Pith. sign in

REVIEW 8 cited by

What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12334 v4 pith:OSYBTQOA submitted 2024-06-18 cs.LG cs.SE

classification cs.LGcs.SE
keywords consistencyllmspromptsensitivityacrosstasksclassificationengineering
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) changed the way we design and interact with software systems. Their ability to process and extract information from text has drastically improved productivity in a number of routine tasks. Developers that want to include these models in their software stack, however, face a dreadful challenge: debugging LLMs' inconsistent behavior across minor variations of the prompt. We therefore introduce two metrics for classification tasks, namely sensitivity and consistency, which are complementary to task performance. First, sensitivity measures changes of predictions across rephrasings of the prompt, and does not require access to ground truth labels. Instead, consistency measures how predictions vary across rephrasings for elements of the same class. We perform an empirical comparison of these metrics on text classification tasks, using them as guideline for understanding failure modes of the LLM. Our hope is that sensitivity and consistency will be helpful to guide prompt engineering and obtain LLMs that balance robustness with performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.

  2. Reinforcement Learning for Machine Learning Engineering Agents

    cs.LG 2025-09 conditional novelty 6.0 of 10

    RL-trained Qwen2.5-3B outperforms prompted Claude-3.5-Sonnet and GPT-4o on 12 MLEBench tasks by an average of 22% and 24%, using two targeted RL modifications.

  3. When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning

    cs.SE 2025-09 conditional novelty 5.0 of 10

    LLM-based sketch generation for embedded ML is fragile, with success rates below 40%, and prompt structure alone can swing outcomes from 15% to 30%.

  4. Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs

    cs.CY 2025-05 conditional novelty 5.0 of 10

    Across 156 configurations on Persian medical board questions, Chain-of-Thought prompting raised accuracy while increasing overconfidence, and emotional prompting inflated confidence without accuracy gains.

  5. Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets

    cs.AI 2025-05 conditional novelty 5.0 of 10

    AI agents in future labor markets will need metacognitive and strategic reasoning because incomplete information creates adverse selection, moral hazard, and reputation effects.

  6. Position: Intelligent Coding Systems Should Write Programs with Justifications

    cs.SE 2025-08 conditional novelty 4.0 of 10

    A position paper advocating that intelligent coding systems should accompany code with justified explanations that are cognitively aligned and semantically faithful.

  7. A Conceptual Framework for Requirements Engineering of Pretrained-Model-Enabled Systems

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A conceptual framework reorganizes requirements engineering for pretrained-model-enabled systems into six activities, based on identified challenges of opaque capabilities, context sensitivity, and continuous evolution.

  8. CEA-LIST at CheckThat! 2025: Evaluating LLMs as Detectors of Bias and Opinion in Text

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Few-shot prompted LLMs rivaled fine-tuned smaller models in multilingual subjectivity detection, winning the Arabic and Polish tracks of CheckThat! 2025.

Pith tools