Pith. sign in

REVIEW 8 cited by

A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12994 v2 pith:EWH2GZOB submitted 2024-07-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsengineeringdifferentpromptlanguagetasksknowledgemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown remarkable performance on many different Natural Language Processing (NLP) tasks. Prompt engineering plays a key role in adding more to the already existing abilities of LLMs to achieve significant performance gains on various NLP tasks. Prompt engineering requires composing natural language instructions called prompts to elicit knowledge from LLMs in a structured way. Unlike previous state-of-the-art (SoTA) models, prompt engineering does not require extensive parameter re-training or fine-tuning based on the given NLP task and thus solely operates on the embedded knowledge of LLMs. Additionally, LLM enthusiasts can intelligently extract LLMs' knowledge through a basic natural language conversational exchange or prompt engineering, allowing more and more people even without deep mathematical machine learning background to experiment with LLMs. With prompt engineering gaining popularity in the last two years, researchers have come up with numerous engineering techniques around designing prompts to improve accuracy of information extraction from the LLMs. In this paper, we summarize different prompting techniques and club them together based on different NLP tasks that they have been used for. We further granularly highlight the performance of these prompting strategies on various datasets belonging to that NLP task, talk about the corresponding LLMs used, present a taxonomy diagram and discuss the possible SoTA for specific datasets. In total, we read and present a survey of 44 research papers which talk about 39 different prompting methods on 29 different NLP tasks of which most of them have been published in the last two years.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution

    cs.MA 2026-08 conditional novelty 6.0 of 10

    An eight-agent question-asking system that front-loads intent clarification produced more complete prompts, higher-rated outputs, and single-turn task completion in a four-person pilot, with unstable effect sizes.

  2. MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval

    cs.IR 2026-01 unverdicted novelty 6.0 of 10

    MCERF delivers a 41.1% relative accuracy gain on the DesignQA benchmark by combining ColPali vision-language retrieval with four specialized reasoning modes and dynamic routing.

  3. Revisiting Prompt Engineering: A Comprehensive Evaluation for LLM-based Personalized Recommendation

    cs.IR 2025-07 conditional novelty 6.0 of 10

    For cost-efficient LLMs, rephrasing, step-back, and structured reasoning prompts raise ranking accuracy; for high-performance LLMs, a simple baseline prompt matches complex prompts at a fraction of the cost.

  4. Large Language Models in the Task of Automatic Validation of Text Classifier Predictions

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLM-based annotators using token-probability thresholds, RAG, and reasoning fine-tuning matched or exceeded human annotator quality on a proprietary 250-class intent-validation task.

  5. Demystifying Feature Requests: Leveraging LLMs to Refine Feature Requests in Open-Source Software

    cs.SE 2025-07 conditional novelty 5.0 of 10

    GPT-4o with in-context learning can flag ambiguity and incompleteness in GitHub feature requests and draft clarification questions, though moderate annotator agreement and a small sample limit the strength of the evidence.

  6. Extracting Research Instruments from Educational Literature Using LLMs

    cs.IR 2025-05 conditional novelty 5.0 of 10

    A multi-step LLM pipeline extracts research instruments and their attributes from education literature, reporting moderate F1 scores but no released code or baseline statistics.

  7. Fine-Tuning and Prompt Engineering of LLMs, for the Creation of Multi-Agent AI for Addressing Sustainable Protein Production Challenges

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A proof-of-concept multi-agent GPT system for microbial protein literature extraction shows both fine-tuning and prompt engineering improve cosine-similarity scores, with fine-tuning slightly ahead but more variable.

  8. Incorporating Token Usage into Prompting Strategy Evaluation

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Prompting strategies show sharply diminishing accuracy returns as token usage increases, and the paper formalizes this with Big-O_tok and Token Cost metrics.

Pith tools