REVIEW 5 cited by
Sensitivity and Robustness of Large Language Models to Prompt Template in Japanese Text Classification Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Prompt engineering relevance research has seen a notable surge in recent years, primarily driven by advancements in pre-trained language models and large language models. However, a critical issue has been identified within this domain: the inadequate of sensitivity and robustness of these models towards Prompt Templates, particularly in lesser-studied languages such as Japanese. This paper explores this issue through a comprehensive evaluation of several representative Large Language Models (LLMs) and a widely-utilized pre-trained model(PLM). These models are scrutinized using a benchmark dataset in Japanese, with the aim to assess and analyze the performance of the current multilingual models in this context. Our experimental results reveal startling discrepancies. A simple modification in the sentence structure of the Prompt Template led to a drastic drop in the accuracy of GPT-4 from 49.21 to 25.44. This observation underscores the fact that even the highly performance GPT-4 model encounters significant stability issues when dealing with diverse Japanese prompt templates, rendering the consistency of the model's output results questionable. In light of these findings, we conclude by proposing potential research trajectories to further enhance the development and performance of Large Language Models in their current stage.
Forward citations
Cited by 5 Pith papers
-
Re-evaluating LLM-based Heuristic Search: A Case Study on the 3D Packing Problem
LLM-guided evolutionary search, aided by scaffolding and self-correction, discovered a 3D packing scoring function competitive with human heuristics, but the model invented no new algorithm structures and its results ...
-
Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech
GPT-4 ratings of five dementia-related language symptoms, added to 40 standard linguistic features, improve automatic Alzheimer's detection from spontaneous speech transcripts, reaching AUROC 0.931 on ADReSS.
-
Enhancing Semantic Consistency of Large Language Models through Model Editing: An Interpretability-Oriented Approach
Steering a handful of attention heads with a bias vector derived from consistent prompt pairs improves semantic consistency of LLaMA-2-7B on paraphrased NLU and NLG tasks.
-
LF-Steering: Latent Feature Activation Steering for Enhancing Semantic Consistency in Large Language Models
LF-Steering steers sparse-autoencoder features instead of whole layers or attention heads to improve the semantic consistency of Llama-2-7B-Chat on paraphrase benchmarks, but the steering formula discards direction.
-
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
Multimodal language models trained with a medium number of instruction templates (5,000 for 7B, 100 for 13B) outperform both fewer and many more templates, with gains up to 10 points on small benchmark samples.
Discussion (0). Continue with ORCID to comment.