REVIEW 3 cited by
Language models and Automated Essay Scoring
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we present a new comparative study on automatic essay scoring (AES). The current state-of-the-art natural language processing (NLP) neural network architectures are used in this work to achieve above human-level accuracy on the publicly available Kaggle AES dataset. We compare two powerful language models, BERT and XLNet, and describe all the layers and network architectures in these models. We elucidate the network architectures of BERT and XLNet using clear notation and diagrams and explain the advantages of transformer architectures over traditional recurrent neural network architectures. Linear algebra notation is used to clarify the functions of transformers and attention mechanisms. We compare the results with more traditional methods, such as bag of words (BOW) and long short term memory (LSTM) networks.
Forward citations
Cited by 3 Pith papers
-
Adversarial Topic-aware Prompt-tuning for Cross-topic Automated Essay Scoring
ATOP uses shared and topic-specific soft prompts with adversarial training and pseudo-labels to improve cross-topic automated essay scoring, reporting better QWK scores than nine baselines on ASAP++.
-
Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
Fine-tuned LLMs can reconstruct item characteristic curves from multiple-choice item text, giving useful estimates of IRT difficulty and discrimination without live student response data.
-
Automated Essay Scoring Incorporating Annotations from Automated Feedback Systems
On the PERSUADE corpus, adding generated argument-component tags to essay text raised automated scoring agreement from a QWK of 0.860 to 0.868, while error-only tags lowered it.
Discussion (0). Continue with ORCID to comment.