Pith. sign in

REVIEW 2 cited by

Automated Question Generation for Science Tests in Arabic Language Using NLP Techniques

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08520 v1 pith:WEGBKG6O submitted 2024-06-11 cs.CL cs.CY

classification cs.CLcs.CY
keywords arabicgenerationquestionquestionsassessmenteducationeducationalefficiency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Question generation for education assessments is a growing field within artificial intelligence applied to education. These question-generation tools have significant importance in the educational technology domain, such as intelligent tutoring systems and dialogue-based platforms. The automatic generation of assessment questions, which entail clear-cut answers, usually relies on syntactical and semantic indications within declarative sentences, which are then transformed into questions. Recent research has explored the generation of assessment educational questions in Arabic. The reported performance has been adversely affected by inherent errors, including sentence parsing inaccuracies, name entity recognition issues, and errors stemming from rule-based question transformation. Furthermore, the complexity of lengthy Arabic sentences has contributed to these challenges. This research presents an innovative Arabic question-generation system built upon a three-stage process: keywords and key phrases extraction, question generation, and subsequent ranking. The aim is to tackle the difficulties associated with automatically generating assessment questions in the Arabic language. The proposed approach and results show a precision of 83.50%, a recall of 78.68%, and an Fl score of 80.95%, indicating the framework high efficiency. Human evaluation further confirmed the model efficiency, receiving an average rating of 84%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity

    cs.CL 2025-01 conditional novelty 4.0 of 10

    XGBoost and Random Forest can distinguish ChatGPT-written cybersecurity paragraphs from human Wikipedia paragraphs with 81 to 83% accuracy, and a narrow XGBoost model beat GPTZero in a three-class test, yet the benchm...

  2. Leveraging Explainable AI for LLM Text Attribution: Differentiating Human-Written and Multiple LLMs-Generated Text

    cs.CL 2025-01 reject novelty 3.0 of 10

    On a 600-essay, two-topic dataset, TF-IDF features let a Random Forest identify which of five LLMs or a human wrote a text with about 97% accuracy, though the comparison to GPTZero is inconsistent.

Pith tools