{"id":"e641a4f3-2b7d-4131-82d8-c41035750691","arxiv_id":"2506.15598","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GPT-4o and Gemma-2 generate Portuguese reading-comprehension MCQs whose expert and psychometric quality is comparable to human-authored items, while a two-step small-model pipeline underperforms.","lead":"This study tested whether large language models can write Portuguese reading-comprehension multiple-choice questions for elementary students, and compared them with questions written by human educators. It found that the best models produce questions that experts and students judge as comparable to human-authored ones, though with weaknesses in clarity, answerability, and distractor appeal.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'current models' claim is internally contradicted by the paper's own two-step result (55.6% vs 73.3% answerable); the headline must be scoped to the one-step GPT-4o/Gemma-2 pipelines.","rationale":"The paper's central empirical contribution is a careful expert-plus-psychometric evaluation of LLM-generated MCQs in European Portuguese, and the evidence that GPT-4o and Gemma-2 produce expert-acceptable items is genuine. I agree with the reader's conditional verdict: the paper needs revision, particularly to narrow the headline claim and strengthen the statistical support for equivalence. However, I do not think the reader's identified weakest assumption — selection bias in the psychometric analysis — is the most load-bearing concern. The filter that produced the 124 validated items is equally selective for human-authored, GPT-4o, and Gemma-2 (each 33/45), so the comparison 'among validated items' does not disproportionately advantage the one-step LLMs relative to the human benchmark. The Ptt5-v2+Gemma-2 pipeline does lose more items, but the paper already treats that pipeline as inferior. The load-bearing issue is instead the abstract's unqualified 'current models' claim, which is directly contradicted by the paper's own Table 5 unless it is read as 'some current models'. The paper's Section 4.6 itself carefully distinguishes the one-step LLMs before stating the 33/45 equality, yet the abstract and conclusion do not. This is not a matter of outside consensus; it is an internal inconsistency between the headline and the reported acceptance rates. A secondary, related concern is that the 'comparable' verdict in Section 5.7 is based solely on non-significant differences, which do not establish equivalence; effect sizes, confidence intervals, or an explicit equivalence test would be needed to support the wording. Both issues can be fixed without new data collection, so I keep the reader's conditional verdict unchanged.","tokens_in":26893,"tokens_out":10878,"duration_ms":119587,"concrete_test":"Recompute answerability acceptance two ways: (a) per pipeline — GPT-4o 33/45, Gemma-2 33/45, Ptt5-v2+Gemma-2 25/45; (b) aggregate of all LLM pipelines — 91/135 = 67.4% versus human 33/45 = 73.3%. If the abstract's 'current models' is meant collectively, the claim fails and must be narrowed to 'one-step LLM-generated MCQs (GPT-4o and Gemma-2)'. Additionally, run a TOST equivalence test on the human-vs-one-step differences in difficulty (1−P) and discrimination (D) with pre-specified bounds, e.g., ±0.05 for difficulty and ±0.10 for discrimination; if the 90% confidence intervals exceed these bounds, replace 'comparable' with 'not statistically distinguishable'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim — 'current models can generate MCQs of comparable quality to human-authored ones' — is not supported by the full results in Section 4.5.4 / Table 5. The equal acceptance rate (33/45, 73.3%) holds only for the one-step GPT-4o and Gemma-2 models. The third pipeline, Ptt5-v2+Gemma-2, achieves 25/45 (55.6%), a shortfall the paper itself calls 'substantially lower than that of human-authored MCQs'. If 'current models' is read collectively, the aggregate LLM acceptance rate is (33+33+25)/135 = 91/135 = 67.4%, below the human 73.3%. The conclusion in Section 7 and the abstract repeat the unqualified 'comparable' language. Because this headline is the main takeaway, it must be restricted to one-step LLMs or explicitly to GPT-4o and Gemma-2 to avoid internal inconsistency. A related but distinct issue: even for these one-step models, the psychometric comparability claim rests on non-significant Kruskal-Wallis/ANOVA results (Section 5.5, H=1.249, p=0.741; H=0.854, p>0.05) without effect sizes or confidence intervals, so 'comparable' is not formally established as equivalence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an empirical evaluation of automatic multiple-choice question generation for European Portuguese reading comprehension, targeting elementary-school students. Three generation pipelines are compared with human-authored items: zero-shot one-step generation with GPT-4o, zero-shot one-step generation with Gemma-2-27B, and a two-step pipeline combining fine-tuned Ptt5-v2 (question generation) with Gemma-2 (options and difficulty). Quality is assessed through expert review (18 experts, 180 items) on well-formedness, narrative alignment, option clarity, answerability, distractor plausibility, and difficulty, and through classical test theory applied to 284 students' responses on 124 pre-filtered items, yielding difficulty and discrimination indices. The paper also analyzes model-assigned difficulty against expert and student judgments. The main finding is that one-step LLMs can generate MCQs of comparable quality to human-authored ones, while the two-step pipeline performs substantially worse on answerability.","tokens_in":27005,"tokens_out":5377,"duration_ms":54623,"significance":"If its claims hold, this is a valuable contribution to MCQ generation for low-resource languages, providing the first evaluation of narrative- and difficulty-controlled generation in European Portuguese with real classroom data. The study is carefully designed: expert evaluation uses majority voting, provenance is blinded, student data was collected under exam conditions, and the authors include model prompts and a detailed limitation section. The negative result for the two-step pipeline and the multi-perspective difficulty analysis are useful for practitioners. However, the headline claim of 'comparable quality' needs to be scoped to the specific one-step models tested and supported with effect-size or equivalence evidence, since the psychometric comparisons rely on small, pre-filtered samples and non-significant test results.","major_comments":[{"comment":"The claim that 'current models can generate MCQs of comparable quality to human-authored ones' is not supported for the two-step pipeline and is internally inconsistent. Section 4.5.4/Table 5 shows Ptt5-v2+Gemma-2 achieves 25/45 (55.6%) answerable items versus 33/45 (73.3%) for human-authored and both one-step LLMs, which the paper itself calls 'substantially lower.' Aggregated across all three pipelines, the LLM acceptance rate is 91/135 (67.4%), below the human 73.3%. Please revise the abstract, summary of findings, and conclusions to scope the comparability claim to the GPT-4o and Gemma-2 one-step pipelines.","section":"Abstract, Section 4.6, Section 7"},{"comment":"The psychometric comparison is conducted only on the 124 MCQs that survived expert filtering (well-formed, clear, answerable), so the conclusion in Section 5.7 that generated MCQs are 'generally comparable' applies only to items that already passed a quality gate. Because the two-step pipeline produced many more unanswerable items (67.6% answerable vs. 82.5% for human), the filter disproportionately removes low-quality generated items, potentially biasing the comparison toward comparability. The authors should either add an analysis of the unfiltered output (e.g., treating unanswerable items as incorrect) or explicitly state in the conclusions that the psychometric comparability holds only for expert-validated items.","section":"Section 5.2"},{"comment":"The equivalence claim for difficulty and discrimination rests on non-significant Kruskal-Wallis tests (H=1.249, p=0.741 for difficulty; H=0.854, p>0.05 for discrimination) with small group sizes (33 to 25 items per provenance). Absence of a statistically significant difference is not evidence of equivalence. Please report effect sizes (e.g., epsilon-squared), confidence intervals for group means, and ideally an equivalence test with pre-specified bounds, so readers can judge whether 'comparable' is actually supported.","section":"Section 5.5"},{"comment":"The claim in Section 6.4 that models can 'effectively assign difficulty values' is only supported by statistically significant differences for GPT-4o and Ptt5-v2+Gemma-2; for Gemma-2 the differences are not significant for either expert (p=0.2502) or student (p=0.2475) perspectives. This mixed result should be reported as such rather than being subsumed into a general statement about all models.","section":"Section 6.2, Table 9"}],"minor_comments":[{"comment":"The typo 'plausability' appears in Section 4.4, Figure 4, and elsewhere; it should be 'plausibility.'","section":"Throughout"},{"comment":"The text says 'as indicated by the one-way ANOVA test (H=0.854, p>0.05)'; H is the Kruskal-Wallis statistic, so the test name should be corrected for consistency with the difficulty test.","section":"Section 5.5"},{"comment":"The 'Form ID' labels in Table 6 appear to be the same as those used for the expert review forms in Table 1; please clarify whether these are the same forms converted to paper sheets or a separate set of forms.","section":"Table 6"},{"comment":"The provenance codes in the panel labels (G7, GE12, P25) are not defined; explain the naming convention so readers can map them to the provenances in Table 8.","section":"Figure 8"},{"comment":"The semantic-similarity features mention the Serafim encoder (reference [63]) but do not specify the similarity measure (presumably cosine) or how the averages are computed over option pairs; please add these details for reproducibility.","section":"Section 6.3.1"},{"comment":"References [4] and [16] are the same Alhazmi et al. paper and should be consolidated to avoid duplicate entries.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a solid empirical paper with transparent reporting and a useful real-classroom evaluation. The main issue is the overbroad headline claim and the inferential gap in the comparability argument; both are fixable with scoped wording and additional statistical support. The paper is within scope for the journal and the authors have been appropriately cautious in the limitations section. I would encourage resubmission after the requested revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is one of the few studies that actually takes LLM-generated MCQs into a classroom: 284 students, 124 expert-validated items, CTT difficulty and discrimination, plus a 3-rule distractor analysis. That is real evidence, and it is rare in this literature. Second, the headline \"current models can generate MCQs of comparable quality\" is too broad. The two-step Ptt5-v2+Gemma-2 pipeline clearly underperforms (55.6% vs 73.3% answerable), and the paper itself says so; the comparable-quality claim only holds for the one-step GPT-4o and Gemma-2 pipelines.\n\nWhat is actually new: first evaluation of LLM-generated MCQs in European Portuguese with narrative-element and difficulty control, validated by expert review and student responses. The design is careful: three provenances, balanced forms, majority-vote expert ratings, clear RQs. The psychometric analysis using CTT on real responses is a genuine step up from the usual n-gram or expert-only evaluations.\n\nWhere the soft spots are. The abstract overclaim is the main one; it is internally contradicted by the paper's own Table 5. That should be fixed by scoping to GPT-4o and Gemma-2. Second, the psychometric comparison in Section 5 is run only on the 124 MCQs that experts deemed well-formed, clear, and answerable. The paper is transparent about this, but it means the 'comparable psychometric properties' conclusion applies only to items that survive expert screening, not to the unfiltered output of the models. Since the two-step model is filtered harder, the comparison is not apples-to-apples across provenances. This is a real limitation, but it does not sink the paper's core accept/reject claim, which is based on the full 45-item sample. Third, the non-significant Kruskal-Wallis/ANOVA results are used to support 'no difference,' but no effect sizes or confidence intervals are reported; with 33-45 items per group, a null result is weak evidence for equivalence. Fourth, no inter-rater reliability is reported for the expert majority votes.\n\nThe citation pattern is fine; the related work is current and relevant. The paper ships no code or data, but the prompt is in an appendix and the evaluation protocol is detailed enough to replicate. If the authors release the item pools and student responses, that would substantially increase the value.\n\nWho this is for: anyone working on multilingual question generation, automated assessment, or practical deployment of LLMs in education for lower-resourced languages. It deserves a serious referee, not a desk reject. For peer review, I'd ask for the abstract rewrite, effect sizes/CIs, a discussion of the filtering bias (and ideally an analysis on unfiltered samples or a correction), and IRR. With those, it would be a solid publication.","headline":"A solid empirical study of LLM-generated Portuguese MCQs with real classroom data; the headline claim needs scoping to one-step models and the psychometric comparison carries a selection-bias caveat.","tokens_in":27689,"tokens_out":2818,"would_cite":true,"duration_ms":30419,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4o and Gemma-2 can generate Portuguese reading-comprehension MCQs that expert reviewers and student-response analysis find comparable in quality to human-authored items.","keywords":["multiple-choice question generation","reading comprehension","European Portuguese","large language models","psychometric evaluation","difficulty control","narrative elements","classroom assessment"],"falsifier":"Run the same Classical Test Theory analysis on the full unfiltered output of each generator—or on a proportional sample that includes rejected items—and check whether the Kruskal-Wallis tests for difficulty and discrimination remain non-significant; if significant differences appear, the comparability finding is an artifact of expert pre-selection.","tokens_in":26552,"feed_emoji":"📝","tokens_out":5900,"duration_ms":63760,"temperature":0.7,"pith_summary":"This paper asks whether current large language models can produce multiple-choice reading-comprehension questions for Portuguese elementary students that are good enough for real classrooms, and answers yes for one-step generation. Expert reviewers found that MCQs generated zero-shot by GPT-4o and Gemma-2 were as acceptable as human-authored items, with both achieving the same 73.3% overall acceptance rate on the strictest quality gate, answerability. A two-step pipeline that first generates the question with a fine-tuned small model and then generates options with Gemini-2 did not: only 55.6% of its items were accepted. The paper also shows that models can control narrative focus (character, feeling, setting, action, causality) and produce difficulty ratings that line up with expert judgment, particularly when difficulty is assigned after the full question is written. The attraction of the result is that it moves the question from 'can models generate text' to 'can models generate usable assessment items' in a language and age group where manual item writing is costly.","feed_headline":"AI-written test questions match teacher quality in Portuguese","feed_subtitle":"Expert review and 284 student responses show GPT-4o and Gemma-2 items are usable; distractors lag behind.","key_machinery":"The carrying mechanism is a dual evaluation pipeline. First, an expert review protocol with majority voting scores each MCQ on well-formedness, narrative alignment, option clarity, answerability, distractor plausibility, and difficulty; answerability—whether the text contains the answer and whether any option matches it—is the strictest gate. Second, a psychometric analysis based on Classical Test Theory uses student responses to compute item difficulty (1 − P), discrimination (D, the difference between top- and bottom-27% performers), distractor engagement, and a three-rule option-quality test. The generation side pairs zero-shot prompting (GPT-4o, Gemma-2) with a two-step modular pipeline (Ptt5-v2 question generator plus Gemma-2 options), and difficulty is annotated either during generation or after the full MCQ exists.","core_discovery":"On the paper's own terms, the central discovery is that zero-shot prompting of current LLMs yields reading-comprehension MCQs in European Portuguese whose quality is indistinguishable from human-authored items by both expert review and psychometrics. Concretely, 33 of 45 (73.3%) LLM-generated items and 33 of 45 (73.3%) human items were accepted as answerable, well-formed, and clearly written; student-response analysis over 284 participants found no statistically significant differences in difficulty (1 − P) or discrimination (D) between human and one-step LLM MCQs. Human-authored items retained an edge in distractor engagement—students selected all three distractors in 57.6% of human items versus 45.5% for GPT-4o and 51.5% for Gemma-2—and in adherence to option-selection rules, so the paper concludes models are approaching, not yet surpassing, human benchmarks. The two-step method, by contrast, produced items that were significantly less answerable (67.6%) and less discriminative, and the authors attribute this to a bottleneck in the fine-tuned question-generation module.","pith_inferences":["Because the equality of acceptance rates (73.3%) holds on a small sample of 45 items per provenance, the 'comparable quality' conclusion is sensitive to sample size; a larger head-to-head could reveal differences the current study lacks power to detect.","The paper's psychometric comparison ignores expert-rejected items, and since rejection was most frequent for the two-step method, the conclusion that the two-step pipeline is psychometrically inferior is actually stated more cautiously in the paper than the data would allow.","If model-assigned difficulty correlates better with experts than with students, then using LLM difficulty scores to construct homogeneous tests could yield tests that feel calibrated to teachers but not to actual student performance; a test-construction experiment would settle this.","The same evaluation scaffold—expert review gates plus Classical Test Theory indices—could be reused to benchmark MCQ generators in other morphologically rich languages."],"forward_implications":["Teachers and educational platforms can treat one-step LLM output as a draft pool that still requires expert review, since roughly a quarter of items fail answerability.","Two-step generation, at least with a small fine-tuned first-stage model, should be avoided for MCQ production.","Difficulty annotations requested after the full MCQ is generated align better with expert perception than annotations made during generation, giving a practical recipe for calibration.","Narrative control works well enough that generated items can be targeted at specific curriculum elements such as character, feeling, or causality.","Human distractors remain more engaging, so research on distractor design—not just correctness—is where the next quality gain lies."],"supporting_citations":[{"why":"Supplies the five narrative elements (character, feeling, setting, action, causality) used as generation-control targets.","marker":"[14]"},{"why":"Provides the three-rule option-quality test and the prompt-design approach for evaluating ChatGPT-generated MCQs.","marker":"[30]"},{"why":"Supplies the Classical Test Theory framework used to estimate item difficulty and discrimination from student responses.","marker":"[61]"},{"why":"Introduces the ptt5-v2 model fine-tuned for the first step of the two-step generation method.","marker":"[57]"},{"why":"Provides the European Portuguese machine-translated FairytaleQA data used to fine-tune the question-generation model.","marker":"[58]"},{"why":"Offers the Portuguese sentence encoder used to compute semantic-similarity features for difficulty analysis.","marker":"[63]"},{"why":"Establishes the principle that MCQs need thorough validation before administration, framing the answerability findings.","marker":"[60]"}],"fun_headline_variants":["AI quiz questions match human quality in Portuguese","Portuguese MCQs: AI matches humans, distractors lag","LLM-generated tests pass human bar in Portuguese","Portuguese reading tests: AI equals teachers, distractor gap","AI makes Portuguese MCQs as good as teachers'"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The psychometric comparison is run only on the 124 MCQs that experts had already cleared as well-formed, clear, and answerable, which disproportionately removes the lowest-quality generated items before difficulty and discrimination are measured.","fun_headline_variants_meta":{"raw":{"variants":["AI quiz questions match human quality in Portuguese","Portuguese MCQs: AI matches humans, distractors lag","LLM-generated tests pass human bar in Portuguese","Portuguese reading tests: AI equals teachers, distractor gap","AI makes Portuguese MCQs as good as teachers'"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000127,"raw_usage":{"total_tokens":1137,"prompt_tokens":987,"completion_tokens":150,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":75}},"tokens_in":603,"tokens_out":150,"duration_ms":2710,"temperature":1.0,"reasoning_tokens":75,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:51:51.373612+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same Classical Test Theory analysis on the full unfiltered output of each generator—or on a proportional sample that includes rejected items—and check whether the Kruskal-Wallis tests for difficulty and discrimination remain non-significant; if significant differences appear, the comparability finding is an artifact of expert pre-selection.","supporting_citations":[{"cited_title":"ERIC, ??? (1986)","cited_arxiv_id":null,"evidence_quote":"Supplies the Classical Test Theory framework used to estimate item difficulty and discrimination from student responses."},{"cited_title":"In: Brazilian Conference on Intelligent Systems, pp","cited_arxiv_id":null,"evidence_quote":"Introduces the ptt5-v2 model fine-tuned for the first step of the two-step generation method."},{"cited_title":"In: Fer- reira Mello, R., Rummel, N., Jivet, I., Pishtari, G., Ruip´ erez Valiente, J.A","cited_arxiv_id":null,"evidence_quote":"Provides the European Portuguese machine-translated FairytaleQA data used to fine-tune the question-generation model."},{"cited_title":"In: EPIA Conference on Artificial Intelligence, pp","cited_arxiv_id":null,"evidence_quote":"Offers the Portuguese sentence encoder used to compute semantic-similarity features for difficulty analysis."},{"cited_title":"Rout- ledge, ??? (2004)","cited_arxiv_id":null,"evidence_quote":"Establishes the principle that MCQs need thorough validation before administration, framing the answerability findings."}],"review_version":1}