{"id":"3377f406-51fe-43ff-bfa5-236732e8f92f","arxiv_id":"1909.01958","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Aristo is the first system to score above 90 percent on the non-diagram multiple-choice portion of the Grade 8 New York Regents Science Exam, using an ensemble dominated by fine-tuned BERT and RoBERTa models.","lead":"Aristo, an ensemble of eight question-answering solvers built on BERT and RoBERTa, scores over 91 percent on the non-diagram multiple-choice part of the New York Grade 8 science exam. A smart generalist might read this as evidence that modern language models can pass a standardized science test in a limited multiple-choice format, a benchmark milestone for AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unseen-questions claim depends on pretraining contamination, which the paper does not rule out; 2017–19 'unavailable at start' does not control for BERT/RoBERTa web-scale pretraining.","rationale":"Read in good faith, the paper's central claim is an empirical milestone: Aristo's 91.6% on Grade 8 Regents NDMC, with similar scores on 2017–19 exams, demonstrating generalization. The paper provides several credible supporting signals: results are on exam-level partitions (Table 1), answer-only accuracy is low (Figure 5), adversarial additions cause only about a 10% drop (Figure 6), and zero-shot probes show some systematic semantics. These are real evidence against crude annotation artifacts. However, the claim 'on unseen test questions' is load-bearing because the whole milestone is defined by exceeding 90% on genuinely held-out questions. The paper's only direct evidence for the 2017–19 tests being unseen is that they were unavailable at the project's start and not in the datasets. That is insufficient after the switch to BERT/RoBERTa; public web-scale pretraining corpora can contain those exams. The reader's weakest assumption identified the same broad area (leakage) and the small test-set size; this pass sharpens it to a specific, checkable mechanism: pretraining corpus contamination. I do not regard this as established contamination, only as an unsecured condition. The right verdict is therefore the same CONDITIONAL; if an independent contamination check comes back clean, the central claim stands. If contaminated questions are found, the headline could drop below 90% and the claim would need revision. Hence UNCHANGED with respect to the reader's CONDITIONAL, with partial agreement: our concern is the more precise form of the leakage assumption.","tokens_in":14081,"tokens_out":6481,"duration_ms":69206,"concrete_test":"Obtain or reconstruct the exact text of the 2017–19 Grade 8 Regents NDMC questions and answer options. Run exact and near-duplicate substring/n-gram search (e.g., 10-gram overlap with thresholding) against the pretraining corpora actually used for AristoRoBERTa and AristoBERT: RoBERTa's 160GB mix (Books, Wikipedia, CC-News, OpenWebText, Stories) and BERT's Wikipedia+Books, using public snapshots or membership-inference probes. Count questions whose stem or options appear verbatim or near-verbatim, especially with answer-key context. If the count is two or more, the 'unseen' and '>90%' claims need re-estimation after removing contaminated items; if the count is zero, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim: Aristo is first to exceed 90% on Grade 8 Regents NDMC, with results on unseen test questions and robust across test years. The load-bearing assumption is that the test questions, especially 2017–19, were never seen by the model. The paper states in Experiments and Results that 2017–19 were 'unavailable at the start of the project and were not part of our datasets.' This rules out supervised fine-tuning but not pretraining contamination. Aristo's scores are dominated by AristoRoBERTa/AristoBERT, fine-tuned from BERT/RoBERTa pretrained on web-scale text; RoBERTa used Common Crawl, CC-News, OpenWebText, and the NY Regents exams are public PDFs on nysedregents.org. A 2017–19 exam could appear verbatim in RoBERTa's pretraining data while still being 'unavailable at the start of the project' (2014). No contamination check is reported, so the 'unseen test questions' assertion is unsupported in the exact regime used to demonstrate robustness. The paper's answer-only baseline (Figure 5) does not resolve this: low answer-only accuracy shows options alone carry little signal, not that full question stems were absent from pretraining. Because the headline 91.6% is about 109/119, even two or three contaminated questions can decide whether the '>90%' claim survives.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a system-level overview of Project Aristo, an ensemble of eight solvers for multiple-choice science questions, and reports new results on the New York Regents Science Exams. The main empirical claim is that Aristo achieves 91.6% accuracy on the non-diagram multiple-choice (NDMC) portion of the Grade 8 Regents exam and exceeds 83% on Grade 12 NDMC questions, with the ensemble dominated by large pretrained language models (AristoBERT and AristoRoBERTa). The paper also reports robustness checks on 2017-2019 exams, an answer-only baseline, adversarial answer-option experiments, a manual error analysis, and zero-shot probes of semantic skills such as negation, conjunction, polarity, factuality, and counting. The authors position the result as the first system to exceed 90% on Grade 8 Regents NDMC questions and as evidence that modern NLP methods can achieve mastery on this task.","tokens_in":14354,"tokens_out":4421,"duration_ms":43345,"significance":"If the empirical claims hold, this is a landmark result for standardized-test question answering and a useful external-benchmark validation of progress in large pretrained language models. The paper's strengths include the use of exam years not available in the training data (2017-2019), a sensible answer-only control, adversarial option perturbation with retraining, zero-shot probing without fine-tuning on probe data, and a careful manual failure analysis with concrete examples. These design choices make the result substantially more convincing than a bare accuracy number. The main weaknesses are the small test set underlying the headline accuracy, the lack of confidence intervals, and the absence of a check for contamination of the web-scale pretraining corpora; these issues are local and addressable rather than fundamental.","major_comments":[{"comment":"The claim that the 2017-2019 Regents exams are 'unseen test questions' rests on the statement that these exams were 'unavailable at the start of the project,' but this does not rule out contamination during the web-scale pretraining of BERT and RoBERTa, on which AristoRoBERTa and AristoBERT are based; the Regents exams are public PDFs and could in principle appear in Common Crawl or similar pretraining corpora. Because the robustness result (92.8% on Grade 4, 93.3% on Grade 8 for 2017-2019) is exactly the evidence used to support generalization, the paper should report a contamination check (e.g., exact-match or near-duplicate detection of test questions in the pretraining corpus, or perplexity-based membership inference) before the 'unseen test questions' assertion is accepted.","section":"Experiments and Results (paragraph beginning 'To further check...')"},{"comment":"The headline result, 91.6% on the Grade 8 NDMC test set, is computed on only 119 questions (109 correct), and the paper gives no confidence interval; the exact binomial 95% CI around 91.6% for n=119 spans roughly 85% to 96%, so the data do not, at the conventional 95% level, exclude a score below 90%. I recommend reporting exact counts and binomial confidence intervals for all test-set accuracies, and adjusting the 'more than 90%' claim accordingly.","section":"Table 1 and Figure 4"},{"comment":"The answer-only baseline is a useful control, but it does not address the contamination concern: a model that memorized full question-answer pairs during pretraining would still perform poorly when given only the answer options, because the question body is absent. To support the 'not overfit' claim, the paper should pair this baseline with a test-time analysis of whether the full question text appears in the pretraining data (see previous comment).","section":"Answer Only Performance and Figure 5"}],"minor_comments":[{"comment":"The phrase 'for the fast majority of questions' should read 'for the vast majority of questions'.","section":"Analysis, 'Good Support for Correct Answer'"},{"comment":"The footnote 'ARC (Easy+Challenge) includes Regents 4th & 8th as a subset' is confusing because the totals column then double-counts Regents questions; clarify whether the Regents questions are excluded from the ARC rows or whether the reported totals are disjunct.","section":"Table 1 footnote"},{"comment":"In the flipped polarity example, the text reads 'decreasing increasing' and 'lowering raising', which appears to be a typographical corruption of the intended contrast; please correct these options.","section":"A Score Card for Aristo's Semantic Skills, Polarity example"},{"comment":"The statement 'All but 39 of the 9366 questions are 4-way multiple choice' should specify whether the 39 questions are distributed across train, dev, and test partitions, since this affects the interpretation of the reported random baseline.","section":"Experiments and Results, dataset description"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the pretraining-contamination issue is the most important unresolved point, and it is load-bearing for the paper's headline 'unseen test questions' claim. The paper also does not describe a code or data release, so the empirical claims are not independently checkable from the manuscript. The heavy citation of the authors' own prior work is appropriate for an overview article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Interesting paper to know about, but the headline claim is a bit more polished than the evidence. The milestone is genuine: 91.6% on Grade 8 NDMC, and the zero-shot probes are a nice addition. But the 'mastery' language goes beyond what a 119-question test set can support, and the paper never addresses whether the web-scale pretraining of BERT/RoBERTa already saw these public exams.\n\nThe new contribution is empirical, not architectural. The system is an ensemble of known components—BERT/RoBERTa fine-tuning, IR, PMI, entailment, a qualitative reasoning module—each described in earlier papers. That is fine for an overview, and the paper does not oversell the novelty. What it does well: it includes a sensible answer-only baseline, an adversarial option experiment, a zero-shot probe suite for negation, conjunction, polarity, factuality, and counting, and a careful failure analysis. These controls show the authors were thinking about artifacts and systematic behavior. The 'answer-only' result is a particularly useful sanity check, even if it does not rule out memorization of full question stems.\n\nThe soft spots are real but not fatal. First, the Grade 8 test set has 119 questions; 91.6% means 109/119, and the binomial 95% interval is roughly 85–96%. The claim that the system is 'above 90%' is within noise. Second, the robustness check on 2017–19 exams is reassuring, but 'unavailable at the start of the project' only rules out fine-tuning on them. RoBERTa was pretrained on Common Crawl and other public text, and the Regents exams are public PDFs. The paper does not report a contamination check, so the 'unseen test questions' statement is not fully supported. That is a legitimate concern, though it applies to most LM-based benchmark results and is not unique to this paper. Third, no code or exact splits are released, which hurts reproducibility but is common for an overview.\n\nBottom line: this is a useful, readable snapshot of a real milestone, and the authors are unusually candid about the limits. The central finding—that modern LMs push elementary science QA into the 90s—holds up as a trend, even if the exact 'first over 90%' phrasing is fragile. I would send it to peer review, with the expectation that the authors add error bars and a contamination analysis. It deserves a serious referee, but not a pass as-is.","headline":"A genuine empirical milestone, but the 'mastery' claim overreaches a 119-question test set and the paper never addresses possible pretraining contamination.","tokens_in":14930,"tokens_out":2546,"would_cite":true,"duration_ms":25828,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Aristo system is the first to score above 90 percent on unseen multiple-choice questions from the New York Grade 8 science Regents exam.","keywords":["Aristo","New York Regents Science Exam","multiple-choice question answering","large-scale language models","BERT","RoBERTa","standardized test benchmark","science question answering"],"falsifier":"Search the pretraining corpora and retrieved-knowledge sources for sentences from the 2017-2019 Regents exams; finding any would void the held-out claim. Separately, compute the exact binomial 95% confidence interval for 109 correct out of 119 questions: because the interval's lower bound falls below 90 percent, the 'above 90 percent' headline is not statistically decisive on this test set, and a decisive test would be to run Aristo on a newly released, never-seen Regents exam and require that accuracy stays above 90 percent.","tokens_in":13911,"feed_emoji":"🧪","tokens_out":7907,"duration_ms":66693,"temperature":0.7,"pith_summary":"The paper reports that Aristo, a question-answering system built from eight solvers and dominated by large-scale language models, achieves 91.6 percent accuracy on the non-diagram multiple-choice (NDMC) portion of the New York Grade 8 Regents Science Exam, on test questions it has not seen. This is the first time any system has surpassed 90 percent on this external benchmark, and the same system also exceeds 83 percent on the Grade 12 NDMC questions. The results hold across different test years, including exams from 2017-2019 that were withheld from development, suggesting the system is not overfit to its training data. The authors argue this demonstrates that modern NLP methods can achieve mastery of this task, a milestone toward systems that can read and reason about science.","feed_headline":"AI breaks 90 percent on 8th-grade science Regents exam","feed_subtitle":"The Aristo system answers unseen multiple-choice questions at a level the state calls 'with distinction'.","key_machinery":"The load-bearing object is the ensemble of eight solvers, with the language-model solvers carrying most of the weight. AristoBERT and AristoRoBERTa frame each question-option pair as a text-classification input of the form [CLS] background [SEP] question [SEP] option [SEP], fine-tune BERT or RoBERTa on a curriculum of reading-comprehension and science datasets, retrieve background knowledge for each option, and ensemble several model variants. The older solvers—information retrieval, pointwise mutual information, tuple-graph inference via integer linear programming, textual entailment combination, and qualitative reasoning—fill gaps the language models miss. The paper argues that the high scores reflect emergent semantic skills such as handling negation, conjunction, and polarity without fine-tuning on those specific probes.","core_discovery":"On its own terms, the paper establishes that a single, unchanged system—Aristo—can answer more than nine out of ten previously unseen multiple-choice science questions from the Grade 8 New York Regents exam, and more than eight out of ten on the Grade 12 exam. The headline figure is 91.6 percent on 119 Grade 8 test questions, with 92.8 percent and 93.3 percent averages on the held-out 2017-2019 Grade 4 and Grade 8 exams respectively. The system's performance is dominated by two language-model solvers, AristoBERT and AristoRoBERTa, which treat each answer option as a classification problem with optional retrieved background knowledge. The paper also shows that an answer-only baseline scores much lower, and that adversarially adding four extra wrong options lowers accuracy by only about ten percent, evidence that the system is reading the questions rather than exploiting answer-option artifacts.","pith_inferences":["The paper does not report a leakage audit, so an independent check of the web-scale pretraining data for 2017-2019 Regents sentences would be needed before treating the generalization claim as settled.","The 91.6 percent figure is based on 119 questions; a 95 percent confidence interval for 109 successes in 119 trials has a lower bound below 90 percent, so the claim of mastery at the 90 percent threshold would be better tested on a larger held-out set.","The counting probe (6 percent, below chance) is a sharp boundary marker: a future system that can both exceed 90 percent on the Regents NDMC questions and handle simple counting would be strong evidence of more general reasoning."],"forward_implications":["If correct, an AI system can now perform at a level that New York State would count as meeting the standards with distinction on the multiple-choice portion of the 8th grade science exam, giving the field an external, human-comparable benchmark.","The same architecture transfers across grade levels (4th, 8th, 12th) and to the broader ARC science dataset without being re-tuned, suggesting the method is not specific to one exam's style.","The answer-only baseline and adversarial-option experiments indicate that high scores are not an artifact of superficial cues in the answer options, which strengthens the case that the system is engaging with question content.","The system's failure analysis points to concrete next targets: combining diverse evidence, reading-comprehension questions, meta-questions, and counting, which remain far below human performance."],"supporting_citations":[{"why":"Supplies the earlier 59.3 percent baseline on the 8th grade exam that Aristo must beat.","marker":"Schoenick et al., 2016"},{"why":"Provides BERT, the base model for the AristoBERT solver that drives the high scores.","marker":"Devlin et al., 2018"},{"why":"Provides RoBERTa, the base model for the AristoRoBERTa solver, the single strongest component.","marker":"Liu et al., 2019"},{"why":"Contributes the curriculum fine-tuning approach used to train the language-model solvers.","marker":"Sun et al., 2019"},{"why":"Supplies the ARC dataset and challenge that make up a large part of the training and evaluation data.","marker":"Clark et al., 2018"},{"why":"Defines the TupleInference solver that reasons over open information extraction tuples with integer linear programming.","marker":"Khot, Sabharwal, and Clark, 2017"},{"why":"Provides Multee, the entailment-based solver that combines textual entailment systems for multi-hop reasoning.","marker":"Trivedi et al., 2019"},{"why":"Supplies the qualitative reasoning solver and the QuaRTz dataset used to probe polarity understanding.","marker":"Tafjord et al., 2019"}],"fun_headline_variants":["AI tops 90% on 8th-grade science Regents exam","Aristo AI breaks 90% on grade 8 science exams","From F to A: AI masters NY Regents science exam","AI scores 91.6% on unseen 8th-grade science questions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim assumes the 2017-2019 Regents exams were completely absent from every training corpus and pretraining source; if those questions leaked in, the reported 90-plus percent scores would measure memorization rather than generalization.","fun_headline_variants_meta":{"raw":{"variants":["AI tops 90% on 8th-grade science Regents exam","Aristo AI breaks 90% on grade 8 science exams","From F to A: AI masters NY Regents science exam","AI scores 91.6% on unseen 8th-grade science questions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1211,"prompt_tokens":932,"completion_tokens":279,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":202}},"tokens_in":548,"tokens_out":279,"duration_ms":3067,"temperature":1.0,"reasoning_tokens":202,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:03:06.886154+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Search the pretraining corpora and retrieved-knowledge sources for sentences from the 2017-2019 Regents exams; finding any would void the held-out claim. Separately, compute the exact binomial 95% confidence interval for 109 correct out of 119 questions: because the interval's lower bound falls below 90 percent, the 'above 90 percent' headline is not statistically decisive on this test set, and a decisive test would be to run Aristo on a newly released, never-seen Regents exam and require that accuracy stays above 90 percent.","supporting_citations":[{"cited_title":"F.; Tafjord, O.; Turney, P","cited_arxiv_id":null,"evidence_quote":"Supplies the earlier 59.3 percent baseline on the 8th grade exam that Aristo must beat."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides BERT, the base model for the AristoBERT solver that drives the high scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the curriculum fine-tuning approach used to train the language-model solvers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the TupleInference solver that reasons over open information extraction tuples with integer linear programming."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides Multee, the entailment-based solver that combines textual entailment systems for multi-hop reasoning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the qualitative reasoning solver and the QuaRTz dataset used to probe polarity understanding."}],"review_version":1}