Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The Aristo system is the first to score above 90 percent on unseen multiple-choice questions from the New York Grade 8 science Regents exam.

desk verdict A genuine empirical milestone, but the 'mastery' claim overreaches a 119-question test set and the paper never addresses possible pretraining contamination. read the letter →

arxiv 1909.01958 v3 pith:R2NGLJMH submitted 2019-09-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords AristoNewYorkRegentsScienceExammultiple-choicequestionansweringlarge-scalelanguagemodelsBERTRoBERTastandardizedtestbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports that Aristo, a question-answering system built from eight solvers and dominated by large-scale language models, achieves 91.6 percent accuracy on the non-diagram multiple-choice (NDMC) portion of the New York Grade 8 Regents Science Exam, on test questions it has not seen. This is the first time any system has surpassed 90 percent on this external benchmark, and the same system also exceeds 83 percent on the Grade 12 NDMC questions. The results hold across different test years, including exams from 2017-2019 that were withheld from development, suggesting the system is not overfit to its training data. The authors argue this demonstrates that modern NLP methods can achieve mastery of this task, a milestone toward systems that can read and reason about science.

What carries the argument

The load-bearing object is the ensemble of eight solvers, with the language-model solvers carrying most of the weight. AristoBERT and AristoRoBERTa frame each question-option pair as a text-classification input of the form [CLS] background [SEP] question [SEP] option [SEP], fine-tune BERT or RoBERTa on a curriculum of reading-comprehension and science datasets, retrieve background knowledge for each option, and ensemble several model variants. The older solvers—information retrieval, pointwise mutual information, tuple-graph inference via integer linear programming, textual entailment combination, and qualitative reasoning—fill gaps the language models miss. The paper argues that the high scores reflect emergent semantic skills such as handling negation, conjunction, and polarity without fine-tuning on those specific probes.

What would settle it

Search the pretraining corpora and retrieved-knowledge sources for sentences from the 2017-2019 Regents exams; finding any would void the held-out claim. Separately, compute the exact binomial 95% confidence interval for 109 correct out of 119 questions: because the interval's lower bound falls below 90 percent, the 'above 90 percent' headline is not statistically decisive on this test set, and a decisive test would be to run Aristo on a newly released, never-seen Regents exam and require that accuracy stays above 90 percent.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that a single, unchanged system—Aristo—can answer more than nine out of ten previously unseen multiple-choice science questions from the Grade 8 New York Regents exam, and more than eight out of ten on the Grade 12 exam. The headline figure is 91.6 percent on 119 Grade 8 test questions, with 92.8 percent and 93.3 percent averages on the held-out 2017-2019 Grade 4 and Grade 8 exams respectively. The system's performance is dominated by two language-model solvers, AristoBERT and AristoRoBERTa, which treat each answer option as a classification problem with optional retrieved background knowledge. The paper also shows that an answer-only baseline scores much lower, and that adversarially adding four extra wrong options lowers accuracy by only about ten percent, evidence that the system is reading the questions rather than exploiting answer-option artifacts.

Load-bearing premise

The central claim assumes the 2017-2019 Regents exams were completely absent from every training corpus and pretraining source; if those questions leaked in, the reported 90-plus percent scores would measure memorization rather than generalization.

Editorial extensions

If this is right

  • If correct, an AI system can now perform at a level that New York State would count as meeting the standards with distinction on the multiple-choice portion of the 8th grade science exam, giving the field an external, human-comparable benchmark.
  • The same architecture transfers across grade levels (4th, 8th, 12th) and to the broader ARC science dataset without being re-tuned, suggesting the method is not specific to one exam's style.
  • The answer-only baseline and adversarial-option experiments indicate that high scores are not an artifact of superficial cues in the answer options, which strengthens the case that the system is engaging with question content.
  • The system's failure analysis points to concrete next targets: combining diverse evidence, reading-comprehension questions, meta-questions, and counting, which remain far below human performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not report a leakage audit, so an independent check of the web-scale pretraining data for 2017-2019 Regents sentences would be needed before treating the generalization claim as settled.
  • The 91.6 percent figure is based on 119 questions; a 95 percent confidence interval for 109 successes in 119 trials has a lower bound below 90 percent, so the claim of mastery at the 90 percent threshold would be better tested on a larger held-out set.
  • The counting probe (6 percent, below chance) is a sharp boundary marker: a future system that can both exceed 90 percent on the Regents NDMC questions and handle simple counting would be strong evidence of more general reasoning.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper presents a system-level overview of Project Aristo, an ensemble of eight solvers for multiple-choice science questions, and reports new results on the New York Regents Science Exams. The main empirical claim is that Aristo achieves 91.6% accuracy on the non-diagram multiple-choice (NDMC) portion of the Grade 8 Regents exam and exceeds 83% on Grade 12 NDMC questions, with the ensemble dominated by large pretrained language models (AristoBERT and AristoRoBERTa). The paper also reports robustness checks on 2017-2019 exams, an answer-only baseline, adversarial answer-option experiments, a manual error analysis, and zero-shot probes of semantic skills such as negation, conjunction, polarity, factuality, and counting. The authors position the result as the first system to exceed 90% on Grade 8 Regents NDMC questions and as evidence that modern NLP methods can achieve mastery on this task.

Significance. If the empirical claims hold, this is a landmark result for standardized-test question answering and a useful external-benchmark validation of progress in large pretrained language models. The paper's strengths include the use of exam years not available in the training data (2017-2019), a sensible answer-only control, adversarial option perturbation with retraining, zero-shot probing without fine-tuning on probe data, and a careful manual failure analysis with concrete examples. These design choices make the result substantially more convincing than a bare accuracy number. The main weaknesses are the small test set underlying the headline accuracy, the lack of confidence intervals, and the absence of a check for contamination of the web-scale pretraining corpora; these issues are local and addressable rather than fundamental.

major comments (3)
  1. [Experiments and Results (paragraph beginning 'To further check...')] The claim that the 2017-2019 Regents exams are 'unseen test questions' rests on the statement that these exams were 'unavailable at the start of the project,' but this does not rule out contamination during the web-scale pretraining of BERT and RoBERTa, on which AristoRoBERTa and AristoBERT are based; the Regents exams are public PDFs and could in principle appear in Common Crawl or similar pretraining corpora. Because the robustness result (92.8% on Grade 4, 93.3% on Grade 8 for 2017-2019) is exactly the evidence used to support generalization, the paper should report a contamination check (e.g., exact-match or near-duplicate detection of test questions in the pretraining corpus, or perplexity-based membership inference) before the 'unseen test questions' assertion is accepted.
  2. [Table 1 and Figure 4] The headline result, 91.6% on the Grade 8 NDMC test set, is computed on only 119 questions (109 correct), and the paper gives no confidence interval; the exact binomial 95% CI around 91.6% for n=119 spans roughly 85% to 96%, so the data do not, at the conventional 95% level, exclude a score below 90%. I recommend reporting exact counts and binomial confidence intervals for all test-set accuracies, and adjusting the 'more than 90%' claim accordingly.
  3. [Answer Only Performance and Figure 5] The answer-only baseline is a useful control, but it does not address the contamination concern: a model that memorized full question-answer pairs during pretraining would still perform poorly when given only the answer options, because the question body is absent. To support the 'not overfit' claim, the paper should pair this baseline with a test-time analysis of whether the full question text appears in the pretraining data (see previous comment).
minor comments (4)
  1. [Analysis, 'Good Support for Correct Answer'] The phrase 'for the fast majority of questions' should read 'for the vast majority of questions'.
  2. [Table 1 footnote] The footnote 'ARC (Easy+Challenge) includes Regents 4th & 8th as a subset' is confusing because the totals column then double-counts Regents questions; clarify whether the Regents questions are excluded from the ARC rows or whether the reported totals are disjunct.
  3. [A Score Card for Aristo's Semantic Skills, Polarity example] In the flipped polarity example, the text reads 'decreasing increasing' and 'lowering raising', which appears to be a typographical corruption of the intended contrast; please correct these options.
  4. [Experiments and Results, dataset description] The statement 'All but 39 of the 9366 questions are 4-way multiple choice' should specify whether the 39 questions are distributed across train, dev, and test partitions, since this affects the interpretation of the reported random baseline.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: Aristo's claims are empirical evaluation results on held-out exams and zero-shot probes, not derivations from fitted inputs.

full rationale

The paper contains no formal derivation chain to audit: every central claim (91.6% on Grade 8 NDMC, 83% on Grade 12, robustness across 2017-19 exams, answer-only baseline, adversarial options, semantic-skill scorecard) is an experimental measurement on a held-out test partition or a zero-shot probe. The system is trained on train splits (RACE, OpenbookQA, ARC, Regents train) and evaluated on disjoint test splits, with the 2017-19 Regents exams stated as 'unavailable at the start of the project and were not part of our datasets.' The answer-only control is a sensible baseline and does not define the headline result. The semantic probes are explicitly zero-shot and are used as interpretive evidence, not as inputs that force the exam scores. Self-citations to earlier Aristo component papers describe the architecture, but the current results are independently evaluated against external benchmark questions; no uniqueness theorem or equivalent is imported to make a choice forced. The only substantive threat to the 'unseen test questions' claim is possible contamination of BERT/RoBERTa web-scale pretraining with public Regents PDFs, which the paper does not test; that is an external validity and leakage concern, not a circular reduction of outputs to inputs. Accordingly, no step meets the evidentiary standard for circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the integrity of the evaluation setup and the validity of the NDMC proxy. There are no fitted physical constants or invented entities.

assumptions (3)
  • domain assumption The non-diagram multiple-choice (NDMC) subset is a valid proxy for the science exam benchmark.
    The paper explicitly restricts evaluation to NDMC questions and omits diagram and direct-answer questions; this limits the generality of the 'mastery' claim.
  • domain assumption The hidden 2017-2019 Regents exams were not involved in system development, so the robustness scores are uncontaminated.
    The paper asserts these years were 'unavailable at the start of the project', but provides no formal leakage check.
  • domain assumption Fine-tuning BERT/RoBERTa via a curriculum of RACE, OpenbookQA, ARC, and Regents train partitions transfers to the test partitions.
    The method is described in the Large-Scale Language Models section; transfer learning assumptions are standard but unproven here.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project." pith.science (2026). https://pith.science/paper/R2NGLJMH

@misc{pith2026190901958,
  author       = {Pith},
  title        = {Pith review of: From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/R2NGLJMH}},
  note         = {Machine review of arXiv:1909.01958}
}
read the original abstract

AI has achieved remarkable mastery over games such as Chess, Go, and Poker, and even Jeopardy, but the rich variety of standardized exams has remained a landmark challenge. Even in 2016, the best AI system achieved merely 59.3% on an 8th Grade science exam challenge. This paper reports unprecedented success on the Grade 8 New York Regents Science Exam, where for the first time a system scores more than 90% on the exam's non-diagram, multiple choice (NDMC) questions. In addition, our Aristo system, building upon the success of recent language models, exceeded 83% on the corresponding Grade 12 Science Exam NDMC questions. The results, on unseen test questions, are robust across different test years and different variations of this kind of test. They demonstrate that modern NLP methods can result in mastery on this task. While not a full solution to general question-answering (the questions are multiple choice, and the domain is restricted to 8th Grade science), it represents a significant milestone for the field.

Figures

Figures reproduced from arXiv: 1909.01958 by the authors.

Figure 1
Figure 1. Progress over time of Aristo’s scores on Regents [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. A simplified picture of Aristo’s architecture. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. The TupleInference Solver retrieves tuples rele [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: The results of each of the Aristo solvers, as well as the overall Aristo system, on each of the test sets. Most notably, [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Scores when looking at the answer options only, [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Aristo’s scores drop a small amount (average 10%) [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 8
Figure 8. Figure 8: Aristo, with no fine-tuning, passes probes for all [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Teaching Smaller Language Models To Generalise To Unseen Compositional Questions (Full Thesis)

    cs.CL 2024-11 conditional novelty 7.0 of 10

    Smaller language models can generalize to unseen compositional questions when trained and evaluated with retrieval-augmented contexts, and combining Wikipedia retrieval with LLM-generated rationales improves accuracy.

Reference graph

Works this paper leans on

63 extracted references · 56 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...

  2. [2]

    Allen, P. 2012. Idea Man: A memoir by the C ofounder of M icrosoft . Penguin

  3. [3]

    Amini, A.; Gabriel, S.; Lin, P.; Koncel-Kedziorski, R.; Choi, Y.; and Hajishirzi, H. 2019. MathQA : T owards I nterpretable M ath W ord P roblem S olving with O peration- B ased F ormalisms. In NAACL-HLT

  4. [4]

    J.; Soderland, S.; Broadhead, M.; and Etzioni, O

    Banko, M.; Cafarella, M. J.; Soderland, S.; Broadhead, M.; and Etzioni, O. 2007. O pen I nformation E xtraction from the W eb. In IJCAI

  5. [5]

    J., and Levesque, H

    Brachman, R. J., and Levesque, H. J. 1985. Readings in K nowledge R epresentation . Morgan Kaufmann Publishers Inc

  6. [6]

    Brachman, R.; Gunning, D.; Bringsjord, S.; Genesereth, M.; Hirschman, L.; and Ferro, L. 2005. Selected G rand C hallenges in C ognitive S cience. Technical report, MITRE Technical Report 05-1218. Bedford MA: The MITRE Corporation

  7. [7]

    Bringsjord, S., and Schimanski, B. 2003. What is A rtificial I ntelligence? P sychometric AI as an A nswer. In IJCAI , 887--893. Citeseer

  8. [8]

    Cheng, G.; Zhu, W.; Wang, Z.; Chen, J.; and Qu, Y. 2016. Taking up the G aokao C hallenge: A n I nformation R etrieval A pproach. In IJCAI , 2479--2485

Show all 63 references
  1. [9]

    W., and Hanks, P

    Church, K. W., and Hanks, P. 1989. Word A ssociation N orms, M utual I nformation and L exicography. In 27th ACL , 76--83

  2. [10]

    Clark, P., and Etzioni, O. 2016. My C omputer is an H onor S tudent - B ut how I ntelligent is it? S tandardized T ests as a M easure of AI . AI Magazine 37(1):5--12

  3. [11]

    Clark, P.; Balasubramanian, N.; Bhakthavatsalam, S.; Humphreys, K.; Kinkead, J.; Sabharwal, A.; and Tafjord, O. 2014. Automatic C onstruction of I nference- S upporting K nowledge B ases. In 4th Workshop on Automated Knowledge Base Construction (AKBC)

  4. [12]

    D.; and Khashabi, D

    Clark, P.; Etzioni, O.; Khot, T.; Sabharwal, A.; Tafjord, O.; Turney, P. D.; and Khashabi, D. 2016. Combining R etrieval, S tatistics, and I nference to A nswer E lementary S cience Q uestions. In AAAI , 2580--2586

  5. [13]

    Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018. Think you have S olved Q uestion A nswering? T ry ARC , the AI2 R easoning C hallenge. ArXiv abs/1803.05457

  6. [14]

    Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019. B ool Q : E xploring the S urprising D ifficulty of N atural Y es/ N o Q uestions. In NAACL-HLT

  7. [15]

    Clark, P.; Tafjord, O.; and Richardson, K. 2020. Transformers as S oft R easoners over L anguage. In IJCAI

  8. [16]

    Davis, E. 2014. The L imitations of S tandardized S cience T ests as B enchmarks for A rtificial I ntelligence R esearch. ArXiv abs/1411.1629

  9. [17]

    Davis, E. 2016. How to W rite S cience Q uestions that are E asy for P eople and H ard for C omputers. AI Magazine 37:13--22

  10. [18]

    Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018. BERT : P re-training of D eep B idirectional T ransformers for L anguage U nderstanding. In NAACL

  11. [19]

    A.; Lally, A.; Murdock, J

    Ferrucci, D.; Brown, E.; Chu-Carroll, J.; Fan, J.; Gondek, D.; Kalyanpur, A. A.; Lally, A.; Murdock, J. W.; Nyberg, E.; Prager, J.; et al. 2010. Building W atson: A n O verview of the DeepQA P roject. AI magazine 31(3):59--79

  12. [20]

    S.; Allen, P

    Friedland, N. S.; Allen, P. G.; Matthews, G.; Witbrock, M.; Baxter, D.; Curtis, J.; Shepard, B.; Miraglia, P.; Angele, J.; Staab, S.; et al. 2004. Project Halo : T owards a D igital A ristotle. AI magazine 25(4):29--29

  13. [21]

    Fujita, A.; Kameda, A.; Kawazoe, A.; and Miyao, Y. 2014. Overview of Todai R obot P roject and E valuation F ramework of its NLP -based P roblem S olving. In LREC

  14. [22]

    R., and Nilsson, N

    Genesereth, M. R., and Nilsson, N. J. 2012. Logical F oundations of A rtificial I ntelligence . Morgan Kaufmann

  15. [23]

    Guo, S.; Zeng, X.; He, S.; Liu, K.; and Zhao, J. 2017. Which is the E ffective W ay for G aokao: I nformation R etrieval or N eural N etworks? In EACL'17 , 111--120

  16. [24]

    R.; and Smith, N

    Gururangan, S.; Swayamdipta, S.; Levy, O.; Schwartz, R.; Bowman, S. R.; and Smith, N. A. 2018. Annotation A rtifacts in N atural L anguage I nference D ata. In NAACL

  17. [25]

    L.; Herrasti, A.; and Joshi, V

    Hopkins, M.; Petrescu-Prahova, C.; Levin, R.; Bras, R. L.; Herrasti, A.; and Joshi, V. 2017. Beyond S entential S emantic P arsing: T ackling the M ath SAT with a C ascade of T ree T ransducers. In EMNLP

  18. [26]

    S.; and Zettlemoyer, L

    Joshi, M.; Choi, E.; Weld, D. S.; and Zettlemoyer, L. 2017. TriviaQA : A L arge S cale D istantly S upervised C hallenge D ataset for R eading C omprehension. In ACL'17 . Vancouver, Canada: Association for Computational Linguistics

  19. [27]

    Khashabi, D.; Khot, T.; Sabharwal, A.; Clark, P.; Etzioni, O.; and Roth, D. 2016. Question A nswering via I nteger P rogramming over S emi- S tructured K nowledge. In IJCAI

  20. [28]

    Khashabi, D.; Khot, T.; Sabharwal, A.; and Roth, D. 2018. Question A nswering as G lobal R easoning over S emantic A bstractions. In AAAI

  21. [29]

    Khot, T.; Balasubramanian, N.; Gribkoff, E.; Sabharwal, A.; Clark, P.; and Etzioni, O. 2015. Exploring M arkov L ogic N etworks for Q uestion A nswering. In EMNLP

  22. [30]

    Khot, T.; Sabharwal, A.; and Clark, P. F. 2017. Answering C omplex Q uestions using O pen I nformation E xtraction. In ACL

  23. [31]

    Krishnamurthy, J.; Tafjord, O.; and Kembhavi, A. 2016. Semantic P arsing to P robabilistic P rograms for S ituated Q uestion A nswering. In EMNLP

  24. [32]

    Lai, G.; Xie, Q.; Liu, H.; Yang, Y.; and Hovy, E. 2017. RACE : L arge-scale R eading C omprehension D ataset from E xaminations. In EMNLP

  25. [33]

    M., and Baroni, M

    Lake, B. M., and Baroni, M. 2017. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In ICML

  26. [34]

    K., and Dumais, S

    Landauer, T. K., and Dumais, S. T. 1997. A S olution to P lato's problem: T he L atent S emantic A nalysis T heory of A cquisition, I nduction, and R epresentation of K nowledge. Psychological review 104(2):211

  27. [35]

    H.; McDermott, J.; Simon, D

    Larkin, J. H.; McDermott, J.; Simon, D. P.; and Simon, H. A. 1980. Models of C ompetence in S olving P hysics P roblems. Cognitive Science 4:317--345

  28. [36]

    Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. RoBERTa : A R obustly O ptimized BERT P retraining A pproach. arXiv preprint arXiv:1907.11692

  29. [37]

    Matsuzaki, T.; Iwane, H.; Anai, H.; and Arai, N. H. 2014. The most U ncreative E xaminee: a F irst S tep toward W ide C overage N atural L anguage M ath P roblem S olving. In AAAI'14

  30. [38]

    Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018. Can a S uit of A rmor C onduct E lectricity? A N ew D ataset for O pen B ook Q uestion A nswering. In EMNLP

  31. [39]

    M.; Dorr, B

    Mohammad, S. M.; Dorr, B. J.; Hirst, G.; and Turney, P. D. 2013. Computing L exical C ontrast. Computational Linguistics 39(3):555--590

  32. [40]

    Mott, N. 2016. Todai R obot G ives U p on G etting I nto the U niversity of T okyo. Inverse . (https://www.inverse.com/article/23761-todai-robot-gives-up-university-tokyo)

  33. [41]

    E.; Neumann, M.; Iyyer, M.; Gardner, M

    Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M. P.; Clark, C.; Lee, K.; and Zettlemoyer, L. S. 2018. Deep C ontextualized W ord R epresentations. In NAACL

  34. [42]

    Piatetsky-Shapiro, G.; Djeraba, C.; Getoor, L.; Grossman, R.; Feldman, R.; and Zaki, M. 2006. What are the G rand C hallenges for D ata M ining?: KDD -2006 P anel R eport. ACM SIGKDD Explorations Newsletter 8(2):70--77

  35. [43]

    E.; Pavlick, E.; White, A

    Poliak, A.; Haldar, A.; Rudinger, R.; Hu, J. E.; Pavlick, E.; White, A. S.; and Durme, B. V. 2018a. Collecting D iverse N atural L anguage I nference P roblems for S entence R epresentation E valuation. In EMNLP

  36. [44]

    Poliak, A.; Naradowsky, J.; Haldar, A.; Rudinger, R.; and Van Durme, B. 2018b. Hypothesis O nly B aselines in N atural L anguage I nference. In StarSem

  37. [45]

    Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016. SQuAD : 100,000+ Q uestions for M achine C omprehension of T ext. In EMNLP

  38. [46]

    Reddy, R. 1988. Foundations and G rand C hallenges of A rtificial I ntelligence: AAAI P residential A ddress. AI Magazine 9(4)

  39. [47]

    Reddy, R. 2003. Three O pen P roblems in AI . J. ACM 50:83--86

  40. [48]

    S.; and Durme, B

    Rudinger, R.; White, A. S.; and Durme, B. V. 2018. Neural M odels of F actuality. In NAACL-HLT

  41. [49]

    F.; Tafjord, O.; Turney, P

    Schoenick, C.; Clark, P. F.; Tafjord, O.; Turney, P. D.; and Etzioni, O. 2016. Moving beyond the T uring T est with the Allen AI Science Challenge . CACM

  42. [50]

    J.; Hajishirzi, H.; Farhadi, A.; Etzioni, O.; and Malcolm, C

    Seo, M. J.; Hajishirzi, H.; Farhadi, A.; Etzioni, O.; and Malcolm, C. 2015. Solving G eometry P roblems: C ombining T ext and D iagram I nterpretation. In EMNLP

  43. [51]

    J.; Kembhavi, A.; Farhadi, A.; and Hajishirzi, H

    Seo, M. J.; Kembhavi, A.; Farhadi, A.; and Hajishirzi, H. 2016. Bidirectional A ttention F low for M achine C omprehension. ArXiv abs/1611.01603

  44. [52]

    Strickland, E. 2013. Can an AI get into the U niversity of T okyo? IEEE Spectrum 50(9):13--14

  45. [53]

    Sun, K.; Yu, D.; Yu, D.; and Cardie, C. 2019. Improving M achine R eading C omprehension with G eneral R eading S trategies. In NAACL-HLT

  46. [54]

    Tafjord, O.; Gardner, M.; Lin, K.; and Clark, P. 2019. QuaRTz : An O pen- D omain D ataset of Q ualitative R elationship Q uestions. In EMNLP

  47. [55]

    Tainaka, M. 2013. The Todai R obot P roject. NII Today 46. (http://www.nii.ac.jp/userdata/results/pr\_data/NII\_Today/ 60\_en/all.pdf)

  48. [56]

    D.; Grus, J.; Yih, W.-t.; Bosselut, A.; and Clark, P

    Tandon, N.; Mishra, B. D.; Grus, J.; Yih, W.-t.; Bosselut, A.; and Clark, P. 2018. Reasoning about A ctions and S tate C hanges by I njecting C ommonsense K nowledge. In EMNLP

  49. [57]

    Trivedi, H.; Kwon, H.; Khot, T.; Sabharwal, A.; and Balasubramanian, N. 2019. Repurposing E ntailment for M ulti- H op Q uestion A nswering T asks. In NAACL

  50. [58]

    Turney, P. D. 2006. Similarity of S emantic R elations. Computational Linguistics 32(3):379--416

  51. [59]

    Turney, P. D. 2017. Leveraging T erm B anks for A nswering C omplex Q uestions: A C ase for S parse V ectors. arXiv preprint arXiv:1704.03543

  52. [60]

    Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019. GLUE : A M ulti-task B enchmark and A nalysis P latform for N atural L anguage U nderstanding. In ICLR

  53. [61]

    Wang, W.; Yan, M.; and Wu, C. 2018. Multi- G ranularity H ierarchical A ttention F usion N etworks for R eading C omprehension and Q uestion A nswering. In ACL

  54. [62]

    Weston, J.; Bordes, A.; Chopra, S.; and Mikolov, T. 2015. Towards AI - C omplete Q uestion A nswering: A S et of P rerequisite T oy T asks. arXiv 1502.05698

  55. [63]

    Wolfson, T.; Geva, M.; Gupta, A.; Gardner, M.; Goldberg, Y.; Deutch, D.; and Berant, J. 2020. Break I t D own: A Q uestion U nderstanding B enchmark. TACL

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.