REVIEW 4 major objections 5 minor 112 references
HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper introduces HintEval, a unified open-source library for generating and scoring hints for questions, and presents it as the first framework of its kind.
desk verdict A genuinely useful first-of-its-kind toolkit for hint generation and evaluation, but the 'reliable evaluation' headline goes beyond the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing piece is the Dataset object model: a unified schema with Instance, Question, Answer, Hint, Entity, and Metric classes that converts any hint dataset into the same structure. Around this core sit two extensible base classes—Model, which exposes a generate function for answer-aware or answer-agnostic hint production, and Evaluation, which exposes an evaluate function for any metric. The argument works by standardization: once every dataset is in the same schema, every generator can write hints into it and every metric can read from it, which is what lets the framework claim consistent, comparable evaluation.
What would settle it
Take a random sample of hints from TriviaHG's test set, have several human annotators rate each hint on relevance, readability, convergence, familiarity, and answer leakage using the paper's own definitions, and compare the average human ratings with HintEval's automatic scores; if the correlation is near zero or negative for any dimension, that metric does not measure what the framework claims.
Extended reading notes
Core claim
The central claim is that HintEval removes the main obstacle to hint research by putting every dataset, generator, and metric under one interface. The framework defines a Dataset schema in which questions, answers, hints, named entities, and metric scores are first-class objects; it ships with preprocessed versions of TriviaHG, WikiHint, HintQA, and KG-Hint; it includes two built-in model classes, Answer-Aware and Answer-Agnostic, built on instruction-tuned LLMs and extensible through a Model base class; and it implements fifteen evaluation methods across five metrics, including new lightweight variants. The paper reports average scores across all datasets and subsets as baselines, and positions the framework as the first of its kind for hint-related NLP and IR research.
Load-bearing premise
The load-bearing premise is that HintEval's automatic metric scores track how human readers actually judge hint quality; the paper does not include a human-agreement study to confirm this.
Editorial extensions
If this is right
- Hint generation systems can be benchmarked head-to-head on the same dataset splits and scored with the same metrics, making reported improvements directly comparable.
- Answer-agnostic generation can be developed and evaluated without gold answers, opening hinting to open-domain and private question-answering settings.
- New evaluation ideas can be implemented as custom metrics against the same data structure, so the five built-in metrics serve as a starting point rather than a ceiling.
- Researchers without heavy NLP infrastructure can load preprocessed datasets and run baseline evaluations with a few lines of code, lowering the barrier to entry.
- The reported baseline scores across TriviaHG, WikiHint, HintQA, and KG-Hint give any new method a ready-made comparison point.
Reading between the lines
- If HintEval becomes the de facto platform, hint quality could become a routinely measured dimension in QA and tutoring systems, comparable across papers.
- The same five dimensions apply naturally to other progressive-disclosure texts, such as tutorial steps, educational-game clues, or Socratic follow-ups, so the evaluation layer could generalize beyond factoid questions.
- The answer-agnostic mode suggests a testable extension: measuring whether hints preserve user engagement and learning outcomes, not just textual quality.
- A natural next study would collect human ratings on a sample of hints and calibrate which of the fifteen methods best matches human judgment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents HintEval, an open-source Python library that unifies hint-generation datasets (TriviaHG, WikiHint, HintQA, KG-Hint), answer-aware and answer-agnostic generation interfaces, and fifteen evaluation methods organized under five metrics (Relevance, Readability, Convergence, Familiarity, Answer Leakage). The manuscript describes the object model, the workflow, and illustrative code listings, and it reports average evaluation scores across all bundled dataset subsets in Table 2. The central claim is that HintEval is the first framework for hint generation and evaluation and that it enables 'clear, multi-faceted, and reliable evaluation' of hints.
Significance. If the evaluation metrics were validated for the hint-generation setting, HintEval would be a genuinely useful contribution: it lowers the integration barrier in a fragmented area, the package appears to be installable and documented, the dataset loader and custom extension points are concrete assets, and the generic reimplementation of previously scattered metrics addresses a real need. The primary value proposition, however, depends on the claim that the implemented scores measure hint quality; the evidence for this claim is currently indirect and incomplete, so the framework's significance for reliable evaluation is not yet established.
major comments (4)
- [§3.3 (footnotes 13–15); Abstract] The abstract and Section 2.4 claim that HintEval provides 'reliable evaluation,' but no experiment in the paper validates any of the five metrics against human judgments of hint quality. The only evidence offered is surrogate: readability models are validated on OneStopEnglish (footnote 13), relevance embeddings on WikiQA answer selection (footnotes 10–11), convergence on TriviaHG with Pearson correlations of only 56.1% and 61.1% (footnote 15), and Familiarity and Answer Leakage have no external validation at all. These benchmarks concern different tasks and do not establish that the scores track human assessments of whether a hint is relevant, readable, convergent, familiar, or leaky. I would require either a human-annotation agreement study on hints or a substantially weakened statement of the reliability claim.
- [Table 2] Table 2 reports average evaluation scores for every dataset subset but provides no variance, confidence intervals, or significance tests. As a result, the table cannot support comparative claims such as fine-tuned versus vanilla generation quality or differences among models; with roughly 100-question subsets and heterogeneous hints, the differences (e.g., Conv LLM 0.31 vs 0.54 across TriviaHG subsets) may be within noise. Please report per-subset standard deviations and, where comparisons are intended, appropriate significance or effect-size measures.
- [§3.3.3, footnote 15; Table 2] The Convergence Neural Network method is fine-tuned on TriviaHG training annotations and evaluated on TriviaHG test annotations, reaching only 0.56–0.61 Pearson correlation. Since Table 2 then uses this same method to score TriviaHG subsets, including a Training row, it is not clear whether the reported 'Conv NN' values for the training split are produced by a model trained on those very annotations; the paper should state this explicitly and, if so, exclude training-split scores from any reliability argument. Independent of that, a 0.56–0.61 correlation is too modest to underwrite 'reliable' convergence measurement.
- [§3.3.4 and §3.3.5] Familiarity and Answer Leakage are implemented with no validation whatsoever: the Wikipedia page-view method and the C4 word-frequency method are never compared with human familiarity ratings, and the lexical and contextual overlap methods are never compared with human leakage judgments. Since these metrics are part of the advertised 'multi-faceted evaluation,' their inclusion in the current paper is a research claim, not an established result.
minor comments (5)
- [Figure 5] The Hint class attribute is spelled 'entites'; this should be 'entities'.
- [Figure 2] The answer 'Micheal Jackson MJ' contains a spelling error ('Micheal' should be 'Michael').
- [Section 3.1, Listing 3] The statement that pickle files 'inherently provide a level of encryption' is incorrect; pickle serialization is not encrypted and is generally unsafe to load from untrusted sources. Please rephrase.
- [Figure 6 and §3.3.2] Figure 6 uses the label 'G-Fox' for the readability method, while the text uses 'Gunning Fog Index'; please align the labels.
- [Listing 5] The argument string 'e x cl u de _ s t o p _ wo r d s' contains unintended spacing; it should read 'exclude_stop_words'.
Circularity Check
No significant circularity: HintEval is a software/library contribution whose claims are engineering claims, not derived predictions; metric validations rest on external benchmarks and train/test splits.
full rationale
The paper is a systems/library contribution, not a derivation chain. Its central claims are that HintEval unifies hint datasets, models, and evaluation metrics and that its metrics are usable; these are engineering claims that can be independently tested by running the open-source library. The evaluation metrics are not derived from the framework's own outputs. Relevance embeddings are checked against WikiQA MAP/MRR (footnotes 10-11), readability models against OneStopEnglish accuracy/F1 (footnotes 12-13), specificity models against Ko et al. data (footnote 14), and the convergence neural network is fine-tuned on the TriviaHG training set and tested on the TriviaHG test set (footnote 15) - a standard supervised split, not a fitted-input-called-prediction reduction. Familiarity and Answer Leakage are presented as implementations of previously published operational definitions rather than as validated predictions. The heavy self-citation reflects that the authors created the relevant datasets (TriviaHG, WikiHint, HintQA) and earlier metric proposals, but the paper does not invoke any self-cited uniqueness theorem or use its own prior work to forbid alternatives; it simply aggregates that work into a common interface. The absence of a human-agreement study for the hint-specific metrics is a validity or correctness concern, not a circularity concern, because the paper makes no claim that the metrics were validated through a chain that presupposes the framework's own correctness. No equation, fitted parameter, or definition reduces to the paper's target claim, so no circular step can be exhibited.
Assumptions & free parameters
assumptions (4)
- domain assumption The implemented evaluation metrics measure the intended hint qualities (relevance, readability, convergence, familiarity, answer leakage).
- domain assumption The preprocessed datasets (TriviaHG, WikiHint, HintQA, KG-Hint) are accurately converted and representative of hint generation research.
- ad hoc to paper TriviaHG convergence annotations are reliable ground truth for training and testing convergence predictors.
- domain assumption LLM-as-judge outputs are acceptable substitutes for human ratings of relevance, readability, and convergence.
Cite this review
Pith. "Pith review of HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions." pith.science (2026). https://pith.science/paper/ILFP5NE6
@misc{pith2026250200857,
author = {Pith},
title = {Pith review of: HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions},
year = {2026},
howpublished = {\url{https://pith.science/paper/ILFP5NE6}},
note = {Machine review of arXiv:2502.00857}
}
read the original abstract
Large Language Models (LLMs) are transforming how people find information, and many users turn nowadays to chatbots to obtain answers to their questions. Despite the instant access to abundant information that LLMs offer, it is still important to promote critical thinking and problem-solving skills. Automatic hint generation is a new task that aims to support humans in answering questions by themselves by creating hints that guide users toward answers without directly revealing them. In this context, hint evaluation focuses on measuring the quality of hints, helping to improve the hint generation approaches. However, resources for hint research are currently spanning different formats and datasets, while the evaluation tools are missing or incompatible, making it hard for researchers to compare and test their models. To overcome these challenges, we introduce HintEval, a Python library that makes it easy to access diverse datasets and provides multiple approaches to generate and evaluate hints. HintEval aggregates the scattered resources into a single toolkit that supports a range of research goals and enables a clear, multi-faceted, and reliable evaluation. The proposed library also includes detailed online documentation, helping users quickly explore its features and get started. By reducing barriers to entry and encouraging consistent evaluation practices, HintEval offers a major step forward for facilitating hint generation and analysis research within the NLP/IR community.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Man- junath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: a system f...
2016
-
[2]
Abdelrahman Abdallah, Bhawna Piryani, and Adam Jatowt. 2023. Exploring the state of the art in legal QA systems. Journal of Big Data 10, 1 (12 Aug 2023),
2023
-
[3]
Heba Abdel-Nabi, Arafat Awajan, and Mostafa Z Ali. 2023. Deep learning-based question answering: a survey. Knowledge and Information Systems 65, 4 (2023), 1399–1485. https://doi.org/10.1007/s10115-022-01783-5
-
[4]
Kosuke Aigo, Takashi Tsunakawa, Masafumi Nishida, and Masafumi Nishimura
-
[5]
Riordan Alfredo, Vanessa Echeverria, Yueqiao Jin, Lixiang Yan, Zachari Swiecki, Dragan Gašević, and Roberto Martinez-Maldonado. 2024. Human-centred learn- ing analytics and AI in education: A systematic literature review. Computers and Education: Artificial Intelligence 6 (2024), 100215. https://doi.org/10.1016/j. caeai.2024.100215
arXiv 2024
-
[6]
Sheng, Wei Emma Zhang, Munazza Zaib, and Ahoud Alhazmi
Elaf Alhazmi, Quan Z. Sheng, Wei Emma Zhang, Munazza Zaib, and Ahoud Alhazmi. 2024. Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation. arXiv e-prints, Article arXiv:2402.01512 (Feb. 2024), arXiv:2402.01512 pages. https://doi.org/10.48550/arXiv.2402.01512 arXiv:2402.01512 [cs.CL]
-
[7]
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies, Volume 1 (...
2019
-
[8]
Albert Bandura. 2013. The role of self-efficacy in goal-based motiva- tion. New developments in goal setting and task performance (2013), 147–
2013
Show all 112 references
-
[9]
Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Jade Goldstein...
2005
-
[10]
Tiffany Barnes and John Stamper. 2008. Toward Automatic Hint Generation for Logic Proof Tutoring Using Historical Student Data. In Proceedings of the 9th International Conference on Intelligent Tutoring Systems (Montreal, Canada) (ITS ’08). Springer-Verlag, Berlin, Heidelberg,...
2008 doi
-
[11]
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic Parsing on Freebase from Question-Answer Pairs. InProceedings of the 2013 Con- ference on Empirical Methods in Natural Language Processing , David Yarowsky, Timothy Baldwin, Anna Korhonen, Karen Livescu...
2013
- [12]
-
[13]
Leo Breiman. 2001. Random Forests. Machine Learning 45, 1 (01 Oct 2001), 5–32. https://doi.org/10.1023/A:1010933404324
2001 doi
-
[14]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[15]
Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing ...
2022 doi
-
[16]
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking Large Language Models in Retrieval-Augmented Generation. Proceedings of the AAAI Conference on Artificial Intelligence 38, 16 (Mar. 2024), 17754–17762. https://doi.org/10.1609/aaai.v38i16.29728
2024 doi
-
[17]
Tianqi Chen and Carlos Guestrin. 2016. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (San Francisco, California, USA) (KDD ’16). Association for Computing Machinery, New York, NY,...
2016
-
[18]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Se- bastian Gehrmann, Parker Schuh, Kensen Shi, Sashank Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, ...
2024
-
[19]
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the Surprising Diffi- culty of Natural Yes/No Questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for C...
2019
-
[20]
Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Di- verse Languages. Transactions of the Association for Computation...
2020 doi
- [21]
-
[22]
Meri Coleman and T. L. Liau. 1975. A Computer Readability Formula Designed for Machine Scoring. Journal of Applied Psychology 60, 2 (1975), 283–284. https: //doi.org/10.1037/h0076540
1975 doi
-
[23]
Ali Darvishi, Hassan Khosravi, Shazia Sadiq, Dragan Gašević, and George Siemens. 2024. Impact of AI assistance on student agency. Computers & Educa- tion 210 (2024), 104967. https://doi.org/10.1016/j.compedu.2023.104967
2024
-
[24]
Rubel Das, Antariksha Ray, Souvik Mondal, and Dipankar Das. 2016. A rule based question generation framework to deal with simple and complex sentences. In 2016 International Conference on Advances in Computing, Communications and Informatics (ICACCI). 542–548. https://doi.org/...
2016
-
[25]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...
2019 doi
-
[26]
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods...
2021
- [27]
- [28]
-
[29]
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Pro- ceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Nik...
2024
-
[30]
Alexander Fabbri, Patrick Ng, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang
-
[31]
Samuel Fernando and Mark Stevenson. 2008. A semantic similarity approach to paraphrase detection. In Proceedings of the 11th annual research colloquium of the UK special interest group for computational linguistics . 45–52
2008
-
[32]
Rudolf Franz Flesch. 1948. A new readability yardstick. Journal of Applied Psychology 32, 3 (1948), 221–233. https://doi.org/10.1037/H0057532
1948 doi
- [33]
-
[34]
Sumam Francis and Marie-Francine Moens. 2023. Investigating better context representations for generative question answering. Inf. Retr. 26, 1–2 (Oct. 2023), 35 pages. https://doi.org/10.1007/s10791-023-09420-7
2023 doi
-
[35]
Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, et al
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, et al. 2023. Gemini: A Family of Highly Capa- ble Multimodal Models. arXiv e-prints, Article a...
-
[36]
Rupali Goyal, Parteek Kumar, and V. P. Singh. 2024. Automated Question and Answer Generation from Texts using Text-to-Text Transformers.Arabian Journal for Science and Engineering 49, 3 (01 Mar 2024), 3027–3041. https: //doi.org/10.1007/s13369-023-07840-7
2024 doi
-
[37]
Robert Gunning. 1952. The Technique of Clear Writing . McGraw–Hill, New York
1952
-
[38]
Tianyong Hao, Xinxin Li, Yulan He, Fu Lee Wang, and Yingying Qu. 2022. Recent progress in leveraging deep learning methods for question answering. Neural Comput. Appl. 34, 4 (Feb. 2022), 2765–2783. https://doi.org/10.1007/s00521-021- 06748-3
2022 doi
-
[39]
Richard Heersmink. 2024. Use of large language models might affect our cog- nitive skills. Nature Human Behaviour 8, 5 (01 May 2024), 805–806. https: //doi.org/10.1038/s41562-024-01859-y
2024 doi
-
[40]
Gregory Hume, Joel Michael, Allen Rovick, and Martha Evens. 1996. Hinting as a Tactic in One-on-One Tutoring. The Journal of the Learning Sciences 5, 1 (1996), 23–47. http://www.jstor.org/stable/1466758
1996
-
[41]
Anubhav Jangra, Jamshid Mozafari, Adam Jatowt, and Smaranda Muresan
-
[42]
Adam Jatowt, Calvin Gehrer, and Michael Färber. 2023. Automatic Hint Genera- tion. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval (Taipei, Taiwan) (ICTIR ’23). Association for Computing Machinery, New York, NY, USA, 117–123. ht...
2023 doi
-
[43]
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. Trivi- aQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Com- prehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ...
2017 doi
-
[44]
Gregor Jošt, Viktor Taneski, and Sašo Karakatič. 2024. The Impact of Large Language Models on Programming Education and Student Learning Outcomes. Applied Sciences 14, 10 (2024). https://doi.org/10.3390/app14104115
2024 doi
-
[45]
Daniel Jurafsky and James Martin. [n. d.]. Speech and Language Processing Pearson New International Edition
-
[46]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP...
2020 doi
-
[47]
Wei-Jen Ko, Greg Durrett, and Junyi Jessy Li. 2019. Domain Agnostic Real- Valued Specificity Prediction. Proceedings of the AAAI Conference on Artificial Intelligence 33, 01 (Jul. 2019), 6610–6617. https://doi.org/10.1609/aaai.v33i01. 33016610
2019 doi
-
[48]
Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Varun Gupta, Iulian Vlad Serban, and Joelle Pineau. 2022. Automated Data-Driven Generation of Personalized Pedagogical Interventions in Intelligent Tutoring Systems. International Journal of Artificial Intelligence in Education 32,...
2022 doi
-
[49]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Ken- ton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Sl...
2019 doi
-
[50]
Md Tahmid Rahman Laskar, Jimmy Xiangji Huang, and Enamul Hoque. 2020. Contextualized Embeddings based Transformer Encoder for Sentence Similarity Modeling in Answer Selection Task. In Proceedings of the Twelfth Language Re- sources and Evaluation Conference, Nicoletta Calzolar...
2020
-
[51]
Harry Mc Laughlin
G. Harry Mc Laughlin. 1969. SMOG Grading-a New Readability Formula.Journal of Reading 12, 8 (1969), 639–646. http://www.jstor.org/stable/40011226
1969
-
[52]
Lee and Jason Lee
Bruce W. Lee and Jason Lee. 2023. LFTK: Handcrafted Features in Computational Linguistics. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023) , Ekaterina Kochmar, Jill Burstein, Andrea Horbach, Ronja Laarmann-Quante, N...
2023 doi
-
[53]
Junlong Li, Jinyuan Wang, Zhuosheng Zhang, and Hai Zhao. 2024. Self- Prompting Large Language Models for Zero-Shot Open-Domain QA. In Pro- ceedings of the 2024 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technolog...
2024 doi
-
[54]
Xin Li and Dan Roth. 2002. Learning question classifiers. In Proceedings of the 19th International Conference on Computational Linguistics - Volume 1 (Taipei, Taiwan) (COLING ’02). Association for Computational Linguistics, USA, 1–7. https://doi.org/10.3115/1072228.1072378
2002
-
[55]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013
2004
-
[56]
Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations. In Proceedings of the 44th International ACM SIGIR Conference ...
2021
-
[57]
Fengkai Liu and John Lee. 2023. Hybrid Models for Sentence Readability Assess- ment. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), Ekaterina Kochmar, Jill Burstein, Andrea Horbach, Ronja Laarmann-Quante, Nitin Ma...
2023 doi
- [58]
-
[59]
Zhuang Liu, Kaiyu Huang, Degen Huang, and Jun Zhao. 2020. Semantics- reinforced networks for question generation. In ECAI 2020. IOS Press, 2078– 2084
2020
-
[60]
Vaibhav Mavi, Anubhav Jangra, and Adam Jatowt. 2024. Multi-hop Question Answering. Found. Trends Inf. Retr. 17, 5 (2024), 457–586. https://doi.org/10. 1561/1500000102
2024
-
[61]
Jessica McBroom, Irena Koprinska, and Kalina Yacef. 2021. A Survey of Auto- mated Programming Hint Generation: The HINTS Framework. ACM Comput. Surv. 54, 8, Article 172 (Oct. 2021), 27 pages. https://doi.org/10.1145/3469885
2021 doi
-
[62]
Jamshid Mozafari, Abdelrahman Abdallah, Bhawna Piryani, and Adam Jatowt
-
[63]
Jamshid Mozafari, Afsaneh Fatemi, and Parham Moradi. 2020. A Method For Answer Selection Using DistilBERT And Important Words. In 2020 6th Inter- national Conference on Web Research (ICWR) . 72–76. https://doi.org/10.1109/ ICWR49608.2020.9122302
2020
-
[64]
Jamshid Mozafari, Afsaneh Fatemi, and Mohammad Ali Nematbakhsh. 2021. BAS: An Answer Selection Method Using BERT Language Model. Journal of Computing and Security 8, 2 (2021), 1–18. https://doi.org/10.22108/jcs.2021. 128002.1066
2021
- [65]
- [67]
-
[68]
Exploring Hint Generation Approaches for Open-Domain Question An- swering. (Nov. 2024), 9327–9352. https://doi.org/10.18653/v1/2024.findings- emnlp.546 SIGIR ’25, July 13–18, 2025, Padova, IT Mozafari et al
2024 doi
-
[69]
Tarek Naous, Michael J Ryan, Anton Lavrouk, Mohit Chandra, and Wei Xu. 2024. ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment. In Proceedings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing , Yaser Al-Onai...
2024 doi
-
[70]
Minh Nguyen, K. C. Kishan, Toan Nguyen, Ankit Chadha, and Thuy Vu. 2023. Efficient Fine-Tuning Large Language Models for Knowledge-Aware Response Planning. In Machine Learning and Knowledge Discovery in Databases: Research Track: European Conference, ECML PKDD 2023, Turin, Ita...
2023 doi
- [71]
-
[72]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (Philadel- phia, Pennsylvania) (ACL ’02). Association for C...
2002
-
[73]
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu F...
2019
-
[74]
Nikahat Mulla and Prachi Gharpure. 2023. Automatic question generation: a review of methodologies, datasets, evaluation metrics, and applications. Prog. in Artif. Intell. 12, 1 (Jan. 2023), 1–32. https://doi.org/10.1007/s13748-023-00295-9
2023 doi
-
[75]
Bhawna Piryani, Jamshid Mozafari, and Adam Jatowt. 2024. ChroniclingAmeri- caQA: A Large-scale Question Answering Dataset based on Historical American Newspaper Pages. In Proceedings of the 47th International ACM SIGIR Confer- ence on Research and Development in Information Re...
2024
-
[76]
Price, Yihuan Dong, Rui Zhi, Benjamin Paaßen, Nicholas Lytle, Veronica Cateté, and Tiffany Barnes
Thomas W. Price, Yihuan Dong, Rui Zhi, Benjamin Paaßen, Nicholas Lytle, Veronica Cateté, and Tiffany Barnes. 2019. A Comparison of the Quality of Data-Driven Programming Hint Generation Algorithms. International Journal of Artificial Intelligence in Education 29, 3 (01 Aug 201...
2019 doi
-
[77]
Vasin Punyakanok, Dan Roth, and Wen-tau Yih. 2004. Natural language in- ference via dependency tree mapping: An application to question answering. (2004)
2004
- [78]
-
[79]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 1, Article 140 (Jan. 2020), 67 pages
2020
-
[80]
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Alessandro Moschitti, Bo Pang, and Walter Daelemans (Eds.). Asso...
2014 doi
-
[81]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Rad- ford, Mark Chen, and Ilya Sutskever. 2021. Zero-Shot Text-to-Image Generation. In Proceedings of the 38th International Conference on Machine Learning (Pro- ceedings of Machine Learning Research, V...
2021
- [82]
- [83]
-
[84]
Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A Conver- sational Question Answering Challenge. Transactions of the Association for Com- putational Linguistics 7 (2019), 249–266. https://doi.org/10.1162/tacl_a_00266
2019 doi
-
[85]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-I...
2019 doi
-
[86]
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Iryna Gurevych and Yusuke Miyao (Eds.). Associa...
2018 doi
-
[87]
Devendra Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022. Improving Passage Retrieval with Zero-Shot Question Generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav...
2022 doi
-
[88]
R. J. Senter and E. A. Smith. 1967.Automated Readability Index. Technical Report AMRL-TR-66-220. Aerospace Medical Research Laboratories, Wright-Patterson Air Force Base. https://apps.dtic.mil/sti/citations/AD0667273
1967
-
[89]
Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language...
2023 doi
-
[90]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2019 doi
-
[91]
Harish Tayyar Madabushi and Mark Lee. 2016. High Accuracy Rule-based Question Classification using Question Syntax and Semantics. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, Yuji Matsumoto and Rashmi Prasad (...
2016
-
[92]
Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2023. QA Dataset Ex- plosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension. ACM Comput. Surv. 55, 10, Article 197 (Feb. 2023), 45 pages. https://doi.org/10.1145/3560260
2023 doi
-
[93]
Asahi Ushio, Fernando Alva-Manchego, and Jose Camacho-Collados. 2023. A Practical Toolkit for Multilingual Question and Answer Generation. In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), Danushka B...
2023 doi
-
[94]
Sowmya Vajjala and Ivana Lučić. 2018. OneStopEnglish corpus: A new corpus for automatic readability assessment and text simplification. In Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications, Joel Tetreault, Jill Burstein, Ekat...
2018 doi
-
[95]
Cunxiang Wang, Pai Liu, and Yue Zhang. 2021. Can Generative Pre-trained Language Models Serve As Knowledge Bases for Closed-book QA?. In Proceed- ings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Nat...
2021 doi
-
[96]
Jiexin Wang, Adam Jatowt, and Masatoshi Yoshikawa. 2022. ArchivalQA: A Large-scale Benchmark Dataset for Open-Domain Question Answering over Historical News Collections (SIGIR ’22). Association for Computing Machinery, New York, NY, USA, 3025–3035. https://doi.org/10.1145/3477...
2022
-
[97]
Luqi Wang, Kaiwen Zheng, Liyin Qian, and Sheng Li. 2022. A Survey of Extrac- tive Question Answering. In 2022 International Conference on High Performance Big Data and Intelligent Systems (HDIS) . 147–153. https://doi.org/10.1109/ HDIS56859.2022.9991478
2022
-
[98]
Usher and Frank Pajares
Ellen L. Usher and Frank Pajares. 2006. Sources of academic and self-regulatory efficacy beliefs of entering middle school students. Contemporary Educational Psychology 31, 2 (2006), 125–141. https://doi.org/10.1016/j.cedpsych.2005.03.002
2006 doi
- [99]
-
[100]
Yi Yang, Wen-tau Yih, and Christopher Meek. 2015. WikiQA: A Challenge Dataset for Open-Domain Question Answering. In Proceedings of the 2015 Con- ference on Empirical Methods in Natural Language Processing , Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association ...
2015 doi
-
[101]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...
2018
-
[102]
Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023. Generate rather than Retrieve: Large Language Models are Strong Context Genera- tors. In The Eleventh International Conference on Learning Representat...
2023
-
[103]
Xingdi Yuan, Tong Wang, Yen-Hsiang Wang, Emery Fine, Rania Abdelghani, Hélène Sauzéon, and Pierre-Yves Oudeyer. 2023. Selecting Better Samples from Pre-trained LLMs: A Case Study on Question Generation. In Findings of the Association for Computational Linguistics: ACL 2023 , A...
2023 doi
- [104]
-
[105]
Weinberger, and Yoav Artzi
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. InInternational Conference on Learning Representations . https://openreview.net/forum?id= SkeHuCVFDr
2020
-
[106]
Tiancheng Zhao, Xiaopeng Lu, and Kyusong Lee. 2021. SPARTA: Efficient Open-Domain Question Answering via Sparse Transformer Matching Re- trieval. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Languag...
2021
- [107]
-
[110]
Ruqing Zhang, Jiafeng Guo, Lu Chen, Yixing Fan, and Xueqi Cheng. 2021. A Review on Question Generation from Natural Language Text. ACM Trans. Inf. Syst. 40, 1, Article 14 (Sept. 2021), 43 pages. https://doi.org/10.1145/3468889
2021 doi
-
[127]
https://doi.org/10.1186/s40537-023-00802-8
-
[157]
https://www.taylorfrancis.com/chapters/edit/10.4324/9780203082744- 13/role-self-efficacy-goal-based-motivation-albert-bandura
-
[2020]
In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.)
Template-Based Question Generation from Retrieved Sentences for Im- proved Unsupervised Question Answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Assoc...
-
[2021]
In 2021 IEEE 10th Global Conference on Con- sumer Electronics (GCCE)
Question Generation using Knowledge Graphs with the T5 Language Model and Masked Self-Attention. In 2021 IEEE 10th Global Conference on Con- sumer Electronics (GCCE) . 85–87. https://doi.org/10.1109/GCCE53005.2021. 9621874
2021
- [2024]
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.