Pith. sign in

REVIEW 4 major objections 5 minor 112 references

HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions

T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read The paper introduces HintEval, a unified open-source library for generating and scoring hints for questions, and presents it as the first framework of its kind.

desk verdict A genuinely useful first-of-its-kind toolkit for hint generation and evaluation, but the 'reliable evaluation' headline goes beyond the evidence. read the letter →

arxiv 2502.00857 v1 pith:ILFP5NE6 submitted 2025-02-02 cs.CL cs.IR

classification cs.CLcs.IR
keywords hintgenerationevaluationquestionansweringPythonframeworklargelanguagemodelsmetricsanswer-awareanswer-agnostic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces HintEval, which it presents as the first unified, open-source framework for automatic hint generation and hint evaluation. Hint generation creates guidance that helps a person answer a question without revealing the answer, and hint evaluation measures whether such guidance is actually good. The authors argue that existing resources are fragmented: hint datasets live in different formats, and evaluation tools are missing or incompatible. HintEval addresses this by packaging four hint datasets into a common data structure, providing generators that work with or without the ground-truth answer, and implementing five evaluation metrics—relevance, readability, convergence, familiarity, and answer leakage—with methods ranging from cheap lexical formulas to expensive LLM judges. If the framework works as claimed, researchers can generate hints, score them, and compare methods on shared infrastructure.

What carries the argument

The load-bearing piece is the Dataset object model: a unified schema with Instance, Question, Answer, Hint, Entity, and Metric classes that converts any hint dataset into the same structure. Around this core sit two extensible base classes—Model, which exposes a generate function for answer-aware or answer-agnostic hint production, and Evaluation, which exposes an evaluate function for any metric. The argument works by standardization: once every dataset is in the same schema, every generator can write hints into it and every metric can read from it, which is what lets the framework claim consistent, comparable evaluation.

What would settle it

Take a random sample of hints from TriviaHG's test set, have several human annotators rate each hint on relevance, readability, convergence, familiarity, and answer leakage using the paper's own definitions, and compare the average human ratings with HintEval's automatic scores; if the correlation is near zero or negative for any dimension, that metric does not measure what the framework claims.

Watch

Extended reading notes

Core claim

The central claim is that HintEval removes the main obstacle to hint research by putting every dataset, generator, and metric under one interface. The framework defines a Dataset schema in which questions, answers, hints, named entities, and metric scores are first-class objects; it ships with preprocessed versions of TriviaHG, WikiHint, HintQA, and KG-Hint; it includes two built-in model classes, Answer-Aware and Answer-Agnostic, built on instruction-tuned LLMs and extensible through a Model base class; and it implements fifteen evaluation methods across five metrics, including new lightweight variants. The paper reports average scores across all datasets and subsets as baselines, and positions the framework as the first of its kind for hint-related NLP and IR research.

Load-bearing premise

The load-bearing premise is that HintEval's automatic metric scores track how human readers actually judge hint quality; the paper does not include a human-agreement study to confirm this.

Editorial extensions

If this is right

  • Hint generation systems can be benchmarked head-to-head on the same dataset splits and scored with the same metrics, making reported improvements directly comparable.
  • Answer-agnostic generation can be developed and evaluated without gold answers, opening hinting to open-domain and private question-answering settings.
  • New evaluation ideas can be implemented as custom metrics against the same data structure, so the five built-in metrics serve as a starting point rather than a ceiling.
  • Researchers without heavy NLP infrastructure can load preprocessed datasets and run baseline evaluations with a few lines of code, lowering the barrier to entry.
  • The reported baseline scores across TriviaHG, WikiHint, HintQA, and KG-Hint give any new method a ready-made comparison point.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If HintEval becomes the de facto platform, hint quality could become a routinely measured dimension in QA and tutoring systems, comparable across papers.
  • The same five dimensions apply naturally to other progressive-disclosure texts, such as tutorial steps, educational-game clues, or Socratic follow-ups, so the evaluation layer could generalize beyond factoid questions.
  • The answer-agnostic mode suggests a testable extension: measuring whether hints preserve user engagement and learning outcomes, not just textual quality.
  • A natural next study would collect human ratings on a sample of hints and calibrate which of the fifteen methods best matches human judgment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents HintEval, an open-source Python library that unifies hint-generation datasets (TriviaHG, WikiHint, HintQA, KG-Hint), answer-aware and answer-agnostic generation interfaces, and fifteen evaluation methods organized under five metrics (Relevance, Readability, Convergence, Familiarity, Answer Leakage). The manuscript describes the object model, the workflow, and illustrative code listings, and it reports average evaluation scores across all bundled dataset subsets in Table 2. The central claim is that HintEval is the first framework for hint generation and evaluation and that it enables 'clear, multi-faceted, and reliable evaluation' of hints.

Significance. If the evaluation metrics were validated for the hint-generation setting, HintEval would be a genuinely useful contribution: it lowers the integration barrier in a fragmented area, the package appears to be installable and documented, the dataset loader and custom extension points are concrete assets, and the generic reimplementation of previously scattered metrics addresses a real need. The primary value proposition, however, depends on the claim that the implemented scores measure hint quality; the evidence for this claim is currently indirect and incomplete, so the framework's significance for reliable evaluation is not yet established.

major comments (4)
  1. [§3.3 (footnotes 13–15); Abstract] The abstract and Section 2.4 claim that HintEval provides 'reliable evaluation,' but no experiment in the paper validates any of the five metrics against human judgments of hint quality. The only evidence offered is surrogate: readability models are validated on OneStopEnglish (footnote 13), relevance embeddings on WikiQA answer selection (footnotes 10–11), convergence on TriviaHG with Pearson correlations of only 56.1% and 61.1% (footnote 15), and Familiarity and Answer Leakage have no external validation at all. These benchmarks concern different tasks and do not establish that the scores track human assessments of whether a hint is relevant, readable, convergent, familiar, or leaky. I would require either a human-annotation agreement study on hints or a substantially weakened statement of the reliability claim.
  2. [Table 2] Table 2 reports average evaluation scores for every dataset subset but provides no variance, confidence intervals, or significance tests. As a result, the table cannot support comparative claims such as fine-tuned versus vanilla generation quality or differences among models; with roughly 100-question subsets and heterogeneous hints, the differences (e.g., Conv LLM 0.31 vs 0.54 across TriviaHG subsets) may be within noise. Please report per-subset standard deviations and, where comparisons are intended, appropriate significance or effect-size measures.
  3. [§3.3.3, footnote 15; Table 2] The Convergence Neural Network method is fine-tuned on TriviaHG training annotations and evaluated on TriviaHG test annotations, reaching only 0.56–0.61 Pearson correlation. Since Table 2 then uses this same method to score TriviaHG subsets, including a Training row, it is not clear whether the reported 'Conv NN' values for the training split are produced by a model trained on those very annotations; the paper should state this explicitly and, if so, exclude training-split scores from any reliability argument. Independent of that, a 0.56–0.61 correlation is too modest to underwrite 'reliable' convergence measurement.
  4. [§3.3.4 and §3.3.5] Familiarity and Answer Leakage are implemented with no validation whatsoever: the Wikipedia page-view method and the C4 word-frequency method are never compared with human familiarity ratings, and the lexical and contextual overlap methods are never compared with human leakage judgments. Since these metrics are part of the advertised 'multi-faceted evaluation,' their inclusion in the current paper is a research claim, not an established result.
minor comments (5)
  1. [Figure 5] The Hint class attribute is spelled 'entites'; this should be 'entities'.
  2. [Figure 2] The answer 'Micheal Jackson MJ' contains a spelling error ('Micheal' should be 'Michael').
  3. [Section 3.1, Listing 3] The statement that pickle files 'inherently provide a level of encryption' is incorrect; pickle serialization is not encrypted and is generally unsafe to load from untrusted sources. Please rephrase.
  4. [Figure 6 and §3.3.2] Figure 6 uses the label 'G-Fox' for the readability method, while the text uses 'Gunning Fog Index'; please align the labels.
  5. [Listing 5] The argument string 'e x cl u de _ s t o p _ wo r d s' contains unintended spacing; it should read 'exclude_stop_words'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: HintEval is a software/library contribution whose claims are engineering claims, not derived predictions; metric validations rest on external benchmarks and train/test splits.

full rationale

The paper is a systems/library contribution, not a derivation chain. Its central claims are that HintEval unifies hint datasets, models, and evaluation metrics and that its metrics are usable; these are engineering claims that can be independently tested by running the open-source library. The evaluation metrics are not derived from the framework's own outputs. Relevance embeddings are checked against WikiQA MAP/MRR (footnotes 10-11), readability models against OneStopEnglish accuracy/F1 (footnotes 12-13), specificity models against Ko et al. data (footnote 14), and the convergence neural network is fine-tuned on the TriviaHG training set and tested on the TriviaHG test set (footnote 15) - a standard supervised split, not a fitted-input-called-prediction reduction. Familiarity and Answer Leakage are presented as implementations of previously published operational definitions rather than as validated predictions. The heavy self-citation reflects that the authors created the relevant datasets (TriviaHG, WikiHint, HintQA) and earlier metric proposals, but the paper does not invoke any self-cited uniqueness theorem or use its own prior work to forbid alternatives; it simply aggregates that work into a common interface. The absence of a human-agreement study for the hint-specific metrics is a validity or correctness concern, not a circularity concern, because the paper makes no claim that the metrics were validated through a chain that presupposes the framework's own correctness. No equation, fitted parameter, or definition reduces to the paper's target claim, so no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The evaluation loop is almost entirely self-generated: the bundled datasets and several scoring methods come from the same research group, and the neural convergence models are trained on the group's own TriviaHG annotations. The framework's existence and API are independently checkable, which limits the circularity burden, but the reliability claims rest on untested assumptions about metric validity. No free parameters or invented entities are introduced.

assumptions (4)
  • domain assumption The implemented evaluation metrics measure the intended hint qualities (relevance, readability, convergence, familiarity, answer leakage).
    Sec 3.3 describes each metric with no human-agreement validation for the hint setting, so construct validity is assumed.
  • domain assumption The preprocessed datasets (TriviaHG, WikiHint, HintQA, KG-Hint) are accurately converted and representative of hint generation research.
    Sec 3.1 and Table 1: conversions were done by the authors and are not independently audited.
  • ad hoc to paper TriviaHG convergence annotations are reliable ground truth for training and testing convergence predictors.
    Sec 3.3.3: convergence NN methods are fine-tuned using convergence values from the TriviaHG training set, a resource created by the same authors in reference [66].
  • domain assumption LLM-as-judge outputs are acceptable substitutes for human ratings of relevance, readability, and convergence.
    Secs 3.3.1, 3.3.2, 3.3.3: LLM methods are invoked through APIs without calibration against human judgments in the hint context.

how reviews work

0 comments
Cite this review

Pith. "Pith review of HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions." pith.science (2026). https://pith.science/paper/ILFP5NE6

@misc{pith2026250200857,
  author       = {Pith},
  title        = {Pith review of: HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ILFP5NE6}},
  note         = {Machine review of arXiv:2502.00857}
}
read the original abstract

Large Language Models (LLMs) are transforming how people find information, and many users turn nowadays to chatbots to obtain answers to their questions. Despite the instant access to abundant information that LLMs offer, it is still important to promote critical thinking and problem-solving skills. Automatic hint generation is a new task that aims to support humans in answering questions by themselves by creating hints that guide users toward answers without directly revealing them. In this context, hint evaluation focuses on measuring the quality of hints, helping to improve the hint generation approaches. However, resources for hint research are currently spanning different formats and datasets, while the evaluation tools are missing or incompatible, making it hard for researchers to compare and test their models. To overcome these challenges, we introduce HintEval, a Python library that makes it easy to access diverse datasets and provides multiple approaches to generate and evaluate hints. HintEval aggregates the scattered resources into a single toolkit that supports a range of research goals and enables a clear, multi-faceted, and reliable evaluation. The proposed library also includes detailed online documentation, helping users quickly explore its features and get started. By reducing barriers to entry and encouraging consistent evaluation practices, HintEval offers a major step forward for facilitating hint generation and analysis research within the NLP/IR community.

Figures

Figures reproduced from arXiv: 2502.00857 by the authors.

Figure 1
Figure 1. HintEval logo. Keywords Hint Evaluation, Hint Generation, Python Package, Framework ACM Reference Format: Jamshid Mozafari, Bhawna Piryani, Abdelrahman Abdallah, and Adam Jatowt. 2025. HintEval: A Comprehensive Framework for Hint Generation and Evaluation for Questions. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (SIGIR ’25). ACM, New York, NY, USA, 13 pages.… view at source ↗
Figure 2
Figure 2. Example hints for a sample question with scoring metrics. The metrics Relevance, Convergence, Familiarity, and [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Workflow of the HintEval: ○1 Questions are loaded and converted into a structured dataset using the Dataset module. ○2 Users can load preprocessed datasets as a structured dataset.○3 Hints can be generated for each question using the Model module and stored in the dataset object. ○4 The Evaluation module assesses all generated hints and questions using various evaluation metrics, storing the results in the dataset o… view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Schema of the Dataset class, illustrating the objects [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 4
Figure 4. Figure 4: A docstring for the evaluate function of the Wikipedia method within the Familiarity evaluation met￾ric. The docstring begins with: ○1 A detailed description of the function, followed by ○2 Notes specific to the evaluation metric and the method. It includes ○3 a compre…
Figure 6
Figure 6. Figure 6: Evaluation metrics included in the HintEval frame￾work. Dark blue boxes denote the primary evaluation metrics, green boxes indicate methods associated with each metric implemented from scratch in HintEval, and light blue boxes highlight methods adopted from prior studi…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

112 extracted references · 26 canonical work pages

  1. [1]

    Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng

    Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Man- junath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: a system f...

  2. [2]

    Abdelrahman Abdallah, Bhawna Piryani, and Adam Jatowt. 2023. Exploring the state of the art in legal QA systems. Journal of Big Data 10, 1 (12 Aug 2023),

  3. [3]

    Heba Abdel-Nabi, Arafat Awajan, and Mostafa Z Ali. 2023. Deep learning-based question answering: a survey. Knowledge and Information Systems 65, 4 (2023), 1399–1485. https://doi.org/10.1007/s10115-022-01783-5

  4. [4]

    Kosuke Aigo, Takashi Tsunakawa, Masafumi Nishida, and Masafumi Nishimura

  5. [5]

    Riordan Alfredo, Vanessa Echeverria, Yueqiao Jin, Lixiang Yan, Zachari Swiecki, Dragan Gašević, and Roberto Martinez-Maldonado. 2024. Human-centred learn- ing analytics and AI in education: A systematic literature review. Computers and Education: Artificial Intelligence 6 (2024), 100215. https://doi.org/10.1016/j. caeai.2024.100215

  6. [6]

    Sheng, Wei Emma Zhang, Munazza Zaib, and Ahoud Alhazmi

    Elaf Alhazmi, Quan Z. Sheng, Wei Emma Zhang, Munazza Zaib, and Ahoud Alhazmi. 2024. Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation. arXiv e-prints, Article arXiv:2402.01512 (Feb. 2024), arXiv:2402.01512 pages. https://doi.org/10.48550/arXiv.2402.01512 arXiv:2402.01512 [cs.CL]

  7. [7]

    Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies, Volume 1 (...

  8. [8]

    Albert Bandura. 2013. The role of self-efficacy in goal-based motiva- tion. New developments in goal setting and task performance (2013), 147–

Show all 112 references
  1. [9]

    Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Jade Goldstein...

  2. [10]

    Tiffany Barnes and John Stamper. 2008. Toward Automatic Hint Generation for Logic Proof Tutoring Using Historical Student Data. In Proceedings of the 9th International Conference on Intelligent Tutoring Systems (Montreal, Canada) (ITS ’08). Springer-Verlag, Berlin, Heidelberg,...

  3. [11]

    Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic Parsing on Freebase from Question-Answer Pairs. InProceedings of the 2013 Con- ference on Empirical Methods in Natural Language Processing , David Yarowsky, Timothy Baldwin, Anna Korhonen, Karen Livescu...

  4. [12]

    Antoine Bordes, Nicolas Usunier, Sumit Chopra, and Jason Weston. 2015. Large- scale Simple Question Answering with Memory Networks.arXiv e-prints, Article arXiv:1506.02075 (June 2015), arXiv:1506.02075 pages. https://doi.org/10.48550/ arXiv.1506.02075 arXiv:1506.02075 [cs.LG]

  5. [13]

    Leo Breiman. 2001. Random Forests. Machine Learning 45, 1 (01 Oct 2001), 5–32. https://doi.org/10.1023/A:1010933404324

  6. [14]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  7. [15]

    Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing ...

  8. [16]

    Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking Large Language Models in Retrieval-Augmented Generation. Proceedings of the AAAI Conference on Artificial Intelligence 38, 16 (Mar. 2024), 17754–17762. https://doi.org/10.1609/aaai.v38i16.29728

  9. [17]

    Tianqi Chen and Carlos Guestrin. 2016. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (San Francisco, California, USA) (KDD ’16). Association for Computing Machinery, New York, NY,...

  10. [18]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Se- bastian Gehrmann, Parker Schuh, Kensen Shi, Sashank Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, ...

  11. [19]

    Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. BoolQ: Exploring the Surprising Diffi- culty of Natural Yes/No Questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for C...

  12. [20]

    Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki

    Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Di- verse Languages. Transactions of the Association for Computation...

  13. [21]

    Benjamin Clavié. 2024. rerankers: A Lightweight Python Library to Unify Ranking Methods. arXiv e-prints , Article arXiv:2408.17344 (Aug. 2024), arXiv:2408.17344 pages. https://doi.org/10.48550/arXiv.2408.17344 arXiv:2408.17344 [cs.IR]

  14. [22]

    Meri Coleman and T. L. Liau. 1975. A Computer Readability Formula Designed for Machine Scoring. Journal of Applied Psychology 60, 2 (1975), 283–284. https: //doi.org/10.1037/h0076540

  15. [23]

    Ali Darvishi, Hassan Khosravi, Shazia Sadiq, Dragan Gašević, and George Siemens. 2024. Impact of AI assistance on student agency. Computers & Educa- tion 210 (2024), 104967. https://doi.org/10.1016/j.compedu.2023.104967

  16. [24]

    Rubel Das, Antariksha Ray, Souvik Mondal, and Dipankar Das. 2016. A rule based question generation framework to deal with simple and complex sentences. In 2016 International Conference on Advances in Computing, Communications and Informatics (ICACCI). 542–548. https://doi.org/...

  17. [25]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  18. [26]

    Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods...

  19. [27]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, et al. 2024. The Llama 3 Herd of Models. arXiv e-prints, Article arXiv:2407.21783 (July 2024), arXiv:2407.21783 pages. https://doi.org/10.485...

  20. [28]

    Luis Enrico Lopez, Diane Kathryn Cruz, Jan Christian Blaise Cruz, and Chari- beth Cheng. 2020. Simplifying Paragraph-level Question Generation via Transformer Language Models. arXiv e-prints, Article arXiv:2005.01107 (May 2020), arXiv:2005.01107 pages. https://doi.org/10.48550...

  21. [29]

    Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Pro- ceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Nik...

  22. [30]

    Alexander Fabbri, Patrick Ng, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang

  23. [31]

    Samuel Fernando and Mark Stevenson. 2008. A semantic similarity approach to paraphrase detection. In Proceedings of the 11th annual research colloquium of the UK special interest group for computational linguistics . 45–52

  24. [32]

    Rudolf Franz Flesch. 1948. A new readability yardstick. Journal of Applied Psychology 32, 3 (1948), 221–233. https://doi.org/10.1037/H0057532

  25. [33]

    Shima Foolad, Kourosh Kiani, and Razieh Rastgoo. 2024. Recent Ad- vances in Multi-Choice Machine Reading Comprehension: A Survey on Methods and Datasets. arXiv e-prints , Article arXiv:2408.02114 (Aug. 2024), arXiv:2408.02114 pages. https://doi.org/10.48550/arXiv.2408.02114 ar...

  26. [34]

    Sumam Francis and Marie-Francine Moens. 2023. Investigating better context representations for generative question answering. Inf. Retr. 26, 1–2 (Oct. 2023), 35 pages. https://doi.org/10.1007/s10791-023-09420-7

  27. [35]

    Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, et al

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, et al. 2023. Gemini: A Family of Highly Capa- ble Multimodal Models. arXiv e-prints, Article a...

  28. [36]

    Rupali Goyal, Parteek Kumar, and V. P. Singh. 2024. Automated Question and Answer Generation from Texts using Text-to-Text Transformers.Arabian Journal for Science and Engineering 49, 3 (01 Mar 2024), 3027–3041. https: //doi.org/10.1007/s13369-023-07840-7

  29. [37]

    Robert Gunning. 1952. The Technique of Clear Writing . McGraw–Hill, New York

  30. [38]

    Tianyong Hao, Xinxin Li, Yulan He, Fu Lee Wang, and Yingying Qu. 2022. Recent progress in leveraging deep learning methods for question answering. Neural Comput. Appl. 34, 4 (Feb. 2022), 2765–2783. https://doi.org/10.1007/s00521-021- 06748-3

  31. [39]

    Richard Heersmink. 2024. Use of large language models might affect our cog- nitive skills. Nature Human Behaviour 8, 5 (01 May 2024), 805–806. https: //doi.org/10.1038/s41562-024-01859-y

  32. [40]

    Gregory Hume, Joel Michael, Allen Rovick, and Martha Evens. 1996. Hinting as a Tactic in One-on-One Tutoring. The Journal of the Learning Sciences 5, 1 (1996), 23–47. http://www.jstor.org/stable/1466758

  33. [41]

    Anubhav Jangra, Jamshid Mozafari, Adam Jatowt, and Smaranda Muresan

  34. [42]

    Adam Jatowt, Calvin Gehrer, and Michael Färber. 2023. Automatic Hint Genera- tion. In Proceedings of the 2023 ACM SIGIR International Conference on Theory of Information Retrieval (Taipei, Taiwan) (ICTIR ’23). Association for Computing Machinery, New York, NY, USA, 117–123. ht...

  35. [43]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. Trivi- aQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Com- prehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) ...

  36. [44]

    Gregor Jošt, Viktor Taneski, and Sašo Karakatič. 2024. The Impact of Large Language Models on Programming Education and Student Learning Outcomes. Applied Sciences 14, 10 (2024). https://doi.org/10.3390/app14104115

  37. [45]

    Daniel Jurafsky and James Martin. [n. d.]. Speech and Language Processing Pearson New International Edition

  38. [46]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP...

  39. [47]

    Wei-Jen Ko, Greg Durrett, and Junyi Jessy Li. 2019. Domain Agnostic Real- Valued Specificity Prediction. Proceedings of the AAAI Conference on Artificial Intelligence 33, 01 (Jul. 2019), 6610–6617. https://doi.org/10.1609/aaai.v33i01. 33016610

  40. [48]

    Ekaterina Kochmar, Dung Do Vu, Robert Belfer, Varun Gupta, Iulian Vlad Serban, and Joelle Pineau. 2022. Automated Data-Driven Generation of Personalized Pedagogical Interventions in Intelligent Tutoring Systems. International Journal of Artificial Intelligence in Education 32,...

  41. [49]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Ken- ton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Sl...

  42. [50]

    Md Tahmid Rahman Laskar, Jimmy Xiangji Huang, and Enamul Hoque. 2020. Contextualized Embeddings based Transformer Encoder for Sentence Similarity Modeling in Answer Selection Task. In Proceedings of the Twelfth Language Re- sources and Evaluation Conference, Nicoletta Calzolar...

  43. [51]

    Harry Mc Laughlin

    G. Harry Mc Laughlin. 1969. SMOG Grading-a New Readability Formula.Journal of Reading 12, 8 (1969), 639–646. http://www.jstor.org/stable/40011226

  44. [52]

    Lee and Jason Lee

    Bruce W. Lee and Jason Lee. 2023. LFTK: Handcrafted Features in Computational Linguistics. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023) , Ekaterina Kochmar, Jill Burstein, Andrea Horbach, Ronja Laarmann-Quante, N...

  45. [53]

    Junlong Li, Jinyuan Wang, Zhuosheng Zhang, and Hai Zhao. 2024. Self- Prompting Large Language Models for Zero-Shot Open-Domain QA. In Pro- ceedings of the 2024 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technolog...

  46. [54]

    Xin Li and Dan Roth. 2002. Learning question classifiers. In Proceedings of the 19th International Conference on Computational Linguistics - Volume 1 (Taipei, Taiwan) (COLING ’02). Association for Computational Linguistics, USA, 1–7. https://doi.org/10.3115/1072228.1072378

  47. [55]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013

  48. [56]

    Jimmy Lin, Xueguang Ma, Sheng-Chieh Lin, Jheng-Hong Yang, Ronak Pradeep, and Rodrigo Nogueira. 2021. Pyserini: A Python Toolkit for Reproducible Information Retrieval Research with Sparse and Dense Representations. In Proceedings of the 44th International ACM SIGIR Conference ...

  49. [57]

    Fengkai Liu and John Lee. 2023. Hybrid Models for Sentence Readability Assess- ment. In Proceedings of the 18th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2023), Ekaterina Kochmar, Jill Burstein, Andrea Horbach, Ronja Laarmann-Quante, Nitin Ma...

  50. [58]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv e-prints , Article arXiv:1907.11692 (July 2019), arXiv:1907.11692 pages....

  51. [59]

    Zhuang Liu, Kaiyu Huang, Degen Huang, and Jun Zhao. 2020. Semantics- reinforced networks for question generation. In ECAI 2020. IOS Press, 2078– 2084

  52. [60]

    Vaibhav Mavi, Anubhav Jangra, and Adam Jatowt. 2024. Multi-hop Question Answering. Found. Trends Inf. Retr. 17, 5 (2024), 457–586. https://doi.org/10. 1561/1500000102

  53. [61]

    Jessica McBroom, Irena Koprinska, and Kalina Yacef. 2021. A Survey of Auto- mated Programming Hint Generation: The HINTS Framework. ACM Comput. Surv. 54, 8, Article 172 (Oct. 2021), 27 pages. https://doi.org/10.1145/3469885

  54. [62]

    Jamshid Mozafari, Abdelrahman Abdallah, Bhawna Piryani, and Adam Jatowt

  55. [63]

    Jamshid Mozafari, Afsaneh Fatemi, and Parham Moradi. 2020. A Method For Answer Selection Using DistilBERT And Important Words. In 2020 6th Inter- national Conference on Web Research (ICWR) . 72–76. https://doi.org/10.1109/ ICWR49608.2020.9122302

  56. [64]

    Jamshid Mozafari, Afsaneh Fatemi, and Mohammad Ali Nematbakhsh. 2021. BAS: An Answer Selection Method Using BERT Language Model. Journal of Computing and Security 8, 2 (2021), 1–18. https://doi.org/10.22108/jcs.2021. 128002.1066

  57. [65]

    Jamshid Mozafari, Florian Gerhold, and Adam Jatowt. 2024. Using Large Lan- guage Models in Automatic Hint Ranking and Generation Tasks. arXiv e- prints, Article arXiv:2412.01626 (Dec. 2024), arXiv:2412.01626 pages. https: //doi.org/10.48550/arXiv.2412.01626

  58. [67]

    Jamshid Mozafari, Mohammad Ali Nematbakhsh, and Afsaneh Fatemi. 2019. Attention-based Pairwise Multi-Perspective Convolutional Neural Network for Answer Selection in Question Answering. arXiv e-prints , Article arXiv:1909.01059 (Sept. 2019), arXiv:1909.01059 pages. https://doi...

  59. [68]

    Exploring Hint Generation Approaches for Open-Domain Question An- swering. (Nov. 2024), 9327–9352. https://doi.org/10.18653/v1/2024.findings- emnlp.546 SIGIR ’25, July 13–18, 2025, Padova, IT Mozafari et al

  60. [69]

    Tarek Naous, Michael J Ryan, Anton Lavrouk, Mohit Chandra, and Wei Xu. 2024. ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment. In Proceedings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing , Yaser Al-Onai...

  61. [70]

    Minh Nguyen, K. C. Kishan, Toan Nguyen, Ankit Chadha, and Thuy Vu. 2023. Efficient Fine-Tuning Large Language Models for Knowledge-Aware Response Planning. In Machine Learning and Knowledge Discovery in Databases: Research Track: European Conference, ECML PKDD 2023, Turin, Ita...

  62. [71]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Alt- man, et al. 2023. GPT-4 Technical Report.arXiv e-prints, Article arXiv:2303.08774 (March 2023), arXiv:2303.08774 pages. https://doi...

  63. [72]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (Philadel- phia, Pennsylvania) (ACL ’02). Association for C...

  64. [73]

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu F...

  65. [74]

    Nikahat Mulla and Prachi Gharpure. 2023. Automatic question generation: a review of methodologies, datasets, evaluation metrics, and applications. Prog. in Artif. Intell. 12, 1 (Jan. 2023), 1–32. https://doi.org/10.1007/s13748-023-00295-9

  66. [75]

    Bhawna Piryani, Jamshid Mozafari, and Adam Jatowt. 2024. ChroniclingAmeri- caQA: A Large-scale Question Answering Dataset based on Historical American Newspaper Pages. In Proceedings of the 47th International ACM SIGIR Confer- ence on Research and Development in Information Re...

  67. [76]

    Price, Yihuan Dong, Rui Zhi, Benjamin Paaßen, Nicholas Lytle, Veronica Cateté, and Tiffany Barnes

    Thomas W. Price, Yihuan Dong, Rui Zhi, Benjamin Paaßen, Nicholas Lytle, Veronica Cateté, and Tiffany Barnes. 2019. A Comparison of the Quality of Data-Driven Programming Hint Generation Algorithms. International Journal of Artificial Intelligence in Education 29, 3 (01 Aug 201...

  68. [77]

    Vasin Punyakanok, Dan Roth, and Wen-tau Yih. 2004. Natural language in- ference via dependency tree mapping: An application to question answering. (2004)

  69. [78]

    Libo Qin, Qiguang Chen, Xiachong Feng, Yang Wu, Yongheng Zhang, Yinghui Li, Min Li, Wanxiang Che, and Philip S. Yu. 2024. Large Language Models Meet NLP: A Survey. arXiv e-prints , Article arXiv:2405.12819 (May 2024), arXiv:2405.12819 pages. https://doi.org/10.48550/arXiv.2405.12819

  70. [79]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 1, Article 140 (Jan. 2020), 67 pages

  71. [80]

    Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Alessandro Moschitti, Bo Pang, and Walter Daelemans (Eds.). Asso...

  72. [81]

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Rad- ford, Mark Chen, and Ilya Sutskever. 2021. Zero-Shot Text-to-Image Generation. In Proceedings of the 38th International Conference on Machine Learning (Pro- ceedings of Machine Learning Research, V...

  73. [82]

    Sahana Ramnath, Preksha Nema, Deep Sahni, and Mitesh M. Khapra. 2020. Towards Interpreting BERT for Reading Comprehension Based QA. arXiv e- prints, Article arXiv:2010.08983 (Oct. 2020), arXiv:2010.08983 pages. https: //doi.org/10.48550/arXiv.2010.08983 arXiv:2010.08983 [cs.CL]

  74. [83]

    David Rau, Hervé Déjean, Nadezhda Chirkova, Thibault Formal, Shuai Wang, Vassilina Nikoulina, and Stéphane Clinchant. 2024. BERGEN: A Benchmark- ing Library for Retrieval-Augmented Generation. arXiv e-prints , Article arXiv:2407.01102 (July 2024), arXiv:2407.01102 pages. https...

  75. [84]

    Siva Reddy, Danqi Chen, and Christopher D. Manning. 2019. CoQA: A Conver- sational Question Answering Challenge. Transactions of the Association for Com- putational Linguistics 7 (2019), 249–266. https://doi.org/10.1162/tacl_a_00266

  76. [85]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Em- pirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-I...

  77. [86]

    Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018. Know What You Don’t Know: Unanswerable Questions for SQuAD. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , Iryna Gurevych and Yusuke Miyao (Eds.). Associa...

  78. [87]

    Devendra Sachan, Mike Lewis, Mandar Joshi, Armen Aghajanyan, Wen-tau Yih, Joelle Pineau, and Luke Zettlemoyer. 2022. Improving Passage Retrieval with Zero-Shot Question Generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav...

  79. [88]

    R. J. Senter and E. A. Smith. 1967.Automated Readability Index. Technical Report AMRL-TR-66-220. Aerospace Medical Research Laboratories, Wright-Patterson Air Force Base. https://apps.dtic.mil/sti/citations/AD0667273

  80. [89]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language...

  81. [90]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...

  82. [91]

    Harish Tayyar Madabushi and Mark Lee. 2016. High Accuracy Rule-based Question Classification using Question Syntax and Semantics. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, Yuji Matsumoto and Rashmi Prasad (...

  83. [92]

    Anna Rogers, Matt Gardner, and Isabelle Augenstein. 2023. QA Dataset Ex- plosion: A Taxonomy of NLP Resources for Question Answering and Reading Comprehension. ACM Comput. Surv. 55, 10, Article 197 (Feb. 2023), 45 pages. https://doi.org/10.1145/3560260

  84. [93]

    Asahi Ushio, Fernando Alva-Manchego, and Jose Camacho-Collados. 2023. A Practical Toolkit for Multilingual Question and Answer Generation. In Proceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), Danushka B...

  85. [94]

    Sowmya Vajjala and Ivana Lučić. 2018. OneStopEnglish corpus: A new corpus for automatic readability assessment and text simplification. In Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications, Joel Tetreault, Jill Burstein, Ekat...

  86. [95]

    Cunxiang Wang, Pai Liu, and Yue Zhang. 2021. Can Generative Pre-trained Language Models Serve As Knowledge Bases for Closed-book QA?. In Proceed- ings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Nat...

  87. [96]

    Jiexin Wang, Adam Jatowt, and Masatoshi Yoshikawa. 2022. ArchivalQA: A Large-scale Benchmark Dataset for Open-Domain Question Answering over Historical News Collections (SIGIR ’22). Association for Computing Machinery, New York, NY, USA, 3025–3035. https://doi.org/10.1145/3477...

  88. [97]

    Luqi Wang, Kaiwen Zheng, Liyin Qian, and Sheng Li. 2022. A Survey of Extrac- tive Question Answering. In 2022 International Conference on High Performance Big Data and Intelligent Systems (HDIS) . 147–153. https://doi.org/10.1109/ HDIS56859.2022.9991478

  89. [98]

    Usher and Frank Pajares

    Ellen L. Usher and Frank Pajares. 2006. Sources of academic and self-regulatory efficacy beliefs of entering middle school students. Contemporary Educational Psychology 31, 2 (2006), 125–141. https://doi.org/10.1016/j.cedpsych.2005.03.002

  90. [99]

    Peng Xu, Davis Liang, Zhiheng Huang, and Bing Xiang. 2021. Attention-guided Generative Models for Extractive Question Answering. arXiv e-prints, Article arXiv:2110.06393 (Oct. 2021), arXiv:2110.06393 pages. https://doi.org/10.48550/ arXiv.2110.06393 arXiv:2110.06393 [cs.CL]

  91. [100]

    Yi Yang, Wen-tau Yih, and Christopher Meek. 2015. WikiQA: A Challenge Dataset for Open-Domain Question Answering. In Proceedings of the 2015 Con- ference on Empirical Methods in Natural Language Processing , Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association ...

  92. [101]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language...

  93. [102]

    Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, and Meng Jiang. 2023. Generate rather than Retrieve: Large Language Models are Strong Context Genera- tors. In The Eleventh International Conference on Learning Representat...

  94. [103]

    Xingdi Yuan, Tong Wang, Yen-Hsiang Wang, Emery Fine, Rania Abdelghani, Hélène Sauzéon, and Pierre-Yves Oudeyer. 2023. Selecting Better Samples from Pre-trained LLMs: A Case Study on Question Generation. In Findings of the Association for Computational Linguistics: ACL 2023 , A...

  95. [104]

    Jiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai, Zhaoyang Liu, Wenhai Wang, Zhe Chen, Xizhou Zhu, Lewei Lu, Tong Lu, Ping Luo, Yu Qiao, and Jifeng Dai. 2024. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks. arXiv e-pr...

  96. [105]

    Weinberger, and Yoav Artzi

    Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. InInternational Conference on Learning Representations . https://openreview.net/forum?id= SkeHuCVFDr

  97. [106]

    Tiancheng Zhao, Xiaopeng Lu, and Kyusong Lee. 2021. SPARTA: Efficient Open-Domain Question Answering via Sparse Transformer Matching Re- trieval. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Languag...

  98. [107]

    Chang Zong, Yuchen Yan, Weiming Lu, Jian Shao, Eliot Huang, Heng Chang, and Yueting Zhuang. 2024. Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering. arXiv e-prints, Article arXiv:2402.14320 (Feb. 2024), arXiv:2402.14320 pages...

  99. [110]

    Ruqing Zhang, Jiafeng Guo, Lu Chen, Yixing Fan, and Xueqi Cheng. 2021. A Review on Question Generation from Natural Language Text. ACM Trans. Inf. Syst. 40, 1, Article 14 (Sept. 2021), 43 pages. https://doi.org/10.1145/3468889

  100. [127]

    https://doi.org/10.1186/s40537-023-00802-8

  101. [157]

    https://www.taylorfrancis.com/chapters/edit/10.4324/9780203082744- 13/role-self-efficacy-goal-based-motivation-albert-bandura

  102. [2020]

    In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.)

    Template-Based Question Generation from Retrieved Sentences for Im- proved Unsupervised Question Answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Assoc...

  103. [2021]

    In 2021 IEEE 10th Global Conference on Con- sumer Electronics (GCCE)

    Question Generation using Knowledge Graphs with the T5 Language Model and Masked Self-Attention. In 2021 IEEE 10th Global Conference on Con- sumer Electronics (GCCE) . 85–87. https://doi.org/10.1109/GCCE53005.2021. 9621874

  104. [2024]

    arXiv e-prints , Article arXiv:2404.04728 (April 2024), arXiv:2404.04728 pages

    Navigating the Landscape of Hint Generation Research: From the Past to the Future. arXiv e-prints , Article arXiv:2404.04728 (April 2024), arXiv:2404.04728 pages. https://doi.org/10.48550/arXiv.2404.04728

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.