Pith. sign in

REVIEW 3 major objections 5 minor 81 references

Exploiting Intrinsic Duality for Multi-Hop Question Generation

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Multi-hop question generation improves when the generator is trained jointly with a question-answering head, using bidirectional alignment and contrastive losses that keep generated questions and answers mutually predictable.

desk verdict Solid empirical paper with a plausible joint-training idea; the central mechanism is under-specified and may be aligning gold states, not generated outputs. read the letter →

arxiv 2608.00712 v1 pith:RCDM3YLL submitted 2026-08-01 cs.CL

classification cs.CL
keywords multi-hopquestiongenerationansweringtaskdualitybidirectionalalignmentcontrastivelearningHotpotQAMuSiQueanswerability
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that multi-hop question generation (MQG) gets better when the generator is trained inside the same model as a question-answering (QA) head, rather than treating answer information as a fixed input. The proposed QQ framework runs both tasks on one architecture: the same parameters produce a multi-hop question from documents plus a target answer, and produce an answer from documents plus a question. Two auxiliary objectives enforce the claimed duality: a bidirectional alignment loss that makes question states and answer states mutually predictable in a shared latent space, and a contrastive loss that pulls matched question–answer representations together while pushing unmatched pairs apart. On HotpotQA and MuSiQue, across GPT and LLaMA backbones, the paper reports consistent gains in BLEU-4, METEOR, ROUGE-L, and BERTScore over the same backbones trained on MQG alone, and higher Exact Match and F1 when generated questions are answered by the unified model. If the claim holds, it offers a practical way to exploit the interdependence of asking and answering instead of keeping the two tasks separate.

What carries the argument

The load-bearing mechanism is a unified encoder–decoder with two output heads, one for MQG and one for QA, trained jointly with four losses. The alignment loss maps the hidden states of generated questions and answers through a projection module followed by a predictor (both two-layer MLPs with ReLU) into a shared predictive space, then maximizes cosine similarity between positive question–answer pairs and minimizes it against negatives in the batch (Eqs. 5–8). The contrastive loss maps the same hidden states through the projection module into a separate contrastive space, pulling paired representations together and pushing unpaired ones apart (Eqs. 9–10). Both losses operate on the outputs the model itself generated, so the ‘intrinsic duality’ is operationalized as a consistency constraint between the two task directions rather than as an external QA check.

What would settle it

Train the same unified backbone with $L_{Q \leftrightarrow A}$ and $L_{\mathrm{CL}}$ but permute the positive pairing so that each question state is aligned with a random answer state from another example; if BLEU-4, METEOR, ROUGE-L, and answerability Exact Match and F1 stay at the same level or improve on the dev sets, the specific question–answer pairing is not the source of the reported gains.

Watch

Extended reading notes

Core claim

The central claim is that MQG and QA are intrinsically dual, and that jointly training them with mutual-predictability constraints transfers utility from the answer-direction task into the question-generation task. Formally, the authors define alignment losses $L_{Q \to A}$ and $L_{A \to Q}$ (Eqs. 5–6), averaged into $L_{Q \leftrightarrow A}$ (Eq. 8), which treat the question state and the QA-generated answer state from the same training example as a positive pair in a learned predictive space; a contrastive loss $L_{\mathrm{CL}}$ (Eq. 9) does the same in a second contrastive space with temperature $\tau_2$. The full objective is $L = L_{\mathrm{MQG}} + L_{\mathrm{QA}} + L_{Q \leftrightarrow A} + L_{\mathrm{CL}}$, and at inference only the MQG head runs. The authors report that this improves question quality on both datasets and across model scales, and that it improves the answerability of generated questions as measured by Exact Match and F1. They also state a limitation: because one answer can correspond to many valid questions, the positive-pair signal is broader than in standard contrastive learning, and its theoretical basis ‘remains to be further refined.’

Load-bearing premise

The load-bearing premise is that a generated question and the QA head's answer to it, taken from the same training example, form a genuinely aligned positive pair, so forcing their latent states to match improves the question rather than reinforcing whatever superficial agreement the two heads already have.

Editorial extensions

If this is right

  • Because the QA head is trained jointly and not bolted on, the framework yields a single model that both generates and answers questions, so answerability checks on generated questions can be run without a separate QA system.
  • The reported gains appear across GPT and LLaMA backbones of different sizes and on both HotpotQA and MuSiQue, which suggests the alignment signal transfers across architectures and reasoning depths rather than fitting one configuration.
  • Ablations in the paper show that the bidirectional alignment loss contributes more consistently than the contrastive loss, with the contrastive loss occasionally hurting individual metrics, so the alignment term is the primary driver of the reported improvement.
  • On MuSiQue, the largest relative gains in several metrics appear in the 2-hop and 4-hop subsets, which the authors attribute to higher reasoning complexity, indicating the method is aimed at harder multi-hop cases.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: a natural tightening is to define positive pairs by answer verification rather than by training-instance identity, for example sampling several candidate questions, answering each, and keeping the one whose answer matches the target answer, which would address the paper's own caveat that one answer admits many valid questions.
  • Beyond the paper: because the alignment loss acts on generated output states rather than on ground-truth text, the same objective could serve in semi-supervised settings where reference questions are scarce but documents and answers are available, with the QA head acting as a consistency teacher.
  • Beyond the paper: the bidirectional-consistency idea should port to other dual generation tasks, such as summarization or information extraction, where the reverse direction supplies a cheap check on whether the generated output preserves the content it was meant to capture.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes QQ, a framework that jointly trains multi-hop question generation (MQG) and question answering (QA) in a unified architecture, using a bidirectional alignment loss (Eqs. 5-8) and a contrastive loss (Eqs. 9-10) to pull paired question-answer representations together and push unpaired ones apart. The authors evaluate QQ on HotpotQA and MuSiQue across GPT and LLaMA backbones with several automatic metrics, answerability (EM/F1), and human evaluation, reporting consistent gains relative to vanilla backbones and better performance than several published baselines.

Significance. The paper identifies a plausible and under-explored direction: exploiting the task-level duality between MQG and QA rather than using QA only as an auxiliary signal. If the proposed mechanism genuinely aligns the representations of generated questions with generated answers, the framework would be a useful contribution to multi-hop question generation. The evaluation is broad, covering multiple backbone sizes, two datasets, ablation studies, answerability checks, and human judgments, and the paper is honest about some limitations. However, the core mechanism is underspecified in the current write-up, and the empirical evidence, while mostly positive, is not statistically grounded. With clarification and additional validation, the idea has real potential.

major comments (3)
  1. [Bidirectional Alignment and Contrastive Objectives, Eqs. (5)-(8)] The variables hq and ha are never defined; Eqs. (1)-(2) only introduce token-level hidden states h_q_t and h_a_t. As written, the positive pairs (zq, ha_+) and (za, hq_+) appear to be formed from the hidden states of the gold question and gold answer in the teacher-forced training pass, not from states produced by decoding a question from (D,A) and then answering that question. Since the NLL objectives in Eqs. (3)-(4) are teacher-forced, the alignment losses as specified align reference representations. The abstract's claim that the framework enforces "strict mutual correspondence between the questions generated by the MQG model and the answers produced by the QA model" is therefore not supported by the described implementation. Please define hq and ha explicitly, and either describe a free-run or sampling procedure that produces output-level states (with a clear gradient path) or revise the claims to describe reference-level alignment and discuss why that yields the observed transfer.
  2. [Table 2] The claim that QQ "consistently improves" and "significantly improves" question generation is not supported by the reported numbers. Several cells decrease, including LLaMA(8B) SF ROUGE-L by 0.28, MuSiQue LLaMA(1B) 2-hop BLEU-4 by 0.76 and ROUGE-L by 0.62, LLaMA(3B) 3-hop METEOR by 0.55, and LLaMA(8B) 4-hop BLEU-4 by 0.73. No confidence intervals, significance tests, or multiple-seed variance are reported. Please provide paired significance tests or variance estimates, or temper the wording from "significantly" to a more hedged claim with an explicit discussion of the negative cells.
  3. [Eqs. (5)-(9) and Implementation Details] The construction of the negative sets V_a, V_q, and S_a is not specified. If these are in-batch negatives, then with the reported batch sizes (e.g., 2 for LLaMA-8B on MuSiQue), the effective number of negatives is very small and the behavior of the alignment and contrastive losses changes qualitatively. The paper should state the negative sampling strategy, the batch size used for the contrastive/alignment heads, and how negatives are shared across the MQG and QA directions.
minor comments (5)
  1. [Table 3 caption] The caption says "Bold indicates L_Q↔A dominance, while underlined indicates LCL dominance," but the table appears to use bold and underlining to mark improvement magnitudes over the vanilla backbone; please clarify the intended meaning.
  2. [Case Study, main text] The case study in the main text refers to "Table 8," but the first case study table is numbered Table 5; the cross-reference should be fixed.
  3. [Figures 4-6] The legend in Figures 4-6 uses the label "LLaMA-3-8B w/ QQ," which is inconsistent with "LLaMA(8B) w/ QQ" used in the text.
  4. [Table 9] Table 9 uses "LLaMA-3(8B)" while the rest of the paper uses "LLaMA(8B)"; please standardize the notation.
  5. [Prompting Large Language Models (supplementary)] The section introduces Qwen-plus but does not provide a citation; please add a reference for the model.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the QQ gains are empirical results of a joint-training objective; no equation reduces a reported metric to a fitted input or to a self-cited claim.

full rationale

The paper contains no derivation chain in which an output quantity is defined in terms of the quantity it is said to predict. The training objective L = L_MQG + L_QA + L_Q↔A + L_CL (Eq. 11) combines standard negative log-likelihoods (Eqs. 3-4) with auxiliary similarity losses (Eqs. 5-9). The auxiliary losses use hidden states h_q and h_a from the same (D, A, Q) training triple; this makes the positive-pair signal broad, as the authors concede in the conclusion ('a single answer may correspond to multiple valid questions, our predictability-based alignment loss defines a broader positive-pair signal than standard contrastive learning, whose theoretical basis remains to be further refined'), but it does not make any evaluation number a fitted input. BLEU-4, METEOR, ROUGE-L, BERTScore, EM and F1 are all computed on held-out generated text against reference questions or answers; none of these metrics appears in the loss, and the losses are not constructed to optimize them. The self-citations present—the HotpotQA split and DPKG baselines from Li, Zhang, and Kong (2025a,b), and the evaluation convention 'Following (Li, Zhang, and Kong 2025b)'—are data and evaluation conventions, not load-bearing arguments that force the claimed improvement. No uniqueness theorem or ansatz is imported from the authors' prior work to rule out alternatives. A separate validity concern, namely whether Eqs. 5-8 align teacher-forced gold states rather than decoded generated outputs, would affect the accuracy of the mechanism description, but it is not circularity: the reported outcomes are not equal to any training objective by construction.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central claim rests on three domain assumptions: that mutual predictability is a valid training signal, that the chosen automatic metrics capture question quality, and that joint training on MQG and QA does not interfere. The only hand-chosen hyperparameters that directly affect the alignment mechanism are the temperatures tau1 and tau2, both set to 0.07 without sensitivity analysis. No new entities are postulated.

free parameters (3)
  • tau1 (alignment temperature) = 0.07
    Set to 0.07 following standard contrastive learning practice; no sensitivity analysis is reported, so the contribution of the alignment loss could depend on this hand-chosen value.
  • tau2 (contrastive temperature) = 0.07
    Same as tau1, fixed at 0.07 with no grid search or sensitivity analysis.
  • LoRA rank and alpha = r=8, alpha=32
    Used for LLaMA fine-tuning; values are standard, but no study of their effect on the alignment losses is provided.
assumptions (3)
  • domain assumption Mutual predictability between generated question states and answer states is a valid proxy for question quality.
    This is the core motivation of Eqs. (5)-(8); the paper's own conclusion says the theoretical basis remains to be refined.
  • domain assumption BLEU-4, METEOR, ROUGE-L, and BERTScore reflect meaningful improvements in question quality.
    These are used for all headline comparisons; only a 300-sample human eval partially validates them.
  • domain assumption A single shared model can be trained on both MQG and QA objectives without harmful interference.
    The unified architecture assumes joint training is beneficial; the paper evaluates this empirically on selected backbones, but no analysis of interference is given.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploiting Intrinsic Duality for Multi-Hop Question Generation." pith.science (2026). https://pith.science/paper/RCDM3YLL

@misc{pith2026260800712,
  author       = {Pith},
  title        = {Pith review of: Exploiting Intrinsic Duality for Multi-Hop Question Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RCDM3YLL}},
  note         = {Machine review of arXiv:2608.00712}
}
read the original abstract

Multi hop question generation (MQG) aims to generate questions from multiple given documents and target answers, whereas question answering (QA) focuses on deriving answers from documents given specific questions. Although MQG and QA are inherently dual tasks, most existing MQG studies largely overlook this intrinsic duality. To address this limitation, we propose QQ, a novel framework that exploits the duality between Question and answer for multi hop Question generation. Specifically, QQ employs a unified architecture functioning simultaneously as both an MQG and a QA model to fully leverage their interdependence. Our framework is driven by two key mechanisms: (i) enforcing bidirectional alignment constraints to ensure strict mutual correspondence between the questions generated by the MQG model and the answers produced by the QA model; and (ii) applying contrastive learning to pull paired question answer representations closer while pushing unpaired ones apart, thereby reinforcing this correspondence. Extensive automatic and human evaluations on the HotpotQA and MuSiQue datasets demonstrate that the QQ framework significantly improves the quality of generated multi hop questions.

Figures

Figures reproduced from arXiv: 2608.00712 by the authors.

Figure 1
Figure 1. An illustration of a multi-hop question generated via [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Our QQ framework. QQ framework employs a unified architecture that functions simultaneously as both an MQG and a QA model: it generates a multi-hop question given documents and a target answer, and conversely generates an answer given documents and a provided question. During the MQG process, question states are mapped into a predictive space via the projection and predictor modules to evaluate their predictability … view at source ↗
Figure 3
Figure 3. Results of question answering with generated questions on the HotpotQA dataset. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Results of question answering using generated [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Results of question answering using generated [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Results of question answering using generated [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Pairwise human evaluation results. Evaluation Metrics Following prior work (Ding, Hong, and Yao 2024; Li, Zhang, and Kong 2025a), we employ the following automatic met￾rics: • BLEU-4 (Papineni et al. 2002) for lexical precision; • METEOR (Banerjee and Lavie 2005) for s…
Figure 8
Figure 8. Figure 8: Human evaluation form (pairwise human evaluation). Framework A refers to LLaMA [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Human evaluation form (pointwise human evaluation). Model_1, Model_2, Model_3, and Model_4 denote DPKG, [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 39 canonical work pages

  1. [1]

    Banerjee, S.; and Lavie, A. 2005. METEOR : An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Goldstein, J.; Lavie, A.; Lin, C.-Y.; and Voss, C., eds., Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 65--72. Ann Arbor, Michigan: Association ...

  2. [2]

    Chali, Y.; and Hasan, S. A. 2012. Towards Automatic Topical Question Generation. In Kay, M.; and Boitet, C., eds., Proceedings of COLING 2012 , 475--492. Mumbai, India: The COLING 2012 Organizing Committee

  3. [3]

    Chen, Y.; Wu, L.; and Zaki, M. J. 2020. Reinforcement Learning Based Graph-to-Sequence Model for Natural Question Generation. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net

  4. [4]

    Chen, Y.; Wu, L.; and Zaki, M. J. 2023. Toward subgraph-guided knowledge graph question generation with graph neural networks. IEEE Transactions on Neural Networks and Learning Systems

  5. [5]

    Ding, C.; Hong, Y.; and Yao, J. 2024. SGCM : Salience-Guided Context Modeling for Question Generation. In Calzolari, N.; Kan, M.-Y.; Hoste, V.; Lenci, A.; Sakti, S.; and Xue, N., eds., Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 14755--14762. Torino, Italia: ELR...

  6. [6]

    Du, X.; Shao, J.; and Cardie, C. 2017. Learning to Ask: Neural Question Generation for Reading Comprehension. In Barzilay, R.; and Kan, M.-Y., eds., Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1342--1352. Vancouver, Canada: Association for Computational Linguistics

  7. [7]

    Dubey, A.; Jauhri, A.; Pandey, A.; and et al. 2024. The Llama 3 Herd of Models. CoRR, abs/2407.21783

  8. [8]

    Fei, Z.; Zhang, Q.; Gui, T.; Liang, D.; Wang, S.; Wu, W.; and Huang, X. 2022. CQG : A Simple and Effective Controlled Generation Framework for Multi-hop Question Generation. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds., Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6896--6906. Du...

Show all 81 references
  1. [9]

    Fei, Z.; Zhang, Q.; and Zhou, Y. 2021. Iterative GNN -based Decoder for Question Generation. In Moens, M.-F.; Huang, X.; Specia, L.; and Yih, S. W.-t., eds., Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2573--2582. Online and Punta Ca...

  2. [10]

    Grill, J.-B.; Strub, F.; Altch \'e , F.; Tallec, C.; Richemond, P.; Buchatskaya, E.; Doersch, C.; Avila Pires, B.; Guo, Z.; Gheshlaghi Azar, M.; et al. 2020. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems...

  3. [11]

    Guo, S.; Liao, L.; Li, C.; and Chua, T.-S. 2024. A Survey on Neural Question Generation: Methods, Applications, and Prospects. In Larson, K., ed., Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24 , 8038--8047. International Jo...

  4. [12]

    He, K.; Fan, H.; Wu, Y.; Xie, S.; and Girshick, R. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 9729--9738

  5. [13]

    Heilman, M.; and Smith, N. A. 2010. Good Question! Statistical Ranking for Question Generation. In Kaplan, R.; Burstein, J.; Harper, M.; and Penn, G., eds., Human Language Technologies: The 2010 Annual Conference of the North A merican Chapter of the Association for Computatio...

  6. [14]

    J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al

    Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2022. Lora: Low-rank adaptation of large language models. ICLR, 1(2): 3

  7. [15]

    Hwang, S.; Kim, Y.; and Lee, G. G. 2024. Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling. In Calzolari, N.; Kan, M.; Hoste, V.; Lenci, A.; Sakti, S.; and Xue, N., eds., Proceedings of the 2024 Joint International Confere...

  8. [16]

    Kim, K.; Park, S.; Lee, J.; and Lee, J. 2024. Non-Essential Is NE cessary: Order-agnostic Multi-hop Question Generation. In Calzolari, N.; Kan, M.-Y.; Hoste, V.; Lenci, A.; Sakti, S.; and Xue, N., eds., Proceedings of the 2024 Joint International Conference on Computational Li...

  9. [17]

    Li, M.; Zhang, L.; and Kong, F. 2025 a . Multi-Hop Question Generation via Dual-Perspective Keyword Guidance. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Findings of the Association for Computational Linguistics: ACL 2025, 10096--10112. Vienna, Austria: As...

  10. [18]

    Li, M.; Zhang, L.; and Kong, F. 2025 b . Multi-Hop Question Generation via Dual-Perspective Keyword Guidance. arXiv:2505.15299

  11. [19]

    Liang, Y.; Wang, J.; Zhu, H.; Wang, L.; Qian, W.; and Lan, Y. 2023. Prompting Large Language Models with Chain-of-Thought for Few-Shot Knowledge Base Question Generation. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in N...

  12. [20]

    Lin, C.-Y. 2004. ROUGE : A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out, 74--81. Barcelona, Spain: Association for Computational Linguistics

  13. [21]

    Liu, N.; Wang, Z.; and Baraniuk, R. 2024. Synthetic Context Generation for Question Generation. arXiv:2406.13188

  14. [22]

    Mazidi, K.; and Nielsen, R. D. 2014. Linguistic Considerations in Automatic Question Generation. In Toutanova, K.; and Wu, H., eds., Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 321--326. Baltimore, Maryland:...

  15. [23]

    Murakhovs ' ka, L.; Wu, C.-S.; Laban, P.; Niu, T.; Liu, W.; and Xiong, C. 2022. M ix QG : Neural Question Generation with Mixed Answer Types. In Carpuat, M.; de Marneffe, M.-C.; and Meza Ruiz, I. V., eds., Findings of the Association for Computational Linguistics: NAACL 2022, ...

  16. [24]

    Pan, L.; Xie, Y.; Feng, Y.; Chua, T.-S.; and Kan, M.-Y. 2020. Semantic Graphs for Generating Deep Questions. In Jurafsky, D.; Chai, J.; Schluter, N.; and Tetreault, J., eds., Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1463--1475. O...

  17. [25]

    Papineni, K.; Roukos, S.; Ward, T.; and Zhu, W. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA , 311--318. ACL

  18. [26]

    Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8): 9

  19. [27]

    Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016. SQ u AD : 100,000+ Questions for Machine Comprehension of Text. In Su, J.; Duh, K.; and Carreras, X., eds., Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2383--2392. Austin, Te...

  20. [28]

    Song, L.; Wang, Z.; Hamza, W.; Zhang, Y.; and Gildea, D. 2018. Leveraging Context Information for Natural Question Generation. In Walker, M.; Ji, H.; and Stent, A., eds., Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Lin...

  21. [29]

    Su, D.; Xu, P.; and Fung, P. 2022. Qa4qg: Using question answering to constrain multi-hop question generation. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 8232--8236. IEEE

  22. [30]

    Su, D.; Xu, Y.; Dai, W.; Ji, Z.; Yu, T.; and Fung, P. 2020. Multi-hop Question Generation with Graph Convolutional Network. In Cohn, T.; He, Y.; and Liu, Y., eds., Findings of the Association for Computational Linguistics: EMNLP 2020, 4636--4647. Online: Association for Comput...

  23. [31]

    Sun, X.; Liu, J.; Lyu, Y.; He, W.; Ma, Y.; and Wang, S. 2018. Answer-focused and Position-aware Neural Question Generation. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,...

  24. [32]

    Trivedi, H.; Balasubramanian, N.; Khot, T.; and Sabharwal, A. 2022. M u S i Q ue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics, 10: 539--554

  25. [33]

    Ushio, A.; Alva-Manchego, F.; and Camacho-Collados, J. 2023. A Practical Toolkit for Multilingual Question and Answer Generation. In Bollegala, D.; Huang, R.; and Ritter, A., eds., Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume ...

  26. [34]

    Vaswani, A. 2017. Attention is all you need. Advances in Neural Information Processing Systems

  27. [35]

    Xia, Z.; Gou, Q.; Yu, B.; Yu, H.; Huang, F.; Li, Y.; and Nguyen, C. 2023. Improving Question Generation with Multi-level Content Planning. In Bouamor, H.; Pino, J.; and Bali, K., eds., Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6...

  28. [36]

    Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W.; Salakhutdinov, R.; and Manning, C. D. 2018. H otpot QA : A Dataset for Diverse, Explainable Multi-hop Question Answering. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Conference on...

  29. [37]

    Q.; and Artzi, Y

    Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020. BERTScore: Evaluating Text Generation with BERT . In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net

  30. [38]

    Zhao, Y.; Ni, X.; Ding, Y.; and Ke, Q. 2018. Paragraph-level Neural Question Generation with Maxout Pointer and Gated Self-attention Networks. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Conference on Empirical Methods in Natural L...

  31. [39]

    Zhou, Q.; Yang, N.; Wei, F.; Tan, C.; Bao, H.; and Zhou, M. 2018. Neural Question Generation from Text: A Preliminary Study. In Huang, X.; Jiang, J.; Zhao, D.; Feng, Y.; and Hong, Y., eds., Natural Language Processing and Chinese Computing, 662--671. Cham: Springer Internation...

  32. [40]

    2025 , eprint=

    Multi-Hop Question Generation via Dual-Perspective Keyword Guidance , author=. 2025 , eprint=

  33. [41]

    Prompting Large Language Models with Chain-of-Thought for Few-Shot Knowledge Base Question Generation

    Liang, Yuanyuan and Wang, Jianing and Zhu, Hanlun and Wang, Lei and Qian, Weining and Lan, Yunshi. Prompting Large Language Models with Chain-of-Thought for Few-Shot Knowledge Base Question Generation. Proceedings of the 2023 Conference on Empirical Methods in Natural Language...

  34. [42]

    Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence,

    A Survey on Neural Question Generation: Methods, Applications, and Prospects , author =. Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence,. 2024 , month =. doi:10.24963/ijcai.2024/889 , url =

  35. [43]

    H otpot QA : A Dataset for Diverse, Explainable Multi-hop Question Answering

    Yang, Zhilin and Qi, Peng and Zhang, Saizheng and Bengio, Yoshua and Cohen, William and Salakhutdinov, Ruslan and Manning, Christopher D. H otpot QA : A Dataset for Diverse, Explainable Multi-hop Question Answering. Proceedings of the 2018 Conference on Empirical Methods in Na...

  36. [44]

    SGCM : Salience-Guided Context Modeling for Question Generation

    Ding, Chuyao and Hong, Yu and Yao, Jianmin. SGCM : Salience-Guided Context Modeling for Question Generation. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 2024

  37. [45]

    Non-Essential Is NE cessary: Order-agnostic Multi-hop Question Generation

    Kim, Kyungho and Park, Seongmin and Lee, Junseo and Lee, Jihwa. Non-Essential Is NE cessary: Order-agnostic Multi-hop Question Generation. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 2024

  38. [46]

    Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling , booktitle =

    Seonjeong Hwang and Yunsu Kim and Gary Geunbae Lee , editor =. Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling , booktitle =. 2024 , url =

  39. [47]

    Multi-hop Question Generation with Graph Convolutional Network

    Su, Dan and Xu, Yan and Dai, Wenliang and Ji, Ziwei and Yu, Tiezheng and Fung, Pascale. Multi-hop Question Generation with Graph Convolutional Network. Findings of the Association for Computational Linguistics: EMNLP 2020. 2020. doi:10.18653/v1/2020.findings-emnlp.416

  40. [48]

    ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

    Qa4qg: Using question answering to constrain multi-hop question generation , author=. ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=. 2022 , organization=

  41. [49]

    2025 , eprint=

    From System 1 to System 2: A Survey of Reasoning Large Language Models , author=. 2025 , eprint=

  42. [50]

    2024 , eprint=

    The Llama 3 Herd of Models , author=. 2024 , eprint=

  43. [51]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

    He, Kaiming and Fan, Haoqi and Wu, Yuxin and Xie, Saining and Girshick, Ross , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

  44. [52]

    , author=

    Lora: Low-rank adaptation of large language models. , author=. ICLR , volume=

  45. [53]

    Advances in neural information processing systems , volume=

    Bootstrap your own latent-a new approach to self-supervised learning , author=. Advances in neural information processing systems , volume=

  46. [54]

    Weinberger and Yoav Artzi , title =

    Tianyi Zhang and Varsha Kishore and Felix Wu and Kilian Q. Weinberger and Yoav Artzi , title =. 8th International Conference on Learning Representations,. 2020 , url =

  47. [55]

    ROUGE : A Package for Automatic Evaluation of Summaries

    Lin, Chin-Yew. ROUGE : A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out. 2004

  48. [56]

    METEOR : An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments

    Banerjee, Satanjeev and Lavie, Alon. METEOR : An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization. 2005

  49. [57]

    OpenAI blog , volume=

    Language models are unsupervised multitask learners , author=. OpenAI blog , volume=

  50. [58]

    Bleu: a Method for Automatic Evaluation of Machine Translation , booktitle =

    Kishore Papineni and Salim Roukos and Todd Ward and Wei. Bleu: a Method for Automatic Evaluation of Machine Translation , booktitle =. 2002 , url =. doi:10.3115/1073083.1073135 , timestamp =

  51. [59]

    , title=

    Abhimanyu Dubey and Abhinav Jauhri and Abhinav Pandey and et al. , title=. CoRR , volume=. 2024 , cdate=

  52. [60]

    Good Question! Statistical Ranking for Question Generation

    Heilman, Michael and Smith, Noah A. Good Question! Statistical Ranking for Question Generation. Human Language Technologies: The 2010 Annual Conference of the North A merican Chapter of the Association for Computational Linguistics. 2010

  53. [61]

    Towards Automatic Topical Question Generation

    Chali, Yllias and Hasan, Sadid A. Towards Automatic Topical Question Generation. Proceedings of COLING 2012. 2012

  54. [62]

    Linguistic Considerations in Automatic Question Generation

    Mazidi, Karen and Nielsen, Rodney D. Linguistic Considerations in Automatic Question Generation. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2014. doi:10.3115/v1/P14-2053

  55. [63]

    Learning to Ask: Neural Question Generation for Reading Comprehension

    Du, Xinya and Shao, Junru and Cardie, Claire. Learning to Ask: Neural Question Generation for Reading Comprehension. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. doi:10.18653/v1/P17-1123

  56. [64]

    Neural Question Generation from Text: A Preliminary Study

    Zhou, Qingyu and Yang, Nan and Wei, Furu and Tan, Chuanqi and Bao, Hangbo and Zhou, Ming. Neural Question Generation from Text: A Preliminary Study. Natural Language Processing and Chinese Computing. 2018

  57. [65]

    Answer-focused and Position-aware Neural Question Generation

    Sun, Xingwu and Liu, Jing and Lyu, Yajuan and He, Wei and Ma, Yanjun and Wang, Shi. Answer-focused and Position-aware Neural Question Generation. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. doi:10.18653/v1/D18-1427

  58. [66]

    Leveraging Context Information for Natural Question Generation

    Song, Linfeng and Wang, Zhiguo and Hamza, Wael and Zhang, Yue and Gildea, Daniel. Leveraging Context Information for Natural Question Generation. Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language ...

  59. [67]

    Paragraph-level Neural Question Generation with Maxout Pointer and Gated Self-attention Networks

    Zhao, Yao and Ni, Xiaochuan and Ding, Yuanyuan and Ke, Qifa. Paragraph-level Neural Question Generation with Maxout Pointer and Gated Self-attention Networks. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. doi:10.18653/v1/D18-1424

  60. [68]

    Zaki , title =

    Yu Chen and Lingfei Wu and Mohammed J. Zaki , title =. 8th International Conference on Learning Representations,. 2020 , url =

  61. [69]

    IEEE Transactions on Neural Networks and Learning Systems , year=

    Toward subgraph-guided knowledge graph question generation with graph neural networks , author=. IEEE Transactions on Neural Networks and Learning Systems , year=

  62. [70]

    SQ u AD : 100,000+ Questions for Machine Comprehension of Text

    Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy. SQ u AD : 100,000+ Questions for Machine Comprehension of Text. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 2016. doi:10.18653/v1/D16-1264

  63. [71]

    Iterative GNN -based Decoder for Question Generation

    Fei, Zichu and Zhang, Qi and Zhou, Yaqian. Iterative GNN -based Decoder for Question Generation. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. doi:10.18653/v1/2021.emnlp-main.201

  64. [72]

    CQG : A Simple and Effective Controlled Generation Framework for Multi-hop Question Generation

    Fei, Zichu and Zhang, Qi and Gui, Tao and Liang, Di and Wang, Sirui and Wu, Wei and Huang, Xuanjing. CQG : A Simple and Effective Controlled Generation Framework for Multi-hop Question Generation. Proceedings of the 60th Annual Meeting of the Association for Computational Ling...

  65. [73]

    Semantic Graphs for Generating Deep Questions

    Pan, Liangming and Xie, Yuxi and Feng, Yansong and Chua, Tat-Seng and Kan, Min-Yen. Semantic Graphs for Generating Deep Questions. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v1/2020.acl-main.135

  66. [74]

    Advances in Neural Information Processing Systems , year=

    Attention is all you need , author=. Advances in Neural Information Processing Systems , year=

  67. [75]

    M ix QG : Neural Question Generation with Mixed Answer Types

    Murakhovs ' ka, Lidiya and Wu, Chien-Sheng and Laban, Philippe and Niu, Tong and Liu, Wenhao and Xiong, Caiming. M ix QG : Neural Question Generation with Mixed Answer Types. Findings of the Association for Computational Linguistics: NAACL 2022. 2022. doi:10.18653/v1/2022.find...

  68. [76]

    Improving Question Generation with Multi-level Content Planning , booktitle =

    Zehua Xia and Qi Gou and Bowen Yu and Haiyang Yu and Fei Huang and Yongbin Li and Cam. Improving Question Generation with Multi-level Content Planning , booktitle =. 2023 , url =. doi:10.18653/V1/2023.FINDINGS-EMNLP.57 , timestamp =

  69. [77]

    Multi-Hop Question Generation via Dual-Perspective Keyword Guidance

    Li, Maodong and Zhang, Longyin and Kong, Fang. Multi-Hop Question Generation via Dual-Perspective Keyword Guidance. Findings of the Association for Computational Linguistics: ACL 2025. 2025

  70. [78]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    Momentum contrast for unsupervised visual representation learning , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  71. [79]

    ♫ M u S i Q ue: Multihop Questions via Single-hop Question Composition

    Trivedi, Harsh and Balasubramanian, Niranjan and Khot, Tushar and Sabharwal, Ashish. ♫ M u S i Q ue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics. 2022. doi:10.1162/tacl_a_00475

  72. [80]

    A Practical Toolkit for Multilingual Question and Answer Generation

    Ushio, Asahi and Alva-Manchego, Fernando and Camacho-Collados, Jose. A Practical Toolkit for Multilingual Question and Answer Generation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). 2023. doi:10.186...

  73. [81]

    2024 , eprint=

    Synthetic Context Generation for Question Generation , author=. 2024 , eprint=

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.