Pith. sign in

REVIEW 4 cited by

SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09939 v2 pith:L2RY3QMM submitted 2024-05-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords sciencescientificsciqagansweringdatasetframeworkllmsquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language models (LLMs). SciQAG consists of a QA generator and a QA evaluator, which work together to extract diverse and research-level questions and answers from scientific papers. Utilizing this framework, we construct a large-scale, high-quality, open-ended science QA dataset containing 188,042 QA pairs extracted from 22,743 scientific papers across 24 scientific domains. We also introduce SciQAG-24D, a new benchmark task designed to evaluate the science question-answering ability of LLMs. Extensive experiments demonstrate that fine-tuning LLMs on the SciQAG dataset significantly improves their performance on both open-ended question answering and scientific tasks. To foster research and collaboration, we make the datasets, models, and evaluation codes publicly available, contributing to the advancement of science question answering and developing more interpretable and reasoning-capable AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control

    cs.MA 2025-05 conditional novelty 7.0 of 10

    A two-tier private/shared memory system with provenance-based access control reduces redundant queries in multi-user LLM agent teams by up to 61 percent without losing accuracy.

  2. Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery

    cs.CL 2026-02 reject novelty 6.0 of 10

    DBench-Bio builds a dynamic biology benchmark from post-release abstracts, but LLM-generated gold answers and unverified per-model temporal separation undermine its claim to measure knowledge discovery.

  3. MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval

    cs.AI 2025-06 conditional novelty 6.0 of 10

    MM-R5, a 7B multimodal re-ranker trained with SFT and GRPO, achieves state-of-the-art page-level recall on MMDocIR by generating per-page reasoning chains.

  4. Towards a Large Physics Benchmark

    physics.data-an 2025-07 conditional novelty 4.0 of 10

    The paper outlines a multi-format, expert-scored living benchmark for evaluating physics understanding and creativity in LLMs, supported so far only by a small pilot.

Pith tools