Pith. sign in

REVIEW 3 cited by

Huatuo-26M, a Large-scale Chinese Medical QA Dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.01526 v1 pith:HI4LEOLN submitted 2023-05-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords datasetexistingmedicalmodelsgenerationhuatuo-26mlanguagemany
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we release a largest ever medical Question Answering (QA) dataset with 26 million QA pairs. We benchmark many existing approaches in our dataset in terms of both retrieval and generation. Experimental results show that the existing models perform far lower than expected and the released dataset is still challenging in the pre-trained language model era. Moreover, we also experimentally show the benefit of the proposed dataset in many aspects: (i) trained models for other QA datasets in a zero-shot fashion; and (ii) as external knowledge for retrieval-augmented generation (RAG); and (iii) improving existing pre-trained language models by using the QA pairs as a pre-training corpus in continued training manner. We believe that this dataset will not only contribute to medical research but also facilitate both the patients and clinical doctors. See \url{https://github.com/FreedomIntelligence/Huatuo-26M}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pruning General Large Language Models into Customized Expert Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Cus-Prun identifies and removes neurons that are irrelevant to a user's target language, domain, and task, producing specialized expert models without post-training.

  2. ListConRanker: A Contrastive Text Reranker with Listwise Encoding

    cs.CL 2025-01 conditional novelty 5.0 of 10

    ListConRanker combines listwise attention over passage embeddings with Circle Loss to set a new mAP average on the C-MTEB reranking benchmark.

  3. IIMedGPT: Promoting Large Language Model Capabilities of Medical Tasks by Efficient Human Preference Alignment

    cs.CL 2025-01 conditional novelty 4.0 of 10

    IIMedGPT, a Qwen-14B-based Chinese medical chatbot fine-tuned with a new 220k-pair instruction dataset and DPO preference alignment, is claimed to surpass prior Chinese medical LLMs in dialogue quality, pending releas...

Pith tools