Pith. sign in

REVIEW 1 cited by

Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06779 v1 pith:YZOK765J submitted 2024-07-09 cs.CL

classification cs.CL
keywords questionsscoreretrievalsystemanswerbiomedicalengineeringlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our team participated in the BioASQ 2024 Task12b and Synergy tasks to build a system that can answer biomedical questions by retrieving relevant articles and snippets from the PubMed database and generating exact and ideal answers. We propose a two-level information retrieval and question-answering system based on pre-trained large language models (LLM), focused on LLM prompt engineering and response post-processing. We construct prompts with in-context few-shot examples and utilize post-processing techniques like resampling and malformed response detection. We compare the performance of various pre-trained LLM models on this challenge, including Mixtral, OpenAI GPT and Llama2. Our best-performing system achieved 0.14 MAP score on document retrieval, 0.05 MAP score on snippet retrieval, 0.96 F1 score for yes/no questions, 0.38 MRR score for factoid questions and 0.50 F1 score for list questions in Task 12b.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    KnowSum extrapolates from observed LLM outputs to estimate unseen knowledge, and counting that hidden knowledge shifts model rankings in several evaluation tasks.

Pith tools