{"id":"3443468e-af4f-4232-92f6-d2650fb099f1","arxiv_id":"2509.04753","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Decoder-based LLMs tuned with LoRA match or beat encoder-only models on clinical concept and relation extraction, and multi-task instruction tuning sharply improves zero- and few-shot transfer, approaching full fine-tuning with 20% of the data.","lead":"This paper benchmarks nine large-language-model configurations on clinical concept and relation extraction across five medical datasets, comparing encoder-only and decoder-only architectures, full fine-tuning versus lightweight LoRA tuning, and single-task versus multi-task instruction tuning. It reports that decoder LLMs with LoRA tune cheaply to the best accuracy, and that multi-task instruction tuning gives a large zero-shot and few-shot boost, reaching near parity with fu","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The '20% data / <0.005 F1 gap' claim lacks a defined full-data comparator and variance estimate; Figure 3's baseline protocol is unspecified.","rationale":"The reader's CONDITIONAL verdict is appropriate. The paper is a useful empirical benchmark with plausible directional findings: decoder-only models with LoRA are competitive, and multitask instruction tuning helps few-shot transfer. Tables 2 and 3 provide a broad comparison, and the efficiency numbers in Table 4 are informative. However, the single most load-bearing numerical claim—the '<0.005 gap at 20% data'—is supported only by a figure and a protocol description that does not establish that the comparison holds tuning strategy, prompts, and parsing fixed. This is not an accusation of wrongdoing; it is a request for a precise comparator definition and variance reporting. A single run with no error bars cannot support a gap of 0.005, especially when the paper elsewhere reports gains and gaps of similar magnitude as meaningful. The GatorTronLlama pretraining overlap with UF Health is also a real concern for zero-shot generalization claims, but it is secondary because Llama 3.1 also shows strong few-shot performance, and the 20% claim could still be evaluated on Llama alone. I therefore keep the reader's CONDITIONAL verdict unchanged, with the condition that Figure 3's comparator and variance be clarified and reported.","tokens_in":12030,"tokens_out":5077,"duration_ms":50643,"concrete_test":"Locate the exact implementation of the full-data comparator behind Figure 3 and re-run the 20%-data experiment for Llama 3.1-8B on at least one dataset (e.g., 2018 n2c2) with: (a) the same LoRA rank 256 and dropout 0.2, (b) the same instruction prompts as the multi-task condition, (c) the same parsing and evaluation script, and (d) 5 random seeds. Compare the 20%-data model against the same base model trained on 100% of the same dataset with the identical protocol. If the mean F1 gap exceeds 0.005, or if 0.005 lies within the seed-to-seed standard deviation, the headline claim must be qualified. Also require the authors to release the exact numeric values underlying Figure 3 in a table.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central data-efficiency claim—that multitask instruction tuning with only 20% of the full dataset yields F1 within 0.005 of full-size fine-tuning—rests entirely on Figure 3, yet the text never specifies what the full-data comparator is. The few-shot protocol in Methods says multi-task models are compared to 'single-dataset fine-tuned models,' but does not state whether those baselines use the same base model, the same LoRA configuration (rank 256, dropout 0.2), the same prompt templates, the same output parsing, or the same training budget. If the 20% trajectory is compared against a model that uses full-parameter fine-tuning, a different prompt format, or a different architecture, then the '<0.005 gap' is not a clean statement about data efficiency of multitask instruction tuning—it conflates tuning strategy and architecture. Additionally, no seeds or confidence intervals are reported anywhere in the paper, so a 0.005 gap from a single run is not distinguishable from run-to-run variance. This is load-bearing because the paper's practical recommendation—'cost-effective solution' using a fraction of labeled data—depends on this comparison being apples-to-apples. The separate GatorTronLlama pretraining/evaluation overlap on UF Health is a second-order concern, but the unspecified comparator is the first-order blocker.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an empirical benchmark of encoder-only (BERT, GatorTron) and decoder-only (GatorTronGPT, Llama 3.1, GatorTronLlama) LLMs for clinical concept extraction (CE) and relation extraction (RE) on five datasets. It compares traditional full-size fine-tuning with prompt-based PEFT/LoRA, and evaluates multi-task instruction tuning with a leave-one-dataset-out protocol for zero- and few-shot generalization. The headline claims are that decoder-based generative LLMs with PEFT match or exceed the best encoder-only models, that prompt-based PEFT improves RE by up to 15.9% over full fine-tuning, and that multi-task instruction tuning using only 20% of the full dataset achieves F1 within 0.005 of full-size fine-tuning.","tokens_in":12348,"tokens_out":4188,"duration_ms":38977,"significance":"If the claims hold, the paper would provide practical guidance for building clinical IE systems with substantially lower annotation and compute costs: decoder LLMs with LoRA and multi-task instruction tuning could replace expensive full fine-tuning of encoder models. The study covers diverse, widely used datasets and reports an efficiency comparison (Table 4), which is useful for practitioners. The strengths are the breadth of the benchmark, the inclusion of several model families, and the explicit efficiency measurements. However, several load-bearing comparisons are not adequately controlled, and no variance information is reported, so the quantitative conclusions are not yet established at the level claimed.","major_comments":[{"comment":"The abstract's central data-efficiency claim—'using only 20% of the full dataset achieved similar performance comparable to the full-size fine-tuning, with a very small gap less than 0.005 in F1 scores'—is not supported by the reported methodology. The full-data comparator is never defined: it is not stated whether the baseline is the same decoder model with LoRA, the same prompt templates, the same output parsing, or the same training budget. Figure 3 is only described qualitatively; no axis labels, error bars, or per-fold numbers are given. Because all results appear to be single runs with no seeds or confidence intervals, a gap of 0.005 is indistinguishable from run-to-run variance. The authors should specify the comparator precisely and report multiple seeds or confidence intervals for this key claim.","section":""},{"comment":"The claim of an 'F1 improvement up to 15.9% over traditional fine-tuning' for prompt-based PEFT is misleading. In Table 3, the 15.9% gap is between GatorTron-base with full fine-tuning on RadGraph (0.6925) and GatorTronLlama with prompt-based PEFT on RadGraph (0.8514). This comparison varies both the architecture (encoder vs. decoder) and the fine-tuning strategy simultaneously. Within the encoder-only family, the largest improvement from PEFT over full fine-tuning is 2.84 percentage points (GatorTron-large-MRC 0.8661 vs. GatorTron-large 0.8375). The abstract should either report the within-architecture comparison or explicitly state that the 15.9% figure includes architecture differences.","section":""},{"comment":"The few-shot and zero-shot baselines are underspecified. The text states that multi-task models are compared with 'single-dataset fine-tuned models', but it does not state whether those baselines use the same LoRA configuration (rank 256, dropout 0.2), the same prompt templates, the same base model, the same optimizer, or the same decoding/parsing of generated text. Without this information, the reported multi-task improvements (e.g., 0.8–9.7% for CE, 1.8–6.0% for RE, and the zero-shot gains) conflate multi-task instruction tuning with differences in the adaptation protocol. The authors should provide the full baseline protocol or run baselines that differ only in the multi-task training data.","section":""},{"comment":"There is a potential pretraining/evaluation overlap that is not addressed. GatorTronLlama is described as having been continue-pretrained on over 100 billion words of clinical text collected from UF Health, and the UF Health evaluation dataset is also drawn from UF Health IDR. Similarly, GatorTron was pretrained on de-identified UF Health notes. The paper does not report any overlap analysis (e.g., n-gram overlap between pretraining corpora and the UF Health test set). Because the UF Health dataset is not public, readers cannot assess this independently. Please report overlap statistics or otherwise justify that the UF Health test set was not seen during pretraining.","section":""},{"comment":"No measure of variability is reported anywhere. All F1 scores in Tables 2 and 3 and Figures 2 and 3 appear to be from single runs. Given that several headline conclusions involve small differences (e.g., GatorTronLlama 0.8981 vs. Llama 3.1 0.8964 for CE; the <0.005 gap in Figure 3), the absence of seeds, confidence intervals, or significance tests is a major limitation. The authors should report at least three seeds for the key comparisons, or a bootstrap confidence interval, to establish that the reported rankings and gaps are stable.","section":""}],"minor_comments":[{"comment":"'similar performance comparable to' is redundant; consider 'performance comparable to'.","section":""},{"comment":"The definition says decoder-based LLMs are 'trained using the encoder component of the transformer architecture'; this appears to be a typo for 'decoder component'.","section":""},{"comment":"Typo: 'GatoTronGPT-base' should be 'GatorTronGPT-base'.","section":""},{"comment":"Typo in row label: 'GaotTronLlma' should be 'GatorTronLlama'.","section":""},{"comment":"The reported improvements are given in percentage points but phrased as percentages (e.g., 'by 34.41%'); please clarify whether these are relative or absolute differences.","section":""},{"comment":"The figure captions describe content, but the figures themselves are not included in the manuscript text. Ensure the final version includes axis labels, legends, and ideally error bars for the few-shot trajectories.","section":""}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the authors' own GatorTron-family models and the UF Health dataset. The pretraining/evaluation overlap concern (major comment 4) is especially important because the UF Health data are not public and the model was reportedly pretrained on UF Health text. The editor may also wish to consider whether the comparisons with GatorTronGPT and GatorTronLlama, which are not publicly available, can be independently reproduced. The central claims are plausible but require additional methodological detail and variance reporting; the current manuscript is not ready for acceptance without these clarifications."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing to know: this is a broad, mostly sensible benchmark of encoder vs decoder LLMs and full fine-tuning vs LoRA PEFT for clinical concept and relation extraction. The genuinely new piece is the leave-one-dataset-out multi-task instruction tuning, and the 20%-data parity claim is the potentially high-value result. If the paper is right, practitioners can get near-full-data performance with LoRA and a fraction of labeled data, which is a real practical payoff.\n\nWhat it does well: the coverage is extensive—nine model configurations, five datasets, two tasks, and a useful efficiency table. The directional findings are consistent with what the field sees: decoder-only LLMs with prompt-based PEFT are competitive, and multi-task instruction tuning clearly helps zero- and few-shot transfer. The error analysis is a nice addition. The authors give a plausible account of why the results land this way.\n\nThe soft spots are real, though they mainly affect the headline magnitudes rather than the qualitative direction. First, the abstract's '20% data, <0.005 F1 gap' is not supported by a defined comparator. The methods say few-shot experiments use 5, 10, 20, 50 samples, but Figure 3 seems to show percentages of the held-out data, with no statement of what the 'full-size fine-tuning' baseline is—same LoRA rank, same prompts, same parsing, same base model? Without that, the comparison conflates architecture, tuning strategy, and data efficiency. Second, there are no seeds or confidence intervals anywhere; a 0.005 F1 difference from one run is within noise. Third, GatorTronLlama was continued-pretrained on over 100 billion words of UF Health text and then evaluated on the UF Health dataset; the paper never addresses pretraining/evaluation overlap. Fourth, the 15.9% RE improvement headline is cherry-picked, comparing the worst encoder full-fine-tuning result to the best decoder PEFT result, confounding architecture and strategy. There are also minor numerical inconsistencies in the text. These are all addressable in revision; none of them kills the qualitative message.\n\nThis paper is for people building clinical IE systems who want practical guidance. It deserves a serious referee, but only with a requirement to specify comparators, report variance, address the overlap, and fix the inconsistency between the few-shot protocol and Figure 3.","headline":"Useful benchmark with plausible directions, but the headline data-efficiency claim is under-specified and the UF Health overlap is unaddressed.","tokens_in":12884,"tokens_out":2233,"would_cite":false,"duration_ms":21359,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Decoder-style generative LLMs with LoRA and multi-task instruction tuning match full fine-tuning using only 20% of labeled clinical data.","keywords":["clinical concept extraction","clinical relation extraction","large language models","parameter-efficient fine-tuning","LoRA","multi-task instruction tuning","few-shot learning","patient information extraction"],"falsifier":"Run the leave-one-dataset-out experiment with a single base model, holding LoRA rank, prompt templates, optimizer, learning rate, and decoding constant, and compare single-task fine-tuning against multi-task instruction tuning; if the 1.1–37.8% F1 boosts shrink to noise, the multitask claim is refuted. Separately, evaluate the clinically continued-pretrained model on notes from an external institution to test whether pretraining/evaluation overlap explains the high scores.","tokens_in":11862,"feed_emoji":"🏥","tokens_out":6544,"duration_ms":55742,"temperature":0.7,"pith_summary":"This paper asks how best to adapt large language models to extract patient information from clinical notes—finding disease mentions, treatments, medications, and the relations among them. It claims that decoder-style generative models tuned with a parameter-efficient method (LoRA adapters, which train under 1% of weights) match or beat the best encoder-style models on both tasks at a fraction of the training cost. Its larger claim is about generalization: when three generative models are first instruction-tuned on four clinical datasets together, then tested on a held-out dataset with 5, 10, 20, or 50 examples, their zero-shot and few-shot F1 scores jump substantially, and with only 20% of a dataset they match models trained on the full data. If true, clinical NLP teams could spend far less on annotation and compute, reusing a multi-task instruction-tuning pool to adapt to new datasets cheaply.","feed_headline":"Multitask tuning lets 20% of clinical data match full training","feed_subtitle":"Decoder LLMs with LoRA hit full-fine-tuning F1 at a fraction of the training cost.","key_machinery":"The argument runs on two mechanisms. First, prompt-based PEFT: instead of adding dataset-specific classification heads, the model is given natural-language instructions and generates concept spans or relation labels as text; LoRA injects small trainable low-rank matrices into attention layers, so fewer than 1% of parameters are updated. Second, multi-task instruction tuning: the same decoder model is fine-tuned jointly on four datasets spanning two tasks with multiple prompt templates, then tested on a held-out dataset. The leave-one-dataset-out protocol is what lets the paper attribute gains in zero-shot and few-shot performance to the instruction-tuning stage.","core_discovery":"The paper's central claim is that the best recipe for clinical concept and relation extraction is a decoder-only generative LLM adapted with prompt-based LoRA parameter-efficient fine-tuning, followed by multi-task instruction tuning. On five benchmark datasets, the clinical continued-pretrained Llama variant reached average F1 of 0.8981 for concept extraction and 0.8978 for relation extraction, slightly ahead of the largest encoder-only models. Prompt-based PEFT improved relation extraction F1 by up to 15.9% over traditional full fine-tuning. In a leave-one-dataset-out evaluation, multi-task instruction tuning lifted zero-shot F1 from near zero to as high as 0.3596 for concept extraction an","pith_inferences":["A direct ablation that holds the base model, LoRA rank, prompt templates, optimizer, and decoding identical between single-task fine-tuning and multi-task instruction tuning would cleanly isolate the instruction-tuning effect the paper attributes to the multi-task stage.","The 20%-data result suggests a practical deployment recipe: maintain a shared multi-task instruction-tuning pool, then add a small number of target-domain examples for each new annotation schema; whether the pool should include non-clinical datasets is testable but not explored here.","Because one of the best-performing models was continued-pretrained on the same institution's notes used for evaluation, a cross-institution replication on notes from different health systems would show how much of the gain reflects broad method versus domain familiarity."],"forward_implications":["A single generative model can be prompted for both concept extraction and relation extraction, replacing task-specific classification heads with one unified text-to-text interface.","LoRA makes adapting an 8-billion-parameter clinical model practical: about 8 GPU hours versus 48 for full fine-tuning of a 9-billion-parameter model, with inference latency nearly unchanged because adapters merge into the base weights.","Multi-task instruction tuning raises zero-shot F1 from near zero to usable levels (about 0.27–0.40 across models), making annotation-free extraction more realistic for new clinical domains.","For a new dataset, labeling roughly 20% of the available examples after instruction tuning can reach the F1 of a fully fine-tuned model, cutting annotation cost by about 80%.","For small encoder models with abundant in-domain data, full fine-tuning can still beat prompt-based PEFT, so the best architecture depends on model scale and data availability."],"supporting_citations":[{"why":"Supplies LoRA, the parameter-efficient fine-tuning method used for all decoder models.","marker":"[36]"},{"why":"Prior work that established prompt-based PEFT for clinical concept and relation extraction, extended here to multi-task instruction tuning.","marker":"[8]"},{"why":"Supplies the prompt-based machine reading comprehension framework for encoder-only LLMs.","marker":"[25]"},{"why":"Motivates multi-task instruction tuning as a way to turn language models into zero-shot learners.","marker":"[32]"},{"why":"Supplies a multi-task instruction-based generative framework for few-shot named entity recognition.","marker":"[37]"},{"why":"Provides GatorTron, the encoder-only clinical LLM baseline.","marker":"[24]"},{"why":"Provides GatorTronGPT, a decoder-only clinical generative LLM baseline.","marker":"[30]"},{"why":"Supplies the radiology report benchmark for entity and relation extraction.","marker":"[38]"},{"why":"Supplies the medication and adverse-drug-event benchmark used in single-task and multi-task experiments.","marker":"[15]"},{"why":"Supplies the UF Health social-determinants benchmark and serves as one held-out target in leave-one-dataset-out evaluation.","marker":"[39]"}],"fun_headline_variants":["LoRA-tuned decoders match full fine-tuning on clinical NLP","Multi-task tuning lifts zero-shot clinical extraction to 0.36 F1","Prompt-based PEFT beats full fine-tuning by 15.9% on relations","20% clinical data, full F1: LoRA wins over full fine-tuning","Decoder LLMs with LoRA outperform encoder-only models"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The headline few-shot gains assume the single-dataset baselines were trained under identical conditions (same adapter settings, prompts, optimizer, and decoding) as the multi-task models, and that the clinical continued-pretraining corpus does not include the evaluation notes; if either assumption fails, the measured boosts are not solely due to multi-task instruction tuning.","fun_headline_variants_meta":{"raw":{"variants":["LoRA-tuned decoders match full fine-tuning on clinical NLP","Multi-task tuning lifts zero-shot clinical extraction to 0.36 F1","Prompt-based PEFT beats full fine-tuning by 15.9% on relations","20% clinical data, full F1: LoRA wins over full fine-tuning","Decoder LLMs with LoRA outperform encoder-only models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2714,"prompt_tokens":772,"completion_tokens":1942,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":516,"completion_tokens_details":{"reasoning_tokens":1846}},"tokens_in":516,"tokens_out":1942,"duration_ms":14185,"temperature":1.0,"reasoning_tokens":1846,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:57:26.249479+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the leave-one-dataset-out experiment with a single base model, holding LoRA rank, prompt templates, optimizer, learning rate, and decoding constant, and compare single-task fine-tuning against multi-task instruction tuning; if the 1.1–37.8% F1 boosts shrink to noise, the multitask claim is refuted. Separately, evaluate the clinically continued-pretrained model on notes from an external institution to test whether pretraining/evaluation overlap explains the high scores.","supporting_citations":[{"cited_title":"All prompt-based PEFT was conducted within NeMo","cited_arxiv_id":null,"evidence_quote":"Prior work that established prompt-based PEFT for clinical concept and relation extraction, extended here to multi-task instruction tuning."},{"cited_title":"Identifying social determinants of health from clinical narratives: A study of performance, documentation ratio, and potential bias","cited_arxiv_id":null,"evidence_quote":"Supplies the medication and adverse-drug-event benchmark used in single-task and multi-task experiments."}],"review_version":1}