Pith. sign in

REVIEW 4 major objections 5 minor 90 references

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read The paper claims that a 2.8-million-sample synthetic instruction corpus, Hum, raises NLU scores of six LLMs by an average of 3.1 points without significantly hurting their other abilities.

desk verdict Useful corpus and synthesis recipes, but the paper's own tables contradict the 'no significant decline' claim, so the abstract overstates what is established. read the letter →

arxiv 2502.03843 v1 pith:KLZE2VHL submitted 2025-02-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords instructionsynthesisnaturallanguageunderstandinginformationextractionmachinereadingcomprehensiontextclassificationsupervisedfine-tuninglargemodelsdatadiversity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the shortage and narrowness of NLU instruction data causes open LLMs to overfit information-extraction templates and lose general language understanding. To fix this, it builds Hum, roughly 2.8 million instructions spanning named entity recognition, relation/event extraction, machine reading comprehension, text classification, open IE, knowledge-graph extraction, and a generalist chat set. The instructions are diversified by three synthesis strategies: guidelines, preference rules, and output-format variants. Fine-tuning six base LLMs on Hum yields an average 3.1% F1 gain across five zero-shot NLU benchmarks, while 28 broader-capability evaluations show no systematic decline. A sympathetic reader would take the claim as: diverse, guideline-rich synthetic instructions can improve NLU without the catastrophic forgetting seen in prior IE-only corpora.

What carries the argument

The load-bearing mechanism is the compound instruction synthesis pipeline. Starting from the structured 'instruction, schema, input' format used in earlier IE corpora, the framework (1) adds guidelines—descriptions, representative positive/negative examples, format specifications, and label-name variants; (2) generates preference rules that alter the semantic annotation rule, such as whether to keep monetary symbols or extract only the highest degree; and (3) varies output formats across JSON, text, markdown, and code while also varying null-value and multi-value handling. These three strategies produce the 45% of Hum that is compound instructions, and ablation studies show removing any one strategy lowers downstream NLU performance, with preference-rule removal hurting the most.

What would settle it

Take a Hum-tuned model and evaluate it on the same five NLU tasks but with instructions rewritten into novel phrasings that never appear in Hum's training data or its standard evaluation templates; if the average improvement over the base model drops below statistical noise, the 3.1% gain is format familiarity, not general NLU.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that a large-scale synthetic instruction corpus covering multiple NLU task families and deliberately varying instruction semantics and output formats can improve the zero-shot NLU performance of open LLMs. The central result is the average 3.1% improvement over six base models on five NLU test sets (CrossNER, FewRel, CCF Law, C3, IMDB), with the strongest gains appearing when test instructions themselves contain the same kinds of guidelines, preference rules, and format variants used during training. The paper attributes this to the three diversification strategies, which it says prevent overfitting to a single instruction style and preserve the model's in-context learning.

Load-bearing premise

The reported NLU gains could come from the models becoming familiar with Hum's own instruction styles, because the evaluation uses formats B and C that Hum itself produces, rather than from a general improvement in language understanding that transfers to unseen instruction phrasings.

Editorial extensions

If this is right

  • If the claim holds, a general-purpose NLU instruction corpus can be synthesized at scale without human annotation of the final instructions, lowering the cost of alignment for NLU-heavy applications.
  • The 3.1% average gain on NLU tasks and the absence of a systematic drop on 28 general-capability datasets suggests that instruction diversity, rather than raw volume alone, is what preserves the base model's broader behavior.
  • Because the gains appear largest on extraction-style tasks when test formats resemble training formats, practitioners should expect the benefit to be format-sensitive and should match inference-time instruction style to training-time style.
  • The robustness experiments show even 10K instructions yield measurable NLU gains, so the method is applicable in low-resource settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If format familiarity explains the gains, then the 3.1% would shrink when test instructions are written in novel styles; the paper does not test this, so the real-world portability of Hum remains open.
  • The preference-rules strategy is essentially a cheap way to multiply one annotated sample into multiple semantically distinct training examples; this could be applied to other structured-output tasks beyond NLU, such as code generation or data wrangling.
  • The paper's evaluation is zero-shot on five tasks, but not on all seven NLU dimensions it claims; a test on truly unseen task types (e.g., dialogue state tracking or semantic parsing) would be a stronger generalization check.
  • The instruction generalist set (8% of Hum) might be doing the work of preventing capability decline; ablating its presence would isolate whether the NLU-specific strategies alone preserve general skills.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces Hum, a 2.8 million-example synthetic instruction corpus for natural language understanding (NLU), covering information extraction (IE), machine reading comprehension (MRC), text classification (TC), and general instruction-following tasks. The authors propose three diversity-enhancing synthesis strategies — guidelines synthesis, preference rules synthesis, and format variants synthesis — and fine-tune six open LLMs with LoRA on Hum. They evaluate on five NLU benchmarks and 28 general-capability datasets across seven dimensions, reporting an average 3.1% improvement on NLU and claiming no significant decline in other capabilities. The paper also includes ablation studies on synthesis strategies, dataset scale, and basic-versus-compound instruction formats.

Significance. If the central claims hold, Hum would be a valuable resource for instruction tuning: it addresses a real gap in NLU instruction diversity, and the proposed synthesis strategies are technically plausible and backed by large-scale experiments. Strengths include the breadth of the evaluation (6 base models, 28 general datasets), the use of multiple synthesis strategies, and the scale/robustness experiments. However, the claim of 'no significant decline' is not supported by the reported numbers, and the evaluation protocols may inflate gains through format familiarity. The paper's contribution is therefore promising but needs substantial revision of its empirical evidence and claims.

major comments (4)
  1. [Abstract; Tables 5, 8, 9, 10] The abstract's claim 'no significant decline observed in other general capabilities' is contradicted by the paper's own tables. Table 9 shows HumMistral OCNLI falling from 48.31 to 12.14 and Table 4 shows HumMistral C3 falling from 67.29 to 47.29; Table 8 shows HumQwen2 MATH/GSM8K average falling from 65.03 to 59.90; Table 10 shows HumQwen2 T-Eval falling from 76.03 to 71.51; Table 7 shows HumQwen2 coding average falling from 62.46 to 60.48. Because only seven-dimension averages are reported, large positive outliers such as HumMistral DROP (7.13 to 55.04) offset the OCNLI collapse. No significance tests, confidence intervals, or per-seed variance are provided, so the blanket 'no significant decline' statement is not merely unsupported but conflicts with the reported numbers.
  2. [Experimental Settings; Table 1] The NLU evaluation uses instruction formats 'B' and 'C' that are generated by the same framework used to synthesize Hum's training data. Hum-fine-tuned models are therefore tested on instruction styles they have seen during training, while train-free baselines such as GPT-4 are tested on these same styles without prior exposure. This setup conflates format familiarity with improved language understanding. To establish that Hum improves general NLU capability, the authors should evaluate on held-out, independently written instruction formats (for example, using instruction templates from prior IE works or human-authored prompts) and show that the gains transfer beyond the synthesized formats.
  3. [Overall Results; Table 5] The reported 'average improvement of 3.1%' is computed from the Language Understanding row of Table 5, but the paper does not state the exact computation or provide per-model confidence intervals. More importantly, the surrounding text in the 'Hum For Natural Language Understanding' section says results are 'mixed with some showing improvements and others showing declines,' and then concludes no comprehensive decline. Given the substantial per-dataset declines listed above, the conclusion needs a statistical or at least a quantile-based definition of 'significant decline.' Without this, the central claim is overstated.
  4. [Dataset Robustness Analysis; Tables 12-13] The robustness analysis shows a positive correlation between data volume and NLU performance, but it is conducted only on Qwen2 and only on one training run per scale. Since LoRA fine-tuning is stochastic, and given that the main claim concerns six different base models, the absence of repeated runs or significance testing is a load-bearing gap for both the robustness and the cross-model conclusions.
minor comments (5)
  1. [Abstract and Experimental Settings] The abstract says '5 NLU tasks,' but Table 4 lists six datasets for the language understanding dimension (C3, WSC, XSum, Lambda, Lcsts, Race), while Table 1 uses five datasets. The relationship between these two evaluations should be clarified.
  2. [Tables 4 and 9] Several dataset names are inconsistent or misspelled: 'Lambda' should be 'Lambada' in Table 4, 'Drop' should be 'DROP' and 'Ocnli' should be 'OCNLI' in Table 9.
  3. [Table 14] In Table 14, the basic-style average for Hum-B (64.67) is higher than for full Hum (64.15), so the claim in the Instruction Ablation Analysis that 'mixing basic and compound instructions yields even better overall performance' is not consistently supported for the B evaluation column; this deserves a brief qualification.
  4. [Experimental Settings] The CCF Law dataset is used for event extraction but is not cited or described beyond the name; a citation or a short description of its construction and license would improve reproducibility.
  5. [General] The paper does not state whether the Hum dataset, the synthetic instructions, or the fine-tuning code will be released; for a dataset-centric paper this availability statement is important.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the claimed NLU gains are measured on external benchmark labels and source texts, not derived from the Hum corpus by construction, though the compound evaluation prompts share the paper's own synthesis machinery.

full rationale

This is an empirical instruction-synthesis paper with no fitted-parameter derivation chain, so the circularity patterns that apply to theoretical or curve-fitting papers do not apply. The central claim (3.1% average NLU improvement) is obtained by supervised fine-tuning on the Hum corpus and then zero-shot evaluation on five external benchmark datasets (CrossNER, FewRel, CCF Law, C3, IMDB) and on 28 general-capability datasets; the benchmark source texts and gold labels do not come from Hum, and Table 16 shows that the evaluation datasets are not listed among the training sources. Hum's own ablations and robustness tests compare against held-out or external test sets rather than against quantities fitted during construction. The paper's self-citations (IEPile, InstructIE, ChatUIE) are used as design templates and baselines, not as proof of the central claim, so they are not load-bearing. One mild external-validity caveat is that the 'C' evaluation prompts are constructed with the same guidelines and format-variants machinery used to synthesize Hum, so part of the observed gain may reflect format familiarity rather than purely generalizable NLU; however, this is a distribution-matching concern, not a circular reduction, because the evaluation labels and passages are external and the models must still produce correct answers or extractions. Separately, the abstract's 'no significant decline' clause is not established by the paper's own data: large drops appear in the reported tables (e.g., Mistral Ocnli 48.31 to 12.14 in Table 9; Mistral C3 67.29 to 47.29 in Table 4; Qwen2 MATH 65.03 to 59.90 in Table 8), and no significance tests or confidence intervals are provided. This is a correctness and overclaim risk, not circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The paper's central empirical claim rests on the corpus composition, the quality of machine-generated labels, and the fairness of the evaluation format, rather than on a mathematical derivation.

free parameters (4)
  • Task distribution in Hum = NER 23%, RE 29%, SPO 11%, EE 5%, EET 3%, EEA 2%, OpenIE 4%, KGE 12%, MRC 2%, TC 1%, IG 8%
    The relative amounts of each task are chosen by the authors and are not optimized or ablated; the central claim depends on this mix.
  • Synthesis strategy mix = 55% basic, 45% compound; among compound: 1,152,470 with guidelines, 34,770 with preference rules, 108,091 with format…
    The proportions of guidelines, preference rules, and format variants are hand-selected; ablations only remove strategies entirely, not tune ratios.
  • LoRA rank and alpha = 64
    Fine-tuning hyperparameters are fixed at 64 without reported tuning; results may shift with different ranks.
  • Learning rate and batch size = 5e-5, batch 320
    Chosen training configuration; no sensitivity analysis is provided.
assumptions (4)
  • domain assumption GPT-4 generated annotations (descriptions, examples, preference rules, and relabeled outputs) are correct and consistent with the original gold labels.
    Preference Rules Synthesis (Figure 3) uses GPT-4 to create new labeling rules and new labels; no human validation or agreement statistics are reported.
  • domain assumption Self-built datasets (marked * in Table 16) are accurately annotated.
    Fourteen self-built datasets are used; annotation procedures are not described in sufficient detail.
  • domain assumption Evaluating on five NLU datasets with the authors' own instruction templates measures general NLU capability fairly across models.
    The B and C instruction formats are products of the same synthesis framework used for training, so fine-tuned models are in-distribution.
  • domain assumption OpenCompass 'gen' mode scores on 28 datasets provide a valid measure of general capabilities.
    The paper relies on OpenCompass defaults without analyzing per-dataset prompt sensitivity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis." pith.science (2026). https://pith.science/paper/KLZE2VHL

@misc{pith2026250203843,
  author       = {Pith},
  title        = {Pith review of: Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KLZE2VHL}},
  note         = {Machine review of arXiv:2502.03843}
}
read the original abstract

High-quality, large-scale instructions are crucial for aligning large language models (LLMs), however, there is a severe shortage of instruction in the field of natural language understanding (NLU). Previous works on constructing NLU instructions mainly focus on information extraction (IE), neglecting tasks such as machine reading comprehension, question answering, and text classification. Furthermore, the lack of diversity in the data has led to a decreased generalization ability of trained LLMs in other NLU tasks and a noticeable decline in the fundamental model's general capabilities. To address this issue, we propose Hum, a large-scale, high-quality synthetic instruction corpus for NLU tasks, designed to enhance the NLU capabilities of LLMs. Specifically, Hum includes IE (either close IE or open IE), machine reading comprehension, text classification, and instruction generalist tasks, thereby enriching task diversity. Additionally, we introduce a human-LLMs collaborative mechanism to synthesize instructions, which enriches instruction diversity by incorporating guidelines, preference rules, and format variants. We conduct extensive experiments on 5 NLU tasks and 28 general capability evaluation datasets for LLMs. Experimental results show that Hum enhances the NLU capabilities of six LLMs by an average of 3.1\%, with no significant decline observed in other general capabilities.

Figures

Figures reproduced from arXiv: 2502.03843 by the authors.

Figure 1
Figure 1. The existing information extraction instructions [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of the natural language understanding instruction synthesis framework. The framework consists of two [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Prompt template for preference rule annotation. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: The source of the Hum dataset and the distribution [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: The performance of different large language mod [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

90 extracted references · 44 canonical work pages

  1. [1]

    I.; Bosma, M.; Michalewski, H.; Dohan, D.; Jiang, E.; Cai, C

    Austin, J.; Odena, A.; Nye, M. I.; Bosma, M.; Michalewski, H.; Dohan, D.; Jiang, E.; Cai, C. J.; Terry, M.; Le, Q. V.; and Sutton, C. 2021. Program Synthesis with Large Language Models. CoRR, abs/2108.07732

  2. [2]

    Bai, Y.; Du, X.; Liang, Y.; Jin, Y.; Liu, Z.; Zhou, J.; Zheng, T.; Zhang, X.; Ma, N.; Wang, Z.; Yuan, R.; Wu, H.; Lin, H.; Huang, W.; Zhang, J.; Chen, W.; Lin, C.; Fu, J.; Yang, M.; Ni, S.; and Zhang, G. 2024. COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning. CoRR, abs/2403.18058

  3. [3]

    L.; Gao, J.; and Choi, Y

    Bisk, Y.; Zellers, R.; Bras, R. L.; Gao, J.; and Choi, Y. 2020. PIQA: Reasoning about Physical Commonsense in Natural Language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificia...

  4. [4]

    Carreras, X.; and M \` a rquez, L. 2004. Introduction to the CoNLL-2004 Shared Task: Semantic Role Labeling. In Ng, H. T.; and Riloff, E., eds., Proceedings of the Eighth Conference on Computational Natural Language Learning, CoNLL 2004, Held in cooperation with HLT-NAACL 2004, Boston, Massachusetts, USA, May 6-7, 2004 , 89--97. ACL

  5. [5]

    Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; de Oliveira Pinto, H. P.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; Ray, A.; Puri, R.; Krueger, G.; Petrov, M.; Khlaaf, H.; Sastry, G.; Mishkin, P.; Chan, B.; Gray, S.; Ryder, N.; Pavlov, M.; Power, A.; Kaiser, L.; Bavarian, M.; Winter, C.; Tillet, P.; Such, F. P.; Cummings, D.; Plappert, M.; Ch...

  6. [6]

    Chen, P.; Xu, H.; Zhang, C.; and Huang, R. 2022. Crossroads, Buildings and Neighborhoods: A Dataset for Fine-grained Location Recognition. In Carpuat, M.; de Marneffe, M.; and Ru \' z, I. V. M., eds., Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2022, ...

  7. [7]

    Chen, Z.; Du, W.; Zhang, W.; Liu, K.; Liu, J.; Zheng, M.; Zhuo, J.; Zhang, S.; Lin, D.; Chen, K.; and Zhao, F. 2024. T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step. arXiv:2312.14033

  8. [8]

    Cheng, D.; Gu, Y.; Huang, S.; Bi, J.; Huang, M.; and Wei, F. 2024. Instruction Pre-Training: Language Models are Supervised Multitask Learners. CoRR, abs/2406.14491

Show all 90 references
  1. [9]

    Cheng, D.; Huang, S.; and Wei, F. 2023. Adapting Large Language Models via Reading Comprehension. CoRR, abs/2309.09530

  2. [10]

    Clark, C.; Lee, K.; Chang, M.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chapter of t...

  3. [11]

    Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. CoRR, abs/1803.05457

  4. [12]

    Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve Math Word Problems. CoRR, abs/2110.14168

  5. [13]

    Contributors, O. 2023. OpenCompass: A Universal Evaluation Platform for Foundation Models. https://github.com/open-compass/opencompass

  6. [14]

    Cui, Y.; Liu, T.; Che, W.; Xiao, L.; Chen, Z.; Ma, W.; Wang, S.; and Hu, G. 2019. A Span-Extraction Dataset for Chinese Machine Reading Comprehension. In Inui, K.; Jiang, J.; Ng, V.; and Wan, X., eds., Proceedings of the 2019 Conference on Empirical Methods in Natural Language...

  7. [15]

    Cui, Y.; Yang, Z.; and Yao, X. 2023. Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca. CoRR, abs/2304.08177

  8. [16]

    Deng, H.; Zhang, Y.; Zhang, Y.; Ying, W.; Yu, C.; Gao, J.; Wang, W.; Bai, X.; Yang, N.; Ma, J.; Chen, X.; and Zhou, T. 2022. Title2Event: Benchmarking Open Event Extraction with a Large-scale Chinese Title Dataset. In Goldberg, Y.; Kozareva, Z.; and Zhang, Y., eds., Proceeding...

  9. [17]

    I.; Leaman, R.; and Lu, Z

    Dogan, R. I.; Leaman, R.; and Lu, Z. 2014. NCBI disease corpus: A resource for disease name recognition and concept normalization. J. Biomed. Informatics, 47: 1--10

  10. [18]

    Dong, G.; Lu, K.; Li, C.; Xia, T.; Yu, B.; Zhou, C.; and Zhou, J. 2024. Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models. CoRR, abs/2406.13542

  11. [19]

    Dua, D.; Wang, Y.; Dasigi, P.; Stanovsky, G.; Singh, S.; and Gardner, M. 2019. DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North American Chap...

  12. [20]

    Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; Goyal, A.; Hartshorn, A.; Yang, A.; Mitra, A.; Sravankumar, A.; Korenev, A.; Hinsvark, A.; Rao, A.; Zhang, A.; et al. 2024. The Llama 3 Herd of Models. arXiv:2...

  13. [21]

    L.; Chen, F.; Yao, S.; Hu, R.; Zhu, X.; Smith, J

    Guan, R.; Man, K. L.; Chen, F.; Yao, S.; Hu, R.; Zhu, X.; Smith, J. S.; Lim, E. G.; and Yue, Y. 2024. FindVehicle and VehicleFinder: a NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval system. Multim. Tools Appl., 83(8):...

  14. [22]

    Gui, H.; Qiao, S.; Zhang, J.; Ye, H.; Sun, M.; Liang, L.; Chen, H.; and Zhang, N. 2023. InstructIE: A Bilingual Instruction-based Information Extraction Dataset. CoRR, abs/2305.11527

  15. [23]

    Gui, H.; Yuan, L.; Ye, H.; Zhang, N.; Sun, M.; Liang, L.; and Chen, H. 2024. IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus. CoRR, abs/2402.14710

  16. [24]

    M.; and Toldo, L

    Gurulingappa, H.; Rajput, A. M.; and Toldo, L. 2012. Extraction of Adverse Drug Effects from Medical Case Reports. J. Biomed. Semant., 3: 15

  17. [25]

    Han, C.; Zhang, J.; Li, X.; Xu, G.; Peng, W.; and Zeng, Z. 2022. DuEE-Fin: A Large-Scale Dataset for Document-Level Event Extraction. In Lu, W.; Huang, S.; Hong, Y.; and Zhou, X., eds., Natural Language Processing and Chinese Computing - 11th CCF International Conference, NLPC...

  18. [26]

    Han, X.; Zhu, H.; Yu, P.; Wang, Z.; Yao, Y.; Liu, Z.; and Sun, M. 2018. FewRel: A Large-Scale Supervised Few-shot Relation Classification Dataset with State-of-the-Art Evaluation. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Confere...

  19. [27]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 a . Measuring Massive Multitask Language Understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net

  20. [28]

    Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 b . Measuring Mathematical Problem Solving With the MATH Dataset. In Vanschoren, J.; and Yeung, S., eds., Proceedings of the Neural Information Processing Systems Track o...

  21. [29]

    H.; Marcus, M

    Hovy, E. H.; Marcus, M. P.; Palmer, M.; Ramshaw, L. A.; and Weischedel, R. M. 2006. OntoNotes: The 90 \ In Moore, R. C.; Bilmes, J. A.; Chu - Carroll, J.; and Sanderson, M., eds., Human Language Technology Conference of the North American Chapter of the Association of Computat...

  22. [30]

    Hu, H.; Richardson, K.; Xu, L.; Li, L.; K \" u bler, S.; and Moss, L. S. 2020. OCNLI: Original Chinese Natural Language Inference. In Cohn, T.; He, Y.; and Liu, Y., eds., Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 ,...

  23. [31]

    Huang, Y.; Bai, Y.; Zhu, Z.; Zhang, J.; Zhang, J.; Su, T.; Liu, J.; Lv, C.; Zhang, Y.; Lei, J.; Fu, Y.; Sun, M.; and He, J. 2023. C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M....

  24. [32]

    Jat, S.; Khandelwal, S.; and Talukdar, P. P. 2018. Improving Distantly Supervised Relation Extraction using Word and Entity Based Attention. CoRR, abs/1804.06987

  25. [33]

    Jiao, Y.; Zhong, M.; Li, S.; Zhao, R.; Ouyang, S.; Ji, H.; and Han, J. 2023. Instruct and Extract: Instruction Tuning for On-Demand Information Extraction. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Language...

  26. [34]

    S.; and Zettlemoyer, L

    Joshi, M.; Choi, E.; Weld, D. S.; and Zettlemoyer, L. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Barzilay, R.; and Kan, M., eds., Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, AC...

  27. [35]

    Kim, J.; Ohta, T.; Tateisi, Y.; and Tsujii, J. 2003. GENIA corpus - a semantically annotated corpus for bio-textmining. In Proceedings of the Eleventh International Conference on Intelligent Systems for Molecular Biology, June 29 - July 3, 2003, Brisbane, Australia, 180--182

  28. [36]

    Kocaman, V.; and Talby, D. 2020. Biomedical Named Entity Recognition at Scale. In Bimbo, A. D.; Cucchiara, R.; Sclaroff, S.; Farinella, G. M.; Mei, T.; Bertini, M.; Escalante, H. J.; and Vezzani, R., eds., Pattern Recognition. ICPR International Workshops and Challenges - Virt...

  29. [37]

    Kolluru, K.; Adlakha, V.; Aggarwal, S.; Mausam; and Chakrabarti, S. 2020. OpenIE6: Iterative Grid Labeling and Coordination Analysis for Open Information Extraction. In Webber, B.; Cohn, T.; He, Y.; and Liu, Y., eds., Proceedings of the 2020 Conference on Empirical Methods in ...

  30. [38]

    Kumar, A.; and Starly, B. 2022. "FabNER": information extraction from manufacturing process science domain literature using named entity recognition. J. Intell. Manuf., 33(8): 2393--2407

  31. [39]

    P.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; Toutanova, K.; Jones, L.; Kelcey, M.; Chang, M.; Dai, A

    Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A. P.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; Toutanova, K.; Jones, L.; Kelcey, M.; Chang, M.; Dai, A. M.; Uszkoreit, J.; Le, Q.; and Petrov, S. 2019. Natural Questions: a Benchmark for Q...

  32. [40]

    Levow, G. 2006. The Third International Chinese Language Processing Bakeoff: Word Segmentation and Named Entity Recognition. In Ng, H. T.; and Kwong, O. O. Y., eds., Proceedings of the Fifth Workshop on Chinese Language Processing, SIGHAN@COLING/ACL 2006, Sydney, Australia, Ju...

  33. [41]

    Li, H.; Dong, Q.; Tang, Z.; Wang, C.; Zhang, X.; Huang, H.; Huang, S.; Huang, X.; Huang, Z.; Zhang, D.; Gu, Y.; Cheng, X.; Wang, X.; Chen, S.; Dong, L.; Lu, W.; Sui, Z.; Wang, B.; Lam, W.; and Wei, F. 2024. Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning f...

  34. [42]

    Li, H.; Zhang, Y.; Koto, F.; Yang, Y.; Zhao, H.; Gong, Y.; Duan, N.; and Baldwin, T. 2023 a . CMMLU: Measuring massive multitask language understanding in Chinese. CoRR, abs/2306.09212

  35. [43]

    Li, P.; Sun, T.; Tang, Q.; Yan, H.; Wu, Y.; Huang, X.; and Qiu, X. 2023 b . CodeIE: Large Code Generation Models are Better Few-Shot Information Extractors. In Rogers, A.; Boyd - Graber, J. L.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association fo...

  36. [44]

    Li, S.; He, W.; Shi, Y.; Jiang, W.; Liang, H.; Jiang, Y.; Zhang, Y.; Lyu, Y.; and Zhu, Y. 2019. DuIE: A Large-Scale Chinese Dataset for Information Extraction. In Tang, J.; Kan, M.; Zhao, D.; Li, S.; and Zan, H., eds., Natural Language Processing and Chinese Computing - 8th CC...

  37. [45]

    Li, X.; Li, F.; Pan, L.; Chen, Y.; Peng, W.; Wang, Q.; Lyu, Y.; and Zhu, Y. 2020. DuEE: A Large-Scale Dataset for Chinese Event Extraction in Real-World Scenarios. In Zhu, X.; Zhang, M.; Hong, Y.; and He, R., eds., Natural Language Processing and Chinese Computing - 9th CCF In...

  38. [46]

    Liu, J.; Pasupat, P.; Cyphers, S.; and Glass, J. R. 2013. Asgard: A portable architecture for multilingual dialogue systems. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2013, Vancouver, BC, Canada, May 26-31, 2013 , 8386--8390. IEEE

  39. [47]

    Liu, Z.; Xu, Y.; Yu, T.; Dai, W.; Ji, Z.; Cahyawijaya, S.; Madotto, A.; and Fung, P. 2021. CrossNER: Evaluating Cross-Domain Named Entity Recognition. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of A...

  40. [48]

    Lu, Y.; Liu, Q.; Dai, D.; Xiao, X.; Lin, H.; Han, X.; Sun, L.; and Wu, H. 2022. Unified Structure Generation for Universal Information Extraction. In Muresan, S.; Nakov, P.; and Villavicencio, A., eds., Proceedings of the 60th Annual Meeting of the Association for Computationa...

  41. [49]

    Luan, Y.; He, L.; Ostendorf, M.; and Hajishirzi, H. 2018 a . Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph Construction. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Conference on E...

  42. [50]

    Luan, Y.; He, L.; Ostendorf, M.; and Hajishirzi, H. 2018 b . Multi-Task Identification of Entities, Relations, and Coreference for Scientific Knowledge Graph Construction. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Conference on E...

  43. [51]

    Luo, L.; Li, N.; Li, S.; Yang, Z.; and Lin, H. 2018. DUTIR at the CCKS-2018 Task1: A Neural Network Ensemble Approach for Chinese Clinical Named Entity Recognition. In CCKS tasks, 7--12

  44. [52]

    L.; Daly, R

    Maas, A. L.; Daly, R. E.; Pham, P. T.; Huang, D.; Ng, A. Y.; and Potts, C. 2011. Learning Word Vectors for Sentiment Analysis. In Lin, D.; Matsumoto, Y.; and Mihalcea, R., eds., The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologi...

  45. [53]

    Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018. Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering. In Riloff, E.; Chiang, D.; Hockenmaier, J.; and Tsujii, J., eds., Proceedings of the 2018 Conference on Empirical Methods in Natu...

  46. [54]

    H.; Abdalla, M.; Abdulmumin, I.; Ahmad, I

    Ousidhoum, N.; Muhammad, S. H.; Abdalla, M.; Abdulmumin, I.; Ahmad, I. S.; Ahuja, S.; Aji, A. F.; Araujo, V.; Beloucif, M.; de Kock, C.; Hourrane, O.; Shrivastava, M.; Solorio, T.; Surange, N.; Vishnubhotla, K.; Yimam, S. M.; and Mohammad, S. M. 2024. SemEval Task 1: Semantic ...

  47. [55]

    Pyysalo, S.; and Ananiadou, S. 2014. Anatomical entity mention recognition at literature scale. Bioinform., 30(6): 868--875

  48. [56]

    Qi, Y.; Peng, H.; Wang, X.; Xu, B.; Hou, L.; and Li, J. 2024. ADELIE: Aligning Large Language Models on Information Extraction. CoRR, abs/2405.05008

  49. [57]

    Ren, J.; Wang, S.; Song, R.; Wu, Y.; Gao, Y.; An, B.; Cheng, Z.; and Xu, G. 2022. IREE: A Fine-Grained Dataset for Chinese Event Extraction in Investment Research. In Sun, M.; Qi, G.; Liu, K.; Ren, J.; Xu, B.; Feng, Y.; Liu, Y.; and Chen, Y., eds., Knowledge Graph and Semantic...

  50. [58]

    Riedel, S.; Yao, L.; and McCallum, A. 2010. Modeling Relations and Their Mentions without Labeled Text. In Balc \' a zar, J. L.; Bonchi, F.; Gionis, A.; and Sebag, M., eds., Machine Learning and Knowledge Discovery in Databases, European Conference, ECML PKDD 2010, Barcelona, ...

  51. [59]

    L.; Rigau, G.; and Agirre, E

    Sainz, O.; Garc \' a - Ferrero, I.; Agerri, R.; de Lacalle, O. L.; Rigau, G.; and Agirre, E. 2023. GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction. CoRR, abs/2310.03668

  52. [60]

    Sang, E. F. T. K.; and Meulder, F. D. 2003. Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. In Daelemans, W.; and Osborne, M., eds., Proceedings of the Seventh Conference on Natural Language Learning, CoNLL 2003, Held in cooperation w...

  53. [61]

    Satyapanich, T.; Ferraro, F.; and Finin, T. 2020. CASIE: Extracting Cybersecurity Event Information from Text. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020...

  54. [62]

    Sun, K.; Yu, D.; Yu, D.; and Cardie, C. 2020. Investigating Prior Knowledge for Challenging Chinese Machine Reading Comprehension. Trans. Assoc. Comput. Linguistics, 8: 141--155

  55. [63]

    Sun, M. 2022. weibo \_ senti \_ 100k and THUCNews. https://doi.org/10.21227/abj8-y636. Accessed on YYYY-MM-DD

  56. [64]

    Sun, X.; Li, X.; Li, J.; Wu, F.; Guo, S.; Zhang, T.; and Wang, G. 2023. Text Classification via Large Language Models. In Bouamor, H.; Pino, J.; and Bali, K., eds., Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 , 8990--90...

  57. [65]

    C.; John, B.; Greene, N.; Kim, J.; and He, Y

    Sun, Z.; Li, J.; Pergola, G.; Wallace, B. C.; John, B.; Greene, N.; Kim, J.; and He, Y. 2022. PHEE: A Dataset for Pharmacovigilance Event Extraction from Text. In Goldberg, Y.; Kozareva, Z.; and Zhang, Y., eds., Proceedings of the 2022 Conference on Empirical Methods in Natura...

  58. [66]

    W.; Chowdhery, A.; Le, Q

    Suzgun, M.; Scales, N.; Sch \" a rli, N.; Gehrmann, S.; Tay, Y.; Chung, H. W.; Chowdhery, A.; Le, Q. V.; Chi, E. H.; Zhou, D.; and Wei, J. 2023. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. In Rogers, A.; Boyd - Graber, J. L.; and Okazaki, N., eds.,...

  59. [67]

    Takanobu, R.; Zhang, T.; Liu, J.; and Huang, M. 2019. A Hierarchical Framework for Relation Extraction with Reinforcement Learning. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Co...

  60. [68]

    Tedeschi, S.; and Navigli, R. 2022. MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation). In Carpuat, M.; de Marneffe, M.; and Ru \' z, I. V. M., eds., Findings of the Association for Computational Linguistics: NAACL...

  61. [69]

    Walker, C.; and Consortium, L. D. 2005. ACE 2005 Multilingual Training Corpus. LDC corpora. Linguistic Data Consortium. ISBN 9781585633760

  62. [70]

    Wang, X.; Zhou, W.; Zu, C.; Xia, H.; Chen, T.; Zhang, Y.; Zheng, R.; Ye, J.; Zhang, Q.; Gui, T.; Kang, J.; Yang, J.; Li, S.; and Du, C. 2023 a . InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction. CoRR, abs/2304.08085

  63. [71]

    A.; Khashabi, D.; and Hajishirzi, H

    Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2023 b . Self-Instruct: Aligning Language Models with Self-Generated Instructions. In Rogers, A.; Boyd - Graber, J. L.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the...

  64. [72]

    Wang, Z.; Pang, Y.; and Lin, Y. 2023. Large Language Models Are Zero-Shot Text Classifiers. CoRR, abs/2312.01044

  65. [73]

    Xia, Y.; and Wang, Q. 2017. Clinical named entity recognition: ECUST in the CCKS-2017 shared task 2. In CEUR workshop proceedings, volume 1976, 43--48

  66. [74]

    Xiao, X.; Wang, Y.; Xu, N.; Wang, Y.; Yang, H.; Wang, M.; Luo, Y.; Wang, L.; Mao, W.; and Zeng, D. 2023. YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information Extraction. CoRR, abs/2312.15548

  67. [75]

    Xu, J.; Sun, M.; Zhang, Z.; and Zhou, J. 2024 a . ChatUIE: Exploring Chat-based Unified Information Extraction Using Large Language Models. In Calzolari, N.; Kan, M.; Hoste, V.; Lenci, A.; Sakti, S.; and Xue, N., eds., Proceedings of the 2024 Joint International Conference on ...

  68. [76]

    Xu, L.; Tong, Y.; Dong, Q.; Liao, Y.; Yu, C.; Tian, Y.; Liu, W.; Li, L.; and Zhang, X. 2020. CLUENER2020: Fine-grained Named Entity Recognition Dataset and Benchmark for Chinese. CoRR, abs/2001.04351

  69. [77]

    Xu, Z.; Jiang, F.; Niu, L.; Deng, Y.; Poovendran, R.; Choi, Y.; and Lin, B. Y. 2024 b . Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing. CoRR, abs/2406.08464

  70. [78]

    Yang, A.; Xiao, B.; Wang, B.; Zhang, B.; Bian, C.; Yin, C.; Lv, C.; Pan, D.; Wang, D.; Yan, D.; Yang, F.; Deng, F.; Wang, F.; Liu, F.; Ai, G.; Dong, G.; Zhao, H.; Xu, H.; Sun, H.; Zhang, H.; Liu, H.; Ji, J.; Xie, J.; Dai, J.; Fang, K.; Su, L.; Song, L.; Liu, L.; Ru, L.; Ma, L....

  71. [79]

    Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Tang, J.; Wang, J.; Yang, J.; Tu, J.; Zhang, J.; Ma, J.; Yang, J.; Xu, J.; Zhou, J.; Bai, J.; He, J.; Lin, J.; Dang, K.; Lu, K.; Chen, K.; Yang, K.; Li, M....

  72. [80]

    Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. HellaSwag: Can a Machine Really Finish Your Sentence? In Korhonen, A.; Traum, D. R.; and M \` a rquez, L., eds., Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, ...

  73. [81]

    Zeng, W.; Xu, C.; Zhao, Y.; Lou, J.; and Chen, W. 2024. Automatic Instruction Evolving for Large Language Models. CoRR, abs/2406.00770

  74. [82]

    Zhang, D.; and Wang, D. 2015. Relation Classification via Recurrent Neural Network. CoRR, abs/1508.01006

  75. [83]

    Zhang, S.; Cheng, H.; Gao, J.; and Poon, H. 2022. Optimizing Bi-Encoder for Named Entity Recognition via Contrastive Learning. CoRR, abs/2208.14565

  76. [84]

    Zhang, X.; Li, C.; Zong, Y.; Ying, Z.; He, L.; and Qiu, X. 2023. Evaluating the Performance of Large Language Models on GAOKAO Benchmark. CoRR, abs/2305.12474

  77. [85]

    Zhang, Y.; and Yang, J. 2018. Chinese NER Using Lattice LSTM . In Gurevych, I.; and Miyao, Y., eds., Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers , 1554--1564. A...

  78. [86]

    Zheng, Y.; Zhang, R.; Zhang, J.; Ye, Y.; Luo, Z.; and Ma, Y. 2024. LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models. CoRR, abs/2403.13372

  79. [87]

    Zhong, W.; Cui, R.; Guo, Y.; Liang, Y.; Lu, S.; Wang, Y.; Saied, A.; Chen, W.; and Duan, N. 2023. AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models. CoRR, abs/2304.06364

  80. [88]

    Zhou, W.; Zhang, S.; Gu, Y.; Chen, M.; and Poon, H. 2024. UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net

  81. [89]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  82. [90]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.