REVIEW 13 cited by
Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We describe the CoNLL-2003 shared task: language-independent named entity recognition. We give background information on the data sets (English and German) and the evaluation method, present a general overview of the systems that have taken part in the task and discuss their performance.
Forward citations
Cited by 13 Pith papers
-
Soft Head Selection for Injecting ICL-Derived Task Embeddings
SITE applies soft gradient-based head selection to inject ICL-derived task embeddings, outperforming prior embedding adaptation and few-shot ICL across generation, reasoning, and NLU tasks on 12 LLMs from 4B to 70B pa...
-
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
LLaMA-Adapter turns frozen LLaMA 7B into a capable instruction follower using only 1.2M new parameters and zero-init attention, matching Alpaca while extending to image-conditioned reasoning on ScienceQA and COCO.
-
BCL: Bayesian In-Context Learning Framework for Information Extraction
BCL introduces a particle-filtering Bayesian update framework to systematically refine label representations in in-context learning for information extraction, claiming consistent gains over prior methods.
-
A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining
LLM-written pipelines and LLM-generated labels are distilled into one small instruction-following model that performs classification and span extraction cheaply at corpus scale.
-
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
Bias in GPT-2 and Llama-2 is localized to a small set of edges, and ablation of those edges reduces bias while impairing unrelated NLP tasks.
-
MPL: Multiple Programming Languages with Large Language Models for Information Extraction
Using multiple programming languages as code-style prompts during fine-tuning improves LLM information extraction accuracy over single-language prompting.
-
LIMO: Less is More for Reasoning
LIMO achieves 63.3% on AIME24 and 95.6% on MATH500 via supervised fine-tuning on roughly 1% of the data used by prior models, supporting the claim that minimal strategic examples suffice when pre-training has already ...
-
DICOM De-Identification via Hybrid AI and Rule-Based Framework for Scalable, Uncertainty-Aware Redaction
A rule-based and AI hybrid with uncertainty-aware detection reports 99.88% de-identification pass rate on its own DICOM, HIPAA, and TCIA checks.
-
DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models
DPO after SFT plus constrained template-filling decoding improves LLM sequence-labeling accuracy and cuts inference time by reusing KV cache for non-label tokens.
-
Extracting Structured Requirements from Unstructured Building Technical Specifications for Building Information Modeling
A study showing that CamemBERT and Fr_core_news_lg achieve over 90% F1 for named entity recognition and Random Forest achieves over 80% F1 for relation extraction on French building technical specifications.
-
PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
Instruction-tuned open-source LLMs, especially DeepSeek-Q1, outperform fine-tuned, RAG, and NER baselines on PII redaction accuracy and leakage in this benchmark.
-
The Science of Evaluating Foundation Models
A survey-and-checklist proposal that organizes LLM evaluation into an ABCD framework (Algorithm, Big Data, Computation, Domain Expertise) for context-aware, documented assessment.
-
To Tune or Not To Tune? How About the Best of Both Worlds?
A sequential fine-tuning strategy for pre-trained language models reports modest accuracy gains of 4.7%, 0.99%, and 0.72% on semantic similarity, sequence labeling, and text classification tasks.
Discussion (0). Continue with ORCID to comment.