Pith. sign in

REVIEW 1 cited by

Developing a Named Entity Recognition Dataset for Tagalog

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07161 v1 pith:SGRGSYOP submitted 2023-11-13 cs.CL

classification cs.CL
keywords datasetentitytagalogacrossnamedrecognitionwereagreement
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We present the development of a Named Entity Recognition (NER) dataset for Tagalog. This corpus helps fill the resource gap present in Philippine languages today, where NER resources are scarce. The texts were obtained from a pretraining corpora containing news reports, and were labeled by native speakers in an iterative fashion. The resulting dataset contains ~7.8k documents across three entity types: Person, Organization, and Location. The inter-annotator agreement, as measured by Cohen's $\kappa$, is 0.81. We also conducted extensive empirical evaluation of state-of-the-art methods across supervised and transfer learning settings. Finally, we released the data and processing code publicly to inspire future work on Tagalog NLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Simple knowledge distillation from mBERT produces smaller, faster Tagalog-only transformers that match the teacher on some tasks and lag on NER.

Pith tools