Pith. sign in

REVIEW

MuLMS: A Multi-Layer Annotated Text Corpus for Information Extraction in the Materials Science Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.15569 v1 pith:MVGVWWSS submitted 2023-10-24 cs.CL

classification cs.CL
keywords domainmaterialsscienceannotatedbeencorpusdatasetsextraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Keeping track of all relevant recent publications and experimental results for a research area is a challenging task. Prior work has demonstrated the efficacy of information extraction models in various scientific areas. Recently, several datasets have been released for the yet understudied materials science domain. However, these datasets focus on sub-problems such as parsing synthesis procedures or on sub-domains, e.g., solid oxide fuel cells. In this resource paper, we present MuLMS, a new dataset of 50 open-access articles, spanning seven sub-domains of materials science. The corpus has been annotated by domain experts with several layers ranging from named entities over relations to frame structures. We present competitive neural models for all tasks and demonstrate that multi-task training with existing related resources leads to benefits.

Discussion (0). Continue with ORCID to comment.

Pith tools