Pith. sign in

REVIEW 1 cited by

Semantic Labeling Using a Deep Contextualized Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.16037 v1 pith:CAYKK6X5 submitted 2020-10-30 cs.LG cs.DBcs.IR

classification cs.LGcs.DBcs.IR
keywords columndatalabelinglabelssemanticvalueslanguageschema
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating schema labels automatically for column values of data tables has many data science applications such as schema matching, and data discovery and linking. For example, automatically extracted tables with missing headers can be filled by the predicted schema labels which significantly minimizes human effort. Furthermore, the predicted labels can reduce the impact of inconsistent names across multiple data tables. Understanding the connection between column values and contextual information is an important yet neglected aspect as previously proposed methods treat each column independently. In this paper, we propose a context-aware semantic labeling method using both the column values and context. Our new method is based on a new setting for semantic labeling, where we sequentially predict labels for an input table with missing headers. We incorporate both the values and context of each data column using the pre-trained contextualized language model, BERT, that has achieved significant improvements in multiple natural language processing tasks. To our knowledge, we are the first to successfully apply BERT to solve the semantic labeling task. We evaluate our approach using two real-world datasets from different domains, and we demonstrate substantial improvements in terms of evaluation metrics over state-of-the-art feature-based methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Time Series Language Model for Descriptive Caption Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    TSLM combines a tagged textual view and a reprogrammed embedding view of a time series with LLM-generated, scorer-filtered training data to produce state-of-the-art time series captions on the STOCK and SYNTH benchmarks.

Pith tools