Pith. sign in

REVIEW 2 cited by

Glyce: Glyph-vectors for Chinese Character Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.10125 v7 pith:6GON3NQF submitted 2019-01-29 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords chinesecharactertasksclassificationglycemodelsabilityable
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

It is intuitive that NLP tasks for logographic languages like Chinese should benefit from the use of the glyph information in those languages. However, due to the lack of rich pictographic evidence in glyphs and the weak generalization ability of standard computer vision models on character data, an effective way to utilize the glyph information remains to be found. In this paper, we address this gap by presenting Glyce, the glyph-vectors for Chinese character representations. We make three major innovations: (1) We use historical Chinese scripts (e.g., bronzeware script, seal script, traditional Chinese, etc) to enrich the pictographic evidence in characters; (2) We design CNN structures (called tianzege-CNN) tailored to Chinese character image processing; and (3) We use image-classification as an auxiliary task in a multi-task learning setup to increase the model's ability to generalize. We show that glyph-based models are able to consistently outperform word/char ID-based models in a wide range of Chinese NLP tasks. We are able to set new state-of-the-art results for a variety of Chinese NLP tasks, including tagging (NER, CWS, POS), sentence pair classification, single sentence classification tasks, dependency parsing, and semantic role labeling. For example, the proposed model achieves an F1 score of 80.6 on the OntoNotes dataset of NER, +1.5 over BERT; it achieves an almost perfect accuracy of 99.8\% on the Fudan corpus for text classification. Code found at https://github.com/ShannonAI/glyce.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Query-Based Named Entity Recognition

    cs.CL 2019-08 conditional novelty 6.0 of 10

    Named entity recognition can be reformulated as answering one natural-language question per entity type with a BERT span extractor, and the paper reports state-of-the-art results on five datasets.

  2. Quantum-Evolutionary Neural Networks for Multi-Agent Federated Learning

    cs.NE 2025-05 reject novelty 2.0 of 10

    Quantum-Evolutionary Neural Networks combine phase-shifted sine activations, evolutionary selection, and federated averaging, but the convergence and privacy proofs are asserted rather than derived, and the empirical ...

Pith tools