REVIEW 2 cited by
Visualizing and Measuring the Geometry of BERT
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformer architectures show significant promise for natural language processing. Given that a single pretrained model can be fine-tuned to perform well on many different tasks, these networks appear to extract generally useful linguistic features. A natural question is how such networks represent this information internally. This paper describes qualitative and quantitative investigations of one particularly effective model, BERT. At a high level, linguistic features seem to be represented in separate semantic and syntactic subspaces. We find evidence of a fine-grained geometric representation of word senses. We also present empirical descriptions of syntactic representations in both attention matrices and individual word embeddings, as well as a mathematical argument to explain the geometry of these representations.
Forward citations
Cited by 2 Pith papers
-
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
Word Confusion measures semantic similarity as classifier confusion between contextual embeddings, matching human judgments as well as or better than cosine similarity, and enables analyst-chosen feature dimensions.
-
Emergent Stack Representations in Modeling Counter Languages Using Transformers
A small transformer trained on counter languages encodes the current stack depth in its final-layer activations, recoverable by simple probing classifiers.
Discussion (0). Continue with ORCID to comment.