REVIEW 2 cited by
ETC: Encoding Long and Structured Inputs in Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformer models have advanced the state of the art in many Natural Language Processing (NLP) tasks. In this paper, we present a new Transformer architecture, Extended Transformer Construction (ETC), that addresses two key challenges of standard Transformer architectures, namely scaling input length and encoding structured inputs. To scale attention to longer inputs, we introduce a novel global-local attention mechanism between global tokens and regular input tokens. We also show that combining global-local attention with relative position encodings and a Contrastive Predictive Coding (CPC) pre-training objective allows ETC to encode structured inputs. We achieve state-of-the-art results on four natural language datasets requiring long and/or structured inputs.
Forward citations
Cited by 2 Pith papers
-
TyDi QA-WANA: A Benchmark for Information-Seeking Question Answering in Languages of West Asia and North Africa
TyDi QA-WANA is a new 28,000-example QA benchmark covering 10 under-represented languages with long-context, information-seeking questions and baseline evaluations.
-
LM2: Large Memory Models
LM2 adds a cross-attention memory bank with input, forget, and output gates to every decoder block, reporting large BABILong gains and no MMLU drop.
Discussion (0). Continue with ORCID to comment.