Pith. sign in

REVIEW 1 cited by

Pay Attention when Required

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.04534 v3 pith:JWDWNIKR submitted 2020-09-09 cs.LG cs.CL

classification cs.LGcs.CL
keywords blockscapturefeed-forwardmeaningself-attentiontransformerachievedarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer-based models consist of interleaved feed-forward blocks - that capture content meaning, and relatively more expensive self-attention blocks - that capture context meaning. In this paper, we explored trade-offs and ordering of the blocks to improve upon the current Transformer architecture and proposed PAR Transformer. It needs 35% lower compute time than Transformer-XL achieved by replacing ~63% of the self-attention blocks with feed-forward blocks, and retains the perplexity on WikiText-103 language modelling benchmark. We further validated our results on text8 and enwiki8 datasets, as well as on the BERT model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Language Modeling for Low-Resource Settings with Hybrid RNN-Transformer Architectures

    cs.CL 2025-02 conditional novelty 4.0 of 10

    A hybrid architecture with two QRNN layers followed by a PAR Transformer reaches 1.013 BPC on enwik8 and 20.91 PPL on Wikitext-103 with 41-60M parameters.

Pith tools