REVIEW 3 cited by
Blockwise Self-Attention for Long Document Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present BlockBERT, a lightweight and efficient BERT model for better modeling long-distance dependencies. Our model extends BERT by introducing sparse block structures into the attention matrix to reduce both memory consumption and training/inference time, which also enables attention heads to capture either short- or long-range contextual information. We conduct experiments on language model pre-training and several benchmark question answering datasets with various paragraph lengths. BlockBERT uses 18.7-36.1% less memory and 12.0-25.1% less time to learn the model. During testing, BlockBERT saves 27.8% inference time, while having comparable and sometimes better prediction accuracy, compared to an advanced BERT-based model, RoBERTa.
Forward citations
Cited by 3 Pith papers
-
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
Small transformers naturally encode task information in specific layers under limited conditions; a new auxiliary loss places a strong task vector at a chosen layer and improves out-of-distribution robustness.
-
Surface Vision Mamba: Leveraging Bidirectional State Space Model for Efficient Spherical Manifold Representation
Surface Vision Mamba (SiM) applies a bidirectional state space model to icosphere patches of cortical surfaces, matching or outperforming transformer- and GDL-based models with up to 4.8x faster inference and 91.7% lo...
-
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
MSWA assigns exponentially increasing attention window sizes across heads and layers, improving sliding-window attention's language modeling and reasoning performance at lower relative cost.
Discussion (0). Continue with ORCID to comment.