Pith. sign in

REVIEW 3 cited by

Blockwise Self-Attention for Long Document Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.02972 v2 pith:D7DPK2NN submitted 2019-11-07 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelblockberttimeattentionbertbetterinferenceless
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present BlockBERT, a lightweight and efficient BERT model for better modeling long-distance dependencies. Our model extends BERT by introducing sparse block structures into the attention matrix to reduce both memory consumption and training/inference time, which also enables attention heads to capture either short- or long-range contextual information. We conduct experiments on language model pre-training and several benchmark question answering datasets with various paragraph lengths. BlockBERT uses 18.7-36.1% less memory and 12.0-25.1% less time to learn the model. During testing, BlockBERT saves 27.8% inference time, while having comparable and sometimes better prediction accuracy, compared to an advanced BERT-based model, RoBERTa.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Task Vectors in In-Context Learning: Emergence, Formation, and Benefit

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Small transformers naturally encode task information in specific layers under limited conditions; a new auxiliary loss places a strong task vector at a chosen layer and improves out-of-distribution robustness.

  2. Surface Vision Mamba: Leveraging Bidirectional State Space Model for Efficient Spherical Manifold Representation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Surface Vision Mamba (SiM) applies a bidirectional state space model to icosphere patches of cortical surfaces, matching or outperforming transformer- and GDL-based models with up to 4.8x faster inference and 91.7% lo...

  3. MSWA: Refining Local Attention with Multi-ScaleWindow Attention

    cs.CL 2025-01 conditional novelty 4.0 of 10

    MSWA assigns exponentially increasing attention window sizes across heads and layers, improving sliding-window attention's language modeling and reasoning performance at lower relative cost.

Pith tools