Pith. sign in

REVIEW 2 cited by

HDT: Hierarchical Document Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08330 v1 pith:Y27SP73J submitted 2024-07-11 cs.LG

classification cs.LG
keywords documentshierarchicaldocumentstructureattentionsparsetransformerefficiency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we propose the Hierarchical Document Transformer (HDT), a novel sparse Transformer architecture tailored for structured hierarchical documents. Such documents are extremely important in numerous domains, including science, law or medicine. However, most existing solutions are inefficient and fail to make use of the structure inherent to documents. HDT exploits document structure by introducing auxiliary anchor tokens and redesigning the attention mechanism into a sparse multi-level hierarchy. This approach facilitates information exchange between tokens at different levels while maintaining sparsity, thereby enhancing computational and memory efficiency while exploiting the document structure as an inductive bias. We address the technical challenge of implementing HDT's sample-dependent hierarchical attention pattern by developing a novel sparse attention kernel that considers the hierarchical structure of documents. As demonstrated by our experiments, utilizing structural information present in documents leads to faster convergence, higher sample efficiency and better performance on downstream tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Overlapping Schwarz Attention: Hierarchical Attention via Domain Decomposition

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    A two-level overlapping Schwarz domain decomposition constructs a hierarchical attention operator that trains faster and approximates the inverse of a discretized 1D diffusion operator more accurately than global low-...

  2. From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    GRACE pairs motion-weighted video tokens with AI-refined emotion text tokens using optimal transport, reporting new UAR and WAR records on DFEW, FERV39k, and MAFW.

Pith tools