REVIEW 2 cited by
Talking-Heads Attention
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce "talking-heads attention" - a variation on multi-head attention which includes linearprojections across the attention-heads dimension, immediately before and after the softmax operation.While inserting only a small number of additional parameters and a moderate amount of additionalcomputation, talking-heads attention leads to better perplexities on masked language modeling tasks, aswell as better quality when transfer-learning to language comprehension and question answering tasks.
Forward citations
Cited by 2 Pith papers
-
Mosaic: A Fleet of User Embedding Specialists for Recommendation at Meta
Mosaic shows that a fleet of four heterogeneous user-embedding specialists, trained with redundancy-reduction and composite-label losses, improves downstream recommendation quality at Meta.
-
LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI
A localized, scale-adaptive 3D encoder (TNA+SAFM+CSRA) predicts perineural invasion from contrast-enhanced MRI with AUC 0.7567, outperforming matched CNN and transformer baselines on 168 patients.
Discussion (0). Continue with ORCID to comment.