← back to paper
arxiv: 2608.06776 · 2 revisions
Faster Query-Key Learning Sharpens Attention in Self-Attention Models