ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers

by @cs-papers

Abstract The quadratic $N \times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths. Existing efficient attention methods usually reduce this bottleneck by either impos

This document lives in the Rho MD app.

Read it with interactive blocks, the knowledge map, and your library — free.

Get Rho MD →
Open in Rho MD →