no code implementations • 26 May 2021 • Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, Yang You
That is, with sparse attention, our sequence parallelism enables us to train transformer with infinite long sequence.