Person
1 entry · 27 May 2022
Reordering attention around GPU memory rather than approximating it cut training time and unlocked longer sequences — and became default infrastructure.
Ideas & essays · Compute & infrastructure