Person
1 entry · 30 November 2022
Leviathan, Kalman and Matias showed a small draft model verified by the large model can cut inference latency 2-3x with identical outputs.
Ideas & essays