QLoRA makes fine-tuning fit on one GPU
Quantised low-rank adaptation let a 65-billion-parameter model be fine-tuned on a single consumer card, and the resulting Guanaco model claimed 99% of ChatGPT's quality after 24 hours' training.
Open weights & ecosystem · Ideas & essays