Papers to read:

  1. LoRA: Low-Rank Adaptation of Large Language Models. 
  2. QLoRA: Efficient Finetuning of Quantized LLMs. 
  3. Accurate LoRA-Finetuning Quantization of LLMs via Information Retention.