Discussion about LISA
Author: caoshuai03Created May 20, 2024Updated Jan 26, 2025
In the article, only the comparison of the average weight paradigm of each layer during lora fine-tuning is given.
- But what if their weights are different before fine-tuning?
- Using the weighted averaging method, is it possible that there are differences in the intermediate iteration process?
Source: OptimalScale/LMFlow