Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool/Back to issues
#193·LoRA

Why not merge BA during training?

Author: rp7svCreated Mar 17, 2025Updated Mar 6, 2026

During training, the low-rank matrices create a bypass, which introduces additional latency. So why not merge the low-rank weights directly into **_W_** during training, just like we do at inference time? I'm not sure whether this would affect gradient flow?

Source: microsoft/LoRA

View original on GitHubView discussion on GitHub