#7652·textgen

How to stream model from Disk to VRAM?

Author: averageaidudeCreated Aug 27, 2026Updated Aug 27, 2026
Labelsenhancement

There are a view new models wich need more VRAM than my 72GB VRAM. Does anybody has experience to run a gguf model from nvme? And yes i have a lack of RAM just 32GB. But i remember that llama.ccp can do this and just needs a small buffer for caching. Any idea how to set this up in Oobabooga?