There are a view new models wich need more VRAM than my 72GB VRAM. Does anybody has experience to run a gguf model from nvme? And yes i have a lack of RAM just 32GB. But i remember that llama.ccp can do this and just needs a small buffer for caching. Any idea how to set this up in Oobabooga?
There are a view new models wich need more VRAM than my 72GB VRAM. Does anybody has experience to run a gguf model from nvme? And yes i have a lack of RAM just 32GB. But i remember that llama.ccp can do this and just needs a small buffer for caching. Any idea how to set this up in Oobabooga?