We are trying kvcached, but sometimes failed to start both 2 qwen3/0.6B instances with vllm/vllm-openai:v0.8.5.post1 (can't use v0.11.1 since cuda version < 12.8) due to memory allocation. Could you give suggestions to start mutiple models? Like set gpu-memory-utilization or any other parameters? We set gpu-memory-utilization 0.4, but looked not always work. Thanks!
We are trying kvcached, but sometimes failed to start both 2 qwen3/0.6B instances with vllm/vllm-openai:v0.8.5.post1 (can't use v0.11.1 since cuda version < 12.8) due to memory allocation. Could you give suggestions to start mutiple models? Like set gpu-memory-utilization or any other parameters? We set gpu-memory-utilization 0.4, but looked not always work. Thanks!