Skip to content
Spencer Bull edited this page Jul 21, 2026 · 2 revisions

Yokai

One terminal to deploy, monitor, and manage LLM inference across all your GPUs.

Yokai is an open-source, terminal-first GPU fleet manager for running vLLM, llama.cpp, and ComfyUI across one or more machines. It combines device onboarding, curated Best-Known-Configs, guided deployment, live fleet monitoring, service lifecycle controls, and AI coding-tool configuration in one cross-platform application.

Official documentation: spencerbull.github.io/yokai-docs

Get started

  1. Install Yokai
  2. Run the quick start
  3. Configure Yokai

Common tasks

Reference

Project links

Clone this wiki locally