-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathrun.sh
More file actions
executable file
·207 lines (191 loc) · 10.3 KB
/
Copy pathrun.sh
File metadata and controls
executable file
·207 lines (191 loc) · 10.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
#!/usr/bin/env bash
# inferhost local dev wrapper.
# Reads .env (if present), ensures a venv exists, and routes high-level commands
# to the inferhost package. The user-facing inferhost command is TUI-only —
# this wrapper exists for repo-development convenience (install / start / stop / status).
set -euo pipefail
PROJECT_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
VENV_DIR="${PROJECT_ROOT}/.venv"
# Load .env if present, ignoring comments/blank lines.
if [[ -f "${PROJECT_ROOT}/.env" ]]; then
set -o allexport
# shellcheck disable=SC1091
source <(grep -E '^[A-Z_][A-Z0-9_]*=' "${PROJECT_ROOT}/.env")
set +o allexport
fi
usage() {
cat <<'EOF'
inferhost run.sh — repo wrapper around the TUI-only inferhost package.
Usage: ./run.sh <command>
Commands:
install Create the project venv and pip install -e ".[dev]".
Runtime binaries (llama.cpp, llama-swap) are auto-downloaded
on the first launch of the TUI. inferhost pulls llama-server
straight from upstream ggml-org/llama.cpp releases — no
cmake required. The variant (Vulkan / ROCm / SYCL / OpenVINO /
CPU / Metal) is picked by hardware probe and overridable via
INFERHOST_LLAMACPP_BACKEND or the TUI Settings screen. Set
INFERHOST_LLAMA_SERVER_PATH to use a custom binary instead
(e.g. a self-built CUDA llama-server).
Every engine (llama-server, llama-swap, sd-server,
llama-tts) is prebuilt-binary-only — no compiling. Kokoro
text-to-speech runs in-process via ONNX Runtime, installed
with inferhost's Python dependencies.
start Launch the TUI (alias of `run`). The TUI is the only UI.
run Launch the TUI.
start-bg Start llama-swap + LiteLLM gateway as background daemons,
no TUI. They survive your shell session — kill them with
`./run.sh stop`. Requires at least one model already
registered (use the TUI to add one the first time). If a
text-to-speech model is registered, the inferhost-tts daemon
(serves /v1/audio/speech) is started too. Image-generation
models (stable-diffusion.cpp) ride llama-swap and serve
/v1/images/generations automatically — no extra daemon.
The inferhost-pinwatch daemon rides along with llama-swap and
re-loads pinned models into VRAM whenever they get evicted
(by a swap, a crash, or a restart) and the GPU is idle again.
stop Stop llama-swap, pinwatch, the LiteLLM gateway, and
inferhost-tts.
restart Stop + start-bg in one shot. Picks up any config edits.
prune `./run.sh prune [--yes] [repo ...]` — list Hugging Face cache
repos that no registered model uses, with sizes. Deleting a
model now removes its weights too, but installs that predate
that carry stranded downloads; this reclaims them. Lists only
unless you pass --yes. Name repos to delete just those —
the HF cache is shared with any other tool on the box that
uses huggingface_hub (ComfyUI, vLLM, Whisper, ...), so
"unused by inferhost" does not mean unused.
update `./run.sh update [bNNNN]` — re-download the runtime binaries
(llama.cpp's llama-server + llama-tts, llama-swap, and
sd-server if image generation is installed) from upstream,
then bring back whatever daemons were running. Use this when
a just-released model refuses to load with "unknown model
architecture": that means the llama.cpp on disk is older than
the model. Binaries are otherwise only fetched on first
launch, so they go stale. Pass an explicit upstream tag
(e.g. `./run.sh update b10353`) to install one specific
build; with no argument it follows INFERHOST_LLAMACPP_VERSION
(default "latest"). Skipped for llama-server when
INFERHOST_LLAMA_SERVER_PATH points at your own build — add
`--rebuild` (e.g. `./run.sh update --rebuild`) to recompile
that build in place instead: inferhost checks out the target
tag in the llama.cpp checkout it came from and runs cmake
with the flags already in its CMakeCache, so a CUDA build
stays a CUDA build. Needed on NVIDIA/Linux, where upstream
publishes no prebuilt CUDA binary.
autostart `./run.sh autostart on|off|status` — install/remove a systemd
user unit so the daemons start automatically at boot (enables
user lingering so no login is needed). `status` shows whether
it is installed and enabled.
status Print daemon + endpoint status (no UI). Note: llama-swap
(swap) listens on 127.0.0.1 by default (internal); the
user-visible gateway is the LiteLLM port (INFERHOST_GATEWAY_PORT).
To expose llama-swap externally, set INFERHOST_SWAP_HOST=0.0.0.0.
uninstall Remove the venv and runtime data dir (keeps HF model cache).
reset Stop daemons and clear the model registry / generated configs.
test Run pytest.
lint Run ruff over src/ and tests/.
shell Open a shell with the venv activated.
docker-build Build the inferhost-test:v0.5 image (CUDA 12.4 base, GPU-ready).
docker-smoke In-container static smoke (imports, pytest, GPU visibility,
release availability). Fast — no model download. Requires
the NVIDIA Container Toolkit on the host (--gpus all).
docker-functional Full end-to-end test: downloads a tiny GGUF, registers it,
starts llama-swap, sends a real chat completion through
llama-server, and verifies the live process has the default
-ctk q8_0 -ctv q8_0 KV flags. Model is cached in a named
volume so subsequent runs are offline.
docker-test Run pytest inside the container.
docker-shell Drop into a bash shell in the running container.
docker-clean Stop the test container and remove its named volumes.
help Show this help.
Configuration lives in .env (see .env.example). Variables include:
INFERHOST_SWAP_PORT INFERHOST_GATEWAY_PORT
INFERHOST_DATA_DIR INFERHOST_CONFIG_DIR INFERHOST_HF_CACHE
INFERHOST_GPU_LAYERS INFERHOST_DEFAULT_CTX INFERHOST_FLASH_ATTENTION
INFERHOST_LLAMACPP_VERSION INFERHOST_LLAMASWAP_VERSION
INFERHOST_LLAMACPP_BACKEND (vulkan|cuda|rocm|sycl|openvino|cpu; auto by default)
EOF
}
ensure_venv() {
if [[ ! -d "${VENV_DIR}" ]]; then
echo ">>> Creating venv at ${VENV_DIR}"
if command -v uv >/dev/null 2>&1; then
uv venv --python 3.12 "${VENV_DIR}"
else
python3 -m venv "${VENV_DIR}"
fi
fi
# shellcheck disable=SC1091
source "${VENV_DIR}/bin/activate"
}
install_dev() {
ensure_venv
if command -v uv >/dev/null 2>&1; then
uv pip install -e ".[dev]"
else
pip install -e ".[dev]"
fi
echo ">>> Install complete. Run './run.sh start' to launch the TUI."
echo " (Runtime binaries are downloaded automatically on first launch.)"
}
uninstall_local() {
echo ">>> Stopping any running daemons"
ensure_venv 2>/dev/null && python -m inferhost._ops stop 2>/dev/null || true
echo ">>> Removing venv: ${VENV_DIR}"
rm -rf "${VENV_DIR}"
echo ">>> Removing runtime data: ${INFERHOST_DATA_DIR:-$HOME/.local/share/inferhost}"
rm -rf "${INFERHOST_DATA_DIR:-$HOME/.local/share/inferhost}"
echo ">>> Removing config: ${INFERHOST_CONFIG_DIR:-$HOME/.config/inferhost}"
rm -rf "${INFERHOST_CONFIG_DIR:-$HOME/.config/inferhost}"
echo "Done. (Hugging Face model cache at ${INFERHOST_HF_CACHE:-$HOME/.cache/huggingface} was kept.)"
}
reset_state() {
ensure_venv
python -m inferhost._ops stop || true
rm -f "${INFERHOST_CONFIG_DIR:-$HOME/.config/inferhost}/models.toml"
rm -f "${INFERHOST_CONFIG_DIR:-$HOME/.config/inferhost}/llama-swap.yaml"
rm -f "${INFERHOST_CONFIG_DIR:-$HOME/.config/inferhost}/litellm.yaml"
echo "Registry cleared. (Model files remain in HF cache.)"
}
COMPOSE_FILE="${PROJECT_ROOT}/docker-compose.test.yml"
DOCKER_SVC="inferhost"
docker_compose() {
if docker compose version >/dev/null 2>&1; then
docker compose -f "${COMPOSE_FILE}" "$@"
else
docker-compose -f "${COMPOSE_FILE}" "$@"
fi
}
docker_ensure_running() {
if ! docker_compose ps --status running --services 2>/dev/null | grep -qx "${DOCKER_SVC}"; then
echo ">>> Starting test container"
docker_compose up -d
fi
}
cmd="${1:-help}"
shift || true
case "${cmd}" in
install) install_dev ;;
uninstall) uninstall_local ;;
start|run|tui) ensure_venv; inferhost ;;
start-bg) ensure_venv; python -m inferhost._ops start ;;
stop) ensure_venv; python -m inferhost._ops stop ;;
restart) ensure_venv; python -m inferhost._ops restart ;;
update) ensure_venv; python -m inferhost._ops update "$@" ;;
prune) ensure_venv; python -m inferhost._ops prune "$@" ;;
status) ensure_venv; python -m inferhost._ops status ;;
autostart) ensure_venv; python -m inferhost._ops autostart "$@" ;;
reset) reset_state ;;
test) ensure_venv; pytest -v "$@" ;;
lint) ensure_venv; ruff check src tests "$@" ;;
shell) ensure_venv; exec "${SHELL:-bash}" ;;
docker-build) docker_compose build "$@" ;;
docker-smoke) docker_ensure_running; docker_compose exec "${DOCKER_SVC}" inferhost-smoke ;;
docker-functional) docker_ensure_running; docker_compose exec "${DOCKER_SVC}" inferhost-functional ;;
docker-test) docker_ensure_running; docker_compose exec "${DOCKER_SVC}" pytest tests/ -v "$@" ;;
docker-shell) docker_ensure_running; docker_compose exec "${DOCKER_SVC}" bash ;;
docker-clean) docker_compose down -v --remove-orphans ;;
help|-h|--help|"") usage ;;
*) echo "Unknown command: ${cmd}" >&2; usage; exit 2 ;;
esac