You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Ship 1.4.0-beta: real Multi-GPU pipeline/tensor and API prompt fail-fast.
Enable Settings-driven GPU split, combined VRAM fit in Admin, CUDA remap, and surface prompt_too_long instead of client 408 hangs.
Co-authored-by: Cursor <cursoragent@cursor.com>
Double-click the Setup.exe → allow UAC → PyTorch CUDA downloads into the venv, ExLlamaV3 extension installs from the package → open **http://127.0.0.1:14563**.
@@ -261,7 +261,7 @@ By design (not product gaps for the EXL3 text + vision chat product):
261
261
-**OpenAI images / audio generation** — **501**; separate Media track
262
262
-**Native DLL generate** — worker-only for production text; DLL remains CI / scheduler ABI
263
263
-**MCP / hosted ReAct agent** — not embedded; use OpenAI `tools` / `tool_calls` with your own agent loop
264
-
-**Multi-GPU TP/PP/MP** — not supported by the EXL3 worker; Settings reject those modes. Multi-GPU visibility via `CUDA_VISIBLE_DEVICES` only
264
+
-**Multi-GPU model-parallel (MP)** — not supported. **Tensor** and **layer autosplit** (`pipeline`) are supported via the EXL3 worker on N NVIDIA GPUs (VRAM-proportional split, optional `GpuSplitGb`).
265
265
266
266
A/B tests: create via `/api/v1/ab`, then send `X-Ab-Test-Id` (or `model: "ab:<guid>"`) on chat/completions. The server assigns A/B via consistent hash, may load the selected model when it differs from the one currently on the GPU, tags audit, and returns `X-Ab-Variant`.
After changing GPU / parallelism / bind address, restart the Windows service so the process picks up the new host environment.
51
+
Saving GPU / parallelism settings recycles the Python worker and reloads the model. Bind address / port still need a Server restart.
52
52
53
53
## Backup & restore
54
54
@@ -60,16 +60,26 @@ Backups do **not** include multi‑GB weight files — back up the `models\` fol
60
60
61
61
## Multi-GPU
62
62
63
-
1. Confirm GPUs with `nvidia-smi` and **About** / Diagnostics.
64
-
2. Set `CudaVisibleDevices` to the device indices to expose (e.g. `0,1`).
65
-
3. Set `ParallelismMode`:
66
-
-`none` — single GPU
67
-
-`tensor` (TP) — split layers across GPUs (latency / large models)
68
-
-`pipeline` (PP) — pipeline stages across GPUs
69
-
-`model` (MP) — whole models on different devices (routing / multi-model)
70
-
4. Restart service and load the model again.
63
+
Works with any mix of NVIDIA GPUs (equal or different VRAM). No SKU is hardcoded.
71
64
72
-
Server helpers: `MultiGpuPlanner` validates device lists and maps modes for the native engine config. NCCL / full TP-PP production paths depend on the native CUDA build (not stub).
65
+
1. Confirm cards on **About** (PCI index, name, VRAM, UUID) or `nvidia-smi`.
66
+
2. Set `CudaVisibleDevices` to the **PCI / nvidia-smi** indices to use (`0,1` or `0,2`, …). Empty = all.
67
+
3. The worker process is started with `CUDA_DEVICE_ORDER=PCI_BUS_ID` and `CUDA_VISIBLE_DEVICES`**reordered** so `cuda:0` is the highest-VRAM GPU in that set.
5. Split: auto `VRAM[i] × GpuMemoryUtilization`, or override `GpuSplitGb` in **remapped** order (e.g. `10,4.5` for a 12 GB + 6 GB pair). Count must match visible GPUs.
74
+
6. Save Settings (or PATCH `/api/v1/settings`) — the worker is recycled and the loaded model is queued again.
75
+
76
+
`tensor` / `pipeline` require at least two indices. OOM usually hits the smallest card first — lower util or set `GpuSplitGb` with more reserve on the display GPU.
77
+
78
+
Examples (docs only — no SKU is hardcoded):
79
+
80
+
- Two equal 12 GB cards: `CudaVisibleDevices=0,1`, `tensor`, empty `GpuSplitGb`, util `0.85` → about `10.2,10.2`.
81
+
- 12 GB + 6 GB: same devices, or `GpuSplitGb=10,4.5` in **remapped** order (`cuda:0` is the 12 GB card even if `nvidia-smi` lists it as index 1).
82
+
- Three cards: `0,1,2` and a 3-value split. Same code path.
Copy file name to clipboardExpand all lines: docs/architecture.md
+14-12Lines changed: 14 additions & 12 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -16,18 +16,18 @@
16
16
│ ├─ EF Core + SQLite (ProgramData) │
17
17
│ ├─ Auth (API keys), rate limit, audit │
18
18
│ └─ ExLlamaSharp C# library │
19
-
│ LibraryImport → exllamasharp.dll │
19
+
│ EXL3 Python worker (production) │
20
20
└─────────────────────────────────────────┘
21
-
│ C ABI / P/Invoke
21
+
│ JSONL stdin/stdout
22
22
┌─────────────────────────────────────────┐
23
-
│ native/exllamasharp (C++ / optional CUDA)│
24
-
│ ├─ Scheduler (continuous batching) │
25
-
│ ├─ PageTable (KV pages / prefix cache) │
26
-
│ └─ Kernels / EXL3 path (non-stub) │
23
+
│ tools/exl3_worker + ExLlamaV3 │
24
+
│ ├─ Model.load (tensor_p / autosplit) │
25
+
│ ├─ Cache / Generator / Tokenizer │
26
+
│ └─ CUDA kernels from the venv │
27
27
└─────────────────────────────────────────┘
28
28
```
29
29
30
-
Optional later: `third_party/exllamav3` git submodule for upstream EXL3 kernels (see `third_party/README.md`).
30
+
Production inference is the **EXL3 Python worker**, not `exllamasharp_native.dll` (that stub is disabled). Multi-GPU TP / layer autosplit goes through `model.load`.
31
31
32
32
## .NET 10 host
33
33
@@ -46,26 +46,28 @@ Performance knobs: Server GC, sustained low-latency mode at startup, in-memory k
46
46
| Mode | CMake | Behavior |
47
47
|------|-------|----------|
48
48
| Stub |`-DEXL_STUB=ON`| No CUDA/LibTorch; deterministic fake generate; real scheduler ABI |
C ABI (`exllamasharp.h`) is the stability boundary — .NET uses source-generated `LibraryImport` (`NativeMethods`).
53
+
The native DLL is **disabled** at runtime (`EngineHostService`). Multi-GPU TP / autosplit is ExLlamaV3 `model.load` in the Python worker.
54
+
55
+
C ABI (`exllamasharp.h`) is the leftover native boundary — .NET still has `LibraryImport` (`NativeMethods`) but the Server does not load that DLL for inference.
54
56
55
57
## Request path (chat)
56
58
57
59
1. Client → `POST /v1/chat/completions` with Bearer key.
Copy file name to clipboardExpand all lines: docs/troubleshooting.md
+25-4Lines changed: 25 additions & 4 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,4 +1,4 @@
1
-
# ExLlamaSharp Troubleshooting
1
+
# ExLlamaSharp Troubleshooting
2
2
3
3
Aligns with the UI page **Diagnostics** (`/diagnostics`) and `GET /health`.
4
4
@@ -69,6 +69,20 @@ Components reported by `HealthService`:
69
69
- Ensure only intended devices in `CudaVisibleDevices`.
70
70
- Close other GPU apps (browsers with HW accel, games).
71
71
72
+
### Tensor parallel: Timed out waiting for worker
73
+
74
+
**Symptom:** Load with `parallelism_mode=tensor` fails with `TimeoutError: Timed out waiting for worker` (ExLlamaV3 `model_tp.py`). nvidia-smi stays almost idle; leftover `python ... spawn_main` processes sit at ~8 MB.
75
+
76
+
**Cause:** Tensor parallel starts extra Python processes. On Windows those children re-enter the worker script and never become TP workers when the host is the long-lived JSONL process. A console `python -c` load can succeed while the Admin/Server load fails. Pipeline mode does not spawn those children.
77
+
78
+
**Fix:** Use **pipeline** (and a KV cache that fits) to load across both GPUs. For Qwen3-32B 4.0bpw on 12 GB + 8 GB use `GpuSplitGb=10.8,6.2` and keep **Max batched tokens** around 2048–4096 — 16384 plus the 32B weights does not fit. Recycle the worker after a failed load (Save on Settings, or restart the Server) so zombie `spawn_main` processes are gone.
79
+
80
+
### VLM + multi-GPU: vision skipped
81
+
82
+
**Symptom:** Models such as `Qwen3.8-27B-exl3` (`Qwen3_5ForConditionalGeneration` + `vision_config`) used to fail load under tensor/pipeline with `vision models are not supported…`.
83
+
84
+
**Behaviour now:** The language model still loads across GPUs; the vision tower is skipped (`vision_capable=false`). Text chat works. Image/video inputs need **ParallelismMode=none** (and enough VRAM on one card), then reload.
@@ -83,9 +97,16 @@ Components reported by `HealthService`:
83
97
84
98
### Multi-GPU not used
85
99
86
-
- Mode still `none`, or only one device in `CudaVisibleDevices`.
87
-
- Stub/mock builds ignore real TP/PP — need CUDA native `exllamasharp.dll`.
88
-
- Restart after Settings changes.
100
+
-`ParallelismMode` still `none`, or only one index in `CudaVisibleDevices`.
101
+
-`nvidia-smi` index is **not**`cuda:N` after remap — check worker log (`cuda:0` = highest VRAM).
102
+
- Production path is the **EXL3 Python worker**, not `exllamasharp_native.dll` / mock.
103
+
- After Settings save the worker should recycle automatically; if VRAM is still on one UUID only, reload the model and read `use_per_device` in the worker log.
104
+
105
+
### OOM on the smaller GPU
106
+
107
+
- Auto-split is `VRAM[i] × GpuMemoryUtilization`. A 6 GB card with util 0.9 only has ~5.4 GB for weights.
108
+
- Display GPU keeps 1.5 GB (others 0.5 GB) folded into `use_per_device` — ExLlamaV3 does not accept use and reserve together. Lower util or set `GpuSplitGb` (remapped order, e.g. `10,4.5`).
109
+
- Tensor / pipeline need ≥2 devices. Speculative, vision, and LoRA are rejected under those modes.
0 commit comments