Skip to content

Commit 8a6b6ae

Browse files
Ship 1.4.0-beta: real Multi-GPU pipeline/tensor and API prompt fail-fast.
Enable Settings-driven GPU split, combined VRAM fit in Admin, CUDA remap, and surface prompt_too_long instead of client 408 hangs. Co-authored-by: Cursor <cursoragent@cursor.com>
1 parent 9a2ab94 commit 8a6b6ae

38 files changed

Lines changed: 1597 additions & 241 deletions

‎README.md‎

Lines changed: 6 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -1,4 +1,4 @@
1-
# ExLlamaSharp
1+
# ExLlamaSharp
22

33
**Local LLM server for Windows with NVIDIA GPUs.**
44

@@ -11,7 +11,7 @@ Inspired by:
1111
- **ExLlamaV3** — fast EXL3 inference on NVIDIA
1212
- **Open WebUI** — browser-based administration
1313

14-
**Current release: 1.2.1** — product-final (4 bars at 100%): Core EXL3 OpenAI chat, Models/Jobs/Keys, Admin avançado (LoRA, speculative, multi-GPU honesty, webhooks, tenants), Agentic tools + vision multimodal chat. OpenAI **images/audio generation** remain **501** (separate media track). Setup.exe bundles the ExLlamaV3 CUDA `.pyd`, worker deps, Python installer and VC++. PyTorch CUDA is downloaded during install. Admin → Models shows a VRAM fit badge (Fits / Tight / Too large; estimate only).
14+
**Current release: 1.4.0-beta** — real Multi-GPU (pipeline / tensor), combined VRAM fit in Admin, CUDA device remap (strongest GPU first), and faster fail for oversized prompts (`prompt_too_long` instead of hanging to 408). Core EXL3 OpenAI chat, Models/Jobs/Keys, LoRA, speculative, webhooks, tenants, agentic tools. Vision models skip the vision tower under multi-GPU (text-only). OpenAI **images/audio generation** remain **501**. Setup.exe bundles the ExLlamaV3 CUDA `.pyd`, worker deps, Python installer and VC++.
1515

1616
Default after install: **http://127.0.0.1:14563**
1717

@@ -24,12 +24,12 @@ Default after install: **http://127.0.0.1:14563**
2424

2525
## Download (Windows x64)
2626

27-
[**ExLlamaSharp-Setup-win-x64.exe**](https://github.com/vitorcastro78/ExLlamaSharp/releases/latest/download/ExLlamaSharp-Setup-win-x64.exe) — latest GitHub Release.
27+
[**ExLlamaSharp-Setup-win-x64.exe**](https://github.com/Kortexio/ExLlamaSharp/releases/latest/download/ExLlamaSharp-Setup-win-x64.exe) — latest GitHub Release.
2828

2929
One-liner (downloads Setup and launches UAC):
3030

3131
```powershell
32-
irm https://raw.githubusercontent.com/vitorcastro78/ExLlamaSharp/main/packaging/install-web.ps1 | iex
32+
irm https://raw.githubusercontent.com/Kortexio/ExLlamaSharp/main/packaging/install-web.ps1 | iex
3333
```
3434

3535
Double-click the Setup.exe → allow UAC → PyTorch CUDA downloads into the venv, ExLlamaV3 extension installs from the package → open **http://127.0.0.1:14563**.
@@ -220,7 +220,7 @@ Persisted server settings:
220220
|-----|----------|
221221
| Network | Bind address, port, CORS, TLS cert path |
222222
| Performance | Max sequences, chunk size, batched tokens, GPU memory util, request timeout |
223-
| Multi-GPU | `CUDA_VISIBLE_DEVICES`, parallelism mode (validated; worker gets device list) |
223+
| Multi-GPU | PCI devices, `none` / `tensor` / `pipeline`, GPU memory util, optional `GpuSplitGb`. Save recycles the worker. |
224224
| Speculative | Enable + draft model + draft K (forwarded to worker) |
225225
| Startup | Load last model on startup, models path |
226226
| Hugging Face | Optional `hf_…` token (also reads `HF_TOKEN`) |
@@ -261,7 +261,7 @@ By design (not product gaps for the EXL3 text + vision chat product):
261261
- **OpenAI images / audio generation** — **501**; separate Media track
262262
- **Native DLL generate** — worker-only for production text; DLL remains CI / scheduler ABI
263263
- **MCP / hosted ReAct agent** — not embedded; use OpenAI `tools` / `tool_calls` with your own agent loop
264-
- **Multi-GPU TP/PP/MP** — not supported by the EXL3 worker; Settings reject those modes. Multi-GPU visibility via `CUDA_VISIBLE_DEVICES` only
264+
- **Multi-GPU model-parallel (MP)** — not supported. **Tensor** and **layer autosplit** (`pipeline`) are supported via the EXL3 worker on N NVIDIA GPUs (VRAM-proportional split, optional `GpuSplitGb`).
265265

266266
A/B tests: create via `/api/v1/ab`, then send `X-Ab-Test-Id` (or `model: "ab:<guid>"`) on chat/completions. The server assigns A/B via consistent hash, may load the selected model when it differs from the one currently on the GPU, tags audit, and returns `X-Ab-Variant`.
267267

‎docs/admin-guide.md‎

Lines changed: 22 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -1,4 +1,4 @@
1-
# ExLlamaSharp Admin Guide
1+
# ExLlamaSharp Admin Guide
22

33
Operational guide for administrators of the Windows service and Blazor UI.
44

@@ -44,11 +44,11 @@ Important fields:
4444
| Backup | `AutoBackupSchedule` (`disabled` / `daily` / `weekly`) |
4545
| Webhooks | `WebhookUrl`, `WebhookSecret` |
4646
| Features | content moderation, multi-tenancy, advanced metrics |
47-
| GPU | `CudaVisibleDevices` (e.g. `0` or `0,1`), `ParallelismMode` (`none` / `tensor` / `pipeline` / `model`) |
47+
| GPU | `CudaVisibleDevices` (PCI indices, e.g. `0,1`), `ParallelismMode` (`none` / `tensor` / `pipeline`), `GpuMemoryUtilization`, optional `GpuSplitGb` |
4848
| Speculative | `SpeculativeEnabled`, `DraftModelId`, `DraftK` |
4949
| Paths | `ModelsPath` |
5050

51-
After changing GPU / parallelism / bind address, restart the Windows service so the process picks up the new host environment.
51+
Saving GPU / parallelism settings recycles the Python worker and reloads the model. Bind address / port still need a Server restart.
5252

5353
## Backup & restore
5454

@@ -60,16 +60,26 @@ Backups do **not** include multi‑GB weight files — back up the `models\` fol
6060

6161
## Multi-GPU
6262

63-
1. Confirm GPUs with `nvidia-smi` and **About** / Diagnostics.
64-
2. Set `CudaVisibleDevices` to the device indices to expose (e.g. `0,1`).
65-
3. Set `ParallelismMode`:
66-
- `none` — single GPU
67-
- `tensor` (TP) — split layers across GPUs (latency / large models)
68-
- `pipeline` (PP) — pipeline stages across GPUs
69-
- `model` (MP) — whole models on different devices (routing / multi-model)
70-
4. Restart service and load the model again.
63+
Works with any mix of NVIDIA GPUs (equal or different VRAM). No SKU is hardcoded.
7164

72-
Server helpers: `MultiGpuPlanner` validates device lists and maps modes for the native engine config. NCCL / full TP-PP production paths depend on the native CUDA build (not stub).
65+
1. Confirm cards on **About** (PCI index, name, VRAM, UUID) or `nvidia-smi`.
66+
2. Set `CudaVisibleDevices` to the **PCI / nvidia-smi** indices to use (`0,1` or `0,2`, …). Empty = all.
67+
3. The worker process is started with `CUDA_DEVICE_ORDER=PCI_BUS_ID` and `CUDA_VISIBLE_DEVICES` **reordered** so `cuda:0` is the highest-VRAM GPU in that set.
68+
4. `ParallelismMode`:
69+
- `none` — load on `cuda:0` only
70+
- `tensor` — ExLlamaV3 tensor parallelism (`tensor_p=True`, backend `native`)
71+
- `pipeline` — layer autosplit (`tensor_p=False` + `use_per_device`)
72+
- `model` is **rejected** (not implemented)
73+
5. Split: auto `VRAM[i] × GpuMemoryUtilization`, or override `GpuSplitGb` in **remapped** order (e.g. `10,4.5` for a 12 GB + 6 GB pair). Count must match visible GPUs.
74+
6. Save Settings (or PATCH `/api/v1/settings`) — the worker is recycled and the loaded model is queued again.
75+
76+
`tensor` / `pipeline` require at least two indices. OOM usually hits the smallest card first — lower util or set `GpuSplitGb` with more reserve on the display GPU.
77+
78+
Examples (docs only — no SKU is hardcoded):
79+
80+
- Two equal 12 GB cards: `CudaVisibleDevices=0,1`, `tensor`, empty `GpuSplitGb`, util `0.85` → about `10.2,10.2`.
81+
- 12 GB + 6 GB: same devices, or `GpuSplitGb=10,4.5` in **remapped** order (`cuda:0` is the 12 GB card even if `nvidia-smi` lists it as index 1).
82+
- Three cards: `0,1,2` and a 3-value split. Same code path.
7383

7484
## Webhooks
7585

‎docs/architecture.md‎

Lines changed: 14 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -16,18 +16,18 @@
1616
│ ├─ EF Core + SQLite (ProgramData) │
1717
│ ├─ Auth (API keys), rate limit, audit │
1818
│ └─ ExLlamaSharp C# library │
19-
│ LibraryImport → exllamasharp.dll │
19+
│ EXL3 Python worker (production) │
2020
└─────────────────────────────────────────┘
21-
│ C ABI / P/Invoke
21+
│ JSONL stdin/stdout
2222
┌─────────────────────────────────────────┐
23-
│ native/exllamasharp (C++ / optional CUDA)│
24-
│ ├─ Scheduler (continuous batching) │
25-
│ ├─ PageTable (KV pages / prefix cache) │
26-
│ └─ Kernels / EXL3 path (non-stub) │
23+
│ tools/exl3_worker + ExLlamaV3 │
24+
│ ├─ Model.load (tensor_p / autosplit) │
25+
│ ├─ Cache / Generator / Tokenizer │
26+
│ └─ CUDA kernels from the venv │
2727
└─────────────────────────────────────────┘
2828
```
2929

30-
Optional later: `third_party/exllamav3` git submodule for upstream EXL3 kernels (see `third_party/README.md`).
30+
Production inference is the **EXL3 Python worker**, not `exllamasharp_native.dll` (that stub is disabled). Multi-GPU TP / layer autosplit goes through `model.load`.
3131

3232
## .NET 10 host
3333

@@ -46,26 +46,28 @@ Performance knobs: Server GC, sustained low-latency mode at startup, in-memory k
4646
| Mode | CMake | Behavior |
4747
|------|-------|----------|
4848
| Stub | `-DEXL_STUB=ON` | No CUDA/LibTorch; deterministic fake generate; real scheduler ABI |
49-
| CUDA | `-DEXL_STUB=OFF` + LibTorch + toolkit | Full path toward EXL3 GEMM / multi-GPU |
49+
| CUDA | `-DEXL_STUB=OFF` + LibTorch + toolkit | Optional native experiment — **not** the production multi-GPU path |
5050

5151
Build helper: `packaging/build-native-stub.ps1`. Details: `native/exllamasharp/README.md`.
5252

53-
C ABI (`exllamasharp.h`) is the stability boundary — .NET uses source-generated `LibraryImport` (`NativeMethods`).
53+
The native DLL is **disabled** at runtime (`EngineHostService`). Multi-GPU TP / autosplit is ExLlamaV3 `model.load` in the Python worker.
54+
55+
C ABI (`exllamasharp.h`) is the leftover native boundary — .NET still has `LibraryImport` (`NativeMethods`) but the Server does not load that DLL for inference.
5456

5557
## Request path (chat)
5658

5759
1. Client → `POST /v1/chat/completions` with Bearer key.
5860
2. Middleware: auth + rate limit; optional moderation.
5961
3. Chat template formats messages → token ids.
60-
4. `EngineHostService` submits a job to `ExLlamaEngine` (native or mock).
61-
5. Native scheduler batches; tokens stream back as SSE chunks if requested.
62+
4. `EngineHostService` submits a job to `ExLlamaV3WorkerEngine` (or mock in Development).
63+
5. The Python worker iterates the ExLlamaV3 Generator; tokens stream back as SSE chunks if requested.
6264
6. Audit / webhooks / metrics updated asynchronously.
6365

6466
## Multi-GPU & advanced
6567

6668
Server-side helpers prepare config for the engine:
6769

68-
- `MultiGpuPlanner` — TP / PP / MP from `CudaVisibleDevices` + `ParallelismMode` (validated; worker receives `CUDA_VISIBLE_DEVICES`)
70+
- `MultiGpuPlanner` — validates `none` / `tensor` / `pipeline` + device list + optional `GpuSplitGb`; worker spawn remaps `CUDA_VISIBLE_DEVICES` strongest-first and `model.load` gets `tensor_p` / `use_per_device`
6971
- `SpeculativeDecodingOptions` — draft model + `DraftK` (forwarded to EXL3 worker)
7072
- `ArchitectureDetector` — llama / qwen / mixtral / llava from `config.json`
7173
- `QuantizationModes` — EXL3 convert via `exllamav3.conversion.convert_model`

‎docs/comparison.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -9,7 +9,7 @@ Positioning: **Ollama’s ease + strong local NVIDIA EXL3 serving + Windows admi
99
| Web UI admin | Yes (Blazor) | No | No (API only) | Yes (desktop) |
1010
| OpenAI-compatible API | Yes (`/v1` + Ollama-style `options`) | Yes | Yes | Yes |
1111
| No Docker required | Yes | Yes* | Typically containers/Linux | Yes |
12-
| Multi-GPU TP / PP / MP | **No** (rejected in Settings); `CUDA_VISIBLE_DEVICES` only | Limited | Yes | Limited |
12+
| Multi-GPU TP / layer autosplit | **Yes** (ExLlamaV3 `tensor` / `pipeline`; N NVIDIA, split by VRAM). MP not supported | Limited | Yes | Limited |
1313
| Non-technical friendly | Yes (wizard + UI) | Yes | No | Yes |
1414
| Team / tenant management | Yes (optional MultiTenancy) | No | DIY | No |
1515
| API keys, quotas, audit | Yes | Basic | DIY / gateway | Basic |

‎docs/troubleshooting.md‎

Lines changed: 25 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -1,4 +1,4 @@
1-
# ExLlamaSharp Troubleshooting
1+
# ExLlamaSharp Troubleshooting
22

33
Aligns with the UI page **Diagnostics** (`/diagnostics`) and `GET /health`.
44

@@ -69,6 +69,20 @@ Components reported by `HealthService`:
6969
- Ensure only intended devices in `CudaVisibleDevices`.
7070
- Close other GPU apps (browsers with HW accel, games).
7171

72+
### Tensor parallel: Timed out waiting for worker
73+
74+
**Symptom:** Load with `parallelism_mode=tensor` fails with `TimeoutError: Timed out waiting for worker` (ExLlamaV3 `model_tp.py`). nvidia-smi stays almost idle; leftover `python ... spawn_main` processes sit at ~8 MB.
75+
76+
**Cause:** Tensor parallel starts extra Python processes. On Windows those children re-enter the worker script and never become TP workers when the host is the long-lived JSONL process. A console `python -c` load can succeed while the Admin/Server load fails. Pipeline mode does not spawn those children.
77+
78+
**Fix:** Use **pipeline** (and a KV cache that fits) to load across both GPUs. For Qwen3-32B 4.0bpw on 12 GB + 8 GB use `GpuSplitGb=10.8,6.2` and keep **Max batched tokens** around 2048–4096 — 16384 plus the 32B weights does not fit. Recycle the worker after a failed load (Save on Settings, or restart the Server) so zombie `spawn_main` processes are gone.
79+
80+
### VLM + multi-GPU: vision skipped
81+
82+
**Symptom:** Models such as `Qwen3.8-27B-exl3` (`Qwen3_5ForConditionalGeneration` + `vision_config`) used to fail load under tensor/pipeline with `vision models are not supported…`.
83+
84+
**Behaviour now:** The language model still loads across GPUs; the vision tower is skipped (`vision_capable=false`). Text chat works. Image/video inputs need **ParallelismMode=none** (and enough VRAM on one card), then reload.
85+
7286
### Slow tokens / queue buildup
7387

7488
- Check `GET /metrics` (`jobs_waiting`, `tokens_per_second`).
@@ -83,9 +97,16 @@ Components reported by `HealthService`:
8397

8498
### Multi-GPU not used
8599

86-
- Mode still `none`, or only one device in `CudaVisibleDevices`.
87-
- Stub/mock builds ignore real TP/PP — need CUDA native `exllamasharp.dll`.
88-
- Restart after Settings changes.
100+
- `ParallelismMode` still `none`, or only one index in `CudaVisibleDevices`.
101+
- `nvidia-smi` index is **not** `cuda:N` after remap — check worker log (`cuda:0` = highest VRAM).
102+
- Production path is the **EXL3 Python worker**, not `exllamasharp_native.dll` / mock.
103+
- After Settings save the worker should recycle automatically; if VRAM is still on one UUID only, reload the model and read `use_per_device` in the worker log.
104+
105+
### OOM on the smaller GPU
106+
107+
- Auto-split is `VRAM[i] × GpuMemoryUtilization`. A 6 GB card with util 0.9 only has ~5.4 GB for weights.
108+
- Display GPU keeps 1.5 GB (others 0.5 GB) folded into `use_per_device` — ExLlamaV3 does not accept use and reserve together. Lower util or set `GpuSplitGb` (remapped order, e.g. `10,4.5`).
109+
- Tensor / pipeline need ≥2 devices. Speculative, vision, and LoRA are rejected under those modes.
89110

90111
## Quick CLI checks
91112

‎packaging/Build-Installer.ps1‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -256,7 +256,7 @@ Run Uninstall.bat as Administrator.
256256
Setup-Exl3Python.bat - reinstall PyTorch into %ProgramData%\ExLlamaSharp\venv
257257
"@ | Set-Content -Path (Join-Path $Stage "README.txt") -Encoding UTF8
258258

259-
$version = "1.3.2.1"
259+
$version = "1.4.0-beta"
260260
$info = @{
261261
product = "ExLlamaSharp"
262262
version = $version

‎packaging/ExLlamaSharp.iss‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@
33
; & "${env:LocalAppData}\Programs\Inno Setup 6\ISCC.exe" packaging\ExLlamaSharp.iss
44

55
#define MyAppName "ExLlamaSharp"
6-
#define MyAppVersion "1.3.2.1"
6+
#define MyAppVersion "1.4.0-beta"
77
#define MyAppPublisher "ExLlamaSharp"
88
#define MyAppURL "http://127.0.0.1:14563"
99
; Stage folder produced by Build-Installer.ps1 (relative to this .iss)
@@ -13,8 +13,8 @@
1313
AppId={{8F3E2A91-6C4B-4D7E-9A12-E5B8C0D4F617}
1414
AppName={#MyAppName}
1515
AppVersion={#MyAppVersion}
16-
VersionInfoVersion=1.3.2.1
17-
VersionInfoProductVersion=1.3.2.1
16+
VersionInfoVersion=1.4.0.0
17+
VersionInfoProductVersion=1.4.0.0
1818
AppMutex=Global\ExLlamaSharp.Server
1919
AppPublisher={#MyAppPublisher}
2020
AppPublisherURL={#MyAppURL}

‎packaging/Redeploy-Local.ps1‎

Lines changed: 6 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -78,12 +78,14 @@ $elevLines = @(
7878
'if ($LASTEXITCODE -ge 8) { throw "robocopy failed $LASTEXITCODE" }'
7979
'$tray = Join-Path "' + $PubTray + '" "ExLlamaSharp.Tray.exe"'
8080
'if (Test-Path $tray) { Copy-Item $tray (Join-Path $dest "ExLlamaSharp.Tray.exe") -Force; Log "copied Tray.exe" }'
81-
'$workerSrc = Join-Path "' + $Root + '" "tools\exl3_worker\worker.py"'
81+
'$workerSrcDir = Join-Path "' + $Root + '" "tools\exl3_worker"'
8282
'$workerDstDir = Join-Path $dest "tools\exl3_worker"'
83-
'if (Test-Path $workerSrc) {'
83+
'if (Test-Path $workerSrcDir) {'
8484
' New-Item -ItemType Directory -Force -Path $workerDstDir | Out-Null'
85-
' Copy-Item $workerSrc (Join-Path $workerDstDir "worker.py") -Force'
86-
' Log "copied worker.py"'
85+
' Copy-Item (Join-Path $workerSrcDir "worker.py") (Join-Path $workerDstDir "worker.py") -Force'
86+
' $launch = Join-Path $workerSrcDir "launch.py"'
87+
' if (Test-Path $launch) { Copy-Item $launch (Join-Path $workerDstDir "launch.py") -Force }'
88+
' Log "copied exl3_worker scripts"'
8789
'}'
8890
'$data = "' + $DataRoot + '"'
8991
'New-Item -ItemType Directory -Force -Path $data | Out-Null'

‎src/ExLlamaSharp.Server/Components/Pages/About.razor‎

Lines changed: 19 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -40,9 +40,25 @@ else
4040
@if (_info.Gpu.Available)
4141
{
4242
<ul class="mb-0">
43-
<li>@_info.Gpu.Name</li>
44-
<li>VRAM: @(_info.Gpu.VramTotalMb?.ToString("0") ?? "?") MB</li>
45-
<li>Compute: @(_info.Gpu.ComputeCapability ?? "—")</li>
43+
@if (_info.Gpu.Devices.Count > 0)
44+
{
45+
@foreach (var g in _info.Gpu.Devices)
46+
{
47+
<li>
48+
PCI @g.Index
49+
@(g.CudaIndex is int cuda ? $" — cuda:{cuda}" : " — not in CUDA_VISIBLE_DEVICES")
50+
— @g.Name — @g.VramUsedMb.ToString("0") / @g.VramTotalMb.ToString("0") MB
51+
@(string.IsNullOrEmpty(g.Uuid) ? "" : $" — {g.Uuid}")
52+
@(g.DisplayActive == true ? " — display" : "")
53+
</li>
54+
}
55+
}
56+
else
57+
{
58+
<li>@_info.Gpu.Name</li>
59+
<li>VRAM: @(_info.Gpu.VramTotalMb?.ToString("0") ?? "?") MB</li>
60+
}
61+
<li>Primary (highest VRAM): @_info.Gpu.Name</li>
4662
</ul>
4763
}
4864
else

0 commit comments

Comments
 (0)