Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
112 changes: 66 additions & 46 deletions docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,17 +5,68 @@

# Gthulhu

## Protect critical workloads from CPU contention

Gthulhu uses eBPF and Linux `sched_ext` to make workload scheduling observable and controllable across Kubernetes nodes.

**Measured results under CPU contention:**

| Workload | Baseline | With Gthulhu | Result |
|---|---:|---:|---:|
| free5GC / GTP data path — average UE ping latency | **88.98 ms** | **2.079 ms** | **97.66% lower** |
| free5GC / GTP data path — maximum UE ping latency | **130.95 ms** | **8.45 ms** | **93.55% lower** |
| vLLM Qwen2.5-0.5B decode (`tg128`) under CPU pressure | **~6.7 t/s** | **~21.3 t/s** with tiered policy | **~3.2× throughput** |

The free5GC numbers are published in the free5GC community blog. The vLLM result comes from a reproducible community benchmark currently proposed to the vLLM project blog; treat it as an experiment under upstream review rather than a project-wide guarantee.

[Read the free5GC case study](https://free5gc.org/blog/20251126/20251126/){: .md-button .md-button--primary }
[vLLM benchmark PR](https://github.com/vllm-project/vllm-project.github.io/pull/300){: .md-button }
[Get Started](k8s.md){: .md-button }

> **DRA chooses what and where; Gthulhu controls how it actually runs.**

Gthulhu is a cloud-native runtime scheduling platform that connects Kubernetes workload intent to Linux task scheduling with eBPF and `sched_ext`.
## Why this matters

Kubernetes can admit a workload, place it on a node, and allocate devices/topology. Gthulhu focuses on the execution gap that follows: identify the Linux tasks that belong to the workload, apply bounded CPU scheduling policy, and verify whether the allocation is actually delivering the workload SLO.
A workload can already own a GPU, NIC, CPU set, or Kubernetes placement and still miss its latency or throughput target because its host-side Linux tasks are delayed by CPU contention.

[Get Started](k8s.md){: .md-button .md-button--primary }
[How It Works](how-it-works.md){: .md-button }
[Claim2Core Roadmap](claim2core.md){: .md-button }
Gthulhu focuses on that execution gap:

```text
Kubernetes admission / placement / allocation
Gthulhu runtime plane
Pod / cgroup / TGID / TID resolution
sched_ext + eBPF
latency / throughput / jitter / SLO
```

## Proven use cases

### 5G user-plane latency

The free5GC community published a GTP-driven scheduling experiment that combines `gtp5g-tracer`, a userspace operator, and Gthulhu. Under the same CPU stress, average UE ping latency dropped from **88.98 ms to 2.079 ms**, while maximum latency dropped from **130.95 ms to 8.45 ms**.

[Read: Implementing GTP-driven Automatic Scheduling Optimization with eBPF-based Scheduler](https://free5gc.org/blog/20251126/20251126/)

A separate earlier free5GC case study also documents using Gthulhu to reduce RTT by combining application/domain knowledge with custom `sched_ext` policy.

[Read: Improving Network Performance with Custom eBPF-based Schedulers](https://free5gc.org/blog/20250726/index.en/)

### vLLM inference under CPU pressure

A reproducible DGX Spark / GB10 experiment uses MicroK8s, vLLM, `stress-ng`, and Gthulhu to isolate the effect of CPU scheduling on GPU inference. In the submitted benchmark, decode throughput under CPU pressure is around **6–7 t/s** with the default scheduler and reaches roughly **21 t/s** on `tg128` with Gthulhu plus tiered policies targeting GPU-related work and vLLM's `EngineCore` thread.

## Current Capabilities
This result is linked here as an **upstream-reviewing community benchmark**, not as a generalized performance guarantee.

[Review the benchmark and methodology in vLLM blog PR #300](https://github.com/vllm-project/vllm-project.github.io/pull/300)

## What Gthulhu provides today

- **Pod-level scheduling observability** with eBPF.
- **Prometheus / Grafana / KEDA integration** for scheduler-aware operations and scaling.
Expand All @@ -24,7 +75,10 @@ Kubernetes can admit a workload, place it on a node, and allocate devices/topolo
- **TID-aware node-policy matching** so non-leader worker threads can be targeted directly.
- **Explicit priority semantics** across user-space and kernel scheduler modes.

## Claim2Core Direction
[How It Works](how-it-works.md){: .md-button }
[Claim2Core Roadmap](claim2core.md){: .md-button }

## Claim2Core: from allocation to delivered performance

The next architecture step is to connect actual Kubernetes allocation to runtime task scheduling:

Expand Down Expand Up @@ -53,18 +107,7 @@ Gthulhu should not reimplement kube-scheduler, DRA, or Kueue. It should consume

Read [Claim2Core](claim2core.md) for the implementation phases and safety boundaries.

## Why This Matters

Allocated resources do not automatically become delivered performance. A workload may own a GPU or NIC while its host-side feeder, tokenizer, NCCL/RDMA progress, DPDK, or packet-processing threads are still delayed by CPU contention.

Gthulhu makes this gap observable and controllable.

Immediate validation paths include:

- **CPU DRA × Gthulhu × free5GC/UPF** for p99/p99.9 latency and jitter;
- **GPU + RDMA + CPU DRA × phase-aware LLM scheduling** for TTFT, ITL, GPU idle, and communication progress.

## Architecture at a Glance
## Architecture at a glance

```text
User / Web UI / CRD
Expand All @@ -89,34 +132,11 @@ Decision Maker DaemonSet
Linux scheduler
```

## Recent Scheduler Semantics

Node-policy matching is thread-aware. The Decision Maker scans `/proc/<tgid>/task/<tid>` so a policy can target a named non-leader worker thread rather than only the process leader.

Priority handling is also explicit:

- `Priority > 0` means boost;
- `Priority == 0` is non-boosting;
- user-space mode supports TID-first lookup with TGID fallback;
- kernel mode does not insert non-boosting strategies into the priority BPF map, avoiding accidental priority-0 promotion.

See [How It Works](how-it-works.md) for the details and current limitations.

## Demo

<iframe width="560" height="315" src="https://www.youtube.com/embed/Cyjrh9cW1a8?si=0TL20Cd084wEoEVv" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>

## Next Steps
## Get involved

- [Deploy Gthulhu with Kubernetes](k8s.md)
- [Understand the architecture and current scheduler semantics](how-it-works.md)
- [Understand the architecture and scheduler semantics](how-it-works.md)
- [Read the Claim2Core roadmap](claim2core.md)
- [Configure pod scheduling metrics](pod-metrics.md)
- [Contribute](contributing.md)

## Community

- **GitHub**: [Gthulhu/Gthulhu](https://github.com/Gthulhu/Gthulhu)
- **Roadmap**: [Issue #141](https://github.com/Gthulhu/Gthulhu/issues/141)
- **Framework**: [Gthulhu/qumun](https://github.com/Gthulhu/qumun)
- **License**: Apache License 2.0
- [GitHub repository](https://github.com/Gthulhu/Gthulhu)
- [Roadmap issue #141](https://github.com/Gthulhu/Gthulhu/issues/141)
122 changes: 71 additions & 51 deletions docs/index.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,26 +5,80 @@

# Gthulhu

## 在 CPU contention 下保護關鍵工作負載

Gthulhu 使用 eBPF 與 Linux `sched_ext`,讓 Kubernetes 節點上的 workload scheduling 可以被觀測、控制與驗證。

**CPU contention 下的實測結果:**

| Workload | Baseline | 使用 Gthulhu | 結果 |
|---|---:|---:|---:|
| free5GC / GTP data path — UE 平均 ping latency | **88.98 ms** | **2.079 ms** | **降低 97.66%** |
| free5GC / GTP data path — UE 最大 ping latency | **130.95 ms** | **8.45 ms** | **降低 93.55%** |
| vLLM Qwen2.5-0.5B decode (`tg128`) under CPU pressure | **~6.7 t/s** | **~21.3 t/s**(tiered policy) | **約 3.2× throughput** |

free5GC 數據已發表於 free5GC 社群部落格。vLLM 數據來自可重現的 community benchmark,目前正在 vLLM 官方 blog PR 審核中,因此應視為 upstream-reviewing experiment,而不是適用所有環境的效能保證。

[閱讀 free5GC Case Study](https://free5gc.org/blog/20251126/20251126/){: .md-button .md-button--primary }
[vLLM Benchmark PR](https://github.com/vllm-project/vllm-project.github.io/pull/300){: .md-button }
[開始使用](k8s.md){: .md-button }

> **DRA chooses what and where; Gthulhu controls how it actually runs.**

Gthulhu 是一個雲原生 runtime scheduling 平台,透過 eBPF 與 Linux `sched_ext`,把 Kubernetes 工作負載意圖連接到真正的 Linux task scheduling。
## 為什麼這件事重要

Kubernetes 可以決定 Workload 是否能開始、放在哪個 Node,以及取得哪些裝置或 topology。Gthulhu 專注於 allocation 之後的 execution gap:找出真正屬於該工作負載的 Linux tasks、套用有邊界的 CPU 排程策略,並驗證 allocation 是否真的轉化成 workload SLO。
Workload 即使已經取得 GPU、NIC、CPU set 或 Kubernetes placement,也可能因 host-side Linux tasksCPU contention 下被延遲,而無法達成 latency / throughput SLO。

[開始使用](k8s.md){: .md-button .md-button--primary }
[了解運作原理](how-it-works.md){: .md-button }
[Claim2Core Roadmap](claim2core.md){: .md-button }
Gthulhu 專注於這個 execution gap:

```text
Kubernetes admission / placement / allocation
Gthulhu runtime plane
Pod / cgroup / TGID / TID resolution
sched_ext + eBPF
latency / throughput / jitter / SLO
```

## 已驗證的使用案例

### 5G User Plane latency

free5GC 社群已發表 GTP-driven scheduling 實驗,將 `gtp5g-tracer`、userspace operator 與 Gthulhu 串起來。在相同 CPU stress 下,UE 平均 ping latency 從 **88.98 ms 降到 2.079 ms**,最大 latency 從 **130.95 ms 降到 8.45 ms**。

[閱讀:Implementing GTP-driven Automatic Scheduling Optimization with eBPF-based Scheduler](https://free5gc.org/blog/20251126/20251126/)

另一篇較早的 free5GC case study 也展示如何利用 network-domain knowledge 搭配 Gthulhu 的 `sched_ext` policy 降低 RTT。

[閱讀:利用 Custom eBPF-based Schedulers 改善網路效能](https://free5gc.org/blog/20250726/)

### vLLM 在 CPU pressure 下的 inference

另一個可重現實驗使用 DGX Spark / GB10、MicroK8s、vLLM、`stress-ng` 與 Gthulhu,刻意隔離 CPU scheduling 對 GPU inference 的影響。在目前提交給 vLLM blog 的 benchmark 中,default scheduler 在 CPU pressure 下 decode throughput 約為 **6–7 t/s**;使用 Gthulhu 並對 GPU-related work 與 vLLM `EngineCore` 套用 tiered policy 後,`tg128` 約可達 **21 t/s**。

## 目前已具備的能力
這裡將它標示為 **upstream 審核中的 community benchmark**,而不是廣泛化的 performance guarantee。

- **Pod 層級排程可觀測性**:使用 eBPF 收集 scheduler signals。
- **Prometheus / Grafana / KEDA 整合**:讓操作與 autoscaling 能看見真實 scheduler pressure。
[查看 vLLM blog PR #300 的 benchmark 與方法](https://github.com/vllm-project/vllm-project.github.io/pull/300)

## Gthulhu 目前已具備的能力

- **Pod 層級 scheduling observability**:使用 eBPF 收集 scheduler signals。
- **Prometheus / Grafana / KEDA 整合**:支援 scheduler-aware operations 與 scaling。
- **分散式 scheduling intent**:Manager 搭配每節點 Decision Maker。
- **自訂 CPU scheduling**:Linux 6.12+ 可使用 `sched_ext`。
- **TID-aware node policy matching**:可以直接命中非 leader 的 worker thread。
- **明確的 priority semantics**:user-space 與 kernel mode 不再把 non-boosting rule 誤解成最高優先級。
- **TID-aware node policy matching**:可以直接命中非 leader worker thread。
- **明確的 priority semantics**:user-space 與 kernel mode 都有清楚的 boosting 行為。

[了解運作原理](how-it-works.md){: .md-button }
[Claim2Core Roadmap](claim2core.md){: .md-button }

## Claim2Core 方向
## Claim2Core:從 allocation 到 delivered performance

下一個架構階段,是把 Kubernetes 的實際 allocation 接到 runtime task scheduling:

Expand All @@ -49,21 +103,10 @@ Delivered workload SLO
- `ResourceSlice` 是 **inventory**。
- `ResourceClaim.status.allocation` 才是 workload 的 **actual allocation**。

Gthulhu 不應重新實作 kube-scheduler、DRA 或 Kueue,而是消費它們的決策,並且只在 Kubernetes/cgroup 已允許的 resource envelope 內控制 CPU execution
Gthulhu 不重新實作 kube-scheduler、DRA 或 Kueue,而是在 Kubernetes/cgroup 建立的 resource envelope 內,控制 workload 真正取得 CPU service 的方式

完整 implementation phases 與安全邊界請看 [Claim2Core](claim2core.md)。

## 為什麼重要

Allocated resource 不等於 delivered performance。Workload 即使已拿到 GPU 或 NIC,host-side 的 feeder、tokenizer、NCCL/RDMA progress、DPDK 或 packet-processing threads 仍可能因 CPU contention 被延遲。

Gthulhu 的價值就是讓這個落差可以被觀測、控制與驗證。

近期最值得驗證的兩條路線:

- **CPU DRA × Gthulhu × free5GC/UPF**:觀察 p99/p99.9 latency 與 jitter;
- **GPU + RDMA + CPU DRA × phase-aware LLM scheduling**:觀察 TTFT、ITL、GPU idle 與 communication progress。

## 架構一覽

```text
Expand All @@ -75,7 +118,7 @@ Manager API ───────▶ MongoDB / Kubernetes API
Decision Maker DaemonSet
├── eBPF 排程指標收集器 ──▶ Prometheus / Grafana / KEDA
├── eBPF scheduling metrics collector ──▶ Prometheus / Grafana / KEDA
└── task resolution / scheduling intent
Expand All @@ -89,34 +132,11 @@ Decision Maker DaemonSet
Linux scheduler
```

## 最近已合併的 scheduler semantics

Node policy 已經是 thread-aware。Decision Maker 會掃描 `/proc/<tgid>/task/<tid>`,因此策略可以直接命中具名的非 leader worker thread,而不是只能命中 process leader。

Priority semantics 也已明確化:

- `Priority > 0` 代表 boost;
- `Priority == 0` 是 non-boosting;
- user-space mode 使用 TID-first lookup,必要時 fallback 到 TGID;
- kernel mode 不會把 non-boosting strategy 塞進 priority BPF map,避免 priority 0 被誤當成最高 preemptive priority。

詳細行為與目前限制請看 [運作原理](how-it-works.md)。

## Demo

<iframe width="560" height="315" src="https://www.youtube.com/embed/Cyjrh9cW1a8?si=0TL20Cd084wEoEVv" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>

## 下一步
## 參與 Gthulhu

- [在 Kubernetes 部署 Gthulhu](k8s.md)
- [理解系統架構與目前 scheduler semantics](how-it-works.md)
- [理解系統架構與 scheduler semantics](how-it-works.md)
- [閱讀 Claim2Core roadmap](claim2core.md)
- [配置 Pod 排程指標](pod-metrics.md)
- [參與貢獻](contributing.md)

## 社群

- **GitHub**: [Gthulhu/Gthulhu](https://github.com/Gthulhu/Gthulhu)
- **Roadmap**: [Issue #141](https://github.com/Gthulhu/Gthulhu/issues/141)
- **Framework**: [Gthulhu/qumun](https://github.com/Gthulhu/qumun)
- **授權**: Apache License 2.0
- [GitHub repository](https://github.com/Gthulhu/Gthulhu)
- [Roadmap issue #141](https://github.com/Gthulhu/Gthulhu/issues/141)