Warning
Pre-alpha. OpenEverest v2 and this provider are under active development. CRD schemas, chart values and defaults change frequently, including in breaking ways, and there is no supported upgrade path between versions yet. Not for production use.
Serve LLMs on Kubernetes through OpenEverest, backed by KubeAI — vLLM on GPU, Ollama on CPU for local demos.
OpenEverest providers translate a single, technology-agnostic Instance custom resource into
the native custom resources of an upstream Kubernetes operator — for databases, but equally
for caches, message queues, object storage, or model-serving runtimes. This repository is the
provider for KubeAI: it owns the technology-specific knowledge — topologies, versions,
parameters — so that users, the API server, and the UI stay technology-agnostic.
Important
This provider is not standalone. It requires an OpenEverest installation (core CRDs and controller) in the cluster. Installing this chart on its own does nothing. See Install OpenEverest.
flowchart LR
U([User / API / UI]) -->|creates| I["Instance<br/>core.openeverest.io"]
I --> P["provider-kubeai<br/>(this repository)"]
P -->|reconciles into| O["Model<br/>kubeai.org/v1"]
O --> W["KubeAI"]
W --> R[("Workloads, Services,<br/>model cache")]
P -->|status, endpoints| I
The provider watches Instance resources whose spec.providerRef.name is provider-kubeai,
and reports workload health back onto Instance.status. It never manages pods directly — all
lifecycle work is delegated to KubeAI.
This provider has not been released yet — the table describes main.
| provider-kubeai | OpenEverest | KubeAI | Kubernetes |
|---|---|---|---|
main |
>= 2.0.0 |
latest release | 1.30 – 1.34 |
What you can do to a running instance through the Instance API. Upgrading the
provider itself is covered under Installation.
| Capability | Status | Notes |
|---|---|---|
| Provisioning | ✅ | |
| Horizontal scaling | ✅ | request-based autoscaling; minReplicas / maxReplicas on the topology |
| Vertical scaling (CPU / memory) | ✅ | through the component's resourceProfile (e.g. nvidia-gpu-l4:1, cpu:1) |
| Version upgrades | ✅ | of the deployed inference server version — change spec.version; see Versions |
| Custom configuration | ✅ | engine args and env on the server component |
| Monitoring | vLLM exposes Prometheus metrics; see Observability | |
| TLS | ❌ | not exposed through the Instance API |
Model artefacts are pulled from the URI given in model.source (hf://, s3://, pvc://,
ollama://), so this provider manages no persistent volumes and has no backup story.
Note
There is no published chart yet. Until the first release, install from a checkout.
git clone https://github.com/openeverest/provider-kubeai.git
cd provider-kubeai
helm install provider-kubeai charts/provider-kubeai --namespace everest-systemmake helm-install does the same thing against your current kube context.
-
KubeAI is not bundled. Install it before or alongside the provider, in the same namespace as your Instances:
helm repo add kubeai https://www.kubeai.org && helm repo update helm upgrade --install kubeai kubeai/kubeai -n default --wait --timeout 10mOn GPU clusters use the checked-in values: deploy/kubeai/values-gpu.yaml.
Uninstall:
helm uninstall provider-kubeai --namespace everest-systemUninstalling the chart does not delete running Instance resources.
Verify that the provider registered itself:
kubectl get providers.core.openeverest.io provider-kubeaiCreate an instance:
apiVersion: core.openeverest.io/v1alpha1
kind: Instance
metadata:
name: qwen2-05b-cpu
spec:
providerRef:
name: provider-kubeai
topology:
type: autoscaled
parameters:
minReplicas: 1
maxReplicas: 1
components:
server:
parameters:
model:
source: ollama://qwen2:0.5b
resourceProfile: cpu:1Component names are defined by this provider — see definition/provider.yaml.
spec.version and spec.topology are optional; the provider defaults apply.
More examples live in examples/ — instance-simple.yaml
runs on CPU, instance-gpu.yaml needs an NVIDIA cluster.
Watch it come up:
kubectl get instance qwen2-05b-cpu -w
kubectl get model # the KubeAI Model the provider created
kubectl get pods -l model=qwen2-05b-cpu # the serving pods KubeAI scheduledThen call it:
kubectl port-forward svc/kubeai 8000:80
curl -s http://127.0.0.1:8000/openai/v1/models | jq
curl http://127.0.0.1:8000/openai/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen2-05b-cpu","messages":[{"role":"user","content":"hi"}],"max_tokens":32}'The API is OpenAI-compatible and served under /openai/v1/... (not /v1/...). The model
field is the Instance name — for the GPU example that is llama-3-8b.
| Topology | Default | Description |
|---|---|---|
autoscaled |
✅ | Request-based autoscaling of the inference server, including scale-to-zero |
Topology parameters: minReplicas, maxReplicas, targetRequests, scaleDownDelaySeconds.
| Version bundle | Default | server |
|---|---|---|
vllm-0.11.2 |
✅ | 0.11.2 |
vllm-0.10.1 |
0.10.1 |
|
ollama-cpu |
0.11.11-ollama (Ollama, CPU only) |
Source of truth: definition/versions.yaml.
- Chart values: charts/provider-kubeai/values.yaml
- Instance parameters: per-component and per-topology
parametersschemas, defined under definition/ and published on theProviderresource (kubectl get provider provider-kubeai -o yaml). The API server and the UI validate user input against these schemas.
The technology-specific knobs worth knowing about, all on the server component:
| Parameter | Purpose |
|---|---|
model.source |
Where the model comes from (hf://, s3://, pvc://, ollama://) |
model.estimatedParamBillions, model.quantization |
Hints KubeAI uses for scheduling and profile selection |
resourceProfile |
KubeAI resource profile, e.g. cpu:1 or nvidia-gpu-l4:1 |
cacheProfile |
KubeAI cache profile for model artefacts |
args, env |
Extra engine arguments and environment variables |
vLLM exposes Prometheus metrics, and KubeAI ships a PodMonitor for them (enabled by
deploy/kubeai/values-gpu.yaml). On a GPU cluster, scrape
them with kube-prometheus-stack:
helm upgrade --install prometheus prometheus-community/kube-prometheus-stack \
-n monitoring --create-namespace \
-f deploy/observability/values-prometheus.yamlImport examples/observability/vllm-grafana-dashboard.json into Grafana for the request, latency and KV-cache panels. Full steps: docs/observability.md.
Requires Go (see go.mod), Docker, Helm, kubectl, and a Kubernetes cluster you can
reach. dev/README.md covers the environment end to end: the recommended
local k3d setup, running against a cluster you already have, and every dev/.env setting.
Step-by-step runbooks: k3d, kind,
GPU.
make dev-up # local cluster + Tilt dev environment (see dev/README.md)
make generate # RBAC, provider spec, Helm chart sync
make run # run the provider locally against the cluster
make test # unit tests
make test-integration # chainsaw suites
make dev-downmake help lists every target. make verify fails when generated files are stale — run
make generate and commit the result.
The provider contract (Validate / Sync / Status / Cleanup), RBAC markers, watches,
and code generation are documented once for all providers in
PROVIDER_DEVELOPMENT.md.
| Path | Purpose |
|---|---|
cmd/provider/ |
Entry point |
internal/provider/ |
ProviderInterface implementation, RBAC markers |
internal/common/ |
Component name constants |
definition/ |
Provider identity, component types, versions, topologies |
charts/provider-kubeai/ |
Helm chart (generated/ is produced by make generate) |
config/rbac/role.yaml |
Generated ClusterRole — do not edit |
deploy/ |
KubeAI and Prometheus values used by the runbooks |
docs/ |
k3d, kind, GPU and observability runbooks |
examples/ |
Example Instance resources and a Grafana dashboard |
dev/ |
Tilt dev environment, .env configuration, k3d cluster config |
.github/workflows/ |
CI: lint, build, unit and integration tests, release |
- Unit tests —
make test. - Integration tests —
make test-integrationruns the chainsaw suites. - CI — .github/workflows/ci.yaml runs lint, build, unit tests, generated-file verification, and Helm lint on every pull request.
kubectl logs -n everest-system deploy/provider-kubeai -f| Symptom | Where to look |
|---|---|
Instance stuck in Creating |
kubectl describe instance <name> conditions, then the provider logs |
No Provider resource in the cluster |
Is the chart installed? Check the provider deployment logs |
Instance ignored entirely |
spec.providerRef.name must be provider-kubeai |
Model created but no pods |
KubeAI must run in the same namespace as the Instance; check the KubeAI controller logs |
| Requests return 404 | Use /openai/v1/..., and make sure the model field matches the Instance name |
| CPU model takes minutes to answer | Ollama cold starts are slow; pin minReplicas: 1 to avoid scale-to-zero |
Issues and pull requests are welcome. See CONTRIBUTING.md, PROVIDER_DEVELOPMENT.md and the OpenEverest Code of Conduct.
Report vulnerabilities per the OpenEverest security policy. Please do not open public issues for security reports.
Apache License 2.0 — see LICENSE for details.