This guide provides detailed instructions for installing Red Hat OpenShift AI (RHOAI) on a Red Hat OpenShift Service on AWS (ROSA) cluster.
The RHOAI version is determined by the CatalogSource image tag used in Step 5. The image follows the pattern:
quay.io/rhoai/rhoai-fbc-fragment:rhoai-<VERSION>
For example:
quay.io/rhoai/rhoai-fbc-fragment:rhoai-3.4for RHOAI 3.4quay.io/rhoai/rhoai-fbc-fragment:rhoai-3.5for RHOAI 3.5
Replace the image tag in Step 5 with the version you want to install. The rest of the installation steps remain the same regardless of version.
Note: The examples in this guide use RHOAI 3.4 as a reference.
- Choosing a Version
- Prerequisites
- Architecture Overview
- Installation Steps
- Step 1: Verify Cluster Access
- Step 2: Create Pull Secret for Red Hat Registries
- Step 3: Install Kyverno Policy Engine
- Step 4: Configure Kyverno Policies
- Step 5: Create RHOAI CatalogSource
- Step 6: Install RHOAI Operator
- Step 7: Create DataScienceCluster
- Step 8: Configure Components
- Step 9: Wait for DataScienceCluster Readiness
- Step 10: Enable Dashboard Features
- Verification
- Troubleshooting
- Resource Requirements
- Post-Installation
ocCLI (OpenShift command-line tool)jq(JSON processor)- Cluster admin access to a ROSA cluster
-
Minimum Resources:
- 2 worker nodes with 4 vCPU and 16 GB RAM each
- For full installation including dashboard: 3+ worker nodes with 8 vCPU each recommended
-
OpenShift Version: 4.14+ recommended
- Container registry credentials at
~/.config/containers/auth.jsonor~/.docker/config.json- Must include credentials for:
registry.redhat.io(required for RHOAI operator bundle images)quay.io(required for RHOAI catalog and some component images)brew.registry.redhat.io(optional, only needed for internal/development builds)
- Must include credentials for:
RHOAI on ROSA requires special handling for image pull secrets because ROSA HCP's platform-managed pull secret gets periodically reconciled (reverted). The solution uses Kyverno to:
- Sync brew registry credentials across all namespaces
- Automatically inject imagePullSecrets into pods cluster-wide
Key Components:
- Kyverno: Policy engine for secret distribution
- RHOAI Operator: Manages DataScienceCluster lifecycle
- CatalogSource: Provides RHOAI operator packages
- DataScienceCluster: Custom resource defining enabled components
First, verify you have cluster admin access and check for existing installations.
# Verify cluster connectivity
oc whoami
oc cluster-info
# Check platform type
oc get infrastructure cluster -o json | jq -r '.status.platformStatus.type'
# Verify no existing RHOAI/ODH installations
oc get csv -A | grep -E 'opendatahub|rhods' || echo "No existing installations found"Expected Output: Should confirm you're logged in as cluster-admin and no existing RHOAI installations are present.
Create a pull secret in the openshift-config namespace using your container registry credentials, typically from ~/.config/containers/auth.json (Podman default) or ~/.docker/config.json (Docker default).
# Create the pull secret from your container registry credentials
# Use ~/.config/containers/auth.json OR ~/.docker/config.json depending on where your credentials are stored
oc create secret generic pull-secret-redhat \
--from-file=.dockerconfigjson=~/.config/containers/auth.json \
--type=kubernetes.io/dockerconfigjson \
-n openshift-configVerification:
oc get secret pull-secret-redhat -n openshift-configInstall the latest version of Kyverno using the official installation manifest.
# Install Kyverno
oc apply -f https://github.com/kyverno/kyverno/releases/latest/download/install.yamlNote: You may see errors about CRDs with annotations being too long. These can be safely ignored as we'll create minimal CRDs next.
If the ClusterPolicy CRDs fail to install, create minimal versions:
# Create ClusterPolicy CRD
cat << 'EOF' | oc apply -f -
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: clusterpolicies.kyverno.io
labels:
app.kubernetes.io/part-of: kyverno
spec:
group: kyverno.io
names:
kind: ClusterPolicy
listKind: ClusterPolicyList
plural: clusterpolicies
shortNames:
- cpol
singular: clusterpolicy
categories:
- kyverno
scope: Cluster
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
x-kubernetes-preserve-unknown-fields: true
subresources:
status: {}
EOF
# Create Policy CRD
cat << 'EOF' | oc apply -f -
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: policies.kyverno.io
labels:
app.kubernetes.io/part-of: kyverno
spec:
group: kyverno.io
names:
kind: Policy
listKind: PolicyList
plural: policies
shortNames:
- pol
singular: policy
categories:
- kyverno
scope: Namespaced
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
x-kubernetes-preserve-unknown-fields: true
subresources:
status: {}
EOFOpenShift requires explicit SCC permissions for pods to run as specific UIDs. Kyverno containers run as UID 65534 (nobody), which is outside the default restricted range. Grant the nonroot-v2 SCC to all Kyverno service accounts:
oc adm policy add-scc-to-user nonroot-v2 -z kyverno-admission-controller -n kyverno
oc adm policy add-scc-to-user nonroot-v2 -z kyverno-background-controller -n kyverno
oc adm policy add-scc-to-user nonroot-v2 -z kyverno-cleanup-controller -n kyverno
oc adm policy add-scc-to-user nonroot-v2 -z kyverno-reports-controller -n kyvernoAfter granting SCCs (or if the admission controller is in CrashLoopBackOff), restart all deployments:
# Restart all Kyverno deployments
oc rollout restart deployment -n kyverno
# Wait for pods to be ready
sleep 15
oc wait --for=condition=ready pod -l app.kubernetes.io/instance=kyverno -n kyverno --timeout=90sVerification:
# All Kyverno pods should be Running
oc get pods -n kyverno
# Expected output:
# NAME READY STATUS RESTARTS AGE
# kyverno-admission-controller-xxxxx 1/1 Running 0 2m
# kyverno-background-controller-xxxxx 1/1 Running 0 2m
# kyverno-cleanup-controller-xxxxx 1/1 Running 0 2m
# kyverno-reports-controller-xxxxx 1/1 Running 0 2mCreate RBAC permissions and policies for Kyverno to manage secrets.
cat << 'EOF' | oc apply -f -
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: kyverno:generate-secrets
rules:
- apiGroups: [""]
resources: ["secrets"]
verbs: ["get", "list", "create", "update", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: kyverno:generate-secrets
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: kyverno:generate-secrets
subjects:
- kind: ServiceAccount
name: kyverno-admission-controller
namespace: kyverno
- kind: ServiceAccount
name: kyverno-background-controller
namespace: kyverno
EOFThis policy clones the brew secret to all namespaces:
cat << 'EOF' | oc apply -f -
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: sync-brew-secret
spec:
background: true
rules:
- name: sync-secret-to-namespaces
match:
any:
- resources:
kinds:
- Namespace
generate:
apiVersion: v1
kind: Secret
name: pull-secret-redhat
namespace: "{{request.object.metadata.name}}"
synchronize: true
clone:
namespace: openshift-config
name: pull-secret-redhat
EOFThis policy adds the pull secret to all pods:
cat << 'EOF' | oc apply -f -
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
name: add-brew-imagepullsecret
spec:
background: false
rules:
- name: add-imagepullsecret
match:
any:
- resources:
kinds:
- Pod
mutate:
patchStrategicMerge:
spec:
imagePullSecrets:
- name: pull-secret-redhat
EOFVerification:
# Check policies are created
oc get clusterpolicy
# Expected output:
# NAME AGE
# add-brew-imagepullsecret 1m
# sync-brew-secret 1mCreate the namespace and catalog source for RHOAI operators.
# Create the operator namespace
oc create namespace redhat-ods-operator
# Create the CatalogSource
cat << 'EOF' | oc apply -f -
apiVersion: operators.coreos.com/v1alpha1
kind: CatalogSource
metadata:
name: rhoai-catalog
namespace: openshift-marketplace
spec:
sourceType: grpc
image: quay.io/rhoai/rhoai-fbc-fragment:rhoai-3.4
displayName: Red Hat OpenShift AI 3.4
publisher: Red Hat
updateStrategy:
registryPoll:
interval: 10m
secrets:
- pull-secret-redhat
EOFIMPORTANT: The Kyverno sync policy only triggers when new namespaces are created. Since openshift-marketplace already exists, you must manually copy the secret:
oc get secret pull-secret-redhat -n openshift-config -o yaml | \
sed 's/namespace: openshift-config/namespace: openshift-marketplace/' | \
oc apply -f -
# Verify the secret contains credentials for registry.redhat.io (required for operator bundle)
oc get secret pull-secret-redhat -n openshift-marketplace -o jsonpath='{.data.\.dockerconfigjson}' | \
base64 -d | jq -r '.auths | keys[]'
# Expected output should include at least:
# - registry.redhat.io (required for RHOAI operator bundle)
# - quay.io (required for RHOAI catalog)# Wait up to 2 minutes for the catalog to become ready
oc wait --for=jsonpath='{.status.connectionState.lastObservedState}'=READY \
catalogsource/rhoai-catalog \
-n openshift-marketplace \
--timeout=120sVerification:
# Check catalog status
oc get catalogsource -n openshift-marketplace rhoai-catalog
# Check catalog pod is running
oc get pods -n openshift-marketplace | grep rhoai-catalog
# Verify RHOAI packages are available
oc get packagemanifest -n openshift-marketplace -o json | \
jq -r '.items[] | select(.status.catalogSource=="rhoai-catalog") | .metadata.name' | \
grep rhodsCreate the OperatorGroup and Subscription to install the RHOAI operator.
cat << 'EOF' | oc apply -f -
apiVersion: operators.coreos.com/v1
kind: OperatorGroup
metadata:
name: redhat-ods-operator-group
namespace: redhat-ods-operator
spec:
targetNamespaces: []
---
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: rhods-operator
namespace: redhat-ods-operator
spec:
channel: beta
installPlanApproval: Automatic
name: rhods-operator
source: rhoai-catalog
sourceNamespace: openshift-marketplace
EOFThe channel beta contains the early access for development/testing builds. To set it to a stable version, use channel: stable-3.4. To pin to a specific version within the channel, use startingCSV.
oc apply -f - <<'EOF'
apiVersion: operators.coreos.com/v1alpha1
kind: Subscription
metadata:
name: rhods-operator
namespace: redhat-ods-operator
spec:
channel: stable-3.4
installPlanApproval: Automatic
name: rhods-operator
source: rhoai-catalog
sourceNamespace: openshift-marketplace
startingCSV: rhods-operator.3.4.0
EOF# Wait for CSV to succeed (up to 10 minutes)
oc wait --for=jsonpath='{.status.phase}'=Succeeded \
csv -l operators.coreos.com/rhods-operator.redhat-ods-operator \
-n redhat-ods-operator \
--timeout=600sVerification:
# Check operator is installed
oc get csv -n redhat-ods-operator
# Expected output:
# NAME DISPLAY VERSION PHASE
# rhods-operator.3.4.0-ea.2 Red Hat OpenShift AI 3.4.0-ea.2 SucceededCreate the DataScienceCluster custom resource to deploy RHOAI components.
cat << 'EOF' | oc apply -f -
apiVersion: datasciencecluster.opendatahub.io/v2
kind: DataScienceCluster
metadata:
labels:
app.kubernetes.io/name: datasciencecluster
name: default-dsc
spec:
components:
aipipelines:
managementState: Managed
dashboard:
managementState: Managed
feastoperator:
managementState: Managed
kserve:
managementState: Managed
nim:
managementState: Managed
wva:
managementState: Removed
kueue:
managementState: Removed
llamastackoperator:
managementState: Managed
mlflowoperator:
managementState: Managed
modelregistry:
managementState: Managed
registriesNamespace: rhoai-model-registries
ray:
managementState: Managed
sparkoperator:
managementState: Removed
trainer:
managementState: Removed
trainingoperator:
managementState: Removed
trustyai:
managementState: Managed
workbenches:
managementState: Managed
EOFNote: This configuration enables most components. See Step 8 for resource-constrained configurations.
For resource-constrained clusters, disable non-essential components:
oc patch datasciencecluster default-dsc --type=merge -p '{
"spec": {
"components": {
"aipipelines": {
"managementState": "Removed"
},
"feastoperator": {
"managementState": "Removed"
},
"llamastackoperator": {
"managementState": "Removed"
},
"mlflowoperator": {
"managementState": "Removed"
},
"modelregistry": {
"managementState": "Removed"
},
"ray": {
"managementState": "Removed"
},
"trustyai": {
"managementState": "Removed"
}
}
}
}'This keeps only:
- Dashboard
- KServe (model serving)
- Workbenches (notebooks)
If resources are very limited, you can further reduce CPU/memory consumption by scaling down replicas and lowering container resource requests.
Scale the operator to a single replica:
oc scale deployment rhods-operator -n redhat-ods-operator --replicas=1Scale the dashboard to a single replica:
oc scale deployment rhods-dashboard -n redhat-ods-applications --replicas=1Lower dashboard container resource requests:
The dashboard deployment runs 9 containers with high default resource requests. On constrained clusters, reduce them:
oc patch deployment rhods-dashboard -n redhat-ods-applications --type='json' -p='[
{"op": "replace", "path": "/spec/template/spec/containers/0/resources/requests/cpu", "value": "100m"},
{"op": "replace", "path": "/spec/template/spec/containers/1/resources/requests/cpu", "value": "100m"},
{"op": "replace", "path": "/spec/template/spec/containers/2/resources/requests/cpu", "value": "10m"},
{"op": "replace", "path": "/spec/template/spec/containers/2/resources/requests/memory", "value": "64Mi"},
{"op": "replace", "path": "/spec/template/spec/containers/3/resources/requests/cpu", "value": "10m"},
{"op": "replace", "path": "/spec/template/spec/containers/3/resources/requests/memory", "value": "64Mi"},
{"op": "replace", "path": "/spec/template/spec/containers/4/resources/requests/cpu", "value": "10m"},
{"op": "replace", "path": "/spec/template/spec/containers/4/resources/requests/memory", "value": "64Mi"},
{"op": "replace", "path": "/spec/template/spec/containers/5/resources/requests/cpu", "value": "10m"},
{"op": "replace", "path": "/spec/template/spec/containers/5/resources/requests/memory", "value": "64Mi"},
{"op": "replace", "path": "/spec/template/spec/containers/6/resources/requests/cpu", "value": "10m"},
{"op": "replace", "path": "/spec/template/spec/containers/6/resources/requests/memory", "value": "64Mi"},
{"op": "replace", "path": "/spec/template/spec/containers/7/resources/requests/cpu", "value": "10m"},
{"op": "replace", "path": "/spec/template/spec/containers/7/resources/requests/memory", "value": "64Mi"},
{"op": "replace", "path": "/spec/template/spec/containers/8/resources/requests/cpu", "value": "10m"},
{"op": "replace", "path": "/spec/template/spec/containers/8/resources/requests/memory", "value": "64Mi"}
]'Verification:
# Check component status
oc get datasciencecluster default-dsc -o jsonpath='{.status.phase}'
# View deployed pods
oc get pods -n redhat-ods-applicationsAfter creating the DataScienceCluster, you can monitor its status until all components are deployed:
# Poll until the phase becomes "Ready" (check every 15 seconds, up to 10 minutes)
while true; do
phase=$(oc get datasciencecluster default-dsc -o jsonpath='{.status.phase}' 2>/dev/null || echo "Unknown")
echo "DataScienceCluster phase: $phase"
[ "$phase" = "Ready" ] && break
sleep 15
doneIf the cluster does not reach the Ready phase, inspect the failing conditions:
oc get datasciencecluster default-dsc -o jsonpath='{range .status.conditions[?(@.status=="False")]}{.type}: {.message}{"\n"}{end}'Enable advanced features in the RHOAI dashboard:
# Wait for OdhDashboardConfig to be created
sleep 30
# Enable feature flags
oc patch odhdashboardconfig odh-dashboard-config \
-n redhat-ods-applications \
--type=merge \
-p '{
"spec": {
"dashboardConfig": {
"disableModelMesh": false,
"disableKServe": false,
"disableProjects": false
},
"notebookController": {
"enabled": true
}
}
}'Note: If the OdhDashboardConfig doesn't exist yet, this step can be done after the dashboard is running.
# Check DataScienceCluster
oc get datasciencecluster default-dsc
# Check operator
oc get csv -n redhat-ods-operator
# Check all components
oc get pods -n redhat-ods-applications
# Check Kyverno policies
oc get clusterpolicyIf the cluster has sufficient resources and the dashboard is running:
# Get the dashboard route
oc get route rhods-dashboard -n redhat-ods-applications
# Example output:
# NAME HOST/PORT
# rhods-dashboard rhods-dashboard-redhat-ods-applications.apps.<cluster-domain>Access the dashboard URL in your browser.
Symptom: Catalog pod shows ImagePullBackOff
Solution:
# Verify secret exists in openshift-marketplace
oc get secret pull-secret-redhat -n openshift-marketplace
# If missing, manually copy it
oc get secret pull-secret-redhat -n openshift-config -o yaml | \
sed 's/namespace: openshift-config/namespace: openshift-marketplace/' | \
oc apply -f -
# Restart catalog pods
oc delete pod -n openshift-marketplace -l olm.catalogSource=rhoai-catalogSymptom: Kyverno admission controller fails to start with error about missing CRDs
Solution:
# Create minimal CRDs (see Step 3)
# Then delete the pod to restart
oc delete pod -n kyverno -l app.kubernetes.io/component=admission-controllerSymptom: Dashboard pods stuck in Pending state with "Insufficient CPU" or "Insufficient memory"
Root Cause: Cluster has insufficient resources
Solutions:
-
Scale the cluster (recommended):
- Add more worker nodes
- Upgrade to larger instance types
-
Reduce dashboard replicas:
oc scale deployment rhods-dashboard -n redhat-ods-applications --replicas=1
-
Disable dashboard entirely:
oc patch datasciencecluster default-dsc --type=merge -p '{ "spec": { "components": { "dashboard": { "managementState": "Removed" } } } }'
# View detailed status
oc get datasciencecluster default-dsc -o yaml | grep -A 5 "conditions:"
# Check specific component status
oc get datasciencecluster default-dsc -o jsonpath='{.status.conditions[?(@.type=="KserveReady")]}'- Nodes: 2 worker nodes
- CPU: 4 vCPU per node (3.5 allocatable)
- Memory: 16 GB per node
- Components: KServe, Workbenches, Model Controller
Resource Usage:
- CPU: ~95% utilization
- Memory: ~70% utilization
- Dashboard: Cannot run (insufficient resources)
- Nodes: 3+ worker nodes
- CPU: 8 vCPU per node
- Memory: 32 GB per node
- Components: All enabled (Dashboard, KServe, Workbenches, Ray, MLflow, etc.)
Resource Breakdown:
- Dashboard: ~4 CPU cores (9 containers)
- KServe: ~500m CPU
- Service Mesh: ~500m CPU
- Auth Proxy: ~500m CPU
- Other RHOAI components: ~2 CPU cores
- Platform services: ~2 CPU cores
# List available notebook images
oc get imagestreams -n redhat-ods-applications# Create a namespace for your project
oc create namespace my-datascience-project
# Label it as a data science project
oc label namespace my-datascience-project \
opendatahub.io/dashboard=true \
modelmesh-enabled=true# Example: Deploy a simple model
cat << 'EOF' | oc apply -f -
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: example-model
namespace: my-datascience-project
spec:
predictor:
model:
modelFormat:
name: sklearn
storageUri: s3://my-bucket/my-model
EOF# Watch all RHOAI pods
watch oc get pods -n redhat-ods-applications
# Monitor DataScienceCluster events
oc get events -n redhat-ods-applications --sort-by='.lastTimestamp'
# Check operator logs
oc logs -n redhat-ods-operator deployment/rhods-operatorTo upgrade to a newer version:
# Update the subscription channel
oc patch subscription rhods-operator -n redhat-ods-operator \
--type=merge \
-p '{"spec":{"channel":"stable-3.x"}}'This guide walked through installing RHOAI on ROSA using:
- Kyverno for automatic pull secret distribution
- Custom CatalogSource pointing to the desired RHOAI version (see Choosing a Version)
- DataScienceCluster for component configuration
Key considerations:
- ROSA HCP requires Kyverno for persistent pull secret management
- Resource requirements vary based on enabled components
- Dashboard requires significant CPU resources (~4 cores)
- Core functionality available without dashboard via CLI/API
For production deployments, ensure adequate cluster resources and consider enabling additional components based on your use case.