Skip to main content

Deploy the platform

The Stacklok Enterprise platform runs in your Kubernetes cluster. You install it as a single umbrella Helm chart that deploys the ToolHive Operator, the Enterprise Manager, the Enterprise Cloud UI, and the Registry Server in one release, along with the custom resource definitions (CRDs) the operator needs.

Distributed and multi-cluster deployments

The umbrella chart can also run a subset of components, so you can spread the platform across clusters or maintain separate registries per environment. See Distributed deployments.

Air-gapped or strict-egress clusters

This page installs the chart directly from Replicated, which requires outbound access at install time. If your cluster can't reach Replicated, or your security posture requires every artifact to come from an internal registry, see Install from a private registry (air-gapped) instead.

Prerequisites

Before deploying, ensure you have:

  • A Kubernetes cluster (1.28 or later)
  • An ingress or gateway controller to publish the components, plus DNS records and TLS certificates for the hostnames you expose the Cloud UI, Enterprise Manager, and Registry Server on. The chart creates Services but no ingress, DNS, or certificates; see Step 6.
  • An OIDC-compatible identity provider configured per Configure platform identity
  • A PostgreSQL database for the Registry Server. The Registry Server serves the MCP and skills catalog the Enterprise Cloud UI reads, and stores it in an external PostgreSQL instance you provide.
  • Your Stacklok Enterprise license, available from the Stacklok install portal at install.stacklok.com. The license grants access to the umbrella chart and the container images it references. Stacklok sends portal access instructions during onboarding.

What the chart includes

The chart bundles each platform component and gives it an enable flag in a single values.yaml, so you turn on only the components you want. Each flag maps to a platform component:

Enable flagWhat it deploys
toolhiveOperatorThe ToolHive Operator and its custom resource definitions (MCPServer, VirtualMCPServer, and others)
enterpriseManagerEnterprise Manager service that serves configuration to Stacklok clients
cloudUiEnterprise Cloud UI Next.js application
registryServerRegistry Server that serves the MCP and skills catalog to the Enterprise Cloud UI, backed by an external PostgreSQL database.
aiGatewayAI Gateway operator that exposes large language models behind a managed gateway. Disabled by default.

The toolhiveOperator subchart deploys the enterprise build of the ToolHive Operator: the same operator codebase as ToolHive Community, repackaged as a hardened, signed, and versioned image. It manages the same workload types as the Community operator, so the existing guides apply unchanged. Use Run MCP servers in Kubernetes for MCP servers and remote proxies, and the Virtual MCP Server guides for Virtual MCP Server (vMCP) gateways.

Namespaces

This guide installs the platform into stacklok-system and uses that namespace throughout. You can choose a different one; adjust the commands to match.

The operator can manage MCP server workloads in any namespace. Running them in their own namespace, separate from the platform components in stacklok-system, keeps the platform and the workloads it manages apart, and is the recommended setup. The Kubernetes guides linked above use toolhive-system in their examples, so substitute your own namespaces as you follow them.

Run preflight checks

Before you authenticate and install, run the platform's preflight spec against your target cluster. The umbrella chart ships a Preflight custom resource as a labeled Secret, so rendering the chart and piping it to the kubectl-preflight plugin analyzes real prerequisites automatically. It catches misconfigurations before they surface later as failed pods or a stuck rollout: a Kubernetes version below the chart floor, tight node capacity, an unreachable OIDC issuer, a missing signing-key Secret.

Install the CLI plugin

Install the preflight plugin with krew, the kubectl plugin manager:

kubectl krew install preflight

Any recent release works; the chart's Preflight spec doesn't depend on a specific CLI version.

Alternative: install a specific release directly

If you don't use krew, or need a specific release (for example, in an air-gapped environment with no krew index access), download the binary from replicatedhq/troubleshoot releases instead:

PREFLIGHT_VERSION=v0.131.1
os=$(uname -s | tr '[:upper:]' '[:lower:]')
if [ "$os" = "darwin" ]; then
asset="preflight_darwin_all.tar.gz"
else
arch=$(uname -m); [ "$arch" = "x86_64" ] && arch=amd64
asset="preflight_${os}_${arch}.tar.gz"
fi
curl -fsSL -o preflight.tar.gz \
"https://github.com/replicatedhq/troubleshoot/releases/download/${PREFLIGHT_VERSION}/${asset}"
tar -xzf preflight.tar.gz preflight
sudo install -m 0755 preflight /usr/local/bin/kubectl-preflight

Check the releases page for the current version rather than assuming v0.131.1 above is still latest.

Run it against your real values

Use the values.yaml you build in Step 3, not an empty or minimal render. The opt-in analyzers (OIDC issuer, signing-key Secret, registry database) only render when their corresponding values are set, so preflighting against an empty file silently skips them, and a pass result verifies less than your real install actually needs.

helm template stacklok-enterprise \
oci://oci.stacklok.com/stacklok-enterprise/<CHANNEL>/stacklok-enterprise-platform \
--version <VERSION> \
--values values.yaml \
| kubectl preflight -

<CHANNEL> and <VERSION> are the channel slug and chart version the install portal at install.stacklok.com generates for your release, in its Existing cluster with Helm instructions. helm template prompts for the same registry credentials you use in Step 1: your license email as the username and your License ID as the password.

What it checks

CheckRenders whenOutcome
Kubernetes versionAlwaysfail below 1.28 (the chart's floor)
Node capacity (CPU and memory)Alwayswarn if tight, evaluated on the smallest node
DistributionAlwayswarn on OpenShift OCP
Egress reachabilitypreflight.checkEgress=truewarn only; off by default, always fails in air gap
Registry database reachabilityregistryServer.enabled and a run-time toolhive-registry-server.preflightDatabaseUri is suppliedfail if unreachable
OIDC issuer reachabilityenterpriseManager.enabled=true and enterprise-manager.idpConfig.issuer is setfail if unreachable or misconfigured
Signing-key Secret existsenterpriseManager.enabled=truefail if the Secret named in enterprise-manager.signingConfig.existingSecret is missing
Registry database credential

The registry database check needs a DB URI that includes a password. Supply it as a one-off flag at preflight time, never in a permanent or GitOps-tracked values file:

helm template stacklok-enterprise \
oci://oci.stacklok.com/stacklok-enterprise/<CHANNEL>/stacklok-enterprise-platform \
--version <VERSION> \
--set registryServer.enabled=true \
--set 'toolhive-registry-server.preflightDatabaseUri=postgres://<USER>:<PASSWORD>@<HOST>:5432/<DATABASE>?sslmode=require' \
| kubectl preflight -

The preflight spec ships as a Kubernetes Secret. A persistently configured URI leaves the database password in-cluster indefinitely, readable with kubectl get secret ... -o yaml, rather than existing only for the duration of one preflight run.

Preflight is advisory by default

On the Helm CLI install path, kubectl preflight can't block helm install. A fail result is a signal to act on, not an automatic gate: don't proceed to Step 4 while any check reports fail. Fix the underlying cluster condition first.

Optional: enforce preflight checks in-cluster

Set preflight.enforce: true in your values (or pass --set preflight.enforce=true) to opt into an in-cluster pre-install and pre-upgrade Helm hook Job that runs the same spec and genuinely fails helm install or helm upgrade on a fail outcome, real blocking, not advisory.

This is off by default because it carries cost you should accept deliberately:

  • The Job needs global.replicated.dockerconfigjson set. It pulls a license-gated preflight-runner image before the Replicated SDK provisions the ordinary pull secret; without it, the Job's pod can't pull its image and the install hangs.
  • The Job runs under its own ServiceAccount with cluster-scoped read RBAC (nodes, namespaces, storage classes, CRDs), plus a namespaced read on the signing-key Secret when enterpriseManager.enabled is set.

If enforcement blocks an install, read the Job's log to see why:

kubectl logs job/stacklok-enterprise-preflight-check -n stacklok-system

To bypass a known false positive for one run, re-run the same helm command with --set preflight.enforce=false.

What preflight can't check

Two prerequisites need manual verification; troubleshoot.sh has no analyzer for either:

  • Installer RBAC rights. No SelfSubjectAccessReview-style analyzer exists. Verify against the same kubeconfig context you install with:

    kubectl auth can-i create customresourcedefinitions.apiextensions.k8s.io
    kubectl auth can-i create clusterroles.rbac.authorization.k8s.io
    kubectl auth can-i create clusterrolebindings.rbac.authorization.k8s.io
    kubectl auth can-i '*' '*' -n stacklok-system

    Each must print yes. If any prints no, helm install fails partway through with a CRD, RBAC, or namespaced-resource creation error instead of upfront.

  • CRD collisions with a prior install. No CRD analyzer ships, since troubleshoot's customResourceDefinition analyzer can only express present as pass or absent as fail, not "already present is a problem." If you're reinstalling over a previous release, check kubectl get crd | grep toolhive.stacklok.dev for version conflicts by hand.

Deploy with Helm

1. Authenticate to the Replicated registry

Stacklok distributes the platform through Replicated. Your license, the umbrella chart, and per-release install instructions all live in the install portal at install.stacklok.com. Log in with the credentials Stacklok provides during onboarding.

The portal serves the chart from an OCI registry (oci.stacklok.com), not a classic Helm chart repository, so you authenticate with helm registry login rather than helm repo add. Use your license email as the username and your License ID as the password:

helm registry login oci.stacklok.com \
--username <YOUR_EMAIL> \
--password <LICENSE_ID>

In the portal, the Existing cluster with Helm instructions generate the exact login and install commands for your release, including your channel slug and the current chart version. Note those values; you reference them when you install the chart.

Image pulls during install

You don't create an image pull secret for this online install. The chart's Replicated integration provisions the pull credentials from your license, so the cluster pulls component images at install time. The air-gapped path differs: you mirror the images and create the pull secret yourself.

2. Prepare secrets

The chart expects a few Secrets to already exist. They live in the namespace the platform installs into, so create that namespace first:

kubectl create namespace stacklok-system

Then prepare the Secrets for the components you're enabling before you configure values.

Enterprise Manager signing key. Generate the key and create the Secret as described in Generate a signing key. The values file references it by name through enterprise-manager.signingConfig.existingSecret.

Better Auth session secret. The Cloud UI needs a secret of at least 32 characters to encrypt its sessions. Generate one:

openssl rand -base64 32

Use the output as toolhive-cloud-ui.betterAuth.secret in the values file.

Registry Server database passwords. Create a Secret for the database password (and a second one if you use a separate migration user), as described in Create the database credential Secrets. The values file references them through secretKeyRef, as the example below shows.

3. Configure values

Create a values.yaml file that enables the components you want and supplies their settings. Each component has an enable flag, and its configuration goes under that component's own key, as the example shows. The example below installs the operator, Enterprise Manager, Cloud UI, and Registry Server, with the AI Gateway disabled:

values.yaml
# Replicated SDK subchart. Required for the online install: it turns the license
# credentials injected at chart-pull time into the enterprise-pull-secret image
# pull secret the platform components reference.
replicated:
enabled: true

# Enable only the components you want.
toolhiveOperator:
enabled: true
enterpriseManager:
enabled: true
cloudUi:
enabled: true
registryServer:
enabled: true
aiGateway:
enabled: false

# Enterprise Manager configuration. See the Enterprise Manager deployment page
# for the full reference of fields under this key.
enterprise-manager:
idpConfig:
issuer: 'https://idp.example.com'
audience: 'enterprise-manager'
requiredScope: 'toolhive:config:read'
idpType: 'generic'
signingConfig:
# Secret created in "Prepare secrets" above
existingSecret: 'enterprise-manager-signing-key'
resourceURL: 'https://config.example.com'
clientID: '<STACKLOK_CLI_CLIENT_ID>'

# Registry Server configuration. Serves the MCP and skills catalog the Cloud
# UI reads. Requires an external PostgreSQL database; supply the password from
# a Secret, never inline.
toolhive-registry-server:
upstream:
config:
database:
host: 'postgres.example.com'
port: 5432
user: 'thv_user'
database: 'toolhive_registry'
sslMode: 'require'
extraEnv:
- name: THV_REGISTRY_DATABASE_PASSWORD
valueFrom:
secretKeyRef:
name: registry-db-credentials
key: password

# Enterprise Cloud UI configuration. See the Cloud UI deployment page for the
# full reference of fields under this key.
toolhive-cloud-ui:
# URL of the Registry Server. When it runs in this cluster, point at its
# in-cluster Service (named registry-api on port 8080).
apiBaseUrl: 'http://registry-api.stacklok-system.svc.cluster.local:8080'
# Cloud UI backend's URL for the Enterprise Manager; matches resourceURL
# above.
enterpriseManagerUrl: 'https://config.example.com'
oidc:
issuerUrl: 'https://idp.example.com'
clientId: '<CLOUD_UI_CLIENT_ID>'
clientSecret: '<CLOUD_UI_CLIENT_SECRET>'
betterAuth:
# Generated in "Prepare secrets" above (openssl rand -base64 32)
secret: '<BETTER_AUTH_SECRET>'
url: 'https://cloud-ui.example.com'

For the full reference of fields each component accepts, see Configure the Enterprise Manager, Configure the Enterprise Cloud UI, and Configure the Registry Server.

Global Redis/Valkey defaults

Several platform components can share an external Redis or Valkey instance for distributed session storage. Rather than repeating the configuration on every component, set it once under global.redis and let the components that support it inherit the value. Valkey is a drop-in replacement for Redis. This block is optional: when global.redis.host is empty, the global default is inactive and each component falls back to its own configuration.

values.yaml (global Redis/Valkey addition)
global:
redis:
# When empty, the global default is inactive.
host: 'redis.example.com'
port: 6379
# Enable the Redis Cluster protocol when connecting to the host.
clusterMode: false
# Reference a pre-created Secret; omit for a passwordless instance.
existingSecret: 'valkey-auth'
existingSecretKey: 'redis-password'
tls:
enabled: false

If passwordless, omit existingSecret and skip the Secret creation step. If your instance requires authentication, create the Secret before installing the chart:

kubectl create secret generic valkey-auth \
--namespace stacklok-system \
--from-literal=redis-password="<YOUR_REDIS_PASSWORD>"

The following components inherit global.redis:

Enable flagHow it uses the global default
toolhiveOperatorDefault session storage for MCPServer, MCPRemoteProxy, and VirtualMCPServer (vMCP) workloads that have no explicit spec.sessionStorage. The per-resource setting always overrides.
aiGatewayRate limiting and virtual API key storage, when the bundled Valkey instance is disabled.

Note that the embedded auth server's token storage is configured separately via MCPExternalAuthConfig. See Configure session storage in the operator guide.

note

The operator subchart also exposes toolhive-operator.upstream.operator.defaultRedis.addr for an operator-specific override. When set, it takes precedence over global.redis.host and global.redis.port; when empty, the operator falls back to the global block.

Global PostgreSQL defaults

The Registry Server reads its database host, port, and SSL mode from global.postgres when its own toolhive-registry-server.upstream.config.database.host is empty. This is the same pattern as global.redis above, intended for umbrella deployments that share a PostgreSQL instance across several backing services. The block is optional: when global.postgres.host is empty, the default is inactive and each subchart uses its own config.database block.

values.yaml (global PostgreSQL addition)
global:
postgres:
# When empty, the global default is inactive.
host: 'postgres.example.com'
port: 5432
sslMode: 'require'

Only host, port, and sslMode are inherited from the global block. The database user, database name, and credentials still go under the subchart's own config.database block (and extraEnv for the password Secret reference, as the example above shows). A locally set toolhive-registry-server.upstream.config.database.host always wins over global.postgres.host.

4. Install the chart

Install the chart into the stacklok-system namespace you created earlier. Reference the chart by its oci:// URL, using the channel slug and version from Step 1:

helm install stacklok-enterprise \
oci://oci.stacklok.com/stacklok-enterprise/<CHANNEL>/stacklok-enterprise-platform \
--version <VERSION> \
--namespace stacklok-system \
--values values.yaml

5. Verify the install

Wait for the platform pods to reach the Running state:

kubectl get pods -n stacklok-system

Confirm the ToolHive CRDs registered:

kubectl get crd | grep toolhive.stacklok.dev

6. Expose the platform endpoints

The chart creates ClusterIP Services for the components but no ingress. Publish these three through your ingress or gateway controller so browsers and clients outside the cluster can reach them at the hostnames you set in Step 3. List the Services to get their names (two are prefixed with your Helm release name, stacklok-enterprise):

kubectl get svc -n stacklok-system

Route each external hostname to its Service with the Ingress, HTTPRoute, or Gateway resources your controller uses:

ComponentService (port)Reached byHostname to route
Enterprise Cloud UIstacklok-enterprise-toolhive-cloud-ui (80)BrowsersbetterAuth.url
Enterprise Managerstacklok-enterprise-enterprise-manager (80)Stacklok Desktop and CLI clientsresourceURL
Registry Serverregistry-api (8080)Stacklok Desktop and CLI clientsthe registry's public API URL

The in-cluster URLs (apiBaseUrl, enterpriseManagerUrl) stay as Service DNS and need no routing. Once the routes resolve, confirm the Enterprise Manager answers at its external hostname:

curl -sf https://config.example.com/.well-known/toolhive-configuration | jq .

7. Prepare workload namespaces

If you run MCP server and vMCP workloads in a namespace other than stacklok-system (the recommended setup, see the Namespaces note above), copy the image pull secret into that namespace. The operator stamps the enterprise-pull-secret secret onto every workload pod it spawns, and the kubelet resolves it in the pod's own namespace. The Replicated integration creates it only in stacklok-system:

kubectl create namespace <WORKLOAD_NAMESPACE>

# Copy the pull secret from the platform namespace.
kubectl get secret enterprise-pull-secret -n stacklok-system \
-o jsonpath='{.data.\.dockerconfigjson}' | base64 -d \
| kubectl create secret generic enterprise-pull-secret \
--namespace <WORKLOAD_NAMESPACE> \
--type kubernetes.io/dockerconfigjson \
--from-file=.dockerconfigjson=/dev/stdin

Repeat for any namespace that hosts operator-managed workloads. If you run everything in stacklok-system, skip this step.

Next steps