From 83923a588afb1b1a550c1621760e7f65eb4e34cc Mon Sep 17 00:00:00 2001 From: Rajiv Shah Date: Fri, 2 Oct 2026 12:26:42 -0500 Subject: [PATCH 1/7] Add Azure AKS Enterprise installation documentation --- docs.json | 3 +- enterprise/k8s-install/aks.mdx | 164 +++++++++++++++++++++++++++++++++ llms-full.txt | 162 ++++++++++++++++++++++++++++++++ llms.txt | 1 + 4 files changed, 329 insertions(+), 1 deletion(-) create mode 100644 enterprise/k8s-install/aks.mdx diff --git a/docs.json b/docs.json index 9babff37f..5b9959207 100644 --- a/docs.json +++ b/docs.json @@ -624,7 +624,8 @@ "enterprise/k8s-install/dns-and-tls", "enterprise/k8s-install/resource-limits", "enterprise/k8s-install/upgrade-guidance", - "enterprise/k8s-install/eks" + "enterprise/k8s-install/eks", + "enterprise/k8s-install/aks" ] } ] diff --git a/enterprise/k8s-install/aks.mdx b/enterprise/k8s-install/aks.mdx new file mode 100644 index 000000000..d897a443a --- /dev/null +++ b/enterprise/k8s-install/aks.mdx @@ -0,0 +1,164 @@ +--- +title: Azure AKS +description: Prepare an Azure Kubernetes Service cluster to run OpenHands Enterprise +icon: microsoft +--- + +Prepare Azure Kubernetes Service (AKS) for OpenHands Enterprise, then follow the +[Helm installation](/enterprise/k8s-install/installation) for application configuration. +This page covers Azure permissions, node pools, persistent storage, and ingress. + +## Cluster Requirements + +| Requirement | Recommendation | +| --- | --- | +| Access | An Azure subscription and permissions to create the resource group, cluster, node pools and networking | +| AKS version | A supported AKS version compatible with [Sysbox](/enterprise/k8s-install/sysbox); verify the actual Ubuntu image and containerd version | +| Sandbox OS | Ubuntu nodes with a working Sysbox runtime | +| Storage class | Azure Disk CSI, such as `managed-csi`, with expansion enabled | +| Capacity | Sufficient regional and VM-family vCPU quota for application nodes, sandbox nodes and upgrades | + +Entra or Azure DevOps authentication alone does not establish subscription access. +Check `az account list --all` and the selected subscription's permissions. +Contributor was sufficient for the evaluated default managed-network deployment; +Azure role assignments require additional permissions. Existing networks, private +cluster access and workload identities need their own access checks. + +Check VM SKU restrictions as well as quota before selecting a region or availability +zones. Use a dedicated resource group and explicit subscription and kubeconfig. + +## Node Pools + +Use separate pools for application services and sandbox workloads: + +- **General pool:** runs OpenHands services and cluster add-ons. Keep application + workloads here through node affinity or selectors. +- **Sysbox pool:** an Ubuntu user node pool for agent sandboxes. Configure + `sysbox-install=yes` as a persistent node-pool label so new nodes receive the installer. + +The evaluation used Azure CNI overlay with Calico, two platform nodes and a sandbox +pool with autoscaler bounds of one to two nodes. Size the pools using the +[Sizing Guide](/enterprise/sizing-guide) and +[Resource Limits](/enterprise/k8s-install/resource-limits), including node overhead, +image storage and warm sandbox capacity. Validate scale-down behavior for active sessions. + +### Set Up the Sandbox Runtime + +1. Follow [Installing Sysbox](/enterprise/k8s-install/sysbox) to install Sysbox on + your Ubuntu sandbox pool. +2. Before installing OpenHands, run a Sysbox test pod with an Azure Disk workspace + using the [AKS installation skill and companion checks](https://github.com/OpenHands/OpenHands-Cloud/tree/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/aks-install). + Confirm the pod starts and can write and read a file. +3. After installing OpenHands, start a conversation and ask it to run `pwd`. + Confirm it returns the workspace directory. + + +Some AKS node images need an additional Sysbox setup step. If the test pod fails +with an error mentioning `sysbox-runc`, contact OpenHands support before continuing. +The AKS installation skill includes the workaround used for the evaluation. + + +Before enabling sandbox-pool autoscaling, confirm that each newly created node can +start a sandbox. A setup correction applied to existing nodes must also be applied +to replacement and new nodes. Validate this with your support team before relying +on automatic scale-out or node upgrades. + + +The failure was observed on AKS 1.35.8, Ubuntu 24.04.5 LTS and containerd 2.3.3-2 +with Sysbox installer v0.7.1-0. The installer registered the runtime under the old +containerd plugin key, so its success message did not establish a working runtime. +See [upstream issue 997](https://github.com/nestybox/sysbox/issues/997). + +The evaluation workaround corrected only the Sysbox registration and restarted +containerd on each affected sandbox node. It requires a reviewed per-node operation +and does not automatically cover nodes added by autoscaling, reimaging or upgrades. +The [AKS install skill](https://github.com/OpenHands/OpenHands-Cloud/blob/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/.agents/skills/aks-install.md) +contains the diagnostic and correction procedure. Verify a real test pod afterward. + + +## Persistent Storage + +Inspect the Azure Disk CSI StorageClass: + +```bash +kubectl get storageclass managed-csi -o yaml +``` + +The evaluated class uses `disk.csi.azure.com`, `WaitForFirstConsumer`, +`ReadWriteOnce` and volume expansion. Configure stateful components and workspace +storage explicitly; the GKE `standard-rwo` and AWS gp3 examples do not apply. + +```yaml +runtime-api: + env: + STORAGE_CLASS: managed-csi +postgresql: + primary: + persistence: + storageClass: managed-csi +``` + +Account for node disk-attachment limits and availability-zone topology. Validate +mounting, persisted content after reattachment, and expansion on a disposable PVC. +The evaluation passed these checks, including expansion from 1 GiB to 2 GiB. +See [Azure Disk CSI provisioning](https://learn.microsoft.com/en-us/azure/aks/create-volume-azure-disk). + +## Object Storage + +Conversation/session storage is separate from workspace PVCs. Helm chart 0.74.0 +supports S3-compatible and GCS filestore configuration; it does not expose an Azure +Blob backend. Do not substitute Azure Blob credentials into the S3 settings. + +The evaluation used the chart's optional RustFS store on Azure Disk CSI. For production, +select a supported durable object store and validate backup, restore and availability. +When enabling automations, also provision the separate automation package bucket. + +## Database + +Use the [External PostgreSQL](/enterprise/external-postgres) guide for production +database requirements and Helm values. If choosing Azure Database for PostgreSQL, +verify compatibility, TLS and network access against those requirements before deployment. +The AKS evaluation used bundled PostgreSQL; a managed Azure database was not validated. + +## Ingress + +Run an ingress controller on the general pool. Traefik was validated with an Azure +public LoadBalancer: + +```yaml +service: + type: LoadBalancer +``` + +Read the Service's external IP and create a wildcard A record for your base domain. +DNS can remain with another provider. For private ingress, select the appropriate +Azure load-balancer configuration and verify client access separately. + +Follow [DNS and TLS](/enterprise/k8s-install/dns-and-tls) for trusted certificates +and hostname configuration. Use flat runtime hostnames with Traefik standard Ingress. +If provisioning certificates manually, assign renewal and Secret-update ownership. + +## Installation Skill + +For agent-assisted evaluation setup, the +[AKS installation skill](https://github.com/OpenHands/OpenHands-Cloud/blob/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/.agents/skills/aks-install.md) +and [companion files](https://github.com/OpenHands/OpenHands-Cloud/tree/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/aks-install) +provide Azure commands, values templates and storage/automation checks. +The skill is currently proposed in [PR #1339](https://github.com/OpenHands/OpenHands-Cloud/pull/1339). + +## Next Steps + + + + Configure and verify the sandbox runtime. + + + Configure hostnames and trusted certificates. + + + Deploy the application and validate a conversation. + + + Configure automation storage and verify a completed run. + + diff --git a/llms-full.txt b/llms-full.txt index e5697256d..c7ce514a9 100644 --- a/llms-full.txt +++ b/llms-full.txt @@ -52811,6 +52811,168 @@ If you can't use a wildcard certificate, obtain one with SANs for the `app`, `au +### Azure AKS +Source: https://docs.openhands.dev/enterprise/k8s-install/aks.md + +Prepare Azure Kubernetes Service (AKS) for OpenHands Enterprise, then follow the +[Helm installation](/enterprise/k8s-install/installation) for application configuration. +This page covers Azure permissions, node pools, persistent storage, and ingress. + +## Cluster Requirements + +| Requirement | Recommendation | +| --- | --- | +| Access | An Azure subscription and permissions to create the resource group, cluster, node pools and networking | +| AKS version | A supported AKS version compatible with [Sysbox](/enterprise/k8s-install/sysbox); verify the actual Ubuntu image and containerd version | +| Sandbox OS | Ubuntu nodes with a working Sysbox runtime | +| Storage class | Azure Disk CSI, such as `managed-csi`, with expansion enabled | +| Capacity | Sufficient regional and VM-family vCPU quota for application nodes, sandbox nodes and upgrades | + +Entra or Azure DevOps authentication alone does not establish subscription access. +Check `az account list --all` and the selected subscription's permissions. +Contributor was sufficient for the evaluated default managed-network deployment; +Azure role assignments require additional permissions. Existing networks, private +cluster access and workload identities need their own access checks. + +Check VM SKU restrictions as well as quota before selecting a region or availability +zones. Use a dedicated resource group and explicit subscription and kubeconfig. + +## Node Pools + +Use separate pools for application services and sandbox workloads: + +- **General pool:** runs OpenHands services and cluster add-ons. Keep application + workloads here through node affinity or selectors. +- **Sysbox pool:** an Ubuntu user node pool for agent sandboxes. Configure + `sysbox-install=yes` as a persistent node-pool label so new nodes receive the installer. + +The evaluation used Azure CNI overlay with Calico, two platform nodes and a sandbox +pool with autoscaler bounds of one to two nodes. Size the pools using the +[Sizing Guide](/enterprise/sizing-guide) and +[Resource Limits](/enterprise/k8s-install/resource-limits), including node overhead, +image storage and warm sandbox capacity. Validate scale-down behavior for active sessions. + +### Set Up the Sandbox Runtime + +1. Follow [Installing Sysbox](/enterprise/k8s-install/sysbox) to install Sysbox on + your Ubuntu sandbox pool. +2. Before installing OpenHands, run a Sysbox test pod with an Azure Disk workspace + using the [AKS installation skill and companion checks](https://github.com/OpenHands/OpenHands-Cloud/tree/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/aks-install). + Confirm the pod starts and can write and read a file. +3. After installing OpenHands, start a conversation and ask it to run `pwd`. + Confirm it returns the workspace directory. + + +Some AKS node images need an additional Sysbox setup step. If the test pod fails +with an error mentioning `sysbox-runc`, contact OpenHands support before continuing. +The AKS installation skill includes the workaround used for the evaluation. + + +Before enabling sandbox-pool autoscaling, confirm that each newly created node can +start a sandbox. A setup correction applied to existing nodes must also be applied +to replacement and new nodes. Validate this with your support team before relying +on automatic scale-out or node upgrades. + + +The failure was observed on AKS 1.35.8, Ubuntu 24.04.5 LTS and containerd 2.3.3-2 +with Sysbox installer v0.7.1-0. The installer registered the runtime under the old +containerd plugin key, so its success message did not establish a working runtime. +See [upstream issue 997](https://github.com/nestybox/sysbox/issues/997). + +The evaluation workaround corrected only the Sysbox registration and restarted +containerd on each affected sandbox node. It requires a reviewed per-node operation +and does not automatically cover nodes added by autoscaling, reimaging or upgrades. +The [AKS install skill](https://github.com/OpenHands/OpenHands-Cloud/blob/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/.agents/skills/aks-install.md) +contains the diagnostic and correction procedure. Verify a real test pod afterward. + + +## Persistent Storage + +Inspect the Azure Disk CSI StorageClass: + +```bash +kubectl get storageclass managed-csi -o yaml +``` + +The evaluated class uses `disk.csi.azure.com`, `WaitForFirstConsumer`, +`ReadWriteOnce` and volume expansion. Configure stateful components and workspace +storage explicitly; the GKE `standard-rwo` and AWS gp3 examples do not apply. + +```yaml +runtime-api: + env: + STORAGE_CLASS: managed-csi +postgresql: + primary: + persistence: + storageClass: managed-csi +``` + +Account for node disk-attachment limits and availability-zone topology. Validate +mounting, persisted content after reattachment, and expansion on a disposable PVC. +The evaluation passed these checks, including expansion from 1 GiB to 2 GiB. +See [Azure Disk CSI provisioning](https://learn.microsoft.com/en-us/azure/aks/create-volume-azure-disk). + +## Object Storage + +Conversation/session storage is separate from workspace PVCs. Helm chart 0.74.0 +supports S3-compatible and GCS filestore configuration; it does not expose an Azure +Blob backend. Do not substitute Azure Blob credentials into the S3 settings. + +The evaluation used the chart's optional RustFS store on Azure Disk CSI. For production, +select a supported durable object store and validate backup, restore and availability. +When enabling automations, also provision the separate automation package bucket. + +## Database + +Use the [External PostgreSQL](/enterprise/external-postgres) guide for production +database requirements and Helm values. If choosing Azure Database for PostgreSQL, +verify compatibility, TLS and network access against those requirements before deployment. +The AKS evaluation used bundled PostgreSQL; a managed Azure database was not validated. + +## Ingress + +Run an ingress controller on the general pool. Traefik was validated with an Azure +public LoadBalancer: + +```yaml +service: + type: LoadBalancer +``` + +Read the Service's external IP and create a wildcard A record for your base domain. +DNS can remain with another provider. For private ingress, select the appropriate +Azure load-balancer configuration and verify client access separately. + +Follow [DNS and TLS](/enterprise/k8s-install/dns-and-tls) for trusted certificates +and hostname configuration. Use flat runtime hostnames with Traefik standard Ingress. +If provisioning certificates manually, assign renewal and Secret-update ownership. + +## Installation Skill + +For agent-assisted evaluation setup, the +[AKS installation skill](https://github.com/OpenHands/OpenHands-Cloud/blob/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/.agents/skills/aks-install.md) +and [companion files](https://github.com/OpenHands/OpenHands-Cloud/tree/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/aks-install) +provide Azure commands, values templates and storage/automation checks. +The skill is currently proposed in [PR #1339](https://github.com/OpenHands/OpenHands-Cloud/pull/1339). + +## Next Steps + + + + Configure and verify the sandbox runtime. + + + Configure hostnames and trusted certificates. + + + Deploy the application and validate a conversation. + + + Configure automation storage and verify a completed run. + + + ### Amazon EKS Source: https://docs.openhands.dev/enterprise/k8s-install/eks.md diff --git a/llms.txt b/llms.txt index 638b3c95d..4f37b8143 100644 --- a/llms.txt +++ b/llms.txt @@ -259,6 +259,7 @@ from the OpenHands Software Agent SDK. - [Admin Console Configuration](https://docs.openhands.dev/enterprise/vm-install/admin-console-configuration.md): Configure an OpenHands Enterprise VM deployment from the Replicated Admin Console. - [Amazon EKS](https://docs.openhands.dev/enterprise/k8s-install/eks.md): Prepare an Amazon EKS cluster to run OpenHands Enterprise +- [Azure AKS](https://docs.openhands.dev/enterprise/k8s-install/aks.md): Prepare an Azure Kubernetes Service cluster to run OpenHands Enterprise - [Analytics](https://docs.openhands.dev/enterprise/analytics.md): Deploy Laminar for trace analysis in OpenHands Enterprise. - [Authentik](https://docs.openhands.dev/enterprise/integrations/saml-providers/authentik.md): Configure Authentik as a SAML identity provider for OpenHands Enterprise. - [Azure DevOps](https://docs.openhands.dev/enterprise/integrations/azure-devops.md): Configure Azure DevOps authentication and automation triggers for OpenHands Enterprise. From 0cf2523dd48166c1ca7592f84ba600b4ad2811aa Mon Sep 17 00:00:00 2001 From: Rajiv Shah Date: Fri, 2 Oct 2026 12:27:25 -0500 Subject: [PATCH 2/7] Use Azure logo for AKS documentation --- enterprise/images/azure-logo.svg | 25 +++++++++++++++++++++++++ enterprise/k8s-install/aks.mdx | 2 +- 2 files changed, 26 insertions(+), 1 deletion(-) create mode 100644 enterprise/images/azure-logo.svg diff --git a/enterprise/images/azure-logo.svg b/enterprise/images/azure-logo.svg new file mode 100644 index 000000000..fc4c5caee --- /dev/null +++ b/enterprise/images/azure-logo.svg @@ -0,0 +1,25 @@ + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/enterprise/k8s-install/aks.mdx b/enterprise/k8s-install/aks.mdx index d897a443a..b65d6dc13 100644 --- a/enterprise/k8s-install/aks.mdx +++ b/enterprise/k8s-install/aks.mdx @@ -1,7 +1,7 @@ --- title: Azure AKS description: Prepare an Azure Kubernetes Service cluster to run OpenHands Enterprise -icon: microsoft +icon: /enterprise/images/azure-logo.svg --- Prepare Azure Kubernetes Service (AKS) for OpenHands Enterprise, then follow the From c30d7b89a2d60a16e60ae4aee522efb0f496c4cc Mon Sep 17 00:00:00 2001 From: Rajiv Shah Date: Tue, 6 Oct 2026 16:59:06 -0500 Subject: [PATCH 3/7] Document runc as the evaluated AKS sandbox configuration --- enterprise/k8s-install/aks.mdx | 181 +++++++++++++++++++++--------- enterprise/k8s-install/sysbox.mdx | 11 +- 2 files changed, 136 insertions(+), 56 deletions(-) diff --git a/enterprise/k8s-install/aks.mdx b/enterprise/k8s-install/aks.mdx index b65d6dc13..94ace8123 100644 --- a/enterprise/k8s-install/aks.mdx +++ b/enterprise/k8s-install/aks.mdx @@ -4,17 +4,30 @@ description: Prepare an Azure Kubernetes Service cluster to run OpenHands Enterp icon: /enterprise/images/azure-logo.svg --- -Prepare Azure Kubernetes Service (AKS) for OpenHands Enterprise, then follow the -[Helm installation](/enterprise/k8s-install/installation) for application configuration. -This page covers Azure permissions, node pools, persistent storage, and ingress. +Running OpenHands Enterprise on Azure Kubernetes Service (AKS) follows the standard +[Helm installation](/enterprise/k8s-install/installation), with provider-specific +choices for node pools, storage, ingress and the sandbox runtime. This guide covers +preparing the cluster with standard runc sandboxes. Once it is ready, follow the +Helm guide to deploy using the runtime values below. Skip +[Installing Sysbox](/enterprise/k8s-install/sysbox). Wherever the Helm guide installs +or checks Sysbox, use the runc values and checks on this page instead. + + +Sysbox is not yet supported for OpenHands Enterprise on AKS. This guide uses the +node’s default runc runtime, evaluated with ordinary coding workflows (see Validation Scope). Docker builds +and Docker Compose inside the sandbox are unavailable in this configuration. +Other system-container workloads, such as systemd or nested containers, were not validated. +Runc does not provide Sysbox’s additional system-container isolation; review the +isolation requirements for your workloads before production adoption. + ## Cluster Requirements | Requirement | Recommendation | | --- | --- | | Access | An Azure subscription and permissions to create the resource group, cluster, node pools and networking | -| AKS version | A supported AKS version compatible with [Sysbox](/enterprise/k8s-install/sysbox); verify the actual Ubuntu image and containerd version | -| Sandbox OS | Ubuntu nodes with a working Sysbox runtime | +| AKS version | An AKS-supported version; verify the actual Ubuntu image and containerd version against your target Enterprise release | +| Sandbox OS | Ubuntu nodes using the default containerd/runc runtime | | Storage class | Azure Disk CSI, such as `managed-csi`, with expansion enabled | | Capacity | Sufficient regional and VM-family vCPU quota for application nodes, sandbox nodes and upgrades | @@ -33,8 +46,9 @@ Use separate pools for application services and sandbox workloads: - **General pool:** runs OpenHands services and cluster add-ons. Keep application workloads here through node affinity or selectors. -- **Sysbox pool:** an Ubuntu user node pool for agent sandboxes. Configure - `sysbox-install=yes` as a persistent node-pool label so new nodes receive the installer. +- **Sandbox pool:** an Ubuntu user node pool for agent sandboxes. Set + `workload=openhands-sandbox` as a persistent node-pool label so new nodes match + the Runtime API selector. Do not install the Sysbox DaemonSet or add its installer label. The evaluation used Azure CNI overlay with Calico, two platform nodes and a sandbox pool with autoscaler bounds of one to two nodes. Size the pools using the @@ -42,39 +56,84 @@ pool with autoscaler bounds of one to two nodes. Size the pools using the [Resource Limits](/enterprise/k8s-install/resource-limits), including node overhead, image storage and warm sandbox capacity. Validate scale-down behavior for active sessions. -### Set Up the Sandbox Runtime +AKS rejected a manual node-pool scale-down with `UnsatisfiablePDB` because of a +sandbox PodDisruptionBudget with `maxUnavailable: 0`. Node-image and Kubernetes +upgrades, which drain nodes, were not tested. Validate maintenance before relying +on it; do not delete or patch PodDisruptionBudgets to force it. -1. Follow [Installing Sysbox](/enterprise/k8s-install/sysbox) to install Sysbox on - your Ubuntu sandbox pool. -2. Before installing OpenHands, run a Sysbox test pod with an Azure Disk workspace - using the [AKS installation skill and companion checks](https://github.com/OpenHands/OpenHands-Cloud/tree/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/aks-install). - Confirm the pod starts and can write and read a file. -3. After installing OpenHands, start a conversation and ask it to run `pwd`. - Confirm it returns the workspace directory. +### Configure runc Sandboxes - -Some AKS node images need an additional Sysbox setup step. If the test pod fails -with an error mentioning `sysbox-runc`, contact OpenHands support before continuing. -The AKS installation skill includes the workaround used for the evaluation. - +Override the Sysbox defaults in the Enterprise chart through Helm values: -Before enabling sandbox-pool autoscaling, confirm that each newly created node can -start a sandbox. A setup correction applied to existing nodes must also be applied -to replacement and new nodes. Validate this with your support team before relying -on automatic scale-out or node upgrades. +```yaml +runtime-api: + env: + RUNTIME_CLASS: "" + SET_HOST_USERS: "false" + RUNTIME_NODE_SELECTOR: '{"workload":"openhands-sandbox"}' + RUNTIME_TOLERATIONS: '[]' +``` - -The failure was observed on AKS 1.35.8, Ubuntu 24.04.5 LTS and containerd 2.3.3-2 -with Sysbox installer v0.7.1-0. The installer registered the runtime under the old -containerd plugin key, so its success message did not establish a working runtime. -See [upstream issue 997](https://github.com/nestybox/sysbox/issues/997). +An empty `RUNTIME_CLASS` omits `runtimeClassName` from new sandbox pods, using +containerd’s default runtime. `SET_HOST_USERS: "false"` omits the `hostUsers` override; +it does not request the Sysbox user-namespace configuration. Save these values in +`values-aks-runc.yaml` and pass that file as the **last** `-f` argument on every +install and upgrade, after your base values and feature overlays. For example: -The evaluation workaround corrected only the Sysbox registration and restarted -containerd on each affected sandbox node. It requires a reviewed per-node operation -and does not automatically cover nodes added by autoscaling, reimaging or upgrades. -The [AKS install skill](https://github.com/OpenHands/OpenHands-Cloud/blob/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/.agents/skills/aks-install.md) -contains the diagnostic and correction procedure. Verify a real test pod afterward. - +```bash +helm upgrade --install openhands --version 0.74.0 \ + --namespace openhands --create-namespace \ + -f values.yaml -f values-automation.yaml -f values-aks-runc.yaml --timeout 10m +``` + +Include `values-automation.yaml` only when enabling automations. Follow the Helm +guide for namespaces, Secrets and other prerequisites. Omitting these overrides +can restore the chart’s Sysbox defaults and prevent new sandboxes from starting. +If you add a sandbox-pool taint, configure a matching toleration instead of the empty list. + +After deployment, create a conversation and verify that its pod is in the +sandbox pool with no `runtimeClassName` or `hostUsers` override: + +```bash +kubectl get pod -n \ + -o jsonpath='{.spec.nodeName}{"\t"}{.spec.runtimeClassName}{"\t"}{.spec.hostUsers}{"\n"}' +``` + +Expect a sandbox-pool node name followed by two empty fields. Use the configured +runtime namespace (`openhands-runtimes` in the evaluation). The evaluated +sandboxes ran without privileged mode, a host Docker socket or added capabilities. +Confirm the agent executes commands and writes and rereads a workspace marker. +Python, Node/npm, Git (public clone; authenticated pushes not tested), browser +navigation/screenshots, HTTP previews and workspace +persistence passed the runc evaluation. Docker-in-sandbox failed; an installed +Docker CLI or Compose binary does not establish daemon functionality. +See [Docker in the agent sandbox](/enterprise/docker-in-sandbox) for that feature’s requirements. + +### Keep Startup Capacity Ready + +Keep enough Ready sandbox nodes with the agent image cached to accommodate the +expected concurrent sessions. Pool minimum counts alone do not establish available +capacity: account for existing pods, memory requests, disk limits and node overhead. +The tested defaults requested 500m CPU and 3 GiB memory per sandbox; a tested +4-vCPU/16-GiB node accommodated three such sandboxes after Kubernetes reservations. +Size production capacity using your own resources and workload measurements. + +In the October 6 evaluation, six starts on ready/cached capacity all succeeded in +38–51 seconds. A separate three-start burst forcing a zero-node pool to scale up +failed around 120 seconds: the new node became Ready at 143 seconds and sandbox +pods at approximately 234–247 seconds. The failed startup tasks did not recover +automatically after their pods became healthy. Three fresh starts on that same node +with its image cached then all succeeded in 53–55 seconds. + +This locates the observed failure before sandbox execution, in cold provisioning +and startup orchestration. It is not a runc-versus-Sysbox performance benchmark. +Nodes that were already Ready with the image cached avoided this failure in testing. +Pre-pulling images onto newly added nodes was not tested and would not remove the +node-readiness delay. No startup deadline or recovery fix was validated. Do not +rely on the autoscaler adding nodes, including scale-from-zero, to absorb interactive +startups in the tested release. Validate scale-out, node replacement and maintenance +for your target release. Pause orphaned sandboxes from failed startups through +supported OpenHands APIs or the UI’s Stop Runtime action before retrying. ## Persistent Storage @@ -84,9 +143,12 @@ Inspect the Azure Disk CSI StorageClass: kubectl get storageclass managed-csi -o yaml ``` -The evaluated class uses `disk.csi.azure.com`, `WaitForFirstConsumer`, -`ReadWriteOnce` and volume expansion. Configure stateful components and workspace -storage explicitly; the GKE `standard-rwo` and AWS gp3 examples do not apply. +The evaluated class uses `disk.csi.azure.com`, `WaitForFirstConsumer`, reclaim +policy `Delete` and volume expansion. The evaluated Azure Disk PVCs use +`ReadWriteOnce`. Configure stateful components and workspace +storage explicitly, including persistence for any in-cluster object store; the GKE +`standard-rwo` and AWS gp3 examples do not apply. Add `STORAGE_CLASS` to the same +`runtime-api.env` map as the runc values above; do not create duplicate YAML keys. ```yaml runtime-api: @@ -100,7 +162,9 @@ postgresql: Account for node disk-attachment limits and availability-zone topology. Validate mounting, persisted content after reattachment, and expansion on a disposable PVC. -The evaluation passed these checks, including expansion from 1 GiB to 2 GiB. +The runc evaluation passed workspace read/write, stop/reopen persistence and +cross-node disk reattachment. Expansion was not validated on the runc configuration; +test it on a disposable PVC. See [Azure Disk CSI provisioning](https://learn.microsoft.com/en-us/azure/aks/create-volume-azure-disk). ## Object Storage @@ -109,21 +173,23 @@ Conversation/session storage is separate from workspace PVCs. Helm chart 0.74.0 supports S3-compatible and GCS filestore configuration; it does not expose an Azure Blob backend. Do not substitute Azure Blob credentials into the S3 settings. -The evaluation used the chart's optional RustFS store on Azure Disk CSI. For production, -select a supported durable object store and validate backup, restore and availability. -When enabling automations, also provision the separate automation package bucket. +The tested configuration used the chart's optional RustFS store on Azure Disk CSI. +Select an object store that meets your durability and availability requirements, +and verify backup and restore for that configuration. +When enabling automations, configure a separate automation package bucket. The +automation service can create it at startup when bucket creation is enabled. ## Database -Use the [External PostgreSQL](/enterprise/external-postgres) guide for production +Use the [External PostgreSQL](/enterprise/external-postgres) guide for managed database requirements and Helm values. If choosing Azure Database for PostgreSQL, verify compatibility, TLS and network access against those requirements before deployment. -The AKS evaluation used bundled PostgreSQL; a managed Azure database was not validated. +The tested configuration used bundled PostgreSQL. ## Ingress Run an ingress controller on the general pool. Traefik was validated with an Azure -public LoadBalancer: +public LoadBalancer. In the Traefik Helm values (chart `41.6.0` was tested): ```yaml service: @@ -138,20 +204,27 @@ Follow [DNS and TLS](/enterprise/k8s-install/dns-and-tls) for trusted certificat and hostname configuration. Use flat runtime hostnames with Traefik standard Ingress. If provisioning certificates manually, assign renewal and Secret-update ownership. -## Installation Skill +## Validation Scope -For agent-assisted evaluation setup, the -[AKS installation skill](https://github.com/OpenHands/OpenHands-Cloud/blob/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/.agents/skills/aks-install.md) -and [companion files](https://github.com/OpenHands/OpenHands-Cloud/tree/832e57028298f3ad564e6ee7461ee1d9d2ee6d62/aks-install) -provide Azure commands, values templates and storage/automation checks. -The skill is currently proposed in [PR #1339](https://github.com/OpenHands/OpenHands-Cloud/pull/1339). + +Validated October 6, 2026 with Helm chart `0.74.0`, OpenHands `1.67.0`, Runtime API +`0.10.0` and agent-server `1.49.6-python` on AKS `1.35.8`, Ubuntu `24.04.5 LTS` +and containerd `2.3.3-2`. The configuration used bundled PostgreSQL and RustFS. +GitHub login, trusted HTTPS, ordinary coding/browser workflows, workspace persistence, +cross-node disk movement and execution on an autoscaler-added node, after it +became Ready, passed. Newly added nodes ran sandboxes without per-node runtime setup. + +Cold autoscaling exceeded the startup deadline. An earlier resume-related HTTP 401 +and a PodDisruptionBudget issue blocking manual pool scale-down remain unresolved. +Automations and PVC expansion were not validated on this runc configuration. Managed Azure PostgreSQL, Azure Blob, private +endpoints, backup/restore, cross-zone recovery, tenant isolation and network +isolation between sandboxes were not covered. No NetworkPolicy targeted the +sandbox namespace in the evaluation. This is evaluation evidence, not production acceptance. + ## Next Steps - - Configure and verify the sandbox runtime. - Configure hostnames and trusted certificates. diff --git a/enterprise/k8s-install/sysbox.mdx b/enterprise/k8s-install/sysbox.mdx index d16399d9e..70659f62f 100644 --- a/enterprise/k8s-install/sysbox.mdx +++ b/enterprise/k8s-install/sysbox.mdx @@ -4,8 +4,15 @@ description: Install the Sysbox runtime so agent sandboxes can run securely icon: cube --- -OpenHands runs each agent session in a sandbox that uses [Sysbox](https://github.com/nestybox/sysbox) -for isolation. This guide covers installing Sysbox. +OpenHands Enterprise can use [Sysbox](https://github.com/nestybox/sysbox) +for sandbox isolation. This guide covers installing Sysbox. + + +Sysbox is not yet supported for OpenHands Enterprise on **Azure Kubernetes Service +(AKS)**. Do not use this procedure for OpenHands Enterprise on AKS. See the +[Azure AKS guide](/enterprise/k8s-install/aks) for the evaluated runc configuration +and its Docker-in-sandbox, isolation and startup limitations. + ## Node Requirements From 2d81473a805d593bbeba45ff24fdb52c6f8c2ec9 Mon Sep 17 00:00:00 2001 From: Rajiv Shah Date: Tue, 6 Oct 2026 17:06:18 -0500 Subject: [PATCH 4/7] Keep AKS guide focused on customer setup --- enterprise/k8s-install/aks.mdx | 149 ++++++++++++--------------------- 1 file changed, 54 insertions(+), 95 deletions(-) diff --git a/enterprise/k8s-install/aks.mdx b/enterprise/k8s-install/aks.mdx index 94ace8123..dfb6bc3d5 100644 --- a/enterprise/k8s-install/aks.mdx +++ b/enterprise/k8s-install/aks.mdx @@ -13,14 +13,17 @@ Helm guide to deploy using the runtime values below. Skip or checks Sysbox, use the runc values and checks on this page instead. -Sysbox is not yet supported for OpenHands Enterprise on AKS. This guide uses the -node’s default runc runtime, evaluated with ordinary coding workflows (see Validation Scope). Docker builds -and Docker Compose inside the sandbox are unavailable in this configuration. -Other system-container workloads, such as systemd or nested containers, were not validated. -Runc does not provide Sysbox’s additional system-container isolation; review the -isolation requirements for your workloads before production adoption. +Sysbox is not yet supported for OpenHands Enterprise on AKS. Use the runc +configuration below. Docker builds and Docker Compose inside the sandbox are +unavailable, and runc does not provide Sysbox's additional isolation. Confirm that +these limits meet your workload requirements before choosing this configuration. + +The values below apply to Enterprise Helm chart `0.74.0` and Runtime API `0.10.0`. +Check the values for your licensed release before installing or upgrading. + + ## Cluster Requirements | Requirement | Recommendation | @@ -33,8 +36,8 @@ isolation requirements for your workloads before production adoption. Entra or Azure DevOps authentication alone does not establish subscription access. Check `az account list --all` and the selected subscription's permissions. -Contributor was sufficient for the evaluated default managed-network deployment; -Azure role assignments require additional permissions. Existing networks, private +Grant the deployment identity access to create the cluster and its dependencies. +Creating Azure role assignments requires additional permissions. Existing networks, private cluster access and workload identities need their own access checks. Check VM SKU restrictions as well as quota before selecting a region or availability @@ -50,16 +53,16 @@ Use separate pools for application services and sandbox workloads: `workload=openhands-sandbox` as a persistent node-pool label so new nodes match the Runtime API selector. Do not install the Sysbox DaemonSet or add its installer label. -The evaluation used Azure CNI overlay with Calico, two platform nodes and a sandbox -pool with autoscaler bounds of one to two nodes. Size the pools using the -[Sizing Guide](/enterprise/sizing-guide) and -[Resource Limits](/enterprise/k8s-install/resource-limits), including node overhead, -image storage and warm sandbox capacity. Validate scale-down behavior for active sessions. +Size both pools using the [Sizing Guide](/enterprise/sizing-guide) and +[Resource Limits](/enterprise/k8s-install/resource-limits). Leave capacity for +Kubernetes services, existing workloads and concurrent sandboxes. Check +PodDisruptionBudgets before draining nodes for maintenance or upgrades. -AKS rejected a manual node-pool scale-down with `UnsatisfiablePDB` because of a -sandbox PodDisruptionBudget with `maxUnavailable: 0`. Node-image and Kubernetes -upgrades, which drain nodes, were not tested. Validate maintenance before relying -on it; do not delete or patch PodDisruptionBudgets to force it. + +Keep sufficient sandbox capacity available before users start conversations. +Provisioning a new AKS node can exceed the sandbox startup timeout. Do not rely +on cold autoscaling, including scale-from-zero, for immediate conversation startup. + ### Configure runc Sandboxes @@ -83,57 +86,34 @@ install and upgrade, after your base values and feature overlays. For example: ```bash helm upgrade --install openhands --version 0.74.0 \ --namespace openhands --create-namespace \ - -f values.yaml -f values-automation.yaml -f values-aks-runc.yaml --timeout 10m + -f values.yaml -f values-aks-runc.yaml --timeout 10m ``` -Include `values-automation.yaml` only when enabling automations. Follow the Helm -guide for namespaces, Secrets and other prerequisites. Omitting these overrides +Use the licensed chart source from the [Helm guide](/enterprise/k8s-install/installation), +and complete its namespace and Secret setup first. When enabling optional features, +place their values files before `values-aks-runc.yaml`. Omitting these overrides can restore the chart’s Sysbox defaults and prevent new sandboxes from starting. If you add a sandbox-pool taint, configure a matching toleration instead of the empty list. -After deployment, create a conversation and verify that its pod is in the -sandbox pool with no `runtimeClassName` or `hostUsers` override: +### Verify a Sandbox -```bash -kubectl get pod -n \ - -o jsonpath='{.spec.nodeName}{"\t"}{.spec.runtimeClassName}{"\t"}{.spec.hostUsers}{"\n"}' -``` +1. Sign in to OpenHands and start a conversation. Ask the agent to run `pwd`, + write a file in its workspace and read it back. +2. Confirm the sandbox pod runs on the sandbox pool using the default runtime: -Expect a sandbox-pool node name followed by two empty fields. Use the configured -runtime namespace (`openhands-runtimes` in the evaluation). The evaluated -sandboxes ran without privileged mode, a host Docker socket or added capabilities. -Confirm the agent executes commands and writes and rereads a workspace marker. -Python, Node/npm, Git (public clone; authenticated pushes not tested), browser -navigation/screenshots, HTTP previews and workspace -persistence passed the runc evaluation. Docker-in-sandbox failed; an installed -Docker CLI or Compose binary does not establish daemon functionality. -See [Docker in the agent sandbox](/enterprise/docker-in-sandbox) for that feature’s requirements. - -### Keep Startup Capacity Ready - -Keep enough Ready sandbox nodes with the agent image cached to accommodate the -expected concurrent sessions. Pool minimum counts alone do not establish available -capacity: account for existing pods, memory requests, disk limits and node overhead. -The tested defaults requested 500m CPU and 3 GiB memory per sandbox; a tested -4-vCPU/16-GiB node accommodated three such sandboxes after Kubernetes reservations. -Size production capacity using your own resources and workload measurements. - -In the October 6 evaluation, six starts on ready/cached capacity all succeeded in -38–51 seconds. A separate three-start burst forcing a zero-node pool to scale up -failed around 120 seconds: the new node became Ready at 143 seconds and sandbox -pods at approximately 234–247 seconds. The failed startup tasks did not recover -automatically after their pods became healthy. Three fresh starts on that same node -with its image cached then all succeeded in 53–55 seconds. - -This locates the observed failure before sandbox execution, in cold provisioning -and startup orchestration. It is not a runc-versus-Sysbox performance benchmark. -Nodes that were already Ready with the image cached avoided this failure in testing. -Pre-pulling images onto newly added nodes was not tested and would not remove the -node-readiness delay. No startup deadline or recovery fix was validated. Do not -rely on the autoscaler adding nodes, including scale-from-zero, to absorb interactive -startups in the tested release. Validate scale-out, node replacement and maintenance -for your target release. Pause orphaned sandboxes from failed startups through -supported OpenHands APIs or the UI’s Stop Runtime action before retrying. + ```bash + kubectl get pod -n \ + -o jsonpath='{.spec.nodeName}{"\t"}{.spec.runtimeClassName}{"\t"}{.spec.hostUsers}{"\n"}' + ``` + + Expect a sandbox-pool node name followed by two empty fields. Use the runtime + namespace configured in your Helm values, such as `openhands-runtimes`. +3. Stop the conversation runtime, reopen the conversation and ask the agent to + read the same file to confirm workspace persistence. + +If a sandbox fails to start, follow the +[Troubleshooting guide](/enterprise/troubleshooting) and collect the pod events +before contacting OpenHands support. ## Persistent Storage @@ -143,12 +123,11 @@ Inspect the Azure Disk CSI StorageClass: kubectl get storageclass managed-csi -o yaml ``` -The evaluated class uses `disk.csi.azure.com`, `WaitForFirstConsumer`, reclaim -policy `Delete` and volume expansion. The evaluated Azure Disk PVCs use -`ReadWriteOnce`. Configure stateful components and workspace -storage explicitly, including persistence for any in-cluster object store; the GKE -`standard-rwo` and AWS gp3 examples do not apply. Add `STORAGE_CLASS` to the same -`runtime-api.env` map as the runc values above; do not create duplicate YAML keys. +Confirm that the class uses the Azure Disk CSI provisioner (`disk.csi.azure.com`) +and supports the binding mode and expansion your workloads need. Azure Disk +workspace PVCs use `ReadWriteOnce`. Configure storage explicitly for workspaces, +PostgreSQL and any in-cluster object store. Merge the following settings into your +base values file: ```yaml runtime-api: @@ -162,9 +141,6 @@ postgresql: Account for node disk-attachment limits and availability-zone topology. Validate mounting, persisted content after reattachment, and expansion on a disposable PVC. -The runc evaluation passed workspace read/write, stop/reopen persistence and -cross-node disk reattachment. Expansion was not validated on the runc configuration; -test it on a disposable PVC. See [Azure Disk CSI provisioning](https://learn.microsoft.com/en-us/azure/aks/create-volume-azure-disk). ## Object Storage @@ -173,9 +149,10 @@ Conversation/session storage is separate from workspace PVCs. Helm chart 0.74.0 supports S3-compatible and GCS filestore configuration; it does not expose an Azure Blob backend. Do not substitute Azure Blob credentials into the S3 settings. -The tested configuration used the chart's optional RustFS store on Azure Disk CSI. -Select an object store that meets your durability and availability requirements, -and verify backup and restore for that configuration. +For an in-cluster store, chart `0.74.0` includes optional RustFS. Configure its +persistence with Azure Disk CSI, provide the object-store credential Secret, and +create the conversation bucket before starting conversations. Select an object +store that meets your durability requirements and configure backup and restore. When enabling automations, configure a separate automation package bucket. The automation service can create it at startup when bucket creation is enabled. @@ -184,12 +161,12 @@ automation service can create it at startup when bucket creation is enabled. Use the [External PostgreSQL](/enterprise/external-postgres) guide for managed database requirements and Helm values. If choosing Azure Database for PostgreSQL, verify compatibility, TLS and network access against those requirements before deployment. -The tested configuration used bundled PostgreSQL. +For bundled PostgreSQL, set its persistence storage class as shown above. ## Ingress -Run an ingress controller on the general pool. Traefik was validated with an Azure -public LoadBalancer. In the Traefik Helm values (chart `41.6.0` was tested): +Run an ingress controller on the general pool. For Traefik with an Azure public +load balancer, use the following Traefik Helm values: ```yaml service: @@ -204,24 +181,6 @@ Follow [DNS and TLS](/enterprise/k8s-install/dns-and-tls) for trusted certificat and hostname configuration. Use flat runtime hostnames with Traefik standard Ingress. If provisioning certificates manually, assign renewal and Secret-update ownership. -## Validation Scope - - -Validated October 6, 2026 with Helm chart `0.74.0`, OpenHands `1.67.0`, Runtime API -`0.10.0` and agent-server `1.49.6-python` on AKS `1.35.8`, Ubuntu `24.04.5 LTS` -and containerd `2.3.3-2`. The configuration used bundled PostgreSQL and RustFS. -GitHub login, trusted HTTPS, ordinary coding/browser workflows, workspace persistence, -cross-node disk movement and execution on an autoscaler-added node, after it -became Ready, passed. Newly added nodes ran sandboxes without per-node runtime setup. - -Cold autoscaling exceeded the startup deadline. An earlier resume-related HTTP 401 -and a PodDisruptionBudget issue blocking manual pool scale-down remain unresolved. -Automations and PVC expansion were not validated on this runc configuration. Managed Azure PostgreSQL, Azure Blob, private -endpoints, backup/restore, cross-zone recovery, tenant isolation and network -isolation between sandboxes were not covered. No NetworkPolicy targeted the -sandbox namespace in the evaluation. This is evaluation evidence, not production acceptance. - - ## Next Steps @@ -232,6 +191,6 @@ sandbox namespace in the evaluation. This is evaluation evidence, not production Deploy the application and validate a conversation. - Configure automation storage and verify a completed run. + Configure optional automation services and storage. From 84d39a4bc632d3dedf3e4b956bc9304a6906a587 Mon Sep 17 00:00:00 2001 From: Rajiv Shah Date: Tue, 6 Oct 2026 17:07:07 -0500 Subject: [PATCH 5/7] Remove runtime warning from AKS setup guide --- enterprise/k8s-install/aks.mdx | 7 ------- 1 file changed, 7 deletions(-) diff --git a/enterprise/k8s-install/aks.mdx b/enterprise/k8s-install/aks.mdx index dfb6bc3d5..d544e5156 100644 --- a/enterprise/k8s-install/aks.mdx +++ b/enterprise/k8s-install/aks.mdx @@ -12,13 +12,6 @@ Helm guide to deploy using the runtime values below. Skip [Installing Sysbox](/enterprise/k8s-install/sysbox). Wherever the Helm guide installs or checks Sysbox, use the runc values and checks on this page instead. - -Sysbox is not yet supported for OpenHands Enterprise on AKS. Use the runc -configuration below. Docker builds and Docker Compose inside the sandbox are -unavailable, and runc does not provide Sysbox's additional isolation. Confirm that -these limits meet your workload requirements before choosing this configuration. - - The values below apply to Enterprise Helm chart `0.74.0` and Runtime API `0.10.0`. Check the values for your licensed release before installing or upgrading. From 65ffe3af6da44b9988022a4a9913475a0a74fc7e Mon Sep 17 00:00:00 2001 From: Rajiv Shah Date: Tue, 6 Oct 2026 17:07:18 -0500 Subject: [PATCH 6/7] Remove startup capacity warning from AKS guide --- enterprise/k8s-install/aks.mdx | 6 ------ 1 file changed, 6 deletions(-) diff --git a/enterprise/k8s-install/aks.mdx b/enterprise/k8s-install/aks.mdx index d544e5156..e45416b89 100644 --- a/enterprise/k8s-install/aks.mdx +++ b/enterprise/k8s-install/aks.mdx @@ -51,12 +51,6 @@ Size both pools using the [Sizing Guide](/enterprise/sizing-guide) and Kubernetes services, existing workloads and concurrent sandboxes. Check PodDisruptionBudgets before draining nodes for maintenance or upgrades. - -Keep sufficient sandbox capacity available before users start conversations. -Provisioning a new AKS node can exceed the sandbox startup timeout. Do not rely -on cold autoscaling, including scale-from-zero, for immediate conversation startup. - - ### Configure runc Sandboxes Override the Sysbox defaults in the Enterprise chart through Helm values: From 6cfe40c9b5ab2fea57b76c92d707b3ec86702fe6 Mon Sep 17 00:00:00 2001 From: Rajiv Shah Date: Wed, 7 Oct 2026 05:39:33 -0500 Subject: [PATCH 7/7] Link AKS install guide to the installation skill --- enterprise/k8s-install/aks.mdx | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/enterprise/k8s-install/aks.mdx b/enterprise/k8s-install/aks.mdx index e45416b89..946288b30 100644 --- a/enterprise/k8s-install/aks.mdx +++ b/enterprise/k8s-install/aks.mdx @@ -17,6 +17,10 @@ The values below apply to Enterprise Helm chart `0.74.0` and Runtime API `0.10.0 Check the values for your licensed release before installing or upgrading. +For an agent-assisted installation on a dedicated test cluster, use the +[AKS installation skill](https://github.com/OpenHands/OpenHands-Cloud/blob/main/.agents/skills/aks-install.md) +and its [setup scripts and templates](https://github.com/OpenHands/OpenHands-Cloud/tree/main/aks-install). + ## Cluster Requirements | Requirement | Recommendation |