A Kubernetes operator for running Memgraph high-availability clusters. It exposes a MemgraphCluster custom resource (API group memgraph.com/v1alpha1, short name mgc): declare the cluster topology in a single resource and the operator provisions the workloads, bootstraps HA registration, and continuously reconciles registration state.
Status: early development (v1alpha1). The API and its guarantees can still change between releases. Read what v1alpha1 does and does not do before running it anywhere that matters. The previous attempt at this operator is preserved on the
archive/pre-operator-mvpbranch.
The operator replaces the memgraph-high-availability Helm chart's fire-and-forget registration Job with a controller that continuously drives the cluster toward its declared topology: one StatefulSet per role (coordinators, data instances), automatic bootstrap and MAIN promotion, and automatic re-registration of instances that lose their registration state. See docs/ for the design of each feature.
From an empty cluster to a registered, MAIN-elected Memgraph HA cluster.
You need: a Kubernetes cluster (v1.25 or newer — the CRD uses CEL validation), kubectl, helm v3, and a Memgraph enterprise license, which high availability requires. Three coordinators and two data instances need five schedulable pods and ten PersistentVolumeClaims of 1Gi each.
The install chart ships the MemgraphCluster CRD, a least-privilege RBAC set, and the controller Deployment. One release per cluster is enough — the operator watches every namespace.
helm repo add memgraph https://memgraph.github.io/helm-charts
helm repo update
helm install memgraph-operator memgraph/memgraph-operator \
--namespace memgraph-operator-system --create-namespace --waitThe cluster reads its license from a Secret you own; no license material ever goes into the MemgraphCluster resource, which keeps it safe to commit to git.
kubectl create namespace memgraph
kubectl create secret generic memgraph-secrets \
--namespace memgraph \
--from-literal=MEMGRAPH_ENTERPRISE_LICENSE='<your-license-key>' \
--from-literal=MEMGRAPH_ORGANIZATION_NAME='<your-organization-name>'This is the whole resource — counts, image, and the Secret from the previous step. Everything else (storage, ports, probes, resources, cluster domain) takes its default:
apiVersion: memgraph.com/v1alpha1
kind: MemgraphCluster
metadata:
name: memgraph
spec:
coordinators: 3
dataInstances: 2
image:
repository: docker.io/memgraph/memgraph
tag: "3.13.0"
secrets:
name: memgraph-secrets
licenseKey: MEMGRAPH_ENTERPRISE_LICENSE
organizationKey: MEMGRAPH_ORGANIZATION_NAMEIt is examples/minimal-cluster.yaml in this repository, and the end-to-end suite applies that file unmodified on every pull request, so it stays a manifest that works:
kubectl apply -n memgraph \
-f https://raw.githubusercontent.com/memgraph/kubernetes-operator/main/examples/minimal-cluster.yamlIf your Secret has a different name, or stores the license under different keys, change the secrets block to match — that is the only edit the example needs.
The operator creates one StatefulSet and one headless Service per role, waits for the pods to become ready, then registers the coordinators and data instances with each other and promotes the initial MAIN. On a cluster that has to pull the Memgraph image, expect a few minutes.
kubectl get mgc -n memgraph -wWhile the pods are still starting, MAIN is empty and both conditions are False; the converged cluster looks like this:
NAME COORDINATORS DATA MAIN READY CONVERGED AGE
memgraph 3 2 instance_0 True True 4m12s
MAINis the data instance the coordinators elected as MAIN — the one that accepts writes. It is observed, not decided by the operator, so it changes on failover.READYis True once a MAIN is elected, i.e. the cluster serves writes.CONVERGEDis True once every declared coordinator and data instance is registered and reported healthy, and both StatefulSets run the declared number of replicas.
kubectl get mgc -n memgraph -o wide adds how many of each role's declared members the cluster actually has registered (REGISTERED-COORDINATORS, REGISTERED-DATA), which is what a scale is watched through.
To block a script or a GitOps step on the cluster being usable:
kubectl wait --namespace memgraph --for=condition=Converged \
memgraphcluster/memgraph --timeout=10mIf it does not converge, kubectl describe mgc memgraph -n memgraph gives the condition messages (which pods are not ready, whether a coordinator is unreachable, or — with reason ApplyFailed — what the API server refused about the workloads), and the operator logs the rest:
kubectl logs -n memgraph-operator-system deploy/memgraph-operator-controller-managerThe resource's identities follow the pod ordinals. For a cluster named memgraph:
| Pod | Registered as | Role |
|---|---|---|
memgraph-coordinator-0, -1, -2 |
coordinator_0, coordinator_1, coordinator_2 |
Raft coordinators |
memgraph-data-0, -1 |
instance_0, instance_1 |
data instances (one MAIN, the rest replicas) |
Ask a coordinator for the cluster's own view of itself:
kubectl exec -n memgraph memgraph-coordinator-0 -c memgraph -- \
bash -c "echo 'SHOW INSTANCES;' | mgconsole"Every pod runs Bolt on port 7687, and the Memgraph image ships mgconsole, so the shortest path to a query is to pipe Cypher into the MAIN pod (instance_0 above is memgraph-data-0):
kubectl exec -i -n memgraph memgraph-data-0 -c memgraph -- mgconsole <<'EOF'
CREATE (:Greeting {text: "hello from the operator"});
MATCH (n:Greeting) RETURN n;
EOFkubectl exec -it -n memgraph memgraph-data-0 -c memgraph -- mgconsole opens the same client interactively.
Writes only succeed against the MAIN — the other data instances are replicas and accept reads. Check .status.main to find it:
kubectl get mgc memgraph -n memgraph -o jsonpath='{.status.main}'Applications inside the cluster reach an instance at its stable DNS name in the role's headless Service:
memgraph-data-0.memgraph-data.memgraph.svc.cluster.local:7687
To load large CSV files, put them on a shared volume every data instance mounts, so the import works on whichever instance is the MAIN — see Importing data.
To reach the cluster from outside Kubernetes, expose it with spec.externalAccess — see External access below. To point a local client such as Memgraph Lab at the cluster while evaluating, forwarding a port is enough:
kubectl port-forward -n memgraph pod/memgraph-data-0 7687:7687kubectl delete mgc memgraph -n memgraphDeleting the resource removes the StatefulSets and Services through garbage collection, but the PersistentVolumeClaims are kept — the default retention policy protects data against an accidental delete, and is what a volume-snapshot restore builds on. Remove them (and with them the data) explicitly, or set spec.storage.retentionPolicy: Delete on dev clusters that should clean up after themselves:
kubectl delete pvc -n memgraph --all
kubectl delete namespace memgraphUninstall the operator with helm uninstall memgraph-operator --namespace memgraph-operator-system. Helm never deletes CRDs it installed, so kubectl delete crd memgraphclusters.memgraph.com once no cluster needs it — see the chart README for the values and the upgrade caveat.
Beyond the quickstart's fields, v1alpha1 exposes storage (PVC size, access mode, storage class, whether Memgraph also writes log files onto the claim, retention) per role, with every size growable on a live cluster without restarting a pod, optional core dump collection with an uploader sidecar of your choice per role, the HA chart's sysctlInitContainer raising the node's vm.max_map_count from a privileged init container (unlike the chart it runs only when the block is written; the quickstart leaves it out because it is proven under the restricted Pod Security Standard, so add it wherever privileged containers are allowed), the pod identity (securityContext: runAsUser, runAsGroup, fsGroup, the chart's memgraphUserId and memgraphGroupId; absent, the images' 101 and 103, and an empty block leaves all three to the platform, see OpenShift), the chart's fixOwnershipInitContainer chowning every pod's volumes to that identity from a root init container for storage drivers that ignore fsGroup (off when absent, as in the chart), resource requests and limits per role, readiness probe timings, custom labels on pods, StatefulSets and Services, scheduling (an anti-affinity rule the operator writes and, per role, node selector, tolerations, topology spread constraints, extra anti-affinity and priority class), the cluster domain used in advertised addresses, external access through LoadBalancers or a Gateway API Gateway, a ServiceMonitor, a Grafana dashboard ConfigMap and the vmagent and Vector sidecars that push the metrics and the logs to Memgraph for remote monitoring, TLS for clients and between members from certificates you supply, Memgraph flags per role (a flag Memgraph can change live is applied to every instance without a restart; any other flag change rolls the pods) and the cluster-wide coordinator settings written to Raft once, and a freeform passthrough per role for environment variables, extra volumes and volume mounts, containers of your own beside Memgraph (the HA chart's userContainers), and init containers of your own ahead of it (the chart's initContainers). Internal ports are fixed: Bolt 7687, management 10000, replication 20000, coordinator 12000, and metrics 9091.
config/samples/v1alpha1_memgraphcluster.yaml spells the full surface out with every default and the reasoning behind it. kubectl explain mgc.spec --recursive documents the same fields from the installed CRD.
OpenShift's default restricted-v2 SCC assigns every pod's runAsUser and fsGroup from the namespace's range at admission and rejects a pod naming values outside it. Two settings cover it:
- On the cluster, write
securityContext: {}so the pods name no uid, gid or fsGroup and take the assigned ones. Memgraph needs no particular uid, only to own the data directory it creates, and OpenShift's storage honorsfsGroup, sofixOwnershipInitContaineris neither needed nor admitted there (it runs as root). - On the operator chart, drop the manager's own uid and gid:
--set podSecurityContext.runAsUser=null --set podSecurityContext.runAsGroup=null.
sysctlInitContainer and coreDumps.configureCorePattern need privileged containers, which restricted-v2 forbids as well; leave the first out and set the second to false.
A Memgraph HA client connects to a coordinator, asks for the routing table, and follows the addresses in it — the MAIN for writes, a replica for reads. Those addresses are the ones each member was registered with, so exposing the cluster means two things: a way in, and a routing table clients outside can follow. spec.externalAccess does both. The operator creates the external objects, registers every exposed member at the address they acquire, and moves it with UPDATE CONFIG as that address appears, changes or goes away. Nothing about the address is declared; it is discovered from the cloud, or from an external-dns hostname annotation when you set one.
LoadBalancer: one Service of type LoadBalancer shared by all coordinators, and one per data instance. Three coordinators and two data instances cost three cloud load balancers.
spec:
externalAccess:
type: LoadBalancerGateway: one Gateway API Gateway the operator creates, with a TCP listener shared by all coordinators on port 7687 and one per data instance on dataPortBase + ordinal, each fed by a TCPRoute. One address, one cloud load balancer, a port per instance. It needs Gateway API v1.6 or newer and a running Gateway controller — Envoy Gateway v1.9 is the first release bundling v1.6 — installed before the operator starts; the operator never bundles the CRDs.
spec:
externalAccess:
type: Gateway
gateway:
gatewayClassName: egEither way, wait for Converged and read the addresses off the status:
kubectl wait --namespace memgraph --for=condition=Converged memgraphcluster/memgraph --timeout=10m
kubectl get mgc memgraph -n memgraph -o jsonpath='{.status.externalAccess}'{"coordinators":"203.0.113.10:7687","data":[{"name":"instance_0","address":"203.0.113.11:7687"},{"name":"instance_1","address":"203.0.113.12:7687"}]}Then connect with the routing scheme to the coordinators' address alone; the driver fetches the routing table from there and follows it:
neo4j://203.0.113.10:7687
While a LoadBalancer or the Gateway has no address yet, the members behind it are announced at their pod addresses and the resource reports Converged=False with reason ExternalAddressPending, naming what it is waiting on; Ready is unaffected, the cluster serves in-cluster throughout. Removing the block deletes the external objects and moves every member back to its pod address. For the address rules, per-instance hostnames with {ordinal}, the Gateway API version requirement and how to upgrade the CRDs past Helm, and the lifecycle on scale and type switch, see docs/external-access.md.
The MVP is deliberately "provision, bootstrap, observe". It does:
- provision one StatefulSet and headless Service per role, with per-pod identity derived from the pod ordinal;
- bootstrap HA: add the coordinators, register the data instances, and promote the initial MAIN once;
- re-register continuously: every reconcile compares
SHOW INSTANCESon the coordinator leader against the declared topology and issues only the missing registrations, so an instance that loses its registration state (say, after being rescheduled onto a fresh node) rejoins without human action; - grow a live cluster: raise
coordinatorsordataInstances(both in one edit if you like, in any step size) and the added pods are provisioned and registered by the same diff that restores a lost registration — no manualADD COORDINATORorREGISTER INSTANCE; - shrink a live cluster: lower
dataInstancesorcoordinatorsand the members above the new count are retired before their pods are shed — a data instance has MAIN moved off it if it holds it and is thenUNREGISTER INSTANCEd, a coordinator isREMOVE COORDINATORed out of the Raft cluster — so the coordinators never expect an instance whose pod is gone, and no removed member's pod outlives its vote; - grow storage on a live cluster: raise a claim size and every claim of the role is grown under its running pod, the StatefulSet recreated around the pods with the new claim template, and
Convergedreports when the space is usable (seedocs/storage-resize.md); - expose the cluster outside Kubernetes, through LoadBalancers or a Gateway API Gateway, and keep the routing table pointing at the addresses clients reach it through (see External access);
- serve OpenMetrics from every instance on port 9091, declared on the pods and the headless Services and pinned to that format, so a Prometheus you already run scrapes the cluster with a ServiceMonitor, PodMonitor or scrape config of your own — and, on request, create the ServiceMonitor for a Prometheus Operator and the "Memgraph OpenMetrics" dashboard ConfigMap for a Grafana sidecar, or run the vmagent that pushes every instance's metrics to the VictoriaMetrics Memgraph runs and the Vector sidecars that push every instance's logs to its VictoriaLogs, so Memgraph can monitor the cluster without reaching into your network (any Prometheus remote-write and any Loki endpoint work), with
spec.monitoring(seedocs/monitoring.md); - serve Bolt and the metrics endpoint over TLS from a
kubernetes.io/tlsSecret you name inspec.tls.bolt, on both roles, with the operator's own dials and the ServiceMonitor following suit, and authenticate the members to each other over mutual TLS from a Secret withca.crtyou name inspec.tls.intraCluster, decided at creation (seedocs/tls.md); - keep the pods of a role on distinct nodes, softly by default, as hard as
requiredwith scoperole(the HA chart'sparity) orcluster(itsunique) on request, and pass a node selector, tolerations, topology spread constraints and a priority class through per role withspec.scheduling(seedocs/scheduling.md); - report the observed MAIN, the registered member counts, the external addresses, and the readiness and convergence conditions on the resource's status.
Scaling is one edit, and Converged tells you when it is finished:
kubectl patch mgc memgraph -n memgraph --type=merge -p '{"spec":{"coordinators":5,"dataInstances":3}}'
kubectl wait --namespace memgraph --for=condition=Converged memgraphcluster/memgraph --timeout=10mBoth counts have a floor the schema enforces at creation and on every update: coordinators must stay odd and at or above three, dataInstances at or above one. A scale-down reports Converged=False with reason RetirementInProgress, naming the members on their way out, until their pods are gone. Three things to know about it:
- A retiring pod that cannot become ready blocks its own removal. The operator only touches the cluster when every pod of both StatefulSets is ready, and until the shrink is applied the retiring pods still belong to their StatefulSet. So a member that is stuck (crash-looping, unschedulable, wedged in a snapshot restore) keeps its own retirement waiting, and the resource reports
WorkloadsNotReadyrather than the operator writing to a cluster whose state it only half knows. Fix the pod, or delete it if it is genuinely unrecoverable, and the retirement continues. - The claims of a retired member follow
spec.storage.retentionPolicy, the same knob that decides what happens to storage when the cluster is deleted —Retain(the default) keeps them, so a shrink made by accident loses no data, and re-raising the count reattaches them. A coordinator removed from Raft keeps running and keeps its state on purpose, which is what makes re-growing onto a retained volume safe: it is in the same position as one whose pod crashed and stayed down, and a laterADD COORDINATORbrings it back in. - A coordinator shrink may have to wait for a Raft election. Raft refuses to remove its own leader, and a StatefulSet sheds only its highest ordinals, so a leader sitting in the retiring range is asked to
YIELD LEADERSHIPfirst — which cannot name a successor. The resource reportsConverged=Falsewith reasonLeadershipTransferInProgresswhile that is pending, and the operator asks again if the election happens to pick another retiring coordinator.
What it does not do yet:
- Failover. The operator promotes a MAIN only when the cluster has none: once at bootstrap, and once more when it demotes an instance that is retiring. It never overrides a MAIN that is staying — leadership belongs to the Raft coordinators, so two control systems never fight over which instance is MAIN.
- Other day-2 operations: orchestrated or rolling version upgrades, storage-mode changes, and a backup feature of its own — volume-snapshot backup and restore work through the claims' names and the retention policy without one, see
docs/backup-restore.md. - Deleting storage: the operator owns no finalizer and runs no cleanup of its own — deleting a volume is left entirely to the StatefulSet's own retention policy.
- Liveness or startup probes. Pods carry a readiness probe only. Memgraph opens no port until every database is recovered, so a liveness check could only ever kill a recovery that outlived a guessed budget — and a recovery longer than the guess would never finish, because every kill starts it over. A recovering data instance shows
0/1until it is done and is restarted by nothing but its own exit; restarting a running instance is the rolling restart's job, which knows the cluster's state. - NodePort or ingress exposure, a declared hostname without external-dns, or attaching to a Gateway you already run — see what external access leaves out.
- TLS, for Bolt or intra-cluster traffic — the addresses external access announces are plain Bolt.
- Bolt authentication — the operator connects to the coordinators unauthenticated, so clusters must not enable auth yet.
- The rest of the HA chart's monitoring: the
mg-exporterand its JSON format, the logs dashboard ConfigMap, the Kubernetes infrastructure metrics the chart's vmagent can scrape alongside Memgraph's, and monitoring objects placed in another namespace — see what monitoring leaves out. (The operator itself serves controller-runtime metrics; see the chart README.) - Standalone (non-HA) topology. The API is shaped to grow one without a breaking change, but v1alpha1 provisions HA clusters only.
- Init containers, sidecars, snapshot-restore fields and the rest of the HA chart's surface — parity roadmap, not MVP.
This operator is the successor to the memgraph-high-availability Helm chart. The plan is to grow it to functional parity with the chart, publish a migration guide, and then freeze the chart (security fixes only) with a deprecation timeline. Until then the chart remains the supported way to run HA in production, and this operator is an alpha for evaluating the reconciliation core. The standalone memgraph and memgraph-lab charts are unaffected and continue independently.
Where a concept carries over, the operator borrows the chart's vocabulary — the secrets.name / secrets.licenseKey / secrets.organizationKey block is the chart's block — so translating a values file is mechanical. The topology is where they deliberately differ: the chart's per-instance blocks and StatefulSet-per-instance model become two integers and one StatefulSet per role.
Migration is fresh-cluster only. The operator will never adopt a chart-deployed cluster in place: the resources are shaped differently and the ownership handover cannot be made safe. Moving means standing up a new cluster and transferring the data (backup/restore, or a replication cutover). The step-by-step guide lands when the operator reaches parity — there is nothing to migrate to before then.
make test-unit # unit tests (pure packages)
make test # unit + envtest
make lint # golangci-lint
make run # run the controller locally against the current kubeconfig
make test-e2e # KinD end-to-end suite; creates and deletes its own Kind clusterEvery pull request runs lint, unit, envtest, chart and end-to-end suites. The e2e job boots a licensed Memgraph cluster on a multi-node Kind cluster, with the license coming from repository secrets; set MEMGRAPH_ENTERPRISE_LICENSE and MEMGRAPH_ORGANIZATION_NAME to run it locally. Kind has no cloud controller and no Gateway controller, so make setup-test-e2e also installs MetalLB, which hands LoadBalancer addresses out of the Kind Docker network, and Envoy Gateway with a GatewayClass named eg (hack/kind-metallb.sh, hack/kind-envoy-gateway.sh); the external access scenarios connect to those addresses from the test process, a genuine client outside the cluster. The envtest suite loads the Gateway API CRDs from the sigs.k8s.io/gateway-api module in go.mod.
Run a development build against a cluster with make docker-build docker-push IMG=<registry>/kubernetes-operator:tag followed by make install (CRDs) and make deploy IMG=<registry>/kubernetes-operator:tag, or install the local chart:
helm install memgraph-operator ./charts/memgraph-operator \
--namespace memgraph-operator-system --create-namespace --waitmake docker-build produces an image for the architecture of the machine running Docker. Building on an arm64 workstation (Apple Silicon, or an arm64 Docker VM) for an amd64 cluster such as AKS needs two adjustments, because a plain docker build --platform linux/amd64 fails with exec format error on any daemon without amd64 emulation: the builder stage would pull the amd64 Go image and try to execute it.
The Dockerfile already reads TARGETARCH, so the fix is to run the builder stage natively and let Go cross-compile. make docker-buildx does exactly that through a generated Dockerfile.cross, but it pushes straight from a separate BuildKit container and skips the local image store. To build the same way with the default builder and keep the image locally:
sed -e '1 s/\(^FROM\)/FROM --platform=\${BUILDPLATFORM}/; t' -e ' 1,// s//FROM --platform=\${BUILDPLATFORM}/' Dockerfile > Dockerfile.cross
docker build --platform linux/amd64 -f Dockerfile.cross -t <registry>/kubernetes-operator:<tag> .
rm Dockerfile.cross
docker image inspect <registry>/kubernetes-operator:<tag> --format '{{.Os}}/{{.Architecture}}' # expect linux/amd64
docker push <registry>/kubernetes-operator:<tag>The final stage is distroless/static and runs no commands, so it needs no emulation either.
Deploy the pushed image by digest, not by tag. The manager Deployment uses imagePullPolicy: IfNotPresent, so a node that has already pulled <tag> keeps its cached copy even after the tag is re-pointed at a different image. After an arm64 image has been pushed under a tag once, re-pushing an amd64 image under the same tag will still land the old binary on that node and the pod will crash-loop with exec /manager: exec format error. docker push prints the manifest digest; use it:
make install
make deploy IMG=<registry>/kubernetes-operator@sha256:<digest>
kubectl -n kubernetes-operator-system get pod -o jsonpath='{.items[0].status.containerStatuses[0].imageID}{"\n"}' # must show the same digestPushing under a fresh tag each time works too; the point is that a tag which once held a different image can no longer be trusted to pull.
make deploy rewrites config/manager/kustomization.yaml with the image you pass. Revert that file before committing.
The install chart is maintained in this repository under charts/memgraph-operator, next to the manifests it ships: its CRDs and the manager's RBAC rules are generated from the Go types and the +kubebuilder:rbac markers (make chart-sync, verified in CI by make chart-verify), so the chart can never drift from the controller version it installs. Running the Release workflow on main releases the operator image and tags the commit v<appVersion>; the chart is published to the memgraph.github.io/helm-charts index from memgraph/helm-charts, like every other Memgraph chart. The chart version and the operator version move independently. See docs/releasing.md.
Each feature's design is documented under docs/. Run make help for all targets, and see the Kubebuilder documentation for the scaffolding conventions this project follows.
Copyright 2026.
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.