muster β to assemble scattered members into formation before they act.
muster is a container entrypoint for clustered stateful workloads. Replicas start up not knowing who they are: muster discovers its peers through its cloud's service registry, coordinates with them to settle roles, computes the workload's command line, then starts and supervises the process. Discovery, coordination, argv, and the process lifecycle are all driven by a small Starlark script, so it can express the imperative decisions that clustered stateful systems (e.g. TiKV/PD) need at startup β am I bootstrapping a new cluster, joining an existing one, or restarting an existing member? β rather than just string interpolation.
The script is imperative and async. It defines one required main() function that drives everything and returns a promise representing the workload; the harness awaits that promise and, on SIGTERM/SIGINT, delivers a graceful-stop signal() to it. main() calls spawn(), which supervises the workload β it resolves the argv, runs it, respawns it, and drives script-supplied readiness/liveness promise factories β passing every callback by reference (resolve, pre_start, post_start, pre_stop, post_stop, readiness, liveness). Async primitives β go(), poll(), promise(), join(), select() β let main() coordinate (seed election, ordering, multiple workloads) and react to health imperatively. Builtins also expose service-registry discovery, the host's interface addresses, filesystem checks, HTTP/TCP/gRPC probes, replica-set preconditions, and a conditional-write key/value store. Those four are the cloud-facing ones, and each is provided by a provider chosen at build time: AWS by default, Google Cloud with -tags gcp.
For example, this script builds a --servers argument from the healthy instances of my-service and supervises the workload:
# servers.star
def resolver():
peers = instances("my-service")
urls = ["http://%s:%d" % (p.ipv4, p.port) for p in peers]
return COMMAND + ["--servers=" + ",".join(urls)]
def main():
return spawn(resolve=resolver, respawn=True)muster \
-namespace my-namespace \
-script servers.star \
-- my-executableNote: the previous
text/templateinterpolation mode and the declarativerespawn()/healthcheck()builtins have been removed in favor of this imperative model. See Migrating.
You can install muster using go install:
go install github.com/moriyoshi/muster@latest # AWS (the default)
go install -tags gcp github.com/moriyoshi/muster@latest # Google CloudThe provider is chosen at build time, not at runtime: one binary talks to one cloud. That is what keeps the default image carrying the AWS SDK and nothing else β the Google Cloud build is three times the size, and neither should pay for the other. CI asserts both directions.
Which provider a binary has is in its startup log, in the PROVIDER global, and
in muster -provider-help. Selecting one that was not compiled in says how to
rebuild rather than reporting an unknown name.
| AWS (default) | Google Cloud (-tags gcp) |
|
|---|---|---|
| Build | go build ./... |
go build -tags gcp ./... |
| Target runtime | ECS (Fargate or EC2) | Cloud Run (services, jobs, worker pools) |
instances() |
CloudMap DiscoverInstances |
Service Directory ResolveService |
register() / deregister() |
unsupported β ECS Service Connect registers tasks for you | Service Directory endpoints; worker pools only |
all_replicas_running() |
ECS RunningCount == DesiredCount |
unsupported β see below |
kv_* store |
a DynamoDB table | a Cloud Storage bucket, or a Firestore collection |
SELF |
ECS task metadata | Cloud Run's environment + the metadata server |
SELF.name |
empty β a task has no name that survives replacement | empty β nor does a Cloud Run instance |
SELF.ipv4 |
the task's own address | the instance's VPC address on a worker pool; empty on a service or job |
| Credentials | the AWS SDK default chain | Application Default Credentials |
Whether peers can address each other depends on the Cloud Run runtime, and it is the single thing that decides what a script can do:
- Worker pools support Direct VPC ingress:
each instance gets a private address on your VPC and other resources β other
instances included β can connect to it there.
SELF.ipv4is that address,register()publishes it, and a peer cluster is possible. - Services and jobs support Direct VPC egress only. The address such an
instance sends from is not one anything can send to: inbound requests go
through the Cloud Run frontend and are load-balanced.
SELF.ipv4is empty there rather than advertising an address nothing can connect to, andregister()refuses for the same reason.
So on a service or a job, instances() is for resolving your dependencies β
a database, a broker set, a cluster running elsewhere β not for assembling
peers. On a worker pool it is for either.
One caveat remains even on a worker pool, and it is the same one Fargate has:
nothing survives replacement. SELF.name is empty on all three runtimes, the
address is not preserved, and the ephemeral disk is "permanently deleted when
the instance shuts down" β so a cluster member name has to be derived from the
address and pre_stop has to evict the member, exactly as on ECS. Cloud Run's
shutdown budget for that is a fixed ten seconds, where ECS allows thirty.
That is workable β e2e/tikv/gcp runs a real TiKV cluster this
way β but a workload wanting an identity or a disk that outlives an instance
wants neither of these platforms.
all_replicas_running() raises on Google Cloud: Cloud Run deliberately does not
expose per-instance counts, and the only source is a Cloud Monitoring metric
delayed by minutes. The portable substitute is to count what you can see β
len(instances(svc, health_status="ALL")) >= n.
Selection is -provider, then $MUSTER_PROVIDER, then autodetection β
which covers both clouds, because both runtimes announce themselves in the
environment: ECS sets ECS_CONTAINER_METADATA_URI_V4, and Cloud Run sets
K_SERVICE, CLOUD_RUN_JOB or CLOUD_RUN_WORKER_POOL according to which of
the three you deployed. The Google Cloud end-to-end stack sets no
MUSTER_PROVIDER at all, so that this keeps working is one of the things it
tests.
Detection is environment-only and deliberately never dials the metadata server, which would cost a timeout on every container start on every other platform. That rules out autodetecting a runtime with nothing distinguishing in its environment β a bare Compute Engine VM, for one β and such a deployment has to name its provider explicitly.
A third provider, mem, is always compiled in and has to be asked for by name.
It backs kv_* with an in-process store and supports nothing else, so
muster -provider mem -kv-store x -script s.star -- cmd runs a script with no
cloud account at all β enough for seed election to behave while you iterate.
Region, credentials and endpoint come from the AWS SDK's own default chain:
-
AWS_REGION: The AWS region to use for service discovery. If not set, the default region from the AWS CLI configuration will be used. -
AWS_PROFILE: The AWS profile to use for service discovery. If not set, the default profile from the AWS CLI configuration will be used. -
AWS_ACCESS_KEY_ID: The AWS access key ID to use for service discovery. If not set, the default credentials from the AWS CLI configuration will be used. -
AWS_SECRET_ACCESS_KEY: The AWS secret access key to use for service discovery. If not set, the default credentials from the AWS CLI configuration will be used. -
AWS_SESSION_TOKEN: The AWS session token to use for service discovery. If not set, the default credentials from the AWS CLI configuration will be used. -
AWS_ENDPOINT_URL: The AWS endpoint URL to use. If not set, the default endpoint URL from the AWS CLI configuration will be used. Specifying this will effectively disable the endpoint prefixing behavior. (Thus the actual endpoint will end up being the same as the endpoint URL, in contrast todata-servicediscovery.<region>.amazonaws.comwhere the endpoint isservicediscovery.<region>.amazonaws.com.) The override applies to the ServiceDiscovery (CloudMap), ECS, and DynamoDB clients alike, so all three can be pointed at a single mock endpoint for testing.
IAM. The task role needs servicediscovery:DiscoverInstances on the namespace and ecs:DescribeServices on the service, plus the kv grants below.
Credentials are Application Default Credentials, which on Cloud Run means the
revision's service account with nothing to configure. Everything else is a
-provider-opt:
project=<id>β defaults to$GOOGLE_CLOUD_PROJECT(thenGCLOUD_PROJECT,GCP_PROJECT), then the metadata server.location=<region>β the Service Directory location. Defaults to this instance's own region.kv.backend=gcs|firestoreβ which store backskv_*. Defaultgcs. See below.kv.database=<id>β the Firestore database. Default(default); ignored by thegcsbackend.endpoint.storage=,endpoint.servicedirectory=β endpoint overrides for a fake server. Setting one also disables authentication for that client, since a fake has no credentials to detect and looking for them would hang. The Firestore backend has no equivalent option: its client readsFIRESTORE_EMULATOR_HOSTitself.
An unknown key is a startup error rather than a shrug, so -provider-help is the
list that matters β the one above is these two plus project, location,
kv.backend and kv.database, and nothing else.
IAM, granted on the resource rather than the project:
| Capability | Role |
|---|---|
instances() |
roles/servicedirectory.viewer on the namespace |
register() / deregister() |
roles/servicedirectory.editor on the namespace |
kv_* (kv.backend=gcs) |
roles/storage.objectUser on the bucket |
kv_* (kv.backend=firestore) |
roles/datastore.user, necessarily project-wide β Firestore has no per-database IAM |
-kv-create |
storage.buckets.create; the gcs backend only, see below |
Nothing registers a Cloud Run instance for you. Service Directory
auto-registration covers GKE Services, and Cloud Run has no equivalent of ECS
Service Connect, so a clustered workload on a worker pool has to announce
itself: register(service, port=β¦) from post_start, deregister() from
pre_stop while peers are still reachable.
Endpoints are keyed on the address, not the instance id: a Service Directory
endpoint id is at most 63 characters of [a-z0-9-] and a Cloud Run instance id
is some two hundred hex digits, so the obvious key does not fit. Keying on the
address means a replacement that lands on the same one inherits the entry rather
than adding a second β but an instance killed without teardown still leaves a
stale endpoint behind, because nothing else will withdraw it. Probe discovered
peers before trusting them, which is the portable habit anyway.
muster \
-namespace <namespace> \
-script <path> \
[-provider <name>] [-provider-opt k=v ...] \
[-kv-store <name> [-kv-key-prefix <prefix>] [-kv-create]] \
[-allow-run] \
[-control-socket <path>] \
-- [command...]The trailing command... after -- is optional; it is passed to the script as the COMMAND global so a script can transform a base command rather than build argv from scratch.
Every flag can also be given as an environment variable, MUSTER_<FLAG> with the dashes turned into underscores β -namespace is MUSTER_NAMESPACE, -kv-store is MUSTER_KV_STORE. An explicitly passed flag always wins. muster is a container entrypoint, and a container's entrypoint is baked into its image while its environment is per deployment; an image whose entrypoint ends in -- <workload> cannot take extra flags by appending arguments at all, so on a platform that only lets you set environment and arguments there would otherwise be no way to pass one. The two mode switches (-health-probe, -provider-help) are deliberately excluded, so a stray variable cannot turn an ordinary run into a probe.
-namespace-
Required. The default namespace for the
instances()builtin (overridable per call via itsnamespaceargument).
> Supervision (respawn) and healthchecking (readiness/liveness) are configured **from the script** as `spawn()` arguments β there are no CLI flags for them. Retrying discovery until instances appear is also scripted: `instances()` is a single lookup that returns an empty list when nothing matches; wrap it in `poll` (join it for a synchronous wait). See [Scripting](#scripting).
-control-socket-
Path to a unix-domain socket on which the harness serves
GET /health(JSON). Exposes harness-level state (uptime) and a per-workload array (up/pid, respawn count, current backoff, health) β one entry perspawn(). Empty (the default) disables it. See Control socket and health probe. -health-probe-
Run the binary as a health-probe client instead of the harness: it connects to
-control-socket, and exits0when healthy or non-zero otherwise. Intended for use as a containerHEALTHCHECK CMD. See Control socket and health probe. -script-
Required. Path to the Starlark script that defines
resolve()and the optional lifecycle hooks. See Scripting. -provider-
Which cloud provider to talk to. Empty (the default) reads
$MUSTER_PROVIDER, then autodetects from the environment. Only providers compiled into the binary can be selected; naming one that was not built in says how to rebuild.-provider-helplists what this binary has. -provider-opt-
Provider-specific setting as
k=v, repeatable. Keys the selected provider does not declare are rejected at startup, so a typo fails rather than being silently ignored.-provider-helplists the keys each compiled-in provider accepts. -kv-store-
Name of the provider-side store backing the
kv_*builtins β a DynamoDB table on AWS. Empty (the default) disables the key/value store; calling akv_*builtin then raises an error. See Key/value store. -kv-key-prefix-
Prefix prepended to every
kv_*key, so multiple clusters can share one store without colliding. Default is empty. -kv-create-
Create the kv store with provider defaults if it does not already exist (on AWS: a DynamoDB table with on-demand billing and TTL enabled on
expires_at). Default is off; normally the store is provisioned out of band (Terraform/CDK), and a provider that cannot create its own store says so rather than starting without one. -allow-run-
Expose the
run([...])builtin, which executes an arbitrary command. Off by default to keep the sandbox tight; preferhttp_requestwhere possible. [command...]-
Optional trailing command after
--, exposed to the script as theCOMMANDglobal (a list of strings).
The workload lifecycle is defined by an imperative Starlark script (a small Python dialect). The only required global is main(), which returns a promise; the harness awaits it and, on SIGTERM/SIGINT, delivers a graceful-stop signal() to it. Every other function is ordinary β wire the ones you need into spawn() by reference.
main() does any up-front coordination and returns the workload promise from spawn():
def resolver(): # per-attempt: return argv (list) or {"argv":[...], "env":{...}}
return COMMAND
def readiness(): # a Promise that resolves when the workload is ready
return poll(http_ok("http://127.0.0.1:8080/ready"), "60s")
def on_stop(): # runs once on teardown
kv_delete("my-cluster/seed", if_value=SELF.ipv4)
def main():
# coordinate here (seed election, ordering) β¦
return spawn(resolve=resolver, pre_stop=on_stop, readiness=readiness, respawn=True)spawn(...) owns the respawn loop and the readiness/liveness probes. Its arguments:
resolve(required)-
A function returning the workload argv each attempt β a list of strings, or
{"argv":[...], "env":{...}}. Runs before every (re)spawn, so a clustered workload can re-decide its role (restart / join / bootstrap) after a crash. name-
Optional label for this workload in the control-socket snapshot (auto-assigned
workload-<n>by position if omitted). Give each workload a distinct name whenmain()runs more than one. pre_start/post_start/pre_stop/post_stop-
Optional lifecycle callbacks around each attempt:
pre_startβ before the process starts. The only hard gate: an error aborts the attempt (respawns if enabled, else fails). Use it to wait on dependencies.post_startβ right after the process is up (PID known). Best-effort (a failure is logged, not fatal β gate on being ready withreadiness). Use it to record/notify a just-launched process.pre_stopβ when the workload is torn down (shutdown, liveness-restart, or exit), while the process is still alive. Runs under a detached, time-bounded context (pre_stop_timeout, elseshutdown_grace) so it can deregister while peers are still reachable.post_stopβ after the process has fully exited. Best-effort, detached/time-bounded likepre_stop. Use it for cleanup that needs the process gone (pid files, final deregistration).
readiness/liveness-
Optional functions returning a promise (see below), called by
spawn()per attempt. A truthyreadinessmarks the workload ready (control socket); if it never resolves truthy the attempt is restarted. Whenlivenesssettles truthy (liveness lost) the workload is marked unhealthy and, ifrestart_on_liveness=True(default), restarted. Intervene (register/deregister) as a side effect inside these callbacks. - respawn policy
-
respawn=False,keep_alive=False,max_retries=5(0 = unlimited),initial_interval="1s",max_interval="60s",multiplier=2.0,reset_after="30s",shutdown_grace="10s",pre_stop_timeout=0,resolve_timeout=0,resolve_failure="retry",restart_on_liveness=True. Withrespawn=True, a non-zero exit restarts with jittered exponential backoff up tomax_retries;keep_alive=Truealso restarts a clean exit; the counter resets after the workload stays upreset_after.Under an orchestrator, prefer a bounded
max_retries.max_retries=0is right for a bare process supervisor, but on ECS or Kubernetes the scheduler is itself the outer retry loop, and retrying forever keeps that loop from ever running. Aresolve()that cannot succeed β a denied API call, peers that never appear β then leaves the container up as PID 1 having never started its workload, which reads as a hang rather than a failure: the task staysRUNNING, and nothing acts on it until a health check's grace period expires. Exhausting the retries instead exits non-zero, so the scheduler replaces the task and a persistent fault surfaces promptly.reset_afterclears the counter once an attempt actually stays up, so a bound only fires on consecutive failures.
The returned promise is signallable: w.signal() requests a graceful stop (stop respawning β pre_stop β SIGTERMβshutdown_graceβSIGKILL) and resolves with a {code, respawn_count, reason} struct. Retries exhausted β the promise rejects and the process exits non-zero.
Async primitives let main() coordinate and react. A promise is a settle-once future.
go(fn, *args) -> promise: runfn(*args)on a background task; resolves with its return, rejects on error. The general async primitive. (fnand its args are frozen β read captured state, don't mutate after launch.)poll(check, timeout, interval=1s) -> promise: resolvesTruewhencheck()becomes truthy,Falseon timeout, rejects on acheckerror. The idiomatic way to buildreadiness/liveness. For a synchronous wait, join it:join(poll(check, "60s"))blocks the current task and returns the bool.promise() -> promise: a bare deferred you settle yourself viap.signal(value)/p.reject(err).join(*p) -> value|list: await all; returns the value (or a list for many); raises on any rejection or cancellation.select(*p) -> promise: race; returns the first-settled promise (compare by identity:if select(a, b) == a).any_true(*p) -> bool: race for the first promise to resolve truthy βTrue(cancelling the rest);Falseonce all resolve falsy; raises on rejection. Short-circuiting concurrent predicate β e.g. "is any peer up?":any_true(*[go(http_ok(PD(p))) for p in peers]).
A task that rejects with nothing to join it is logged: join(), select()
and any_true() all count as observing the outcome, and a failure none of them
ever collects would otherwise be discarded in silence. This is what makes a
background go() loop safe to start β a loop that dies on its first iteration
says so, instead of looking exactly like a loop with nothing to report. A task
still running when muster begins shutting down is not reported: it was stopped,
not failed.
Since Starlark cannot catch a raise, containing a failure is structural: run
the fallible part on its own task and await it with select(), which unlike
join() does not re-raise. That is the only way to express "skip this and carry
on", and it is what a pre_stop doing several independent things needs β a peer
that has already gone must not cost you the deregistration after it.
Every promise has p.done(). Cancellable promises (go/poll) add p.cancel() (abort and discard β resolves None). Signallable promises (spawn(), bare promise()) add p.signal(value) / p.reject(err). go/poll accept signallable=True; promise() accepts cancellable=/signallable=.
COMMAND: the trailing--command as a list of strings.PROVIDER: the compiled-in provider's name ("aws"/"gcp"/"mem"). Always set, including whenSELFisNoneβ which is exactly when a script needs it.SELF: this instance's identity β.id(unique, stable for its lifetime; also the kv lease owner),.name(stable across replacement, empty on ECS, which has no such thing β derive a member name from.ipv4there instead),.group,.service,.zone,.region,.network,.ipv4,.ipv6,.created_at, plus a sub-struct named for the provider βSELF.awscarrying.task_arn,.cluster,.family,.revision,.vpc_id;SELF.gcpcarrying.project,.instance_id, and whichever of.service,.revision,.configuration,.job,.executionand.task_indexthe runtime sets. Reading the wrong one raises, so anything under that name is a visible declaration that the script is not portable.Nonewhen the platform gave muster nothing, soif SELF:is a real guard. Read once at startup. Individual fields can be empty even on a supported platform β notably.networkon Fargate, which AWS only populates for tasks on EC2 container instances, and.created_aton Cloud Run, which the metadata server does not carry.
Discovery / networking
instances(service, health_status=None, namespace=None): one service-registry lookup β list of structs with.ipv4,.ipv6,.port. No internal retry; an empty result is an empty list β wrap inpoll(join it) for a quorum.namespaceoverrides-namespace;health_statusoverrides the defaultHEALTHY(HEALTHY/UNHEALTHY/ALL/HEALTHY_OR_ALL); a provider that cannot honour a filter raises rather than widening it.ifaddr(cidr): the host's own address withincidr.SELF.ipv4usually says the same thing without needing the CIDR.register(service, port=None, address=None, namespace=None)/deregister(): publish this instance into the registry, for platforms that do not do it themselves. Both raise where the platform already registers for you (AWS), or where the instance has no address anything can reach (a Cloud Run service or job), so a script is told rather than silently doing nothing. Address defaults toSELF.ipv4. See Providers.
Replica-set preconditions β group/service default to SELF.group/SELF.service:
all_replicas_running(group=None, service=None)β bool: every desired replica is running (on AWS, ECSRunningCount == DesiredCount). Gate on it inpre_start:if not join(poll(all_replicas_running, "60s")): fail(...). It is the scheduler's view, not "the workload inside is up" β treat it as advisory and probe peers yourself.
Filesystem (read-only): path_exists(path), read_file(path) (β€1 MiB).
Health / HTTP
http_ok(url, timeout=2s),tcp_ok(hostport, timeout=2s),grpc_ok(hostport, service="", timeout=2s)β a check factory: each captures its target and returns a zero-arg callable that runs the probe and returns a bool. Pass it straight topoll/go(poll(http_ok(url), "60s")) or call it for the bool inline (if http_ok(url)(): ...).un(fn, *args)β a check factory computingnot fn(*args). Negates a probe without a lambda:poll(un(http_ok(url, "2s")), "24h")resolves when the target stops being live.http_request(method, url, body=None, headers={}, timeout=30s)β{status, body}(e.g. a PD member delete via HTTPDELETE).
JSON: json.decode(str), json.encode(value), json.indent(str, prefix="", indent="\t") β starlark-go's own module. Pair it with http_request when a decision depends on what an API said rather than merely whether it answered: json.decode(r.body)["stores"]. Substring-matching a response body is the alternative and it passes for the wrong reasons β '"state_name":"Up"' in body is true of a healthy cluster and of one where a single store is up and the rest are down.
Key/value store β see below.
Misc: env(name, default=None) β str (reads the harness's own environment), log(msg, **kwargs), sleep(seconds) (cancellable), rand() β float [0.0, 1.0), randint(a, b) (inclusive, like random.randint), and run([...]) β {code, stdout, stderr} (only with -allow-run).
Durations accept a number of seconds or a string like "5s" / "500ms". Blocking builtins are cancelled when the harness shuts down.
When -kv-store is set, the kv_* builtins provide a conditional-write key/value store backed by the provider. Atomic put-if-absent and compare-and-swap with TTL leases make it a distributed lock, which is what lets exactly one node win seed election during a cold start (avoiding split-brain).
kv_put_if_absent(key, val, ttl=None)β bool. Writes only if the key is absent or its lease has expired. The seed-election primitive.kv_compare_and_swap(key, old, new, ttl=None)β bool.kv_get(key)β str orNone. To wait for a key, compose withpoll:join(poll(lambda: kv_get(k) != None, "120s")), thenkv_get(k).kv_delete(key, if_value=None)β bool.kv_renew(key, ttl)β bool. Extends a lease you own;Falsemeans the lease was lost.
A ttl of None/0 is a permanent key; a positive ttl is a lease that expires (and frees the lock) if the holder dies.
Expiry is filtered when a key is read, on every backend. Both DynamoDB TTL and
Cloud Storage lifecycle rules delete lazily β hours or days late β so neither
can be trusted to have reaped a lapsed lease. Use -kv-key-prefix to share one
store across clusters.
Every backend runs the same conformance suite, so they agree on the semantics that matter, down to the corners: a permanent key is never claimable and never renewable, and a conditional delete matches on value alone so a script can release its own lease even if it lapsed while the process was shutting down.
Table schema. Partition key pk (String), value val (String), and expires_at (Number, Unix epoch seconds) configured as the table's TTL attribute. Provision it out of band, or pass -kv-create to create it (on-demand billing, TTL enabled) on first run.
IAM. The task role needs dynamodb:GetItem, PutItem, UpdateItem, and DeleteItem on the table. With -kv-create, also CreateTable, DescribeTable, and UpdateTimeToLive.
Two backends, chosen with -provider-opt kv.backend=. They satisfy the same
conformance suite, so the choice is about what you are willing to provision, not
about semantics.
Cloud Storage (gcs, the default). One object per key under leases/, the
value in the body and the lease in custom metadata (owner, ttl_ms).
Atomicity comes from generation preconditions, which makes compare-and-swap
immune to a value being changed and changed back between a script's read and its
write β something the DynamoDB backend's value comparison cannot see.
-kv-store is the bucket. Provision it out of band, or pass -kv-create (which
additionally needs storage.buckets.create, and a project/location it can
resolve). A created bucket gets uniform access, soft delete switched off β it
defaults to seven billed days, and a lease object is rewritten on every renew β
and a one-day lifecycle rule scoped to leases/ as a janitor.
IAM. roles/storage.objectUser on the bucket.
Firestore (firestore). One document per key, and each operation is one
transaction β so put-if-absent and compare-and-swap read as the semantics are
stated rather than as an encoding of them, and cost one round trip instead of
two. Keys contain /, which a document id may not, so they are percent-escaped.
-kv-store is the collection and kv.database the database. There is no
-kv-create for it: a database is a long-running operation and a project-level
decision, not something a container entrypoint should make on your behalf.
IAM. roles/datastore.user. Firestore has no per-database IAM, so this is
necessarily project-scoped β which is a point in Cloud Storage's favour if least
privilege matters to you.
Expiry is filtered on read by every backend. Firestore's native TTL deletes lazily β typically within 24 hours β as does DynamoDB's, so none of them can be trusted to have reaped a lapsed lease.
# pd.star β run with: -kv-store <table> -allow-run -- /pd-server
PD = lambda ip: "http://%s:2379" % ip
def resolver():
me = SELF.ipv4
peers = [i.ipv4 for i in instances("tikv-pd") if i.ipv4 != me]
if path_exists("/pd/member"): # (A) restart existing member
mode = []
elif peers and any_true(*[go(http_ok(PD(p) + "/pd/api/v1/members")) for p in peers]):
mode = ["--join", ",".join([PD(p) for p in peers])] # (B) cluster exists β join (probes race)
elif kv_put_if_absent("tikv-pd/seed", me, "90s"): # (C) win the lease β bootstrap
mode = []
else: # someone else is seeding β wait, then join
join(poll(lambda: kv_get("tikv-pd/seed") != None, "120s"))
seed = kv_get("tikv-pd/seed")
join(poll(http_ok(PD(seed) + "/pd/api/v1/health"), "120s"))
mode = ["--join", PD(seed)]
return COMMAND + ["--advertise-client-urls", PD(me),
"--advertise-peer-urls", "http://%s:2380" % me] + mode
def liveness(): # settles when PD stops being live
me = SELF.ipv4
return poll(un(http_ok(PD(me) + "/pd/api/v1/health", "2s")), "24h", interval="10s")
def on_stop():
me = SELF.ipv4
run(["/pd-ctl", "-u", PD(me), "member", "delete", "name", SELF.id])
kv_delete("tikv-pd/seed", if_value = me)
def main():
return spawn(resolve=resolver, liveness=liveness, pre_stop=on_stop,
respawn=True, max_retries=5) # bounded: give up β ECS replaces the taskA worked version of this β three PD replicas and three TiKV stores, with the
seed election exercised against a real cold start β is in
e2e/tikv/aws.
muster now runs on more than one cloud, and the script surface no longer names AWS services. The old names are gone, not deprecated: Starlark resolves free names at compile time, so a script using them fails to load β naming the missing symbol β before the workload starts.
| Old | New | |
|---|---|---|
TASK |
SELF |
|
TASK.task_arn |
SELF.id |
the unique, lifetime-stable id of this instance; also the kv lease owner |
| β | SELF.name |
an id that survives replacement, safe to persist as a cluster member name. Empty on ECS, which has none |
| β | SELF.ipv4 / SELF.ipv6 |
this instance's own addresses, so ifaddr(cidr) is no longer needed to find them |
TASK.cluster |
SELF.group |
"cluster" now means only the workload's own cluster |
TASK.service_name |
SELF.service |
matches the service= kwarg |
TASK.availability_zone |
SELF.zone |
SELF.region is new |
TASK.vpc_id |
SELF.network |
|
TASK.created_at |
SELF.created_at |
unchanged |
TASK.family, .revision |
SELF.aws.family, .revision |
AWS-only; absent on other providers |
| β | PROVIDER |
"aws"; set even when SELF is None |
all_ecs_tasks_running(cluster=β¦, service=β¦) |
all_replicas_running(group=β¦, service=β¦) |
argument order is unchanged |
health_status="HEALTHY_OR_ELSE_ALL" |
health_status="HEALTHY_OR_ALL" |
HEALTHY/UNHEALTHY/ALL are unchanged |
-kv-table <name> |
-kv-store <name> |
the store is a table only on AWS |
-kv-create-table |
-kv-create |
-namespace, -kv-key-prefix, instances(), and everything about spawn() and the promise primitives are unchanged.
SELF.idis byte-identical toTASK.task_arnon AWS, so migrating to it changes nothing there. A script that parses it (arn:aws:β¦) is provider-locked and should say so by readingSELF.aws.task_arninstead.- Do not switch a persisted member name to
SELF.nameon a live cluster. etcd, PD and Consul remember member names; changing the derivation orphans the existing member. It is empty on ECS anyway β keep deriving fromSELF.ipv4there. SELFisNonewhen the platform gave muster nothing, exactly asTASKwas, so anif SELF:guard still works.PROVIDERis separate precisely so a script can still branch on the platform in that case.- The script and the flags fail differently, and the flags are the risk. A renamed builtin is a load-time error, so muster exits before spawning anything. A renamed flag is caught by the flag parser, which exits non-zero as PID 1 β but the script ships inside your image while
-kv-tablelives in a task definition, so the flags are the half that can lag the binary. Update the task definition and the image in the same deploy. Passing a removed flag names its replacement rather than reporting an undefined flag.
The text/template interpolation and the declarative resolve()/respawn()/healthcheck() model have been removed. In the current model:
- Define
main()(the only required global); it returns a promise, usually fromspawn(). - Pass the argv provider and lifecycle callbacks to
spawn()by reference (any names):resolve,pre_start,post_start,pre_stop,post_stop,readiness,liveness. - The old
respawn(...)knobs are nowspawn(...)arguments. - The old
healthcheck(probe=...)becomes areadiness/livenessfunction returning a promise (typicallypoll(...));spawn()drives it and restarts on liveness loss (restart_on_liveness). - List comprehensions replace the old
extract/mapprintf/join/excludetemplate funcs (["http://%s:%d" % (p.ipv4, p.port) for p in instances("s")]); env goes in the resolve dict{"argv":[...], "env":{...}}.
With -control-socket <path>, the harness serves GET /health (JSON) over a unix-domain socket. Because main() can spawn() more than one workload, the payload carries harness-level state plus a workloads array β one entry per spawn(), each with its own process, respawn, and health fields:
{
"harness": { "running": true, "started_at": "β¦", "uptime_seconds": 123.4 },
"workloads": [
{ "name": "pd", "up": true, "pid": 4321, "started_at": "β¦", "uptime_seconds": 100.2,
"respawn_count": 2, "current_backoff_seconds": 0, "max_retries": 5,
"last_exit_code": 0, "last_exit_error": "",
"health": { "state": "healthy", "consecutive_ok": 3, "consecutive_fail": 0,
"last_probe_at": "β¦", "last_probe_error": "" } }
]
}Each workload's name is its spawn(name=β¦) argument (auto-assigned workload-<n> by position if omitted). A workload's health.state is healthy/unhealthy/unknown (unknown = no readiness/liveness configured).
The same binary can act as a probe client with -health-probe, which queries the socket and maps state to an exit code (0 = healthy, non-zero = unhealthy or unreachable). The container is reported healthy only when every workload is healthy β a workload with no health probe falls back to being up, and any unhealthy or down workload fails the probe. This makes it usable directly as a container healthcheck command:
ENTRYPOINT ["muster", "-namespace", "my-namespace", \
"-script", "/workload.star", \
"-control-socket", "/run/muster.sock", \
"--", "my-workload"]
HEALTHCHECK CMD ["muster", "-health-probe", "-control-socket", "/run/muster.sock"]The probe client uses a short timeout, so a missing socket or dead harness reports unhealthy promptly rather than hanging.
go test ./... # the default (AWS) build
go test -tags gcp ./... # the Google Cloud buildThese cover seed election and the lifecycle hooks against the in-memory kv store, and run the kv conformance suite β the same set of assertions every backend has to satisfy, so a second implementation cannot quietly disagree about when a lease has expired or who is allowed to renew it. The Google Cloud store runs it against an in-process fake Cloud Storage that implements the generation preconditions its correctness rests on; it caught two real bugs on the way in.
The Firestore store runs the same suite against Google's emulator, which is a
Java process β so it is skipped unless FIRESTORE_EMULATOR_HOST is set. Docker
supplies the runtime if you would rather not install one:
docker run -d --name muster-fs-emu -p 8484:8484 \
gcr.io/google.com/cloudsdktool/google-cloud-cli:emulators \
gcloud emulators firestore start --host-port=0.0.0.0:8484
FIRESTORE_EMULATOR_HOST=127.0.0.1:8484 go test -tags=gcp -race ./internal/provider/gcp/There is no Service Directory emulator, so that path is covered by fakes at the interface only.
The Cloud Storage store can also be run against a real bucket, which is where contention and retries actually happen. It costs a few cents and is the one part of the Google Cloud provider subtle enough to be worth it:
MUSTER_GCP_KV_BUCKET=my-bucket \
go test -tags=gcp,gcp_live ./internal/provider/gcp/An opt-in end-to-end test exercises the DynamoDB-backed kv store and seed election against a real, stateful AWS emulator β WinterbΓ€ume in its standalone winterbaume-server mode, which speaks the DynamoDB, ECS, and CloudMap APIs over one HTTP endpoint. It is build-tag gated and skipped unless an endpoint is provided:
winterbaume-server & # or any DynamoDB-compatible endpoint
AWS_ENDPOINT_URL=http://127.0.0.1:8080 AWS_REGION=us-east-1 \
go test -tags=e2e -run E2E ./...It runs the same conformance suite as the unit tests, against the real store.
A second, heavier pair of end-to-end suites runs a real TiKV cluster β three PD replicas and three stores, with muster as the entrypoint β on throwaway infrastructure each provisions with Terraform.
e2e/tikv/aws runs it as Fargate tasks on ECS. It is
the only test that can see a split brain: it asks every PD replica, on its own
loopback via ECS Exec, and requires them to agree on one cluster id, then stops
a PD task and checks that its replacement joins rather than bootstrapping a
second cluster. Nothing in the stack is reachable from outside the VPC. It
creates billable AWS resources and is gated behind both a build tag and an
environment variable:
cd e2e/tikv/aws && make e2e # provision, assert, tear downe2e/tikv/gcp is the Google Cloud counterpart: the same cluster
on Cloud Run worker pools β the one runtime whose instances are addressable
at their VPC address, and so the only one that can host a Raft group. Nothing in
that stack is reachable from outside the VPC and there is no bastion, so each PD
replica reports its own view on a loop and the test reads those back: asking
every replica about itself is what makes the split-brain check exhaustive,
the same property the ECS suite gets from shelling into each task.
It makes the replacement claim too, by scaling the PD pool to two and back to
three: an instance stops, and its pre_stop has to evict its member from the
Raft group inside Cloud Run's fixed ten-second shutdown budget β the one
platform difference that could have made a clustered workload impossible here β
then the replacement has to join the same cluster rather than bootstrap a new
one. Both halves pass.
cd e2e/tikv/gcp && make e2eBoth suites' Starlark scripts are loaded and checked by the ordinary go test ./... run, so a syntax error surfaces without a cloud account.