reference/workerfleet is the app-local worker monitoring service built on Plumego.
It is not the canonical v1 reference application and is not a reusable stable
surface for Plumego applications. Use reference/standard-service for v1
application structure, route wiring, and stable-root-only onboarding.
Current status:
- workerfleet is a work-in-progress reference, not a production-readiness claim
- worker ingress, heartbeats, list/detail queries, fleet summary, alerts, and history-backed drilldown paths are implemented when their backing runtime dependencies are wired
- some app service methods still return
workerapp.ErrNotImplementedwhen the current reference profile omits the required ingest, query, or case-step dependency - treat the app as a partial reference for worker monitoring ideas, not as the canonical template or a fully complete operations product
Submodule boundary:
reference/workerfleetbuilds as theworkerfleetGo module- app-local packages import each other as
workerfleet/... - Plumego root packages are consumed through
replace github.com/spcent/plumego => ../.. - MongoDB dependencies stay scoped to this submodule and do not modify the repository root
go.mod
Current documents:
- API
- Storage
- Metrics
- Case And Step Metrics Design
- Grafana Dashboard Plan
- Kubernetes Deployment Assets
- Technical Design
- 技术方案设计
Run locally:
cd reference/workerfleet && go run .cd reference/workerfleet && go build .cd reference/workerfleet && WORKERFLEET_HTTP_ADDR=:9090 go run .- env.example includes the current workerfleet env surface for local bootstrapping
make workerfleet-mongo-testruns the optional MongoDB integration gate whenWORKERFLEET_MONGO_TEST_URIis set; CI invokes the same target and skips it unless the repository secret is configured
HTTP entrypoint configuration:
WORKERFLEET_HTTP_ADDRcontrols the listen address, default:8080WORKERFLEET_SHUTDOWN_TIMEOUTcontrols graceful shutdown timeout, default10s/metricsis registered on the same HTTP server as the workerfleet API/healthzreports process liveness without external dependency checks/readyzreports readiness based on initialized runtime dependencies
Worker ingress authentication:
WORKERFLEET_WORKER_AUTH_TOKENenables Bearer-token auth forPOST /v1/workers/registerandPOST /v1/workers/heartbeat- when unset, worker ingress auth is disabled for local development
- when set, missing, malformed, or invalid credentials fail closed with
401 WORKERFLEET_PROFILE=prodrequiresWORKERFLEET_WORKER_AUTH_TOKENat startup
Query API authentication:
WORKERFLEET_ADMIN_AUTH_TOKENenables Bearer-token auth for query endpoints such as/v1/workers,/v1/tasks/:task_id,/v1/fleet/summary, and/v1/alertsWORKERFLEET_QUERY_AUTH_REQUIRED=truerequiresWORKERFLEET_ADMIN_AUTH_TOKENeven outside the production profile- worker ingress auth and query API auth use separate tokens so worker pods cannot query fleet state unless explicitly granted the admin token
/healthz,/readyz, and/metricsare not covered by query API auth in this reference appWORKERFLEET_PROFILE=prodrequiresWORKERFLEET_ADMIN_AUTH_TOKENat startup
Runtime loop configuration:
WORKERFLEET_PROFILE=dev|prod, defaultdevWORKERFLEET_KUBE_SYNC_ENABLED, defaultfalseWORKERFLEET_STATUS_SWEEP_ENABLED, defaultfalseWORKERFLEET_ALERT_EVALUATION_ENABLED, defaultfalseWORKERFLEET_NOTIFICATION_ENABLED, defaultfalseWORKERFLEET_KUBE_SYNC_INTERVAL, default30sWORKERFLEET_STATUS_SWEEP_INTERVAL, default30sWORKERFLEET_ALERT_EVALUATION_INTERVAL, default30sWORKERFLEET_NOTIFIER_DELIVERY_TIMEOUT, default5sWORKERFLEET_LOOP_LEASE_TTL, default90sWORKERFLEET_LOOP_LEASE_OWNER, default host nameWORKERFLEET_EXPERIMENTAL_METRICS_ENABLED, defaulttrueindevandfalseinprod- when
WORKERFLEET_NOTIFICATION_ENABLED=true, at least one notifier URL must be configured WORKERFLEET_KUBE_API_HOSToptionally overrides in-cluster Kubernetes API discoveryWORKERFLEET_KUBE_BEARER_TOKENoptionally overrides service account token discoveryWORKERFLEET_KUBE_NAMESPACEcontrols the namespace watched by Kubernetes syncWORKERFLEET_KUBE_LABEL_SELECTORlimits watched worker podsWORKERFLEET_KUBE_WORKER_CONTAINERselects the worker container, defaultworker- runtime loop errors are exported as
workerfleet_runtime_errors_total - each runtime loop runs with built-in non-overlap, a default
25siteration timeout, a5sinitial failure backoff, and a1mmax failure backoff - when
WORKERFLEET_STORE_BACKEND=mongo, Kubernetes sync, status sweep, alert evaluation, and notification delivery use MongoDB-backed distributed leases through theloop_leasescollection - when
WORKERFLEET_STORE_BACKEND=memory, loop lease behavior remains process-local and the deployment should stay single-replica for enabled loops - the reference Kubernetes deployment stays at
replicas: 1by default even though Mongo-backed loop leases allow safe multi-replica ownership
Status and alert policy configuration:
WORKERFLEET_PROFILE=devuses the app-local development policy defaultsWORKERFLEET_PROFILE=prodswitches to more conservative production defaults for heartbeat, offline, stage-stuck, and restart-burst thresholdsWORKERFLEET_STATUS_STALE_AFTERoverrides worker heartbeat stalenessWORKERFLEET_STATUS_OFFLINE_AFTERoverrides worker offline detectionWORKERFLEET_STATUS_STAGE_STUCK_AFTERoverrides worker stage-stuck detectionWORKERFLEET_STATUS_RESTART_BURST_THRESHOLDoverrides restart-burst status evaluationWORKERFLEET_ALERT_STAGE_STUCK_AFTERoverrides alert-side stage-stuck firingWORKERFLEET_ALERT_RESTART_BURST_THRESHOLDoverrides alert-side restart-burst firing- policy values are validated at startup; unsafe low thresholds fail closed before the HTTP server is exposed
WORKERFLEET_EXPERIMENTAL_METRICS_ENABLEDoverrides the profile default for pod, exec-plan, and case-step experimental metric families
Single-cluster Kubernetes assumptions:
- the service reads pods from one Kubernetes cluster
- pod discovery is namespace-scoped
- label selectors are used to limit the watched worker pod set
- the service expects a service account token or an explicit bearer token
Minimum RBAC expectations:
get,list, andwatchon pods in the target namespace
Current monitoring model:
- one pod maps to one worker
- workers report the full active-task set on each heartbeat
- active cases can include
exec_plan_idandcurrent_step - case step timeline and exec-plan drilldown APIs expose case-level detail for Grafana links
- current state is stored separately from seven-day task, event, and alert history
Storage backend configuration:
WORKERFLEET_STORE_BACKEND=memory|mongomemoryis the default local backendmongorequiresWORKERFLEET_MONGO_URIandWORKERFLEET_MONGO_DATABASE- optional Mongo settings:
WORKERFLEET_MONGO_CONNECT_TIMEOUT,WORKERFLEET_MONGO_OPERATION_TIMEOUT,WORKERFLEET_MONGO_MAX_POOL_SIZE WORKERFLEET_RETENTION_DAYScontrols Mongoexpire_atgeneration for task history, worker events, and alerts; values must be greater than zero and no more than 106751 days- Mongo stores loop ownership in
loop_leaseswith anexpires_atlease expiry - MongoDB integration tests are skip-safe in default gates; set
WORKERFLEET_MONGO_TEST_URIlocally or configure the CI secret and runmake workerfleet-mongo-testbefore changing Mongo store behavior