Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 10 additions & 2 deletions .github/workflows/cd.yml
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,8 @@ jobs:
- run: python -m pip install -e ".[dev]"
- name: Verify the serving configuration matches the catalog
run: python -m llm_router.serving > /tmp/ray-serve.yaml && diff -u config/ray-serve.yaml /tmp/ray-serve.yaml
- name: Verify the Ray topology matches the catalog
run: python -m llm_router.topology > /tmp/ray-service.yaml && diff -u deploy/overlays/ray/ray-service.yaml /tmp/ray-service.yaml
- name: Render the canary and rollback plans
# One plan per track (model, adapter, policy), each naming what it
# rolls back to and the criteria that trigger it.
Expand All @@ -73,13 +75,19 @@ jobs:
run: |
curl -sSLo kubeconform.tar.gz https://github.com/yannh/kubeconform/releases/download/v0.7.0/kubeconform-linux-amd64.tar.gz
tar xf kubeconform.tar.gz kubeconform
./kubeconform -strict -ignore-missing-schemas -summary deploy/kubernetes
# Both topologies are rendered the way they would be applied: the
# single-engine base, and the Ray overlay that replaces the engine.
kubectl kustomize deploy/kubernetes > rendered-base.yaml
kubectl kustomize deploy/overlays/ray > rendered-ray.yaml
./kubeconform -strict -ignore-missing-schemas -summary rendered-base.yaml rendered-ray.yaml
- uses: actions/upload-artifact@v7
with:
name: deployment-plan-${{ env.RELEASE_REF }}
path: |
canary-plan.json
governance-plan.json
config/ray-serve.yaml
deploy/kubernetes
rendered-base.yaml
rendered-ray.yaml
deploy
retention-days: 14
42 changes: 40 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -207,9 +207,46 @@ unprivileged workloads, digest-pinned images, bounded resources, real probes, GP
pinning, and `/metrics` reachable only from monitoring.

```bash
kubectl apply -k deploy/kubernetes
kubectl apply -k deploy/kubernetes # one vLLM engine
kubectl apply -k deploy/overlays/ray # Ray Serve across GPU pools
```

The base runs a single vLLM engine, which serves one model. The
[`deploy/overlays/ray`](deploy/overlays/ray) overlay replaces it with a KubeRay `RayService`
that serves every local model in the catalog behind one OpenAI-compatible endpoint:

- A head that schedules and never runs a model.
- One worker group per accelerator type, pinned by `nvidia.com/gpu.product`, so GPU pools stay
separate. Each group is sized from the autoscaling bounds of the models placed on it, and a
worker holds as many GPUs as the largest replica on its pool needs.
- Model weights are loaded from the location [governance](#governance-in-mlflow) records for
each revision.

The `RayService` is generated from the catalog, never hand-edited, and verified by a test and
by CD:

```bash
python -m llm_router.topology > deploy/overlays/ray/ray-service.yaml
```

Under the overlay, engine metrics are scraped from the Ray pods by Prometheus. The gateway
cannot read a whole cluster from one address, so its own `router_engine_*` and `router_gpu_*`
gauges stay empty there and live load does not influence routing.

GPU support comes from the NVIDIA GPU Operator, installed cluster-wide with
[`deploy/gpu-operator/values.yaml`](deploy/gpu-operator/values.yaml). It provides the
`nvidia.com/gpu` resource, the node label the pools select on, and the DCGM GPU exporter.

Two Grafana dashboards in [`deploy/kubernetes/dashboards`](deploy/kubernetes/dashboards) ship as
a labelled ConfigMap for Grafana's sidecar: one for the gateway and router, one for engines and
GPUs. A `PrometheusRule` alerts on latency, load shedding, a stuck queue, fallback rate, canary
rollback, an open engine circuit, KV-cache pressure, and GPU memory. A test fails if a dashboard
or alert queries a metric the gateway does not publish.

None of this has been applied to a cluster. The manifests are schema-validated and
contract-tested; the Ray image, GPU product labels, and node sizes are placeholders to set for
the hardware you have.

Stateless ingress scales separately from GPU replicas. Set `ROUTER_REDIS_URL` so cache and
quota state are shared once the gateway runs more than one replica; without it both are
in-process and correct for a single replica only. Install the client with the extra:
Expand All @@ -220,7 +257,8 @@ python -m pip install -e ".[redis]"

CD renders the canary plans (one per track, each with its rollback target) and the governance
plan, verifies
`config/ray-serve.yaml` against the catalog, and validates the manifests with kubeconform.
`config/ray-serve.yaml` and the Ray topology against the catalog, and validates both rendered
topologies with kubeconform.
Applying to a cluster stays disabled until a deployment destination is configured.

## Model registry
Expand Down
43 changes: 43 additions & 0 deletions deploy/gpu-operator/values.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
# Values for the NVIDIA GPU Operator chart (nvidia/gpu-operator). The operator
# is cluster-wide infrastructure, installed once and not part of this
# platform's namespace:
#
# helm upgrade --install gpu-operator nvidia/gpu-operator \
# --namespace gpu-operator --create-namespace \
# --values deploy/gpu-operator/values.yaml
#
# It supplies the three things the serving manifests rely on.

# 1. The nvidia.com/gpu resource that engine and Ray worker pods request.
driver:
enabled: true
toolkit:
enabled: true
devicePlugin:
enabled: true

# 2. The nvidia.com/gpu.product node label that pins each pool to one
# accelerator type. Node feature discovery finds the hardware and GPU
# feature discovery turns it into the label.
nfd:
enabled: true
gfd:
enabled: true

# 3. The GPU exporter. DCGM publishes the DCGM_FI_DEV_* series the dashboards
# and alerts read, and the ServiceMonitor hands them to Prometheus.
dcgmExporter:
enabled: true
serviceMonitor:
enabled: true
interval: 15s
additionalLabels:
release: prometheus

# GPU nodes are tainted so only inference lands on them; the operator's own
# daemons have to tolerate that taint to run there at all.
daemonsets:
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
Loading
Loading