Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .github/workflows/cd.yml
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,10 @@ jobs:
run: python -m llm_router.serving > /tmp/ray-serve.yaml && diff -u config/ray-serve.yaml /tmp/ray-serve.yaml
- name: Verify the Ray topology matches the catalog
run: python -m llm_router.topology > /tmp/ray-service.yaml && diff -u deploy/overlays/ray/ray-service.yaml /tmp/ray-service.yaml
- name: Verify the Helm chart matches the manifests
run: python -m llm_router.chart --check && helm lint deploy/helm/llm-routing
- name: Package the Helm chart
run: helm package deploy/helm/llm-routing --destination chart
- name: Render the canary and rollback plans
# One plan per track (model, adapter, policy), each naming what it
# rolls back to and the criteria that trigger it.
Expand All @@ -87,6 +91,7 @@ jobs:
canary-plan.json
governance-plan.json
config/ray-serve.yaml
chart
rendered-base.yaml
rendered-ray.yaml
deploy
Expand Down
26 changes: 25 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -211,6 +211,29 @@ kubectl apply -k deploy/kubernetes # one vLLM engine
kubectl apply -k deploy/overlays/ray # Ray Serve across GPU pools
```

The same manifests install as a Helm chart, [`deploy/helm/llm-routing`](deploy/helm/llm-routing):

```bash
helm upgrade --install llm-routing deploy/helm/llm-routing \
--namespace llm-routing --create-namespace \
--set serving.mode=ray
```

| Value | Default | Purpose |
|---|---|---|
| `serving.mode` | `vllm` | `vllm` for one engine, `ray` for the Ray Serve topology below. |
| `images.*` | digest placeholders | One digest-pinned reference per workload. |
| `gateway.replicas` | `2` | Starting gateway size. |
| `gateway.autoscaling.minReplicas` / `maxReplicas` | `2` / `20` | KEDA bounds. |

The chart is generated from the manifests and never edited by hand; tests check that it renders
exactly what kustomize renders, in both modes.

```bash
python -m llm_router.chart # rebuild after changing a manifest
python -m llm_router.chart --check # exit 1 if the committed chart is stale
```

The base runs a single vLLM engine, which serves one model. The
[`deploy/overlays/ray`](deploy/overlays/ray) overlay replaces it with a KubeRay `RayService`
that serves every local model in the catalog behind one OpenAI-compatible endpoint:
Expand Down Expand Up @@ -258,7 +281,8 @@ python -m pip install -e ".[redis]"
CD renders the canary plans (one per track, each with its rollback target) and the governance
plan, verifies
`config/ray-serve.yaml` and the Ray topology against the catalog, and validates both rendered
topologies with kubeconform.
topologies with kubeconform. It also checks the Helm chart against the manifests, lints it, and
packages it.
Applying to a cluster stays disabled until a deployment destination is configured.

## Model registry
Expand Down
7 changes: 7 additions & 0 deletions deploy/helm/llm-routing/Chart.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
apiVersion: v2
name: llm-routing
description: OpenAI-compatible gateway, policy router and local LLM serving plane.
type: application
version: 0.1.0
appVersion: "0.1.0"
kubeVersion: ">=1.27.0-0"
Loading
Loading