feat: serve every catalog model from Ray worker pools split by accelerator - #37
Merged
Merged
Conversation
…rator Generate a KubeRay RayService from the catalog with a head and one worker group per accelerator type, shipped as an overlay that replaces the single engine. Add GPU Operator values with the DCGM exporter, Grafana dashboards, and alert rules checked against the metrics the gateway publishes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Closes most of the deployment topology gap in the spec (§7.3, §17): Ray head and worker topology, GPU pools separated by accelerator type, NVIDIA GPU Operator support, the GPU exporter, and Grafana. Helm follows in its own PR.
Why an overlay
The base runs one vLLM engine, which can serve one model, while the catalog holds three local models on three accelerator types.
deploy/overlays/rayreplaces that engine with a KubeRayRayServicethat serves all of them behind one OpenAI-compatible endpoint. The base is left as it was, so the single-engine path and its gateway-side engine telemetry keep working.How
llm_router.topologygenerates theRayServicefrom the catalog: a head that runs no model, and one worker group per accelerator type sized from the autoscaling bounds of its models. A worker holds as many GPUs as the largest replica on its pool, so a tensor-parallel model fits on one node.LLMConfigaccepts, and loads weights from the location governance records in MLflow.vllm-serveDeployment, Service and NetworkPolicy, repoints the gateway, and adds a Ray network policy and aPodMonitor.deploy/gpu-operator/values.yaml: device plugin, GPU feature discovery (thegpu.productlabel the pools select on), and the DCGM exporter with a ServiceMonitor.PrometheusRulewith eight alerts.Tests
tests/unit/test_topology.py(20 tests): pool sizing, head and worker placement,LLMConfigfield allow-list, artifact sources matching governance, pod security contract, generated-file drift, overlay structure, a realkubectl kustomizerender of the overlay, and a check that every dashboard and alert expression queries a metric the gateway actually publishes.Local: ruff and mypy clean, 389 tests pass, 98% coverage.
Limits
router_engine_*androuter_gpu_*gauges stay empty, so live load does not influence routing; engine metrics go to Prometheus from the Ray pods instead.<artifact root>/adapters/<adapter id>, without the revision segment governance records; the deploy step has to publish the promoted revision there.🤖 Generated with Claude Code