Skip to content

feat: serve every catalog model from Ray worker pools split by accelerator - #37

Merged
github-actions[bot] merged 1 commit into
mainfrom
feat/24-ray-topology
Oct 4, 2026
Merged

github-actions[bot] merged 1 commit into
mainfrom
feat/24-ray-topology

Conversation

@Yash-Chindam

Copy link
Copy Markdown
Owner

What

Closes most of the deployment topology gap in the spec (§7.3, §17): Ray head and worker topology, GPU pools separated by accelerator type, NVIDIA GPU Operator support, the GPU exporter, and Grafana. Helm follows in its own PR.

Why an overlay

The base runs one vLLM engine, which can serve one model, while the catalog holds three local models on three accelerator types. deploy/overlays/ray replaces that engine with a KubeRay RayService that serves all of them behind one OpenAI-compatible endpoint. The base is left as it was, so the single-engine path and its gateway-side engine telemetry keep working.

How

  • llm_router.topology generates the RayService from the catalog: a head that runs no model, and one worker group per accelerator type sized from the autoscaling bounds of its models. A worker holds as many GPUs as the largest replica on its pool, so a tensor-parallel model fits on one node.
  • The embedded Serve application keeps only fields Ray's LLMConfig accepts, and loads weights from the location governance records in MLflow.
  • The overlay deletes the vllm-serve Deployment, Service and NetworkPolicy, repoints the gateway, and adds a Ray network policy and a PodMonitor.
  • deploy/gpu-operator/values.yaml: device plugin, GPU feature discovery (the gpu.product label the pools select on), and the DCGM exporter with a ServiceMonitor.
  • Two Grafana dashboards shipped as a labelled ConfigMap, and a PrometheusRule with eight alerts.
  • CD now fails on topology drift and validates both rendered topologies with kubeconform.

Tests

tests/unit/test_topology.py (20 tests): pool sizing, head and worker placement, LLMConfig field allow-list, artifact sources matching governance, pod security contract, generated-file drift, overlay structure, a real kubectl kustomize render of the overlay, and a check that every dashboard and alert expression queries a metric the gateway actually publishes.

Local: ruff and mypy clean, 389 tests pass, 98% coverage.

Limits

  • Nothing here has been applied to a cluster. The Ray image digest, GPU product labels and node sizes are placeholders.
  • Under the overlay the gateway's own router_engine_* and router_gpu_* gauges stay empty, so live load does not influence routing; engine metrics go to Prometheus from the Ray pods instead.
  • Ray resolves adapters at <artifact root>/adapters/<adapter id>, without the revision segment governance records; the deploy step has to publish the promoted revision there.
  • Dashboards chart gateway and DCGM series only. Ray's own vLLM series are scraped but not charted, because their exported names were not verified.

🤖 Generated with Claude Code

…rator

Generate a KubeRay RayService from the catalog with a head and one worker
group per accelerator type, shipped as an overlay that replaces the single
engine. Add GPU Operator values with the DCGM exporter, Grafana dashboards,
and alert rules checked against the metrics the gateway publishes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation area/api area/tests area/ci-cd labels Oct 4, 2026
@github-actions
github-actions Bot merged commit 2c119bc into main Oct 4, 2026
6 checks passed
@github-actions
github-actions Bot deleted the feat/24-ray-topology branch October 4, 2026 06:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/api area/ci-cd area/tests documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant