diff --git a/apps/web/PRODUCT.md b/apps/web/PRODUCT.md index 5a8afa95e..f2022d3a8 100644 --- a/apps/web/PRODUCT.md +++ b/apps/web/PRODUCT.md @@ -32,7 +32,7 @@ The console runs beside the administrator's own Core, with execution, files and - **Monitor**: Overview (service status, running Sessions, sandbox slots, Sessions needing attention, 24-hour Session activity, a compact inventory of Core and up to four nodes, prioritizing offline and degraded nodes when the list is full, with a popover glance at each, the attention table, usage by project), Core metrics (the Core process's CPU and memory, execution slots and the Turn queue, connected daemons, the database and background jobs), Agent metrics (requests, errors, duration, tokens, models, tools, Agents and API keys for 1 h / 6 h / 24 h / 7 d), Sandbox metrics (node capacity and hosted Runtimes across projects; a node or a sandbox opens in a dialog with its figures and CPU and memory charts), Session log (every Session, read-only, with a failed Session's reason under its status, opening one Session's history, which jumps to its failed Turns; a self-hosted Session's page also has its environment's executor credentials). Agent metrics' By Agent table opens an Agent's page and, from its failed Turns, its Sessions in the Session log. - **Resources**: Agents, Environment templates, Skills, Files, Vaults. Each list shows one project or all projects, with a Project column when all are shown and a Creator column naming the creating key. Detail pages show the resource's facts and offer Delete. - **Platform**: Projects and keys (projects, their assets and usage, named keys, write history), Nodes (the node list, capacity, host figures, allocations and individual node operations). Add node asks for limits before issuing its one-time command; installers use Core's public URL and require supported node artifacts. Removal offers the host's uninstall command. System owns installation facts, the Domain and HTTPS secondary page, each harness's default model configuration, startup settings, and a link to the Sandbox configuration secondary page. That page owns setup, resource edits, rollout details and reset. Setup selects a backend, size and Runtime, then asks for a deliberate save; own-machine setup continues to Add node. -- A node whose provider is not ready names the reason (Docker unreachable, no Docker limits, missing Runtime image, no KVM, missing microsandbox components, a host too small) and its fix in the help tip beside its status, wherever that status shows. +- A node whose provider is not ready names the reason as one Provider-neutral readiness class (provider unavailable, host unsupported, provider files or Runtime image missing, Runtime download failed, a host too small) and its fix in the help tip beside its status, wherever that status shows. - A node enrolled with an earlier Core address gets no new sandboxes, so on the Nodes list and its page its status is Old address, with "Remove and add again", never Available. - **Sandbox reset** is an explicit administrator operation in System → Sandbox configuration. Auto clear is the default, with a one-hour deadline (5 minutes–24 hours); Force clear requires destructive confirmation. Reset stops new hosted Session admission, clears idle, suspended and pending hosted work, and waits for busy Turns and file writes until Core forces the remaining work. It does not affect self-hosted execution. Histories and persisted Files/Artifacts remain; archived Sessions cannot resume, and unpersisted workspace contents may be lost. Cancel stops further clearing without undoing archives. Core alone reports progress and completion, including resources blocked on named offline nodes; force does not bypass their cleanup. Completion clears the backend configuration and retires old nodes/enrollment credentials. A new configuration is then a separate deliberate save. - **Online sandbox configuration** changes the same backend's resources, Runtime or E2B template without retiring existing nodes or changing existing Sessions' resource ownership. New placement follows Core's qualified capacity; saving a target does not promise immediate placement on it. Configuration rollout shows Core's target preparation and retained previous-generation sandbox count. A settled rollout can still have failed, update-required or unknown nodes and old resources. An offline node stays offline even when it has a recorded serving generation. Node and allocation detail distinguish the serving pin, target preparation and each resource's configuration generation. diff --git a/apps/web/e2e/data/routes.mjs b/apps/web/e2e/data/routes.mjs index dd7572437..b0a7601ac 100644 --- a/apps/web/e2e/data/routes.mjs +++ b/apps/web/e2e/data/routes.mjs @@ -85,7 +85,7 @@ export function buildDemo(now = Math.floor(Date.now() / 1000), publicUrl = "http sessions.sort((a, b) => b.created_at - a.created_at); const nodes = [ { rollout: { state: "ready", ready_generation: 1 }, id: "node-local", name: "core-01", provider: "docker", online: true, provider_ready: true, cpu_count: 16, available_memory_bytes: 38 * 2 ** 30, available_disk_bytes: 410 * 2 ** 30, running: 5, snapshots: 2, last_seen_at: new Date((now - 8) * 1000).toISOString(), max_active: 8, max_retained: 16, active: 5, reserved: 1, retained: 2, cleanup_pending: 0, created_at: new Date((now - 86400 * 30) * 1000).toISOString() }, - { rollout: { state: "failed", ready_generation: null, diagnostic: "docker_limits_unsupported" }, id: "node-gpu", name: "gpu-worker-02", provider: "docker", online: true, provider_ready: false, diagnostic: "docker_limits_unsupported", cpu_count: 32, available_memory_bytes: 12 * 2 ** 30, available_disk_bytes: 96 * 2 ** 30, running: 7, snapshots: 5, last_seen_at: new Date((now - 12) * 1000).toISOString(), max_active: 8, max_retained: 16, active: 7, reserved: 0, retained: 5, cleanup_pending: 1, created_at: new Date((now - 86400 * 12) * 1000).toISOString() }, + { rollout: { state: "failed", ready_generation: null, diagnostic: "host_unsupported" }, id: "node-gpu", name: "gpu-worker-02", provider: "docker", online: true, provider_ready: false, diagnostic: "host_unsupported", cpu_count: 32, available_memory_bytes: 12 * 2 ** 30, available_disk_bytes: 96 * 2 ** 30, running: 7, snapshots: 5, last_seen_at: new Date((now - 12) * 1000).toISOString(), max_active: 8, max_retained: 16, active: 7, reserved: 0, retained: 5, cleanup_pending: 1, created_at: new Date((now - 86400 * 12) * 1000).toISOString() }, { rollout: { state: "unknown", ready_generation: 1 }, id: "node-edge", name: "edge-03", provider: "docker", online: false, provider_ready: false, cpu_count: 8, available_memory_bytes: null, available_disk_bytes: null, running: 0, snapshots: 0, last_seen_at: new Date((now - 5400) * 1000).toISOString(), max_active: 4, max_retained: 8, active: 0, reserved: 0, retained: 0, cleanup_pending: 0, created_at: new Date((now - 86400 * 3) * 1000).toISOString() }, ]; const hosted = sessions.filter((session) => session.environment.type === "openai_hosted"); diff --git a/apps/web/e2e/monitoring.spec.ts b/apps/web/e2e/monitoring.spec.ts index c48b1beea..be9449e51 100644 --- a/apps/web/e2e/monitoring.spec.ts +++ b/apps/web/e2e/monitoring.spec.ts @@ -24,7 +24,7 @@ test("shows the deployment's health on Overview and each monitor page", async ({ await page.getByRole("button", { name: "Sandbox metrics" }).click(); await expect(page.getByRole("table").first()).toContainText("core-01"); // A degraded node names why its provider is not ready. - await expect(page.locator(".status-with-help").filter({ hasText: /^Provider not ready/ }).getByRole("button", { name: "Docker limits unsupported", exact: true })).toBeVisible(); + await expect(page.locator(".status-with-help").filter({ hasText: /^Provider not ready/ }).getByRole("button", { name: "Host unsupported", exact: true })).toBeVisible(); }); test("opens a Session's conversation from the Session log, read-only", async ({ page, request }) => { diff --git a/apps/web/e2e/nodes.spec.ts b/apps/web/e2e/nodes.spec.ts index 5f32831da..6c9bb2c4d 100644 --- a/apps/web/e2e/nodes.spec.ts +++ b/apps/web/e2e/nodes.spec.ts @@ -91,9 +91,9 @@ test("adds a node: host requirements, a root/sudo command, a countdown, the same await expect(problem).toContainText("Not connected yet"); await expect(problem).toContainText("sudo journalctl -u oac-node-7f3c2a90-5b1e-4c2d-9e3f-0a1b2c3d4e5f.service"); await expect(add.getByText("No sudo on this host?")).toHaveCount(0); - // Connected, it reports why Docker isn't ready; once ready, the node is connected. - await setNode(request, { id: "node-new", online: true, diagnostic: "docker_limits_unsupported" }); - await expect(problem).toContainText("Docker limits unsupported"); + // Connected, it reports why its provider isn't ready; once ready, the node is connected. + await setNode(request, { id: "node-new", online: true, diagnostic: "host_unsupported" }); + await expect(problem).toContainText("Host unsupported"); await expect(progress).toContainText("Waiting for Docker"); await setNode(request, { id: "node-new", provider_ready: true, diagnostic: "" }); await expect(progress).toHaveText("edge-04 · Connected"); diff --git a/apps/web/src/features/sandbox/node-enrollment.test.ts b/apps/web/src/features/sandbox/node-enrollment.test.ts index 683c925df..74025f841 100644 --- a/apps/web/src/features/sandbox/node-enrollment.test.ts +++ b/apps/web/src/features/sandbox/node-enrollment.test.ts @@ -31,7 +31,7 @@ describe("node enrollment", () => { const connected = { ...fresh, online: true }; expect(enrollmentProgress(connected, 0, 1_000)).toEqual({ stage: "connected", problem: "" }); expect(enrollmentProgress(connected, 0, NODE_READY_WAIT_MS)).toEqual({ stage: "connected", problem: "provider_unavailable" }); - expect(enrollmentProgress({ ...connected, diagnostic: "kvm_unavailable" }, 0, 1_000)).toEqual({ stage: "connected", problem: "kvm_unavailable" }); + expect(enrollmentProgress({ ...connected, diagnostic: "host_unsupported" }, 0, 1_000)).toEqual({ stage: "connected", problem: "host_unsupported" }); expect(enrollmentProgress({ ...connected, provider_ready: true }, 0, NODE_READY_WAIT_MS * 2)).toEqual({ stage: "ready", problem: "" }); }); diff --git a/apps/web/src/lib/locale-strings.ts b/apps/web/src/lib/locale-strings.ts index fd23a23c3..cd7648aff 100644 --- a/apps/web/src/lib/locale-strings.ts +++ b/apps/web/src/lib/locale-strings.ts @@ -285,21 +285,17 @@ export const chinese = { "Sandbox ownership mismatch": "沙箱归属不一致", "Reconcile the assigned resource and its ownership record before resuming execution.": "请核对已分配资源及其归属记录,再恢复执行。", "Sandbox provider unavailable": "沙箱后端不可用", - "Restore the provider on the assigned node, then refresh. A connected node alone does not confirm that its sandbox provider is ready.": "请恢复已分配节点上的运行后端,然后刷新。节点在线并不代表其沙箱后端已就绪。", + "Restore the provider on the assigned node, then refresh. Running the install command again on the host checks its requirements and names the fix.": "请恢复已分配节点上的运行后端,然后刷新。在该主机上重新运行安装命令,会检查主机要求并指出修复方法。", "Sandbox state needs attention": "沙箱状态需要检查", "Inspect the assigned node and resource, then refresh.": "请检查已分配节点和资源,然后刷新。", - "Docker unavailable": "Docker 不可用", - "The node can't reach the Docker daemon. Check that Docker is running and the node can use its socket.": "节点连不上 Docker 守护进程。请确认 Docker 正在运行,且节点能访问它的 socket。", - "Docker limits unsupported": "Docker 无法限制资源", - "Docker on this host doesn't enforce CPU and memory limits. Enable cgroup limits.": "这台主机上的 Docker 不能限制 CPU 和内存。请启用 cgroup 限制。", + "Host unsupported": "主机不满足要求", + "The host lacks a capability its sandbox provider requires. Running the install command again on the host checks its requirements and names the fix.": "主机缺少沙箱后端所需的能力。在该主机上重新运行安装命令,会检查主机要求并指出修复方法。", + "Provider files missing": "后端文件缺失", + "Pinned provider files are missing or fail their checksum. Run the install command again.": "指定版本的后端文件缺失或校验不通过。请重新运行安装命令。", "Runtime download failed": "Runtime 下载失败", "Runtime files could not be downloaded or verified. Check the node's network access and the configured Runtime release.": "Runtime 文件下载或验证失败。请检查节点网络连接及配置的 Runtime 发布版本。", "Runtime image missing": "缺少 Runtime 镜像", "The pinned Runtime image isn't on the host. Run the install command again.": "主机上没有指定版本的 Runtime 镜像。请重新运行安装命令。", - "KVM unavailable": "KVM 不可用", - "/dev/kvm isn't available to the node. Enable virtualization or use a KVM-capable host.": "节点无法使用 /dev/kvm。请开启虚拟化,或换一台支持 KVM 的主机。", - "microsandbox components missing": "microsandbox 组件缺失", - "microsandbox components are missing or fail their checksum. Run the install command again.": "microsandbox 组件缺失或校验不通过。请重新运行安装命令。", "Host too small": "主机资源不足", "The host has less CPU or memory than one sandbox needs. Use a bigger host or a smaller sandbox size.": "主机的 CPU 或内存不够运行一个沙箱。请换一台更大的主机,或调小沙箱规格。", "Connecting to this console's Core…": "正在连接此控制台的 Core…", diff --git a/apps/web/src/lib/locale.test.ts b/apps/web/src/lib/locale.test.ts index 439cefe4b..1d61ce7e1 100644 --- a/apps/web/src/lib/locale.test.ts +++ b/apps/web/src/lib/locale.test.ts @@ -12,8 +12,8 @@ describe("sandbox localization", () => { expect(sandboxStateLabel("internal-value", "zh")).toBe("未知状态"); }); it("localizes all diagnostic labels and advice", () => { - for (const code of ["node_unavailable", "resource_missing", "compute_unconfirmed", "ownership_mismatch", "provider_unavailable", "docker_unavailable", - "docker_limits_unsupported", "runtime_download_failed", "runtime_image_unavailable", "kvm_unavailable", "microsandbox_artifacts_unavailable", "capacity_insufficient", "unknown"]) { + for (const code of ["node_unavailable", "resource_missing", "compute_unconfirmed", "ownership_mismatch", "provider_unavailable", "host_unsupported", + "artifacts_unavailable", "runtime_download_failed", "runtime_image_unavailable", "capacity_insufficient", "unknown"]) { const message = sandboxDiagnosticMessage(code, "zh"); expect(message?.label).toMatch(/[\u4e00-\u9fff]/); expect(message?.advice).toMatch(/[\u4e00-\u9fff]/); diff --git a/apps/web/src/lib/sandbox-diagnostic.test.ts b/apps/web/src/lib/sandbox-diagnostic.test.ts index 6c26cfb9f..4ead7d8fb 100644 --- a/apps/web/src/lib/sandbox-diagnostic.test.ts +++ b/apps/web/src/lib/sandbox-diagnostic.test.ts @@ -17,11 +17,11 @@ describe("sandbox diagnostics", () => { expect(sandboxDiagnosticMessage("provider_unavailable")?.label).toBe("Sandbox provider unavailable"); }); it("names why a node's provider is not ready, reading an unknown code as provider_unavailable", () => { - expect(sandboxDiagnosticMessage(nodeProviderDiagnostic({ online: true, provider_ready: false, diagnostic: "kvm_unavailable" }))?.label).toBe("KVM unavailable"); + expect(sandboxDiagnosticMessage(nodeProviderDiagnostic({ online: true, provider_ready: false, diagnostic: "host_unsupported" }))?.label).toBe("Host unsupported"); expect(nodeProviderDiagnostic({ online: true, provider_ready: false, diagnostic: "future_code" as SandboxNodeDiagnostic })).toBe("provider_unavailable"); expect(nodeProviderDiagnostic({ online: true, provider_ready: true, diagnostic: "" })).toBe(""); // An offline node's last code may no longer apply. - expect(nodeProviderDiagnostic({ online: false, provider_ready: false, diagnostic: "docker_unavailable" })).toBe(""); + expect(nodeProviderDiagnostic({ online: false, provider_ready: false, diagnostic: "host_unsupported" })).toBe(""); }); it("does not expose an unknown raw error or turn it into a healthy state", () => { const message = sandboxDiagnosticMessage("private-provider-error-with-secret"); diff --git a/apps/web/src/lib/sandbox-diagnostic.ts b/apps/web/src/lib/sandbox-diagnostic.ts index 20caf6db5..67820d9d9 100644 --- a/apps/web/src/lib/sandbox-diagnostic.ts +++ b/apps/web/src/lib/sandbox-diagnostic.ts @@ -22,16 +22,16 @@ const diagnostics: Record = { }, provider_unavailable: { label: "Sandbox provider unavailable", - advice: "Restore the provider on the assigned node, then refresh. A connected node alone does not confirm that its sandbox provider is ready.", + advice: "Restore the provider on the assigned node, then refresh. Running the install command again on the host checks its requirements and names the fix.", }, - // Fixed Runtime preparation and provider readiness diagnostics; Core sends only the code. - docker_unavailable: { - label: "Docker unavailable", - advice: "The node can't reach the Docker daemon. Check that Docker is running and the node can use its socket.", + // Provider-neutral readiness classes; Core sends only the code, and the node keeps the local detail. + host_unsupported: { + label: "Host unsupported", + advice: "The host lacks a capability its sandbox provider requires. Running the install command again on the host checks its requirements and names the fix.", }, - docker_limits_unsupported: { - label: "Docker limits unsupported", - advice: "Docker on this host doesn't enforce CPU and memory limits. Enable cgroup limits.", + artifacts_unavailable: { + label: "Provider files missing", + advice: "Pinned provider files are missing or fail their checksum. Run the install command again.", }, runtime_download_failed: { label: "Runtime download failed", @@ -41,14 +41,6 @@ const diagnostics: Record = { label: "Runtime image missing", advice: "The pinned Runtime image isn't on the host. Run the install command again.", }, - kvm_unavailable: { - label: "KVM unavailable", - advice: "/dev/kvm isn't available to the node. Enable virtualization or use a KVM-capable host.", - }, - microsandbox_artifacts_unavailable: { - label: "microsandbox components missing", - advice: "microsandbox components are missing or fail their checksum. Run the install command again.", - }, capacity_insufficient: { label: "Host too small", advice: "The host has less CPU or memory than one sandbox needs. Use a bigger host or a smaller sandbox size.", diff --git a/contracts/agents-api/core.openapi.yaml b/contracts/agents-api/core.openapi.yaml index 957cba8d8..9a1ef6ea1 100644 --- a/contracts/agents-api/core.openapi.yaml +++ b/contracts/agents-api/core.openapi.yaml @@ -946,12 +946,10 @@ definitions: description: Fixed reason for the last reported unreadiness; absent while the provider is ready. Clients treat an unknown value as provider_unavailable. enum: - provider_unavailable - - docker_unavailable - - docker_limits_unsupported + - host_unsupported + - artifacts_unavailable - runtime_download_failed - runtime_image_unavailable - - kvm_unavailable - - microsandbox_artifacts_unavailable - capacity_insufficient type: string enrollment_id: @@ -1037,12 +1035,10 @@ definitions: description: Fixed reason for the last reported unreadiness; absent while the provider is ready. Clients treat an unknown value as provider_unavailable. enum: - provider_unavailable - - docker_unavailable - - docker_limits_unsupported + - host_unsupported + - artifacts_unavailable - runtime_download_failed - runtime_image_unavailable - - kvm_unavailable - - microsandbox_artifacts_unavailable - capacity_insufficient type: string enrollment_id: @@ -1106,12 +1102,10 @@ definitions: diagnostic: enum: - provider_unavailable - - docker_unavailable - - docker_limits_unsupported + - host_unsupported + - artifacts_unavailable - runtime_download_failed - runtime_image_unavailable - - kvm_unavailable - - microsandbox_artifacts_unavailable - capacity_insufficient type: string ready_generation: diff --git a/contracts/agents-api/node-generation-protocol.md b/contracts/agents-api/node-generation-protocol.md index a7b293f41..14dc79870 100644 --- a/contracts/agents-api/node-generation-protocol.md +++ b/contracts/agents-api/node-generation-protocol.md @@ -17,7 +17,7 @@ Every frame is one JSON text message whose `version` equals `node.ProtocolVersio Core counts a node as online while it is connected under the current owner epoch and its last heartbeat is less than 45 seconds old. A heartbeat establishes provider readiness and the last host measurements, never Session activity. -Health carries `provider_ready`, an optional fixed `diagnostic`, `observed_at`, `active_operations` (at most 32) and the host measurements that the [Runtime telemetry API](./runtime-observability-api.md#node-host-observations-and-history) reports. A node without generation management probes its provider for every report; an unready provider reports one fixed diagnostic code, classified from typed probe errors, and the probe text and host paths stay on the node. Core stores an unknown code as `provider_unavailable`. The [nodes guide](../../docs/getting-started/nodes.md#readiness-codes) lists the codes and their causes. A generation-managing node reports readiness per generation instead, as described below. +Health carries `provider_ready`, an optional fixed `diagnostic`, `observed_at`, `active_operations` (at most 32) and the host measurements that the [Runtime telemetry API](./runtime-observability-api.md#node-host-observations-and-history) reports. A node without generation management probes its provider for every report; an unready provider reports the Provider-neutral readiness class of its first failed check, from the class error the probe wraps, and the probe text, vendor detail and host paths stay on the node. A `diagnostic` in node or generation health that is not a declared class makes the frame invalid. The [nodes guide](../../docs/getting-started/nodes.md#readiness-codes) lists the classes and their causes. A generation-managing node reports readiness per generation instead, as described below. ## Provider requests @@ -119,6 +119,6 @@ A fresh installation also records the verified checksums of its Runtime files se An interrupted download repairs only missing bytes at the original paths. When collection comes before any import attempt, the preparation journal proves that the generation has no imported native image. A generation whose native executable is missing and whose import may have started stays retained: missing files never prove native absence, and an empty native inventory never erases receipt or store history. -The diagnostic codes are authored in `services/core/internal/sandbox/node_diagnostic.go`. The shared `services/core/internal/sandbox/testdata/node-diagnostics.json` fixture checks the Go mapping, OpenAPI source annotations and generated enums, and the TypeScript client declaration. Web uses the client normalizer and checks localized messages for every declared code. Update these projections with a code change; unknown codes normalize to `provider_unavailable`. +The readiness classes, one exported error and one code each, are authored in `services/core/internal/sandbox/node_diagnostic.go`. The shared `services/core/internal/sandbox/testdata/node-diagnostics.json` fixture checks the Go mapping, OpenAPI source annotations and generated enums, and the TypeScript client declaration. Web uses the client normalizer and checks localized messages for every declared code. Update these projections with a code change; clients read an unknown code as `provider_unavailable`. Preparation diagnostics keep fixed typed causes. Only artifact transfer, checksum or release-provenance failures report `runtime_download_failed`; the private preparer signals that class through its exit category, without Core or the node parsing stderr. Provider, ownership, cancellation and unclassified failures keep their typed code or `provider_unavailable`. No raw provider text crosses the protocol. diff --git a/contracts/agents-api/sandbox-deployment.md b/contracts/agents-api/sandbox-deployment.md index bd1ff3424..c13eeba34 100644 --- a/contracts/agents-api/sandbox-deployment.md +++ b/contracts/agents-api/sandbox-deployment.md @@ -177,7 +177,7 @@ Node capacity is approved by the administrator, separately from the deployment s `GET /core/v1/sandbox/nodes` returns `{data: [...]}` with, for each node: `id`, `name`, `provider`, `online`, `last_seen_at`, `created_at`, `max_active`, `max_retained`, the counts `active`, `reserved`, `running`, `retained`, `snapshots` and `cleanup_pending`, `provider_ready`, `diagnostic`, the host measurements `cpu_count`, `available_memory_bytes` and `available_disk_bytes`, `rollout`, `enrollment_id` and `core_url`. `enrollment_id` is the handle of the command that registered the node, or null when Core has none. `core_url` is the installation public URL at enrollment; a node whose `core_url` differs from the current public URL receives no new placements. Work already placed on it finishes there, including a placed Environment that has no allocation yet, and its retained sandboxes can still resume while the old address reaches Core. Remove it and add it again. -A node without generation management reports its provider's readiness itself: `provider_ready`, and when it is unready one fixed `diagnostic` code. A node added with Web's command manages generations, so its readiness follows its serving generation and its fixed code for the target generation appears in `rollout.diagnostic`. The node classifies the first failed readiness check and sends only the code; Core stores any other value as `provider_unavailable` and never stores or returns probe text or host paths. The codes are `docker_unavailable`, `docker_limits_unsupported`, `runtime_download_failed`, `runtime_image_unavailable`, `kvm_unavailable`, `microsandbox_artifacts_unavailable`, `capacity_insufficient` and `provider_unavailable`. `runtime_download_failed` means the exact Runtime artifacts could not be transferred or verified; it never contains artifact URLs, credentials or transport output. [Readiness codes](../../docs/getting-started/nodes.md#readiness-codes) gives causes and operator actions. Core and nodes must come from the same distribution. +A node without generation management reports its provider's readiness itself: `provider_ready`, and when it is unready one fixed `diagnostic` code. A node added with Web's command manages generations, so its readiness follows its serving generation and its fixed code for the target generation appears in `rollout.diagnostic`. The node classifies the first failed readiness check and sends only its code; a frame with any other value is invalid, and Core never stores or returns probe text or host paths. The codes are Provider-neutral classes: `provider_unavailable`, `host_unsupported`, `artifacts_unavailable`, `runtime_download_failed`, `runtime_image_unavailable` and `capacity_insufficient`. `runtime_download_failed` means the exact Runtime artifacts could not be transferred or verified; it never contains artifact URLs, credentials or transport output. [Readiness codes](../../docs/getting-started/nodes.md#readiness-codes) gives causes and operator actions. Core and nodes must come from the same distribution. `PATCH /core/v1/sandbox/nodes/{node_id}` takes `{name, max_active, max_retained}`; lowering a limit stops no running sandbox. `DELETE /core/v1/sandbox/nodes/{node_id}` refuses with 409 `runtime_node_in_use` while the node holds allocations, snapshots, reservations or pending cleanup, including while it is offline. Removal deletes no compute and retires the node's identity; the host can come back only as a new node. There is no node drain. @@ -197,7 +197,6 @@ Some fields keep one name across providers but differ in meaning, or do not appl | Enrollment-token `max_active`, `max_retained` | 409 `sandbox_deployment_conflict`, after the 400 capacity checks; E2B has no nodes | `max_retained` always equals `max_active` | Both limits apply | | Node list and detail | Empty list; detail returns 404 | Enrolled nodes | Enrolled nodes | | Node `retained`, `snapshots`, `max_retained` | Not applicable | Docker never suspends: `retained` equals `active`, `snapshots` is 0 and `max_retained` equals `max_active` | Suspended sandboxes are `retained` minus `active` | -| Node `diagnostic` codes | Not applicable | `docker_unavailable`, `docker_limits_unsupported`, `runtime_download_failed`, `runtime_image_unavailable`, `capacity_insufficient` or `provider_unavailable` | `kvm_unavailable`, `microsandbox_artifacts_unavailable`, `runtime_download_failed`, `capacity_insufficient` or `provider_unavailable` | | Node `host.available_disk_bytes` | Not applicable | Free space on the filesystem of the node state directory, not a container's disk | Free space on the filesystem of the node state directory; sandbox disks have their own quotas | | Allocation `compute_phase`, `compute_phase_changed_at` | Not applicable: no node allocations | Always `disabled`, counted as running until release; the time is the allocation's creation | Includes `suspended`; its time plus `suspension.retention_seconds` tells roughly when Core reclaims the snapshot | | Runtime observation `cpu`, `memory` | From E2B metrics: `cpu.utilization_ratio` and `capacity_cores`, memory usage and limit; no cumulative CPU time | From Docker stats: `cpu.usage_seconds_total`, CPU and memory limits, memory usage | From the VM: `cpu.usage_seconds_total`, CPU and memory limits, memory usage | diff --git a/contracts/agents-api/zh/node-generation-protocol.md b/contracts/agents-api/zh/node-generation-protocol.md index e6e5de8fc..a5c9a658a 100644 --- a/contracts/agents-api/zh/node-generation-protocol.md +++ b/contracts/agents-api/zh/node-generation-protocol.md @@ -1,7 +1,7 @@ --- title: "沙箱节点协议" source: contracts/agents-api/node-generation-protocol.md -source_hash: 1ee43dfcdd0eec0806ea3bc8a4c1227e10bd8ac5e69486505cb113a98e3f548a +source_hash: b7f9a3a2835f65b477b53fcc62b474250eb77c3111f5e4692b561911c0ff067b --- 沙箱节点在其主机上运行 Docker 或 microsandbox Provider,并通过一个 WebSocket 与 Core 相连。Core 通过该连接发送 Provider 操作;节点针对本地 Provider 执行这些操作,并报告就绪状态、主机测量值及其持有的部署代次。Core 始终是唯一的生命周期所有者:节点绝不重试变更操作或调度工作。帧和校验器位于 [`services/core/internal/sandbox/node`](https://github.com/MiniMax-AI/OpenAgentCore/tree/main/services/core/internal/sandbox/node)(`wire.go`、`generation_wire.go`);节点用于注册和读取配置的 HTTP 路由位于[机器连接 API](machine-api.md#node-routes)。 @@ -19,7 +19,7 @@ source_hash: 1ee43dfcdd0eec0806ea3bc8a4c1227e10bd8ac5e69486505cb113a98e3f548a 只要节点在当前所有者 epoch 下保持连接,并且最近一次心跳距今不足 45 秒,Core 就会将该节点计为在线。心跳会确立 Provider 的就绪状态和最近的主机测量值,但绝不表示 Session 活动。 -健康报告包含 `provider_ready`、可选的固定 `diagnostic`、`observed_at`、最多为 32 的 `active_operations`,以及 [Runtime 遥测 API](runtime-observability-api.md#node-host-observations-and-history) 报告的主机测量值。未启用代次管理的节点每次报告时都会探测其 Provider;未就绪的 Provider 会报告一个固定诊断代码,该代码根据类型化探测错误进行分类;探测文本和主机路径保留在节点上。Core 会将未知代码存储为 `provider_unavailable`。[节点指南](../../../docs/zh/getting-started/nodes.md#readiness-codes) 列出了这些代码及其原因。支持代次管理的节点则按下文所述按代次报告就绪状态。 +健康报告包含 `provider_ready`、可选的固定 `diagnostic`、`observed_at`、最多为 32 的 `active_operations`,以及 [Runtime 遥测 API](runtime-observability-api.md#node-host-observations-and-history) 报告的主机测量值。未启用代次管理的节点每次报告时都会探测其 Provider;未就绪的 Provider 会报告其首个失败检查所属的、与 Provider 无关的就绪类别,该类别取自探测所包装的类别错误;探测文本、厂商细节和主机路径保留在节点上。节点或代次健康状态中的 `diagnostic` 若不是已声明的类别,该帧即无效。[节点指南](../../../docs/zh/getting-started/nodes.md#readiness-codes) 列出了这些类别及其原因。支持代次管理的节点则按下文所述按代次报告就绪状态。 ## Provider 请求 {#provider-requests} @@ -121,6 +121,6 @@ Runtime 字节缺失时,绝不将固定的放置实例迁移到当前 Runtime 下载中断后,只会修复原始路径中缺失的字节。如果在任何导入尝试之前执行回收,准备日志会证明该代次没有已导入的原生镜像。如果某代次的原生可执行文件缺失,且导入可能已经开始,该代次仍会保留:文件缺失永远不能证明原生制品不存在,而空的原生清单也永远不能抹除回执或存储历史。 -诊断代码编写于 `services/core/internal/sandbox/node_diagnostic.go`。共享的 `services/core/internal/sandbox/testdata/node-diagnostics.json` 测试夹具检查 Go 映射、OpenAPI 源注释和生成的枚举,以及 TypeScript 客户端声明。Web 使用客户端规范化器,并检查每个已声明代码的本地化消息。代码变更时要同步更新这些投影;未知代码会规范化为 `provider_unavailable`。 +就绪类别编写于 `services/core/internal/sandbox/node_diagnostic.go`,每个类别对应一个导出错误和一个代码。共享的 `services/core/internal/sandbox/testdata/node-diagnostics.json` 测试夹具检查 Go 映射、OpenAPI 源注释和生成的枚举,以及 TypeScript 客户端声明。Web 使用客户端规范化器,并检查每个已声明代码的本地化消息。代码变更时要同步更新这些投影;客户端将未知代码读作 `provider_unavailable`。 准备诊断使用固定的类型化原因。只有制品传输、校验和或版本来源验证失败才会报告 `runtime_download_failed`;私有准备器通过退出类别指示这一类失败,Core 和节点都不解析 stderr。Provider 故障、所有权故障、取消和未分类故障保留其类型化代码,或使用 `provider_unavailable`。协议中不会传输任何 Provider 原始文本。 diff --git a/contracts/agents-api/zh/sandbox-deployment.md b/contracts/agents-api/zh/sandbox-deployment.md index 3bc42898a..9e09f084b 100644 --- a/contracts/agents-api/zh/sandbox-deployment.md +++ b/contracts/agents-api/zh/sandbox-deployment.md @@ -1,7 +1,7 @@ --- title: "沙箱部署" source: contracts/agents-api/sandbox-deployment.md -source_hash: 68aabafdb8983280a3bb29e7398042f9e843310d1784acd4a9a94e9647fd3919 +source_hash: 2a6b114b7b4f324b2a68ee1f5bde5e9f229a2d44c06aa7c0dd4560bb4ed60166 --- 沙箱部署为 Core 管理的 `openai_hosted` 执行选择 Sandbox Provider、每个沙箱的资源以及不可变的 Runtime 发行版。PostgreSQL 为每个安装维护一个当前有效选择;Web 和 Core API 写入同一配置。节点文件保存其已安装副本和特定于主机的路径,且不能覆盖其资源或 Runtime。该选择独立于 Harness;部署可以保持未配置状态,既无节点,也不接受托管准入。 @@ -179,7 +179,7 @@ POST 会在持久保存候选配置之前对其进行验证,并且不会创建 `GET /core/v1/sandbox/nodes` 返回 `{data: [...]}`,其中每个节点包含 `id`、`name`、`provider`、`online`、`last_seen_at`、`created_at`、`max_active`、`max_retained`,计数项 `active`、`reserved`、`running`、`retained`、`snapshots` 和 `cleanup_pending`,以及 `provider_ready`、`diagnostic`、主机测量值 `cpu_count`、`available_memory_bytes` 和 `available_disk_bytes`、`rollout`、`enrollment_id` 和 `core_url`。`enrollment_id` 是注册该节点的命令的句柄;如果 Core 没有该句柄,则为 null。`core_url` 是注册时的安装公开 URL;`core_url` 与当前公开 URL 不同的节点不会收到新放置。已放置到该节点的工作会继续在那里完成,包括已放置但尚未获得分配的 Environment;只要旧地址仍能访问 Core,其保留沙箱仍可恢复。请移除该节点并重新添加。 -不具备代次管理的节点会自行报告其提供商的就绪状态:`provider_ready`,未就绪时则报告一个固定的 `diagnostic` 代码。通过 Web 的命令添加的节点会管理代次,因此其就绪状态取决于服务代次,而目标代次的固定代码会出现在 `rollout.diagnostic` 中。节点会对首次失败的就绪检查进行分类,并且只发送代码;Core 会将任何其他值存储为 `provider_unavailable`,且绝不存储或返回探测文本或主机路径。代码包括 `docker_unavailable`、`docker_limits_unsupported`、`runtime_download_failed`、`runtime_image_unavailable`、`kvm_unavailable`、`microsandbox_artifacts_unavailable`、`capacity_insufficient` 和 `provider_unavailable`。`runtime_download_failed` 表示无法传输或验证精确的 Runtime 制品;它绝不会包含制品 URL、凭据或传输输出。[就绪代码](../../../docs/zh/getting-started/nodes.md#readiness-codes)给出了原因和操作员应采取的措施。Core 和节点必须来自同一发行包。 +不具备代次管理的节点会自行报告其提供商的就绪状态:`provider_ready`,未就绪时则报告一个固定的 `diagnostic` 代码。通过 Web 的命令添加的节点会管理代次,因此其就绪状态取决于服务代次,而目标代次的固定代码会出现在 `rollout.diagnostic` 中。节点会对首次失败的就绪检查进行分类,并且只发送代码;携带任何其他值的帧均无效,且 Core 绝不存储或返回探测文本或主机路径。这些代码是与 Provider 无关的类别:`provider_unavailable`、`host_unsupported`、`artifacts_unavailable`、`runtime_download_failed`、`runtime_image_unavailable` 和 `capacity_insufficient`。`runtime_download_failed` 表示无法传输或验证精确的 Runtime 制品;它绝不会包含制品 URL、凭据或传输输出。[就绪代码](../../../docs/zh/getting-started/nodes.md#readiness-codes)给出了原因和操作员应采取的措施。Core 和节点必须来自同一发行包。 `PATCH /core/v1/sandbox/nodes/{node_id}` 接受 `{name, max_active, max_retained}`;降低限制不会停止任何正在运行的沙箱。当节点仍持有分配、快照、预留或待处理清理时,包括节点离线期间,`DELETE /core/v1/sandbox/nodes/{node_id}` 会拒绝操作并返回 409 `runtime_node_in_use`。移除操作不会删除任何计算资源,并且会停用该节点的身份;该主机只能作为新节点重新加入。不存在节点排空过程。 @@ -199,7 +199,6 @@ POST 会在持久保存候选配置之前对其进行验证,并且不会创建 | 注册令牌 `max_active`、`max_retained` | 先执行 400 容量检查,然后返回 409 `sandbox_deployment_conflict`;E2B 没有节点 | `max_retained` 始终等于 `max_active` | 两个限制均适用 | | 节点列表和详情 | 空列表;详情返回 404 | 已注册节点 | 已注册节点 | | 节点 `retained`、`snapshots`、`max_retained` | 不适用 | Docker 绝不暂停:`retained` 等于 `active`,`snapshots` 为 0,`max_retained` 等于 `max_active` | 已暂停沙箱数为 `retained` 减去 `active` | -| 节点 `diagnostic` 代码 | 不适用 | `docker_unavailable`、`docker_limits_unsupported`、`runtime_download_failed`、`runtime_image_unavailable`、`capacity_insufficient` 或 `provider_unavailable` | `kvm_unavailable`、`microsandbox_artifacts_unavailable`、`runtime_download_failed`、`capacity_insufficient` 或 `provider_unavailable` | | 节点 `host.available_disk_bytes` | 不适用 | 节点状态目录所在文件系统的可用空间,而不是容器的磁盘 | 节点状态目录所在文件系统的可用空间;沙箱磁盘有自己的配额 | | 分配 `compute_phase`、`compute_phase_changed_at` | 不适用:没有节点分配 | 始终为 `disabled`,在释放前计为运行中;该时间为分配创建时间 | 包含 `suspended`;该时间加上 `suspension.retention_seconds` 可大致确定 Core 回收快照的时间 | | Runtime 观测 `cpu`、`memory` | 来自 E2B 指标:`cpu.utilization_ratio` 和 `capacity_cores`、内存使用量和限制;无累计 CPU 时间 | 来自 Docker stats:`cpu.usage_seconds_total`、CPU 和内存限制、内存使用量 | 来自 VM:`cpu.usage_seconds_total`、CPU 和内存限制、内存使用量 | diff --git a/docs/getting-started/nodes.md b/docs/getting-started/nodes.md index fd99451cd..6410138f6 100644 --- a/docs/getting-started/nodes.md +++ b/docs/getting-started/nodes.md @@ -157,20 +157,18 @@ When a node is online but its sandbox provider is not ready, **Nodes** and **Ove - **Nodes added with Web's command** report a failed check of the current sandbox configuration as **Preparation failed** in the node's target status, with the reason in a help tip (`rollout.diagnostic` in `GET /core/v1/sandbox/nodes`). The tip beside **Provider not ready** only says *Sandbox provider unavailable*. - **Manually registered nodes** show the reason in the help tip beside **Provider not ready** (`diagnostic`). -The node's log has the local error behind the code. +Each code is a Provider-neutral class; the local error behind it stays on the node. A manually registered node logs it. For a node added with Web's command, rerun the command on the host: the installer checks the host requirements first and names what to fix ([Installer messages](#installer-messages)). A node reports only its first failed check, in this order: the Docker daemon or KVM, Docker's limit support, host capacity, then the installed Runtime files. An unreachable Docker daemon therefore hides a missing image. The next heartbeat, about ten seconds after a fix, clears or replaces the code. An offline node keeps its last code, which Web hides until the node reconnects. | Code | Help tip | Cause | Fix | | --- | --- | --- | --- | -| `docker_unavailable` | Docker unavailable | The Docker socket is unreachable or not accessible, or Docker fails its info or image request | Start Docker and give the node's user access to `/var/run/docker.sock` | -| `docker_limits_unsupported` | Docker limits unsupported | Docker reports no CPU quota or memory limit support | Use a host whose cgroups enforce CPU and memory limits (cgroup v2) | +| `provider_unavailable` | Sandbox provider unavailable | The provider's service is unreachable or fails a request, such as a stopped Docker daemon, or a failure without a class | Rerun the node's command, or read a manually registered node's log | +| `host_unsupported` | Host unsupported | The host lacks a capability the provider requires, such as Docker CPU and memory limits (cgroup v2) or read-write access to `/dev/kvm` | Rerun the node's command, or read a manually registered node's log | | `capacity_insufficient` | Host too small | The host has fewer CPUs or less memory than one sandbox | Use a larger host, or change the sandbox size | -| `runtime_image_unavailable` | Runtime image missing | Docker does not have the pinned Runtime image | A node added with Web's command downloads it again by itself; otherwise load the image from the matching release | -| `kvm_unavailable` | KVM unavailable | The node can't open `/dev/kvm` for reading and writing | Enable hardware virtualization and give the node's user KVM access, through the `kvm` group | -| `microsandbox_artifacts_unavailable` | microsandbox components missing | The Runtime or firmware is missing or fails its SHA-256 check, or the helper is missing | A node added with Web's command downloads the missing files by itself; otherwise restore them from the matching release | +| `runtime_image_unavailable` | Runtime image missing | The provider does not have the pinned Runtime image | A node added with Web's command downloads it again by itself; otherwise load the image from the matching release | +| `artifacts_unavailable` | Provider files missing | A pinned provider file, such as the microsandbox Runtime, firmware or helper, is missing or fails its SHA-256 check | A node added with Web's command downloads the missing files by itself; otherwise restore them from the matching release | | `runtime_download_failed` | Runtime download failed | While preparing a new configuration, the node could not download or verify the Runtime files | Check the node's HTTPS access to the console and the release. The node retries with growing delays, up to 30 minutes apart | -| `provider_unavailable` | Sandbox provider unavailable | Any other failure | Read the node's log | A new group membership applies only to a new process. Restart the node service: `sudo systemctl restart oac-node-.service`. A node that is registered but never connects usually can't reach Core at the public URL, or its `/api/v1` WebSocket doesn't pass the reverse proxy. diff --git a/docs/web/console-api-usage.md b/docs/web/console-api-usage.md index 09df51a49..60c2b623f 100644 --- a/docs/web/console-api-usage.md +++ b/docs/web/console-api-usage.md @@ -98,7 +98,7 @@ The list carries each harness's configuration, so the console does not read `GET | Deployment | `GET`, `POST`, `PUT /core/v1/sandbox/deployment` | Read the provider, the read-only `core_url` (`OAC_PUBLIC_URL`, shown in the setup review and never sent), reset state, installation ID and specification; a 409 `sandbox_configuration_error` (E2B with a loopback `public_url`) shows the shared client's fixed safe address-configuration message in the setup wizard, with Managed in System leading to System, and leaves nothing to confirm; initialize the deployment with `resources` and the Docker or microsandbox `runtime` release, or with the E2B account and no `resources` (Core adopts the template build's CPU and memory); change its settings with the expected generation. E2B's `metadata.template_build` (status, CPU, memory, disk) shows on System, the Sandbox configuration summary and Sandbox metrics, and sizes each sandbox when `specification.resources` is missing; microsandbox's `suspension` (idle and retention seconds) shows on System and the Nodes summary | | E2B discovery | `POST /core/v1/sandbox/providers/e2b/discovery` | The setup wizard lists the templates the entered E2B key can see, then the selected template's ready builds. The key travels only in these request bodies and the deployment write | | Reset | `POST`, `DELETE /core/v1/sandbox/deployment/reset` | Explicitly clear hosted resources, or cancel the remaining clear at the observed generation; show Core's remaining and offline projection | -| Nodes | `GET /core/v1/sandbox/nodes` | Nodes page; fleet on Overview; node capacity on Sandbox metrics. An online node's `diagnostic` (`docker_unavailable`, `docker_limits_unsupported`, `runtime_image_unavailable`, `kvm_unavailable`, `microsandbox_artifacts_unavailable`, `capacity_insufficient`, `provider_unavailable`; any other value reads as `provider_unavailable`) marks it degraded and names the reason and fix in the help tip beside its status on each of these and on the node's page. A node whose `core_url` (the address it enrolled with) differs from the deployment's `core_url` is named on the Nodes page as bound to an old address, to be removed and added again, and its status there and on its page reads Old address instead of its health. **Add node** follows only the node whose `enrollment_id` equals its command's | +| Nodes | `GET /core/v1/sandbox/nodes` | Nodes page; fleet on Overview; node capacity on Sandbox metrics. An online node's `diagnostic` (a [readiness code](../getting-started/nodes.md#readiness-codes); any other value reads as `provider_unavailable`) marks it degraded and names the reason and fix in the help tip beside its status on each of these and on the node's page. A node whose `core_url` (the address it enrolled with) differs from the deployment's `core_url` is named on the Nodes page as bound to an old address, to be removed and added again, and its status there and on its page reads Old address instead of its health. **Add node** follows only the node whose `enrollment_id` equals its command's | | Node detail | `GET /core/v1/sandbox/nodes/{node_id}?range=1h\|6h\|24h` | Sandbox metrics node dialog: the host's CPU busy share and memory from its last heartbeat, and their history over the page's range. **Edit node** reads `host.effective_cpu_cores` and `host.total_memory_bytes` to show the host beside each sandbox's size and at most how many of those fit | | Allocations | `GET /core/v1/sandbox/nodes/{node_id}/allocations` | Nodes page; Sandbox metrics. Under microsandbox, a node's page shows from `compute_phase_changed_at` how long each allocation has been in its compute phase and, while suspended, about when Core reclaims it (that time plus the deployment's `suspension.retention_seconds`); a null time shows a dash | | Enrollment | `POST /core/v1/sandbox/enrollment-tokens` | **Add node**: the administrator sets the node's sandbox limits (`max_active`; `max_retained` only for microsandbox, equal to `max_active` for Docker) before Core issues a single-use token inside a command that verifies the installer checksum, with the command's `enrollment_id`, which the node it registers reports. The command runs the installer with sudo (a system service) and passes the token on standard input; root runs it directly. No ordinary-user installation or removal entry is exposed, and the log hint always names the system service. The command downloads the installer from the installation's `public_url`. No token is requested until the installation is read, when it cannot be read, when it is `local_only` (or its `public_url` is not an HTTPS origin), or when `/console/config` lists `node_artifacts` without the deployment's provider. The dialog reads both again on opening and when the window regains focus | diff --git a/docs/zh/getting-started/nodes.md b/docs/zh/getting-started/nodes.md index ec1e69d0b..871fed81b 100644 --- a/docs/zh/getting-started/nodes.md +++ b/docs/zh/getting-started/nodes.md @@ -1,7 +1,7 @@ --- title: "添加和管理节点" source: docs/getting-started/nodes.md -source_hash: f5bec50bf81ebf1d08faaa54432da6a9c6e3ddbf88d93e33244276546c74aaab +source_hash: 24b505c131f398dc8340d4f32616152a7964a783cd4fc06cfdeb3ccd422448db --- 节点是一台 Linux 主机,在沙箱后端为 Docker 或 microsandbox 时,为 Core 托管 Session 运行沙箱。Core 将新 Session 分配给有空余容量的节点;节点创建沙箱,沙箱回连 Core。E2B 不需要节点。应用为自己的 Session 连接的机器是[自托管执行器](self-hosted.md),而不是节点。 @@ -159,20 +159,18 @@ root 只准备账号、组和服务单元;其他操作(包括 Docker 网络 - **使用 Web 命令添加的节点**:当前沙箱配置检查失败时,在目标状态显示 **Preparation failed**,帮助提示中给出原因(`GET /core/v1/sandbox/nodes` 的 `rollout.diagnostic`)。**Provider not ready** 旁的提示仅显示 *Sandbox provider unavailable*。 - **手动注册的节点**:在 **Provider not ready** 旁的帮助提示展示原因(`diagnostic`)。 -节点日志包含状态码背后的本地错误。 +每个状态码都是与 Provider 无关的类别;其背后的本地错误留在节点上。手动注册的节点会把它写入日志。对于使用 Web 命令添加的节点,请在主机上重新运行该命令:安装程序会先检查主机要求,并指出需要修复的问题([安装程序消息](#installer-messages))。 节点按如下顺序仅报告首个失败检查:Docker 守护进程或 KVM、Docker 限制支持、主机容量、已安装的 Runtime 文件。因此无法访问 Docker 守护进程时,会隐藏镜像缺失问题。修复后约十秒的下一次心跳会清除或替换状态码。离线节点保留最后状态码,Web 在节点重连前隐藏它。 | 状态码 | 帮助提示 | 原因 | 解决方法 | | --- | --- | --- | --- | -| `docker_unavailable` | Docker unavailable | Docker 套接字不可达、无权访问,或 Docker info/镜像请求失败 | 启动 Docker 并赋予节点用户访问 `/var/run/docker.sock` 的权限 | -| `docker_limits_unsupported` | Docker limits unsupported | Docker 报告不支持 CPU 配额或内存限制 | 使用 cgroups 强制执行 CPU 与内存限制的主机(cgroup v2) | +| `provider_unavailable` | Sandbox provider unavailable | 提供商的服务不可达或请求失败(例如 Docker 守护进程已停止),或失败没有类别 | 重新运行节点的命令,或阅读手动注册节点的日志 | +| `host_unsupported` | Host unsupported | 主机缺少提供商所需的能力,例如 Docker 的 CPU 与内存限制(cgroup v2)或对 `/dev/kvm` 的读写权限 | 重新运行节点的命令,或阅读手动注册节点的日志 | | `capacity_insufficient` | Host too small | 主机 CPU 或内存不足以运行一个沙箱 | 使用更大主机或修改沙箱规格 | -| `runtime_image_unavailable` | Runtime image missing | Docker 中没有固定版本的 Runtime 镜像 | Web 命令添加的节点自动重新下载;其他节点加载匹配发行版镜像 | -| `kvm_unavailable` | KVM unavailable | 节点无法读写 `/dev/kvm` | 启用硬件虚拟化,通过 `kvm` 组赋予节点用户 KVM 访问权限 | -| `microsandbox_artifacts_unavailable` | microsandbox components missing | Runtime 或固件缺失、SHA-256 检查失败,或辅助程序缺失 | Web 命令添加的节点自动下载缺失文件;其他节点从匹配发行版恢复 | +| `runtime_image_unavailable` | Runtime image missing | 提供商中没有固定版本的 Runtime 镜像 | Web 命令添加的节点自动重新下载;其他节点加载匹配发行版镜像 | +| `artifacts_unavailable` | Provider files missing | 固定版本的提供商文件(例如 microsandbox 的 Runtime、固件或辅助程序)缺失或 SHA-256 检查失败 | Web 命令添加的节点自动下载缺失文件;其他节点从匹配发行版恢复 | | `runtime_download_failed` | Runtime download failed | 准备新配置时无法下载或验证 Runtime 文件 | 检查节点到控制台和发行下载地址的 HTTPS 访问。节点以递增间隔重试,最长间隔 30 分钟 | -| `provider_unavailable` | Sandbox provider unavailable | 其他失败 | 阅读节点日志 | 新用户组成员关系仅对新进程生效。重启节点服务:`sudo systemctl restart oac-node-.service`。已注册但从未连接的节点通常无法通过公开 URL 访问 Core,或 `/api/v1` WebSocket 无法通过反向代理。 diff --git a/docs/zh/web/console-api-usage.md b/docs/zh/web/console-api-usage.md index 1b4890fd5..02d8896f6 100644 --- a/docs/zh/web/console-api-usage.md +++ b/docs/zh/web/console-api-usage.md @@ -1,7 +1,7 @@ --- title: "控制台 API 使用" source: docs/web/console-api-usage.md -source_hash: 7318d1d082d19cc1681a2bd91556c4dc7e16b211ae7cd3936b99d63c26949900 +source_hash: 071c5b80262cdc89ed3517cf6a63d1c657ee8d0f98133630bad0850595003a2c --- 本页列出各控制台页面读取和写入的 Core 路由,以及控制台如何限定读取范围。[administrator API contract](../../../contracts/agents-api/zh/admin-api.md) 定义了路由、响应结构、分页和审计记录;[API namespaces and credentials](../api/index.md) 定义了本文使用的术语。 @@ -100,7 +100,7 @@ source_hash: 7318d1d082d19cc1681a2bd91556c4dc7e16b211ae7cd3936b99d63c26949900 | 部署 | `GET`、`POST`、`PUT /core/v1/sandbox/deployment` | 读取提供商、只读 `core_url`(即 `OAC_PUBLIC_URL`,会显示在设置审核中且绝不发送)、重置状态、安装 ID 和规范;409 `sandbox_configuration_error`(E2B 搭配回环地址形式的 `public_url`)会在设置向导中显示共享客户端固定的安全地址配置消息,并通过 Managed in System 前往 System,且无需确认;使用 `resources` 以及 Docker 或 microsandbox 的 `runtime` release 初始化部署,或者使用 E2B 账户且不提供 `resources`(Core 采用模板构建的 CPU 和内存);使用预期的 generation 更改设置。E2B 的 `metadata.template_build`(状态、CPU、内存、磁盘)会显示在 System、Sandbox 配置摘要和 Sandbox metrics 中;当缺少 `specification.resources` 时,它还会确定每个 Sandbox 的大小;microsandbox 的 `suspension`(空闲和保留秒数)会显示在 System 和 Nodes 摘要中 | | E2B 发现 | `POST /core/v1/sandbox/providers/e2b/discovery` | 设置向导先列出输入的 E2B 密钥可见的模板,再列出所选模板的可用构建。该密钥只会通过这些请求体和部署写入请求传输 | | 重置 | `POST`、`DELETE /core/v1/sandbox/deployment/reset` | 显式清除托管资源,或在观测到的 generation 处取消剩余清除;显示 Core 的剩余资源和离线预测 | -| Nodes | `GET /core/v1/sandbox/nodes` | Nodes 页面;Overview 上的机群;Sandbox metrics 中的节点容量。在线节点的 `diagnostic`(`docker_unavailable`、`docker_limits_unsupported`、`runtime_image_unavailable`、`kvm_unavailable`、`microsandbox_artifacts_unavailable`、`capacity_insufficient`、`provider_unavailable`;任何其他值均读取为 `provider_unavailable`)会将其标记为降级,并在上述每个页面及节点页面中,紧邻状态的帮助提示里说明原因和修复方法。如果节点的 `core_url`(其注册时使用的地址)与部署的 `core_url` 不同,Nodes 页面会将其标记为绑定到旧地址,需要移除后重新添加;此时它在该页面和节点页面中的状态会显示 Old address,而不是健康状态。**Add node** 仅跟踪 `enrollment_id` 与其命令所含 `enrollment_id` 相等的节点 | +| Nodes | `GET /core/v1/sandbox/nodes` | Nodes 页面;Overview 上的机群;Sandbox metrics 中的节点容量。在线节点的 `diagnostic`(一个[就绪状态码](../getting-started/nodes.md#readiness-codes);任何其他值均读取为 `provider_unavailable`)会将其标记为降级,并在上述每个页面及节点页面中,紧邻状态的帮助提示里说明原因和修复方法。如果节点的 `core_url`(其注册时使用的地址)与部署的 `core_url` 不同,Nodes 页面会将其标记为绑定到旧地址,需要移除后重新添加;此时它在该页面和节点页面中的状态会显示 Old address,而不是健康状态。**Add node** 仅跟踪 `enrollment_id` 与其命令所含 `enrollment_id` 相等的节点 | | 节点详情 | `GET /core/v1/sandbox/nodes/{node_id}?range=1h\|6h\|24h` | Sandbox metrics 节点对话框:主机自最近一次心跳以来的 CPU 忙碌占比和内存使用量,以及页面所选范围内二者的历史记录。**Edit node** 读取 `host.effective_cpu_cores` 和 `host.total_memory_bytes`,用于在每个 Sandbox 大小旁显示主机容量,以及最多可容纳多少个该大小的 Sandbox | | 分配 | `GET /core/v1/sandbox/nodes/{node_id}/allocations` | Nodes 页面;Sandbox metrics。在 microsandbox 下,节点页面根据 `compute_phase_changed_at` 显示每个分配处于计算阶段的时间,并在分配暂停时估算 Core 回收它的时间(该时间加上部署的 `suspension.retention_seconds`);时间为 null 时显示短横线 | | 注册 | `POST /core/v1/sandbox/enrollment-tokens` | **Add node**:管理员先设置节点的 Sandbox 限制(`max_active`;`max_retained` 仅适用于 microsandbox,在 Docker 下等于 `max_active`),然后 Core 才会把一次性令牌放入命令中;该命令会验证安装程序校验和,并包含命令的 `enrollment_id`,节点注册时会报告此 ID。命令使用 sudo 运行安装程序(作为系统服务),并通过标准输入传递令牌;以 root 运行时则直接执行。界面不提供普通用户安装或移除入口,日志提示始终指明系统服务。命令从安装的 `public_url` 下载安装程序。只有成功读取安装信息后才会请求令牌;如果安装信息无法读取、安装为 `local_only`(或其 `public_url` 不是 HTTPS 来源),或者 `/console/config` 列出的 `node_artifacts` 不包含部署的提供商,则不会请求令牌。对话框在打开时和窗口重新获得焦点时,会再次读取这两项信息 | diff --git a/packages/agents-client/src/sandbox-client.test.ts b/packages/agents-client/src/sandbox-client.test.ts index 7567263b4..9bea39776 100644 --- a/packages/agents-client/src/sandbox-client.test.ts +++ b/packages/agents-client/src/sandbox-client.test.ts @@ -18,7 +18,7 @@ const node = { }; /** Never heard from, enrolled before Core recorded enrollment IDs, with a fixed readiness code. */ const unready = { - ...node, rollout: { state: "unknown", ready_generation: 1 }, id: "7f6e5d4c-3b2a-4190-8f7e-6d5c4b3a2918", online: false, provider_ready: false, diagnostic: "kvm_unavailable", + ...node, rollout: { state: "unknown", ready_generation: 1 }, id: "7f6e5d4c-3b2a-4190-8f7e-6d5c4b3a2918", online: false, provider_ready: false, diagnostic: "host_unsupported", cpu_count: null, available_memory_bytes: null, available_disk_bytes: null, running: 0, last_seen_at: null, active: 0, retained: 0, enrollment_id: null, }; const detail = { @@ -156,8 +156,8 @@ describe("Core sandbox credential boundaries", () => { it("returns the node's fixed readiness diagnostic unchanged and reads an unknown code as provider_unavailable", async () => { const fetch = vi.fn().mockResolvedValue(response({ data: [unready, { ...unready, diagnostic: "future_code" }, node] })); const { data } = await new SandboxAdminClient({ baseUrl: "/core/v1/sandbox", fetch }).listNodes(); - expect(data.map((entry) => entry.diagnostic)).toEqual(["kvm_unavailable", "provider_unavailable", undefined]); - expectTypeOf().toEqualTypeOf(); + expect(data.map((entry) => entry.diagnostic)).toEqual(["host_unsupported", "provider_unavailable", undefined]); + expectTypeOf().toEqualTypeOf(); }); it("requires each node's enrollment ID: a string, or null for nodes enrolled before Core recorded it", async () => { const fetch = vi.fn().mockResolvedValue(response({ data: [node, unready] })); diff --git a/packages/agents-client/src/sandbox-client.ts b/packages/agents-client/src/sandbox-client.ts index 81735f221..ef1cafcb8 100644 --- a/packages/agents-client/src/sandbox-client.ts +++ b/packages/agents-client/src/sandbox-client.ts @@ -8,12 +8,10 @@ export type SandboxDiagnostic = "" | "node_unavailable" | "resource_missing" | " /** Checked against Core's shared node-diagnostics.json fixture. */ export const sandboxNodeDiagnostics = [ "provider_unavailable", - "docker_unavailable", - "docker_limits_unsupported", + "host_unsupported", + "artifacts_unavailable", "runtime_download_failed", "runtime_image_unavailable", - "kvm_unavailable", - "microsandbox_artifacts_unavailable", "capacity_insufficient", ] as const; /** Fixed reason a node's provider is not ready. Core omits the field while the provider is ready, so read it as falsy (undefined) then. The client reads an unknown future value as provider_unavailable. */ diff --git a/services/core/internal/deployment/node.go b/services/core/internal/deployment/node.go index 2509363da..485d464e8 100644 --- a/services/core/internal/deployment/node.go +++ b/services/core/internal/deployment/node.go @@ -46,7 +46,7 @@ type Enrollment struct { type NodeHealth struct { Host *NodeHost `json:"-"` // Fixed reason for the last reported unreadiness; absent while the provider is ready. Clients treat an unknown value as provider_unavailable. - Diagnostic string `json:"diagnostic,omitempty" enums:"provider_unavailable,docker_unavailable,docker_limits_unsupported,runtime_download_failed,runtime_image_unavailable,kvm_unavailable,microsandbox_artifacts_unavailable,capacity_insufficient"` + Diagnostic string `json:"diagnostic,omitempty" enums:"provider_unavailable,host_unsupported,artifacts_unavailable,runtime_download_failed,runtime_image_unavailable,capacity_insufficient"` ProviderReady bool `json:"provider_ready"` CPUCount *int64 `json:"cpu_count"` AvailableMemoryBytes *int64 `json:"available_memory_bytes"` diff --git a/services/core/internal/deployment/rules_test.go b/services/core/internal/deployment/rules_test.go index 1c4819293..84be8e066 100644 --- a/services/core/internal/deployment/rules_test.go +++ b/services/core/internal/deployment/rules_test.go @@ -263,7 +263,7 @@ func TestNodeRollout(t *testing.T) { {"failed without diagnostic", NodeRecord{Online: true, ProtocolVersion: 2, TargetState: "failed", ReadyGeneration: &ready}, "failed", ""}, {"ready ignores a diagnostic", NodeRecord{Online: true, ProtocolVersion: 2, TargetState: "ready", TargetDiagnostic: "boom", ReadyGeneration: &ready}, "ready", ""}, {"unknown target state", NodeRecord{Online: true, ProtocolVersion: 2, TargetState: "other", ReadyGeneration: &ready}, "unknown", ""}, - {"protocol 1 failed on the target", NodeRecord{Online: true, ProtocolVersion: 1, DeploymentGeneration: 2, TargetGeneration: 2, TargetState: "failed", TargetDiagnostic: "kvm_unavailable", ReadyGeneration: &ready}, "failed", "kvm_unavailable"}, + {"protocol 1 failed on the target", NodeRecord{Online: true, ProtocolVersion: 1, DeploymentGeneration: 2, TargetGeneration: 2, TargetState: "failed", TargetDiagnostic: "host_unsupported", ReadyGeneration: &ready}, "failed", "host_unsupported"}, {"protocol 1 without a target state", NodeRecord{Online: true, ProtocolVersion: 1, DeploymentGeneration: 2, TargetGeneration: 2, ReadyGeneration: &ready}, "unknown", ""}, } { got := nodeRollout(c.n) diff --git a/services/core/internal/deployment/view.go b/services/core/internal/deployment/view.go index 228fc8e7e..4c6beb6bc 100644 --- a/services/core/internal/deployment/view.go +++ b/services/core/internal/deployment/view.go @@ -46,7 +46,7 @@ type NodeRollout struct { State string `json:"state" enums:"ready,preparing,failed,update_required,unknown"` // Durable serving-generation pin; online and provider_ready still gate placement. ReadyGeneration *uint64 `json:"ready_generation" extensions:"x-nullable"` - Diagnostic string `json:"diagnostic,omitempty" enums:"provider_unavailable,docker_unavailable,docker_limits_unsupported,runtime_download_failed,runtime_image_unavailable,kvm_unavailable,microsandbox_artifacts_unavailable,capacity_insufficient"` + Diagnostic string `json:"diagnostic,omitempty" enums:"provider_unavailable,host_unsupported,artifacts_unavailable,runtime_download_failed,runtime_image_unavailable,capacity_insufficient"` } type RolloutNodes struct { diff --git a/services/core/internal/persistence/postgres/deploymentpg/nodes_test.go b/services/core/internal/persistence/postgres/deploymentpg/nodes_test.go index a06cb97f0..a3c75e268 100644 --- a/services/core/internal/persistence/postgres/deploymentpg/nodes_test.go +++ b/services/core/internal/persistence/postgres/deploymentpg/nodes_test.go @@ -419,3 +419,45 @@ func TestNodeGenerationDowngradePreservesServingProtocol(t *testing.T) { }) } } + +// Stored vendor codes from before the readiness classes, including an offline +// node's last report, are rewritten to their class. +func TestNodeReadinessClassMigrationRewritesStoredCodes(t *testing.T) { + f := newFixture(t) + changes, _ := f.execution(t) + _, view := f.initialize(t, changes, sandbox.Selection{Provider: "docker", DeploymentSpec: testSpecification("docker")}) + node := f.enroll(t, view, deployment.Capacity{MaxActive: 1, MaxRetained: 1}) + connection := f.connect(t, node.NodeID) + failed := []sandbox.GenerationStatus{{Generation: view.Generation, SpecificationDigest: view.SpecificationDigest, State: "failed", Diagnostic: "provider_unavailable"}} + if err := f.service.HeartbeatGenerations(t.Context(), node.NodeID, connection, view.OwnerEpoch, deployment.NodeHealth{Diagnostic: "provider_unavailable"}, failed); err != nil { + t.Fatal(err) + } + db := sql.OpenDB(stdlib.GetConnector(*f.pool.Config().ConnConfig)) + defer db.Close() + migration, err := goose.NewProvider(goose.DialectPostgres, db, os.DirFS("../../../../migrations"), goose.WithTableName("agents_api_schema_version")) + if err != nil { + t.Fatal(err) + } + var version int64 + for _, source := range migration.ListSources() { + if strings.HasSuffix(source.Path, "_node_readiness_classes.sql") { + version = source.Version + } + } + if _, err := migration.DownTo(t.Context(), version-1); err != nil { + t.Fatal(err) + } + if _, err := f.pool.Exec(t.Context(), `UPDATE runtime_nodes SET health=jsonb_set(health,'{diagnostic}','"kvm_unavailable"') WHERE id=$1`, node.NodeID); err != nil { + t.Fatal(err) + } + if _, err := f.pool.Exec(t.Context(), "UPDATE runtime_node_generation_status SET diagnostic='microsandbox_artifacts_unavailable' WHERE node_id=$1", node.NodeID); err != nil { + t.Fatal(err) + } + if _, err := migration.Up(t.Context()); err != nil { + t.Fatal(err) + } + detail, err := f.service.NodeDetail(t.Context(), node.NodeID, "1h") + if err != nil || detail.Diagnostic != "host_unsupported" || detail.Rollout.State != "failed" || detail.Rollout.Diagnostic != "artifacts_unavailable" { + t.Fatal(detail.Diagnostic, detail.Rollout, err) + } +} diff --git a/services/core/internal/persistence/postgres/deploymentpg/presence_test.go b/services/core/internal/persistence/postgres/deploymentpg/presence_test.go index 7787174dd..7c8fd3905 100644 --- a/services/core/internal/persistence/postgres/deploymentpg/presence_test.go +++ b/services/core/internal/persistence/postgres/deploymentpg/presence_test.go @@ -270,13 +270,13 @@ func TestNodeDiagnosticReachesListAndDetail(t *testing.T) { reported, want string ready bool }{ - {reported: "docker_unavailable", want: "docker_unavailable"}, + {reported: "host_unsupported", want: "host_unsupported"}, {reported: "capacity_insufficient", want: "capacity_insufficient"}, // Older nodes send provider_unavailable or nothing; unknown text is never stored. {reported: "provider_unavailable", want: "provider_unavailable"}, {reported: "", want: ""}, {reported: "dial unix /var/run/docker.sock: permission denied", want: "provider_unavailable"}, - {reported: "kvm_unavailable", want: "", ready: true}, + {reported: "artifacts_unavailable", want: "", ready: true}, } { if err := f.service.Heartbeat(t.Context(), node.NodeID, connection, epoch, deployment.NodeHealth{ProviderReady: tc.ready, Diagnostic: tc.reported}); err != nil { t.Fatal(tc.reported, err) diff --git a/services/core/internal/sandbox/node/generations.go b/services/core/internal/sandbox/node/generations.go index 047bbe4e1..42bd03f8c 100644 --- a/services/core/internal/sandbox/node/generations.go +++ b/services/core/internal/sandbox/node/generations.go @@ -369,7 +369,7 @@ func (m *GenerationManager) probeLoop() { if err != nil { g.state = "failed" g.diagnostic = sandbox.NodeDiagnostic(err) - if errors.Is(err, sandbox.ErrRuntimeImageUnavailable) || errors.Is(err, sandbox.ErrMicrosandboxArtifactsUnavailable) { + if errors.Is(err, sandbox.ErrRuntimeImageUnavailable) || errors.Is(err, sandbox.ErrArtifactsUnavailable) { g.repairing = true } } diff --git a/services/core/internal/sandbox/node/generations_test.go b/services/core/internal/sandbox/node/generations_test.go index 8831d5321..bce4d88cd 100644 --- a/services/core/internal/sandbox/node/generations_test.go +++ b/services/core/internal/sandbox/node/generations_test.go @@ -364,7 +364,7 @@ func TestInterruptedCollectionNeverPreparesOrServesAfterRestart(t *testing.T) { } func TestPreparationDiagnosticPreservesTypedCause(t *testing.T) { - for _, cause := range []error{sandbox.ErrRuntimeDownloadFailed, sandbox.ErrDockerUnavailable, sandbox.ErrKVMUnavailable, sandbox.ErrOwnership, context.Canceled, errors.New("raw secret provider text")} { + for _, cause := range []error{sandbox.ErrRuntimeDownloadFailed, sandbox.ErrProviderUnavailable, sandbox.ErrHostUnsupported, sandbox.ErrOwnership, context.Canceled, errors.New("raw secret provider text")} { t.Run(sandbox.NodeDiagnostic(cause)+cause.Error(), func(t *testing.T) { m, err := NewGenerationManager(t.Context(), GenerationManagerOptions{ Prepare: func(context.Context, uint64, string) (GenerationProvider, error) { return GenerationProvider{}, cause }, diff --git a/services/core/internal/sandbox/node/node_test.go b/services/core/internal/sandbox/node/node_test.go index 9aca8cd5a..946470389 100644 --- a/services/core/internal/sandbox/node/node_test.go +++ b/services/core/internal/sandbox/node/node_test.go @@ -302,7 +302,7 @@ func TestHealthSendsOnlyFixedDiagnosticCode(t *testing.T) { err error want string }{ - {fmt.Errorf("%w: dial unix /home/operator/private/docker.sock", sandbox.ErrDockerUnavailable), "docker_unavailable"}, + {fmt.Errorf("%w: open /home/operator/private/kvm", sandbox.ErrHostUnsupported), "host_unsupported"}, {errors.New("open /home/operator/private/runtime: permission denied"), "provider_unavailable"}, {nil, ""}, } { diff --git a/services/core/internal/sandbox/node_diagnostic.go b/services/core/internal/sandbox/node_diagnostic.go index 00b1ed898..b13fbe34b 100644 --- a/services/core/internal/sandbox/node_diagnostic.go +++ b/services/core/internal/sandbox/node_diagnostic.go @@ -2,38 +2,42 @@ package sandbox import "errors" -// Node readiness failures. A node probe returns or wraps the one matching its -// first failed check. Only the fixed code crosses the node transport; the error -// text and any wrapped detail, such as host paths or daemon messages, stay local. +// Node readiness classes. A node probe returns or wraps the class of its first +// failed check and keeps the Provider's detail, such as host paths or daemon +// messages, in the local error text. Only the class code crosses the node +// transport. var ( - ErrRuntimeDownloadFailed = errors.New("Runtime preparation failed") - ErrDockerUnavailable = errors.New("Docker daemon is unavailable") - ErrDockerLimitsUnsupported = errors.New("Docker host does not enforce CPU and memory limits") - ErrRuntimeImageUnavailable = errors.New("pinned Runtime image is unavailable") - ErrKVMUnavailable = errors.New("KVM is unavailable to sandbox node") - ErrMicrosandboxArtifactsUnavailable = errors.New("pinned microsandbox artifacts are unavailable") - ErrCapacityInsufficient = errors.New("node cannot provide one sandbox of the deployment specification") + // The Provider's native service is unreachable or does not answer. + ErrProviderUnavailable = errors.New("sandbox provider unavailable") + // The host lacks a capability the Provider requires. + ErrHostUnsupported = errors.New("host lacks a capability the sandbox provider requires") + // A pinned native artifact is missing or fails its integrity check. + ErrArtifactsUnavailable = errors.New("pinned provider artifacts are unavailable") + // The exact Runtime artifacts could not be transferred or verified. + ErrRuntimeDownloadFailed = errors.New("Runtime preparation failed") + // The pinned Runtime image is not available to the Provider. + ErrRuntimeImageUnavailable = errors.New("pinned Runtime image is unavailable") + // The host cannot hold one sandbox of the deployment specification. + ErrCapacityInsufficient = errors.New("node cannot provide one sandbox of the deployment specification") ) -// NodeProviderUnavailable reports every readiness failure without a fixed cause. +// NodeProviderUnavailable also reports every readiness failure without a class. const NodeProviderUnavailable = "provider_unavailable" -const NodeRuntimeDownloadFailed = "runtime_download_failed" var nodeDiagnostics = []struct { err error code string }{ - {ErrRuntimeDownloadFailed, NodeRuntimeDownloadFailed}, - {ErrDockerUnavailable, "docker_unavailable"}, - {ErrDockerLimitsUnsupported, "docker_limits_unsupported"}, + {ErrProviderUnavailable, NodeProviderUnavailable}, + {ErrHostUnsupported, "host_unsupported"}, + {ErrArtifactsUnavailable, "artifacts_unavailable"}, + {ErrRuntimeDownloadFailed, "runtime_download_failed"}, {ErrRuntimeImageUnavailable, "runtime_image_unavailable"}, - {ErrKVMUnavailable, "kvm_unavailable"}, - {ErrMicrosandboxArtifactsUnavailable, "microsandbox_artifacts_unavailable"}, {ErrCapacityInsufficient, "capacity_insufficient"}, } -// NodeDiagnostic maps a readiness probe result to its fixed code: empty when -// ready and provider_unavailable when no classified cause is wrapped. +// NodeDiagnostic maps a readiness probe result to its class code: empty when +// ready and provider_unavailable when no class is wrapped. func NodeDiagnostic(err error) string { if err == nil { return "" @@ -49,7 +53,7 @@ func NodeDiagnostic(err error) string { // NormalizeNodeDiagnostic keeps an empty or known code. Any other reported value // becomes provider_unavailable, so Core never stores node-supplied text. func NormalizeNodeDiagnostic(code string) string { - if code == "" || code == NodeProviderUnavailable { + if code == "" { return code } for _, d := range nodeDiagnostics { diff --git a/services/core/internal/sandbox/node_diagnostic_test.go b/services/core/internal/sandbox/node_diagnostic_test.go index e7a9602eb..84b199cd6 100644 --- a/services/core/internal/sandbox/node_diagnostic_test.go +++ b/services/core/internal/sandbox/node_diagnostic_test.go @@ -18,7 +18,7 @@ func TestNodeDiagnosticContract(t *testing.T) { if err := json.Unmarshal(raw, &fixture); err != nil { t.Fatal(err) } - codes := []string{NodeProviderUnavailable} + var codes []string for _, diagnostic := range nodeDiagnostics { codes = append(codes, diagnostic.code) if got := NodeDiagnostic(fmt.Errorf("private probe detail: %w", diagnostic.err)); got != diagnostic.code { diff --git a/services/core/internal/sandbox/providers/generation_probe_test.go b/services/core/internal/sandbox/providers/generation_probe_test.go index 83e9dfe28..0c0004e09 100644 --- a/services/core/internal/sandbox/providers/generation_probe_test.go +++ b/services/core/internal/sandbox/providers/generation_probe_test.go @@ -73,7 +73,7 @@ cat "$MSB_HOME/output" } func TestMicrosandboxGenerationProbeRetainsEarlierFailures(t *testing.T) { - for _, failure := range []error{context.Canceled, sandbox.ErrKVMUnavailable, sandbox.ErrCapacityInsufficient, sandbox.ErrMicrosandboxArtifactsUnavailable, errors.New("microsandbox state directory is unavailable")} { + for _, failure := range []error{context.Canceled, sandbox.ErrHostUnsupported, sandbox.ErrCapacityInsufficient, sandbox.ErrArtifactsUnavailable, errors.New("microsandbox state directory is unavailable")} { probe := microsandboxGenerationProbe(Microsandbox{RuntimePath: "/missing"}, func(context.Context) error { return failure }) if got := probe(t.Context()); got != failure { t.Fatalf("earlier failure changed: got %v, want %v", got, failure) diff --git a/services/core/internal/sandbox/providers/probe.go b/services/core/internal/sandbox/providers/probe.go index 4ead40c4b..c7e34316a 100644 --- a/services/core/internal/sandbox/providers/probe.go +++ b/services/core/internal/sandbox/providers/probe.go @@ -19,11 +19,12 @@ import ( "github.com/moby/moby/client" ) -// Probes report the first failed check. Precedence runs from the provider -// platform (Docker daemon or KVM), through Docker limit support and host -// capacity for one sandbox of the deployment specification, to the installed -// Runtime content (image or microsandbox artifacts). Unclassified failures stay -// provider_unavailable. The returned text is local; only its code is reported. +// Probes report the first failed check as its readiness class. Precedence runs +// from the provider platform (Docker daemon or KVM), through Docker limit +// support and host capacity for one sandbox of the deployment specification, to +// the installed Runtime content (image or microsandbox artifacts). Unclassified +// failures stay provider_unavailable. The returned text is local; only its class +// is reported. // kvmDevice is replaceable only by tests. var kvmDevice = "/dev/kvm" @@ -31,14 +32,14 @@ var kvmDevice = "/dev/kvm" func dockerProbe(c *client.Client, image string, resources sandbox.Resources) func(context.Context) error { return func(ctx context.Context) error { if _, err := c.Ping(ctx, client.PingOptions{}); err != nil { - return sandbox.ErrDockerUnavailable + return fmt.Errorf("%w: Docker daemon is unreachable", sandbox.ErrProviderUnavailable) } host, err := c.Info(ctx, client.InfoOptions{}) if err != nil { - return fmt.Errorf("%w: cannot inspect Docker host resource support", sandbox.ErrDockerUnavailable) + return fmt.Errorf("%w: cannot inspect Docker host resource support", sandbox.ErrProviderUnavailable) } if !host.Info.MemoryLimit || !host.Info.CPUCfsQuota { - return sandbox.ErrDockerLimitsUnsupported + return fmt.Errorf("%w: Docker does not enforce CPU and memory limits", sandbox.ErrHostUnsupported) } if host.Info.MemTotal <= 0 { return errors.New("Docker host memory capacity is unavailable") @@ -50,7 +51,7 @@ func dockerProbe(c *client.Client, image string, resources sandbox.Resources) fu return sandbox.ErrRuntimeImageUnavailable } else if err != nil { // The daemon did not answer; the image may still be present. - return fmt.Errorf("%w: cannot inspect the pinned Runtime image", sandbox.ErrDockerUnavailable) + return fmt.Errorf("%w: cannot inspect the pinned Runtime image", sandbox.ErrProviderUnavailable) } return nil } @@ -68,11 +69,11 @@ func microsandboxProbe(entry Microsandbox, resources sandbox.Resources) func(con return err } if runtime.GOOS != "linux" { - return fmt.Errorf("%w: microsandbox requires a Linux KVM node", sandbox.ErrKVMUnavailable) + return fmt.Errorf("%w: microsandbox requires a Linux KVM node", sandbox.ErrHostUnsupported) } kvm, err := os.OpenFile(kvmDevice, os.O_RDWR, 0) if err != nil { - return sandbox.ErrKVMUnavailable + return fmt.Errorf("%w: KVM is unavailable to sandbox node", sandbox.ErrHostUnsupported) } _ = kvm.Close() if err := hostCapacity(resources); err != nil { @@ -89,7 +90,7 @@ func microsandboxProbe(entry Microsandbox, resources sandbox.Resources) func(con integrity.Unlock() helper, err := os.Stat(entry.HelperPath) if err != nil || !helper.Mode().IsRegular() || helper.Mode().Perm()&0111 == 0 { - return fmt.Errorf("%w: microsandbox helper is unavailable", sandbox.ErrMicrosandboxArtifactsUnavailable) + return fmt.Errorf("%w: microsandbox helper is unavailable", sandbox.ErrArtifactsUnavailable) } home, err := os.Lstat(entry.RuntimeHome) if err != nil || !home.IsDir() || home.Mode().Perm() != 0700 { @@ -103,13 +104,13 @@ func verifyMicrosandboxArtifacts(entry Microsandbox) error { for _, artifact := range []struct{ path, hash string }{{entry.RuntimePath, entry.RuntimeSHA256}, {entry.FirmwarePath, entry.FirmwareSHA256}} { f, err := os.Open(artifact.path) if err != nil { - return sandbox.ErrMicrosandboxArtifactsUnavailable + return fmt.Errorf("%w: microsandbox Runtime or firmware is missing", sandbox.ErrArtifactsUnavailable) } h := sha256.New() _, err = io.Copy(h, f) _ = f.Close() if err != nil || hex.EncodeToString(h.Sum(nil)) != artifact.hash { - return fmt.Errorf("%w: artifact integrity check failed", sandbox.ErrMicrosandboxArtifactsUnavailable) + return fmt.Errorf("%w: microsandbox artifact integrity check failed", sandbox.ErrArtifactsUnavailable) } } return nil diff --git a/services/core/internal/sandbox/providers/probe_test.go b/services/core/internal/sandbox/providers/probe_test.go index e8c25f02e..b258aaf39 100644 --- a/services/core/internal/sandbox/providers/probe_test.go +++ b/services/core/internal/sandbox/providers/probe_test.go @@ -120,9 +120,9 @@ func TestMicrosandboxProbeDiagnostics(t *testing.T) { want string }{ // KVM is reported before capacity and missing artifacts. - {name: "kvm", resources: sandbox.Resources{CPUs: 255, MemoryMiB: 1048576}, want: "kvm_unavailable"}, + {name: "kvm", resources: sandbox.Resources{CPUs: 255, MemoryMiB: 1048576}, want: "host_unsupported"}, {name: "capacity", kvm: true, resources: sandbox.Resources{CPUs: 255, MemoryMiB: 1048576}, want: "capacity_insufficient"}, - {name: "artifacts", kvm: true, resources: smallSandbox, want: "microsandbox_artifacts_unavailable"}, + {name: "artifacts", kvm: true, resources: smallSandbox, want: "artifacts_unavailable"}, } { t.Run(tc.name, func(t *testing.T) { useKVM(t, tc.kvm) @@ -143,13 +143,13 @@ func TestDockerProbeDiagnostics(t *testing.T) { want string unreachable, infoFails bool }{ - {name: "unreachable", unreachable: true, want: "docker_unavailable"}, - {name: "info", infoFails: true, want: "docker_unavailable"}, + {name: "unreachable", unreachable: true, want: "provider_unavailable"}, + {name: "info", infoFails: true, want: "provider_unavailable"}, // The pinned image is also missing below; earlier checks take precedence. - {name: "limits", cpus: 8, imageStatus: 404, want: "docker_limits_unsupported"}, + {name: "limits", cpus: 8, imageStatus: 404, want: "host_unsupported"}, {name: "capacity", limits: true, cpus: 1, imageStatus: 404, want: "capacity_insufficient"}, {name: "image", limits: true, cpus: 8, imageStatus: 404, want: "runtime_image_unavailable"}, - {name: "image_inspect_fails", limits: true, cpus: 8, imageStatus: 500, want: "docker_unavailable"}, + {name: "image_inspect_fails", limits: true, cpus: 8, imageStatus: 500, want: "provider_unavailable"}, {name: "ready", limits: true, cpus: 8, imageStatus: 200}, } { t.Run(tc.name, func(t *testing.T) { @@ -201,7 +201,7 @@ func TestMicrosandboxProbeRecoversRepairedArtifacts(t *testing.T) { t.Fatal(err) } probe := microsandboxProbe(Microsandbox{HelperPath: artifact, RuntimePath: artifact, FirmwarePath: artifact, RuntimeSHA256: hex.EncodeToString(digest[:]), FirmwareSHA256: hex.EncodeToString(digest[:]), RuntimeHome: home}, smallSandbox) - if got := sandbox.NodeDiagnostic(probe(t.Context())); got != "microsandbox_artifacts_unavailable" { + if got := sandbox.NodeDiagnostic(probe(t.Context())); got != "artifacts_unavailable" { t.Fatalf("missing artifact diagnostic = %q", got) } if err := os.WriteFile(artifact, content, 0700); err != nil { diff --git a/services/core/internal/sandbox/testdata/node-diagnostics.json b/services/core/internal/sandbox/testdata/node-diagnostics.json index 109529c5f..4c885eec2 100644 --- a/services/core/internal/sandbox/testdata/node-diagnostics.json +++ b/services/core/internal/sandbox/testdata/node-diagnostics.json @@ -1,10 +1,8 @@ [ "provider_unavailable", - "docker_unavailable", - "docker_limits_unsupported", + "host_unsupported", + "artifacts_unavailable", "runtime_download_failed", "runtime_image_unavailable", - "kvm_unavailable", - "microsandbox_artifacts_unavailable", "capacity_insufficient" ] diff --git a/services/core/migrations/000095_node_readiness_classes.sql b/services/core/migrations/000095_node_readiness_classes.sql new file mode 100644 index 000000000..36cc27afb --- /dev/null +++ b/services/core/migrations/000095_node_readiness_classes.sql @@ -0,0 +1,19 @@ +-- +goose Up +-- Node readiness diagnostics are Provider-neutral classes. An offline node keeps +-- its last report, so rewrite every stored vendor code to its class. +UPDATE runtime_nodes SET health = jsonb_set(health, '{diagnostic}', to_jsonb(CASE health->>'diagnostic' + WHEN 'docker_unavailable' THEN 'provider_unavailable' + WHEN 'microsandbox_artifacts_unavailable' THEN 'artifacts_unavailable' + ELSE 'host_unsupported' END::text)) +WHERE health->>'diagnostic' IN ('docker_unavailable', 'docker_limits_unsupported', 'kvm_unavailable', 'microsandbox_artifacts_unavailable'); +UPDATE runtime_node_generation_status SET diagnostic = CASE diagnostic + WHEN 'docker_unavailable' THEN 'provider_unavailable' + WHEN 'microsandbox_artifacts_unavailable' THEN 'artifacts_unavailable' + ELSE 'host_unsupported' END +WHERE diagnostic IN ('docker_unavailable', 'docker_limits_unsupported', 'kvm_unavailable', 'microsandbox_artifacts_unavailable'); + +-- +goose Down +UPDATE runtime_nodes SET health = jsonb_set(health, '{diagnostic}', '"provider_unavailable"') +WHERE health->>'diagnostic' IN ('host_unsupported', 'artifacts_unavailable'); +UPDATE runtime_node_generation_status SET diagnostic = 'provider_unavailable' +WHERE diagnostic IN ('host_unsupported', 'artifacts_unavailable');