From 1d5f9759e7e1fc17a04d4200dc2b4349e718172a Mon Sep 17 00:00:00 2001 From: Manuel de Brito Fontes Date: Tue, 8 Sep 2026 01:03:59 -0300 Subject: [PATCH] Ask QEMU whether it accepts the command line this package writes MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Nothing did. The device set is decided in qemu/devices.mak and the device and property names are typed out by hand in machine/machine.go — virtio-balloon-pci with free-page-reporting and deflate-on-oom, virtio-mem-pci with requested-size, vhost-vsock-pci, vmgenid, virtio-net-pci with romfile=, the ICH9-LPC globals, the four -sandbox restrictions — and the two halves move independently. The package's tests assert on the strings it produces, which proves it agrees with itself. A device dropped from the allowlist, a property renamed at the next version bump and a typo added here therefore all produce the same thing: a VM that does not start, found days later by whoever booted one, in an error that reads like a fault in whatever launched it. `task verify:args` builds nine Specs — with and without a memory file, with and without a memory ceiling, with and without a vsock, with disks, with a NIC, with a named CPU model, and one with all of it — and starts each under the TCG binary with -S, then requires an answer on QMP. "It parsed" would have been the weaker claim: a process that replies to query-status has already created and realized every device on its command line. Nine machines start, answer and are killed in 0.23 s together, which is less than one boot. Deliberately broken to check it goes red and says where: free-page-reporting misspelt gives `-device virtio-balloon-pci,free-page-reportng=on,...: Property 'virtio-balloon-pci.free-page-reportng' not found`, and virtio-mem-pci as virtio-mem-pcie gives `'virtio-mem-pcie' is not a valid device model name`, each under the full command line, one argument per line. **Three arguments are rewritten, and no others.** CI has no /dev/kvm and the binary a tenant's host runs refuses to start without it, so the question can only be put to the TCG build — and `accel=kvm` becomes `accel=tcg`, `-cpu host` becomes `max` (a KVM-only value by construction: "CPU model 'host' requires KVM or HVF"; max is the other model that derives its features from the silicon, so migratable=on stays under test), and a NIC's tap backend becomes a hub port. Each rewrite is printed, so how far the claim reaches is visible in the log rather than buried here. What that leaves uncovered, stated rather than hidden: - `tap,id=netN,fd=N,vhost=on`. The backend is a file descriptor the caller opened and a test cannot open one without CAP_NET_ADMIN. The device that references it is checked verbatim, romfile= and slot included, which is the half a release can get wrong. - A real named CPU model. enforce=on is precisely what stops one being checked under emulation — Skylake-Server-v4 gives ten warnings and then "TCG doesn't support requested features", exit 1, which is the flag doing its job. qemu64 is used instead, so what is checked is the package's own contribution: that enforce is a property this binary has. - vhost-vsock-pci on a machine that will not give up /dev/vhost-vsock. Both workflows modprobe it *and chmod it*, which is not decoration: the first run of this on a GitHub runner loaded the module, found the node, and failed with "Could not open '/dev/vhost-vsock': Permission denied" — indistinguishable, in a stat, from having it. The test opens the device rather than stat'ing it for the same reason, and a host that refuses names the device as unchecked instead of passing. The kernel is a 197-byte ELF with a PVH entry note, written by the test. Not the real one, deliberately: what is under test is the command line, and requiring vmlinux would tie this to the kernel build — tens of minutes — so it could only run in a lane that had one, which is the lane a machine/*.go change never reaches. QEMU still enters it through pvh.bin out of the firmware directory, so -L is exercised either way. Without the note it says "Error loading uncompressed kernel without PVH ELF Note", which is how the stub was checked. It runs in ci.yml, on every push, which is where machine/*.go is changed and where qemu.yml — filtered to qemu/** — never fires. ci.yml still builds nothing: `task qemu:fetch` unpacks the published runtime image for the version qemu/Dockerfile pins into the same _output/ tree a build writes, in seconds, and ends in the same qemu:verify. The one case that cannot work is a pull request bumping the pin, since publishing happens from main — and that pull request changes qemu/**, so qemu.yml builds the binary and runs the identical check against it, which is the other place this was added. `task test` is unchanged: with no binary in _output/ the test skips and says so. `task verify:args` refuses instead, because a gate that passes for want of a binary is worse than no gate. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_019GHRRFvbnZi5q3KSTvyW5a --- .github/workflows/ci.yml | 54 ++++ .github/workflows/qemu.yml | 31 ++ README.md | 8 +- Taskfile.yml | 23 ++ machine/qemu_accepts_test.go | 529 +++++++++++++++++++++++++++++++++++ qemu/Taskfile.yml | 33 +++ 6 files changed, 677 insertions(+), 1 deletion(-) create mode 100644 machine/qemu_accepts_test.go diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index c57e3d5..9394a24 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -5,6 +5,10 @@ name: CI # apt. Each has its own workflow, triggered by the files that can change it. What is left # here is fast and is the half that catches most mistakes — the machine definition, and # whether the scripts and Taskfiles parse at all. +# +# It does pull one of them: the published QEMU of the pinned version, seconds rather than +# minutes, so that the command line this repository builds can be handed to the binary that +# has to accept it. See the step. # A branch is tested through its pull request, and main when something lands on it. Not # both: an unfiltered `push:` fires alongside `pull_request:` for every push to a branch @@ -28,6 +32,8 @@ concurrency: permissions: contents: read + # To read back the QEMU this commit pins, rather than build it. See the step. + packages: read jobs: gate: @@ -62,3 +68,51 @@ jobs: # boot first found itself with no /init. - name: Build the tools run: task tools + + # `task test` proves the machine package agrees with itself. The steps below ask the + # QEMU binary, which is the half that can disagree: the device set is decided in + # qemu/devices.mak and the device and property names are written out by hand in + # machine/machine.go, and nothing connected the two. A dropped device, a property + # renamed at a version bump, or a typo all produced the same thing — a VM that does + # not start, found days later by whoever booted one. + # + # Here and not only in qemu.yml, which is the workflow that has a QEMU: qemu.yml + # fires on qemu/** alone, so a change to machine/machine.go — the likeliest way to + # get this wrong — never reaches it. + - name: Log in to GitHub Packages + uses: docker/login-action@v4 + with: + registry: ghcr.io + username: ${{ github.actor }} + password: ${{ secrets.GITHUB_TOKEN }} + + # Seconds, against the tens of minutes this workflow deliberately does not spend: + # the runtime image qemu.yml published for the version qemu/Dockerfile pins, unpacked + # into the same _output/ tree a build would write. + # + # continue-on-error, for one case that is not a fault: a pull request that bumps the + # pin asks for a version no one has published yet, because publishing happens from + # main. That pull request changes qemu/**, so qemu.yml builds the binary and runs the + # very same check against it — which is why this may say so and stop, and the step + # after it is skipped rather than run against some other QEMU. Any other pull failure + # skips the check too, and shows as a failed step in the run. + - name: Fetch the pinned QEMU + id: qemu + continue-on-error: true + run: task qemu:fetch + + # vhost-vsock is a kernel module, and the machine's vsock device cannot be created + # without /dev/vhost-vsock. The runner has the node and gives the user no access to + # it, which is why the chmod is not decoration: without it QEMU says "Could not open + # '/dev/vhost-vsock': Permission denied" and the device goes unchecked. Loaded here rather than stood in for: there is no backend + # that keeps the device line and drops the requirement, and a runner without it is + # told what went unchecked instead of being shown a pass. + - name: Load vhost_vsock + if: steps.qemu.outcome == 'success' + run: | + sudo modprobe vhost_vsock && sudo chmod 0666 /dev/vhost-vsock || + echo "no vhost_vsock here; the check will name vhost-vsock-pci as unchecked" + + - name: Check QEMU accepts the machine's arguments + if: steps.qemu.outcome == 'success' + run: task verify:args diff --git a/.github/workflows/qemu.yml b/.github/workflows/qemu.yml index 8723fc0..5e65a0b 100644 --- a/.github/workflows/qemu.yml +++ b/.github/workflows/qemu.yml @@ -47,6 +47,14 @@ jobs: - name: Set up Task uses: go-task/setup-task@v2 + # For the acceptance check at the end, which is a Go test. The runner's own Go is + # whatever the image happens to carry, and go.mod names the one this builds with. + - name: Set up Go + uses: actions/setup-go@v7 + with: + go-version-file: go.mod + cache: true + - name: Set up Buildx uses: docker/setup-buildx-action@v4 @@ -70,6 +78,29 @@ jobs: QEMU_CACHE_TO=type=gha,mode=max \ QEMU_JOBS=4 + # vhost-vsock is a kernel module, and the machine's vsock device cannot be created + # without /dev/vhost-vsock. The runner has the node and gives the user no access to + # it, which is why the chmod is not decoration: without it QEMU says "Could not open + # '/dev/vhost-vsock': Permission denied" and the device goes unchecked. Loaded rather than stood in for: there is no backend that + # keeps the device line and drops the requirement, and a runner without it is told + # what went unchecked instead of being shown a pass. + - name: Load vhost_vsock + run: | + sudo modprobe vhost_vsock && sudo chmod 0666 /dev/vhost-vsock || + echo "no vhost_vsock here; the check will name vhost-vsock-pci as unchecked" + + # The other half of what a device allowlist decides. `task qemu:build` ends in + # qemu:verify, which asks this binary which accelerators and firmware files it has; + # this asks whether it accepts the command line the machine package builds, which is + # where a device removed from qemu/devices.mak actually surfaces — as a VM that does + # not start, weeks later, reading like a fault in whatever launched it. + # + # ci.yml runs the same check on every push, against the published binary. This is the + # lane that catches a change made on *this* side, before the binary it would break is + # published at all. + - name: Check the built QEMU accepts the machine's arguments + run: task verify:args + - name: Upload the extracted binaries uses: actions/upload-artifact@v4 with: diff --git a/README.md b/README.md index 5f09d37..8c2efbe 100644 --- a/README.md +++ b/README.md @@ -34,6 +34,7 @@ task build # everything, into _output/ task shell # boot the machine and look around inside it task lint # gofmt, vet, and whether the scripts and Taskfiles parse task test # the machine definition +task verify:args # and whether the QEMU in _output/ accepts what it produces task fingerprint # this machine's identity task release # one tarball, one version ``` @@ -334,7 +335,12 @@ pins the toolchain, `image/build.sh` is the build, and the task runs it with Five workflows, and the split is about cost. `ci.yml` runs on every push and builds none of the three artefacts — it is `task lint` and `task test`, which is fast and catches most -mistakes. Each artefact has a path-triggered workflow of its own, because QEMU and the +mistakes. It also *pulls* one: `task qemu:fetch` unpacks the published QEMU of the pinned +version in seconds, and `task verify:args` hands it the command line `machine.Spec.Args` +builds. The device set is decided in `qemu/devices.mak` and the device and property names +are written out by hand in `machine/machine.go`; nothing else connects the two, and a +device dropped from the allowlist reads exactly like a typo added here — a VM that does not +start, found by whoever boots one. Each artefact has a path-triggered workflow of its own, because QEMU and the kernel are tens of minutes each and the base image is ~1.5 GB of apt: `qemu.yml`, `kernel.yml`, `image.yml`. Each ends in the verification that belongs to it, so a build that lost a device, a kernel that lost its PVH notes, or an image that grew an identity diff --git a/Taskfile.yml b/Taskfile.yml index c433a26..7d790e4 100644 --- a/Taskfile.yml +++ b/Taskfile.yml @@ -182,6 +182,29 @@ tasks: done echo "OK: the shell scripts parse" + verify:args: + desc: >- + Ask the QEMU in _output/ whether it accepts the command line the machine package + builds. Every interesting Spec is started under the TCG binary, stopped before its + first instruction, and required to answer on QMP. + cmds: + # It crosses two parts — the machine definition and the binary — so it is here rather + # than in qemu/Taskfile.yml, and it refuses rather than skipping: the test itself + # skips when there is no binary, which is right for `go test ./...` in a source + # checkout and would be a gate that passes for the wrong reason here. + - | + set -euo pipefail + test -x {{.OUTPUT_ABS}}/bin/qemu-system-x86_64-tcg || { + echo "no {{.OUTPUT_ABS}}/bin/qemu-system-x86_64-tcg to ask." >&2 + echo " task qemu:fetch the published build of the pinned version, seconds" >&2 + echo " task qemu:build from source, tens of minutes" >&2 + exit 1; } + # -count=1 because the answer depends on a file the test cache does not know about: + # a rebuilt QEMU with a device removed would otherwise be met with a cached pass. + # -v because what was rewritten for the TCG binary, and what a host without + # /dev/vhost-vsock left unchecked, is the part a reader has to see. + - SPIN_MACHINE_OUTPUT={{.OUTPUT_ABS}} go test ./machine -run TestQEMUAcceptsEveryArgument -count=1 -v + fingerprint: desc: >- Print this machine's identity: the hash of the QEMU binary, the kernel and the initrd diff --git a/machine/qemu_accepts_test.go b/machine/qemu_accepts_test.go new file mode 100644 index 0000000..7602a4c --- /dev/null +++ b/machine/qemu_accepts_test.go @@ -0,0 +1,529 @@ +// SPDX-License-Identifier: Apache-2.0 + +package machine + +import ( + "cmp" + "encoding/binary" + "encoding/json" + "fmt" + "net" + "os" + "os/exec" + "path/filepath" + "strings" + "testing" + "time" +) + +// The rest of this package's tests assert on the strings Args produces, which +// proves the package agrees with itself. This one asks the binary. +// +// Every device and property named in machine.go — free-page-reporting, +// deflate-on-oom, requested-size, romfile=, the ICH9-LPC globals, the four +// -sandbox restrictions — is a string agreed with a QEMU built from +// qemu/devices.mak, and the two halves move independently: a device dropped from +// the allowlist, a property renamed at the next version bump, or a typo added +// here all produce the same thing, a VM that does not start, found days later by +// whoever booted one and reading like a fault in whatever launched it. +// +// It starts the machine rather than parsing it, because "it parsed" is a weaker +// claim than "it started": -S stops it before the first instruction, and a +// process that answers on the QMP monitor has already created and realized every +// device on its command line. Nine machines start, answer and are killed in +// 0.23 s together (measured 2026-09-08) — less than one boot. +// +// It is not part of `task test`: it needs a QEMU binary, and a source checkout +// has none. `task verify:args` is the gate, and it refuses rather than skips +// when the binary is missing. A developer running `go test ./...` with an empty +// _output/ gets the skip below. + +// tcgOnly names the three arguments that cannot mean on the TCG binary what they +// mean on the one a tenant's host runs, and is the whole of what this check +// changes about a command line before handing it over. Everything else is passed +// byte for byte as Args produced it. +// +// Each rewrite is reported and printed with -v, so a reader can see exactly how +// far the claim reaches — and each is here because the alternative is worse: +// +// - accel=kvm -> accel=tcg. The KVM-only binary refuses to start without +// /dev/kvm and CI has none, so the TCG build is the only one that can answer +// the question at all. Nothing else in the machine string moves; +// kernel-irqchip=on, hpet=off and acpi=on are accepted by both, which is +// itself worth knowing. +// +// - -cpu host -> max. "host" is a KVM-only value by construction — QEMU says +// "CPU model 'host' requires KVM or HVF" and exits — and max is the other +// model that derives its features from the silicon, so migratable=on, the +// property that only exists on those two, stays under test. +// +// - the tap netdev -> a hub port. A NIC's backend is a TAP file descriptor the +// caller opened, and a test cannot open one without CAP_NET_ADMIN, so +// `tap,id=netN,fd=N,vhost=on` is the one line here that is not checked. What +// is checked is the device that references it, verbatim: virtio-net-pci with +// romfile= (which is a NIC that cannot start if the option ROM is missing and +// the property is misspelt), the MAC, disable-legacy and the slot. +func tcgOnly(args []string) ([]string, []string, error) { + out := append([]string(nil), args...) + var rewrites []string + accel := false + + for i := 0; i+1 < len(out); i++ { + value := out[i+1] + switch out[i] { + case "-machine": + if !strings.Contains(value, "accel=kvm") { + continue + } + out[i+1] = strings.Replace(value, "accel=kvm", "accel=tcg", 1) + rewrites = append(rewrites, "-machine accel=kvm -> accel=tcg") + accel = true + case "-cpu": + model, rest, _ := strings.Cut(value, ",") + if model != "host" { + continue + } + out[i+1] = strings.TrimSuffix("max,"+rest, ",") + rewrites = append(rewrites, fmt.Sprintf("-cpu %s -> %s", value, out[i+1])) + case "-netdev": + if !strings.HasPrefix(value, "tap,") { + continue + } + id := "" + for _, field := range strings.Split(value, ",") { + if v, ok := strings.CutPrefix(field, "id="); ok { + id = v + } + } + if id == "" { + return nil, nil, fmt.Errorf("a tap netdev with no id: %q", value) + } + out[i+1] = fmt.Sprintf("hubport,id=%s,hubid=0", id) + rewrites = append(rewrites, fmt.Sprintf("-netdev %s -> %s", value, out[i+1])) + } + } + + // Loudly, rather than by testing a machine nobody runs: if the machine string + // stops saying accel=kvm, this check has been quietly answering a different + // question. + if !accel { + return nil, nil, fmt.Errorf("no accel=kvm in the machine string, so there is nothing for the TCG binary to stand in for") + } + return out, rewrites, nil +} + +// TestQEMUAcceptsEveryArgument starts each interesting Spec under the TCG binary +// and requires it to reach the monitor. +func TestQEMUAcceptsEveryArgument(t *testing.T) { + qemu, firmware := qemuTCG(t) + kernel := pvhStub(t) + + // A memory file, and one large enough: memory-backend-file takes the size + // from the object and maps the file, so a short one is an error about the + // file rather than about the command line. + memFile := filepath.Join(t.TempDir(), "memory") + if err := os.WriteFile(memFile, nil, 0o600); err != nil { + t.Fatal(err) + } + if err := os.Truncate(memFile, 512<<20); err != nil { + t.Fatal(err) + } + + base := func() Spec { + return Spec{ + QEMU: qemu, + Kernel: kernel, + Firmware: firmware, + BootCPUs: 2, + Memory: Memory{SizeMB: 512}, + Cmdline: DefaultCmdline().String(), + } + } + + cases := []struct { + name string + // needsVsock is the one host resource this check cannot stand in for. + needsVsock bool + spec func(s Spec) Spec + }{{ + // The plainest machine there is: anonymous RAM, no growth, no vsock, no + // disk, no NIC. Everything unconditional in Args is in this one. + name: "plain", + spec: func(s Spec) Spec { return s }, + }, { + // The machine a template is taken from: RAM in a file, mapped shared so + // the pages the guest dirties reach the file the restores will read. + name: "template source", + spec: func(s Spec) Spec { + s.Memory.File, s.Memory.Shared = memFile, true + return s + }, + }, { + // And the machine one is restored into: the same file mapped private, + // with no machine state until QMP says where to load it from. + name: "restore target", + spec: func(s Spec) Spec { + s.Memory.File = memFile + s.IncomingDefer = true + return s + }, + }, { + // A memory ceiling, which is what adds virtio-mem and its second memory + // backend, and a vCPU ceiling, which is what adds maxcpus=. + name: "ceilings", + spec: func(s Spec) Spec { + s.Memory.MaxMB = 2048 + s.MaxCPUs = 8 + return s + }, + }, { + name: "vsock", + needsVsock: true, + spec: func(s Spec) Spec { + s.VsockCID = 12345 + return s + }, + }, { + // Both disks a VM here gets: the read-only base, opened by every VM on + // the host at once, and a writable overlay with a lock and a serial. + name: "disks", + spec: func(s Spec) Spec { + s.Disks = []Disk{ + {Path: rawDisk(t, "base.raw"), Format: "raw", Readonly: true, Serial: "base"}, + {Path: qcow2Disk(t, qemu, "overlay.qcow2"), Format: "qcow2", + Serial: "overlay", Locking: true, Cache: "writeback"}, + } + return s + }, + }, { + name: "nic", + spec: func(s Spec) Spec { + s.NICs = []NIC{{TapFD: 3, MAC: "52:54:00:12:34:56"}} + return s + }, + }, { + // The other half of Shape: a named model, which is enforce=on rather than + // migratable=on. + // + // qemu64 and not a microarchitecture anybody would deploy, because + // enforce=on is exactly what stops one from being checked here — it + // refuses to start a model the accelerator cannot fully provide, which is + // the whole reason it is on the command line, and TCG cannot provide any + // of them. Skylake-Server-v4 was tried: ten warnings and then "TCG doesn't + // support requested features", exit 1, which is the flag doing its job. + // What is left to check is the package's own contribution — that enforce + // is a property this binary has, on a model it knows — since which model + // is named is the caller's. + name: "named cpu", + spec: func(s Spec) Spec { + s.CPU = "qemu64" + return s + }, + }, { + // The shape a real VM has, all at once: file-backed RAM, a ceiling, a + // vsock, two disks, a NIC and a console. + name: "everything", + needsVsock: true, + spec: func(s Spec) Spec { + s.Memory.File, s.Memory.Shared = memFile, true + s.Memory.MaxMB, s.MaxCPUs = 2048, 8 + s.VsockCID = 12345 + s.Disks = []Disk{ + {Path: rawDisk(t, "everything.raw"), Format: "raw", Readonly: true, Serial: "base"}, + } + s.NICs = []NIC{{TapFD: 3, MAC: "52:54:00:12:34:56"}} + s.Serial = "file:" + filepath.Join(t.TempDir(), "console.log") + return s + }, + }} + + for _, c := range cases { + t.Run(c.name, func(t *testing.T) { + if c.needsVsock { + // The one device with a host dependency this check cannot stand + // in for: vhost-vsock is a kernel module, and there is no + // backend that keeps the device line and drops the requirement. + // The workflows load it and open it up; a machine that cannot is + // told what went unchecked rather than shown a pass. + // + // Opened and not stat'd, because the two ways this fails are + // different and only one of them is "no module": a GitHub runner + // has the device node and hands the user no access to it, and a + // stat cannot tell the difference — measured 2026-09-08, where it + // reached QEMU as "Could not open '/dev/vhost-vsock': Permission + // denied" and read like a rejected argument. + dev, err := os.OpenFile("/dev/vhost-vsock", os.O_RDWR, 0) + if err != nil { + t.Skipf("vhost-vsock-pci went unchecked, because this host will not"+ + " give the device up: %v\n\t(modprobe vhost_vsock, and the node"+ + " has to be openable by whoever runs QEMU)", err) + } + _ = dev.Close() + } + + s := c.spec(base()) + s.QMPSocket = qmpSocket(t) + + args, err := s.Args() + if err != nil { + t.Fatal(err) + } + args, rewrites, err := tcgOnly(args) + if err != nil { + t.Fatal(err) + } + for _, r := range rewrites { + t.Logf("rewritten for the TCG binary: %s", r) + } + + // -S, and it is the only argument added: it stops the machine before + // the first instruction, so what is measured is device creation and + // not a guest booting under emulation. + start(t, qemu, append(args, "-S"), s.QMPSocket) + }) + } +} + +// start runs QEMU and requires it to answer on its monitor. +// +// Every failure here has to name the argument that was rejected, because a check +// whose failure says "exit status 1" costs more than it saves. QEMU names it +// itself — "-device virtio-balloon-pci,...: Property '...' not found" — so its +// stderr is reproduced whole, under the command line that produced it, one +// argument per line. +func start(t *testing.T, qemu string, args []string, socket string) { + t.Helper() + + var stderr strings.Builder + cmd := exec.Command(qemu, args...) // #nosec G204 -- the binary and arguments this test built + cmd.Stderr = &stderr + cmd.Stdout = &stderr + if err := cmd.Start(); err != nil { + t.Fatalf("starting %s: %v", qemu, err) + } + + // Closed rather than sent to, so that both the loop below and the deferred + // kill can wait on it. Sent to, the first receive took the value and the + // second blocked for the whole test timeout — which is how a rejected + // argument presented as a hang instead of the error QEMU had already printed. + var exit error + done := make(chan struct{}) + go func() { exit = cmd.Wait(); close(done) }() + + fail := func(what string, err error) { + t.Helper() + var b strings.Builder + fmt.Fprintf(&b, "%s: %v\n\n%s", what, err, qemu) + for _, a := range args { + fmt.Fprintf(&b, " \\\n %s", a) + } + if out := strings.TrimSpace(stderr.String()); out != "" { + fmt.Fprintf(&b, "\n\nQEMU said:\n%s", out) + } else { + b.WriteString("\n\nQEMU said nothing.") + } + t.Fatal(b.String()) + } + + defer func() { + _ = cmd.Process.Kill() + <-done + }() + + // The socket exists as soon as the chardev is created, which is before the + // CPU is created and long before any device is realized, so its existence + // proves nothing — a machine whose CPU model the accelerator refused left a + // listening socket behind and had already exited. An answer on it proves it. + deadline := time.Now().Add(20 * time.Second) + var conn net.Conn + for { + select { + case <-done: + // The usual failure: QEMU rejected an argument and exited. + fail("QEMU exited before answering on the monitor", exit) + default: + } + c, err := net.Dial("unix", socket) + if err == nil { + conn = c + break + } + if time.Now().After(deadline) { + fail("QEMU never answered on the monitor", err) + } + time.Sleep(20 * time.Millisecond) + } + defer func() { _ = conn.Close() }() + + if err := conn.SetDeadline(time.Now().Add(20 * time.Second)); err != nil { + t.Fatal(err) + } + dec := json.NewDecoder(conn) + + var greeting struct { + QMP *struct{} `json:"QMP"` + } + if err := dec.Decode(&greeting); err != nil { + fail("reading the QMP greeting", err) + } + if greeting.QMP == nil { + fail("the monitor did not greet", fmt.Errorf("no QMP field")) + } + + // query-status after the handshake, and not the greeting alone: the greeting + // is written by the chardev when the socket is accepted, whereas a reply to a + // command comes from the main loop, which is only reached once every device + // on the command line has been created and realized. That is the claim this + // check makes. + for _, command := range []string{"qmp_capabilities", "query-status"} { + if err := json.NewEncoder(conn).Encode(map[string]string{"execute": command}); err != nil { + fail("sending "+command, err) + } + var reply struct { + Return json.RawMessage `json:"return"` + Error *struct { + Desc string `json:"desc"` + } `json:"error"` + } + if err := dec.Decode(&reply); err != nil { + fail("reading the reply to "+command, err) + } + if reply.Error != nil { + fail(command+" failed", fmt.Errorf("%s", reply.Error.Desc)) + } + if command == "query-status" { + t.Logf("running, monitor answered query-status with %s", reply.Return) + } + } +} + +// qemuTCG is the emulating build and the firmware beside it. +// +// The TCG binary and not the one a host runs: that one has no TCG compiled in +// (qemu:verify asserts it) and refuses to start without /dev/kvm, which is a +// device CI does not have and a check must not require. +// +// The two paths out of a release tree, and not machine.Open, because Open asks +// for a whole machine and this needs half of one: the lane that runs this check +// on every push has a QEMU and deliberately no kernel and no base image, both of +// which are tens of minutes to build. +func qemuTCG(t *testing.T) (qemu, firmware string) { + t.Helper() + + // The default is the tree beside this package; the variable is what `task + // verify:args` passes, because OUTPUT_DIR can move and a check that looked at + // the old location would answer about a QEMU nobody is building any more. + out, err := filepath.Abs(cmp.Or(os.Getenv("SPIN_MACHINE_OUTPUT"), "../_output")) + if err != nil { + t.Fatal(err) + } + qemu = filepath.Join(out, qemuTCGName) + if _, err := os.Stat(qemu); err != nil { + t.Skipf("no QEMU at %s: %v\n\t`task verify:args` is the gate — it fetches one"+ + " and refuses rather than skipping", qemu, err) + } + return qemu, filepath.Join(out, firmwareDir) +} + +// pvhStub writes a kernel QEMU will load and never execute: an ELF with a PVH +// entry note, which is the only way into this machine. +// +// A stub and not the kernel from _output/kernel, deliberately. What is under +// test is the command line, and requiring the real kernel would tie this check +// to the kernel build — tens of minutes — so it could only ever run in a lane +// that had one, which is the lane a machine/*.go change does not reach. QEMU +// still loads it through pvh.bin out of the firmware directory, so -L and the +// option ROM that boot depends on are exercised either way. +// +// Without the note QEMU says "Error loading uncompressed kernel without PVH ELF +// Note" and exits, which is how this was checked. +func pvhStub(t *testing.T) string { + t.Helper() + + const entry = 0x100000 + le := binary.LittleEndian + + // One Xen ELF note: XEN_ELFNOTE_PHYS32_ENTRY (18), the 32-bit entry point. + note := make([]byte, 0, 20) + note = le.AppendUint32(note, 4) // namesz, "Xen\0" + note = le.AppendUint32(note, 4) // descsz, one address + note = le.AppendUint32(note, 18) // XEN_ELFNOTE_PHYS32_ENTRY + note = append(note, 'X', 'e', 'n', 0) + note = le.AppendUint32(note, entry) + + const phoff, phentsize, phnum = 64, 56, 2 + noteOff := uint64(phoff + phentsize*phnum) + textOff := noteOff + uint64(len(note)) + + elf := make([]byte, 0, 256) + elf = append(elf, 0x7f, 'E', 'L', 'F', 2, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0) + elf = le.AppendUint16(elf, 2) // ET_EXEC + elf = le.AppendUint16(elf, 62) // EM_X86_64 + elf = le.AppendUint32(elf, 1) // EV_CURRENT + elf = le.AppendUint64(elf, entry) + elf = le.AppendUint64(elf, phoff) + elf = le.AppendUint64(elf, 0) // no sections + elf = le.AppendUint32(elf, 0) // no flags + elf = le.AppendUint16(elf, 64) + elf = le.AppendUint16(elf, phentsize) + elf = le.AppendUint16(elf, phnum) + elf = le.AppendUint16(elf, 64) + elf = le.AppendUint16(elf, 0) + elf = le.AppendUint16(elf, 0) + + segment := func(typ, flags uint32, offset, addr, size uint64) { + elf = le.AppendUint32(elf, typ) + elf = le.AppendUint32(elf, flags) + elf = le.AppendUint64(elf, offset) + elf = le.AppendUint64(elf, addr) // vaddr + elf = le.AppendUint64(elf, addr) // paddr + elf = le.AppendUint64(elf, size) // filesz + elf = le.AppendUint64(elf, size) // memsz + elf = le.AppendUint64(elf, 4) // align + } + segment(4, 4, noteOff, 0, uint64(len(note))) // PT_NOTE + segment(1, 5, textOff, entry, 1) // PT_LOAD, r-x + elf = append(elf, note...) + elf = append(elf, 0xf4) // hlt, which -S means is never reached + + path := filepath.Join(t.TempDir(), "vmlinux-stub") + if err := os.WriteFile(path, elf, 0o600); err != nil { + t.Fatal(err) + } + return path +} + +func rawDisk(t *testing.T, name string) string { + t.Helper() + path := filepath.Join(t.TempDir(), name) + if err := os.WriteFile(path, make([]byte, 1<<20), 0o600); err != nil { + t.Fatal(err) + } + return path +} + +// qcow2Disk creates one with qemu-img, which is the tool a caller creates a +// VM's overlay with and sits beside the emulator in every release. +func qcow2Disk(t *testing.T, qemu, name string) string { + t.Helper() + path := filepath.Join(t.TempDir(), name) + img := filepath.Join(filepath.Dir(qemu), "qemu-img") + cmd := exec.Command(img, "create", "-f", "qcow2", path, "16M") // #nosec G204 -- beside the binary under test + if out, err := cmd.CombinedOutput(); err != nil { + t.Fatalf("%s create: %v\n%s", img, err, out) + } + return path +} + +// qmpSocket keeps the path short. A Unix socket address is 108 bytes including +// the terminator, and a test's temporary directory plus a socket name is close +// enough to it that QEMU has failed with "UNIX socket path is too long". +func qmpSocket(t *testing.T) string { + t.Helper() + dir, err := os.MkdirTemp("", "qmp") + if err != nil { + t.Fatal(err) + } + t.Cleanup(func() { _ = os.RemoveAll(dir) }) + return filepath.Join(dir, "s") +} diff --git a/qemu/Taskfile.yml b/qemu/Taskfile.yml index aa0f3b0..49eb8e5 100644 --- a/qemu/Taskfile.yml +++ b/qemu/Taskfile.yml @@ -11,6 +11,11 @@ vars: QEMU_VERSION: sh: sed -n 's/^ARG QEMU_VERSION=//p' qemu/Dockerfile + # Where the runtime image is published, which `fetch` below reads back. The tag is the + # pinned version, so asking for this repository's QEMU and asking for the one this commit + # says it runs are the same question. + QEMU_PUBLISHED_IMAGE: '{{.QEMU_PUBLISHED_IMAGE | default "ghcr.io/spin-stack/spin-machine/qemu"}}' + tasks: build: @@ -77,6 +82,34 @@ tasks: echo "OK: both qemu-system binaries run here, qemu-img runs, five firmware files," echo " and only the CI binary can emulate" + fetch: + desc: >- + Put the published QEMU of the pinned version into _output/, instead of building it. + Seconds against tens of minutes, and the same binaries — they are static, so the + image around them carries nothing they need. + cmds: + # The same tree `build` writes, so whatever reads _output/ cannot tell which of the + # two filled it. Emptied first for the reason build empties it: --output writes files + # and removes none, and a firmware blob left over from another version is + # indistinguishable from one this QEMU shipped. + - rm -rf {{.OUTPUT_ABS}}/bin/qemu-* {{.OUTPUT_ABS}}/qemu + - mkdir -p {{.OUTPUT_ABS}}/bin + - | + set -euo pipefail + image="{{.QEMU_PUBLISHED_IMAGE}}:{{.QEMU_VERSION}}" + docker pull --quiet --platform linux/amd64 "$image" + # create/cp and not `docker run cat`: the image's entrypoint is qemu-system-x86_64, + # and nothing has to start for a file to be copied out of a container that never ran. + cid=$(docker create --platform linux/amd64 "$image") + trap 'docker rm --force "$cid" > /dev/null' EXIT + for b in qemu-system-x86_64 qemu-system-x86_64-tcg qemu-img; do + docker cp "$cid:/usr/local/bin/$b" {{.OUTPUT_ABS}}/bin/$b + done + docker cp "$cid:/usr/share/spin-stack/qemu" {{.OUTPUT_ABS}}/qemu + # Which asserts the fetched tree is whole and that both binaries run here — the same + # question asked of a built one, and worth asking of a downloaded one twice over. + - task: verify + image: desc: >- Build the QEMU runtime image locally (--load, not pushed). A convenience for