You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On the Azure Windows/WHP runners, a microVM guest can lose its TSC clocksource during cold boot, before any snapshot is requested. Linux's clocksource watchdog finds too much skew between tsc-early and refined-jiffies and marks the TSC unstable. The guest then stays on refined-jiffies with a periodic tick permanently.
On dev, nothing checks the clocksource for these guests, so they are snapshotted on refined-jiffies. #285 makes the #253 capture contract mandatory: the clocksource must be tsc or kvm-clock, and every online CPU must run a one-shot tick. A guest in this state can never meet it, so its capture is refused, and all three Windows/WHP jobs on #285 fail.
This is neither #211 nor a #253 recurrence. No restore happens, the two reproducible failures have one online CPU, and the skew builds up during boot.
Guest kernel: Linux 6.18.38, the same vmlinux as devrun 36794657146 (SHA-256 e2f4c0e0425b6992ec0d42078777953f3e6e428a52babcc52fff2c36cd2cb792). Xeon Platinum 8370C, Booting paravirtualized kernel on bare hardware, tsc_early_khz=2793439.
wait_for_capture_clock gives up after 10 s. The guest powers off with status 0xff at 11.7 s, and the test reports portb closed before the expected marker
nvx-snapshot: timed out waiting for a tsc or kvm-clock clocksource and one-shot ticks on every online CPU (clocksource: refined-jiffies), then snapshot was not captured within 40s
Artifacts:
nvx-microvm-tests-whp, which includes scratch-fresh-capture.log (expires 2026-12-30)
Ticks lost while interrupts are disabled during boot
Root cause
Azure WHP guests run a periodic tick under the watchdog for about a second
No TSC-deadline timer. The NVX LAPIC frequency patch prints APIC timer: using supplied frequency after calibrate_APIC_clock has already returned for TSC-deadline. Linux prints TSC deadline timer available only when it has one, and this log lacks that line. The OpenVMM pin made no difference: the timer was missing both when the pin hid it and when it exposed it.
The TSC watchdog stays enabled. Linux turns it off only when the TSC is constant and nonstop and TSC_ADJUST is present. The watchdog checked tsc-early here, so at least one of those features is missing.
tsc registers late.tsc_early_khz doesn't mark the TSC frequency as known (tsc.c). So init_tsc_clocksource schedules a refinement, which registers tsc only HZ jiffies later. Lost ticks stretch that delay.
Until then, the tick stays periodic.tsc-early is watched by refined-jiffies, which can't make it valid for high-resolution mode. Jiffies therefore count only the timer interrupts that are actually delivered.
The allowed skew is 64 ms:32 ms for tsc-early plus 32 ms for jiffies.
This is the tsc-early window that #253 analyzed. #253 keeps captures out of it, but every cold boot still passes through it.
intel_pstate MSR faults print with interrupts disabled
intel_pstate reads MSR_TURBO_RATIO_LIMIT (0x1ad) in core_get_turbo_pstate and writes MSR_IA32_PERF_CTL (0x199) in intel_pstate_set_pstate. Both go through rdmsrq_on_cpu() and wrmsrq_on_cpu(). On the local CPU, generic_exec_single runs the access with interrupts disabled.
WHP faults both accesses. OpenVMM logs invalid msr read msr=0x1ad and invalid msr write msr=0x199. ex_handler_msr then prints a warning and a full stack trace from that context.
The kernel log goes to hvc0. The xe9 HVC driver writes each byte with an outb to port 0xe9, and each outb exits to OpenVMM.
No tick ran while either trace printed. All lines of each trace share a timestamp. Until sched_clock is marked stable, sched_clock_local can't run more than one 10 ms tick past the last tick it saw. The first warning line is stamped 1.613812, and the rest of that trace is frozen at 1.623507.
The traces kept interrupts masked for about 200 ms.
The 898-byte RDMSR trace took 85.4 ms, from 1.613812 to the next message at 1.699255. The same trace takes 36.6 ms on a bare-metal WHP host.
At the runner's rate, the 1,291-byte WRMSR trace takes about 123 ms. That matches the 117.9 ms between its frozen timestamp and the next message.
That is about 20 tick periods, but the LAPIC holds only one pending timer interrupt per masked span. Roughly 18 ticks are lost. At 1.90 s, the watchdog sees tsc-early advance 674.5 ms while jiffies advance 500 ms: a 174.5 ms skew.
#253 and #211 call these MSR traces unrelated, and in quiet boots they are: with loglevel=0, they never reach the console. In these two console-logging boots, they cause the skew.
The contract can never be met afterward
Once the TSC is unstable, the refinement never registers tsc. The guest stays on refined-jiffies with a periodic tick for the rest of its life.
The OpenVMM test waits 10 s and powers off with status 0xff.
Both refusals are correct. A direct capture request would also be rejected by OpenVMM's new periodic-LAPIC check.
Why dev is green
The boot is the same.dev and d541ad2 use the same vmlinux. The OpenVMM change 885fde3...5801894 doesn't touch virt_whp, because 5801894 reverted the TSC-deadline hiding from d984976. On the runners, both boots are identical up to the shell.
Nothing checked these guests' clocksource:
dev's nvx-snapshot has no clock check, and the fresh-scratch capture calls it directly.
Why #285's WHP change and bare-metal validation missed it
The pin change made no difference on the runners.d541ad2 blamed 37ec91d's failure on hiding TSC-deadline, and pinned 5801894 to keep TSC-deadline exposed on WHP. The runners never offer TSC-deadline. The same boot showed 173.8 ms of skew with it hidden and 174.5 ms with it exposed, at the same point in the boot.
The bare-metal host can't reach this state. The bare-metal WHP host used for validation (Xeon Silver 4114) gives the guest tsc_adjust, constant_tsc, nonstop_tsc, and tsc_deadline_timer. Linux turns off the TSC watchdog, so tsc-early is valid for high-resolution mode and ticks one-shot from early boot. A probe saw a one-shot lapic-deadline tick on tsc-early at 0.78 s. The targeted 10/10 runs and the full suite couldn't exercise this failure.
The design document assumes otherwise.snapshot-and-restore.md L549-L600 assumes that WHP retains TSC-deadline and can expose TSC_ADJUST. Neither is true on the Azure runners.
Open: the 2-vCPU shell-snapshot-restore capture
This guest also ended on refined-jiffies. Its boot is quiet, though, so the MSR traces never reach the console, and CI keeps no dmesg. It failed on e74a40e and d541ad2 and passed on 37ec91d, all on azure-windows-4, so an intermittent cause is involved. Candidates:
Ticks lost to something else in the same window, such as host descheduling, or hvc0 output with interrupts masked while the harness streams the SMP probe into the console.
These measurements used a bare-metal Windows/WHP host and the CI artifacts from run 36796290820: OpenVMM 5801894 with the same kernel and initramfs. clearcpuid=tsc_adjust lapic=notscdeadline approximates the runner's guest CPU features.
Guest
Clock state before tsc
tsc registered
Default
tsc_adjust, constant_tsc, nonstop_tsc, and tsc_deadline_timer; one-shot lapic-deadline tick on tsc-early
1.14 s
clearcpuid=tsc_adjust lapic=notscdeadline, quiet
Periodic lapic tick on tsc-early
1.22 s
Same, with the kernel log on the console
Periodic. About 6 ticks (~60 ms) were lost and then caught up at the one-shot switch, just under the 64 ms margin. The TSC stayed stable
1.76 s
Same, quiet, plus setcpuid=tsc_known_freq
tsc_known_freq; tsc registers at device_initcall, and the tick is one-shot before userspace
0.135 s
scratch-snapshot passed on the default guest, with NVX-SNAPSHOT-CLOCK: source=tsc waited_us=270000.
Stop the boot-time MSR faults that cause the two reproducible failures. Either drop intel_pstate and cpufreq from kernel/config-microvm, or have OpenVMM complete these MSR accesses without a #GP (mshv: incorrect msr filtering on intel openvmm#3462). This alone leaves the window open.
Validate WHP clock changes on the runners, or on bare metal with clearcpuid=tsc_adjust lapic=notscdeadline. A bare-metal WHP guest with TSC_ADJUST and TSC-deadline can't exercise this failure.
After choosing a fix, correct the WHP statements in the snapshot design document and the rationale in d541ad2's commit message.
Summary
On the Azure Windows/WHP runners, a microVM guest can lose its TSC clocksource during cold boot, before any snapshot is requested. Linux's clocksource watchdog finds too much skew between
tsc-earlyandrefined-jiffiesand marks the TSC unstable. The guest then stays onrefined-jiffieswith a periodic tick permanently.On
dev, nothing checks the clocksource for these guests, so they are snapshotted onrefined-jiffies. #285 makes the #253 capture contract mandatory: the clocksource must betscorkvm-clock, and every online CPU must run a one-shot tick. A guest in this state can never meet it, so its capture is refused, and all three Windows/WHP jobs on #285 fail.This is neither #211 nor a #253 recurrence. No restore happens, the two reproducible failures have one online CPU, and the skew builds up during boot.
Current incident
CI, run 36796290820, attempt 1,pull_requestfor Enforce safe microVM snapshot clock state #285d541ad289e7b29ee147d966df6b801bcc3bcafe75801894061f8b0fecc7f5a748596d8dc2c196427vmlinuxasdevrun 36794657146 (SHA-256e2f4c0e0425b6992ec0d42078777953f3e6e428a52babcc52fff2c36cd2cb792). Xeon Platinum 8370C,Booting paravirtualized kernel on bare hardware,tsc_early_khz=2793439.scratch-snapshotazure-windows-2tsc-earlyunstable at 1.90 s. Thennvx-snapshottimes out, and the harness reportsOpenVMM output did not reach EOF within 60sttrpc::test_ttrpc_microvm_linux_direct_lifecycle_and_snapshotazure-windows-1maxcpus=1 rdinit=/microvm-test, the kernel log on the console, and the NVX kernel viaOPENVMM_MICROVM_TEST_KERNELwait_for_capture_clockgives up after 10 s. The guest powers off with status0xffat 11.7 s, and the test reportsportb closed before the expected markershell-snapshot-restorewith 2 vCPUsazure-windows-4quiet loglevel=0nvx-snapshot: timed out waiting for a tsc or kvm-clock clocksource and one-shot ticks on every online CPU (clocksource: refined-jiffies), thensnapshot was not captured within 40sArtifacts:
nvx-microvm-tests-whp, which includesscratch-fresh-capture.log(expires 2026-12-30)openvmm-vmm-tests-Windows-whp-36796290820-1(expires 2026-10-08)benchmark-diagnostics-windows-whp-virtual-machine-36796290820-attempt-1(expires 2026-10-02)The first two failures reproduced on every #285 head. The third reproduced on two of the three:
scratch-fresh(azure-windows-2)shell-snapshot-restore(azure-windows-4)e74a40e, run 367841267407078c32azure-windows-1)37ec91d, run 36789320098d984976azure-windows-3)d541ad2, run 367962908205801894azure-windows-1)Decisive guest log
Excerpt from
scratch-fresh-capture.logatd541ad2:The log has no
TSC deadline timer availableline. The same guest on a bare-metal WHP host prints that line at 0.005 s.Why this is not #211 or #253
check_tsc_sync_sourcewhile APs come onlineRoot cause
Azure WHP guests run a periodic tick under the watchdog for about a second
APIC timer: using supplied frequencyaftercalibrate_APIC_clockhas already returned for TSC-deadline. Linux printsTSC deadline timer availableonly when it has one, and this log lacks that line. The OpenVMM pin made no difference: the timer was missing both when the pin hid it and when it exposed it.TSC_ADJUSTis present. The watchdog checkedtsc-earlyhere, so at least one of those features is missing.tscregisters late.tsc_early_khzdoesn't mark the TSC frequency as known (tsc.c). Soinit_tsc_clocksourceschedules a refinement, which registerstsconlyHZjiffies later. Lost ticks stretch that delay.tsc-earlyis watched byrefined-jiffies, which can't make it valid for high-resolution mode. Jiffies therefore count only the timer interrupts that are actually delivered.tsc-earlyplus 32 ms for jiffies.This is the
tsc-earlywindow that #253 analyzed. #253 keeps captures out of it, but every cold boot still passes through it.intel_pstateMSR faults print with interrupts disabledintel_pstatereadsMSR_TURBO_RATIO_LIMIT(0x1ad) incore_get_turbo_pstateand writesMSR_IA32_PERF_CTL(0x199) inintel_pstate_set_pstate. Both go throughrdmsrq_on_cpu()andwrmsrq_on_cpu(). On the local CPU,generic_exec_singleruns the access with interrupts disabled.invalid msr read msr=0x1adandinvalid msr write msr=0x199.ex_handler_msrthen prints a warning and a full stack trace from that context.hvc0. The xe9 HVC driver writes each byte with anoutbto port0xe9, and eachoutbexits to OpenVMM.sched_clockis marked stable,sched_clock_localcan't run more than one 10 ms tick past the last tick it saw. The first warning line is stamped 1.613812, and the rest of that trace is frozen at 1.623507.RDMSRtrace took 85.4 ms, from 1.613812 to the next message at 1.699255. The same trace takes 36.6 ms on a bare-metal WHP host.WRMSRtrace takes about 123 ms. That matches the 117.9 ms between its frozen timestamp and the next message.tsc-earlyadvance 674.5 ms while jiffies advance 500 ms: a 174.5 ms skew.#253 and #211 call these MSR traces unrelated, and in quiet boots they are: with
loglevel=0, they never reach the console. In these two console-logging boots, they cause the skew.The contract can never be met afterward
Once the TSC is unstable, the refinement never registers
tsc. The guest stays onrefined-jiffieswith a periodic tick for the rest of its life.nvx-snapshotwaits 5 s and fails.0xff.Both refusals are correct. A direct capture request would also be rejected by OpenVMM's new periodic-LAPIC check.
Why
devis greendevandd541ad2use the samevmlinux. The OpenVMM change885fde3...5801894doesn't touchvirt_whp, because5801894reverted the TSC-deadline hiding fromd984976. On the runners, both boots are identical up to the shell.dev'snvx-snapshothas no clock check, and the fresh-scratch capture calls it directly.dev's WHP waits in the capture controller andsnapshot-corereject onlytsc-early, sorefined-jiffiespasses.885fde3's ttrpc test has no clock wait.devalmost certainly hits the same fallback. With an identical boot, these two guests most likely snapshot onrefined-jiffiesondevtoo. That is harmless for ci: Linux/MSHV restore-processors marks restored tsc-early unstable via clocksource watchdog #253, because Linux stops watching a TSC it has marked unstable, but it violates the contract that Enforce safe microVM snapshot clock state #285 enforces. Passing jobs don't upload guest logs, so CI history can't confirm this.Why #285's WHP change and bare-metal validation missed it
d541ad2blamed37ec91d's failure on hiding TSC-deadline, and pinned5801894to keep TSC-deadline exposed on WHP. The runners never offer TSC-deadline. The same boot showed 173.8 ms of skew with it hidden and 174.5 ms with it exposed, at the same point in the boot.tsc_adjust,constant_tsc,nonstop_tsc, andtsc_deadline_timer. Linux turns off the TSC watchdog, sotsc-earlyis valid for high-resolution mode and ticks one-shot from early boot. A probe saw a one-shotlapic-deadlinetick ontsc-earlyat 0.78 s. The targeted 10/10 runs and the full suite couldn't exercise this failure.TSC_ADJUST. Neither is true on the Azure runners.Open: the 2-vCPU
shell-snapshot-restorecaptureThis guest also ended on
refined-jiffies. Its boot is quiet, though, so the MSR traces never reach the console, and CI keeps nodmesg. It failed one74a40eandd541ad2and passed on37ec91d, all onazure-windows-4, so an intermittent cause is involved. Candidates:hvc0output with interrupts masked while the harness streams the SMP probe into the console.TSC_ADJUST, Linux runs that check at cold boot too, and ci: investigate intermittent Linux/MSHV TSC warp in restore-tsc-sync #211's latest update found cold-boot warps onazure-kvm-4.Bare-metal measurements
These measurements used a bare-metal Windows/WHP host and the CI artifacts from run 36796290820: OpenVMM
5801894with the same kernel and initramfs.clearcpuid=tsc_adjust lapic=notscdeadlineapproximates the runner's guest CPU features.tsctscregisteredtsc_adjust,constant_tsc,nonstop_tsc, andtsc_deadline_timer; one-shotlapic-deadlinetick ontsc-earlyclearcpuid=tsc_adjust lapic=notscdeadline, quietlapictick ontsc-earlysetcpuid=tsc_known_freqtsc_known_freq;tscregisters atdevice_initcall, and the tick is one-shot before userspacescratch-snapshotpassed on the default guest, withNVX-SNAPSHOT-CLOCK: source=tsc waited_us=270000.Suggested next steps
Close the cold-boot
tsc-earlywindow on hosts that lack TSC-deadline orTSC_ADJUST. Iftscregisters atdevice_initcall, lost ticks become harmless, because one-shot ticks recompute jiffies from the clocksource. It also removes the capture wait. The candidates come from ci: Linux/MSHV restore-processors marks restored tsc-early unstable via clocksource watchdog #253's alternatives:tsc_early_khzas a known frequency, as in ci: Linux/MSHV restore-processors marks restored tsc-early unstable via clocksource watchdog #253's prototype. The last row of the measurements shows the same effect throughsetcpuid.tsc_early_khz, which works only on Intel.The current ci: Linux/MSHV restore-processors marks restored tsc-early unstable via clocksource watchdog #253 design avoids kernel patches, so this needs a decision.
Stop the boot-time MSR faults that cause the two reproducible failures. Either drop
intel_pstateand cpufreq fromkernel/config-microvm, or have OpenVMM complete these MSR accesses without a#GP(mshv: incorrect msr filtering on intel openvmm#3462). This alone leaves the window open.Make capture refusals explain themselves. When
nvx-snapshottimes out, it should print the kernel's clocksource and TSC messages. That would classify the next 2-vCPU failure as watchdog skew or a ci: investigate intermittent Linux/MSHV TSC warp in restore-tsc-sync #211 warp.Validate WHP clock changes on the runners, or on bare metal with
clearcpuid=tsc_adjust lapic=notscdeadline. A bare-metal WHP guest withTSC_ADJUSTand TSC-deadline can't exercise this failure.After choosing a fix, correct the WHP statements in the snapshot design document and the rationale in
d541ad2's commit message.Related
unchecked MSR access errortraces.