Skip to content

core: LocalBackend::cancelPending settles its calls before stopping their handlers - #883

Merged
Yaraslaut merged 1 commit into
masterfrom
fix/local-cancel-settles-before-stop
Oct 8, 2026
Merged

Yaraslaut merged 1 commit into
masterfrom
fix/local-cancel-settles-before-stop

Conversation

@Yaraslaut

Copy link
Copy Markdown
Member

Fixes #876.

What the nightly failure actually was

The Valgrind job reported no memory errors on any of the failing nights: ERROR SUMMARY: 0 errors in every binary. It failed because Catch2 tests failed in tests/test_coroutine_model.cpp:

  • 2026-10-07: line 879 and line 1460
  • 2026-10-05: line 854

Each of these is a 2 s wait for a backend switch, or a cancelPending, to give the caller a specific error type (BackendChangedError or DisconnectedError).

That error sometimes never arrived, because of an ordering race in LocalBackend::cancelPending:

  1. It requested stop on every running Task handler first.
  2. Only then did it settle the pending calls with the error it was given.

A stopped handler resumes on its strand thread and settles its own call with OperationCancelled. When that settle landed first, the caller got OperationCancelled, and the later settle was a no-op. docs/spec/concurrency_and_lifetimes.md says these verbs settle the call with the given error, so this is a library defect, not a test-timing problem.

The fix

  • Settle first, then request the stops. Whatever the stopped handler settles afterwards is now the ignored second settle.
  • docs/spec/concurrency_and_lifetimes.md gains a sentence saying why the order matters.

Evidence

Measured locally with clang 20.1.2 (Debug) and valgrind 3.22.0. Each run executed the three nightly-failing tests together, 40 runs, 4 at a time:

failed
before 23/40. Every failing run, and no passing one, logged one completion settled with OperationCancelled (temporary instrumentation in SwitchedFlag, not committed)
before, --fair-sched=yes 16/40, so Valgrind's thread scheduler is not the cause
before, wait budget raised to 60 s runs still gave up at the full 60 s, so it's a lost error, not a slow one
after 0/40

New regression test: "cancelPending answers a running Task handler's caller with its error and not with the stop". It runs the backend's strands on morph::testing::InlineExecutor, so the stopped handler settles inside request_stop() itself. That makes the race deterministic:

  • backend.hpp reverted to master: fails 50/50, holds<DisconnectedError>(answered) is false and the caller got OperationCancelled
  • with the fix: passes 50/50

Full morph_tests, natively, with the nightly job's filters (~[oom-injector] ~[issue108]): 1810 test cases, all pass. The one "failed as expected" is the existing [!shouldfail] replay-ledger control.

Natively, without Valgrind, the old order never failed in 200 single-CPU runs. So this showed up only on the Valgrind leg, where threads are serialised.

Not verified locally:

  • GCC 15, which the Valgrind job uses
  • clang 22
  • clang-tidy
  • clang-format 22 (clang-format 18 reports nothing on the changed files)

🤖 Generated with Claude Code

https://claude.ai/code/session_01WeDnphcJ7EZBLwVfE6pqZu


Generated by Claude Code

…heir handlers

cancelPending requested stop on every running Task handler and only then
settled the pending calls with the error it was given. A stopped handler
resumes on its strand and settles its own call with OperationCancelled;
when that settle won the race, the caller was answered with the stop
instead of BackendChangedError / DisconnectedError, which
concurrency_and_lifetimes.md says it gets. Settle first, then stop: the
handler's own settle is then the ignored second one.

This is what the nightly Valgrind job has been failing on. Three tests in
test_coroutine_model.cpp wait for that error type and timed out when the
other one arrived. Measured locally under valgrind 3.22, the three tests
together, 40 runs each:

  before: 23/40 failed; every failing run, and no passing one, logged a
          completion settled with OperationCancelled
  after:   0/40 failed

Raising the wait budget to 60 s did not help (runs still gave up at 60 s)
and neither did --fair-sched=yes (16/40), so it was neither slowness nor
valgrind's scheduler.

The new test runs the backend's strands on an inline executor, so the
stopped handler settles inside request_stop() itself. With the old order
it fails 50/50; with the fix it passes 50/50.

Fixes #876

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WeDnphcJ7EZBLwVfE6pqZu
@codecov

codecov Bot commented Oct 7, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@Yaraslaut
Yaraslaut merged commit e281559 into master Oct 8, 2026
39 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ci: nightly Valgrind memcheck failed

2 participants