Repository navigation
core, wire, net, qt: a hello-negotiated cancel envelope, honoured by RemoteServer and sent by both client backends (#864, #865) - #867
Merged
Conversation
… onto the run's stop source (#864) A `cancel` names the execute it stops by `cancelCallId` and carries its own `callId`, so the server's `ok` to it can never be matched against the still pending execute. The server advertises the kind in its `hello` reply (`ProtocolRange::capabilities`, "cancel"); a client sends it only after such a reply, so a server that predates it never sees one. Both members are additive, so this is a capability, not a kProtocolVersion bump. RemoteServer files each admitted Task execute that arrived on a connection scope, with what its admission decided on: the verified principal, the model and action types, the instance and its owner. A cancel is stamped (stampVerifiedPrincipal), looked up only under the connection it arrived on, and honoured only when its verified principal is the execute's and it passes the execute's own authorize and authorizeInstance. Every outcome — stopped, unknown, finished, another connection's, another principal's, refused, unscoped — answers the same `ok`, so a cancel cannot probe for calls. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…'s stop is requested (#865) Both backends learn from the opt-in negotiateProtocolVersion whether the server honours "cancel" (SocketBackend gains the verb), re-send hello after every reconnect and forget the capability on every disconnect. Only then do they send `cancel {cancelCallId}`, fire-and-forget under a fresh callId: - when a call's stop is requested (an execute deadline, or any holder of ActionCall::stopSource) — a stop callback registered once the call has its id and its frame is queued, which hands the cancel to the backend's owner; - for every execute cancelPending sweeps (~Bridge, switchBackend). The cancel carries the execute's own session, since the server honours it only under the execute's verified principal. Teardown lets the cancels out: SocketBackend's close is a close-after-flush on the loop, and ~QtWebSocketBackend flushes before its abort. Both discarded the frames ~Bridge's cancelPending had just queued. SocketBackend::cancelPending is posted to the I/O loop (G0 on return); the backend spec states why a cancel queued behind earlier loop work is acceptable: it can never overtake its execute, the delay is only to the server's handler, and the caller's own completion is settled in that task. Tests over real loopbacks (SocketServer + SocketBackend, QtWebSocketServer + QtWebSocketBackend): a client deadline, and destroying the Bridge, each make the server's Task handler observe the stop; a client that did not negotiate, or negotiated with a legacy server, sends none. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Yaraslaut
force-pushed
the
lane/wire-cancel
branch
from
October 4, 2026 10:58
f1c5003 to
c526c2c
Compare
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #864
Closes #865
Two commits, one per ticket, server side first; either can be dropped and the branch force-pushed without redoing the other.
#864 — server side (
core, wire)New
"cancel"kind. The target travels in a newEnvelope::cancelCallId; the cancel has its owncallId. Were the target incallId, the server'sokto the cancel would be matched by the client against the still-pending execute and settle it with an empty result.Negotiated through
hello:ProtocolRangegainscapabilities(the server lists"cancel"), withwire::helloAdvertises(reply, wire::kCapabilityCancel)to read it. Both members are additive, so this is a capability and not akProtocolVersionbump. A bump would have refused every v1 client under the default range. A legacy server never receives a cancel.RemoteServerfiles every admitted Task execute that arrived on a connection scope, together with the stop source and the facts its admission decided on: the verified principal, model and action type, instance, and owner. Acancelis handled in this order:stampVerifiedPrincipal.cancelCallId.authorizeandauthorizeInstance.Only then is
request_stop()called. Every outcome (stopped, unknown, finished, foreign connection, foreign principal, refused, unscoped) gets the sameok, so a cancel cannot probe for calls.Specs:
wire.md("Cancelling a call"),backend.md(envelope table, execute flow),security.md("Oncancel").Security tests (
tests/test_remote_cancel.cpp)cancel. The execute replieserr, and the cancel's reply carries its own callId.okwith its value.ok.Mutation checks (real output, from a probe binary built from
test_remote_cancel.cpp)stop->request_stop()→ no-op):stampVerifiedPrincipalremoved from the cancel branch: "bob's token claiming to be alice" goes red (lines 278 and 279), and so does the control (295).test cases: 8 | 7 passed | 1 failed.#865 — client side (
net, qt)negotiateProtocolVersion, which is opt-in.SocketBackendgains that verb, mirroring Qt's. After every reconnect they re-sendhello, and on every disconnect they forget the capability. They sendcancelonly when it was advertised. The cancel is fire-and-forget, under a fresh callId that is filed nowhere.ActionCall::stopSource). AStopCallbackis registered once the call has its id and its frame is queued. It only postsrequestCancel(callId)to the owner: the loop forSocketBackend, the Qt thread for Qt. The owner sends the cancel only if the call is still waiting for its reply.cancelPending(~Bridge,switchBackend). A cancel goes out for every drained execute.SocketBackend::closecleared the outbox, and~QtWebSocketBackendaborted the socket with the frames still unsent. Either way, the cancels~Bridge'scancelPendinghad just queued were dropped, and the~Bridgetests failed on both backends.SocketBackendnow uses close-after-flush, still bounded bysendTimeout. Qt now calls_socket.flush()beforeabort(), which does not block.SocketBackend::cancelPendingis posted to the loop (G0 on return), and that is acceptable:scripts/branch_partial_allowlist.json: thesocket_backend.hppdefault:hint moves from line 797 to 966. The source text is unchanged and unique.Tests (real loopback)
tests/net/test_socket_backend.cpp[cancel], withSocketServerandSocketBackend:cancelPendingis the markerderegister.tests/qt/test_qt_websocket.cpp[cancel], against a realQtWebSocketServer: the deadline test and the~Bridgetest. Qt 6 (Homebrew) was available locally, and both were run.Mutation checks (real output)
cancelPendingsend removed:~Bridgetest failed,0 == 1.~Bridgetest failed,0 == 1.Qt stop path: no templated
invokeMethod(clang-tidy-diff fix)CI's clang-tidy 22 with Qt 6.8.1 reported
clang-analyzer-cplusplus.NewDeleteLeaksatQtCore/qobjectdefs.h:624. The path ran through the functor overloadQMetaObject::invokeMethod(context, lambda, Qt::QueuedConnection)inCancelOnStop.The leak report is a Qt false positive:
invokeMethodCallableHelperputs itsNOLINTNEXTLINEon thenew QCallableObjectline, but the analyzer reports at thereturn invokeMethodImpl(...)two lines later.So the functor call is gone from the changed code instead of suppressed:
CancelOnStopsets a per-callstd::shared_ptr<std::atomic<bool>>flag.QMetaObject::invokeMethod(&_cancelWake, "start", Qt::QueuedConnection). In 6.8.1 that reachesinvokeMethodImpl(QObject*, const char*, ...)with nonew; I readqobjectdefs.hat v6.8.1, lines 372–389._cancelWakeis a zero-interval single-shotQTimer. Itstimeoutis connected once, in the constructor, todrainCancels(), the same lambda-connect pattern as the constructor's existing connects._pendingstays owner-only, with no lock.Why this shape: it is the smallest one that allocates no callable in changed code. A
postEvent(new QEvent)alternative would have needed an owning-memory NOLINT.Not reproduced locally. Homebrew Qt here is 6.11.1, whose headers already widen the suppression, so local clang-tidy 23 never showed the finding. It shows no
NewDeleteLeaksbefore or after this change. The claim that this fixes CI rests on the new code making no templatedinvokeMethod/singleShotcall.Re-measured after the change:
[cancel](deadline and~Bridge): green, 25 of 25 repeats.morph_qt_tests: 85 of 85 passed.test_qt_websocket.cpp:2857: FAILED: CHECK( wsParkProbe().stopped.load() == 1 ) with expansion: 0 == 1.0 == 1.WsParkServer parkis nowconstin both Qt tests, the other two findings.Cancellation-policy rows to move in
docs/spec/concurrency_and_lifetimes.mdonce both PRs landThat file belongs to the other lane, so this PR does not touch it. The proposed rows are in
backend.mdunder "Cancellation-policy rows for the remote backends":SocketBackend::cancelPending: Work becomes "a Task handler on a server that advertisedcancelis asked to stop; otherwise the server keeps executing". The work column is measured (net suite, via~Bridge). The G-level is unchanged and still read.QtWebSocketBackend::cancelPending: the same, measured in the Qt suite via~Bridge. The G-level is still read.Bridge::setExecuteDeadline: add that over either remote backend, the deadline's stop reaches the server's Task handler through a cancel (measured, net and Qt).Verification
morph_tests: 1752 test cases. 1751 passed and 1 failed as expected; that is the suite's pre-existing[!shouldfail]case.morph_net_tests: all passed, 214 cases.morph_qt_tests: all passed, 85 cases.morph_net_qt_interop_tests: all passed.[cancel]: 475 runs with one failure, which happened in the first batch of 25. That failure's output was not captured: it ran while another lane was compiling. The next 450 runs had none. Weak evidence of a timing flake (2 swaitUntilbudgets under load), not reproduced.[cancel]: 50 runs, 0 failures.[cancel]: 50 runs, 0 failures.readability-trailing-commaandlifetime-safety-*. Existing code on master has the same patterns, and CI 22 passes them.doctarget builds with no warnings.Review notes (done inline)
executeTimeoutis set). An ordinary handler's execute is unchanged.callFinished, which is posted with the reply, or by the cancel that stops it. A reused callId on one connection replaces the entry, andcallFinishederases only its own entry, identified by its stop source._shuttingDownduring the destructor'sprocessEvents.Not verified
branch_partial_allowlisthint was updated by hand, andcheck_branch_coverage.pywas not run.RemoteServer::closeConnectionstill does not stop a dropped connection's running Task handlers. That is a possible follow-up and is not filed here.🤖 Generated with Claude Code