ARTEMIS-6179 Fix deleteReference/depage deadlock - #6606
Closed
shivantaher-oviva wants to merge 3 commits into
Closed
ARTEMIS-6179 Fix deleteReference/depage deadlock#6606shivantaher-oviva wants to merge 3 commits into
shivantaher-oviva wants to merge 3 commits into
Conversation
Adds testRemoveMessageWhilstPagingAndConsuming, mirroring the existing testMoveMessageWhilstPagingAndConsuming/ManagementCopyThread pattern that caught the equivalent copyReference() deadlock (ARTEMIS-5376). QueueImpl#deleteReference() is still synchronized and calls iterQueue(), which locks depageLock, while QueueImpl#depage() locks depageLock first and then enters a synchronized(this) block. Racing QueueControl#removeMessage() against depaging while consuming can deadlock the two threads against each other. Detection uses the JVM's own ThreadMXBean deadlock detector instead of a fixed timeout, since a hang here is the failure itself.
Removes synchronized from QueueImpl#deleteReference(). It calls iterQueue(), which acquires depageLock internally, while depage() acquires depageLock first and then enters a synchronized(this) block. The reversed lock order deadlocks QueueControl#removeMessage() against the paging executor whenever they race on an actively paging queue. This mirrors the fix already applied to copyReference() in ARTEMIS-5376: iterQueue() already provides its own synchronization via depageLock, so the outer synchronized on deleteReference() is redundant and unsafe. Verified with testRemoveMessageWhilstPagingAndConsuming, which reliably deadlocks without this change and passes cleanly with it.
Removes testRemoveMessageWhilstPagingAndConsuming and ManagementRemoveThread, added earlier on this branch to demonstrate the deadlock before the fix. The reproduction is a probabilistic race rather than a deterministic test: it reliably caught the deadlock without the fix and passed cleanly with it, but relies on winning a narrow timing window rather than a guaranteed interleaving. Kept out of the permanent suite for that reason; it remains in branch history as evidence for the fix in the preceding two commits.
shivantaher-oviva
marked this pull request as ready for review
August 11, 2026 08:52
faizankazi-oviva
approved these changes
Aug 11, 2026
Contributor
|
I have brought your changes into #6607 I also did a little refactoring on the test.. and I'm keeping the test (I don't want to remove the good test you wrote.. although I made it run faster, while still reproduced the issue). (I want to keep it faster to we can run it on the CI) |
Contributor
|
I'm closing this as I have created a separated PR for this.. will merge it shortly.. thanks a lot for this. |
Author
|
Thank you for taking this up quickly! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
https://issues.apache.org/jira/browse/ARTEMIS-6179
QueueImpl#deleteReference()issynchronizedand callsiterQueue(), which blocks ondepageLock.lock().depage()does the reverse: it takesdepageLockfirst (viatryLock()), then needs to entersynchronized(this). Two threads acquiring the same two locks in opposite order deadlock as soon as they interleave: aremoveMessage()call holds the object monitor and blocks ondepageLock, whiledepage()holdsdepageLockand blocks on the monitor. Neither ever releases.This is the same bug class already fixed for
copyReference()in ARTEMIS-5376 (30c8fc7).deleteReference()wasn't touched by that fix because it didn't calliterQueue()yet at the time — it was moved ontoiterQueue()two weeks later and keptsynchronized, reintroducing the same bug.The fix
Drop
synchronizedfromdeleteReference().iterQueue()already synchronizes internally viadepageLock, so the outer lock is redundant and is exactly what causes the reversed order above.Checked this doesn't introduce a new race:
deleteReference()'s callback (incDelivering()+acknowledge()) now runs in the exact synchronization context thatmoveReference()andsendMessageToDeadLetterAddress()already use unsynchronized since ARTEMIS-5376, calling the same underlyingacknowledge()/move()machinery.Testing
Verified with a dedicated reproduction test (testRemoveMessageWhilstPagingAndConsuming, racing
removeMessage()againstdepage()while consuming) that reliably deadlocks without this change and passes cleanly with it, plus the existingQueueControlTest,QueueControlUsingCoreTest,ManagementWithPagingServerTest, andLVQTestsuites, all green. The reproduction test is in the first commit on this branch and removed again in the third, since it's a probabilistic race rather than a deterministic test and isn't suited to the permanent suite.