Skip to content

A JVM consumer's parked page allocation fails when another consumer of the task empties its balance #6304

Description

@dwsmith1983

Describe the bug

Follow-up to #6224, which covers native acquires through CometTaskMemoryManager. JVM consumers of the same task have the same exposure and are not covered there.

Spark's ExecutionMemoryPool.acquireMemory registers the task's memoryForTask entry once, before its wait loop, and reads it with memoryForTask(taskAttemptId) on every pass. releaseMemory removes the entry when the task's balance reaches zero. A caller parked in the loop that wakes after the removal gets NoSuchElementException: key not found: <taskAttemptId>. TaskMemoryManager.releaseExecutionMemory takes no monitor, so a release can run while an acquire of the same task is parked.

CometUnifiedShuffleMemoryAllocator.allocate calls allocatePage, and Spark's own operators in the task (sorters, aggregates) do the same. When one of them is parked and a native CometTaskMemoryManager.releaseMemory empties the task's balance, the parked allocatePage fails with that exception instead of completing or returning a page of zero size. Any off-heap consumer's freeMemory can do the same to another parked consumer of the task.

Steps to reproduce

Component-level, in the style of CometTaskMemoryManagerSuite: a 100-byte off-heap UnifiedMemoryManager, one TaskMemoryManager. A CometTaskMemoryManager acquires 10 bytes for the task and another task takes 90. On a second thread, CometUnifiedShuffleMemoryAllocator.allocate(20) parks below the task's minimum share. The main thread calls CometTaskMemoryManager.releaseMemory(10). The parked allocate throws NoSuchElementException: key not found: <taskAttemptId>.

Expected behavior

A parked page allocation completes once memory is free, or fails with Spark's usual out-of-memory path. It does not fail because another consumer of the same task released memory.

Additional context

The root cause is in Spark: reading the entry with getOrElseUpdate(taskAttemptId, 0L) inside the wait loop would remove the window for every consumer. Until that lands, Comet can only guard its own callers. CometTaskMemoryManager retries the acquire (#6224); the allocator could do the same in allocate, and Spark's own operators cannot be guarded from Comet.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:shuffleShuffle (JVM and native)bugSomething isn't workingpriority:mediumFunctional bugs, performance regressions, broken features

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions