[JSC] Less memory per function that never runs, and code that can be dropped and decoded again - #628
[JSC] Less memory per function that never runs, and code that can be dropped and decoded again#628Jarred-Sumner wants to merge 18 commits into
Conversation
Preview Builds
|
209144a to
affb84c
Compare
Pins the preview release autobuild-preview-pr-628-affb84c2 (affb84c2ee68fe4d838fc819aef2830de30f4deb); to be replaced by the merged sha before this lands. What the new JavaScriptCore does differently: a module's function declarations are instantiated when their binding is first read and stay in the embedded bytecode until then; interpreter and Baseline call sites get their link record on their second execution, in either tier (a tail call on its first), and the collector is told about those records; value profile predictions and an unlinked code block's value and array profiles are only allocated for code that warms up; module and program code is released once it has run; code decoded from an embedded bytecode cache can be dropped and decoded again (JSC::VM::shrinkFootprintNow with flags); Heap::totalBytesAllocated() and sizeAfterLastCollection() for embedders; the thunks of the call slow paths clear the stack their C++ function's frame is going to occupy, so that what the last callee left there is not kept alive by a later conservative scan. The bytecode cache format revision goes from 5 to 6.
| #if ASSERT_ENABLED | ||
| #define ASSERT_CALL_SLOW_PATH_RUNS_IN_CLEARED_STACK(calleeFrame) \ | ||
| ASSERT(std::bit_cast<uintptr_t>(currentStackPointer()) + maxFrameExtentForSlowPathCall + stackBytesClearedForCallSlowPath >= std::bit_cast<uintptr_t>(calleeFrame)) |
There was a problem hiding this comment.
🔴 On C_LOOP debug builds this assertion compares the native C++ stack pointer against calleeFrame, which lives in the heap-allocated CLoopStack, so the inequality is between unrelated address spaces and can fire spuriously on the very first LLInt call slow path. Fix: guard the assertion (or make it a no-op) when ENABLE(C_LOOP), matching the if not C_LOOP guard around the stack-clearing loop in linkFor (LowLevelInterpreter.asm) that the assertion is meant to verify. Same pattern at 4 sites (llint/LLIntSlowPaths.cpp: llint_default_call, llint_unlinked_call, llint_virtual_call, llint_polymorphic_call).
Extended reasoning...
LowLevelInterpreter.asm's linkFor macro wraps the new stack-clearing loop in if not C_LOOP … end, so on the C loop the trampolines call llint_default_call/llint_unlinked_call/llint_virtual_call/llint_polymorphic_call without clearing anything. Each of those functions now does ASSERT_CALL_SLOW_PATH_RUNS_IN_CLEARED_STACK(calleeFrame), which expands to ASSERT(bit_cast<uintptr_t>(currentStackPointer()) + maxFrameExtentForSlowPathCall + stackBytesClearedForCallSlowPath >= bit_cast<uintptr_t>(calleeFrame)). With ENABLE(C_LOOP) the interpreter stack is a heap block (CLoopStack) and calleeFrame points into it, while currentStackPointer() (wtf/StackPointer.h) returns the native thread stack pointer; depending on where the heap block lands relative to the native stack, the comparison can be false and the ASSERT fires on the first JS call in a debug C_LOOP build (which the base handled fine because none of these asserts existed). The JIT thunks (ThunkGenerators.cpp emitCallSlowPath) are not affected because they are compiled only with ENABLE(JIT).
Verification: normal — On C_LOOP debug builds the new assertion compares unrelated address spaces and can fire on any call. The macro added by this diff (Source/JavaScriptCore/assembler/MaxFrameExtentForSlowPathCall.h:79-84) is: ``` #if ASSERT_ENABLED #define ASSERT_CALL_SLOW_PATH_RUNS_IN_CLEARED_STACK(calleeFrame) \ ASSERT(std::bit_cast<uintptr_t>(currentStackPointer()) + maxFrameExtentForSlowPathCall +…
| noInline(callExported); | ||
| noInline(callPrivate); | ||
|
|
||
| for (let i = 0; i < 100000; ++i) { |
There was a problem hiding this comment.
🟡 nit (optional): new JSTests use hardcoded iteration counts (100000, 200000) instead of testLoopCount, violating JSTests/README.md rule 2 (referenced by JSTests/CLAUDE.md), so they cannot exit early in configs where tier-up doesn't matter and risk breaking the 200 ms rule. sweep:for \(let i = 0; i < [12]00000; Fix: replace hardcoded warm-up bounds with testLoopCount (e.g. for (let i = 0; i < testLoopCount; ++i)), keeping small fixed inner counts like 200 as-is.
Extended reasoning...
JSTests/README.md (imported by JSTests/CLAUDE.md) requires new tests to use testLoopCount/wasmTestLoopCount so the harness can scale iterations per configuration, and to run in under 200 ms in all configurations. At least a dozen new tests in this PR (JSTests/modules/lazy-function-declarations-dfg.js:77,97,100,111,114; JSTests/modules/lazy-function-declarations-dfg-slow-path.js; JSTests/stress/baseline-lazy-call-link-info.js; JSTests/stress/llint-lazy-call-link-info*.js; JSTests/stress/get-from-scope-global-variable-watchpoint-lookup.js; and others) hard-code 100000 or 200000 loop bounds. In slow configurations (no-cjit, eager, cloop, collect-continuously) these fixed counts can push individual tests well past 200 ms and cannot be scaled down by the runner, and in configs where the JIT is disabled the extra iterations serve no purpose. Base branch has no such tests; this is introduced by the diff. Structural fix: use testLoopCount for the tier-up warm-up loops across all new test files.
Verification: nit — JSTests/README.md:20 (imported by JSTests/CLAUDE.md via @ README.md) states rule 2: "Use testLoopCount or wasmTestLoopCount to control how many iterations a test runs. The jsc CLI sets these based on the configuration of the test, so tests iterate enough to tier up where that matters and exit early where it doesn't." At the candidate's location,… | nit — JSTests/README.md:20…
…able for its whole life Records of several module loaders share a ModuleProgramExecutable, the FunctionExecutables of its function declarations and with them their optimized code, while each record has its own environment. That is sound because every such environment is made from the executable's one symbol table: the optimizing tiers treat the scope of a table that has only seen one environment as a constant (SymbolTable::singleton(), for closure variables through the table in resolve_scope's metadata, for imports through topLevelExecutable->moduleEnvironmentSymbolTable()), and the second environment made from the table invalidates that. ScriptExecutable::clearCode() cleared the table, and getUnlinkedCodeBlock() made a new one (and a new vector of declaration executables) whenever it generated the code again. A record that had adopted the executable and made its environment but not run yet (waiting for a dependency suspended at a top-level await, or suspended itself) when all code is deleted generates the code again to run; the executable then has unlinked code and is adopted by the next loader, whose environment comes from the new table. Functions of the module compiled after that fold the newest loader's exporting environment into the code that the earlier loaders' functions run as well: an imported binding read in loader a returns loader d's value. The table and the vector are now made from the executable's first unlinked code and stay (an environment keeps its table alive anyway). Code fetched again is asked for in the code generation mode of the first, since the layout of the module environment depends on it; JSModuleRecord::getOrMakeExecutable therefore only lets a record adopt an executable whose mode is the one the record would ask for. An executable whose code was deleted is still not adopted; it no longer sits in the clearable-code set for the sake of its table. * JSTests/stress/module-loaders-share-one-environment-symbol-table.js: Added.
…t first reaches the Baseline JIT Every UnlinkedCodeBlock allocated one UnlinkedValueProfile per value profile and one UnlinkedArrayProfile per array profile up front, so that its CodeBlocks can fold their predictions into them at each collection and the next CodeBlock linked from it starts from those. Only the DFG reads what ends up in a linked profile that way, and nearly all functions never leave the LLInt. Keep the number of array profiles on the UnlinkedCodeBlock, derive the number of value profiles from the parameter count and the metadata table, and allocate the two arrays as one ButterflyArray when a BaselineJITPlan is first created for one of its CodeBlocks (on the main thread; the pointer is published with a release store and the collector and compiler threads that fold profiles load it once, with acquire). Until then folding finds nothing to fold into. Builtin functions, which never fold, no longer get the arrays at all. The bytecode cache no longer stores the number of value profiles (cachedTypesFormatRevision 6, which also covers the two format changes that follow). Options::useLazyUnlinkedValueAndArrayProfiles() (default true) restores the allocation at generation and decode when false.
…Data Few code blocks have a jump whose target does not fit its operand, but each one carried an empty hash map for them. With the previous change this brings sizeof(UnlinkedCodeBlock) from 208 to 192, one size class down for every kind of unlinked code block. The bytecode cache writes them with the rest of the rare data (still cachedTypesFormatRevision 6). A static_assert keeps the size where it is in release builds.
…er character class The table that resolves two-character inline strings had an entry for every pair of Latin-1 characters: 512 KB per VM, all of it touched when it is zeroed. Minified names only use $, _, digits and ASCII letters, which is 64 classes, so a 64 x 64 table (32 KB) covers them. Any other pair shares the direct-mapped, verified cache the three-character strings use.
InitializeEnvironment creates a JSFunction and a FunctionExecutable for every heap allocated function declaration of a module before any of its code runs. A large bundled CLI application declares ~30,000 top-level functions in the chunks it evaluates before it is ready for input and has read ~17% of them by then; each unread one costs a 128-byte FunctionExecutable, a 32-byte function object and, when the code was parsed rather than decoded from the bytecode cache, an UnlinkedFunctionExecutable. With Options::useLazyModuleFunctionDeclarations() (default true) InitializeEnvironment leaves the module environment slot of such a declaration empty and hands the module record the list of them. The function object is created when the binding is first read: - get_from_scope with the new LazyClosureVar resolve type. BytecodeGenerator emits ResolvedLazyClosureVar for a module's own declarations; the linker turns ModuleVar and closure variables that are function declaration slots of a module environment into LazyClosureVar. LLInt, Baseline and LOL test the loaded value for empty and take the slow path; the DFG has GetLazyClosureVar (clobberizes as a read of the slot that can fire watchpoints and allocate), lowered by the FTL with a slow path call. A slot that holds a value is read exactly like ClosureVar. - JSModuleEnvironment::readVariable() for everything else: by-name lookup on the environment, module namespace object property access, WebAssembly imports of the export, the debugger. An empty slot that is not a function declaration is still a TDZ binding. Type and control flow profilers keep the eager path (they want every function's range up front). Records of several module loaders that share a ModuleProgramExecutable share the declarations' FunctionExecutables (ModuleProgramExecutable::functionDeclaration): the record that reads a declaration first links it, every record fills its own environment's slot with its own function object. Optimized code of one loader therefore meets empty slots when the next loader runs it, which JSTests/modules/module-loaders-lazy-function-declarations.js exercises in the optimizing tiers without the testing option. On that application, 12 s after the prompt: FunctionExecutable 55.1k -> 30.0k, function objects 156k -> 131k, GC heap -6.3 MB, RssAnon -7.2 MB. CPU: LLInt call-heavy loop -0.8%, Baseline-only +1.6% (test and branch per lazy read), DFG/FTL unchanged to the instruction. The bytecode cache stores the resolve type and the slot table: still cachedTypesFormatRevision 6.
…s from the closure ones Of the get_from_scope instructions linked by the time a large bundled ESM application is ready for input, 89% are ClosureVar or LazyClosureVar and 11% are global accesses (ClosureVar 54.8%, LazyClosureVar 34.4%, GlobalProperty 9.8%, GlobalVar 1.0%); in a script it is the other way round. The chain of compares tested GlobalProperty, GlobalVar and GlobalLexicalVar before ClosureVar, so a closure variable read waited for three compares that a module never needs. The resolve type is now read as the low byte of the GetPutInfo (one load instead of a load and a mask; GetPutInfo.h asserts that it fits), and one compare sends everything above GlobalLexicalVar past the three global types, the last of which needs no compare of its own any more. Instructions up to the access itself: GlobalProperty 4 -> 5, GlobalVar 6 -> 7, GlobalLexicalVar 8 -> 7, ClosureVar 10 -> 5, LazyClosureVar (new, with its empty check) 7. Interpreter-only instruction counts: closure-scope loops -4..-5% (infer-one-time-closure-ten-vars, infer-closure-const-then-mov), a loop that calls a global function and accumulates into a global: 45.11 G -> 45.01 G, global-heavy microbenchmarks within +-1% (function-call +0.35%, array-prototype-indexOf-empty +0.46%, array-of-contiguous-large -1.1%).
…tecode cache payload With Options::useLazyModuleFunctionDeclarations(), a module decoded from a bytecode cache payload that stays around (Decoder::canDeferIntoPayload()) does not create the UnlinkedFunctionExecutables of the function declarations InitializeEnvironment would have instantiated: their slots in UnlinkedCodeBlock::m_functionDecls stay null and the module's ModuleFunctionDeclarationSlots remembers the Decoder and the record array. UnlinkedCodeBlock::functionDecl() decodes one when it is first asked for it (mutator only; collector and compiler threads use functionDeclIfDecoded()), which is when its binding is first read. The module record refers to the unlinked code block instead of a vector of executables, and an empty slot is all that says a declaration is uninstantiated. A never-read module function goes from 256 bytes of cells (UnlinkedFunctionExecutable, FunctionExecutable, JSFunction) to an 8-byte empty slot and a 4-byte table entry. On a large bundled CLI application: UnlinkedFunctionExecutable 56.4k -> 31.3k at the prompt; `--help` runs 505 M -> 470 M user instructions (-7%) since most of the declarations of the chunks it loads are never touched.
detect_stack_use_after_return (on by default) moves address-taken locals to a heap-backed fake stack that the conservative root scan does not visit, so a cell only such a local refers to is collected at the next GC: with --slowPathAllocsBetweenGCs or --collectContinuously every test that imports more than one module died in JSPromise::pipeFrom on the promise JSModuleLoader::hostLoadImportedModule had just got back from fetch(). Same default as the embedder's binary.
…and ArrayProfile on their second execution
Most call sites of most functions run at most once, and most functions never
leave the LLInt, but every op_call carried an 80 byte DataOnlyCallLinkInfo and a
16 byte ArrayProfile in its metadata (96 bytes per call site).
The metadata of the call opcodes without varargs now holds a LazyCallLinkInfo,
which is a pointer to a CallSiteData { DataOnlyCallLinkInfo, ArrayProfile }.
Until a site has run twice it points at one of two CallSiteDatas that the VM
owns and all such sites share ("never executed" and "executed once"). Their
CallLinkInfo looks like a polymorphic call whose destination is a new thunk,
llint_unlinked_call, so the LLInt call fast path is what it was plus one load
of the pointer, and it needs no branch for the unallocated case. The thunk finds
the site through the caller's frame. On the first execution it calls the callee
without a CallLinkInfo (a temporary one on the stack for the error paths and for
callees that are not JS functions) and moves the site to "executed once". On the
second it allocates the site's own CallSiteData and hands it to linkFor(), which
links it, exactly when a CallLinkInfo used to be linked (the first slow path
trip only set the seen bit). op_tail_call has given up the caller's frame by
the time the thunk would run, so it checks for a CallLinkInfo without owner
and takes a slow path that allocates. The varargs opcodes keep their inline
DataOnlyCallLinkInfo; their slow path needed it on the first execution anyway.
The ArrayProfile of op_call / op_call_ignore_result / op_tail_call moves into the
CallSiteData. The LLInt stores |this|'s structure id through the pointer it has
already loaded, into the shared CallSiteData when the site has none of its own;
nobody reads the ArrayProfile of a shared one.
Baseline code loads the pointer as well and expects a CallSiteData that belongs
to the site: CodeBlock::setupWithUnlinkedBaselineCode() gives every site one
before it installs the code, so neither the Baseline nor the optimizing JITs
ever meet a shared CallSiteData. Compiler threads and the collector get at the
CallLinkInfos through CodeBlock::forEachLLIntOrBaselineCallLinkInfo() and at the
ArrayProfiles through CodeBlock::getArrayProfile(); both skip sites without a
CallSiteData of their own, which reads as "no information". A new CallSiteData
is initialized before a store-store fence publishes it and it lives as long as
the MetadataTable. Nothing may write to the CallLinkInfo of a shared
CallSiteData: it is flagged, and every mutator of CallLinkInfo asserts.
sizeof(Metadata): op_call, op_call_ignore_result, op_tail_call 96 -> 8, op_construct
80 -> 8, op_super_construct 88 -> 16, op_iterator_open 120 -> 48, op_iterator_next
136 -> 64. Options::useLazyLLIntCallLinkInfos() = false allocates every
CallSiteData when the CodeBlock is linked.
… line, and only for code that has warmed up A ValueProfile in front of a MetadataTable was a bucket and a prediction, 16 bytes per profiled instruction. The LLInt and the Baseline JIT only ever store to the bucket. The prediction is written when a collection or a compiler folds the bucket into it, which for code that never tiers up buys nothing. The table is now preceded by the buckets alone (8 bytes each; the LLInt's store is indexed by the negated profile offset instead of multiplying it, the Baseline JIT's is the same instruction with another displacement). The predictions are an array that hangs off the table's LinkingData and is allocated by the first ValueProfileRef::mergePrediction() that has something to record; it is published with a compare-and-swap since the mutator, a marker thread and a compiler thread can all get there first. ValueProfileRef is what CodeBlock::valueProfileForOffset() and friends now return; without a predictions array it predicts SpecNone. A collection used to fold every bucket of every live LLInt / Baseline CodeBlock into its prediction, because nothing marks the cell in a bucket. So that this does not allocate the predictions of everything that ever ran, a CodeBlock that is still interpreted, has no predictions yet and whose LLInt execution counter is below Options::thresholdForValueProfilePredictions() (40: it has been called twice at most and has not looped much) keeps its samples in the buckets: CodeBlock::reconcileWeakReferencesAtGCEnd() only clears the buckets whose cell died, and visitChildren() leaves them alone. Once the code is warmer than that, or has Baseline code, or a compiler looks at it as an inlining candidate, predictions are computed as before, starting from whatever the buckets hold. What is lost is the type of an object that died before its CodeBlock's third call. Options::useLazyValueProfilePredictions() = false allocates the predictions with the table and always folds.
…t set in its metadata The LLInt and the Baseline JIT never look at it. The DFG bytecode parser uses it to constant-fold reads of global variables that have been written once; it can find the set in the SymbolTableEntry of the global object's / global lexical environment's symbol table, under the table's lock, which is where linking got it from. With the StructureID moved out of the union into what was padding, sizeof(OpGetFromScope::Metadata) goes from 24 to 16. m_getPutInfo, m_structureID and m_operand keep their names, types and meaning.
Counts the live linked CodeBlocks by tier and kind, the unlinked code blocks by kind, the bytes of their metadata tables, JIT code and instruction streams, the UnlinkedFunctionExecutables whose code is still in the bytecode cache they were decoded from, the RegExps holding compiled code and the module executables that still hold code. For embedders that want to see where code memory goes, and for tests about code lifetime.
…d decoded again
An UnlinkedFunctionExecutable decoded from a bytecode cache lets go of its Decoder and record offsets when its code is
first decoded, so the only way to get its unlinked code back after dropping it was to parse the function again, and
such code was therefore never dropped (Heap::deleteAllUnlinkedCodeBlocks does not even see it). In a process whose cache
payload stays mapped for its whole life a dropped block can instead be decoded again for a few thousand instructions.
- A code block decoded from a persistent payload remembers its payload (an index into the new per-VM
PersistentBytecodePayloads, two bytes in existing padding) and its record's offset (four bytes in existing padding).
sizeof(UnlinkedCodeBlock) and sizeof(UnlinkedFunctionExecutable) are unchanged.
- UnlinkedFunctionExecutable::returnCodeToCache() puts the executable back into its m_isCached state, naming the code
blocks' records (negative offsets) and a Decoder for the payload: the payload's live one if there still is one, so
that what it decoded stays shared. The live child executables of a dropped block are remembered weakly, in position
order under the parent's record offset (which the block carries, so dropping never reads the payload: at deep idle
the embedder has paged it out, and faulting it back cost ~28 MB of file pages for ~20 MB of heap released), and
adopted by the block decoded from the same record later, so closures made before and after share their unlinked
code. The registry prunes itself after each full collection like a WeakGCMap.
Only with the lean decoder: the other one remembers cells by record offset and would hand out a dead one.
- Heap::deleteAllUnlinkedCodeBlocks takes which kinds of unlinked code to drop (generated, recoverable from a cache,
only what nothing links against); Heap::deleteAllCodeBlocks and ScriptExecutable::clearCode can keep what would have
to be parsed again.
Compilations that are ready are finished (Heap::completeAllJITPlans allocates: DFG::LazyJSValue) and the set of unlinked
code that linked code refers to is collected before the heap is prepared for iteration, as in deleteAllCodeBlocks().
- VM::shrinkFootprintNow / shrinkFootprintWhenIdle take flags: KeepCodeThatNeedsParsing (linked code, recoverable
unlinked code, RegExp code and parser caches go; nothing that needs a re-parse, including the builtins),
KeepCodeInUse (additionally keeps all linked and RegExp code and the unlinked code of every function that still has
linked code) and LeaveCollectionToCaller. Without flags they behave as before.
- VM::entryCountFromOutside() and a public Heap::lastActiveCollectionTime() let an embedder tell whether the program
is really at rest before it asks for any of this.
- A slot of PersistentBytecodePayloads stands for one tree of unlinked code: the payload's bytes as decoded for one
SourceProvider (an embedder may wrap the same bytes in a new CachedBytecode and SourceProvider per load; what is decoded
for a provider names it, e.g. as the provider of class sources). Children remembered under one slot are never handed to
code decoded under another, and code that is private to one executable (a module whose loader has a module scope of its
own) is not registered at all: its Baseline code assumes one resolution of that scope. Every Decoder of a slot and every
code block that carries its index hold the slot; with the last of them the payload and provider references go and the
index is used again, so a process that loads the same file over and over does not accumulate them (was: one entry per
load until the VM died).
- shrinkFootprintNow also refuses when called from inside the collector, documents that the Keep modes wait for a running
collection while the flagless mode returns false instead, clears lastException (its stack kept dropped code alive in a
second generation), and clears the SourceProvider caches in every mode. VM::deleteAllCode (the flagless mode) now deletes
the code a cache can hand back as well: on a synthetic application the heap after shrink + GC is 5.4 MB instead of 31.7.
- Heap::deleteAllUnlinkedCodeBlocks defers collection from before it finishes the compiler threads' plans until it is done:
nothing marks while an executable's code block slots turn into {Decoder, offsets}. The two sides publish and read those
slots with m_isCached in the middle and fences on both sides. Heap::lastActiveCollectionTime() is atomic.
Options::useCodeRecoveryFromBytecodeCache (default on) gates the bookkeeping and with it the recovery.
On a large bundled CLI application (idle 250 s after a 20-turn session): RssAnon 152-154 MB -> 135-138 MB with
KeepCodeInUse, 131 MB when everything is dropped, RssFile unchanged; decoding a dropped block again costs ~2.5k
instructions against ~111k for a re-parse. Synthetic (2000 modules): unlinked code blocks 19,252 -> 2,383, GC heap
33 -> 23 MB.
A module body runs once per module record, but its ModuleProgramCodeBlock, UnlinkedModuleProgramCodeBlock and CodeCache entry stayed alive for as long as any function the module created (through FunctionExecutable::topLevelExecutable and the CodeCache's Strong), the linked code until a full collection found it past its TTL. Once JSModuleRecord::evaluate sees the body finished, ModuleProgramExecutable::didFinishEvaluation() drops the linked code, and the unlinked code together with its CodeCache entry if it was decoded from a persistent bytecode cache and can be decoded again; the environment's symbol table and the function declarations' executables stay, and a later getUnlinkedCodeBlock() keeps them, asks for the code generation mode of the code it dropped and refuses code that did not come from the payload. Records of several module loaders share one ModuleProgramExecutable, so the executable counts the records that have adopted it and not finished yet (JSModuleRecord::getOrMakeExecutable, didFinishEvaluation): the code goes when the last one is done, which covers a loader suspended at a top-level await while another runs the same module to its end. A loader that comes later adopts an executable that has only let go of recoverable unlinked code and has it decoded again. A body suspended at a top-level await is not finished and keeps its unlinked code, also in VM::shrinkFootprintNow. Interpreter::executeProgram, which creates its ProgramExecutable itself, clears it after the run. Code that left the interpreter, or has a baseline compile queued, is left to age out: a compiler thread may be looking at it (GlobalExecutable::canReleaseLinkedCodeNow). With useLazyModuleFunctionDeclarations a module record referred to the module's unlinked code so that a declaration nobody has read yet can be instantiated later; since most never are, that alone kept nearly every module's unlinked code alive. When the declarations were left in a persistent payload the record's ModuleFunctionDeclarationSlots can decode any of them on its own, so in that case the record does not reference the code block: a read links from the executable's code block while it has one and from the payload otherwise. Options::useRunOnceCodeRelease (default on). Options::useSharedModuleFunctionExpressionExecutables (default off): the FunctionExecutables of the function expressions and classes in a module's top-level code belong to the ModuleProgramExecutable instead of to each linked ModuleProgramCodeBlock, so their CodeBlocks and JIT code survive the module's own linked code and are shared by every evaluation of the module. Synthetic application of 2,000 modules: ModuleProgramCodeBlock 2,001 -> 1, UnlinkedModuleProgramCodeBlock 2,564 -> 2, GC heap 39.4 -> 24.3 MB, extra memory 11.8 -> 1.5 MB, startup instructions -4.4%. A large bundled CLI application: RssAnon 10 s after launch -3 MB.
…ated, for embedders An embedder that wants to tell an idle program from a busy one that happens not to grow its heap needs two numbers Heap keeps but does not hand out: - sizeAfterLastCollection(): m_sizeAfterLastCollect, the live size (cells and extra memory) as of the last finished collection, eden or full. - totalBytesAllocated(): everything the mutator has allocated, cells and reported extra memory, the current cycle included. The per-cycle counters are reset in updateAllocationLimits(); they are added to a running total right before that. Mutator thread only. Both are read once per embedder GC timer tick.
…to the collector Since call sites in LLInt / Baseline metadata keep their CallLinkInfo and ArrayProfile out of line (LazyCallLinkInfo), the 100-odd bytes per call site that the table used to hold inline, and report with its size, are separate allocations the collector does not know about: for a function with 20,000 call sites that is 2.2 MB per CodeBlock that neither paces eden collections nor counts towards the next full one. MetadataTable::LinkingData counts the CallSiteDatas its call sites own (in what is padding on 64-bit targets; only the mutator writes it, so a relaxed load and store). On the allocation side one record is below what Heap::reportExtraMemoryAllocated() takes note of, so LazyCallLinkInfo::ensureSlow() reports them 32 at a time, against the owning CodeBlock, which also covers a CodeBlock that is already marked. On the visiting side CodeBlock::visitChildren() and estimatedSize() add them for the CodeBlock the table was linked for; the optimized CodeBlocks that share the table of the one they replace do not, or a hot function's records would be counted three or four times in every collection. While here: MetadataTable::sizeInBytesForGC() took a reference on the UnlinkedMetadataTable (an atomic increment and decrement) from every marker thread for every CodeBlock; it reads it through the LinkingData now.
…+ function's frame is going to occupy The slow paths for calls (llint_default_call, llint_virtual_call, llint_polymorphic_call, llint_unlinked_call; operationDefaultCall, operationVirtualCall, operationPolymorphicCall) run with the stack pointer at the callee's frame: their own frame lies over whatever the last callee at that depth left there, and sanitizeStackForVM(), which they call first thing, can only clear what is below the function that calls it. What such a frame does not write stays in reach of the conservative scan of a later, shallower callee's native frames. How much stays depends on the size of that frame, which nothing controls: llint_unlinked_call()'s is 216 bytes against llint_default_call()'s 120, and that was enough to show. Seen as: `rewrite(); rewrite(); Bun.gc(true)` at the top level of a script keeps the HTMLRewriter that the last rewrite() made alive through that collection (Bun's html-rewriter-leak test). The wrapper's address sits in an unwritten slot of the host function's frame (bindgen_BunObject_jsGc), 176 bytes below the frame pointer of rewrite()'s and Bun.gc's common caller; a heap snapshot taken at that point has the wrapper with no incoming edge and no root; SlotVisitor::append(ConservativeRoots) is what marks it; clearing that one word in a debugger lets it die. The first call of the `pf()` site went through llint_unlinked_call(), whose frame covers that address; through llint_default_call() the address is below the frame and cleared. Rather than keep one of eight frames small, which only the optimizer decides (and which a different ABI, or no optimization, decides differently), the thunks that call these functions now clear the stackBytesClearedForCallSlowPath bytes that the C++ frame is going to occupy before it exists: the linkFor macro in the LLInt, and emitCallSlowPath() for the JIT's thunks, which is what the default, virtual and polymorphic call thunks had four copies of. Nothing is written below the stack pointer (no ABI we target promises that memory, and a signal handler's frame goes there): the stack pointer moves down over the window, a multiple of 16 bytes, the window is cleared above it, and it moves back. 256 bytes in a release build (the largest of these frames is 152 bytes on x86_64; everything the function calls is below what it clears itself), 32 straight-line stores, 16 stp on ARM64; 2 KB, in a loop, with assertions or ASan, where ASSERT_CALL_SLOW_PATH_RUNS_IN_CLEARED_STACK() checks in each slow path that its frame does not reach below the window. The LLInt's linkFor clears 64 bytes per iteration of its loop. That is 36 instructions per call that takes one of these slow paths with the JIT, 49 without. sanitizeStackForVM(), which each of them calls, spends about 55 on two thread-local lookups to find the bounds of the current thread's stack and to check that the thread holds the API lock; these slow paths run in the middle of JS, on the thread that holds the lock, so sanitizeStackForVMInCallSlowPath() takes the bounds from the lock's owner and makes the same two checks of lastStackTop against them. A call site that takes a slow path on every call (3 M calls of Proxy callables): 3.715 G -> 3.652 G user instructions with the JIT, 4.889 G -> 4.862 G without; a plain call loop does not change.
…ns for the second time, as the LLInt does A CodeBlock that was set up with Baseline code got a CallSiteData for every call site that did not own one (CodeBlock::ensureCallLinkInfos), and that is every new CodeBlock of a function whose Baseline code is shared through its UnlinkedCodeBlock: a module or a CommonJS wrapper that is evaluated again, a function linked in another realm. For exactly the functions the lazy records are for, big ones whose call sites mostly run once per evaluation, that is a pass over the instruction stream, an allocation, an initialization and a free per call site per evaluation. A CommonJS module with 20,000 calls required 500 times: 7.2 G user instructions before call sites kept their records out of line, 10.8 G since; resident memory of the loop swings by 100-150 MB with it (the records are small allocations the allocator purges late, where one metadata allocation of several MB used to be unmapped at once). Nothing in Baseline code needs the record to be the site's own. It loads the pointer from the metadata on every execution, and the CallLinkInfo the unowned sites share looks like a polymorphic call whose destination is the unlinked call thunk. That was the LLInt's trampoline; with the JIT it is now a thunk like the default call thunk (LLInt::unlinkedCall() chooses, next to LLInt::defaultCall()), calling operationUnlinkedCall(), which returns the exception-throwing stub when the call threw. It and llint_unlinked_call() share LLInt::handleUnlinkedCall(). All the slow path gets is the shared CallLinkInfo, so the site is what the caller left in its frame (CallFrame::bytecodeIndex(); the LLInt and Baseline code store it the same way). handleUnlinkedCall() checks what that rests on: the caller's CodeBlock is LLInt / Baseline code, the instruction is one of FOR_EACH_OPCODE_WITH_LAZY_CALL_LINK_INFO (CodeBlock::lazyCallLinkInfoAt() crashes otherwise), it is not a tail call, and the site does point at a shared CallLinkInfo. A tail call has given up the caller's frame by then, so a tail call site gets its own record before it runs for the first time. The LLInt tested for that already (prepareCallSiteForTailCall). Baseline code pays nothing per tail call for it: tail call sites start out with a third shared record, whose CallLinkInfo looks unlinked rather than polymorphic, so their first execution takes the path of the inline cache that an unlinked call takes anyway, before the frame is given up, and that path now tests for a CallLinkInfo without an owner and calls operationEnsureCallLinkInfoForTailCall(). A direct eval whose callee is not eval makes a virtual call with the site's CallLinkInfo in Baseline code's slow case, which does the same test first. Nothing looks at the whole instruction stream any more. What compiler threads see: more CodeBlocks now have sites that get their record while an optimizing compile reads them. LazyCallLinkInfo::ownData() orders what it reads through the pointer after the pointer (the publisher already has a storeStoreFence), and CallLinkStatus::computeFor() looks the site up when the map it was given predates the record, so that what DFG::InliningPlan saw and what the parser sees agree. The ArrayProfile of |this|: both tiers note the structure in whatever CallSiteData the site has, so the first two executions went into the shared one, which nobody reads (arrayProfile() only returns a site's own); the record is now seeded with the structure of |this| when it is made. One CodeBlock::ensureCallLinkInfoAt() replaces four copies of the lookup, and $vm.numberOfOwnCallLinkInfos(function) lets a test see which sites own a record. User instructions, a standalone runtime built on this, 500 requires of a module (before: the current WebKit main; then this series without and with this change): - 20,000 plain calls in the CommonJS wrapper: 7.23 G 10.8 G 6.64 G - 20,000 tail calls in a switch, 4 of them run: 10.7 G 15.0 G 8.37 G - 2,000 strict functions, each with one tail call: 7.87 G - 7.29 G (with a small heap, --smol: 7.18 G, 8.57 G, 7.71 G; the difference to the line above is collections that find more CodeBlocks alive, not these paths) A large bundled CLI application: --help 554.6 M -> 494.4 M, a 20-request session 23.0-23.6 G -> 22.7-23.1 G.
affb84c to
72ea768
Compare
15 commits on
cf1b36ec8703. Each commit is one feature, builds on its own, and carries itstests; they are ordered by dependency and can be reviewed (and reverted) one at a time.
Why
A large bundled CLI application (ESM, ~1,800 modules in ~900 chunks, standalone executable with an embedded bytecode cache)
sits at an idle prompt with ~100 MB of anonymous memory of which JSC code structures are the largest part, and at ~170 MB
after a session. Almost all of it belongs to code that never runs or ran once:
ran at most once, 23% value-profile predictions nobody reads.
against ~111k for a re-parse; module bodies keep their linked and unlinked code although they run once.
Nothing here changes behaviour observable from JavaScript. Every feature has an option (all default on unless stated) so it can
be switched off individually.
Commits
1. A ModuleProgramExecutable keeps one module environment symbol table for its whole life
A fix for the base, found while testing this series against module loaders (#522), and the invariant the rest relies on.
sound only because every environment is made from the executable's ONE symbol table, so that the second environment
invalidates
SymbolTable::singleton()(the DFG folds the scope ofresolve_scopeclosure variables through the table in themetadata, and of imports through
topLevelExecutable->moduleEnvironmentSymbolTable()).ScriptExecutable::clearCode()cleared the table and
getUnlinkedCodeBlock()made a new one (and a new declarations vector) when it generated the codeagain. A record that had made its environment but not run yet when all code was deleted (waiting on a dependency suspended
at a top-level await) regenerates the code; the next loader adopts the executable with the new table; functions compiled
after that bake the newest loader's exporting environment into code the earlier loaders run too: an imported binding read
in loader a returns loader d's value (
JSTests/stress/module-loaders-share-one-environment-symbol-table.jsfails oncf1b36ec8703). With lazy declarations (commit 5) the same split would also hand a record executables specialised onanother record's environment, hence a
RELEASE_ASSERT(environment->symbolTable() == executable->moduleEnvironmentSymbolTable())on the lazy read path there.
its table alive anyway). Code fetched again is asked for in the code generation mode of the first (the environment's layout
depends on it);
getOrMakeExecutableonly adopts an executable whose mode is the one the record would ask for. An executablewhose code was deleted is still not adopted, and no longer sits in the clearable-code set for the sake of its table.
resolve_scopefolding,
tryGetConstantClosureVar,GetLazyClosureVar's fast case,FunctionExecutable::singleton()): all reduce to theone-table rule; details in the commit message.
2. Allocate an UnlinkedCodeBlock's value and array profiles when it first reaches the Baseline JIT
Options::useLazyUnlinkedValueAndArrayProfilesm_valueProfiles/m_arrayProfiles(two FixedVectors) become oneButterflyArray(ValueAndArrayProfiles) allocated when aBaselineJITPlanis first created for a CodeBlock of the unlinked code. The number of value profiles is derived(
numParameters + metadata.numValueProfiles), the number of array profiles is kept as a 32-bit count in existing padding.Builtins never get them. The bytecode cache no longer stores the value profile count (format revision 5 -> 6; this is the only
bump in the series, commits 3 and 5-7 change the format too and stay on 6).
and compiler threads that fold profiles (
CodeBlock::updateAll*Predictions) load it once per fold with acquire and skipfolding when it is null. Freed only in
~UnlinkedCodeBlock.3. Move UnlinkedCodeBlock's out-of-line jump targets into its RareData
An empty hash map per code block for a case few blocks have. With commit 2:
sizeof(UnlinkedCodeBlock)208 -> 192, one sizeclass down for all unlinked code blocks (
static_assertin release, 64-bit, non-Windows builds). No new RareData on thesynthetic app (902/3009 before and after).
4. Index the bytecode cache's two-character atom table by identifier character class
512 KB per VM (zeroed, so all touched) -> 32 KB: 64 x 64 classes (
$ 0-9 A-Z _ a-z); other pairs use the verifieddirect-mapped cache the three-character names use.
Commits 2-4 together, synthetic: RssAnon 184.4 -> 178.4 MB; startup instructions -0.5..-0.8%, exercise -1.0..-1.6%.
5. Instantiate module function declarations on first read
Options::useLazyModuleFunctionDeclarations(predictFunctionForUnprofiledLazyClosureVarForTestingfor tests)object and its FunctionExecutable are created when the binding is first read:
get_from_scopewith the newLazyClosureVarresolve type (LLInt, Baseline, LOL: load, test for empty, slow path; DFG
GetLazyClosureVar, FTL lowering with a slow pathcall), or
JSModuleEnvironment::readVariable()for by-name lookups, module namespace objects, WebAssembly imports and thedebugger.
ResolvedLazyClosureVaris what BytecodeGenerator emits for a module's own declarations; the linker turnsModuleVarand closure variables that are declaration slots intoLazyClosureVar. An empty slot that is not a declarationis still TDZ. The type/control-flow profilers keep the eager path.
ModuleProgramExecutableshare the declarations' FunctionExecutables(
linkedFunctionDeclaration/linkFunctionDeclarationnext tofunctionDeclaration); the record that reads a declarationfirst links it, each record fills its own environment with its own function objects. Whether a slot is a declaration slot
depends only on the module's code, so every CodeBlock of an UnlinkedCodeBlock links a given
get_from_scopethe same way.Optimized code that one loader warmed up meets empty slots when the next loader runs it:
JSTests/modules/module-loaders-lazy-function-declarations.jsdrives that through DFG/FTL code without the testing option.Graph::tryGetConstantClosureVargets no constant.GetLazyClosureVarclobberizes as a read of the slot that can firewatchpoints (
symbolTablePutTouchWatchpointSeton first instantiation) and allocate. The record's list of uninstantiateddeclarations is visited and released under
cellLock().-6.3 MB, RssAnon -7.2 MB. Synthetic (2,000 modules x 46 declarations): FunctionExecutable 154k -> 66k, heap 72.6 -> 59.1 MB,
startup instructions -6.5%.
6. LLInt get_from_scope: test for the closure variable resolve types first
Site mix at the prompt of the CLI application: ClosureVar 54.8%, LazyClosureVar 34.4%, GlobalProperty 9.8%, GlobalVar 1.0%.
Three compares fewer per closure variable read, two more per global one; pays for the empty check of commit 5.
7. Leave a module's uninstantiated function declarations in the bytecode cache payload
With commit 5, a module decoded from a payload that stays around does not create the UnlinkedFunctionExecutables of
declarations nobody read:
m_functionDecls[i]stays null,ModuleFunctionDeclarationSlotsremembers the Decoder and the recordarray,
UnlinkedCodeBlock::functionDecl()decodes on first request (mutator only; other threads usefunctionDeclIfDecoded()). A never-read module function: 256 B of cells -> an 8-byte empty slot + a 4-byte table entry.--helpof the CLI application: 505 M -> 470 M user instructions.8. jsc shell: ASAN builds keep locals on the real stack
detect_stack_use_after_return=0as the shell's default ASAN option: the fake stack is not scanned conservatively, so everymulti-module test died under
--collectContinuously/--slowPathAllocsBetweenGCsin ASAN builds.9. Call sites in LLInt / Baseline metadata get their CallLinkInfo and ArrayProfile on their second execution
Options::useLazyLLIntCallLinkInfosCallSiteData { DataOnlyCallLinkInfo, ArrayProfile }(96 B) allocated on the site's second execution; until then it points at one of two per-VM shared records ("never" /
"once") that look like a polymorphic call to the new
llint_unlinked_callthunk, so the LLInt fast path is the old one plusone load and no branch.
sizeof(Metadata): op_call / call_ignore_result / tail_call 96 -> 8, construct 80 -> 8, super_construct88 -> 16, iterator_open 120 -> 48, iterator_next 136 -> 64.
is fully initialized before a store-store fence publishes it and lives as long as the MetadataTable. Baseline code never
meets a shared record (
setupWithUnlinkedBaselineCodeallocates all of them first). Compiler threads and the collector gothrough
forEachLLIntOrBaselineCallLinkInfo()/getArrayProfile(), which skip sites without their own record ("noinformation").
10. Keep the predictions of a MetadataTable's value profiles out of line, and only for code that has warmed up
Options::useLazyValueProfilePredictions,thresholdForValueProfilePredictions(40)LinkingData, allocated by the firstmergePrediction()with something to record. A still-interpreted CodeBlock below the threshold keeps samples in its buckets: GCend only clears buckets whose cell died.
readers without an array see SpecNone.
11. op_get_from_scope does not need the global variable's watchpoint set in its metadata
24 -> 16 bytes; the DFG parser finds the set in the symbol table entry under the table's lock. Field names kept.
Commits 9-11, synthetic (31.8k interpreter-only blocks): metadata 28.2 -> 15.8 MB (+0.2 out of line), 887 -> 496 B per block,
RssAnon at startup 127.2 -> 113.8 MB. Per call site: never run -73, first run +48, second run +445 instructions.
12. VMInspector::codeBlockCensus and $vm.codeBlockCensus()
Counts used by the tests of commits 13-14 (and by embedders that want to see where code memory goes).
13. Code decoded from a persistent bytecode cache can be dropped and decoded again
Options::useCodeRecoveryFromBytecodeCachePersistentBytecodePayloads) and record offset (32 bits), both in existing padding (sizeofunchanged).UnlinkedFunctionExecutable::returnCodeToCache()puts an executable back into itsm_isCachedstate.VM::shrinkFootprintNow(flags)/shrinkFootprintWhenIdle:KeepCodeThatNeedsParsing,KeepCodeInUse,LeaveCollectionToCaller;Heap::deleteAllUnlinkedCodeBlockstakes which kinds to drop.VM::entryCountFromOutside()andHeap::lastActiveCollectionTime()let the embedder decide when the program is at rest.weakly under the parent's record offset and adopted by the block decoded from that record later, so closures made before and
after a drop share code.
PersistentBytecodePayloadsis one tree of unlinked code = the payload's bytes as decoded forone SourceProvider (embedders wrap the same bytes in a new CachedBytecode + SourceProvider per load). Nothing remembered under one
slot is handed to code decoded under another; code private to one executable (a module whose loader has its own module
scope) is not registered, since its Baseline code assumes one resolution of that scope. Every Decoder and every code block of
a slot hold it (
retain/release,~UnlinkedCodeBlock,~Decoder); with the last one the payload/provider references go and theindex is reused: reloading one file 50 times leaves 3 slots, not 52.
shrinkFootprintNowreturns false (nothing done) under JS or from inside the collector; the Keep modes wait for arunning collection, the flagless mode returns false instead;
lastExceptionis cleared (its stack kept dropped code alive twice);VM::deleteAllCode(flagless) now also deletes what a cache can hand back (heap after shrink + GC 5.4 MB instead of 31.7 MB on asynthetic application). A CodeCache entry decoded from a payload is committed and dropped in every mode: a provider that has
the payload decodes again, one with equal source text but no payload parses.
Heap::deleteAllUnlinkedCodeBlockscompletes the compiler threads' ready plans(that allocates) BEFORE its
HeapIterationScopeand underDeferGC, then re-asserts that no collection runs, so nothing markswhile
returnCodeToCacheturns an executable's two code block slots into {Decoder, offsets}; both directions publish withm_isCachedin the middle (fences), andvisitChildrenreads flag, slots, flag.OnlyWithoutLinkedCode(restricts therecoverable kind) skips blocks anything links against; otherwise all plans are completed first, so no compiler thread holds a
block that is dropped.
Heap::lastActiveCollectionTime()is atomic (zero until the first collection that saw the mutator busy).Lean decoder only.
33 -> 23 MB.
14. Module and program code that has run is released right away
Options::useRunOnceCodeRelease;useSharedModuleFunctionExpressionExecutables(default off)JSModuleRecord::evaluatesees the body finished, the executable drops its linked code, and its unlinked code andCodeCache entry if they can be decoded again; the environment symbol table and the declarations' executables stay.
Interpreter::executeProgramclears the ProgramExecutable it created.one (covers a loader suspended at a top-level await while another finishes). A later loader adopts an executable that only
released recoverable code and has it decoded again (the symbol table and
m_functionDeclarationsare the executable's forlife, commit 1;
ClearCode::Allwithdraws a released executable from adoption). With commit 5/7 the module record no longer referencesreleasable unlinked code for the sake of unread declarations (it decodes them from the payload).
(
GlobalExecutable::canReleaseLinkedCodeNow).extra memory 11.8 -> 1.5 MB, startup instructions -4.4%.
15. Heap: live size after the last collection and total bytes allocated, for embedders
Heap::sizeAfterLastCollection()andHeap::totalBytesAllocated()(running total kept inupdateAllocationLimits()right beforethe per-cycle counters are reset; mutator thread only). For an embedder's idle detection: a busy program whose heap does not grow
still allocates.
Not in this PR: dropping cold LLInt CodeBlocks in any collection and block sealing were written and measured; neither moved
resident memory on the application, both were off by default, so they were left out.
Numbers on the CLI application (RssAnon MB, RssFile 46-51 MB throughout; two to three runs each; rows 1-4 were measured with the series on the previous base
dfd696443b, row 5 on this branch; runs of one round are comparable, rounds differ by up to ~8 MB)shrinkFootprintWhenIdleat deep idle)cf1b36ec8703(same-run baseline 124-127 / 105-106 / 173 / 163-169)CPU (instruction counts; wall clock on the measuring box is noise):
--help, user instructions (like-for-like builds on the previous base)--help, this branch, same binary with every option of the series off / onuseConcurrentJIT=0)Tests
New:
JSTests/modules/lazy-function-declarations*.js(6, run under 32 option sets: tiers, eager thresholds, no-cjit validation,collectContinuously, bytecode cache fill/hit incl. persistent payload),
module-loaders-lazy-function-declarations.js;JSTests/stress/bytecode-cache-unlinked-value-and-array-profiles.js,llint-lazy-call-link-info{,-async,-relink}.js,value-profile-predictions-{on-demand,after-cold-collection}.js,get-from-scope-global-variable-watchpoint-lookup.js,shrink-footprint-recovers-code-from-bytecode-cache.js,module-code-released-after-evaluation.js,module-function-expression-executables-shared.js,module-lazy-function-declarations-outlive-released-code.js,module-loaders-share-released-code.js,module-loaders-share-one-environment-symbol-table.js,shrink-footprint-while-compilations-are-ready.js,shrink-footprint-decodes-code-again-repeatedly.js(also with--destroy-vm),bytecode-cache-persistent-payloads-are-released.js.Results on this branch, against a base list taken with
jscbuilt fromcf1b36ec8703itself (release, and Debug+ASANwhere noted); every difference from base was re-run on both binaries:
cf1b36ec8703)jsc) and passes the tests it adds or changes, at that commitJSTests/modules(122 files), default and no-cjit-validateJSTests/stressx {default, no-cjit-validate, no-JIT, collectContinuously, low thresholds, eager jettison, no concurrent JIT}module-loaders-share-one-environment-symbol-table.js, which fails on the base because of the bug commit 1 fixes); the only others are three tests that flip between runs on both binaries (recursive-try-catch.js,int8-repeat-in-then-out-of-bounds.js,codeblock-aging-ftl-idle.js)bytecode-cache-*.jsand the new cache tests, all their option variantsJSTests/modules, plain and with a persistent payloadmodule-loaders-changed-dependency.js, which writes new source files on every run and so cannot pass a forced cache hit on either binarySeen while testing, not caused by this series (reproduces on the base with no shrink):
ASSERTION FAILED: addResult.isNewEntryinCachedBytecode::copyLeafExecutableson the cache-filling run with--useJIT=0 --forceCodeBlockToJettisonDueToOldAge=1when generatedunlinked code is regenerated after a collection (stale raw-pointer keys in
m_leafExecutables).LowLevelInterpreter32_64.asmis not updated for commits 5 and 9-11 (no 32-bit target is built from this repository).