You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
On a warm start, clangd builds every module's BMI again instead of reusing the ones the cold start left on disk. The engine plan's compile arguments differ between a model that comes from the producer (mcpp emit build-database) and the same model read back from the model cache. clangd keys its persistent module cache on the full command, so none of the cold start's BMIs match.
On the mcpp repository (fixture self-mcpp, 176 modules) this costs a full rebuild on every start after the first: 55–78 s to first navigation on a 4-core runner, with no module primed (the primer sees BMIs on disk and skips them, but clangd itself rebuilds each one).
Evidence
Measured with the temporary instrumentation in draft #29 (not for merge), on ubuntu-24.04, 4 cores. Each start ran on the same workspace and cache directory, cold and then warm.
87.8 s (preparation window 158.7 s, 174 modules primed)
(ring truncated)
-
176
warm
55.5 s (0 modules primed)
176
0
352
An earlier run (36415909922) measured the same pattern: cold 126.0 s, warm 78.5 s to first navigation, and 0 modules primed on the warm start.
Same unit, different cache key. clangd's module cache is modules/<unit>-<source-path hash>/<command hash>/, and the command hash covers the working directory and every argument except the output (getCompileCommandStringHash in ModulesBuilder.cpp). For std.cc the source directory is the same on both starts (std.cc-AAAEC3D741C2BE82), but the command hash is 25864D8C218D24EC on the cold start and 1A91DE07102C408E on the warm one.
The arguments differ. Here is main.cpp's command as the server log shows it, first 12 arguments:
cold (model from the producer) warm (model from the cache)
clang++ clang++
--driver-mode=g++ --driver-mode=g++
-I<ws>/modules/libs/src/json -std=c++23
-I<gtest>/googletest/include -D__mcpp_target_linux__=1
-I<gtest>/googlemock/include -I<ws>/modules/libs/src/json
-std=c++23 -I<gtest>/googletest/include
-O2 -I<gtest>/googlemock/include
--sysroot=<subos> --sysroot=<subos>
-D__mcpp_target_linux__=1 --no-default-config
--no-default-config --target=x86_64-linux-gnu
--target=x86_64-linux-gnu -stdlib=libstdc++
-stdlib=libstdc++ --gcc-install-dir=...
On the warm start the arguments are reordered, and -O2 is gone, which also changes predefined macros such as __OPTIMIZE__ compared with the cold start. The warm start's log says the cached mcpp model ... matches its inputs; planning with it while the producer confirms it, followed by project model ... loaded again: unchanged. Because the producer confirms the model as unchanged, no replan happens, and the whole session keeps the cache-derived arguments.
Likely location.src/project/modelcache.cpp stores the model's database through spec::to_json(model.database) (S1 structured form). Reading it back does not reproduce the producer's argument list: the order is normalized and -O2 is dropped. I have not pinned down which step drops -O2.
Impact
Every time the model source switches between the producer and the cache, clangd rebuilds all modules. That happens on the first reopen after a cold start, and again when the cache goes stale and the producer answers. For the mcpp repository that is 176 BMIs and 1–4 minutes of CPU-bound work on 4 cores.
The two paths give clangd semantically different commands (-O2 and __OPTIMIZE__).
The existing check module-cache-reused does not catch this. It looks at std only, and in the timing fixture both paths happen to produce the same arguments.
Suggested fix
Make both paths produce byte-identical engine arguments. Either keep the producer's raw arguments in the cache so the round trip is the identity, or normalize both paths through one canonicalization before the engine plan (fixed order, and a stated rule for -O* and other flags that do not affect semantics).
Add a regression check: on a real-project fixture, start cold and then warm, and require zero Built module lines in clangd's log on the warm start (or zero new .pcm files). Draft [DO NOT MERGE] temporary: time clangd's module BMI rebuild on real projects #29 has the pieces: the clangd log ring written out once preparation is idle, and a .pcm snapshot per start.
Context: found while measuring how much of a cold start is BMI rebuilding, for a design study of a universal module interface format.
Summary
On a warm start, clangd builds every module's BMI again instead of reusing the ones the cold start left on disk. The engine plan's compile arguments differ between a model that comes from the producer (
mcpp emit build-database) and the same model read back from the model cache. clangd keys its persistent module cache on the full command, so none of the cold start's BMIs match.On the mcpp repository (fixture
self-mcpp, 176 modules) this costs a full rebuild on every start after the first: 55–78 s to first navigation on a 4-core runner, with no module primed (the primer sees BMIs on disk and skips them, but clangd itself rebuilds each one).Evidence
Measured with the temporary instrumentation in draft #29 (not for merge), on
ubuntu-24.04, 4 cores. Each start ran on the same workspace and cache directory, cold and then warm.Run 36419266307, artifact
bmi-timing:Built moduleReusing persistent module.pcmfiles afterAn earlier run (36415909922) measured the same pattern: cold 126.0 s, warm 78.5 s to first navigation, and 0 modules primed on the warm start.
Same unit, different cache key. clangd's module cache is
modules/<unit>-<source-path hash>/<command hash>/, and the command hash covers the working directory and every argument except the output (getCompileCommandStringHashinModulesBuilder.cpp). Forstd.ccthe source directory is the same on both starts (std.cc-AAAEC3D741C2BE82), but the command hash is25864D8C218D24ECon the cold start and1A91DE07102C408Eon the warm one.The arguments differ. Here is
main.cpp's command as the server log shows it, first 12 arguments:On the warm start the arguments are reordered, and
-O2is gone, which also changes predefined macros such as__OPTIMIZE__compared with the cold start. The warm start's log saysthe cached mcpp model ... matches its inputs; planning with it while the producer confirms it, followed byproject model ... loaded again: unchanged. Because the producer confirms the model as unchanged, no replan happens, and the whole session keeps the cache-derived arguments.Likely location.
src/project/modelcache.cppstores the model's database throughspec::to_json(model.database)(S1 structured form). Reading it back does not reproduce the producer's argument list: the order is normalized and-O2is dropped. I have not pinned down which step drops-O2.Impact
-O2and__OPTIMIZE__).module-cache-reuseddoes not catch this. It looks atstdonly, and in thetimingfixture both paths happen to produce the same arguments.Suggested fix
-O*and other flags that do not affect semantics).Built modulelines in clangd's log on the warm start (or zero new.pcmfiles). Draft [DO NOT MERGE] temporary: time clangd's module BMI rebuild on real projects #29 has the pieces: the clangd log ring written out once preparation is idle, and a.pcmsnapshot per start.Context: found while measuring how much of a cold start is BMI rebuilding, for a design study of a universal module interface format.