Skip to content

Ship parallel engines on every platform at a fixed -j 8 (macOS static libomp, Windows vcomp140 beside the engine) (#1923) - #1925

Merged
swapnilpaliwal-sd merged 2 commits into
devfrom
fix/parallel-engines-everywhere
Oct 11, 2026
Merged

swapnilpaliwal-sd merged 2 commits into
devfrom
fix/parallel-engines-everywhere

Conversation

@swapnilpaliwal-sd

@swapnilpaliwal-sd swapnilpaliwal-sd commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #1923 — every platform ships PARALLEL (OpenMP) engines, solving at a fixed -j 8, with no switch.

Before this, only Linux shipped parallel: macOS engines were built without OpenMP, and Windows was made serial in #1914 because /openmp needs vcomp140.dll, absent without the Visual C++ Redistributable.

What changes (.github/workflows/build-engines.yml, graph/pipeline/run-souffle.sh, packaging/engine-package.json)

  • macOS (arm64, x64): OpenMP with libomp.a (built for macOS 12) linked statically; otool -L shows only /usr/lib/libc++ and libSystem. A guard fails the build if an engine links anything outside /usr/lib and /System.
  • Windows: /openmp again, with vcomp140.dll (and whatever VC runtime DLL it imports; none for MSVC 14.44) copied beside each engine, so a machine with no VC++ Redistributable runs it. The guard now requires every DLL an engine (or a shipped DLL) imports to be part of Windows or shipped beside it; the engine package carries the DLLs.
  • Linux: unchanged (static libgomp, build-engines: engines start on a clean machine — static libgomp on Linux, no OpenMP on Windows (#1913) #1914).
  • Run time: a packaged parallel engine always solves at -j 8; no environment variable changes it. A local source compile keeps its probe and AXIOM_SOLVE_PARALLEL / AXIOMCODE_SOLVE_THREADS.

Validation (packaged engines + packed CLI, npm install -g, clean environment)

platform full suite parallel = serial (direct re-solve, 5 runs × java/py/ts/js/c#) crashes
darwin-arm64 (this Mac, Homebrew/souffle off PATH) e2e 222/227 0 fail · MCP 40/40 · extras 80/81 0 fail 0 differing relations, every run 0
win32-x64 (stock Windows Server 2022, no VC++ redist, 8 vCPU) e2e 222/227 0 fail · MCP all pass · extras 80/81 0 fail 0 differing relations, every run 0

Solve time, serial → -j 8 (avg of 5):

project (facts) macOS arm64 Windows
Python, larger (200 MB) — 31.4 s → 15.3 s
Java, larger (113 MB) — 13.8 s → 9.9 s
JavaScript (33 MB) 4.73 → 3.48 s 7.25 → 4.70 s
C# (11 MB) 0.81 → 0.61 s 2.75 → 1.49 s
Python (16 MB) 1.45 → 1.15 s 4.69 → 2.8 s
Java (10 MB) 0.65 → 0.54 s 2.74 → 2.60 s
TypeScript (6 MB) 0.50 → 2.13 s 1.87 → 1.29 s

Known trade-off: on a very small input (6 MB of facts) macOS TypeScript is slower at -j 8 than serial; accepted — the gain is on large repositories.

darwin-x64 (x64 Node under Rosetta on Apple Silicon, clean environment): e2e 222/227, 0 failed · MCP 40/40 · extras 80/81, 0 failed · crashscan 0 crash markers. Parallel = serial: 0 differing relations at index level (java 39, python 95, ts 58, js 51, c# 44 relations) and in all 5 direct parallel re-solves per language; 0 crashes. Timing under Rosetta, with the case suites running on the same machine, so it is noisy: serial build → -j 8 median: java 2.1 → 2.5 s, python 2.7 → 4.2 s, ts 1.8 → 9.3 s, js 7.8 → 8.7 s, c# 2.6 → 1.1 s. On these small inputs -j 8 under translation does not pay; native Intel Macs were not available to measure.

Case suites on this branch (macOS arm64, local compile): java (--oracle --no-torture), typescript (--oracle), python, javascript (--oracle), csharp: all pass. Query cases (tests/run.py): all five languages pass. The first local run had 2 java checks fail in library-auto-discovers-imported-dependencies because no Vineflower decompiler was installed locally (CI installs it). The csharp suite needs the Roslyn oracle built. With both in place everything passes, and engine-package-test.sh passes too. A wrapper around the installed engine confirmed the packaged path passes -j 8 whatever AXIOM_SOLVE_PARALLEL (unset/0/1) or AXIOMCODE_SOLVE_THREADS says.

swapnilpaliwal-sd and others added 2 commits October 10, 2026 20:10
macOS: build a static libomp from the pinned LLVM release with a macOS 12
deployment target and link it by path (-Xclang -fopenmp), so an engine loads
only /usr/lib/libc++ and libSystem; write the .parallel marker. A new guard
fails the build on anything outside /usr/lib and /System, a minos above 12.0,
or a .parallel engine without the OpenMP runtime in it.

Windows: compile the language engines with /openmp again and copy
vcomp140.dll (plus any VC runtime DLL it imports) from the MSVC redist folder
beside each engine; write the .exe.parallel marker. The guard now requires
every DLL an engine or a shipped DLL imports to be part of Windows or beside
it. The engine package's files list carries */*.dll.

run-souffle.sh: AXIOM_SOLVE_PARALLEL=0 now holds a packaged parallel engine
to -j 1 too, and the PARALLEL SOLVE comments describe every platform.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
A packaged engine is the parallel flavor on every published platform, so
run-souffle.sh passes a fixed -j 8 whenever the .parallel marker is beside it:
no AXIOM_SOLVE_PARALLEL or thread-count variable changes the shipped path, and
every user's solve is the configuration the release validated. A local compile
keeps its OpenMP probe, AXIOM_SOLVE_PARALLEL and AXIOMCODE_SOLVE_THREADS.

Co-authored-by: axiomcode-bot[bot] <334110751+axiomcode-bot[bot]@users.noreply.github.com>
@swapnilpaliwal-sd
swapnilpaliwal-sd merged commit 297f6db into dev Oct 11, 2026
11 checks passed
@swapnilpaliwal-sd
swapnilpaliwal-sd deleted the fix/parallel-engines-everywhere branch October 11, 2026 04:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

macOS and Windows engines ship serial: package OpenMP so every platform solves in parallel

1 participant