Repository navigation
Release v0.1 - #28
Open
ThrudPrimrose wants to merge 1046 commits into
Open
Release v0.1#28ThrudPrimrose wants to merge 1046 commits into
ThrudPrimrose wants to merge 1046 commits into
Conversation
added 30 commits
October 3, 2026 16:08
…penmp:, MAGMA libomp), offered in that family only; the prompt lists the catalog as the task's language family may link it, and discovery condenses compilers only
…eeps it on NVIDIA where its host code is gcc's
…eep their man pages, so nothing is reinstalled and fontconfig's postinst no longer fails the build
…ces (task text, not package code); the lock resolves CPython 3.12 to 3.14 until the GPU extras ship 3.15 wheels
…typed config accessors, narrowed optionals, shared BuildResult and Sandbox accessors
…mit.sh ENV_ONLY stages a setup's .env alone, the worklist reads a setup name back into its knobs (studies.yaml base = the setups.yaml experiment), and a setup it cannot stage is a problem, not an item graded under defaults
…torch reference and CPF bridge; drop invalid __slots__ from TypedDicts
… gate, and the prompt no longer points agents at man
…ree, native CPU builds, MAGMA in the llvm context only
…rict pyright project
… package, unit constants, HTTPStatus names, strict-typing fixes
… the image's /opt/rocm: PyTorch's ROCm wheel bundled a second libamdhip64 and librccl and crashed every ML grade; the image build asserts it and verify_image gates one HIP runtime and RCCL per process
…ndices, Owned, PromptFold, ResolvedOutput; task_token_totals returns the EpisodeTotals it reads
… string: sizing, scoring, curves, recording, the baseline curve and Harbor take the member; mpi.mode and the DB keep the spelling, parsed or written at that boundary, so a misspelled law fails at the config
… sandbox BuildResult, run_compiled_reference and the C reference a ReferenceRun (outputs, best_ns, hidden, samples_ns)
…l (status, payload), containers.Backends (spellings, passthrough)
…(one failing test: regrade replay StopIteration)
# Conflicts: # hpcagent_bench/stats/figures/scaling.py
…regrade worklist stages each setup through submit.sh
…t; env_spec's Model enum (members from the layer files) carries a reasoned ignore
…uncalled static inline in C++, and a kernel with no heap allocation never calls it (18 warnings broke the -Wall -Wextra ratchet)
…on errors carry their SDFG, fp16 mixed operands, framecode allocation placement)
…extended now refuses a gap in tuple return names (__return_1 with no __return_0)
…omic and a running job keeps the inode it mounted, so only the in-place writes (build, export, pull) refuse a mounted path
…alar version (two arrays sized by one runtime scalar combine again)
…ror, detail) NamedTuple; callers read .ok. measurement_statistics documents the ML final grade's 4 aligned inputs in one launch
…unner owns its cores: every leg saw all 128 and sized its OpenMP, BLAS and numba pools off them, so twelve legs at once oversubscribed the node until serial steps read as hangs (Phase 4b: 67 s alone, past its 7.5 min budget in 669803); -n auto gets one worker per CPU of the slice
…and numba on two workqueue threads: in the judge image numpy shares the kernels' OpenMP OpenBLAS, sized off 128 cores before the oracle's *_NUM_THREADS defaults, and a parallel region in a forked child of the parent's libgomp pool waits forever (gdb: GOMP_parallel(num_threads=128) under the kernel's cblas GEMM). gpt2_block c/cpp hung to FAIL:timeout, cegterg numba (REQUIRE_OK) and nine more numba kernels to skip:too-long; in the image all are ok in 3-101 s. Comments state the reasons in the present tense
… one included (triton, an offload setup's host grade); only a CPU-track grade keeps fork/forkserver. A device runtime the parent started does not survive a fork: in the gpu step test_an_offload_kernel_is_timed_without_its_transfers' host-resident cupy transfer read 'Out of memory allocating 256,000,000 bytes (allocated so far: 0)' once earlier tests had touched the GPU in the pytest process; with the spawn the step's device tests pass in the image
…olumn's device-resident contract: ppcg's mirrors are stripped (device_resident_host), so the host half carries no hipMalloc/Memcpy/Free to find and the kernel takes device pointers; the test stages A, B, C with cupy as copy_func does, checks the hip error API marks the translation and no mirror survives, and reads C back. Docstrings in the file state the failures they pin in the present tense
…he description-length bound of the skill index, the one-clause and imperative-verb shape of a when: trigger, and the CPF tool description, skill page and reminder phrase checks. The parse, non-empty, unique-trigger and reminder-attached checks stay
…om (test_every_cache_is_typed)
…ide any Slurm job: inside one the episode takes SLURM_JOB_ID and the job IS NULL row it reads back does not exist
…et, untouched but one page: a 4x scratch was served from heap the forked child inherited free from the parent (mimalloc keeps freed segments mapped), so the kernel-budget call did not fail whenever an earlier test in the worker had freed a gigabyte
…d started from forkserver (GPU runtime loaded) pickles them Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e copies it; *.csv was ignored, so a clean build had no file) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…test correct submission; a passing /score is promoted only when the slot has no correct submission of the kernel Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s follow the one rule: a correct-but-slower latest submission is owed, and a slot with a correct submission promotes no /score Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ource: an unsourced one cannot be regraded, so it never blocks the slot's grade Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…5 20 slots, temperature3); register the temperature3 study Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e judge node (one 11 h pass) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ts newest episode (a relaunched slot was owed one per episode) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing (out, x, batch_size, dim); env.sh and the suite floor only the soft core limit; the suite records the host checkout's commit; the env-staging test clears the suite's mi200 hardware Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lver containers/images/submit.sh <job.sbatch> [--system S] [--partition P] [--account A] [--time T] [--gpus-per-node N] [--nice N] [--dry-run] [-- job args] reads the partition, account and GPUs per node from `hpcagent-bench job options` (flag over variable or site layer over systems.yaml), parsing its flags with submit_common.sh's parse_job_flags. A build_and_verify.sbatch or verify_image.sbatch role runs on its images.env partition unless --partition or --system names one. docs/jobs samples no longer hardcode --gpus-per-node=4: `hpcagent-bench job submit` adds the system's count; the docs that started them with plain sbatch now use it (or `job options`). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Container and helper jobs take sbatch options from the one resolver
…ader over the system `hpcagent-bench job submit` now starts every job script, the container jobs included, and containers/images/submit.sh is gone: its one extra, an image build or verify job's images.env partition, is a pin of the job itself in systems.py. A field is its flag, else Slurm's own SBATCH_* variable (which beats #SBATCH under plain sbatch too), else the script's leading #SBATCH line (long or short form), else its HPCAGENT_BENCH_* variable, else the system's entry; the system fills only what the script does not pin, so build_and_verify.sbatch's 1 task x 96 cores stays 1 x 96 on Beverin. The GPU pair and the new task pair (--ntasks / --ntasks-per-node) each come whole from the highest source. Without --system, a partition picks the entry of the same cluster that names it (--partition mi200 on Beverin is beverin-mi200: 8 GPUs, 4 x 16). --ntasks and --nice are fields now. The docs/jobs samples drop their Beverin task shape from the header and srun lines, so the system's shape reaches them (Daint gets 4 x 72 instead of a header-pinned 4 x 24). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
One Slurm launcher: job submit for every job, the script's #SBATCH header over the system
…LLM_ROCM_USE_AITER=1) on mi300" This reverts commit a4c225f. Under real agent load (solver14 670522, temperature3 670523/670524) the vLLM ROCM_AITER_FA line garbles qwen3.8's output: random multilingual tokens in the reasoning and invented tool names in 39-45 of 60 agents, no judge call in solver14. The synthetic serving gates (300-500 output tokens) did not catch it. qwen3.8 serves on SGLang again until the vLLM line passes a long-reasoning coherence check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… has the judge tools The MCP server's command is the driver's interpreter, the image's launch venv, which lands under /tmp wherever the image binds no /opt/node-shm (the agent images). The seal mounts a private tmpfs over /tmp, so claude could not start the server, reported it failed, and every sealed agent ran with Read/Edit/Bash only (solver14 669854-670629, temperature3 670523-670628). seal_worker --keep holds an O_PATH handle on a directory before the private /tmp lands and binds it back read-only at its own path; the driver keeps its own venv when it lives under /tmp. Verified in the agent image: hpcagent_bench MCP status failed without --keep, connected with it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…assed comm, unsealed draws spy - scripts/host_python.sh is sourced (env.sh), so it floors the soft core limit only, like site_env.sh and cache_env.sh; the hard 0 it set made tests/test_core_dump_guard.py's precondition fail in the judge image. - The gloo real-launch kernels use the mpi4py communicator the driver passes a python kernel. - The fresh-draws final-grade test runs unsealed: a sealed host call from a parent that mapped a GPU runtime starts from the forkserver, which neither sees the patched variant_for nor unpickles a closure. - test_runner_stub_gemm_ok names the failing row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Latest main merged (#2623, #2659, #2515, #2661) and the DaCe-frontend fixes for the HB CI failures at dea0b39c6 (port lowering, numeric agreement, frontend validity). No public API change; new knob compiler.cpu.simd_maps (default true). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…that score is unlimited; temperature3 is single-submission - start_agent ends an agent whose MCP server never connects (McpUnavailable) instead of running it with Read/Edit/Bash only: such a run finishes rc 0 and grades nothing it was meant to. - submission-single.md opens with score UNLIMITED, then submit exactly ONE. - temperature3 takes the single-submission policy like solver14. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… restored) 558522989 shipped dace/external/cub and dace/external/moodycamel as symlinks instead of gitlinks, so a fresh install had no moodycamel headers and every DaCe compile failed. ba01a0b7b restores the gitlinks and changes nothing else. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.