Repository navigation
Conversation
Replace hardcoded cluster paths, partitions, email and study-specific inputs in the SLURM .sh runners with placeholders and a required PIPELINE_ROOT env var. Move commands that preceded the #SBATCH block below it so sbatch parses the directives. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
generate_slurm.py now derives the repo root from its own location (or PIPELINE_ROOT), takes partition/email from args or env with no baked-in default, and only writes GPU feature constraints when asked. Skill docs ask the user for partition, email and reference paths instead of assuming one cluster. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace sys.path entries pointing at one cluster checkout with paths computed from each script's location, load evaluation resources from the repo's Resources/ dir, make the mini-dataset builder take its input as a CLI argument, and drop cluster paths from test docstrings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add a public-repository rule section to CLAUDE.md, replace the cluster environment and resource-path blocks with generic PIPELINE_ROOT / Resources wording, and make setup_resources.sh take its source dir from RESOURCES_SRC. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Delete skill evaluation/workspace outputs that recorded study runs, and gitignore tasks/, .baton/ and skill workspace/output dirs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace cell-type-specific defaults in the annotation code (prompt role and context, PubMed keyword, evidence-scoring keywords) with generic ones; the PubMed keyword is now optional. Vertex AI project, location and bucket come from env vars or CLI with no built-in project. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tools/check_no_lab_specific_content.py scans tracked files (or staged / ref-relative added lines) for forbidden patterns, with exceptions in tools/lab_specific_allowlist.txt. Runs in CI on push/PR and as a local pre-commit hook. CLAUDE.md points to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--Conditions (plotting) and --Sample (excel summary) no longer default to one study's labels; when omitted, labels are the unique values of the categorical key in the h5mu. Runners stop passing hardcoded labels and docs use placeholder labels. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Excel summary's --Sample flag is now --Conditions (--Sample kept as a deprecated alias with a notice). --categorical_key and --Conditions help texts are unified across stages, docs and the runner skill use the condition wording, and the drift check understands flag aliases. CHANGELOG records the rename and the scrub changes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Stage 1: expect loading/cNMF_{scores,loadings}_{K}_{thresh with _}.txt,
matching compile_results; parallel tests now use the
{run}_{K}/Inference/cnmf_tmp layout that rename_all_NMF reads.
- Stage 2 U-test: call U_test functions with explicit
out_dir/run_name/K/sel_threshs instead of a module-global args.
- Stage 3 KSelection: load_perturbation_data(conditions=...), 'condition'
column, *_per_condition / *_all_conditions plot names.
- Stage 3 Gene: pass gene_name_key='symbol'; assert LazyGeneCorr /
LazyPerturbCorr and check rows against the dense Pearson matrix.
- get_significant_programs_df: keep columns when nothing is significant
instead of raising KeyError on set_index; add unit test.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Compile_excel_sheet.py: summary functions take conditions= (matching load_perturbation_data and --Conditions); callers, README, notebook and tests updated. simple_Summary_cols locals renamed so they no longer shadow the parameter. - get_significant_programs_df: '# programs <condition>' is always int (was str counts mixed with int 0); test checks dtype and values. - Gene corr test: compare only uniquely named genes. - KSelection conftest: drop unused synthetic_test_stats_df fixture. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Base:
scrub-lab-specific-content(not merged yet; it rewrites paths in these files). Retarget tomainonce it lands.Problem
A full test run on an HPC cluster (main @ d5289bf, 2026-09-27) failed in four suites. The tests had drifted from the code, and one real bug turned up.
Fix
loading/cNMF_scores_5_2.0.txt..._5_2_0.txt(whatcompile_resultswrites)rename_all_NMFfound no spectra{run}_{K}/Inference/cnmf_tmplayout that the pipeline usesout_dir, run_name, K, sel_threshsU_testfunctions with explicit args (no module-globalargs)load_perturbation_data(samples=)conditions=,conditioncolumn,*_per_condition/*_all_conditionsplot namesisinstance(DataFrame), DID NOT RAISELazyGeneCorr/LazyPerturbCorrand check rows against the dense Pearson matrix; passgene_name_key='symbol'get_significant_programs_dfKeyError: 'target_name'when nothing is significantResult (cluster, SLURM)
* Skips are CRT tests (skipped on main as well) plus tests that need untracked resources (GWAS file, guide annotation). Rerun with those resources present: test_metrics + test_utest 61 passed, 0 skipped; KSelection 20 passed, 0 skipped. The only remaining skips are the CRT tests.
🤖 Generated with Claude Code