Skip to content

Remove lab-specific paths and defaults; add a public-content guard - #14

Merged
ymo6 merged 11 commits into
mainfrom
scrub-lab-specific-content
Sep 29, 2026
Merged

ymo6 merged 11 commits into
mainfrom
scrub-lab-specific-content

Conversation

@engreitz

Copy link
Copy Markdown
Contributor

Problem

PerturbNMF is a public repo, but main carries lab- and study-specific details: cluster paths and partitions, personal emails, dated run directories, hardcoded condition labels (D0 sample_D1 …), and tracked skill output folders. New users can't run the SLURM scripts without editing them, and study details leak into a public tool.

Fix

  • SLURM runners: PIPELINE_ROOT is required; <partition> / <your_email> placeholders; study paths → /path/to/.... Moved set -euo pipefail below #SBATCH in six runners (SLURM was silently ignoring their directives).
  • Python entry points: hardcoded sys.path → paths relative to __file__; eval resources from the repo's Resources/.
  • Skills, CLAUDE.md, READMEs, setup_resources.sh, environment.yml, notebooks: site-specific defaults removed; notebook outputs cleared.
  • Conditions: labels auto-detected from obs[categorical_key] when --Conditions is omitted; excel summary --Sample → --Conditions (alias kept, with a deprecation note).
  • Guard: tools/check_no_lab_specific_content.py runs in CI and as a pre-commit hook, with an allowlist for authorship metadata.

Result

Check Result
Full test suite, main vs this branch (cluster, CPU + GPU) Identical pass/fail/skip on every test
Condition auto-detect / --Sample alias smoke tests Pass
Guard CI Pass
bash -n on all runners; drift check Clean

The failures that already exist on main (stale test signatures, 2.0 vs 2_0 output names) are fixed in a separate PR.

Recommendation

Merge before the annotation-v3 and CRT PRs; both touch the same runner and skill files.

🤖 Generated with Claude Code

engreitz and others added 11 commits September 27, 2026 22:12
Replace hardcoded cluster paths, partitions, email and study-specific
inputs in the SLURM .sh runners with placeholders and a required
PIPELINE_ROOT env var. Move commands that preceded the #SBATCH block
below it so sbatch parses the directives.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
generate_slurm.py now derives the repo root from its own location (or
PIPELINE_ROOT), takes partition/email from args or env with no baked-in
default, and only writes GPU feature constraints when asked. Skill docs
ask the user for partition, email and reference paths instead of
assuming one cluster.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace sys.path entries pointing at one cluster checkout with paths
computed from each script's location, load evaluation resources from the
repo's Resources/ dir, make the mini-dataset builder take its input as a
CLI argument, and drop cluster paths from test docstrings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add a public-repository rule section to CLAUDE.md, replace the cluster
environment and resource-path blocks with generic PIPELINE_ROOT /
Resources wording, and make setup_resources.sh take its source dir from
RESOURCES_SRC.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Delete skill evaluation/workspace outputs that recorded study runs, and
gitignore tasks/, .baton/ and skill workspace/output dirs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace cell-type-specific defaults in the annotation code (prompt role
and context, PubMed keyword, evidence-scoring keywords) with generic
ones; the PubMed keyword is now optional. Vertex AI project, location
and bucket come from env vars or CLI with no built-in project.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tools/check_no_lab_specific_content.py scans tracked files (or staged /
ref-relative added lines) for forbidden patterns, with exceptions in
tools/lab_specific_allowlist.txt. Runs in CI on push/PR and as a local
pre-commit hook. CLAUDE.md points to it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--Conditions (plotting) and --Sample (excel summary) no longer default
to one study's labels; when omitted, labels are the unique values of the
categorical key in the h5mu. Runners stop passing hardcoded labels and
docs use placeholder labels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Excel summary's --Sample flag is now --Conditions (--Sample kept as
a deprecated alias with a notice). --categorical_key and --Conditions
help texts are unified across stages, docs and the runner skill use the
condition wording, and the drift check understands flag aliases.
CHANGELOG records the rename and the scrub changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@engreitz
engreitz requested a review from ymo6 September 29, 2026 03:54
@ymo6
ymo6 merged commit add0212 into main Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants