Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 8 additions & 5 deletions docs/pysceptre-backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,8 @@ them needs R.
| `src/run_power_simulation.R` + `lib/simulate.R` + `lib/pert_input.R` | **ported to Python**; the NB draw, guide-to-guide variability, centering and seeding are WattEG's method and stay in WattEG (§6) |
| `src/consolidate_replicates.R` | ported (pyarrow) |
| `src/compute_power.R` + `lib/stats.R` | ported (Wilson interval) |
| `src/summarize_power.R`, `src/fit_power_curve.R` | ported |
| `src/summarize_power.R` | ported |
| `src/fit_power_curve.R` | **deleted** (2026-09-28): the analytical estimate is PerturbPlan's closed form, `pysceptre.analytical_power` |
| `patches/`, `lib/apply_patch.R`, `src/check_sceptre_api.R`, `src/install_sceptre.R`, `src/install_ondisc.R`, `src/audit_dependencies.R` | **deleted.** All six exist only because the pipeline reaches into unexported sceptre S4 slots and patches its CRT path |
| `lib/sceptre_io.R` | **deleted from the pipeline**; its odm-materialisation logic moves to the one-off export (§4) |
| `src/make_test_data.R` | ported, or replaced by a synthetic fixture generated in Python |
Expand Down Expand Up @@ -655,8 +656,8 @@ one stage with an absolute bar rather than a relative one.
`compute_power` and `summarize_power`, checked against R **on identical input**, which makes
them exact comparisons rather than statistical ones: every column agrees to machine epsilon,
in R's column order. The comparison caught `max_effect_size_tested` missing entirely and the
per-gene columns sitting in the wrong place. `fit_power_curve` is **not** ported — it serves
the paper's reduced-design study rather than the pipeline, and nothing in the DAG calls it.
per-gene columns sitting in the wrong place. `fit_power_curve` was not ported, and was
deleted on 2026-09-28; nothing in the DAG called it.
6. **Nextflow rewiring** — done: `FIT_NULL_MODELS` and `MERGE_NULL_MODELS` are gone, the six
remaining processes call the `watteg-*` entry points, the samplesheet column is `dataset`
(a `.h5mu`) rather than `sceptre_object`, and `pixi.toml` holds no R. **Still open: the new
Expand Down Expand Up @@ -825,8 +826,10 @@ figure always clears the 4 GB floor), while the 1,000 Python cis tasks peaked at
The largest, simulating from `exp(X . beta)`, has been taken; the other two are judgement calls
about which cells and which estimator belong, and both change every published number.

**One thing deliberately not ported.** `fit_power_curve.R` serves the paper's reduced-design study
rather than the pipeline, and nothing in the DAG calls it. It stays in R.
**One thing not ported, and since deleted.** `fit_power_curve.R` fitted a per-pair probit power
curve for the paper's reduced-design study; nothing in the DAG called it. It was removed on
2026-09-28: the analytical estimate is PerturbPlan's closed form (`pysceptre.analytical_power`),
and the simulation stays the reference.

**The synthetic fixture.** `src/make_test_data.R` produces a sceptre object; the Python path needs
a `.h5mu`, so `assets/samplesheet_synthetic.csv` points at a file nothing generates yet. The stub
Expand Down
20 changes: 6 additions & 14 deletions docs/status.md
Original file line number Diff line number Diff line change
Expand Up @@ -941,20 +941,12 @@ published since. This one happened to come back clean; the next one may not. Pin
their locked versions would make future re-solves honest, and is worth doing before the next
dependency change rather than during one.

### Step 11 — fit the power curve and run three effect sizes instead of six

`src/fit_power_curve.R` fits each pair's power curve so a sweep needs three effect sizes rather
than six, and makes the minimum detectable effect size continuous rather than snapped to whichever
effect sizes were run. Usage is in [Usage]({{ site.baseurl }}{% link usage.md %}); the model is
derived in [Methods]({{ site.baseurl }}{% link methods.md %}).

**The validation — whether three points really do reproduce six, the covariate model that predicts
the curve for pairs never simulated, and the calibrated prediction intervals — moved to
[broadinstitute/WattEG-paper](https://github.com/broadinstitute/WattEG-paper) on 2026-09-04**, with
the rest of the paper analyses. Short version of what it established, so a pipeline user knows what
the feature is worth: the best three-point design reproduced the six-point per-pair MDES for 92.0 %
of pairs on `day0`, against a reference whose own bootstrapped self-agreement was 91.4 %. Grid
placement is dataset-specific and a one-effect-size pilot is enough to choose it.
### Step 11 — fit the power curve (removed)

`src/fit_power_curve.R` fitted each pair's power curve so a sweep could run three effect sizes
rather than six. It was removed on 2026-09-28, with the rest of the project's own power model:
the analytical estimate is PerturbPlan's closed form (`pysceptre.analytical_power`), and the
simulation stays the reference. The fit and its validation are in git history.

## Running the comparisons on the cluster

Expand Down
281 changes: 0 additions & 281 deletions src/fit_power_curve.R

This file was deleted.

5 changes: 3 additions & 2 deletions src/summarize_power.R
Original file line number Diff line number Diff line change
Expand Up @@ -126,8 +126,9 @@ for (column in shared) {

# Per-gene values from sim_input, joined here rather than carried through the simulation.
#
# WHY HERE. The theory behind the power curve is SE^2 ~ (1/n_pert_cells) * (1/mu + 1/theta), so any
# covariate model of power needs the gene's dispersion as well as its expression. Only
# WHY HERE. The test statistic's variance goes as SE^2 ~ (1/n_pert_cells) * (1/mu + 1/theta), so an
# analytical power estimate (PerturbPlan's closed form) needs the gene's dispersion as well as its
# expression. Only
# `average_expression_all_cells` reaches this table through the per-simulation output, so every such
# analysis has had to load a 16 MB sim_input.rds to find the other half. Joining it here costs one
# read and makes the summary self-sufficient.
Expand Down
Loading
Loading