A Snakemake workflow that takes measured genotypes (a VCF, e.g. from the genotyping workflow) and prepares them for single-cell demultiplexing by quality-controlling the calls and imputing missing variants against a 1000 Genomes reference panel.
It contains two workflows that run in sequence:
- QC (
Snakefile_qc,rules/qc.smk) — filters SNPs, and checks reported sex and assigns ancestry by comparison against a 1000 Genomes reference. - Imputation (
Snakefile_imputation,rules/imputation.smk,rules/postimputation.smk) — lifts the genotypes to GRCh38, harmonises and fixes reference alleles, phases, imputes missing variants against a 1000 Genomes reference panel, and filters the result (MAF, imputation R², exonic regions) to produce SNP profiles.
These are the second and third stages of the SNP-based demultiplexing pipeline:
[genotyping](https://github.com/swarbricklab/genotyping) -> imputation (this repo: QC + imputation) -> [snp_demux](https://github.com/swarbricklab/snp_demux) (Demuxafy)
The final VCF of imputed SNP profiles is consumed by the snp_demux workflow which uses Demuxafy to assign cells to donors in multiplexed 10x Chromium pools.
This workflow was adapted from the sc-eQTLGen consortium WG1 genotype-QC + imputation pipeline (Powell Lab / Garvan; sc-eQTLgen-consortium/WG1-pipeline-QC). See
docs/UPSTREAM_DIFFERENCES.mdfor exactly how this version differs (and how it relates to the Powell Lab's siblingSNP_imputation_1000g_hg38). Citation details are inCITATION.cffand under Attribution below.
- VCF — genotype calls for all samples (hg19 for QC, hg38 for imputation)
- psam — a sample description file for
plink - id map — maps original sample ids to the
plink-safe ids in the.psam
- QC: a post-QC set of
plinkfiles (pgen/pvar/psam) with sex double-checked and ancestry assigned from the 1000 Genomes reference. - Imputation: imputed, filtered SNP profiles ready for Demuxafy.
The rules drive standard population-genetics tooling, each pinned to a specific public version:
| Step | Tool | Version |
|---|---|---|
VCF manipulation, reference fixing (+fixref) |
bcftools |
1.10.2 |
| QC, ancestry, format conversion | plink / plink2 |
1.90b6.21 / 2.00a3.7 |
| Missingness, heterozygosity, site filtering | vcftools |
0.1.16 |
| Liftover hg19 → GRCh38 | CrossMap |
0.6.5 |
| Strand/allele harmonisation | GenotypeHarmonizer |
1.4.23 |
| Phasing | Eagle |
2.4.1 |
| Imputation | Minimac4 |
1.0.2 |
| Het filter + ancestry-PCA plotting (R scripts) | R (r-base) |
4.3 |
| Initial sample reheader | bcftools (biocontainer) |
1.21 |
Versions are pinned to those the upstream monolithic .sif shipped — see
docs/UPSTREAM_DIFFERENCES.md.
Containerised, with a conda fallback. The former monolithic image (
SNP_imputation_1000g_hg38.sif) has been retired. By default every rule runs from a pinned, public per-rule OCI image underghcr.io/swarbricklab(built with [absconda(https://github.com/swarbricklab/absconda)), selected via thecontainers:map in the config; run with--use-singularity(the bundlednciprofile already does). GenotypeHarmonizer and Minimac4 (no bioconda package) are baked into their images; the reheader step uses the publicquay.io/biocontainers/bcftools:1.21image.No Singularity/Apptainer? Every rule also carries a commented
conda:directive pointing atworkflow/envs/*.yaml(the same pinned versions). Uncomment theconda:blocks and run with--use-condainstead of--use-singularityto build the tools from conda — no container runtime required. See issue #23.
All reference data is public and tracked via dvc import-url from public
sources (see the .dvc files under resources/), so it can be fetched without
access to any private registry:
resources/eQTLGenImpRef.tar.gz— sceQTL-Gen / Powell Lab bundle (Minimac4 imputation reference, 1000 Genomes 30x-GRCh38 phasing/panel, GRCh38 QC FASTA)resources/1000G.tar.gz— 1000 Genomes phase-3 plink (ancestry QC)resources/bed/hg38exonsUCSC.bed— hg38 exon BEDresources/liftover/GRCh37_to_GRCh38.chain.gz— Ensembl GRCh37→GRCh38 liftover chainresources/genomes/chr_map/,resources/genomes/GRCh38_chr.fai— small, git-tracked
Fetch, then extract into the paths the config/rules expect:
# If you have access to the project's DVC remote:
dvc pull
# Or, without remote access, re-download straight from the public source URLs:
dvc update resources/eQTLGenImpRef.tar.gz.dvc resources/1000G.tar.gz.dvc \
resources/bed/hg38exonsUCSC.bed.dvc \
resources/liftover/GRCh37_to_GRCh38.chain.gz.dvc
./prep.sh # extract the archives into place; tarballs can then be deletedAdapted from the sc-eQTLGen consortium WG1 genotype-QC + imputation pipeline
(sc-eQTLgen-consortium/WG1-pipeline-QC;
Powell Lab / Garvan) — see docs/UPSTREAM_DIFFERENCES.md for
the differences, including how it relates to the Powell Lab's sibling pipeline
powellgenomicslab/SNP_imputation_1000g_hg38. If you use this workflow, please cite:
- van der Wijst et al. (2020), eLife — the sceQTL-Gen / single-cell eQTL reference approach.
- Neavin et al. (2024), Genome Biology — Demuxafy, the downstream demultiplexing framework these SNP profiles feed.
- The reference panels and tools above per their own citation requirements (1000 Genomes Project; Minimac4; Eagle; plink; bcftools; GenotypeHarmonizer; CrossMap).
See CITATION.cff for how to cite this workflow itself.
Each stage has a driver script that runs from the top of a super-project that mounts this repo as a module (see brca_mega_atlas-data):
./modules/imputation/run_qc.sh # QC stage
./modules/imputation/run_imputation.sh # imputation stageOn NCI the driver uses the shared PBS profile carried as a nested submodule
(profiles/global → swarbricklab/snakemake_config). Check the repo out
recursively:
git submodule update --init --recursivePublic, citable (CITATION.cff), and reproducible from public inputs — ready for the
v1.0.0 release. Runs are deterministic: reruns from identical inputs are byte-identical.
Deliberate divergences from the upstream
sc-eQTLGen pipeline are documented in docs/UPSTREAM_DIFFERENCES.md; planned
post-1.0 improvements are tracked in the
issue tracker.