Skip to content

About

Impute SNPs

Topics

Resources

Stars

0 stars

Watchers

3 watching

Forks

Repository files navigation

imputation

A Snakemake workflow that takes measured genotypes (a VCF, e.g. from the genotyping workflow) and prepares them for single-cell demultiplexing by quality-controlling the calls and imputing missing variants against a 1000 Genomes reference panel.

It contains two workflows that run in sequence:

  1. QC (Snakefile_qc, rules/qc.smk) — filters SNPs, and checks reported sex and assigns ancestry by comparison against a 1000 Genomes reference.
  2. Imputation (Snakefile_imputation, rules/imputation.smk, rules/postimputation.smk) — lifts the genotypes to GRCh38, harmonises and fixes reference alleles, phases, imputes missing variants against a 1000 Genomes reference panel, and filters the result (MAF, imputation R², exonic regions) to produce SNP profiles.

These are the second and third stages of the SNP-based demultiplexing pipeline:

[genotyping](https://github.com/swarbricklab/genotyping)  ->  imputation (this repo: QC + imputation)  ->  [snp_demux](https://github.com/swarbricklab/snp_demux) (Demuxafy)

The final VCF of imputed SNP profiles is consumed by the snp_demux workflow which uses Demuxafy to assign cells to donors in multiplexed 10x Chromium pools.

This workflow was adapted from the sc-eQTLGen consortium WG1 genotype-QC + imputation pipeline (Powell Lab / Garvan; sc-eQTLgen-consortium/WG1-pipeline-QC). See docs/UPSTREAM_DIFFERENCES.md for exactly how this version differs (and how it relates to the Powell Lab's sibling SNP_imputation_1000g_hg38). Citation details are in CITATION.cff and under Attribution below.

Inputs (from genotyping)

  • VCF — genotype calls for all samples (hg19 for QC, hg38 for imputation)
  • psam — a sample description file for plink
  • id map — maps original sample ids to the plink-safe ids in the .psam

Outputs

  • QC: a post-QC set of plink files (pgen/pvar/psam) with sex double-checked and ancestry assigned from the 1000 Genomes reference.
  • Imputation: imputed, filtered SNP profiles ready for Demuxafy.

Rule graphs

Plink QC

QC rule graph

Imputation

Imputation rule graph

Tools

The rules drive standard population-genetics tooling, each pinned to a specific public version:

Step Tool Version
VCF manipulation, reference fixing (+fixref) bcftools 1.10.2
QC, ancestry, format conversion plink / plink2 1.90b6.21 / 2.00a3.7
Missingness, heterozygosity, site filtering vcftools 0.1.16
Liftover hg19 → GRCh38 CrossMap 0.6.5
Strand/allele harmonisation GenotypeHarmonizer 1.4.23
Phasing Eagle 2.4.1
Imputation Minimac4 1.0.2
Het filter + ancestry-PCA plotting (R scripts) R (r-base) 4.3
Initial sample reheader bcftools (biocontainer) 1.21

Versions are pinned to those the upstream monolithic .sif shipped — see docs/UPSTREAM_DIFFERENCES.md.

Containerised, with a conda fallback. The former monolithic image (SNP_imputation_1000g_hg38.sif) has been retired. By default every rule runs from a pinned, public per-rule OCI image under ghcr.io/swarbricklab (built with [absconda(https://github.com/swarbricklab/absconda)), selected via the containers: map in the config; run with --use-singularity (the bundled nci profile already does). GenotypeHarmonizer and Minimac4 (no bioconda package) are baked into their images; the reheader step uses the public quay.io/biocontainers/bcftools:1.21 image.

No Singularity/Apptainer? Every rule also carries a commented conda: directive pointing at workflow/envs/*.yaml (the same pinned versions). Uncomment the conda: blocks and run with --use-conda instead of --use-singularity to build the tools from conda — no container runtime required. See issue #23.

Reference data and provenance

All reference data is public and tracked via dvc import-url from public sources (see the .dvc files under resources/), so it can be fetched without access to any private registry:

  • resources/eQTLGenImpRef.tar.gz — sceQTL-Gen / Powell Lab bundle (Minimac4 imputation reference, 1000 Genomes 30x-GRCh38 phasing/panel, GRCh38 QC FASTA)
  • resources/1000G.tar.gz — 1000 Genomes phase-3 plink (ancestry QC)
  • resources/bed/hg38exonsUCSC.bed — hg38 exon BED
  • resources/liftover/GRCh37_to_GRCh38.chain.gz — Ensembl GRCh37→GRCh38 liftover chain
  • resources/genomes/chr_map/, resources/genomes/GRCh38_chr.fai — small, git-tracked

Fetch, then extract into the paths the config/rules expect:

# If you have access to the project's DVC remote:
dvc pull
# Or, without remote access, re-download straight from the public source URLs:
dvc update resources/eQTLGenImpRef.tar.gz.dvc resources/1000G.tar.gz.dvc \
           resources/bed/hg38exonsUCSC.bed.dvc \
           resources/liftover/GRCh37_to_GRCh38.chain.gz.dvc

./prep.sh    # extract the archives into place; tarballs can then be deleted

Attribution

Adapted from the sc-eQTLGen consortium WG1 genotype-QC + imputation pipeline (sc-eQTLgen-consortium/WG1-pipeline-QC; Powell Lab / Garvan) — see docs/UPSTREAM_DIFFERENCES.md for the differences, including how it relates to the Powell Lab's sibling pipeline powellgenomicslab/SNP_imputation_1000g_hg38. If you use this workflow, please cite:

  • van der Wijst et al. (2020), eLife — the sceQTL-Gen / single-cell eQTL reference approach.
  • Neavin et al. (2024), Genome Biology — Demuxafy, the downstream demultiplexing framework these SNP profiles feed.
  • The reference panels and tools above per their own citation requirements (1000 Genomes Project; Minimac4; Eagle; plink; bcftools; GenotypeHarmonizer; CrossMap).

See CITATION.cff for how to cite this workflow itself.

Running

Each stage has a driver script that runs from the top of a super-project that mounts this repo as a module (see brca_mega_atlas-data):

./modules/imputation/run_qc.sh           # QC stage
./modules/imputation/run_imputation.sh   # imputation stage

On NCI the driver uses the shared PBS profile carried as a nested submodule (profiles/global → swarbricklab/snakemake_config). Check the repo out recursively:

git submodule update --init --recursive

Status

Public, citable (CITATION.cff), and reproducible from public inputs — ready for the v1.0.0 release. Runs are deterministic: reruns from identical inputs are byte-identical. Deliberate divergences from the upstream sc-eQTLGen pipeline are documented in docs/UPSTREAM_DIFFERENCES.md; planned post-1.0 improvements are tracked in the issue tracker.

About

Impute SNPs

Topics

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages