Utilities for preparing whole-slide-image patch features and spatial transcriptomics gene/morphology features used by downstream MIL and graph models.
This repository does not include CLAM patch extraction code. It consumes patch folders generated by the modified CLAM preprocessing code at https://github.com/wwyi1828/CLAM. That CLAM branch aligns patch coordinates to a global 224-step grid instead of starting a separate coordinate grid from each contour's local bounding box.
For WSI feature preprocessing, the full flow is:
- Clone and set up the modified CLAM preprocessing repo:
git clone https://github.com/wwyi1828/CLAM.git-
Use that CLAM repo to extract PNG patches from raw WSIs.
-
Clone this repo and update one example config, such as
examples/configs/camelyon16_clubyol_r18.yaml, so that:src_folderpoints to the CLAM-generatedpatches/folderlabel_filepoints to a CSV withfilename,typexml_folderpoints to annotation XMLs when patch-level ratios are neededdst_fileis the output.pklpath or output H5 folderckpt_pathpoints to the model checkpoint when using a custom ResNet
-
Run
pipeline/extract_pretrain_feats.pyto produce downstream-ready feature files.
For caption preprocessing, use preprocessing/create_labels.py to convert a
SlideInstruct-style caption JSON into a filename,type CSV, then pass that CSV
or the caption JSON in the WSI feature extraction config.
For HEST-style spatial transcriptomics preprocessing, use the gene/morphology
entrypoints directly on a root containing st/, patches/, and metadata/.
Install the Python packages required by the task you run. The WSI feature path uses PyTorch, torchvision, h5py, pandas, shapely, rtree, lxml, transformers, and timm. Gene/morphology preprocessing additionally uses scanpy, anndata, scipy, and pybiomart.
First use the modified CLAM repository, outside this repo, to extract 224 x 224 PNG patches:
python /path/to/CLAM/create_patches_png.py \
--source /path/to/wsi_folder \
--save_dir /path/to/clam_patches \
--patch_size 224 \
--step_size 224 \
--seg \
--patchThe scripts in this repo expect the resulting patch directory to be shaped like:
/path/to/clam_patches/patches/
slide_001/
0_0.png
224_0.png
slide_002/
...
Extract one .pkl file containing one object per slide:
python pipeline/extract_pretrain_feats.py \
--config examples/configs/camelyon16_clubyol_r18.yaml \
--model_type ResNet18 \
--batch_size 256For UNI, set UNI_CKPT_DIR or pass --uni_ckpt_dir. For CONCH and UNIv2,
pass --conch_ckpt_path or --univ2_ckpt_path.
The output slide objects use these fields:
x: patch featurespos: patch coordinates parsed from patch filenamesy: optional patch-level annotation overlap ratiosslide_y: optional slide-level one-hot labelslide_caption: optional caption textslide_index: slide identifier
To write one .h5 per slide, set dst_file in the YAML to a directory path
instead of a .pkl path. H5 outputs use feats, cords, ratios, label,
and caption.
Create a small label CSV from a SlideInstruct-style caption JSON:
python preprocessing/create_labels.py \
--input-json examples/data/slideinstruct_caption.example.json \
--output-csv outputs/slidebench_labels.csvTo attach captions during WSI feature extraction, provide caption_file in the
YAML config, as shown in examples/configs/slidebench_caption_conch.yaml.
For HEST-style data, the expected root contains:
/path/to/HEST1K/
st/<sample_id>.h5ad
patches/<sample_id>.h5
metadata/<sample_id>.json
Run the unified preprocessing path:
python pipeline/extract_molmor_feats_unified.py \
--data_root /path/to/HEST1K \
--subtype_json examples/data/hest_subtypes.example.json \
--dst_file outputs/hest_example \
--gene_prc \
--imge_prc UNI \
--uni_ckpt_dir /path/to/UNIThe older per-dataset and breast-specific entrypoints are kept for compatibility:
python pipeline/extract_molmor_feats.py \
--dataset MISC_brain \
--data_root /path/to/HEST1K \
--dst_file outputs/hest_single \
--gene_prc
python pipeline/extract_molmor_feats_breast.py \
--dataset SPA_breast \
--data_root /path/to/HEST1K \
--dst_file outputs/hest_breast \
--gene_prcConvert a .pkl feature file to one .h5 file per slide:
python maintenance/convert_pkl_to_h5folder.py \
--input outputs/C16_CluBYOL_R18.pkl \
--output outputs/C16_CluBYOL_R18_h5Convert an H5 folder back to a .pkl:
python maintenance/convert_h5folder_to_pkl.py \
--input outputs/C16_CluBYOL_R18_h5 \
--output outputs/C16_CluBYOL_R18_roundtrip.pklRecompute annotation overlap ratios if XML annotations or coordinate scaling need to be revisited:
python maintenance/recompute_ratios.py \
--input outputs/C16_CluBYOL_R18.pkl \
--xml-folder /path/to/lesion_annotations \
--output outputs/C16_CluBYOL_R18_with_ratios.pkl