Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PatchPreprocess

Utilities for preparing whole-slide-image patch features and spatial transcriptomics gene/morphology features used by downstream MIL and graph models.

This repository does not include CLAM patch extraction code. It consumes patch folders generated by the modified CLAM preprocessing code at https://github.com/wwyi1828/CLAM. That CLAM branch aligns patch coordinates to a global 224-step grid instead of starting a separate coordinate grid from each contour's local bounding box.

End-to-end Workflow

For WSI feature preprocessing, the full flow is:

  1. Clone and set up the modified CLAM preprocessing repo:
git clone https://github.com/wwyi1828/CLAM.git
  1. Use that CLAM repo to extract PNG patches from raw WSIs.

  2. Clone this repo and update one example config, such as examples/configs/camelyon16_clubyol_r18.yaml, so that:

    • src_folder points to the CLAM-generated patches/ folder
    • label_file points to a CSV with filename,type
    • xml_folder points to annotation XMLs when patch-level ratios are needed
    • dst_file is the output .pkl path or output H5 folder
    • ckpt_path points to the model checkpoint when using a custom ResNet
  3. Run pipeline/extract_pretrain_feats.py to produce downstream-ready feature files.

For caption preprocessing, use preprocessing/create_labels.py to convert a SlideInstruct-style caption JSON into a filename,type CSV, then pass that CSV or the caption JSON in the WSI feature extraction config.

For HEST-style spatial transcriptomics preprocessing, use the gene/morphology entrypoints directly on a root containing st/, patches/, and metadata/.

Install

Install the Python packages required by the task you run. The WSI feature path uses PyTorch, torchvision, h5py, pandas, shapely, rtree, lxml, transformers, and timm. Gene/morphology preprocessing additionally uses scanpy, anndata, scipy, and pybiomart.

Patch Input

First use the modified CLAM repository, outside this repo, to extract 224 x 224 PNG patches:

python /path/to/CLAM/create_patches_png.py \
  --source /path/to/wsi_folder \
  --save_dir /path/to/clam_patches \
  --patch_size 224 \
  --step_size 224 \
  --seg \
  --patch

The scripts in this repo expect the resulting patch directory to be shaped like:

/path/to/clam_patches/patches/
  slide_001/
    0_0.png
    224_0.png
  slide_002/
    ...

WSI Feature Extraction

Extract one .pkl file containing one object per slide:

python pipeline/extract_pretrain_feats.py \
  --config examples/configs/camelyon16_clubyol_r18.yaml \
  --model_type ResNet18 \
  --batch_size 256

For UNI, set UNI_CKPT_DIR or pass --uni_ckpt_dir. For CONCH and UNIv2, pass --conch_ckpt_path or --univ2_ckpt_path.

The output slide objects use these fields:

  • x: patch features
  • pos: patch coordinates parsed from patch filenames
  • y: optional patch-level annotation overlap ratios
  • slide_y: optional slide-level one-hot label
  • slide_caption: optional caption text
  • slide_index: slide identifier

To write one .h5 per slide, set dst_file in the YAML to a directory path instead of a .pkl path. H5 outputs use feats, cords, ratios, label, and caption.

Caption and Label Preprocessing

Create a small label CSV from a SlideInstruct-style caption JSON:

python preprocessing/create_labels.py \
  --input-json examples/data/slideinstruct_caption.example.json \
  --output-csv outputs/slidebench_labels.csv

To attach captions during WSI feature extraction, provide caption_file in the YAML config, as shown in examples/configs/slidebench_caption_conch.yaml.

Gene and Morphology Preprocessing

For HEST-style data, the expected root contains:

/path/to/HEST1K/
  st/<sample_id>.h5ad
  patches/<sample_id>.h5
  metadata/<sample_id>.json

Run the unified preprocessing path:

python pipeline/extract_molmor_feats_unified.py \
  --data_root /path/to/HEST1K \
  --subtype_json examples/data/hest_subtypes.example.json \
  --dst_file outputs/hest_example \
  --gene_prc \
  --imge_prc UNI \
  --uni_ckpt_dir /path/to/UNI

The older per-dataset and breast-specific entrypoints are kept for compatibility:

python pipeline/extract_molmor_feats.py \
  --dataset MISC_brain \
  --data_root /path/to/HEST1K \
  --dst_file outputs/hest_single \
  --gene_prc

python pipeline/extract_molmor_feats_breast.py \
  --dataset SPA_breast \
  --data_root /path/to/HEST1K \
  --dst_file outputs/hest_breast \
  --gene_prc

Conversion Utilities

Convert a .pkl feature file to one .h5 file per slide:

python maintenance/convert_pkl_to_h5folder.py \
  --input outputs/C16_CluBYOL_R18.pkl \
  --output outputs/C16_CluBYOL_R18_h5

Convert an H5 folder back to a .pkl:

python maintenance/convert_h5folder_to_pkl.py \
  --input outputs/C16_CluBYOL_R18_h5 \
  --output outputs/C16_CluBYOL_R18_roundtrip.pkl

Recompute annotation overlap ratios if XML annotations or coordinate scaling need to be revisited:

python maintenance/recompute_ratios.py \
  --input outputs/C16_CluBYOL_R18.pkl \
  --xml-folder /path/to/lesion_annotations \
  --output outputs/C16_CluBYOL_R18_with_ratios.pkl

About

Convert CLAM-extracted WSI patches into training-ready feature files

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages