Skip to content

Repository files navigation

RaptQA

Annealing-guided RNA aptamer design with a multi-temperature Restricted Boltzmann Machine trained on HT-SELEX data.

RaptQA learns a sequence landscape from changes in sequence frequencies across High-Throughput SELEX (HT-SELEX) rounds and formulates candidate design as a quadratic unconstrained binary optimization (QUBO) problem. The released workflow covers three RNA aptamer targets: Integrin αVβ3, hFGF9, and human transglutaminase 2 (TG2).

The pipeline has four stages:

  1. Extract random regions from FASTQ files and build round-wise read-count tables.
  2. Train a multi-temperature RBM (multi-beta RBM) across selection rounds.
  3. Export or solve the resulting QUBO with an Ising-machine or classical backend.
  4. Evaluate held-out model fit, filter diverse candidate sequences, and analyze surface plasmon resonance (SPR) measurements for experimentally tested candidates.

Installation

git clone https://github.com/hmdlab/RaptQA.git
cd RaptQA
python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip setuptools wheel

Install PyTorch for the intended platform, followed by RaptQA. The paper environment used Python 3.11.3 and PyTorch 2.7.1 with CUDA 12.6. RaptQA requires PyTorch 2.6 or later; earlier releases are affected by GHSA-53q9-r3pm-6pq6, and the released workflows and tests were validated with PyTorch 2.7.1.

# GPU environment used for the paper
pip install --index-url https://download.pytorch.org/whl/cu126 torch==2.7.1

# Alternatively, for CPU-only use
# pip install --index-url https://download.pytorch.org/whl/cpu torch==2.7.1

pip install -e .

constraints/paper-recorded.txt lists the package versions recorded in the publication artifacts. It is a partial record of the paper environment; constraints/README.md identifies the supporting metadata and explains how to use it.

Install the optional Amplify SDK integration for remote QUBO backends, or the development dependencies for testing and linting:

pip install -e ".[qubo]"
pip install -e ".[ci]"

Quick start

The repository includes an author-generated synthetic test fixture: a small toy dataset that exercises training, evaluation, and local simulated QUBO solving on a CPU.

raptqa-train --config config/tests/toy_smoke.toml
raptqa-eval --config config/reference/eval_toy.toml
raptqa-qubo --config config/reference/qubo_toy.toml

Training writes a Result_*/ directory below selex_data/toy_data_10/. The smoke driver repeats this workflow for sequence-only inputs and for dot-bracket secondary-structure features, and can also run an optional multiple-sequence-alignment (MSA) feature pass:

PYTHONPATH=. python tests/run_smoke_suite.py

The MSA pass uses mlocarna from LocARNA and RNAalifold and RNAfold from the ViennaRNA Package. In CI, it runs only when SELEX_RBM_SMOKE_MSA=1 is set and is skipped when any of these executables is unavailable. An error from an installed tool fails the pass. See third-party requirements for installation and licensing information.

Paper data and reproduction

Processed count tables are included for all three targets:

Dataset Target Random region Rounds DRA accession
RAPT3 Integrin αVβ3 40 nt R3-R6 DRA009384
RAPT4 hFGF9 35 nt R1-R6 DRA019577
RAPT1 TG2 30 nt R0-R8 DRA009383

RAPT1 is the historical dataset identifier for TG2. Its count tables retain 28-32 nt extracted reads; the model loader selects the read-count-weighted modal length of 30 nt. The fixed splits, model-input length selection, and byte-level checksums are documented in the dataset notes for RAPT1, RAPT3, and RAPT4.

The reference configurations document the recorded training protocol. Checkpoint provenance, exact file hashes, and the limit on byte-identical retraining are documented in paper_results/README.md.

See docs/REPRODUCTION.md for training and evaluation commands, publication and smoke configurations, external-solver requirements, and artifact provenance.

Released artifacts are organized as follows:

paper_results/
|-- models/       # Representative RBM checkpoints
|-- generated/    # Recorded solver output pools
|-- candidates/   # Candidate tables used in the study
|-- source_data/  # Figure/table values and their source map
|-- source_inputs/ # Minimal inputs used to assemble source_data
|-- qubo_inputs/  # Solver-independent QUBO coefficients and variable maps
`-- SHA256SUMS.txt

The repository contains these study artifacts under paper_results/. They are not included in the Python distributions: the wheel contains the installable package, while the source distribution also contains the documentation, reference configurations, and synthetic toy fixture needed for the quick start. Both omit the study data and trained checkpoints.

Use paper_results/source_data/figure_table_source_map.csv to locate the data for a manuscript item. Current solver comparisons are under paper_results/source_data/solver/. Portable QUBO inputs can be inspected or regenerated without solver credentials.

Raw FASTQ files for the three study datasets are not duplicated in the repository. With SRA Toolkit installed, the public DRA records can be downloaded and processed with:

bash scripts/download_raw_data.sh all

The script writes replay output under each dataset's raw/ directory for comparison and does not overwrite the versioned count tables.

Documentation

The main source package is selex_rbm/; command-line entry points are raptqa-train, raptqa-eval, raptqa-qubo, raptqa-preprocess, and raptqa-sample. Reference protocols are in config/reference/, small test protocols in config/tests/, and analysis utilities in scripts/analysis/.

Testing

pytest -q
PYTHONPATH=. python tests/run_smoke_suite.py

Citation

If you use RaptQA in your research, please cite:

Shumpei Uno, Shigetaka Nakamura, Daiki Kawahara, Yoshikazu Nakamura, Tatsuhiko Shirai, Dai Fujiwara, Tatsuo Adachi, Nozomu Togawa & Michiaki Hamada. "Learning fitness landscapes from HT-SELEX for annealing-guided RNA aptamer design." Manuscript (2026).

Machine-readable metadata is provided in CITATION.cff.

License

These grants apply only to rights held by the RaptQA contributors.

Releases

Packages

Used by

Contributors

Languages