Code, trained checkpoints, and public inputs for the manuscript "Chromatin accessibility of primary cancers informs regional mutagenesis in metastases through multi-scale deep learning".
CAMM is a hierarchical, multi-scale, multi-task neural network for predicting single-nucleotide variant (SNV) and indel density at 1 Mb, 100 kb, and 10 kb resolution from chromatin accessibility (CA) and replication timing (RT). The study uses metastatic whole-genome data from the Hartwig Medical Foundation (HMF) for training and PCAWG primary tumors for external validation.
Code/
step1/ # Hyperparameter search and model training
step2/ # Cross-validation, ablations, baselines, and PCAWG validation
step3/ # Permutation importance and SHAP attribution
step4/ # Mutation-enriched windows and gene annotation
Data/
CA_RT/ # Chromatin accessibility and replication timing features
PCAWG/ # Public validation mutation-count tables
Model/ # Six cancer-specific PyTorch checkpoints
Figure_script/ # Python and R scripts for Figures 1–4
docs/
parameters.md # Required inputs, optional parameters, and implementation notes
requirements.txt # Python dependency inventory
Use Python 3.11 for the following setup. The shell commands use macOS/Linux syntax. Training uses CUDA when available and otherwise runs on CPU.
git clone --depth 1 https://github.com/reimandlab/CAMM.git
cd CAMM
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt "pandas==2.3.3"For R figure scripts, install:
install.packages(c("ggplot2", "dplyr", "readr", "tidyr", "ggnewscale",
"cowplot", "reshape2", "patchwork", "reticulate", "fs"))| Location | Contents |
|---|---|
| Data/CA_RT/ | CA/RT features at three resolutions: 796 TCGA ATAC-seq profiles and 96 ENCODE Repli-seq profiles |
| Data/PCAWG/ | PCAWG SNV and indel validation count tables at three resolutions |
HMF whole-genome data and metastatic sample metadata are controlled access. Requests are submitted through the HMF data-access procedure and require approval and the applicable data access or material transfer agreements. HMF-derived intermediate files are not publicly shared because of data-use restrictions.
Feature tables are tab-separated, optionally gzip-compressed; mutation tables are comma-separated. Both use chr and start coordinates. Mutation targets are selected by cancer-type column.
Reconstruct the chromosome-split 10 kb feature matrix before using the full data:
python Data/CA_RT/atac_with_repliseq_10kb/combine_chr_tsv.py Data/CA_RT/atac_with_repliseq_10kb --output Data/CA_RT/atac_with_repliseq.10kb.tsv.gzModel/ contains PyTorch checkpoints for breast, colorectal, esophagus, lung, prostate, and skin cancers. Scripts with --model_dir, --best_dir, or --model_outdir look for best_model.pt in the supplied directory.
Checkpoint use requires the matching model architecture, feature order, and preprocessing.
Run scripts from the repository root. Main training requires feature, SNV, and indel paths at all three resolutions, --ctype to select the cancer column, and --outdir for outputs.
Inspect command-line help after installation:
python Code/step1/run_model_hier_multi.py --help
python -m Code.step3.feature_importance --helpThe feature-importance script uses module invocation to resolve its model import.
| Stage | Scripts |
|---|---|
| Tune and train | Optuna search; hierarchical multi-task model |
| Cross-validation | Full model; single-task ablation; single-scale ablation |
| Paired comparisons | Multi-task vs single-task; multi-scale vs single-scale |
| Random Forest / XGBoost baselines | Independent task/scale models |
| PCAWG validation | Frozen-model evaluation and calibration |
| Residual analysis | 10 kb SNV predictions; 10 kb indel predictions |
| Feature importance | Permutation importance and SHAP |
| Gene annotation | Mutation-enriched windows and cancer-gene annotation |
| Figure | Scripts in Figure_script/ |
Inputs |
|---|---|---|
| 1 | ATAC-seq heatmaps; Repli-seq heatmaps | CA/RT features and HMF mutation tables |
| 2 | Model comparison; legend | Model comparison and baseline summary tables; legend is self-contained |
| 3 | Feature correlations; SHAP plots | SHAP tables and selected-feature lists |
| 4 | Residual distributions; per-cancer residual z-scores; gene windows; z-scores | Residual, gene annotation, and enrichment summary tables |
Run Python figure scripts with python Figure_script/<filename>.py and R scripts with Rscript Figure_script/<filename>.R, after preparing their required inputs. Some scripts use historical paths, including lowercase data/; set the documented path options or edit local path settings to match your files.
Questions and reproducibility issues can be submitted through GitHub Issues.