Skip to content

Latest commit

 

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CAMM: Chromatin Accessibility to Metastatic Mutagenesis

Code, trained checkpoints, and public inputs for the manuscript "Chromatin accessibility of primary cancers informs regional mutagenesis in metastases through multi-scale deep learning".

CAMM is a hierarchical, multi-scale, multi-task neural network for predicting single-nucleotide variant (SNV) and indel density at 1 Mb, 100 kb, and 10 kb resolution from chromatin accessibility (CA) and replication timing (RT). The study uses metastatic whole-genome data from the Hartwig Medical Foundation (HMF) for training and PCAWG primary tumors for external validation.

Repository layout

Code/
  step1/          # Hyperparameter search and model training
  step2/          # Cross-validation, ablations, baselines, and PCAWG validation
  step3/          # Permutation importance and SHAP attribution
  step4/          # Mutation-enriched windows and gene annotation
Data/
  CA_RT/          # Chromatin accessibility and replication timing features
  PCAWG/          # Public validation mutation-count tables
Model/            # Six cancer-specific PyTorch checkpoints
Figure_script/    # Python and R scripts for Figures 1–4
docs/
  parameters.md   # Required inputs, optional parameters, and implementation notes
requirements.txt  # Python dependency inventory

Installation

Use Python 3.11 for the following setup. The shell commands use macOS/Linux syntax. Training uses CUDA when available and otherwise runs on CPU.

git clone --depth 1 https://github.com/reimandlab/CAMM.git
cd CAMM
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt "pandas==2.3.3"

For R figure scripts, install:

install.packages(c("ggplot2", "dplyr", "readr", "tidyr", "ggnewscale",
                   "cowplot", "reshape2", "patchwork", "reticulate", "fs"))

Data availability

Location Contents
Data/CA_RT/ CA/RT features at three resolutions: 796 TCGA ATAC-seq profiles and 96 ENCODE Repli-seq profiles
Data/PCAWG/ PCAWG SNV and indel validation count tables at three resolutions

HMF whole-genome data and metastatic sample metadata are controlled access. Requests are submitted through the HMF data-access procedure and require approval and the applicable data access or material transfer agreements. HMF-derived intermediate files are not publicly shared because of data-use restrictions.

Feature tables are tab-separated, optionally gzip-compressed; mutation tables are comma-separated. Both use chr and start coordinates. Mutation targets are selected by cancer-type column.

Reconstruct the chromosome-split 10 kb feature matrix before using the full data:

python Data/CA_RT/atac_with_repliseq_10kb/combine_chr_tsv.py Data/CA_RT/atac_with_repliseq_10kb --output Data/CA_RT/atac_with_repliseq.10kb.tsv.gz

Trained checkpoints

Model/ contains PyTorch checkpoints for breast, colorectal, esophagus, lung, prostate, and skin cancers. Scripts with --model_dir, --best_dir, or --model_outdir look for best_model.pt in the supplied directory.

Checkpoint use requires the matching model architecture, feature order, and preprocessing.

Analysis and figure workflow

Run scripts from the repository root. Main training requires feature, SNV, and indel paths at all three resolutions, --ctype to select the cancer column, and --outdir for outputs. Inspect command-line help after installation:

python Code/step1/run_model_hier_multi.py --help
python -m Code.step3.feature_importance --help

The feature-importance script uses module invocation to resolve its model import.

Stage Scripts
Tune and train Optuna search; hierarchical multi-task model
Cross-validation Full model; single-task ablation; single-scale ablation
Paired comparisons Multi-task vs single-task; multi-scale vs single-scale
Random Forest / XGBoost baselines Independent task/scale models
PCAWG validation Frozen-model evaluation and calibration
Residual analysis 10 kb SNV predictions; 10 kb indel predictions
Feature importance Permutation importance and SHAP
Gene annotation Mutation-enriched windows and cancer-gene annotation
Figure Scripts in Figure_script/ Inputs
1 ATAC-seq heatmaps; Repli-seq heatmaps CA/RT features and HMF mutation tables
2 Model comparison; legend Model comparison and baseline summary tables; legend is self-contained
3 Feature correlations; SHAP plots SHAP tables and selected-feature lists
4 Residual distributions; per-cancer residual z-scores; gene windows; z-scores Residual, gene annotation, and enrichment summary tables

Run Python figure scripts with python Figure_script/<filename>.py and R scripts with Rscript Figure_script/<filename>.R, after preparing their required inputs. Some scripts use historical paths, including lowercase data/; set the documented path options or edit local path settings to match your files.

Support

Questions and reproducibility issues can be submitted through GitHub Issues.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages