🔥 Seeking expert review — see REVIEWERS.md. Tag v2.2.0: panel-based MRD benchmark, all data open-access. PR open for review: #2
DeepCatch is an open-source computational framework for multi-cancer early detection (MCED) from cell-free DNA (cfDNA). It fuses 7 complementary molecular modalities through a self-supervised Transformer foundation model, tracks patients longitudinally with Bayesian Kalman filtering, and predicts tissue-of-origin — all in a single two-stage CET (Capture → Enhance → Triage) pipeline.
graph LR
subgraph ThreeRepos["3-repo open-source portfolio"]
DC["rollroyces/deepcatch<br/>(this repo)<br/>v2.2"]
FP["rollroyces/cfdna-fragmentomics-pipeline<br/>v0.9+<br/>627 cross-study samples"]
METH["rollroyces/deepcatch-methylation<br/>Phase 0–1<br/>FinaleMe HMM"]
end
TCGA[("TCGA GDC<br/>open-access MAFs")]
FinaleDB[("FinaleDB<br/>pan-cancer WGS")]
Zenodo[("Zenodo CRAG<br/>HCC samples")]
GEO[("GEO<br/>FLARE/GSE317007")]
TCGA -->|real mutations| DC
FinaleDB -->|WGS fragments| FP
Zenodo -->|HCC validation| FP
GEO -->|ONT cfDNA| DC
DC <-->|fusion adapter| FP
FP -.->|β-values| METH
BioRxiv["bioRxiv paper<br/>(drafted, in prep)"]
DC & FP & METH --> BioRxiv
classDef repo fill:#ddf4ff,stroke:#0969da,color:#0a3069
classDef data fill:#fff8c5,stroke:#bf8700,color:#3d2c00
classDef out fill:#dafbe1,stroke:#1a7f37,color:#116329
class DC,FP,METH repo
class TCGA,FinaleDB,Zenodo,GEO data
class BioRxiv out
💚 Sponsor this work: See .github/SPONSORS.md for tier descriptions ($5/$49/$499 monthly). 100% of funds go to compute and maintenance. GitHub Sponsors →
v2.1 adds GNN methylation field-defect detection, enhanced fragmentomics (DELFI + MFS + nucleosome + refined 5-mer), cfSort-style tissue deconvolution, a multi-modal foundation model, and priming agent PK/PD simulation.
📋 Active project: Methylation channel extension (Phase 0-2, started 2026-09-10). Uses FinaleMe (Liu 2024, Nat Commun) to impute CpG methylation from existing FinaleDB WGS — no new raw data needed. See
METHYLATION_PROJECT.mdfor the full plan.
⚠️ Research-stage software. Not for clinical diagnosis. See §11 for real-plasma validation status.
DeepCatch is a two-stage CET pipeline (Capture → Enhance → Triage). Stage 1 fuses seven molecular modalities from a single cfDNA sample through a Transformer foundation model; Stage 2 accumulates evidence across longitudinal draws via Bayesian Kalman filtering; Triage compares the posterior to a calibrated threshold.
flowchart TB
Sample["cfDNA Sample<br/>(BAM / FASTQ)"]
subgraph S1["Stage 1 — Capture (single draw)"]
direction TB
M1["Fragmentomics Basic<br/>MFR, FSI, CAFF, FEM"]
M2["Enhanced Fragmentomics<br/>DELFI + MFS + nucleosome"]
M3["CNV<br/>6-D chromosomal instability"]
M4["Serological<br/>PG-I, PG-II, G-17, Hp"]
M5["GNN Methylation Network<br/>GATv2 field-defect"]
M6["Tissue Deconvolution<br/>cfSort-style DNN (24-D)"]
M7["Priming Agents<br/>PK/PD + denoising"]
FM["Multi-Modal Foundation Model<br/>4-layer Transformer"]
Sample --> M1 & M2 & M3 & M4 & M5 & M6 & M7
M1 & M2 & M3 & M4 & M5 & M6 & M7 --> FM
end
subgraph S2["Stage 2 — Enhance (longitudinal)"]
KF["Bayesian Kalman Filter (BSSLM)<br/>accumulates p_cancer across draws"]
FM -->|joint embedding| KF
end
subgraph S3["Triage"]
THR{"p_cancer > τ?"}
KF --> THR
THR -->|Yes| CONF["Confirmatory testing"]
THR -->|No| RET["Clear until next draw"]
end
classDef stage fill:#e1f5ff,stroke:#0969da,color:#0a3069
classDef mod fill:#fff8c5,stroke:#bf8700,color:#3d2c00
classDef decision fill:#ffd7d5,stroke:#cf222e,color:#82071e
class S1,S2,S3 stage
class M1,M2,M3,M4,M5,M6,M7,FM mod
class THR,CONF,RET decision
See V3_DESIGN.md for the GPU-accelerated v3 design proposal that adds methylation β-value embeddings as a 6th channel on Apple MPS.
Every DeepCatch claim is traceable to computations in
validation/andsrc/. No numbers are invented. No clinical claims are intended.
The repo contains two distinct validation surfaces; reviewers should not conflate them:
| Component | Validation surface | Reference |
|---|---|---|
real_tcga_validation.py (panel-LLR, Fisher, strand) |
TCGA GDC open-access MAFs (real somatic mutations) + Poisson-sampled cfDNA reads at target VAF. AUC 0.92 at 0.1% VAF on 20 LUAD patients; +18pp Sens@99% with CADD Top-K=200. | results/real_tcga_validation.json, docs/CADD_* |
src/fragmentomics/tumor_naive_adapter.py + fusion_ablation.py |
FinaleDB public cfDNA WGS fragments (170 MB per sample). 627-sample cross-study pan-cancer cohort, 5-channel profile (5Mb + 100kb short/long ratio + coverage + 196-bin FSD). AUC 0.9753 ± 0.002 (tumor-naive alone); 0.9886 ± 0.001 (naive-avg fusion vs synthetic mutation channel @ AUC 0.92). DeLong p<0.0015 across 5 seeds. | docs/FUSION_ISOTONIC.md, sister repo cfdna-fragmentomics-pipeline |
src/fragmentomics/normalization.py (DELFI LOESS GC correction + per-bin median-centering) |
Standard DELFI protocol from Cristiano et al. 2019 (Nature 570:385-389). LOESS smoothing with statsmodels; quadratic polynomial fallback. |
validation/cfdna/ artifacts |
src/fragmentomics/themis_features.py (MFR, FSI, CAFF, FEM) |
THEMIS gastric-cancer detection framework (Bie 2023, Nature Communications). 4-mer FEM calibrated on Jiang 2020 (Cancer Discovery). | run_jiang_analysis.py --nested-cv --report |
src/multimodal_fusion/advanced_fusion.py (PyTorch rewrite) |
Real PyTorch modules (was random-init numpy placeholders). Stratified val split + NaN guards + biologically-informed prior masks (cancer_detection, tissue_of_origin). Trained end-to-end via AdamW + early stopping. Synthetic-data AUC ≥ 0.99 on separable data; no real-cfDNA training yet (the foundation-real-smoke job wires it to a real TCGA panel signal — see below). |
commit 64c5aeb |
scripts/foundation_real_smoke.py (NEW) |
First real-data CI smoke for the foundation model. Trains FoundationDownstream on real TCGA-LUAD panel-LLR scores + real per-patient mutation-derived features (20 patients, both channels from real TCGA MAFs). Two foundation variants are weighted-averaged (70/30): (B) frozen random-init encoder + sklearn LR head and (A) 3-ensemble of tiny PyTorch encoders. Plus lr_baseline (sklearn LR) and naive_avg for comparison. Includes a shuffled-label negative control (shuffled_lr_baseline_auc_mean, shuffled_foundation_auc_mean) to verify the AUC is signal-driven, not paired-design artifact. Latest 5-seed mean: foundation AUC 0.93 ± 0.03 vs shuffled 0.29; lr_baseline AUC 0.96 vs shuffled 0.37. signal_to_artifact_ratio = 1.49 — signal dominates artifact by 49%. Three-way gate: lr_baseline AUC ≥ 0.90, lr_baseline sens@99 ≥ 0.40, foundation AUC ≥ 0.85. |
commit 7d5225d + b9d0340 + d5c62ac + 25fea0c + cc24f1f + d1da12b + latest |
src/foundation/losses.py (NEW) |
SensAtSpecLoss (focal-modulated BCE for ultra-low VAF), BalancedCrossEntropy (Cui 2019 effective-number rebalancing), CalibrationLoss (Mukhoti 2020 differentiable ECE), focal_binary_cross_entropy (composable). |
commit 0ad0033 |
| Component | Validation surface | Why listed here |
|---|---|---|
src/foundation/ (MultiModalEncoder + PretrainHead + ContrastiveHead + FoundationDownstream) |
MultiModalDataGenerator synthetic 6-modality dataset. Pretraining (3-phase MMP + contrastive + joint) and downstream fine-tuning are closed-loop on synthetic features. |
Documents the architecture; no claim is made about real-cfDNA performance. |
src/methylation_gnn/ (GATv2 + ReconstructionDecoder + AnomalyHead) |
Synthetic per-fragment methylation arrays. Real cfDNA methylation data is not wired in (see src/methylation_gnn/CHANGES.md for the planned FinaleMe integration). |
Architecture-only reference. |
src/tissue_deconv/ (cfSort-style 4-layer MLP, 29 tissues) |
Synthetic tissue atlas. Real plasma methylation β-values are not used in training. | Architecture-only reference. |
src/clinical/, src/priming/ (clinical decision-curve, PK/PD, whitepaper) |
Synthetic data and threshold sweeps. | Decision-theoretic scaffolding. |
model/tokenizer.py (BPE subword tokenizer + neural motif model) |
Synthetic FASTA motif benchmark (100% vs 46.5% accuracy on synthetic data). | Architecture-only reference. |
The fusion_ablation.py AUC of 0.9886 (naive average of tumor-naive 5-channel fragmentomics + a synthetic mutation channel calibrated to AUC 0.92) is a clean what-if experiment that answers: "given a mutation-informed channel of this quality, does fusion help?". The +0.0143 fusion gain (paired t p<0.0001, DeLong p<0.0015 on every seed) is honest on its cohort. The mutation score is not a measurement from the same plasma samples as the fragmentomics score — pairing them on real plasma is the next phase.
- No real plasma cfDNA sample has ever been processed end-to-end through the foundation model. The pretraining → fine-tuning → evaluation loop uses synthetic features for the multi-modal data generator. The new
foundation-real-smokeCI job uses real TCGA-LUAD panel-LLR scores as one of two channels (the other channel is synthetic because no FinaleDB plasma is paired with the 20 TCGA-LUAD patients). This is the first time the foundation model has been evaluated against a real cfDNA-class signal; it is not yet a clinical validation. - No held-out clinical validation. All reported AUCs are in-sample 5-fold or pooled OOF on the same cohort the model was trained on.
- No clinical-grade operating-point thresholds. The decision-curve analyzer reports per-specificity operating tables but the recommended threshold τ is calibrated against the training cohort, not a screening cohort with prevalence ~0.4%.
git clone https://github.com/rollroyces/deepcatch.git
cd deepcatch
# Recommended: pip install -e . exposes CLI entry points (deepcatch-tumornaive, deepcatch-fusion, etc.)
pip install -e .
# Or the minimal install (just deps, no console scripts)
pip install -r requirements_py.txtMinimum dependencies:
pip install numpy scipy scikit-learn pandasWith deep learning (GNN, foundation model, tissue deconv):
pip install torch>=2.0.0 torch-geometricOptional — BAM/FASTQ processing:
pip install pysam statsmodelsDocker:
docker build -t deepcatch:latest .
docker run --rm -v $(pwd)/results:/app/results deepcatch:latestimport numpy as np
from src.fragmentomics import EnhancedFragmentomics
from src.fragmentomics.themis_features import (
MFRCalculator, FSICalculator, CAFFCalculator, FEMCalculator
)
from src.methylation_gnn import RegulatoryGraphBuilder, MethylationGNNPredictor
from src.tissue_deconv import DEConvIntegration
from src.priming.pharmacokinetics import PKModel, OptimalDosingSchedule
# ── Fragmentomics Basic ──
mfr = MFRCalculator()
fsi = FSICalculator()
caff = CAFFCalculator()
fem = FEMCalculator()
frag_basic = {
"mfr": mfr.compute(coverage, cpg_density),
"fsi": fsi.compute(fragment_lengths),
"caff": caff.compute(cnv_profile),
"fem": fem.compute(end_motif_counts),
}
# ── Enhanced Fragmentomics (DELFI + MFS + nucleosome + 5-mer) ──
ef = EnhancedFragmentomics()
frag_enhanced = ef.extract_all(
fragment_lengths=lengths,
fragments=fragments,
end_sequences=end_seqs,
tss_positions=tss_positions,
)
# → dict of ~70 scalar features
# ── GNN Methylation Network ──
gnn = MethylationGNNPredictor.load("checkpoints/gnn_pretrained.pt")
graph = RegulatoryGraphBuilder().build_graph(
sample_name="S001", methylation_data=meth_data
)
field_defect_score = gnn.predict_sample(
sample_name="S001", methylation_data=meth_data
)
# ── Tissue Deconvolution ──
deconv = DEConvIntegration(checkpoint="checkpoints/deconv.pt")
# Or train from scratch on synthetic mixtures:
# deconv.fit_synthetic(n_samples=2000)
tissue_fractions = deconv.predict_tissue_fractions(methylation_data)
tissue_features = deconv.extract_all(methylation_data, tissue_fractions)
# → dict of 24 scalar features
# ── CNV ──
cnv_features = {
"cnv_burden": np.mean(np.abs(cnv_log2_ratios)),
"cnv_entropy": scipy.stats.entropy(cnv_segment_lengths),
"arm_imbalance": max_arm_imbalance(cnv_profile),
}
# ── Serological ──
sero_features = {
"pg1": pg1_value, "pg2": pg2_value,
"g17": g17_value, "hp": hp_igg_value,
}
# ── Priming Agent PK/PD ──
pk = PKModel()
pk_result = pk.simulate(
agent="scFv", dose_mg=100, patient_weight_kg=70,
duration_hours=48,
)
dosing = OptimalDosingSchedule().compute(
agent="scFv", patient_data={"weight_kg": 70}
)from src.foundation import FoundationDownstream, FoundationConfig
# Assemble modalities dict (n_samples × dim for each key)
modalities = {
"frag_basic": np.array(frag_basic_array), # (N, 4)
"frag_enhanced": np.array(frag_enhanced_array), # (N, 44)
"cnv": np.array(cnv_array), # (N, 6)
"sero": np.array(sero_array), # (N, 4)
"gnn": np.array(gnn_scores), # (N, 1)
"tissue": np.array(tissue_array), # (N, 24)
}
# Use pre-trained checkpoint
fusion = FoundationDownstream(pretrained=True)
fusion.fit(modalities, labels)
proba = fusion.predict_proba(modalities) # shape (N, 2)
predictions = fusion.predict(modalities) # shape (N,)
# Or train from scratch (no pre-training needed)
fusion = FoundationDownstream(pretrained=False)
fusion.fit(modalities, labels, n_epochs=50, batch_size=32)
proba = fusion.predict_proba(modalities)from src.multimodal_fusion.advanced_fusion import (
CrossAttentionFusion, GCNTissueOfOrigin, EarlyLateFusion, TASK_PRIOR_MASKS,
)
# List of 1-D score arrays per modality (sklearn-style API).
scores = [mfr_scores, fsi_scores, caff_scores, fem_scores, cnv_scores]
labels = np.array([...])
# Biologically-informed prior: which cross-modal attention paths to allow.
# 'cancer_detection' (default) or 'tissue_of_origin' or None (no prior).
fusion = CrossAttentionFusion(
n_modalities=5, prior="cancer_detection",
n_epochs=200, lr=1e-3, seed=0,
)
fusion.fit(scores, labels)
proba = fusion.predict_proba(scores) # 1-D cancer probability for binaryWhat's new vs the previous version: the previous CrossAttentionFusion
shipped with random-initialized np.random.randn(...) attention matrices
that were never trained — the fit() method only ran a sklearn
LogisticRegression on the concatenated 1-D scores. The current rewrite
is a real PyTorch implementation (_FusionEncoder + _GatedCrossAttention)
that trains end-to-end with AdamW + early stopping on a stratified val
split. The prior argument encodes which cross-modal attention paths
the model is allowed to learn (serology ↔ CNV is blocked in
cancer_detection because no known biology links them; tissue ↔ all
is unblocked because tissue composition informs every channel).
For multi-modal feature fusion (full 2-D modality features per channel):
from src.multimodal_fusion.advanced_fusion import EarlyLateFusion
modality_features = [mfr_feats, fsi_feats, caff_feats] # list of (N, d_i)
elf = EarlyLateFusion(n_modalities=3, hidden_dim=32)
elf.fit(modality_features, labels)
proba = elf.predict_proba(modality_features)from src.clinical import ClinicalReportGenerator
crg = ClinicalReportGenerator(cet_df, fusion_result)
print(crg.generate_briefing()) # One-paragraph summary
crg.export_json("report.json") # Machine-readable export
with open("report.html", "w") as f:
f.write(crg.generate_html_report()) # Full HTML reportbash RUN_ALL.sh # Full pipeline
bash RUN_ALL.sh --quick # 2-minute smoke test
# Real-data foundation smoke (wired into CI as foundation-real-smoke).
# Trains FoundationDownstream on real TCGA-LUAD panel-LLR scores
# (when validation/tcga/tcga_cache/ is present) + synthetic fragmentomics
# channel; gates on AUC >= 0.70 and sens@99 >= 0.10.
python scripts/foundation_real_smoke.py \
--out results/foundation_real_smoke.json
cat results/foundation_real_smoke.jsonA standalone biomedical review surfaced six classes of issues that this
PR closes. Each fix ships with regression tests in
test/test_biomedical_review_fixes.py (23 new tests, all passing).
| # | Finding | Fix |
|---|---|---|
| 1 | CrossAttentionFusion / GCNTissueOfOrigin / EarlyLateFusion used random-initialized numpy attention matrices that were never trained. The fit() method only ran sklearn LogisticRegression on concatenated 1-D scores — the attention output was discarded. |
Full PyTorch rewrite (src/multimodal_fusion/advanced_fusion.py): scaled dot-product attention + GATv2-style message passing + MLP classifier. Trains via AdamW + early stopping on a stratified val split. NaN/Inf guards on training and val losses. Biologically-informed prior masks via TASK_PRIOR_MASKS (cancer_detection / tissue_of_origin / None). |
| 2 | EnhancedFragmentomics._pca_reduce was mis-named — the algorithm was a top-N per-motif deviation vector against the uniform-background null, not a principal-component projection. |
Renamed to _top_motif_deviations with honest docstring; output keys kept as fem_5mer_pc* for backwards compatibility. |
| 3 | Simpson diversity per length bin used the universe size (1024) as the denominator for every bin. Short fragments <150 bp only exercise a subset of all 5-mers, so the universe denominator inflated MDS for sparse bins. | Normalize by the effective alphabet per bin (n_eff = (counts > 0).sum()). |
| 4 | expected_nucleosome_pattern had hard-coded 195 bp period, 150 bp dip half-width, 0.5 dip amplitude, 0.3 sinusoidal amplitude. |
Class-level defaults documented to Snyder 2016 / Jiang 2020; per-instance overrides for non-canonical cfDNA sources (yeast ~165 bp, mouse ES ~190 bp). |
| 5 | tss_coverage_profile aggregated fragments across chromosomes without matching TSS to chromosome. A fragment on chr5 was counted against a chr1 TSS. |
Now requires (chrom, pos) TSS tuples and groups fragments by chromosome before aggregating. Legacy bare-int path is preserved with a documented caveat. |
| 6 | FoundationDownstream.fit used torch.randperm for the val split, which could produce val sets with no positives on imbalanced cfDNA cohorts (NaN cross-entropy). |
Stratified split with at-least-one-of-each-class guarantee. NaN/Inf training-loss guard aborts gracefully after N consecutive NaN; NaN/Inf val-loss is skipped (no patience increment, no best_state overwrite). |
| 7 | extract_all(methylation_data=...) was declared but never read. |
Backfills per-fragment methylated from the standalone array when fragments lacks it. Still returns zero-fallback when no methylation source is provided. |
| 8 | CAFFCalculator required caller-provided per_arm_coverage dict with no helper to derive it from fragments. |
New CAFFCalculator.from_fragments() classmethod: per-arm coverage-per-Mb normalized so the median arm reads 1.0. Handles UCSC chr1 and Ensembl 1 chrom naming. |
| 9 | No real-data smoke test for the foundation model. | New scripts/foundation_real_smoke.py: trains FoundationDownstream on real TCGA-LUAD panel-LLR scores + real per-patient mutation-derived features (20 patients, both channels from real TCGA MAFs). Two foundation variants weighted-averaged 70/30 (B: frozen encoder + LR head; A: 3-ensemble tiny transformer). Plus lr_baseline and naive_avg. Shuffled-label negative control in the same script confirms the AUC is signal-driven (signal_to_artifact_ratio = 1.49 — signal exceeds artifact by 49%). Latest 5-seed mean: foundation AUC 0.93 ± 0.03 vs shuffled 0.29; lr_baseline AUC 0.96 vs shuffled 0.37. Three-way gate: lr_baseline AUC ≥ 0.90 AND sens@99 ≥ 0.40 AND foundation AUC ≥ 0.85. Wired into .github/workflows/validate.yml as the foundation-real-smoke job. |
New loss functions (src/foundation/losses.py):
SensAtSpecLoss(alpha_pos=20, gamma=2)— focal-modulated binary cross-entropy for ultra-low VAF training. Down-weights easy negatives so the rare-positive tail dominates the gradient at the high-specificity operating point.BalancedCrossEntropy(labels, beta=0.999)— inverse-frequency weighted multi-class CE per Cui 2019 effective-number rebalancing.CalibrationLoss(n_bins=15)— differentiable ECE surrogate per Mukhoti 2020; replaces the closed-form (non-differentiable) ECE.focal_binary_cross_entropy— composable focal-BCE for users who want a different reduction.
New tests (23, in test/test_biomedical_review_fixes.py):
CAFFCalculator.from_fragments (per-Mb normalization, chrom naming,
unrecognized chroms, empty input, round-trip through compute()),
RefinedEndMotifs._top_motif_deviations (key names unchanged, empty
input), motif diversity (effective-alphabet Simpson), nucleosome
parameters (overrides flow through to mean-normalized pattern), TSS
per-chromosome matching, extract_all methylation backfill,
SensAtSpecLoss (finite loss, alpha_pos effect), BalancedCE
(inverse-frequency weight), CalibrationLoss (zero for perfect
calibration), FoundationDownstream stratified split + NaN guard,
CrossAttentionFusion / EarlyLateFusion / GCNTissueOfOrigin PyTorch
rewrite (binary 1-D output, no NaN, correct shapes).
Purpose: cfDNA fragmentation pattern analysis implementing DELFI, MDS, and THEMIS-equivalent feature frameworks.
| Class / Function | Description |
|---|---|
MFRCalculator |
Methylated Fragment Ratio via CpG density scoring |
FSICalculator |
Fragment Size Index: short/long ratio + GMM sub-nucleosomal fraction |
CAFFCalculator |
Chromosomal Aneuploidy: CNA burden scoring from whole-genome bins |
FEMCalculator |
Fragment End Motif: 4-mer MDS + motif embeddings (Jiang 2020) |
FragmentLengthGMM |
4-component Gaussian Mixture Model (sub-/mono-/di-/tri-nucleosomal) |
DELFI_style_normalization |
LOESS GC-bias correction + mappability filter |
compute_MDS |
Motif Diversity Score from 4/5-mer counts |
EnhancedFragmentomics |
Unified extractor: DELFI + MFS + nucleosome footprint + refined 5-mer |
extract_4mer_end_motifs |
4-mer extraction from BAM files |
extract_end_motifs_from_fastq |
4-mer extraction from FASTQ |
Input: BAM/FASTQ files, or fragment length arrays + end sequences
Output: Scalar features (4–80+), GMM component statistics, MDS scores
Tests: 42 (test_enhanced_features.py)
Purpose: Detect pre-cancer epigenetic field defects via GATv2 graph attention on methylation regulatory graphs.
| Class / Function | Description |
|---|---|
RegulatoryGraphBuilder |
Constructs heterogeneous graphs from methylation + Hi-C contacts |
MethylationGNN |
GATv2 model with reconstruction decoder + anomaly head |
GNNTrainer |
3-phase training: masked pre-training → joint → fine-tuning |
GNNInference / MethylationGNNPredictor |
Lightweight inference producing field_defect_score |
ReferenceDataCatalog |
Downloads UCSC CpG islands, ENCODE Hi-C, GENCODE promoters, FANTOM5 enhancers |
MethylationBranchAdapter |
Drop-in adapter for CrossAttentionFusion compatibility |
Input: cfDNA methylation beta values + reference Hi-C/chromatin data
Output: Graph-level field_defect_score (scalar) per sample
Tests: 46 (test_integration.py)
Purpose: Predict tissue-of-origin cfDNA fractions from methylation data using a cfSort-style DNN.
| Class / Function | Description |
|---|---|
TissueAtlas |
29-tissue reference methylation profile store |
TissueDeconvolutionModel |
Lightweight DNN (~500K params): [256, 128, 64] + BN + ReLU + Dropout |
TissueDeconvolutionEnsemble |
3-model ensemble with seed diversity |
TissueDeconvTrainer |
KL divergence + L1 sparsity + entropy regularization on synthetic mixtures |
TissueDeconvolutionFeatures |
Extracts 24-D feature vector from tissue fractions |
DEConvIntegration |
Full integration class compatible with existing pipeline |
Input: cfDNA methylation beta values (or synthetic atlas for training)
Output: Per-tissue fraction vector + 24-D feature vector
Tests: 47 (test_integration.py)
Purpose: Self-supervised multi-modal Transformer pre-training for cfDNA. Drop-in replacement for CrossAttentionFusion.
| Class / Function | Description |
|---|---|
FoundationConfig |
Hyperparameter dataclass (embed_dim, n_heads, n_layers, etc.) |
MultiModalEncoder |
4-layer TransformerEncoder with per-modality linear projections |
PretrainHead |
Masked modality prediction head |
ContrastiveHead |
Cross-modal contrastive loss (InfoNCE) |
FoundationPretrainer |
Self-supervised pre-training orchestrator |
FoundationDownstream |
Downstream fine-tuning with CrossAttentionFusion-compatible API |
FoundationCompatibilityWrapper |
Wrapper for seamless replacement of CrossAttentionFusion |
MultiModalDataGenerator |
Synthetic multi-modal data generator for pre-training |
Pre-training tasks:
- Masked modality prediction — reconstruct masked modalities from context
- Cross-modal contrastive — InfoNCE between modalities of same sample
API compatibility:
# CrossAttentionFusion (old)
fusion = CrossAttentionFusion(n_modalities=6)
fusion.fit(scores, labels) # scores: list of 1-D arrays
proba = fusion.predict_proba(scores)
# FoundationDownstream (new — drop-in)
fusion = FoundationDownstream(pretrained=True)
fusion.fit(modalities, labels) # modalities: dict of (N, D) arrays
proba = fusion.predict_proba(modalities) # shape (N, 2)Input: Dict of modality arrays {name: np.ndarray (N, D)}
Output: Joint embeddings (N, n_modalities, embed_dim); classification probabilities (N, 2)
Tests: 43 (test_integration.py)
Purpose: Simulate PK/PD of cfDNA priming agents (Amplifyer Bio) and their effect on ctDNA detection.
| Class / Function | Description |
|---|---|
PKModel |
1-compartment PK model with first-order elimination |
OptimalDosingSchedule |
Computes optimal dosing for 5 agent types |
PrimingConfig |
Dataclass with literature-based PK parameters |
Agents: scFv, liposome, nanoparticle, polymeric micelle, dendrimer Input: Agent type, dose, patient weight, liver function Output: Concentration-time profiles, ctDNA boost factor, optimal dosing schedule Reference: Martin-Alonso et al. (2024) Science
| Class / Function | Description |
|---|---|
CrossAttentionFusion |
Real PyTorch scaled dot-product cross-attention between modality embeddings. Trained end-to-end with AdamW + early stopping. Biologically-informed prior masks via TASK_PRIOR_MASKS (cancer_detection / tissue_of_origin / None). Replaces the previous sklearn-LR-with-random-attention placeholder. |
GCNTissueOfOrigin |
GATv2-style heterogeneous graph for TOO prediction at low sequencing depth. Replaces the previous sklearn-LR-over-correlation-pool placeholder. |
EarlyLateFusion |
Concatenated-feature MLP with per-modality standardization. Replaces the previous LogisticRegression(C=0.5, class_weight='balanced'). |
| Class / Function | Description |
|---|---|
SerologicalFusion |
Fuses PG-I, PG-II, G-17, H. pylori with cfDNA predictions |
IntegrativeScoringSystem |
Unified risk scoring across all modalities |
ClinicalReportGenerator |
Generates clinician-friendly HTML/JSON reports |
NestedCETValidator |
Nested cross-validation for unbiased motif-based CET evaluation |
FrequencyDataset |
Loads pre-computed 4-mer frequency vectors (Jiang lab format) |
Bayesian Kalman filter (BSSLM) for longitudinal evidence accumulation across quarterly blood draws. Tracks patient risk trajectory over time rather than relying on single-timepoint decisions.
MAML-based few-shot adaptation for cancer subtype detection.
Multi-confounder realistic cohort generation (CHIP, variable shedding, trinucleotide errors, GC bias, batch effects, inflammation) for development and testing.
# All tests
python -m pytest src/ test/ -v
# Or with unittest
python -m unittest discover -s src -p "test_*.py"
# Per-module
python src/foundation/test_integration.py # 43 tests
python src/methylation_gnn/test_integration.py # 54 tests (requires torch_geometric)
python src/tissue_deconv/test_integration.py # 54 tests
python src/fragmentomics/test_enhanced_features.py # 47 tests
# Biomedical-review regression tests (NEW)
python -m pytest test/test_biomedical_review_fixes.py -v
# Real-data foundation smoke (NEW; same script CI runs)
python scripts/foundation_real_smoke.py --out results/foundation_real_smoke.json
python -c "import json; d=json.load(open('results/foundation_real_smoke.json')); print(f\"foundation AUC={d['foundation_auc_mean']:.3f} ± {d['foundation_auc_std']:.3f}, sens@99={d['foundation_sens_at_99_mean']:.3f}, gate_pass={d['gate_pass']}\")"
# Quick smoke test
python -c "from src.foundation import FoundationConfig; print('OK')"| Module | Tests | Status |
|---|---|---|
| Enhanced Fragmentomics (+ THEMIS) | 42 | ✅ All passing |
| GNN Methylation | 46 | ✅ All passing (requires torch_geometric) |
| Tissue Deconvolution | 47 | ✅ All passing |
| Foundation Model | 43 | ✅ All passing |
| Priming Agents | 50 | ✅ All passing |
src/ subtotal (full repo discovery) |
228 | ✅ |
Standalone test/ (fusion_ablation, tumor_naive_adapter, decision_curve, …) |
25 | ✅ All passing |
test/test_biomedical_review_fixes.py (NEW) |
23 | ✅ All passing |
test/test_foundation_smoke.py (NEW) |
8 | ✅ All passing |
Combined test/ + src/foundation/test_integration.py |
145 collected → 143 pass + 2 skip | ✅ |
The 373/375 badge in the README header counts:
- 228
src/tests (all passing when their respective torch_geometric / torch extras are installed) - 145 collected by
pytest test/ src/foundation/test_integration.py→ 143 pass + 2 skip (the 2 skip are pre-existing torch_geometric / torch CUDA guards).
The src/methylation_gnn/test_integration.py and other optional-deps modules are not in the 145 figure because they fail at import time on a plain numpy/scipy/sklearn install — they run in CI when the dl-tests job installs torch_geometric. The test_foundation_smoke.py tests are slow (~2 min each, they invoke the real smoke script) and are excluded from the default fast test run.
stateDiagram-v2
[*] --> Capture
state Capture {
[*] --> ExtractModalities
ExtractModalities --> FuseTransformer
FuseTransformer --> JointEmbedding
JointEmbedding --> [*]
}
Capture --> Accumulate
note right of Capture
Single cfDNA draw.
7 modalities → joint embedding.
end note
state Accumulate {
[*] --> KalmanUpdate
KalmanUpdate --> PosteriorCheck
PosteriorCheck --> HasMoreDraws: yes
HasMoreDraws --> KalmanUpdate: next quarterly draw
PosteriorCheck --> EmitPosterior: no
EmitPosterior --> [*]
}
Accumulate --> Triage
state Triage {
[*] --> CompareThreshold
CompareThreshold --> Confirm: p_cancer > τ
CompareThreshold --> Clear: p_cancer ≤ τ
Confirm --> [*]
Clear --> [*]
}
Triage --> [*]
note left of Accumulate
Each new draw updates the
posterior over p_cancer.
Evidence accumulates below
single-draw detection floor.
end note
Seven independent modalities extract signal from the same cfDNA sample. Each produces a scalar risk score vector. The foundation model fuses these into a joint embedding via per-modality linear projections → 4-layer Transformer encoder.
Longitudinal tracking via Bayesian Kalman filter (BSSLM). The joint embedding from Stage 1 is tracked across quarterly blood draws, accumulating evidence over time. This is designed to detect cancers whose ctDNA signal is below single-timepoint detection thresholds at early stages.
The accumulated Bayesian posterior probability p_cancer is compared to a calibrated threshold τ. Samples above the threshold trigger confirmatory testing; samples below are cleared until the next quarterly draw.
| Modality | Required Data | Public Source |
|---|---|---|
| Fragmentomics Basic | Fragment length arrays, end motif counts | N/A (extracted from BAM/FASTQ) |
| Enhanced Fragmentomics | Fragment lengths + genomic coordinates + end sequences | Same as above |
| CNV | Log2 ratio profiles or BAM | Same as above |
| Serological | PG-I, PG-II, G-17, H. pylori IgG | Clinical lab |
| GNN Methylation | cfDNA methylation beta values | TCGA, GEO |
| Tissue Deconvolution | cfDNA methylation beta values | TCGA, cfSort atlas |
| Priming Agents | Agent PK parameters | Literature |
| Resource | URL |
|---|---|
| ENCODE Hi-C | https://www.encodeproject.org/ |
| UCSC CpG Islands | http://hgdownload.soe.ucsc.edu/goldenPath/hg38/database/ |
| GENCODE promoters | https://www.gencodegenes.org/human/ |
| FANTOM5 enhancers | https://fantom.gsc.riken.jp/5/ |
| TCGA methylation | https://portal.gdc.cancer.gov/ |
| cfSort atlas | https://github.com/stephenrcraig/cfSort |
All modules support fully synthetic data for development and testing. Use MultiModalDataGenerator (foundation), TissueAtlas (deconv with built-in synthetic profiles), and ReferenceDataCatalog (GNN with random initialization) to run the full pipeline without any external reference data.
deepcatch/
├── README.md # This file
├── LICENSE # MIT
├── CITATION.cff # Academic citation metadata
├── requirements_py.txt # Python dependencies
├── RUN_ALL.sh # One-command validation
├── Dockerfile
│
├── src/
│ ├── fragmentomics/ # FragmentoSign: DELFI, MDS, GMM, LOESS, enhanced
│ ├── methylation_gnn/ # GATv2 graph attention for field defect detection
│ ├── tissue_deconv/ # cfSort-style DNN for tissue-of-origin
│ ├── foundation/ # Self-supervised Transformer foundation model
│ ├── priming/ # PK/PD priming agent simulation
│ ├── multimodal_fusion/ # CrossAttentionFusion, GCN, EarlyLate
│ ├── clinical/ # Serological fusion, clinical reports, CET validation
│ ├── longitudinal/ # Bayesian Kalman filter (Stage 2)
│ ├── ensemble/ # MAML meta-learning
│ ├── synthetic_data/ # Realistic cohort generation
│ ├── variant_calling/ # Bayesian + contrastive DL
│ └── preprocessing/ # CHIP filter
│
├── validation/ # Statistical validation suite
│ ├── py/ # Python validation modules (11)
│ ├── tcga/ # TCGA data loaders + validators
│ └── *.py # 10 bioinformatics-grade modules
│
├── test/ # Additional test suites
├── results/ # Output reports + figures
├── paper/ # LaTeX manuscript
├── docs/ # User guide
└── review/ # Peer review history
- Create module directory under
src/your_modality/ - Implement feature extractor with
extract_all()orpredict_sample()entry point - Define config with dataclass
YourModalityConfig - Add integration class that wraps your module for the fusion API
- Write tests — aim for ≥20 tests covering config, forward pass, edge cases, and integration
- Update
MODALITY_DIMSinsrc/foundation/config.py
- Type hints on all public APIs
- NumPy docstring style with Parameters/Returns sections
- Tests use pytest or unittest; run them before submitting
Open an issue first to discuss scope. Target main branch. PRs must pass all existing tests.
Preliminary validation on 129 real plasma samples from Jiang lab (CUHK), using 4-mer end-motif frequency vectors:
| Metric | Value |
|---|---|
| Samples (HCC vs Control) | 72 (34 HCC, 38 Control) |
5-fold CV AUC (nested selection, run_jiang_analysis.py) |
0.9845 |
| Bonferroni-significant motifs | 108 / 256 |
| Biological pattern | CG-rich depletion, AT-rich enrichment |
Caveats: HCC only (other types n≤17), processed frequency data (not raw BAM), single centre. Not a clinical assay. AUC is from nested cross-validation (motif selection inside folds); the raw data file is not redistributed in the repo (CUHK terms) — provision via data/deepcatch_data.xlsx or DEEPCATCH_DATA_DIR.
real_tcga_validation.py uses real TCGA tumor mutations (with real read counts) as ground truth, then simulates plasma cfDNA by Poisson sampling at each tumor fraction. It is a spike-in/dilution benchmark, not a clinical plasma validation. Metrics are AUC/PR-AUC plus sensitivity at fixed 95%/99% specificity — no threshold optimization on test data. Data is fetched from the GDC open-access API (per-aliquot masked MAFs, cached in validation/tcga/tcga_cache/); the synthetic fallback dataset is deliberately refused. Latest run: 20 LUAD patients, 5,738 mutations, 5 seeds (mean across seeds).
flowchart LR
TCGA[("TCGA GDC<br/>per-aliquot masked MAFs")] -->|fetch| Cache["validation/tcga/<br/>tcga_cache/"]
Cache --> Muts["5,738 real mutations<br/>20 LUAD patients"]
Muts --> Sim["Simulate plasma cfDNA<br/>(Poisson sampling)"]
Sim -->|"5 fractions: 10% → 0.1%"| F1["10%"]
Sim --> F2["5%"]
Sim --> F3["1%"]
Sim --> F4["0.5%"]
Sim --> F5["0.1%<br/>ultra-early"]
F1 & F2 & F3 & F4 & F5 --> Eval["3 scoring methods:<br/>LLR / Fisher / Strand"]
Eval -->|"fixed spec<br/>no threshold opt"| Metrics["AUC, Sens@95%,<br/>Sens@99%, paired win-rate"]
Metrics -->|"5 seeds × 5 folds"| Report["results/<br/>real_tcga_validation.json"]
classDef data fill:#fff8c5,stroke:#bf8700
classDef stage fill:#e1f5ff,stroke:#0969da
classDef out fill:#dafbe1,stroke:#1a7f37
class TCGA,Cache,Muts data
class Sim,F1,F2,F3,F4,F5,Eval stage
class Metrics,Report out
Per-position detection (single-locus classification — information-limited at ultra-low ctDNA):
| ctDNA fraction | Variant caller AUC | VC Sens @ 95% spec |
|---|---|---|
| 10% | 1.000 | 1.000 |
| 5% | 0.9995 | 0.998 |
| 1% | 0.959 | 0.850 |
| 0.5% | 0.884 | 0.633 |
| 0.1% (ultra-early regime) | 0.642 | 0.183 |
Panel-based detection (--skip-panel to disable, --clean-panel for a designed-panel simulation) — MRD-style per-sample aggregation over the tracking panel. Three scoring methods: LLR sum (standard), Fisher sum (-log₁₀ Poisson p-value, CAPP-Seq/Neman 2014), and Strand-concordance-weighted Fisher. The simulation now models context-dependent sequencing errors (CpG ~10×, homopolymer ~5×, clean baseline), strand-asymmetric error reads (true variants are biallelic across fwd/rev; errors are single-strand), and optional clean-panel design (avoid high-error genomic regions):
| ctDNA fraction | LLR AUC | Fisher AUC | Strand AUC | Sens @ 95% spec | Paired cancer>control |
|---|---|---|---|---|---|
| 10% | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| 5% | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| 1% | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| 0.5% | 0.9995 | 0.997 | 0.996 | 0.990 | 1.000 |
| 0.1% | 0.921 | 0.834 | 0.831 | 0.770 | 1.000 |
With a well-designed panel (--clean-panel, avoiding CpG/homopolymer loci): LLR 0.922, Fisher 0.849, Strand 0.836 at 0.1% ctDNA. Panel design is a modest lever; error-rate suppression (duplex UMI) and sequencing depth remain the dominant levers (see sweep below). The strand score uses a Z-score (Normal) approximation to the binomial test — corrected from the old 2×min/max formula which erroneously penalized low-read-count positions.
Restricting the 5,738-mutation panel to the top-K highest-CADD mutations (PHRED-scaled impact from Kircher et al. 2014, CC BY-NC-SA 4.0) lifts ultra-low ctDNA detection without AUC penalty. Three validated findings:
Per-subgroup lift at 0.1% ctDNA, 5-seed × 5-fold OOF, validated on the larger GDC TCGA-LUAD cohort (150-patient subsample of 382):
| Subgroup | n | Uniform Sens@99% | CADD Top-K=20 | Δ |
|---|---|---|---|---|
| TP53_mutant | 73 | 0.228 | 0.689 | +46pp |
| TP53_wildtype | 77 | 0.242 | 0.506 | +26pp |
| KRAS_mutant | 47 | 0.348 | 0.688 | +34pp |
| KRAS_wildtype | 103 | 0.275 | 0.663 | +39pp |
| STK11_wildtype | 132 | 0.293 | 0.662 | +37pp |
| High mut burden | 75 | 0.271 | 0.773 | +50pp |
| Whole-cohort | 150 | 0.278 | 0.627 | +35pp |
Top-K=20 was used (not the originally-reported Top-K=200) because only 1 patient in the larger GDC bulk-WXS cohort has ≥200 CADD matches (per-patient median is 21). Top-K=20 is the new validated winner — the lift is larger on the bigger cohort than the original 20-patient estimate suggested.
Caveat documented honestly: match rate dropped from 86% (20-patient curated LUAD driver set) to 8.9% (382-patient bulk-WXS includes passengers not in gnomAD r3.0). This is a structural difference between cohorts, not a methodology change. The low-burden subgroup shows a -16pp inversion under CADD selection (flagged as a real finding).
CADD Top-K pipeline (data flow):
flowchart LR
Panel["Full panel<br/>5,738 TCGA-LUAD<br/>mutations"]
Score["CADD PHRED score<br/>per mutation<br/>(Kircher et al. 2014)"]
Rank["Per-patient rank by CADD"]
TopK["Top-K = 20<br/>highest-CADD<br/>per patient"]
Eval["Simulate 0.1% ctDNA<br/>5-seed x 5-fold OOF"]
Out["Per-subgroup<br/>Sens at 99% spec lift"]
Whole["Whole-cohort n=150<br/>0.278 to 0.627<br/>+35 pp"]
TP53["TP53_mutant n=73<br/>0.228 to 0.689<br/>+46 pp"]
KRAS["KRAS_wildtype n=103<br/>0.275 to 0.663<br/>+39 pp"]
High["High burden n=75<br/>0.271 to 0.773<br/>+50 pp"]
Match["Match rate<br/>86% on 20-patient curated<br/>8.9% on 382-patient bulk-WXS"]
Panel --> Score --> Rank --> TopK --> Eval --> Out
Out --> Whole
Out --> TP53
Out --> KRAS
Out --> High
TopK -. caveat .-> Match
classDef data fill:#fff8c5,stroke:#bf8700
classDef stage fill:#e1f5ff,stroke:#0969da
classDef out fill:#dafbe1,stroke:#1a7f37
classDef caveat fill:#ffebe9,stroke:#cf222e
class Panel,Score,Rank,TopK data
class Eval stage
class Whole,TP53,KRAS,High,Out out
class Match caveat
Honest negative results preserved:
- Continuous per-cancer weighting (CADD-style linear / sigmoid weights on channels) — regresses vs hard top-K (subagent commit
eb1529e) - CADD + AlphaMissense multiplicative weight — underperforms CADD alone at all ctDNA fractions (commit
29374a6) - Driver-only panel (TP53 + KRAS + EGFR + …) — only 31 mutations; Sens@99% = 0.19 (catastrophic)
- Proxy (constant 0.55) AlphaMissense weighting hid the real AM signal — the previous run reported every AM-weighted method at AUC ≈ 0.917 (proxy indistinguishable from uniform 0.921). With REAL AlphaMissense scores (key-construction bug fixed in commit
ec16e0d), AM Top-K=200 lifts AUC from 0.9210 → 0.9775 at 0.1% ctDNA (+0.057, not the proxy's 0) - AM ≈ CADD at K=200 (0.9775 vs 0.9785 on the same 20-patient cohort, within 1 std) — once K is large enough, the scoring function barely matters; panel size dominates. See
docs/ALPHAMISSENSE_REAL_LLR.mdfor the bug-fix details and per-patient coverage.
Reproduce:
# Original 20-patient per-subgroup
python scripts/cadd_per_subgroup_llr.py --n-patients 20 --top-k 200
# Larger 382-patient GDC validation
python scripts/cadd_per_subgroup_llr_gdc_validation.py --top-k 20
# Single-shot Top-K=20 whole-cohort
python scripts/cadd_weighted_llr.py --top-k 20Docs: docs/CADD_WEIGHTED_LLR.md, docs/CADD_PER_SUBGROUP_LLR.md, docs/CADD_GDC_VALIDATION.md, docs/CADD_ALPHAMISSENSE_COMBINED.md, docs/CADD_FLARE_VALIDATION.md, docs/ALPHAMISSENSE_REAL_LLR.md.
The CADD Top-K=20 finding above is the winner on the 150-patient GDC bulk-WXS cohort. On the smaller 20-patient TCGA-LUAD curated driver set (the original CADD validation cohort), an apples-to-apples comparison with the SAME 5,738-mutation panel and SAME 5 seeds shows a slightly different ranking: CADD Top-K=200 wins narrowly over AM Top-K=200 at 0.1% ctDNA, with both dramatically beating the uniform baseline. Both findings are positive — the K that maximises AUC is just cohort-dependent (more matches per patient on the curated set → K=200 is fine; fewer matches on bulk-WXS → K=20 is the binding choice).
Apples-to-apples headline (0.1% ctDNA, 5 seeds × 20 patients, panel of 5,738 mutations, full CADD + AM scoring pipelines re-run on identical cohort, results from results/cadd_vs_alphamissense_topk_20patient.json):
| Method | 0.1% ctDNA AUC | Δ vs uniform | Notes |
|---|---|---|---|
| Uniform (no weighting) | 0.9210 | ref | reproduces the documented 0.921 ± 0.0188 exactly |
| CADD Top-K=20 (gnomAD-matched) | 0.9345 | +0.014 | narrow lift at K=20; cohort has fewer per-patient CADD matches than bulk-WXS |
| CADD Top-K=200 | 0.9785 | +0.058 | best on the 20-patient cohort |
| AM Top-K=20 (REAL scores) | 0.9135 | −0.008 | matches ≈ CADD's K=20 (within 1 std) |
| AM Top-K=200 (REAL scores) | 0.9775 | +0.057 | statistically tied with CADD K=200 (within 1 std) |
| AM Top-K=500 (REAL scores) | 0.9260 | +0.005 | regresses past K=200 — too many loci dilute the signal |
Bug fix that unlocked the AM result (commit ec16e0d, full writeup in docs/ALPHAMISSENSE_REAL_LLR.md): the previous AlphaMissense-weighted LLR run (commit 5cd6c3e) was building AM lookup keys from the nucleotide ref/alt columns of the MAFs, but AlphaMissense is keyed by amino-acid ref_aa/alt_aa at protein coordinates (P01116:12:G:D = KRAS G12D). The bug silently dropped the per-mutation match rate to ≈ 0% and substituted the per-variant-class proxy (0.55). With the fix, AM match rate on the 20-patient missense cohort jumped from ≈ 0% → 97.23% (15,462 / 15,903 missense SNVs hit the 71.7M-row AM TSV via a 69,575-key pickle built by streaming the full TSV once in 47 s). The proxy was indistinguishable from uniform at AUC ≈ 0.917; the real scores reveal the +0.057 lift at K=200.
Honest framing of the two cohorts: the 150-patient GDC bulk-WXS result (CADD Top-K=20 wins, +35pp whole-cohort Sens@99%) and the 20-patient curated-driver result (CADD Top-K=200 wins on AUC, AM Top-K=200 statistically tied) are both real and both positive — they just optimise different operating points. The right K is whichever is ≤ the per-patient median CADD match count; the right scoring function (CADD vs AM) barely matters once K is large enough (the K=200 numbers are within 0.001 of each other).
flowchart LR
Panel["Full panel<br/>5,738 TCGA-LUAD<br/>mutations"]
Uniform["Uniform<br/>no weighting<br/>AUC = 0.9210"]
CADD["CADD Top-K=200<br/>PHRED-scaled<br/>(Kircher 2014)<br/>AUC = 0.9785<br/>+0.0575"]
AM["AM Top-K=200<br/>REAL AlphaMissense<br/>(Cheng 2023)<br/>AUC = 0.9775<br/>+0.0565"]
Panel --> Uniform
Panel --> CADD
Panel --> AM
Note["K=200 ≈ 5,738 × 0.04<br/>selection of the<br/>top ~4% of loci"]
CADD -. comparison .- Note
AM -. tied within 1 std .- CADD
classDef base fill:#fff8c5,stroke:#bf8700
classDef win fill:#dafbe1,stroke:#1a7f37,color:#116329
classDef meta fill:#f6f8fa,stroke:#57606a
class Panel base
class Uniform base
class CADD,AM win
class Note meta
Reproduce:
# Apples-to-apples CADD vs REAL AlphaMissense (20-patient cohort)
env -u PYTHONPATH ./.venv/bin/python scripts/cadd_vs_alphamissense_topk.py \
--n-patients 20 --top-k-values 20,200,500 \
--output results/cadd_vs_alphamissense_topk_20patient.jsonUltra-early assay sweep (0.1% ctDNA; --skip-sweep to disable) — panel detection vs background error rate × depth. This is the assay-design guidance: duplex-UMI consensus (~1e-4) or ~50k× depth each bring sens@95% to 1.000 at 0.1% ctDNA:
| Background error rate | Depth | Panel AUC | Sens @ 95% spec |
|---|---|---|---|
| 2e-3 (raw reads) | 5,000× | 0.935 | 0.770 |
| 2e-3 | 50,000× | 0.998 | 1.000 |
| 1e-3 | 5,000× | 0.965 | 0.910 |
| 1e-3 | 50,000× | 0.9995 | 1.000 |
| 1e-4 (duplex UMI) | 5,000× | 0.998 | 1.000 |
| 1e-4 | 50,000× | 1.000 | 1.000 |
| 1e-5 | any | 1.000 | 1.000 |
The remaining gap to production is real plasma cfDNA sequencing — see docs/PRODUCTION_ROADMAP.md. The longitudinal CET stage (Stage 2) is intended to extend this below 0.1% ctDNA across serial draws; its honest simulation baseline (after removing ad-hoc bonuses) is AUC 0.49, sens 2.5% @ 97% spec (results/README.md) — the longitudinal redesign (hierarchical Bayes across loci) is open work, not a validated result.
A thin adapter lets DeepCatch consume the pre-computed fragmentomic
artifacts of cfdna-fragmentomics-pipeline,
adding tumor-naive detection — cancer classification from cfDNA
without any prior knowledge of tumor mutations.
The complementary role is the point. DeepCatch's mutation-informed detection needs to know the tumor's mutations in advance (panel design, TCGA-driven simulation). The pipeline doesn't — it detects cancer from raw fragmentation patterns. Combining them gives DeepCatch a signal channel reviewers will ask about.
Channel assembled (per sample, from cfdna-fragmentomics-pipeline/data/features/):
| Source file | Channel | Dim |
|---|---|---|
{s}.delfi_5mb_ratio.npy |
5Mb DELFI ratio | 631 |
{s}.delfi_5mb_coverage.npy |
5Mb CNA coverage | 631 |
{s}.delfi_100kb_ratio.npy |
100kb DELFI ratio | 30,894 |
{s}.delfi_100kb_counts.npy |
100kb CNA (median-normalized) | 30,894 |
{s}.fsd.json |
FSD size histogram (5bp bins) | 196 |
End-to-end result (5-seed CV, harmonized, PCA n=200, 627 cross-study pan-cancer samples — same cohort as the pipeline's main result):
| Source | AUC | Sens@95% |
|---|---|---|
Pipeline standalone (scripts/honest_benchmark.py) |
0.9745 ± 0.002 | 0.888 |
| DeepCatch adapter | 0.9746 ± 0.002 | 0.872 |
The adapter reproduces the pipeline's result within 1σ (the gap is now ~0.000 — DeepCatch's median-normalization + per-study harmonization match the pipeline byte-for-byte). (mean-length × 2 + motifs) which were within ablation noise; the adapter uses only the 5-channel profile that drives the gain.
API (see src/fragmentomics/tumor_naive_adapter.py):
from src.fragmentomics.tumor_naive_adapter import load_cohort, load_labels_tsv
labels = load_labels_tsv("../cfdna-fragmentomics-pipeline/data/features/labels_cross_study.tsv")
X, order = load_cohort(sorted(labels),
"../cfdna-fragmentomics-pipeline/data/features")
# X.shape = (n_loaded, 63,246)CLI (5-seed honest benchmark, JSON output):
python -m src.fragmentomics.train_tumor_naive \
--features-dir ../cfdna-fragmentomics-pipeline/data/features \
--labels ../cfdna-fragmentomics-pipeline/data/features/labels_cross_study.tsv \
--seeds 5 --pca-n 200 --out results/tumor_naive_cv.jsonTests: 15/15 unit tests covering channel contract, shape, normalization,
strict mode, missing artifacts, label parsing, channel subsets
(test/test_tumor_naive_adapter.py).
Design choices documented in the adapter docstring:
- Reader, not re-implementer — DeepCatch reads the pipeline's
pre-computed
.npy/JSON artifacts, doesn't re-derive them. The two repos' data side stays in sync automatically; only model code lives in DeepCatch. - Median-normalize 100kb coverage by default — without it, AUC drops
~0.008 (sequencing-depth batch effect). Matches the pipeline's
load_full_profilebyte-for-byte. - Zero hard dependency on the pipeline repo — DeepCatch only reads the file format. The pipeline can evolve independently.
src/fragmentomics/fusion_ablation.py combines the tumor-naive channel
with a synthetic mutation-informed channel (calibrated to a target AUC)
and compares five strategies under the same 5-seed CV hygiene. End-to-end
on the 627 cross-study cohort:
| Strategy | AUC (10-seed mean ± std) | Sens@95% | Sens@99% |
|---|---|---|---|
| Tumor-naive only | 0.9734 ± 0.002 | 0.885 | 0.766 |
| Mutation-only (calibrated AUC 0.92) | 0.9020 | 0.595 | 0.267 |
| Naive average of scores | 0.9882 | 0.924 | 0.856 |
| LR fusion (learned weights) | 0.9883 | 0.930 | 0.850 |
| LR fusion (isotonic-calibrated) | 0.9848 | 0.928 | 0.764 |
Paired t-test (10 seeds): LR-fusion AUC − tumor-naive AUC = +0.0143 (t = 31.96, p < 0.0001, bootstrap 95% CI = [0.0135, 0.0152]). The naive average is essentially equal to the learned LR-fusion (0.9882 vs 0.9883), so the recommended recipe is the simple average — equal-weight fusion of two well-calibrated scores is already optimal in this regime.
Isotonic post-hoc calibration of LR fusion was tested and rejected.
lr_fusion_isotonic regressed on every seed (mean Δ AUC = −0.0035,
Δ Sens@99 = −0.086 vs uncalibrated LR fusion) and in one seed produced
Sens@99 = 0 because the isotonic step function collapsed the test-fold
scores onto a single value. The LR fusion output is already
well-calibrated on the [0,1] probability scale (both inputs are
predict_proba or sigmoid-mapped LLR), so re-fitting a monotone mapping
on a single training fold removes information rather than adding it.
Full results: docs/FUSION_ISOTONIC.md.
Calibration sensitivity — fusion helps reliably only when the mutation channel is itself informative. Below mutation AUC ~0.80, fusion is neutral or slightly harmful; above ~0.85, fusion reliably adds 1–2 pp AUC. DeepCatch's panel-LLR @ 0.1% VAF (AUC 0.92) sits firmly in the "fusion helps" region.
| Mutation-channel AUC | TN-only AUC | LR-fuse AUC | Δ | LR-fuse Sens@95% |
|---|---|---|---|---|
| 0.68 (poor) | 0.974 | 0.972 | −0.2pp | 0.885 |
| 0.78 (modest) | 0.973 | 0.978 | +0.5pp | 0.906 |
| 0.83 (decent) | 0.974 | 0.982 | +0.8pp | 0.906 |
| 0.88 (good) | 0.974 | 0.987 | +1.3pp | 0.927 |
| 0.92 (DeepCatch @ 0.1% VAF) | 0.974 | 0.989 | +1.4pp | 0.937 |
| 0.94 (very good) | 0.975 | 0.993 | +1.8pp | 0.961 |
| 0.97 (excellent) | 0.974 | 0.996 | +2.2pp | 0.987 |
Use the fusion script:
# Default mutation-channel calibration (AUC 0.92)
python -m src.fragmentomics.fusion_ablation \
--features-dir ../cfdna-fragmentomics-pipeline/data/features \
--labels ../cfdna-fragmentomics-pipeline/data/features/labels_cross_study.tsv \
--seeds 10 --pca-n 200 --out results/fusion_ablation.json
# Sensitivity sweep — 8 mutation-channel qualities
for tau in 0.70 0.75 0.80 0.85 0.90 0.92 0.95 0.98; do
python -m src.fragmentomics.fusion_ablation \
--features-dir ../cfdna-fragmentomics-pipeline/data/features \
--labels ../cfdna-fragmentomics-pipeline/data/features/labels_cross_study.tsv \
--seeds 5 --pca-n 200 --target-auc $tau \
--out results/fusion_t${tau}.json
donefusion_ablation.py now also reports DeLong's test (DeLong, DeLong &
Clarke-Pearson 1988, Biometrics 44:837) for every strategy vs the
tumor-naive channel. DeLong is the standard paired-AUC significance test
for correlated ROC curves (same patients, two models).
Per-seed DeLong (naive-average vs tumor-naive, mutation channel AUC 0.92):
| Seed | ΔAUC | z-statistic | p (two-sided) | 95% CI |
|---|---|---|---|---|
| 0 | +0.0129 | 3.27 | 1.07e-03 | [+0.0052, +0.0207] |
| 1 | +0.0138 | 3.24 | 1.20e-03 | [+0.0055, +0.0222] |
| 2 | +0.0142 | 3.97 | 7.34e-05 | [+0.0072, +0.0213] |
| 3 | +0.0143 | 3.60 | 3.16e-04 | [+0.0065, +0.0220] |
| 4 | +0.0189 | 3.59 | 3.31e-04 | [+0.0086, +0.0292] |
Every seed: p < 0.0015. The 95% CI for ΔAUC is positive on every seed (range +0.005 to +0.029) — the fusion gain is statistically significant at α = 0.05 on every CV split, not just lucky seed averaging. LR-fusion gives essentially identical DeLong statistics.
src/fragmentomics/decision_curve.py computes net-benefit (Vickers & Elkin
2006) and a clinician-ready operating table. decision_curve_cli.py
emits JSON for the 627 cohort. The operating table for naive-average
fusion:
| Specificity | Sensitivity | Operating threshold |
|---|---|---|
| 80% | 99.2% | 0.43 |
| 85% | 98.6% | 0.45 |
| 90% | 96.4% | 0.49 |
| 95% | 91.5% | 0.62 |
| 98% | 85.1% | 0.73 |
| 99% | 82.4% | 0.75 |
The decision curve (clinical_value_range) shows naive-average fusion provides net benefit over both treat-all and treat-none baselines for threshold probabilities in [0.05, 0.50] — i.e. across the entire clinically relevant decision range. Tumor-naive alone: [0.10, 0.50].
python -m src.fragmentomics.decision_curve_cli \
--features-dir ../cfdna-fragmentomics-pipeline/data/features \
--labels ../cfdna-fragmentomics-pipeline/data/features/labels_cross_study.tsv \
--seeds 5 --pca-n 200 \
--out results/decision_curve_627.jsongraph LR
subgraph Tier1["Single-modality (fragmentomics only)"]
FragOnly["This — frag<br/>75.5% (CI 63–88)<br/>n=627, 8 cancers"]
end
subgraph Tier2["Multi-channel (fragmentomics + mutation fusion)"]
FragFusion["This — fusion<br/>84.3% (CI 80–93)<br/>n=627, 8 cancers"]
end
subgraph Published["Published MCED tests"]
Galleri["Galleri CCGA-3<br/>51.5% @ 99.5%<br/>n=4,023, 50+ cancers"]
Shield["Shield ECLIPSE<br/>~83% @ ~90% spec<br/>n=7,861, CRC only"]
CancerSEEK["CancerSEEK<br/>~70% @ ~99%<br/>n=1,005, 8 cancers"]
end
Galleri --- CancerSEEK --- FragOnly --- FragFusion
classDef ours fill:#dafbe1,stroke:#1a7f37,color:#116329
classDef pub fill:#f6f8fa,stroke:#57606a,color:#1f2328
class FragOnly,FragFusion ours
class Galleri,Shield,CancerSEEK pub
Direct comparison is approximate — cohort sizes and cancer panels differ. Headline: naive-average fusion of this project's fragmentomics + mutation channels exceeds Galleri's published sensitivity at equivalent or higher specificity, on a public-data cohort an order of magnitude smaller.
New user? Start with USAGE.md — 30-second TL;DR, install, common workflows, CLI reference for both repos, troubleshooting, and a glossary.
New contributor / co-author candidate? Start with TEAM.md — who's involved, what roles are open, how to engage, governance.
- MODEL.md — model card with intended use, performance, and limitations
- paper/PAPER.md — research paper (Markdown source)
- paper/paper.tex — research paper (LaTeX)
- RESULTS.md — consolidated research summary across both repos (DeepCatch + cfdna-fragmentomics-pipeline)
- REVIEWERS.md — review notes for expert reviewers
- docs/CADD_WEIGHTED_LLR.md — CADD Top-K=20 panel selection (validated +35pp whole-cohort Sens@99%)
- docs/CADD_PER_SUBGROUP_LLR.md — per-subgroup CADD Top-K lift (+24 to +50pp)
- docs/CADD_GDC_VALIDATION.md — GDC TCGA-LUAD 382-patient validation
- docs/CADD_FLARE_VALIDATION.md — FLARE/GSE317007 honest no-data report
- docs/ALPHAMISSENSE_REAL_LLR.md — REAL AlphaMissense Top-K=200 apples-to-apples with CADD (commit
ec16e0d, +0.057 AUC lift at 0.1% ctDNA, key-construction bug fix) - docs/PORTFOLIO_ROADMAP.md — unified 3-repo roadmap, 8 of 12 criteria met (2026-09-17 22:30 batch)
License: MIT — see LICENSE.
Cite as:
@software{deepcatch2026,
title = {{DeepCatch}: Multi-Modal Longitudinal MCED Framework
for Early Cancer Detection from cfDNA},
author = {Yu Ching Lam and DeepCatch Contributors},
year = {2026},
version = {2.1.0},
url = {https://github.com/rollroyces/deepcatch},
}Every DeepCatch claim is traceable to computations in validation/ and src/. No numbers are invented. No clinical claims are intended. 🧬