Skip to content

Latest commit

 

History

28 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DeepCrypt — Cryptographic Function Detection with Graph Neural Networks

DeepCrypt detects and classifies cryptographic functions inside compiled binaries (PE & ELF). It disassembles each function, builds its Control Flow Graph (CFG), and classifies it with an Address-Aware Graph Neural Network (GAT-based, PyTorch Geometric) into one of 9 classes:

AES · RSA · ECC · SHA · XOR · HMAC · DSA · PRNG · Non-Crypto

Results (held-out test set): 83.2% accuracy · 84.0% macro F1 (best validation F1: 85.6%). See docs/paper/ for the research paper, confusion matrix, and training curves.

An earlier iteration of this project used a Random Forest over flat, Ghidra-extracted features. It is kept for comparison — see Legacy: Random Forest below. The GNN is the current implementation and is what the web app serves.


How It Works

binary (.exe / .elf)
        │
        ▼
gnn/extract_cfg.py        pefile / pyelftools + capstone (no Ghidra needed)
        │                 per-basic-block features, CFG edges, graph-level stats
        ▼
gnn/build_graph_dataset.py   labeling, class balancing, train/val/test split
        │
        ▼
gnn/train.py              AddressAwareGNN: node encoder → GAT/GCN layers with
        │                 residuals → mean+max+sum pooling → MLP classifier
        ▼
gnn/predict.py            per-function class + confidence for a new binary

The model sees three feature levels per function:

  • Node (basic block): instruction mix ratios (xor/shift/rotate/mul...), crypto-constant hits, immediate entropy, address alignment, SIMD usage, table-lookup presence
  • Edge (branch): jump distance, direction, conditional/unconditional, loop edges, branch complexity
  • Graph (function): cyclomatic complexity, loop depth, S-box/RCON/SHA constant detection, byte entropy, address span/density

Quick Start

git clone https://github.com/Preygle/crypto-function-detection-ml.git
cd crypto-function-detection-ml/1project

python -m venv venv
# Windows: .\venv\Scripts\activate    Linux/Mac: source venv/bin/activate

pip install -r requirements.txt
pip install -r gnn/requirements.txt   # torch, torch-geometric, capstone, ...

Analyze a binary (CLI)

cd gnn
python predict.py ../test_binaries/aes_sample.exe --show-all
python predict.py <your_binary> --threshold 0.6 --output results.json

A trained checkpoint ships with the repo (gnn/checkpoints/best_gnn_model.pt), so this works out of the box. Sample binaries to try live in test_binaries/.

Run the web app (FastAPI + React)

.\run_web_app.ps1

This launches:

See web_app/README.md for manual setup.

Streamlit demos

  • app.py — dual-model comparison (GNN vs. legacy Random Forest)
  • gnn/app.py — GNN-only analyzer
streamlit run app.py

Live Demo

Important

Cold start: the free Render tier spins the backend down when idle. If the UI shows a connection error, open the backend link first and wait ~50 s for it to wake up.


Retraining the GNN

  1. (Optional) Compile fresh training binaries from C sources in dataset_gen/sources/ (needs GCC/MinGW):
    python dataset_gen/02_compile_all.py
  2. Extract CFG features from every binary (JSON per binary, written to gnn/cfg_features/):
    cd gnn
    python extract_cfg.py --batch --binaries-dir ../dataset_gen/binaries --output-dir cfg_features
  3. Build the graph dataset (graph_dataset.pkl):
    python build_graph_dataset.py
  4. Train:
    python train.py --epochs 200 --conv-type gcn --hidden-dim 64 --num-layers 3
    Best checkpoint and metrics are written to gnn/checkpoints/.

Derived artifacts (gnn/cfg_features/, graph_dataset.pkl, cfg_features_dataset.csv, dataset_gen/features/) are gitignored — they are fully regenerable from the tracked binaries.


Repository Layout

1project/
├── gnn/                    # ★ Main implementation (GNN)
│   ├── extract_cfg.py      #   CFG + feature extraction (capstone)
│   ├── build_graph_dataset.py
│   ├── model.py            #   AddressAwareGNN / HierarchicalGNN
│   ├── train.py            #   training loop, early stopping
│   ├── predict.py          #   CLI inference on new binaries
│   ├── app.py              #   Streamlit demo (GNN only)
│   └── checkpoints/        #   trained model + training metrics
├── web_app/                # FastAPI backend + React frontend (serves the GNN)
├── docs/
│   └── paper/              # Research paper, figures, plot scripts
├── dataset_gen/            # Dataset generation: C sources + compiled binaries
├── test_binaries/          # Sample binaries for quick testing
├── constants.py            # Class list, crypto constants (S-boxes, RCON, ...)
├── app.py                  # Streamlit dual-model comparison (GNN vs RF)
├── crypto_predictor.py     # Legacy RF inference
├── train_model.py          # Legacy RF training
├── ghidra_extract.py       # Legacy Ghidra-based feature extraction
└── models/                 # Legacy RF model artifacts

Legacy: Random Forest (baseline)

The first iteration classified functions with a Random Forest over flat features (entropy, instruction ratios, average jump distance) extracted via Ghidra's headless analyzer, with a lightweight pefile/pyelftools fallback.

  • Inference: crypto_predictor.py · Training: train_model.py · Artifacts: models/
  • Pipeline: dataset_gen/03_extract_ghidra.py → dataset_gen/04_build_dataset.py → train_model.py
  • Requires GHIDRA_HOME pointing at a Ghidra install (JDK 17+) for accurate extraction.

It remains in the repo as the baseline the GNN is compared against (see the comparison branch and app.py).

About

DeepCrypt detects and classifies cryptographic functions inside compiled binaries (PE & ELF). It disassembles each function, builds its Control Flow Graph (CFG), and classifies it with an Address-Aware Graph Neural Network (GAT-based, PyTorch Geometric) into one of 9 classes

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages