Skip to content
Merged

Dev #119

Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -25,3 +25,6 @@ share/python-wheels/
.installed.cfg
*.egg
MANIFEST

# mkdocs build output (regenerate with: docs/build_docs.sh)
site/
9 changes: 7 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,11 +52,16 @@ Easiest is to check [datasets](datasets) examples to see how the above files loo

## Documentation

* [PDF reference manual](https://github.com/bedapub/splicekit/raw/main/docs/splicekit_docs.pdf)
* [Google docs](https://docs.google.com/document/d/15ZRCeK8xyg3klLktZSHZ9k__Xw_BZRn_Q-J4W35JNnQ/edit?usp=sharing) of the above PDF (comment if you like)
Full documentation, including installation, a quick start guide, and reference pages for configuration, sample annotation, features, edgeR, motif/scanRBP analysis, juDGE plots, additional analyses, JBrowse2 and the command line, is available at:

* [bedapub.github.io/splicekit](https://bedapub.github.io/splicekit/)

## Changelog<a name="changelog"></a>

**Docs**: released in September 2026

* migrated documentation to a mkdocs-material site ([bedapub.github.io/splicekit](https://bedapub.github.io/splicekit/)), retiring the PDF/Google Docs manual

**v0.8.1**: released in July 2026

* added `bam_file` column support in `samples.tab` for per-sample BAM paths (subfolder layouts)
Expand Down
22 changes: 22 additions & 0 deletions docs/build_docs.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
#!/bin/bash
# Builds the mkdocs documentation site (source: docs/, config: ../mkdocs.yml)
# into docs/site/. Uses mkdocs from the "pybio" micromamba environment
# directly by path, so this works whether or not that environment is
# currently activated. Run from anywhere -- it cd's to the repo root itself.
#
# Usage:
# docs/build_docs.sh # build the static site into docs/site/
# docs/build_docs.sh serve # serve locally with live-reload for editing

set -e

MKDOCS=/home/gregor/micromamba/envs/pybio/bin/mkdocs
# this script lives in docs/, but mkdocs.yml lives one level up at the repo root
cd "$(dirname "$0")/.."

if [ "$1" == "serve" ]; then
"$MKDOCS" serve
else
"$MKDOCS" build --strict
echo "Built docs/site/ -- open docs/site/index.html or run 'docs/build_docs.sh serve' to preview locally."
fi
15 changes: 15 additions & 0 deletions docs/pages/about.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Authors

[splicekit](https://github.com/bedapub/splicekit) is developed by [Gregor Rot](https://grexor.github.io/) and contributors at BEDA/Roche, together with [pybio](https://grexor.github.io/pybio/) and [scanRBP](https://grexor.github.io/scanRBP/).

# Citing

If you find **splicekit** useful in your work and research, please cite:

Rot, G., Wehling, A., Schmucki, R., Berntenis, N., Zhang, J. D., & Ebeling, M. (2024)<br>
[splicekit: an integrative toolkit for splicing analysis from short-read RNA-seq](https://academic.oup.com/bioinformaticsadvances/article/4/1/vbae121/7735317)<br>
Bioinformatics Advances, 4(1). [doi:10.1093/bioadv/vbae121](https://doi.org/10.1093/bioadv/vbae121)

# Issues and Suggestions

Use the [GitHub repository issues page](https://github.com/bedapub/splicekit/issues) to report issues and leave suggestions, or write to [Gregor Rot](mailto:gregor.rot@gmail.com).
60 changes: 60 additions & 0 deletions docs/pages/additional-analyses.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# Additional analyses

Beyond the core edgeR / motif / juDGE pipeline, splicekit ships several further analyses, each runnable as its own `splicekit` command and included in `splicekit process`.

## Promiscuity: `splicekit promisc`

```bash
splicekit promisc
```

Promiscuity analysis looks at genes and how many of their junctions change across conditions — genes with many independently-regulated junctions ("promiscuous" splicing) versus genes regulated through a single dominant junction.

## Cluster of pairwise logFC: `splicekit clusterlogfc`

```bash
splicekit clusterlogfc
```

Clusters comparisons by the logFC of their significant (FDR < 0.05) changes, computed separately at the junction, exon and gene level. This groups comparisons/treatments that produce similar splicing/expression signatures.

## JUNE: junction-event analysis

```bash
splicekit june
```

JUNE (junction-events) classifies pairs/sets of overlapping junctions into splicing event types — e.g. skipped exons and mutually exclusive exons — by comparing their shared and differing donor/acceptor coordinates, rather than looking at single junctions or exons in isolation.

## rMATS

```bash
splicekit rmats
```

Runs [rMATS](http://rnaseq-mats.sourceforge.net/) on the project's BAM files as an independent, complementary splicing-event caller (skipped exons, alternative 5'/3' splice sites, mutually exclusive exons, retained introns) alongside splicekit's own junction/exon-based analysis.

## DEXSeq

```bash
splicekit dexseq # all feature types
splicekit dexseq junctions # a single feature type
```

An alternative to `splicekit edgeR` for differential feature usage, using [DEXSeq](https://bioconductor.org/packages/release/bioc/html/DEXSeq.html) instead of edgeR. Requires `dexseq_scripts` to be set in `splicekit.config` to the path of the DEXSeq R scripts.

## DonJuAn: `splicekit juan`

```bash
splicekit juan
```

Merges donor/acceptor anchor edgeR results back into the junction results (`results/results_edgeR_junctions.tab`), so each junction's row also reports its anchors' differential usage statistics. Runs automatically as part of `splicekit process`, right after `splicekit edgeR`.

## HTML report: `splicekit report`

```bash
splicekit report
```

Assembles the project's results (edgeR, JUNE, and more) into a browsable HTML report under `report/`, served together with [JBrowse2](jbrowse2.md) when you run `splicekit web`.
Binary file added docs/pages/assets/judge.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/pages/assets/splicekit_logo.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
62 changes: 62 additions & 0 deletions docs/pages/cli.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,62 @@
# Command-line reference

Every `splicekit` invocation prints its version first, then runs the requested command. Run `splicekit -help` for the built-in summary; this page is the fuller reference.

## Processing

| Command | Description |
|---|---|
| `splicekit process` | Run all available analyses in order: setup, annotation, features, edgeR, juan, judge, motifs, promisc, clusterlogfc, june, report, jbrowse2. |

## Initialization

| Command | Description |
|---|---|
| `splicekit setup` | Initialize the project folder structure. |
| `splicekit annotation` | Load `samples.tab` and build comparisons. See [Sample annotation](samples.md). |
| `splicekit features` | Create junction/exon/anchor/gene count tables from BAM files. See [Features & count tables](features.md). |

## Splicing analyses

| Command | Description |
|---|---|
| `splicekit edgeR [junctions\|exons\|anchors\|genes]` | Run edgeR differential usage analysis. All feature types if no sub-command is given. See [Differential splicing (edgeR)](edgeR.md). |
| `splicekit dexseq [junctions\|exons\|anchors\|genes]` | Run DEXSeq as an alternative to edgeR. See [Additional analyses](additional-analyses.md#dexseq). |
| `splicekit juan` | Merge donor/acceptor anchor edgeR results into the junction results. |
| `splicekit judge` | Generate juDGE plots (junction logFC vs. gene logFC). See [juDGE plots](judge.md). |
| `splicekit promisc` | Promiscuity analysis: genes and their junction changes across conditions. |
| `splicekit clusterlogfc` | Cluster comparisons by logFC of significant (FDR < 0.05) changes, at junction/exon/gene level. |
| `splicekit june` | JUNE junction-event analysis (e.g. skipped/mutually exclusive exons). |
| `splicekit rmats` | Process the project's BAM files with rMATS. |
| `splicekit cassettes` | Cassette exon analysis. |
| `splicekit patterns` | Sequence pattern analysis. |

## Motif & scanRBP

| Command | Description |
|---|---|
| `splicekit motifs [dreme\|scanrbp]` | Run motif logos, DREME and scanRBP analysis. All three if no sub-command is given. See [Motif & RNA-binding analysis](motifs.md). |

## JBrowse2 & reporting

| Command | Description |
|---|---|
| `splicekit jbrowse2 [process\|start]` | Process (build) JBrowse2 files and/or start the JBrowse2 web server. Both steps if no sub-command is given. |
| `splicekit jbrowse` | Alias for `splicekit jbrowse2`. |
| `splicekit web` | Start the local web server serving both the HTML report and JBrowse2. See [Exploring results](jbrowse2.md). |
| `splicekit report` | Generate the HTML report under `report/`. |

## Other

| Command | Description |
|---|---|
| `splicekit version` | Print the installed splicekit version. |

## Global options

| Option | Description |
|---|---|
| `-help` | Print the built-in usage summary (or a sub-command's usage, e.g. `splicekit edgeR -help`). |
| `-version` | Print the installed version and exit. |
| `-verbose` | Print more detailed progress output. |
| `-force` | Force recomputation of steps that would otherwise be skipped if their output already exists (used by `process` and `jbrowse2`). |
136 changes: 136 additions & 0 deletions docs/pages/configuration.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
# Configuration

splicekit reads parameters from two files in your project folder:

- **`splicekit.config`** — a Python file (each line is `exec`'d) describing the study, sample annotation, genome and analysis parameters. Start from the [template](https://github.com/bedapub/splicekit/blob/main/splicekit/splicekit.config.template) or one of the [datasets](https://github.com/bedapub/splicekit/tree/main/datasets) examples.
- **`config.yaml`** — a Snakemake config file describing per-rule compute resources (cores/memory/time). Start from the [template config.yaml](https://github.com/bedapub/splicekit/blob/main/config.yaml).

## splicekit.config

### Core parameters

**`study_name`**, default = `"descriptive short study name"`
Descriptive short study name, arbitrary string describing the study.

**`library_type`**, default = `"paired-end"`
Either `"paired-end"` or `"single-end"`.

**`library_strand`**, default = `"NONE"`
Possible values:

- `"SECOND_READ_TRANSCRIPTION_STRAND"`
- `"FIRST_READ_TRANSCRIPTION_STRAND"`
- `"SINGLE_STRAND"`
- `"SINGLE_REVERSE"`
- `"NONE"`

For unstranded data, use `"NONE"`. For paired-end stranded data, the most common value is `"SECOND_READ_TRANSCRIPTION_STRAND"`, meaning the second read of the pair maps in the transcript direction and the first read maps in the reverse direction. For stranded single-end sequencing, `"SINGLE_STRAND"` means the reads map in the transcript direction, and `"SINGLE_REVERSE"` means the reads map on the opposite strand of the transcripts.

**`edgeR_FDR_thr`**, default = `0.05`
FDR threshold used to filter `results/results_edgeR_{feature_type}.tab`. See [Differential splicing (edgeR)](edgeR.md).

**`dexseq_scripts`**, default = `""`
Path to DEXSeq R scripts, if you want to run `splicekit dexseq` as an alternative to `splicekit edgeR`.

### Sample annotation parameters

**`sample_column`**, default = `"sample_id"`
Which column in `samples.tab` holds the sample IDs. BAM files are then expected to be named `sample_id.bam` (see `bam_path`/`bam_column` below).

**`treatment_column`**, default = `"treatment_id"`
Which column in `samples.tab` defines treatment (test and control labels).

**`control_name`**, default = `"DMSO"`
The text identifying control samples in the treatment column. Non-control samples are compared against these.

**`separate_column`**, default = `""`
Group samples by this column and only compare within a group (empty = don't separate). Useful when, say, samples come from two tissues and you only want within-tissue comparisons: add a tissue column to `samples.tab` and set `separate_column` to its name.

**`group_column`**, default = `""`
Only compare a test sample against controls from the same "domain" (e.g. the same sequencing plate).

### Genome parameters

**`species`**, default = `None`
Genome species, resolved through `pybio`/Ensembl. Any species pybio knows about, or a custom species you registered with pybio via a FASTA+GTF pair.

**`genome_version`**, default = `None`
Ensembl genome version, or your custom genome version. `None` takes the latest Ensembl release.

**`bam_path`**, default = `""`
Folder where BAM files are stored, expected to contain `{sample_id}.bam` for every sample. Absolute, or relative to the project folder.

**`bam_column`**, default = `"bam_file"`
Alternative to `bam_path`: the name of a column in `samples.tab` holding the full BAM path for each sample individually (useful when BAMs live in per-sample subfolders rather than one flat folder). At least one of `bam_path` or `bam_column` must resolve to something — splicekit exits with an error otherwise.

### scanRBP parameters

**`scanRBP`**, default = `True`
Should splicekit run the scanRBP RNA-protein binding analysis as part of `splicekit motifs`? See [Motif & RNA-binding analysis](motifs.md).

**`protein`**, default = `"K562.TARDBP.0"`
ID of the protein PWM to scan regulated sequences against. List available IDs with `scanRBP search <term>`, e.g.:

```
$ scanRBP search hnRNPA
scan_id protein tissue description
HNRNPA1.HepG2.00 HNRNPA1 HepG2 heterogeneous nuclear ribonucleoprotein A1
HNRNPA1.HepG2.01 HNRNPA1 HepG2 heterogeneous nuclear ribonucleoprotein A1
...
HNRNPA1.K562.00 HNRNPA1 K562 heterogeneous nuclear ribonucleoprotein A1
```

**`protein_label`**, default = `"tdp43"`
Short descriptive label for the protein, used in file names and plot titles.

### Processing parameters

**`platform`**, default = `"desktop"`
`"desktop"`, `"cluster"` (LSF/`bsub`) or `"SLURM"`. When running under Snakemake, job submission is instead handled by `run_snakemake_local.sh`/`run_snakemake_slurm.sh` and `config.yaml`.

**`container`**, default = `""`
Empty string assumes all software dependencies are installed on the machine/cluster. Set to `"singularity run docker://ghcr.io/bedapub/splicekit:main"` to run non-Python steps from the provided container image instead (pybio and scanRBP are installed via pip regardless).

**`edgeR_memory`**, default = `"16GB"`
Memory reserved for edgeR cluster jobs (only relevant on `"cluster"`/`"SLURM"` platforms; under Snakemake, use `config.yaml` instead).

### Visualization and labeling parameters

**`short_names`**, default = `[]`
By default, no names are shortened or replaced in results or plots. Provide a list of triples to replace/shorten long strings:

```python
# example splicekit.config short_names parameter
# replace cell_line_A with A, cell_line_B with B (only on an exact/"complete" match)
short_names = [("cell_line_A", "A", "complete"), ("cell_line_B", "B", "complete")]
```

Use `"partial"` instead of `"complete"` to also replace the string when it occurs inside a larger one (e.g. `...cell_line_A...` → `...A...`).

## config.yaml

`config.yaml` sizes the Snakemake rules — how many cores, how much memory and how much walltime each step gets, whether running locally or submitting to SLURM.

```yaml
mapping:
perform_mapping: True # should splicekit map FASTQ -> BAM with pybio/STAR? (True/False)
alignIntronMax: 0 # STAR max intron size; 0 lets STAR derive it from its default window parameters

defaults:
cores: 1
mem: 4 # GB
time: "01:00:00"

map_fastq_single: { cores: 8, mem: 8, time: "02:00:00" }
map_fastq_paired: { cores: 8, mem: 8, time: "02:00:00" }
bam_index: { cores: 8, mem: 2, time: "02:00:00" }
bam_bw: { cores: 8, mem: 2, time: "02:00:00" }
feature_counts: { cores: 8, mem: 2, time: "01:00:00" }
edgeR: { cores: 1, mem: 4, time: "04:00:00" }
edgeR_assemble: { cores: 1, mem: 16, time: "04:00:00" }
juan: { cores: 1, mem: 16, time: "01:00:00" }
```

Any rule without its own section falls back to `defaults`. Set `mapping.perform_mapping: False` if you're supplying your own BAM files and don't want splicekit to align FASTQs itself.

After setting up `splicekit.config` and `config.yaml`, run the pipeline as described in [Quick Start](quickstart.md).
13 changes: 13 additions & 0 deletions docs/pages/coordinates.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Genomic coordinates

All genomic coordinates splicekit operates with are **0-based, left+right inclusive**. E.g. the range 100-103 includes coordinates 100, 101, 102 and 103; the first coordinate is 0.

More specifically:

- **Feature-specific**: all feature coordinates (junctions, anchors, exons) are given in numeric sort order regardless of strand — `feature_start` is always `<` `feature_stop`. Example: `chr1+_100_102` represents a junction spanning coordinates `[100, 101, 102]`; `chr1-_100_102` represents a junction spanning coordinates `[102, 101, 100]`.
- **Junction-specific**: junction coordinates cover/overlap 1 nucleotide of each adjoining exon.
- `chr1+_100_200` — a junction on chromosome 1 (`+` strand) from `[100..200]`. 100 is the last nucleotide of the upstream exon, 200 is the first nucleotide of the downstream exon.
- `chr1-_100_200` — a junction on chromosome 1 (`-` strand) from `[100..200]`. 200 is the last nucleotide of the upstream exon, 100 is the first nucleotide of the downstream exon.

!!! important
RefSeq and Ensembl GTF files are 1-indexed. When splicekit reads files from RefSeq/Ensembl, it performs `coordinate -= 1` on every coordinate to keep them consistent with splicekit's internal 0-indexed structures.
21 changes: 21 additions & 0 deletions docs/pages/dependencies.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# Dependencies

## Conda/micromamba environment (`splicekit.yaml`)

Installed by `micromamba create -y -f splicekit.yaml`:

[pigz](https://zlib.net/pigz/), [deeptools](https://deeptools.readthedocs.io/), [samtools](http://www.htslib.org/), [Snakemake](https://snakemake.readthedocs.io/), R (`r-base`, `r-locfit`), [STAR](https://github.com/alexdobin/STAR), [rMATS](http://rnaseq-mats.sourceforge.net/), [Node.js](https://nodejs.org/), Ghostscript, [subread](http://subread.sourceforge.net/) (`featureCounts`, > 2.0.6), [MEME](https://meme-suite.org/meme/) (> 5.5.1, for DREME), Perl's `cpanminus`, and the `snakemake-executor-plugin-cluster-generic` pip package (used by `run_snakemake_slurm.sh` for SLURM submission).

## Installed by `install.sh`

- R packages: `BiocManager`, `data.table`, `statmod`, `R.utils`, and Bioconductor's [edgeR](https://bioconductor.org/packages/release/bioc/html/edgeR.html).
- Perl modules for the JBrowse2/SOAP tooling (`XML::Compile::*`, `Log::Log4perl`, `Math::CDF`, `JSON`, and others).
- [`@jbrowse/cli`](https://www.npmjs.com/package/@jbrowse/cli) via `npm install -g`, used to set up the local [JBrowse2](jbrowse2.md) instance.

## Python dependencies (installed by `pip install .`)

[Levenshtein](https://pypi.org/project/Levenshtein/), [logomaker](https://logomaker.readthedocs.io/), [plotly](https://plotly.com/python/), `python-dateutil`, [pybio](https://grexor.github.io/pybio/), [scanRBP](https://grexor.github.io/scanRBP/), [pysam](https://pysam.readthedocs.io/), [numpy](https://numpy.org/), `psutil`, `beautifulsoup4`, `requests` and `rangehttpserver`.

## Optional

[Singularity](https://sylabs.io/singularity/) — only needed if you set `container = "singularity run docker://ghcr.io/bedapub/splicekit:main"` in `splicekit.config` instead of installing the conda environment directly. See [Installation: Container](installation.md#container).
Loading
Loading