A systematic trading research engine for cryptocurrency.
The finding: the model predicts Bitcoin's direction better than chance — 52.41% accuracy against a 50.43% base rate, properly calibrated, validated walk-forward over five years. Every attempt to trade that edge lost money. Across 36 configurations, alpha was negative in all of them.
That result is the point of the project, not a disappointment in it.
Most amateur backtests are wrong in three specific ways: they ignore transaction costs, they let signals see prices that had not happened yet, and they report results from the same period used to tune the parameters.
TradeMind was built with those three failures designed out from the start, so that when a strategy looks profitable there is some reason to believe it.
Tested on 172,779 fifteen-minute Bitcoin candles from Binance — five years, validated on ingestion (3 missing bars across 2 gaps, caught automatically).
A moving-average crossover, the textbook baseline
| In-sample Sharpe | 0.46 |
| Out-of-sample Sharpe | −0.14 |
| Out-of-sample return | −20.91% |
| Buy and hold, same window | +38.05% |
The 0.60 Sharpe gap is the portion of the backtest that was parameters fitting past noise. A single train/test split would have shown only the 0.46.
A trained directional model
| Accuracy, out of sample | 52.41% |
| Base rate | 50.43% |
| Brier score | 0.24951 (baseline 0.24997) |
| Skill over baseline | +0.185% |
Calibrated: in the 0.60–0.70 bucket the hourly model predicted 61.0% and 60.7% actually happened. When it says 61 out of 100, that is close to what it means.
Trading that model
| Rule | Trades | Net | Time in market | Alpha |
|---|---|---|---|---|
| Single threshold | 11,384 | −100.00% | 51.5% | −113% |
| Hysteresis, 12h hold | 2,459 | −90.26% | 75.6% | −110% |
| Hysteresis, 10d hold | 177 | +14.60% | 98.6% | −12.47% |
The last row is the interesting one. It made money — and it was in the market 98.6% of the time. Holding for that same fraction of the period with no skill whatsoever would have returned +27.08%. The strategy had quietly converged on buy-and-hold and underperformed it by 12 percentage points.
Alpha was negative in all 36 configurations, across both a 15-minute and a 1-hour model. A 2-percentage-point edge is real. It cannot pay a 24 basis point round-turn toll.
- Point-in-time data. A position decided on bar t is applied to the move from t to t+1. No signal can act on the bar it is predicting.
- Costs are not optional. Exchange fee, half the quoted spread and slippage are charged on every change in position.
- Walk-forward everywhere. Parameters are chosen on each training window and measured once on the untouched window that follows.
- Embargoed splits. Rows at the train/test boundary are dropped, because their labels are drawn from the test period.
- Calibration, not just accuracy. A probability that is not calibrated is a score wearing a percentage sign.
- The same-time-in-market benchmark. The column that separates timing skill from simply owning an asset that went up.
- The interface never invents a number. No model, no reading. No connection, no price. The header states its data source at all times.
trademind/
adapters/ venue contract, Binance ingestion, synthetic test venue
storage.py SQLite bar store, idempotent writes, resumable checkpoints
validation.py gap, duplicate, alignment, sanity and lookahead checks
ingest.py historical backfill
live.py live quote / candle / order-book proxy, cached and retrying
server.py read-only local API and static host
paper.py live paper-trading account, simulation only
research/
costs.py transaction cost model
strategy.py baselines, probability rules, hysteresis
engine.py bar-level backtest engine and metrics
walkforward.py walk-forward validation
features.py feature construction, shared by training and live scoring
model.py calibrated logistic regression
train.py training, scoring and the economic test
tearsheet.py self-contained HTML report per run
tests/ 130 tests across four suites
pip install -r requirements.txt
python tests/test_pipeline.py && python tests/test_research.py \
&& python tests/test_model.py && python tests/test_paper.py
# pull history (about 172k bars, a couple of minutes)
python -m trademind.ingest --venue binance --symbol BTC-USD --interval 15m --days 1800 --no-resume
# backtest a baseline strategy, writes a tearsheet to reports/
python -m trademind.research.run --symbol BTC-USD --interval 15m --folds 5
# train the directional model and run the economic test
python -m trademind.research.train --symbol BTC-USD --interval 15m --horizon 4 --folds 5
# live dashboard at http://127.0.0.1:8765/index.html
python -m trademind.serverSQLite, not parquet. Single file, transactional, no dependency, and its primary key makes re-ingestion idempotent. Parquet earns its complexity when research reads become the bottleneck; they are not.
Partial candles are never stored. The bar currently forming looks identical to a finished one but its values change underneath you. Store it once and a backtest stops being reproducible.
Logistic regression, not gradient boosting. Every coefficient is inspectable, which is what feeds the interface's attribution panel. On a problem where the true signal is tiny, a flexible model finds structure in noise and reports high confidence about nothing.
One feature implementation. Live scoring and training call the same function. Two copies of the same feature maths is how live and backtest silently diverge.
No stop-loss, no take-profit. Positions close when the model's view changes, not at a price. Every added rule is another parameter to fit, and the walk-forward has already demonstrated how easily that goes wrong. Whether a stop improves the out-of-sample result is a testable question this repository can answer, and has not yet been asked.
No user accounts. No multi-tenancy. No live order placement — the paper engine has no code path to any venue and needs no exchange key. No foreign exchange yet, though the adapter layer exists so that adding it is a new file rather than a rewrite.
- Order-book features (the data is fetched for display but never recorded)
- Per-symbol models — the shipped model is trained on BTC-USD only
- Stop-loss and take-profit variants, measured properly
- Automated retraining on a schedule
This is a research tool. It estimates probabilities; it does not know what the market will do. Nothing it produces is financial advice or a recommendation to buy or sell any instrument. A simulated result is not evidence of future returns. Any trading decision, and any loss, is the user's own.