Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
124 changes: 124 additions & 0 deletions docs/issue-166-validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
# Issue #166: offline decryption validation

Validated on 2026-09-28 against upstream `1b516cf` plus this fix, on macOS arm64
(macOS 26.3.1, Python 3.11.15). No private databases, keys, account identifiers,
message content, or raw application logs are included in this report.

## Real WeChat data

A stable, private copy of an existing local account was used. The project's
normal scanner selected **20 business databases**; search indexes and key-info
files were excluded by the existing database filter. Main files plus sidecars
copied for these databases totalled **1,332,165,592 bytes**.

No original WeChat files were modified. Each database/sidecar pair was copied
only when its before/after size, modification time and inode were unchanged.
Input-copy SHA-256 checksums were also unchanged after each test. Runtime output
and key persistence used a separate private temporary directory.

| Check | Result |
| --- | --- |
| Production `GET /api/decrypt_stream` router through FastAPI TestClient | HTTP 200; completed; 20 successful, 0 failed |
| Production `POST /api/decrypt` router over localhost HTTP | HTTP 200; completed; 20 successful, 0 failed |
| Key persistence through the actual router | Successful on both routes |
| Independent `PRAGMA integrity_check` on every output | 20/20 passed after each route |
| `session.db` | Authenticated, decrypted and passed integrity checking; 7 tables |
| WAL replay observed | 1 database, 4 committed frames, 3 applied pages |
| Production chat sessions route with `source=decrypted` | HTTP 200; returned 5 sessions |
| Production chat messages route with `source=decrypted` | HTTP 200; returned 5 messages |

The other nonempty WAL files did not contain current committed frames selected
by recovery; file size alone does not establish outstanding transactions. The
unmodified real account did not require index rebuilding. Its successful export
therefore does **not** independently reproduce the reporter's original index
corruption. That failure and recovery are covered by the encrypted regression
fixtures below.

A second local account was discovered but not decrypted because none of its
stored keys passed cross-database authentication. This is an untested account,
not a successful test or a decryption regression.

## Desktop startup follow-up

The original expired-runtime blocker was resolved upstream in `1b516cf`, merged
into this branch. The new official macOS source runtime is
`macos-source-runtime-20260928-1790594550`, with expiry
`2026-11-12T11:22:30Z`. The standard `npm run dev` entry point downloaded it,
verified its pinned artifacts, accepted the `source-public` profile, and launched
Nuxt and Electron. A subsequent launch verified and reused the cache. No expiry,
signature or native-runtime checks were bypassed.

The first launch exceeded the desktop's default 30-second readiness timeout
while setting up a fresh Python environment. After dependencies were installed,
a retry with a longer desktop startup timeout exposed a separate local wait:
the native broker remained inside macOS `SecItemCopyMatching`, called by
`load_or_create_device_identity`. The backend did not reach its health endpoint
before the broker startup timeout. A process sample established this wait;
it is not evidence that the renewed runtime is expired or incompatible.

The user completed the macOS keychain prompt and the application's first-use
agreement. The standard Electron-launched backend then returned HTTP 200 with
`status=healthy`, and the actual Electron window displayed the application and
existing chat data.

### Issue-specific desktop backend regression

The existing encrypted regression fixture generator was reused for two
snapshots with the exact named-index entry-count error from #166. The same
input bytes were tested against the unfixed code and the running desktop
backend. The baseline checkout is `b70da36`; its decrypt implementation,
SQLite diagnostics and decrypt router are unchanged through upstream `1b516cf`.

| Same-input comparison | Unfixed code | Fixed desktop backend |
| --- | --- | --- |
| Page authentication | 3/3 pages, no HMAC warnings | Authenticated |
| Named-index mismatch without WAL | `wrong # of entries in index SessionUnreadListTable_1_NameId_CreateTime`; session fails | Rebuilds only the affected index; full integrity check passes |
| Same mismatch with a correcting committed WAL page | Same index failure; WAL omitted | 1 committed frame replayed; full integrity check passes without REINDEX |
| Key persistence | Rejected because session did not verify | Saved after session/message authentication |
| Records in repaired session fixture | Output rejected | Both expected records present |

Each fixture included an authenticated message database. Both POST requests to
the actual desktop backend completed with 2 successes and 0 failures. The WAL
case also passed the actual SSE endpoint used by the decrypt page, ending in
`complete`, 2 successes, 0 failures and `db_key_persisted=true`. Input hashes
were unchanged. These are encrypted SQLite fixtures, not the reporter's
unavailable 734-page original database.

The already prepared real-account snapshot was then tested through this same
Electron-launched backend: **20/20 successful, 0 failures, 20/20 independent
integrity checks passed, key persisted, and source-copy hashes unchanged**.
The POST request took 6.82 seconds. Explicit `source=decrypted` chat requests
returned 5 sessions and 5 messages from these outputs. The visible Electron
chat page was also inspected; private content and screenshots are not published.

This is targeted regression testing against the complete desktop application's
backend plus a visible desktop startup/chat check. It does not claim automated
click-through of the entire first-use/key-capture/decrypt wizard. The earlier
minimal-router HTTP tests remain supplementary evidence. No new account-key
capture or modification of original WeChat databases was needed.

The frontend production static build (`npm run generate`) passed, generating
34 routes. After merging `1b516cf`, all 76 focused Python tests and 3 frontend
feedback tests passed again.

## Regression tests

- 76 focused Python tests passed across offline WAL recovery, key modes, SSE,
decryption UI contracts, decrypted fallback and account validation.
- 3 frontend feedback tests passed (`node --test frontend/tests/decrypt-feedback.test.mjs`).
- `git diff --check` passed.

`tests/test_offline_wal_recovery.py` uses real SQLite pages and encrypted
fixtures, without mocking integrity diagnostics. Coverage includes:

- The same named-index entry-count error reported in #166, with all encrypted
pages authenticating; recovery both from a committed WAL page and by rebuilding
only the affected indexes on a disposable output.
- Little- and big-endian WAL checksums, repeated page writes, committed versus
uncommitted/partial tails, recycled salts, growth, truncation and VACUUM.
- Invalid WAL headers, frame checksums, page HMACs and missing growth pages.
- Refusing table corruption, changing sources and source/output aliases.
- Preserving the prior output when validation or atomic publication fails, and
preventing an old output WAL from overriding the new snapshot.
- Keeping key authentication independent of output integrity without accepting
a key that authenticates only one of the required database roles.
30 changes: 27 additions & 3 deletions frontend/pages/decrypt.vue
Original file line number Diff line number Diff line change
Expand Up @@ -1026,6 +1026,16 @@
</div>
</transition>

<details v-if="databaseFailures.length" class="mt-4 rounded-lg border border-amber-200 bg-amber-50 p-4">
<summary class="cursor-pointer text-sm font-semibold text-amber-800">数据库解密详情({{ databaseFailures.length }} 个未完成)</summary>
<ul class="mt-3 space-y-2 text-sm text-amber-900">
<li v-for="item in databaseFailures" :key="item.id">
<strong>{{ item.name }}</strong>:{{ item.error }}
<span v-if="item.authenticated">(密钥认证已通过)</span>
</li>
</ul>
</details>

<!-- 错误提示 -->
<transition name="fade">
<ErrorNotice v-if="error" :message="error" class="mt-6 animate-shake" />
Expand Down Expand Up @@ -1117,7 +1127,7 @@ const {
const loading = ref(false)
const error = ref('')
const warning = ref('') // 警告,用于密钥提示
const warningIsError = computed(() => /失败|错误|异常|中断/.test(String(warning.value || '')))
const warningIsError = computed(() => !String(warning.value || '').startsWith('解密部分成功:') && /失败|错误|异常|中断/.test(String(warning.value || '')))
const currentStep = ref(0)
const mediaAccount = ref('')
const activeKeyAccount = ref('')
Expand All @@ -1144,7 +1154,7 @@ const imageKeyMemoryScanNote = computed(() => String(
|| platformCapabilities.value?.image_key_memory_scan_note
|| '图片密钥扫描原生资源缺失或安装不完整,请重新安装完整发行包。'
))
const DB_KEY_PERSISTENCE_WARNING = '数据库密钥未通过完整实时库校验或无法安全保存;请重新获取并确认主要数据库解密成功,仍失败请检查数据目录权限。'
const DB_KEY_PERSISTENCE_WARNING = '数据库密钥未通过 session/message 跨库认证或保存失败;请查看失败详情,确认账号密钥及数据目录写入权限。'
const guideDialog = reactive({
open: false,
eyebrow: '操作提示',
Expand Down Expand Up @@ -2012,8 +2022,15 @@ const cancelDbKeyAcquisition = () => {
}

const showDbKeyPersistenceWarning = (result) => {
if (result?.failure_count > 0 && result?.success_count > 0) {
warning.value = `解密部分成功:${result.success_count}/${result.total_databases} 个数据库可用;失败文件请查看解密详情,已成功的数据可继续使用。`
}
const repaired = Object.values(result?.account_results || {}).flatMap(account =>
Object.values(account.db_diagnostics || {}).filter(db => db.success && db.index_repair?.success).map(db => db.db_name)
)
if (repaired.length) warning.value = [warning.value, `已在解密输出副本中重建索引并通过完整性检查:${repaired.join('、')}。`].filter(Boolean).join(' ')
if (result?.db_key_persisted !== false) return
warning.value = DB_KEY_PERSISTENCE_WARNING
warning.value = [warning.value, DB_KEY_PERSISTENCE_WARNING].filter(Boolean).join(" ")
logDecryptDebug('decrypt:db-key-persistence-warning', {
error_count: Array.isArray(result?.db_key_persistence_errors)
? result.db_key_persistence_errors.length
Expand Down Expand Up @@ -2525,6 +2542,12 @@ const getMediaDecryptConcurrency = () => {

// 解密结果存储
const decryptResult = ref(null)
const databaseFailures = computed(() => Object.entries(decryptResult.value?.account_results || {}).flatMap(([account, result]) =>
Object.entries(result.db_diagnostics || {}).filter(([, db]) => db.success === false).map(([name, db]) => ({
id: `${account}/${name}`, name, authenticated: db.key_authenticated === true,
error: db.error === 'key_mismatch' ? '密钥与此数据库不匹配' : (db.error || '数据库完整性检查未通过')
}))
))

// 验证表单
const validateForm = () => {
Expand Down Expand Up @@ -2657,6 +2680,7 @@ const handleDecrypt = async () => {
db_key_length: String(formData.key || '').trim().length
})
loading.value = true
decryptResult.value = null
error.value = ''
warning.value = ''

Expand Down
29 changes: 29 additions & 0 deletions frontend/tests/decrypt-feedback.test.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
import test from 'node:test'
import assert from 'node:assert/strict'
import { readFileSync } from 'node:fs'
import vm from 'node:vm'

const source = readFileSync(new URL('../pages/decrypt.vue', import.meta.url), 'utf8')
function feedback(result) {
const start = source.indexOf('const showDbKeyPersistenceWarning =')
const end = source.indexOf('\nconst runMacosLldbFallback', start)
const context = { result, warning: { value: '' }, DB_KEY_PERSISTENCE_WARNING: 'KEY_SAVE_WARNING', logDecryptDebug: () => {} }
vm.runInNewContext(`${source.slice(start, end)}\nshowDbKeyPersistenceWarning(result)`, context)
return context.warning.value
}

test('partial success with an authenticated key does not ask for key recapture', () => {
const message = feedback({ success_count:27, failure_count:1, total_databases:28, db_key_persisted:true })
assert.match(message, /解密部分成功:27\/28/)
assert.doesNotMatch(message, /KEY_SAVE_WARNING/)
})

test('recovered indexes are disclosed and key saving failures stay visible', () => {
const message = feedback({ db_key_persisted:false, account_results:{ account:{ db_diagnostics:{ session:{ db_name:'session.db', success:true, index_repair:{ success:true } } } } } })
assert.match(message, /重建索引.*session\.db/)
assert.match(message, /KEY_SAVE_WARNING/)
})

test('normal success has no warning', () => {
assert.equal(feedback({ db_key_persisted:true, success_count:28, failure_count:0 }), '')
})
13 changes: 9 additions & 4 deletions src/wechat_decrypt_tool/routers/decrypt.py
Original file line number Diff line number Diff line change
Expand Up @@ -211,12 +211,17 @@ def _database_diagnostic_role(diagnostic: dict[str, Any]) -> str:

def _database_diagnostic_verified(diagnostic: dict[str, Any]) -> bool:
return bool(
diagnostic.get("success") is True
and not bool(diagnostic.get("copied_as_sqlite"))
not bool(diagnostic.get("copied_as_sqlite"))
and str(diagnostic.get("key_mode") or "").strip()
in {"raw_enc_key", "sqlcipher_passphrase"}
and int(diagnostic.get("failed_pages") or 0) == 0
and str(diagnostic.get("diagnostic_status") or "").strip() == "ok"
# Key authentication is independent of WAL/page/index integrity.
# Retain compatibility with results produced before key_authenticated.
and (diagnostic.get("key_authenticated") is True or (
"key_authenticated" not in diagnostic
and diagnostic.get("success") is True
and int(diagnostic.get("failed_pages") or 0) == 0
and str(diagnostic.get("diagnostic_status") or "").strip() == "ok"
))
)


Expand Down
38 changes: 38 additions & 0 deletions src/wechat_decrypt_tool/sqlite_diagnostics.py
Original file line number Diff line number Diff line change
Expand Up @@ -182,3 +182,41 @@ def format_sqlite_diagnostics(diagnostics: Mapping[str, Any]) -> str:
continue
compact[str(key)] = value
return json.dumps(compact, ensure_ascii=False, sort_keys=True)


def repair_sqlite_indexes(path: str | Path) -> dict[str, Any]:
"""Repair index-only damage on a disposable output, never on the source.

Do not attempt salvage of table/page corruption. A full integrity check must
pass after rebuilding every affected named index before output is accepted.
"""
import re

result: dict[str, Any] = {"attempted": False, "success": False}
conn = None
try:
conn = sqlite3.connect(str(path))
conn.execute("PRAGMA trusted_schema=OFF")
errors = [str(row[0]) for row in conn.execute("PRAGMA integrity_check(1000)")]
if not errors or errors == ["ok"] or len(errors) >= 1000:
return result
indexes = set()
for error in errors:
match = re.fullmatch(r"(?:wrong # of entries in index|row \d+ missing from index) (.+)", error)
if match is None:
return result
indexes.add(match.group(1))
names = {row[0] for row in conn.execute("SELECT name FROM sqlite_master WHERE type='index'")}
if not indexes or not indexes <= names:
return result
result.update(attempted=True, indexes=sorted(indexes), original_errors=errors[:5])
for name in sorted(indexes):
conn.execute(f"REINDEX {_quote_ident(name)}")
conn.commit()
result['success'] = conn.execute("PRAGMA integrity_check").fetchall() == [('ok',)]
except sqlite3.Error as exc:
result['error'] = _clean_error(exc)
finally:
if conn is not None:
conn.close()
return result
89 changes: 89 additions & 0 deletions src/wechat_decrypt_tool/sqlite_wal.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
"""Recover a stable SQLite/SQLCipher WAL snapshot without touching source files.

Format: https://sqlite.org/fileformat2.html#walformat. SQLCipher computes WAL
checksums over the encrypted page payload, before writing the frame to disk.
"""
from __future__ import annotations

import struct
from collections.abc import Callable


def _checksum(data: bytes, endian: str, state: tuple[int, int] = (0, 0)) -> tuple[int, int]:
a, b = state
for x, y in struct.iter_unpack(endian + 'II', data):
a = (a + x + b) & 0xffffffff
b = (b + y + a) & 0xffffffff
return a, b


def merge_wal_snapshot(
database: bytes,
wal: bytes,
page_size: int,
*,
verify_page: Callable[[bytes, int], bool] | None = None,
) -> tuple[bytes, dict]:
"""Apply only checksum-valid committed frames, respecting database truncation.

A salt mismatch marks the recycled tail of a WAL. Complete frames with a
bad checksum are rejected rather than silently exporting an older snapshot.
Incomplete/uncommitted trailing frames never become part of the output.
"""
info = {'wal_bytes': len(wal), 'committed_frames': 0, 'applied_pages': 0,
'ignored_tail_bytes': 0, 'database_pages': len(database) // page_size}
if not wal:
return database, info
if len(wal) < 32:
raise ValueError('WAL 文件头不完整,请退出微信后重试')
magic, version, wal_page_size = struct.unpack('>III', wal[:12])
if magic not in (0x377f0682, 0x377f0683) or version != 3007000:
raise ValueError('WAL 文件头或版本无效')
if wal_page_size != page_size or len(database) % page_size:
raise ValueError('WAL 与数据库页大小不匹配,或主数据库页不完整')
endian = '<' if magic == 0x377f0682 else '>'
state = _checksum(wal[:24], endian)
if state != struct.unpack('>II', wal[24:32]):
raise ValueError('WAL 文件头校验失败')
salt = wal[16:24]
frame_size = 24 + page_size
pending: dict[int, bytes] = {}
committed: dict[int, bytes] = {}
base_pages = len(database) // page_size
size = base_pages
commit_end = 32
for offset in range(32, len(wal) - frame_size + 1, frame_size):
header = wal[offset:offset + 24]
if header[8:16] != salt:
break # old frames left after WAL reset / preallocated zero tail
pgno, db_size = struct.unpack('>II', header[:8])
if pgno == 0 or pgno > 0xfffffffe:
raise ValueError('WAL 页号无效')
page = wal[offset + 24:offset + frame_size]
state = _checksum(page, endian, _checksum(header[:8], endian, state))
if state != struct.unpack('>II', header[16:24]):
raise ValueError('WAL 帧校验失败,请重新获取稳定的数据库副本')
pending[pgno] = page
if db_size:
committed.update(pending)
pending.clear()
committed = {n: p for n, p in committed.items() if n <= db_size}
base_pages = min(base_pages, db_size)
# Check before allocating: malformed sizes must not cause huge allocations.
extension = sum(n > base_pages for n in committed)
if db_size - base_pages != extension:
raise ValueError('WAL 提交缺少数据库扩展页')
size = db_size
commit_end = offset + frame_size
info['committed_frames'] = (commit_end - 32) // frame_size
info.update(applied_pages=len(committed), ignored_tail_bytes=len(wal) - commit_end,
database_pages=size)
if not info['committed_frames']:
return database, info
output = bytearray(database[:base_pages * page_size])
output.extend(b'\0' * (size * page_size - len(output)))
for pgno, page in committed.items():
if verify_page is not None and not verify_page(page, pgno):
raise ValueError(f'WAL 第 {pgno} 页 HMAC 校验失败')
output[(pgno - 1) * page_size:pgno * page_size] = page
return bytes(output), info
Loading