Add Chinese literary milestones and correct Han dictation counts - #120
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8dca54b79f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: adbbfaf522
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5e3344c6ed
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| await SessionWordCountMigration.run(modelContainer: resolvedContainer) | ||
| // Retention remains active even if statistics backfill must retry. | ||
| TranscriptionAutoCleanupService.shared.startMonitoring(modelContext: mainContext) |
There was a problem hiding this comment.
Export backfilled counts before retention cleanup
When transcript retention is enabled but Usage Sync is disabled during this upgrade, the backfill sets wordCountNeedsSync, then this immediately starts a sweep that can delete the corresponding Transcription; if sync is enabled later, the outbox scan requeues the ID, but appendLocalChanges skips it because it requires a local transcription, so the corrected metric is never exported and its outbox flag is never cleared. Preserve enough data to export the metric independently, or defer cleanup until these durable outbox entries have been acknowledged.
AGENTS.md reference: AGENTS.md:L46-L46
Useful? React with 👍 / 👎.
| guard version < WordCounter.currentVersion | ||
| || (recountCurrentVersion && version == WordCounter.currentVersion), |
There was a problem hiding this comment.
Recount metrics after local re-enhancement
When a user re-enhances a saved transcription while Usage Sync is off, AudioPlayerView.reEnhanceOnly saves the changed enhancedText and posts .transcriptionCompleted, but no local caller requests recountCurrentVersion. Because metrics created by this patch are already version 2, this guard also skips them during every startup backfill, leaving dashboard totals and saved-time estimates based on the previous text even though synchronized text changes are recounted; invoke the same recount path for local completed-text changes.
Useful? React with 👍 / 👎.
The dashboard compares cumulative dictation with familiar Chinese works, all Four Great Classical Novels, combined milestones and repeated complete sets. It shows next-target progress with Chinese/English translated titles and identical approximate Chinese-original thresholds. German translations are also supplied.
Han characters now count individually for session statistics. Historical counts are recomputed when retained text is available; missing text keeps legacy estimates. Additive count-version and local-only outbox fields preserve old data, prevent count downgrades and recover saved backfill after interruption. Duplicate history uses the existing canonical-row selector. Cleanup continues after a failed migration; remote text updates recount locally without echo writes. The short-enhancement threshold retains its original word segmentation.
Validation:
Risk and rollback: benchmark lengths are rounded estimates, and deleted historical text cannot be reconstructed. Chinese totals and the derived time-saved estimate may rise. Roll back with the previous release or revert this PR. No ASR/model behavior changes, real-model benchmark or /Applications installation. See docs/chinese-literary-milestones.md. Intended formal release: 2.4.0.