Skip to content

Commit c2c44ca

Browse files
authored
round-24 self-review: spoken-form time weights, any-year gate, explicit yearGate flag, evals wording (1.9.12)
1 parent 35ca1a3 commit c2c44ca

4 files changed

Lines changed: 32 additions & 15 deletions

File tree

‎plugins/Wzdhehe/html2video-for-mcode/CHANGELOG.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@
77
- **New optional field `clauses[].say`** — the spoken form of that sentence. `text` remains the on-screen/subtitle form and keeps the natural written form (Arabic numerals, original orthography — spelled-out digits like `八十二点三` belong in `say`, never on screen). **TTS input and the ASR checklist expectation are always `say ?? text`**; the burned-in subtitles and `out/subs.srt` always come from `text`. `say` propagates through `plan-timings` into `timings.json` (which is what `build-video --asr` reads); `say` applies to the main line only (`text2` stays display-only).
88
- **Year-reading gate (plan-timings, zh/yue):** a clause whose effective spoken text (`say ?? text`) contains a four-digit year followed by 年 gets a warning that names the digit-by-digit broadcast form and hands over the ready-to-paste `say` value (`加 say:"二零二六年…"`). The gate is deliberately narrow — whether a number reads digit-by-digit or as an integer is intent (`2026年` digit-by-digit, `2026点` as an integer), so only the unambiguous year shape is gated and the rest of the number discipline lives in the docs (SKILL.md schema section + tts-and-timing.md, with all three reported examples).
99
- **Tests +1 → 274 in fifteen files:** the year gate (no-`say` named with the paste-ready form / with-`say` silent / `en` not over-warned / `say` passed through to timings verbatim and never invented for `say`-less clauses), and the smoke pipeline now runs `build-video --asr` on a `say`-bearing clause and asserts the checklist expects the spoken form while the srt keeps the display form (`say` must not leak into subtitles). Each guard red-proofed: disabling the year gate, dropping the `say` propagation, or reverting the checklist to `c.text` reddens exactly its own case.
10+
- **Round-24 self-review corrections (same release):** the two-axis review of this increment found four real items in the fix itself, all closed here — ① **clause-start allocation weighted by the display form**: the audio speaks `say` ("82.3%" shows 5 chars, "百分之八十二点三" speaks 8), so time weights, per-clause `chars` and the pacing check now use `say ?? text` (subtitle-wrap geometry still uses `text` — the pill is on screen); the year-gate test now pins the weighted start against both candidate formulas. ② the year gate's regex only matched `19xx/20xx` while the comments claimed "four-digit year" — now any `\d{4}年` (1897 warns too). ③ the zh-guard used the layout metric `emPer===1` as a language proxy — replaced by an explicit `yearGate` flag on `zh`/`yue` (unknown-language fallback stays ungated). ④ two `evals.json` cases still taught the old "write the spoken form into the narration text" move (the EN number-swap line and the ASR-mismatch fix) — both now teach the `say` mechanism. Each correction red-proofed (weight reverted → weighted-start assert red; regex reverted → 1897 red; gate off → warn red).
1011

1112
## 1.9.11 — 2026-09-28
1213

‎plugins/Wzdhehe/html2video-for-mcode/skills/html2video-for-mcode/evals/evals.json‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -33,7 +33,7 @@
3333
"id": 4,
3434
"name": "asr-mismatch",
3535
"prompt": "ASR 校验发现第 5 张转写里 '92%' 听成了 '九二', 数字对不上, 怎么处理?",
36-
"expected_output": "改口播写法(如 '九成二')→ 只重做该段 TTS → 重跑 plan-timings → 删 build/frames/05 与 out/slide-05.mp4 → 该张重 capture + build-video, 不整片重做。若确认是同音字级微小差异可接受, 数字/产品名不一致必须返工。",
36+
"expected_output": "text(字幕)保持 '92%' 不动, 给该句加 say:'…百分之九十二…'(显示/发音分离: 发音形态只进 say, 不上屏)→ 只重做该段 TTS → 重跑 plan-timings → 删 build/frames/05 与 out/slide-05.mp4 → 该张重 capture + build-video, 不整片重做。若确认是同音字级微小差异可接受, 数字/产品名不一致必须返工。",
3737
"files": []
3838
},
3939
{
@@ -110,7 +110,7 @@
110110
"id": 15,
111111
"name": "english-narration",
112112
"prompt": "这条视频要发给海外同事, 口播改成英语。要改哪些地方?",
113-
"expected_output": "开工对齐时就该问清语言(铁律 4 的语言行: 中文/英语/粤语/其他)。改动四处: ① script.json 设 lang:'en'; ② 口播稿改写英语并按英文基准控量 —— plan-timings 自动切换(约 14 字符/秒、常见 9–18; 字幕单句建议 ≤42 字符(可读性惯例), 胶囊实际行宽 ≈55 字符@1080 / ≈75@1920, 折 3 行起才警告; 每张约 14–26 词); ③ 音色换成 English_* 前缀(Chinese (Mandarin)_* 配英文会出怪腔调), 做完用 asr.mjs --language en 验语种; ④ 若开双语, 主行=英语口播、text2 放中文。排版、脚本、渲染流程都不用动; 数字写法反过来(口播 nine hundred million, 画面 900M)。",
113+
"expected_output": "开工对齐时就该问清语言(铁律 4 的语言行: 中文/英语/粤语/其他)。改动四处: ① script.json 设 lang:'en'; ② 口播稿改写英语并按英文基准控量 —— plan-timings 自动切换(约 14 字符/秒、常见 9–18; 字幕单句建议 ≤42 字符(可读性惯例), 胶囊实际行宽 ≈55 字符@1080 / ≈75@1920, 折 3 行起才警告; 每张约 14–26 词); ③ 音色换成 English_* 前缀(Chinese (Mandarin)_* 配英文会出怪腔调), 做完用 asr.mjs --language en 验语种; ④ 若开双语, 主行=英语口播、text2 放中文。排版、脚本、渲染流程都不用动; 数字写法分离(text 写 900M 上字幕, 加 say:'nine hundred million' 控制发音 —— TTS/ASR 吃 say ?? text, 字幕恒显示形态)。",
114114
"files": []
115115
},
116116
{

‎plugins/Wzdhehe/html2video-for-mcode/skills/html2video-for-mcode/scripts/plan-timings.mjs‎

Lines changed: 16 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -11,9 +11,11 @@ import { positionalDir, probeDuration, requireTool, safeId, safeOut, safeRel, su
1111
// 语种相关的计量基准。中文按"字", 英文按"字符"(含词间节奏, 与音节时长大致成正比)。
1212
// pacing = 每单位每秒的常见语速; emPer = 单字宽按 em 折算(行宽估算用 —— 行宽公式在
1313
// tools.subLineCap, 字幕胶囊几何的唯一来源也在 tools.SUB_GEOMETRY); pace 区间用于语速异常预警。
14+
// yearGate = 年份逐位读的闸门只对中/粤语面开(第 24 轮: 用 emPer===1 充当语族判别是排版
15+
// 属性代理语义, 未来 emPer:1 的语种会误触); 未知语种 fallback 不开闸。
1416
const LANG_CFG = {
15-
zh: { unit: '字', pacing: 4.8, paceMin: 3, paceMax: 6.5, emPer: 1, word: null },
16-
yue: { unit: '字', pacing: 4.8, paceMin: 3, paceMax: 6.5, emPer: 1, word: null },
17+
zh: { unit: '字', pacing: 4.8, paceMin: 3, paceMax: 6.5, emPer: 1, word: null, yearGate: true },
18+
yue: { unit: '字', pacing: 4.8, paceMin: 3, paceMax: 6.5, emPer: 1, word: null, yearGate: true },
1719
en: { unit: '字符', pacing: 14, paceMin: 9, paceMax: 18, emPer: 0.44, word: 'words' },
1820
};
1921
const langCfg = code => LANG_CFG[code] ?? { unit: '字符', pacing: 14, paceMin: 8, paceMax: 20, emPer: 0.44, word: null };
@@ -54,7 +56,11 @@ for (const s of script.slides) {
5456
const tts = probeDur(audioPath);
5557
if (tts == null) { console.error(`✗ ffprobe 读不出时长: ${audioPath}`); process.exit(1); }
5658

57-
const total = s.clauses.reduce((n, c) => n + charCount(c.text), 0);
59+
// 时间权重按**发音形态**(say ?? text)算 —— 音频念的是 say(第 24 轮: "82.3%"显示 5 字、
60+
// 发音"百分之八十二点三"8 字, 按显示形态分摊会把长 say 句的开口系统性估早, 字幕窗与
61+
// ASR 切分跟着偏)。字幕行宽(:93)仍按显示形态 text —— 胶囊几何是屏幕上的事。
62+
const spokenLen = c => charCount(c.say ?? c.text);
63+
const total = s.clauses.reduce((n, c) => n + spokenLen(c), 0);
5864
if (total === 0) warns.push(`${s.id}: clauses 为空, 将整张静止`);
5965

6066
// 第 k 句开口时刻 ≈ 实测时长 × (前 k-1 句字数占比); stage 取该层最早一句, 再提前 LEAD
@@ -63,11 +69,11 @@ for (const s of script.slides) {
6369
let cum = 0;
6470
for (const c of s.clauses) {
6571
const start = total > 0 ? (tts * cum) / total : 0;
66-
// say = 发音形态(只喂 TTS/ASR), text = 显示形态(字幕/画面) —— 原样传递给下游
67-
clauses.push({ stage: c.stage ?? null, start: r3(start), chars: charCount(c.text), text: c.text, ...(c.text2 ? { text2: c.text2 } : {}), ...(c.say ? { say: c.say } : {}) });
72+
// say = 发音形态(只喂 TTS/ASR), text = 显示形态(字幕/画面) —— 原样传递给下游; chars = 发音字数(计时簿记)
73+
clauses.push({ stage: c.stage ?? null, start: r3(start), chars: spokenLen(c), text: c.text, ...(c.text2 ? { text2: c.text2 } : {}), ...(c.say ? { say: c.say } : {}) });
6874
const t = Math.max(0, start - LEAD);
6975
if (c.stage != null) stageTime[c.stage] = c.stage in stageTime ? Math.min(stageTime[c.stage], t) : t;
70-
cum += charCount(c.text);
76+
cum += spokenLen(c);
7177
}
7278
clauses.forEach((c, i) => { c.dur = r3((clauses[i + 1]?.start ?? tts) - c.start); });
7379
Object.assign(stageTime, s.stageTimes ?? {}); // 显式 stageTimes 覆盖优先
@@ -96,10 +102,11 @@ for (const s of script.slides) {
96102
if (c.text2 && c.text2.length > 60) warns.push(`${s.id} 第 ${clauses.indexOf(c) + 1} 句双语第二行 ${c.text2.length} 字符 > 60 — 建议精简译文`);
97103
// 年份读法闸门(2026-09-28 用户实测: TTS 把 2026年 念成"两千零二十六年"): 显示与发音
98104
// 分离 —— text 保持 2026年 上字幕, 加 say:'二零二六年' 控制发音。判别只钉无歧义的
99-
// "四位数字+年"(整读/逐位读是意图问题, 2026点 就该整读 —— 其余数字写法走文档纪律)。
100-
if (CFG.emPer === 1) { // 中文/粤语面; 英语 TTS 自己会读 twenty twenty-six
105+
// "四位数字+年"(任意年份都逐位读, 不止 19xx/20xx —— 第 24 轮: 旧正则漏 1897年 这类;
106+
// 整读/逐位读是意图问题, 2026点 就该整读 —— 其余数字写法走文档纪律)。
107+
if (CFG.yearGate) {
101108
const spoken = c.say ?? c.text;
102-
const ym = /(?:19|20)\d{2}\s*年/.exec(spoken);
109+
const ym = /\d{4}\s*年/.exec(spoken);
103110
if (ym) {
104111
const digitsRead = ym[0].replace(/\D/g, '').split('').map(d => '零一二三四五六七八九'[+d]).join('');
105112
warns.push(`${s.id} 第 ${clauses.indexOf(c) + 1} 句 "${ym[0]}" 会被 TTS 按整数读(2026年→"两千零二十六年") — 年份按播报惯例逐位读: text 原样上字幕, 加 say:"${digitsRead}年…"`);

‎plugins/Wzdhehe/html2video-for-mcode/skills/html2video-for-mcode/tests/render-smoke.test.mjs‎

Lines changed: 13 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -421,20 +421,29 @@ test('plan-timings 年份读法闸门: 无 say 必须点名(给逐位写法),
421421
const s = JSON.parse(fs.readFileSync(sp, 'utf8'));
422422
s.slides = [{ id: '01', html: '01-title.html', audio: '01.mp3', clauses: [
423423
{ stage: 1, text: '回顾2026年,七条要闻速览。' }, // 无 say → 必须点名
424-
{ stage: 2, text: '展望2027年,继续同行。', say: '展望二零二七年,继续同行。' }, // 有 say → 静默
424+
{ stage: 2, text: '展望2027年,增长82.3%。', say: '展望二零二七年,增长百分之八十二点三。' }, // 有 say → 静默; 且 say(19) 比 text(16) 长 → 时间权重的判别样本
425+
{ stage: 3, text: '远溯1897年,一切开始。' }, // 19xx/20xx 之外的四位年份同样要报
425426
] }];
426427
fs.writeFileSync(sp, JSON.stringify(s, null, 2));
427428
fs.writeFileSync(path.join(proj, 'slides', '01-title.html'),
428429
'<!doctype html><html><head><meta charset="utf-8"></head><body><div class="stage"></div></body></html>');
429430
const r = runSkill('plan-timings.mjs', [proj]);
430431
assert.equal(r.status, 0, r.stdout + r.stderr);
431432
const out = r.stdout + r.stderr;
432-
assert.match(out, /第 1 句.*2026年.*逐位读.*say:"二零二六年/, '无 say 的年份必须点名并给出逐位写法: ' + out.slice(-400));
433-
assert.ok(!/第 2 句.*(整读|逐位)/.test(out), '有 say 的年份不得再报: ' + out.slice(-400));
433+
assert.match(out, /第 1 句.*2026年.*逐位读.*say:"二零二六年/, '无 say 的年份必须点名并给出逐位写法: ' + out.slice(-500));
434+
assert.ok(!/第 2 句.*(整读|逐位)/.test(out), '有 say 的年份不得再报: ' + out.slice(-500));
435+
assert.match(out, /第 3 句.*1897年.*say:"一八九七年/, '闸门不得只认 19xx/20xx —— 任意四位年份都逐位读(第 24 轮): ' + out.slice(-500));
434436
// say 原样传递(build-video 的 ASR 清单要消费它, 丢了清单就回退显示形态)
435437
const timings = JSON.parse(fs.readFileSync(path.join(proj, 'build', 'timings.json'), 'utf8'));
436-
assert.equal(timings.slides[0].clauses[1].say, '展望二零二七年,继续同行。');
438+
assert.equal(timings.slides[0].clauses[1].say, '展望二零二七年,增长百分之八十二点三。');
437439
assert.ok(!('say' in timings.slides[0].clauses[0]), '无 say 的句子不得凭空造出 say');
440+
// 时间权重按**发音形态**分摊(第 24 轮): 用实测 tts 反推两个候选 —— 语音字数 15/19/13 vs 显示 15/16/13,
441+
// 第 2 句开口只有按 say 加权才对得上; 按显示形态分摊会系统性估早(字幕窗/ASR 切分跟着偏)。
442+
const tts0 = timings.slides[0].tts;
443+
const expSpoken = (tts0 * 15) / (15 + 19 + 13), expText = (tts0 * 15) / (15 + 16 + 13);
444+
const start2 = timings.slides[0].clauses[1].start;
445+
assert.ok(Math.abs(start2 - expSpoken) < 0.02, `开口应按 say 加权≈${expSpoken.toFixed(3)}(实测 ${start2})`);
446+
assert.ok(Math.abs(start2 - expText) > 0.04, `与显示形态加权(${expText.toFixed(3)})必须可区分, 否则判据失效`);
438447
// 英语面不误报(英语 TTS 自己读 twenty twenty-six, 且正则本身要求"年"字)
439448
s.lang = 'en';
440449
s.slides[0].clauses = [{ stage: 1, text: 'Looking back at 2026, seven stories.' }];

0 commit comments

Comments
 (0)