fix(chat): reconcile events before retrying failed append - #5606
Conversation
Signed-off-by: jackie-cqz <2557911191@qq.com>
1b7837e to
eed044d
Compare
CI Follow-up (2026-10-05)
The shared failure comparison uses upstream main run 37247859756, not an assumption from unrelated test names.
The PR remains scoped to its original repair. Rebase removes obsolete baseline/fixture drift but does not establish a green merge gate. Shared repairs require their own review and qualification; no checks were weakened or suppressed. |
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh
Exact head: eed044de9a2e9bf4bf448a5fdb7877d459734da1; immutable merge base: 796cb29610336bcde0924ffd551102c2457aece2.
动机
在 Chat 中接收流式回答和进度的用户,以及会自动重试事件落盘的执行器。 两个内容和时间完全相同的合法片段已写入,但落盘响应丢失:旧基线重试后从两条变成四条;新 head 仍保留两条独立事件。 同进程可确认或可重新读取的追加重试不会重复已写片段;正常输出、历史行和重连游标保持。 不承诺跨进程重启或全局 exactly-once,不关闭完整 A22/M3、真实模型或 live Lark 验收。
用户看到的是重复的回答片段,而执行器并没有产生新内容。后台每次重试都会把这个问题延续下去,因此修复有实际价值。本次交付是有界改进:确认读回也失败、随后另一进程压缩掉证据时,旧片段仍可能复活;全局恰好一次、进程重启和完整产品旅程未验收。 此处的事件身份只区分某一次追加,不表示用户权限或模型执行权。
改动思路
继续使用既有 Chat 文件存储、文件锁、尾部修复、缓存和后台事件 buffer。每次 append 先产生一个私有 UUID;发生报错后,在仍持有同一个文件锁时确认磁盘里哪些 UUID 已存在。下一次只补没有确认的事件,并从当前磁盘末尾分配序号。内容和时间相同的两个合法调用不能用内容去重,否则会丢掉真实信息;序号也不能充当稳定身份,因为部分写入后另一 writer 可能占用下一个序号。
最强反对理由是新增磁盘元数据、异常路径扫描与返回字典复制成本,而且并未解决全部复合故障。但不做修复会留下实测的重复回答;直接清空 pending 会丢未写内容,盲目全批重试会重复已写内容。当前机制仍在 Python 原生 File IO owner 内,不添加第二个 authority、全局 ledger、CLI 参数或配置步骤,改动与这个真实问题相称。
具体改动
关键代码讲解
ChatSessionStore.append_event(loopx/chat_store.py:1676)把 pending 改为(event, append_id),每次调用独立产生 UUID;公开 event 字典不会因此获得新权限字段。flush_events(1698)在既有 file lock 内识别 uncertain 批次,从磁盘精确匹配_append_id。已确认的 event 保留原event_id/sequence,未确认的才 append;另一 writer 插入后仍按新尾序号排序。错误处理在释放锁前移除已确认 pending,剩余项放在后来入队项前,并使缓存失效,原异常继续抛出。_read_jsonl(83)新增局部 strict-error 选项:普通旧 reader 行为保持;uncertain retry 读回失败会抛出 OSError,不能把不可读当空文件再次写。FileNotFound 仍为空,既有 incomplete EOF 修复继续由原 append helper 负责。events_after(1788)复用原 cursor/bisect/cache,并在公开结果删除 store-only_append_id;HTTP SSE 和前端data/chat.ts仍读取原 public fields。旧无 metadata 行可读,正常 cursor gap 和 reconnect 行为未改变。tests 文件增加五个 real IO/HTTP 负例与 legacy 读回断言,没有新的测试框架。
规范依据为 docs/architecture/rfcs/capable-manager-semantic-handoff-v0.md,固定 revision 796cb29610336bcde0924ffd551102c2457aece2,先于 diff 读取。A22:deferred,本 PR 交付其中可确认/可重新读取的同进程事件重试修复;完整 host/progress/SSE/final-answer 产品旅程和复合故障恢复仍未验收。M3 complete exchange:out_of_scope,该表明确允许 early format/retry fixes 保留原 owner 独立推进,不等于新的 exchange schema、完整结果返回或 M2/M3 qualification 已交付。
对主干的风险
未发现阻塞这个有界改进的回归。独立 Mac base/head 五个实际 suite 为 68 /73 passed;Ruff、规定19文件 mypy、control-plane TS typecheck、actual-diff advisory、full semantic smoke 与 whitespace 通过。唯一 commit 有 DCO sign-off。首次 review 命令猜错测试/semantic 路径和私有 fixture 错用了 store.root(额外 chat 层)均保留,按真实路径和 runtime root 修正后完成对照,未改 PR、删除断言或把初次失败当产品缺陷。没有查询或等待 CI;作者的跨平台/browser/CI 记录只作为历史声明,没有冒充本次独立资格。
同一独立 fixture(SHA 8adba3a002a81be6d0e7225f8679ed6c80f2264fb1eca464c1b4c14e3bb1e71b)通过真实 file lock、真实 append/fsync、真实 _TurnEventBuffer 和 localhost HTTP SSE 执行:九项有界路径在 head 满足独立 oracle,其中七项在 base 失败。两条相同事件丢响应后的 HTTP replay 为 base4/head2;真实后台 retry 为 base2/head1;半条 JSON 记录+另一 writer 后,head 为 first/other/second,base 多出 first。确认连续两次不可读时,head 保留 pending、保持文件大小且不再次 append,恢复后只有一个事件。cache 更新故障、后来入队项和已确认记录被另一 store 压缩后的续接也得到独立读回。正常、旧行、另一个 session 与未来 turn 的完整 public 观测保持,HTTP 重连返回原 cursor 后的正确子集。异常通过 wrapper 在真实落盘后注入,不用 mock store 提供恢复结论。
剩余边界必须保留:同时丢失 append 响应和锁内确认读回,再由另一进程压缩掉旧记录时,head 的 HTTP 输出仍可出现 terminal sequence2 后的旧 delta sequence3;base 也有复活并额外重写 terminal。这与 PR 主文的明确限制一致,不能拿九项通过来宣称全局 exactly-once、进程重启去重或完整 A22 已满足。该复合 fault 为实测 retained gap,不是未执行、被删断言或“全部通过”。既有 Chat event-store/A22 owner 在更强恢复承诺前应解决保留身份与压缩证据寿命;完整产品旅程仍归 M1/M3,不在本 PR 强行增建一套全局 ledger。
语义与 CI 对齐
这是默认 File IO 错误恢复行为的明确修复,不是 default-off capability。新增 metadata 为 local-only optional persisted field,无新 shared vocabulary/权限状态;公开 event v1 不改。每行编码增加48字节(32hex UUID加字段标点),异常路径增加当前 turn 文件读回,公开结果复制仅返回的 row,不扫描 cached cursor 前缀。没有测持续吞吐、模型 tokens 或总体 latency,因此不声称性能数值收益。既有 TS Core work/authority owner 和 quota/scheduler 均未修改;source-qualified 本地检查与远端 CI 状态是分开的事实。
我的整体评价
APPROVE,problem_context=justified_increment。long_horizon 和 user_experience 对这个有界路径均 improved:已经写入的片段不再因普通可确认重试反复追加,真实后台重试及 HTTP 续接更可靠;用户没有新增参数、重复输入或确认步骤。效率的预期正向来自少写、少回放重复事件,代价是48字节/行与局部读回/复制,未量化长期吞吐。剩余复合故障和完整 A22/M3 明确保留,因此不是 whole-RFC/Goal completion。
相邻可维护性审查已落实为复用 existing reader 的 strict option、已有锁/EOF/cache owner,不新增抽象。考虑过提取两处 UUID reconciliation 循环:两个阶段分别承担 readable retry 和 error-confirmation/requeue,本次保持局部可读 IO seam;没有必须先增加的新模块或语言迁移。normal/legacy readback 与 negative sensitivity 支持当前范围,未经证实的 host/model/browser长期结果仍不予认证。
English verdict: APPROVE - eed044d. This useful bounded IO retry repair preserves per-append identity: independent real file/background-buffer/HTTP SSE comparison removes duplicates across seven former failing paths while normal and legacy observations remain equal. Head73/base68 tests and configured static/semantic checks pass. Compound confirmation loss followed by another process's compaction still revives a delta; full A22/M3, restart/global exactly-once and sustained efficiency remain explicitly unqualified.
Goal And Delivered Outcome
mainat796cb29610336bcde0924ffd551102c2457aece2.Author Declaration
docs/architecture/rfcs/capable-manager-semantic-handoff-v0.mdA22 event-identity replay clause, at the base revision. This repairs that store boundary without claiming the whole RFC acceptance.Scope And Continuation
Complete for the demonstrated same-process pending-flush retry. Placement remains in the existing Python native File IO store; there is no new capability, decision owner, ledger or public vocabulary. JSONL adds optional store-only
_append_idmetadata, stripped from publicevents_afterand SSE. Older rows remain readable. Retry flush counts reflect only work settled in that call: a fully confirmed earlier attempt leaves zero pending; partial/no-write attempts retain the missing events. The original append error remains visible.The bounded refactor review keeps the stable identity and strict uncertain-read handling in this owner. There is no global exactly-once or process-restart guarantee. If both append confirmation and lock-held readback fail, and another process compacts away the evidence before a successful retry read, deduplication is still not guaranteed across that compound failure.
Original Validation
1b7837eb3c26865c4d8a0d6803f757e6d312ac2e; synthetic inputs; finished local validation.git diff --checkpass.conversation-activityandconversation-return-continuityscenarios pass using the Node event-transport fixture. These consumer checks preceded the final error-reconciliation refinement; frontend sources/public stream shape are unchanged. The final Python store/SSE backend is independently exercised above.CI Follow-up (2026-10-05)
eed044de9a2e9bf4bf448a5fdb7877d459734da1;mainbase:796cb29610336bcde0924ffd551102c2457aece2. One signed repair commit remains above that base.The shared failure comparison uses upstream main run 37247859756, not an assumption from unrelated test names.
staleinstead ofready. Both failures also reproduce on main at the new base. The completion-receipt digest repair is proposed in #5587, which is still open.unsupported replan_context schema_versionfailure is not covered by those candidate file lists; they must not be treated as a complete CI repair.The PR remains scoped to its original repair. Rebase removes obsolete baseline/fixture drift but does not establish a green merge gate. Shared repairs require their own review and qualification; no checks were weakened or suppressed.
Frontend / Visual Evidence
UI impact: none. Existing event consumers retain the same public fields and cursor semantics. No model/host invocation, Turn restart, authority grant or configuration activation is introduced.
Boundary Checklist