Repository navigation
fix(runtime): preserve Node probe failures and accept slow compatible startup - #5997
Conversation
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.
Request changes conclusion (author-owned PR; GitHub blocks formal self-review)
REQUEST_CHANGES:慢兼容 Node 的可用性改进成立,但 deep doctor 的后续启动检查遗漏了共享恢复建议。 独立审查完整 9 文件,head d3b347a108922f35cd1b9df05e8ce24582431bd6,base 3ed5d6bc88ffaaa05f2b1bc6a98ed5d921379fd1。
动机
使用 CLI 或 App 的用户,在宿主繁忙或 Node launcher 启动较慢时,原来的两秒检查会把兼容 Node 误报为不可用,要求升级,并让重试重复同一失败。真实对照中,同一个延迟 2.6 秒的 Node 24.21.0 在 base 约 4.08 秒后拒绝,在 head 约 5.97 秒后完成实际请求。该切片改善启动可用性及失败解释;它不证明整项工作吞吐、持续运行或宿主负载根因。尚未解决的是 deep doctor 后续版本检查失败时的准确恢复建议。
改动思路
保留现有按需启动的 TypeScript 运行时和 Python 物理进程适配边界,复用已有 15 秒预算和 22.22.3 门槛,将观察与恢复文案集中到 runtime/node_probe.py。没有增加持久缓存、配置开关、服务管理步骤或第二个 Goal/Todo 决策源。新增的 NodeProbeOutcome 是进程事实的本地枚举,旧公开 readiness 状态仍由它投影;共享诊断码则明确扩展了协议。
该边界比单纯把 timeout=2 改大合理:仅增加等待仍会误报权限、启动、退出和输出失败;把 Node 探测交给尚未启动的 TS 运行时又会循环依赖。当前 PR 的合理边界是启动探测、准确诊断及既有入口恢复,保留 runtime fingerprint、SQLite 资格、request identity 和 authority owner。相关小重构已经移走重复版本与恢复知识,维护成本可接受;但 doctor 的 deep catch 仍保留第二套泛化恢复文案,需要同 PR 收拢。
具体改动
审查依据是改动前的 TypeScript migration RFC,spec_revision=3ed5d6bc88ffaaa05f2b1bc6a98ed5d921379fd1,没有用本 PR 新增规范反证自身正确。采用原章节编号作 criterion_id:
- 2.2:自动发现、启动且无需用户管理服务,implemented。快、慢真实请求以及 warm 同 PID 复用均通过。
- 2.3:迁移后的控制面 effects 保留 typed owner,implemented。新 Python 模块只观察宿主进程;18 项真实 SQLite CLI、发布和请求范围测试通过,未变更权限、writer 或 receipt。
- 7:启动前检测 Node、给出准确 remediation,并维持统一 lifecycle/公开安全投影,not_met。初始探测和 guided 恢复通过,但下面的 deep 第二次探测仍给出错误的重装建议;Node 门槛、原 readiness 状态和 raw-output 隔离已验证。
关键代码讲解
probe_node用实际退出/异常事实分类,只有缺失或已解析的低版本推荐安装。新增 130 行涵盖本地结果、稳定诊断、共享恢复和真实版本调用;既有版本匹配规则保留。_node_executable把权限拒绝映射到原 host permission 错误,其余 Node 失败成为EffectRuntimeNodeProbeError;effect_runtime_request不再自动重复这些 pre-dispatch 失败,其他 startup 和 safe transport 重试仍在原 owner 中。collect_effect_runtime_readiness初始检查使用共享诊断,成功后 deep 模式会实际启动运行时。后者还能触发第二次版本检查,正是当前遗漏分支。_effect_runtime_startup_recommended_action先复用 Node 恢复规则,再保留已有非 Node 启动建议。
完整 diff 为 399 additions / 69 deletions:3 个生产路径、4 个测试路径、2 个文档路径。新 186 行 Node 测试覆盖分类、一次拒绝、取消、真实慢启动及直接 probe child 回收;其他三个测试文件适配新返回结构。协议和 RFC 披露等待预算与诊断变化,没有改写 provider 默认或业务验收。当前测试缺少“第一次成功、deep 后续启动失败”的组合。
对主干的风险
[P2] deep doctor 必须复用后续 Node probe 的恢复建议。 代码位置:真实 launcher 第一次 --version 返回 v24.21.0,第二次等待超过 15 秒。doctor --deep 的实际冷启动抛出 EffectRuntimeNodeProbeError;catch 保留 node_probe_timeout,却返回“若仍失败,reinstall LoopX and verify Node.js”,没有使用 node_probe_remediation 的 host load/launcher 建议。这让兼容性未知再次变成安装排查,也违反本 PR startup/doctor/guided 统一恢复的目标。最小修复:该 catch 在保留 host-permission 优先级后,对 Node 诊断复用同一 remediation,再回落到其他 runtime 错误文案。补首检成功、后续 timeout/launch/exit/invalid-version/unsupported 的组合回归,并验证恢复后原请求成功;不能仅继续测试首次探测失败。
独立执行:163 项 runtime/Node/permission/readiness/compile-cache/start-goal/host 测试通过,另 18 项真实 SQLite CLI/publication/request-scope 测试通过;慢 launcher 对照、六类错误的 base/head 诊断对照通过。语义 diff advisory 先于全树语义检查;全树语义、Ruff、maintainability 以及 canary premerge --from-git-diff 的 19 项选定检查和 5 项直接检查通过。质量检查通过没有覆盖上述组合故障,因此不能取代可行动 finding。
初次 wheel 构建因旧 frontend bundle 与当前源码不符失败,保留记录;按原 build:chat 流程重建后 fresh wheel 成功,3 个生产文件的安装字节与精确 head SHA256 一致。安装态 CLI 的快、慢、缺失、低版本、无效输出、非零退出、损坏 launcher 和直接超时 8 个场景符合预期;第 9 个“首检成功、deep 后续超时”场景再次复现相同错误建议,直接 probe child 已回收且无运行时启动。作者报告的两个环境相关红项也保留:本次不可变 base/head 的相同脚本均通过,当前 exact-scope premerge 通过;未据此猜测旧环境根因或声称修复了那些旧失败。没有读取、轮询或等待远端 CI。
直接 probe child 的 timeout/cancel 回收、startup lock 释放和无请求 dispatch 已验证。另一个真实 wrapper 会留下它自行派生的后代:base 与 head 都如此,属于原有 process-tree 清理限制,不列为本 PR 新回归,也不声称它已解决。Windows 真实进程、packaged App/Lark 交互及长期成本未重验;该 PR没有新增其设置或操作入口。
语义与 CI 对齐
本地 NodeProbeOutcome 不应复用仅因同有 permission_denied 文本而被 advisory 提示的 settlement 词表,两者 producer 和含义不同。公开诊断码是明确的 vocabulary extension,readiness schemas/status 和 Node floor 保持;故障文案按 Goal 无关的物理观察描述。探测拒绝是机器执行的 precondition,恢复建议是 advisory,均不授予能力或 authority。当前未新增 opt-in/default-off 能力;warm/正常启动和原公开结构的成对对照通过,明确的默认变化是冷探测容忍度及失败重试。
我的整体评价
long_horizon:improved(有界条件),慢兼容运行时现在能完成请求,warm 复用保持,失败探测只拒绝一次,避免重复无效启动;15 秒等待是明确披露的成本,未证明全任务吞吐提升。user_experience:not_yet_proven,首检失败的解释改善,但 deep 后续失败仍误导恢复,需要上面的最小修复及组合测试。交付判断是有实际价值的增量,当前准确恢复旅程仍有缺口;不以测试数、规范新增或 recall 文本代替体验验收。
仓库经验建议仅采纳“沿普通任务核验恢复后的可用结果”,本次已实际请求、读回 SQLite 资格、检查后续失败并停止自有运行时;历史评审结论及长期收益不继承。架构 assessment 为 retain,当前边界限定启动诊断和恢复;将 deep catch 的共享 remediation 补齐即可,不需要新增框架、策略开关或并行 Python 决策。Runtime/control-plane 合并交维护者,当前结论不授予 merge 或本机升级权限。
English verdict: REQUEST_CHANGES - head d3b347a. Slow compatible startup is independently verified; preserve the shared Node remediation when a later cold-start probe fails inside deep doctor. 181 focused/real-SQLite tests and current risk-selected premerge pass, but do not exercise that missing recovery branch.
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
Signed-off-by: LoopX Agent <337587101+loopx-agent@users.noreply.github.com>
|
Addressed the later-probe recovery finding from review 5464121263 in head The previous head fails five combination regressions and a real initial-success/later-timeout case. The updated source passes 217 focused tests and five additional real SQLite CLI tests. Five fresh installed CLI cases exercise initial success followed by timeout, launch, exit, invalid-version or unsupported-version failure; all recover to the original real request and deep semantic readiness with qualified SQLite after launcher repair. Probe reaping, startup-lock release and owned runtime shutdown pass. Risk-selected premerge is 19/19 plus five direct checks; Ruff and the public-boundary scan pass. No manual holds or skipped checks in this current risk set; real Windows, packaged App/Lark and sustained whole-task performance remain untested. Earlier environment-sensitive failures are retained without claiming this patch fixes them. Changed surfaces: Effect transport/readiness, regression tests, host diagnostic contract and the existing runtime/install RFC checkpoint. The bounded future-facing pass reuses the existing remediation owner across the previously missed catch. Requesting independent review of the full updated head; this comment is a repair/validation report, not approval or merge authorization. |
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewed exact head 2ed1eff7ea9f7360979f87654b9173689e167fd1 against base 3ed5d6bc88ffaaa05f2b1bc6a98ed5d921379fd1. APPROVE,本次完整重新判断,不继承旧批准。
动机
CLI 或 App 用户启动任务时,繁忙宿主或慢 Node launcher 会让兼容运行时被误报为不可用。
原先两秒探测可能误判慢 Node;首次探测成功、随后超时时还建议重装 LoopX。现在冷探测使用既有十五秒预算,按真实失败给出恢复建议,修复 launcher 后能重新完成深度检查。
本次 217 项测试通过;真实 CLI 对照复现旧版本的重装建议,当前版本提示检查宿主负载和 launcher,超时子进程及锁被清理,修复后 doctor --deep 和 SQLite 资格均通过。
本 PR 不提升宿主权限、不改变 Node 版本门槛、Goal/Todo/额度规则或现有账户,也不证明长期吞吐、Windows 进程行为或任意 launcher 后代已清理。
改动思路
Node launch observations and remediation belong in one Python physical-process adapter because the TypeScript runtime cannot inspect the prerequisite needed to launch itself. Goal, Todo, quota and Effect decision authority remain in their existing TypeScript owners.
The complete current slice is bounded Node startup observation, accurate failure projection through startup/guided/deep doctor, and recovery to a real running SQLite-qualified runtime; it adds no persistent version cache, optional strategy, service manager or capability.
单改 timeout 更小,但仍会把权限、启动、退出及输出故障混为版本不足。当前小重构收拢已有重复知识;十五秒拒绝延迟是披露的有界成本,不是吞吐提升证明。仓库经验建议仅用于实际恢复旅程,没有继承历史 verdict 或宣称模型记忆效果。
具体改动
完整 B..H 为 9 文件 +499/-70,新增本地进程观察 Enum/不可变结果和共享安全恢复文案,既有 runtime/start-goal 消费它;没有缓存、服务管理器、配置开关或并行领域决策源。对照旧 review head d3b347a108922f35cd1b9df05e8ce24582431bd6,新增修复在 deep catch 保留已知 Node cause 的 remediation,并加入首次 ready/后续失败的组合与真实恢复测试。
关键代码讲解
probe_node:物理子进程使用既有十五秒预算;只有缺失或解析出的低版本支持安装建议,timeout/launch/exit/invalid 明确表示兼容性未知。直接子进程 timeout/cancel 由现有 subprocess 生命周期清理。collect_effect_runtime_readiness:首检和 deep 后续 cold start 都保留同一 cause-specific 建议,host permission 优先;非 Node startup error 仍走原 fallback。_effect_runtime_startup_recommended_action:复用诊断 owner,保留 guided retry 和同 registry/Goal/Agent/Turn、guard 前不改变 authority/不 spend 的完整条件。
规格依据:docs/architecture/rfcs/typescript-control-plane-migration-v0.md,spec_revision=3ed5d6bc88ffaaa05f2b1bc6a98ed5d921379fd1,在读实现前读取其接受文本。2.2 implemented:自动启动既有运行时,不要求新服务;2.3 implemented:Python 仅做物理进程适配,TS 领域决策/effect owner 不变;7 implemented:无法资格化时在工作前拒绝并提供准确恢复。PR 新增 checkpoint 是待评报告,不用它定义自己的验收。
对主干的风险
当前九个相关/相邻测试文件 217 passed,包括真实慢兼容 Node 的请求、warm PID 复用、直接 child 超时/取消清理、startup lock 释放和 real SQLite readiness。Ruff、TS typecheck、DCO 四 commits、B..H diff check、advisory 后的完整 semantic smoke 通过;Pydantic 的既有 unresolved-annotation warning 保留。
本轮用当前组合测试运行在不可变旧 head:5 failed, 2 passed,失败均是首次成功、后续 timeout/launch/exit/invalid/unsupported 后丢失建议。另用相同真实 loopx doctor --installation-only --deep 和临时 PATH launcher 在旧/current 两版本执行:首检成功,第二次版本 child 超时。旧版说 reinstall LoopX,当前说检查 host load/launcher,诊断保持 node_probe_timeout;直接 child 已退出,锁释放。随后只修复 launcher,同一命令在两版本均 top-level ok=true,semantic probe passed、SQLite qualified。这里验证的是修复后的实际可用结果,不把 refusal receipt 当恢复。
正常/warm/Node floor 和公开 schema/status 保留;冷探测 2→15 秒、准确故障投影及失败不重复 dispatch 是明确的默认修复。NodeProbeOutcome 是本地进程事实,不能仅因 permission_denied 同名而复用 settlement 词表;共享诊断码在协议中扩展。无新增 opt-in/default-off 功能。探测前置条件是机器执行,恢复建议是 advisory,不授予宿主或 Goal 权限。
最初验证命令引用不存在的测试文件,零测试执行;已按实际文件清单修正,完整失败记录保留,没有记为产品回归。旧报告的 wrapper 后代清理限制仍不声称已解决,真实 Windows、长期成本、packaged App/Lark 交互未重验;本 PR 未新增这些入口。未查询、轮询或等待 CI。
我的整体评价
APPROVE。原慢兼容启动问题和旧 review 的后续准确恢复缺口,现在均有实际入口和消辨对照。未来重构检查已在同一 owner 收拢重复探测与文案,无需更多框架或领域迁移。该审查不授予 runtime/control-plane 合并或本机升级权限;平台 review aggregate 和维护者合并资格另行读回。
English verdict: APPROVE — exact head 2ed1eff. Fresh whole-PR review, 217 local tests and static/semantic checks pass; five old-head later-probe failures and the same physical public CLI timeout/recovery workload independently distinguish the repaired remediation. No CI or long-soak/all-platform claim.
|
Follow-up on exact head I inspected the two completed failing checks:
No source change, CI restart, check override or merge was performed. The existing runtime/diagnostics boundary remains the owner; no new abstraction or policy was added. Preserve the independent review, the remaining CI distinctions and the maintainer merge/adoption gate separately. |
A compatible Node launcher that took longer than two seconds was rejected as
node_unavailableand told to upgrade. Startup now uses the existing 15-second readiness budget for its version observation and distinguishes timeout, permission denial, launch failure, unsuccessful exit and invalid version output. Startup, doctor and guided start-goal share the diagnosis and recovery; only missing or parsed unsupported versions recommend installation. Deep doctor also preserves recovery when the initial check succeeds but its later cold-start probe fails, with host permission taking precedence and other runtime errors retaining their existing fallback.The physical observation stays beside the existing Effect transport. Its enum is local process observation; public readiness states and the Node 22.22.3 floor remain. Diagnostic codes are explicitly extended in the host contract. The companion refactor removes duplicate probe/remediation knowledge; the review repair reuses that same owner. Failed pre-dispatch Node probes return once for caller recovery. Warm reuse, cancellation, request identity, SQLite qualification and authority fencing keep their existing owners.
Validation:
This is bounded startup availability and recovery evidence. It does not establish shared-host contention, sustained runtime acceptance, or whole-task throughput improvement. Real Windows and packaged App/Lark journeys are not newly qualified. Independent review of the updated exact head and maintainer merge remain required.