Repository navigation
Help wanted: capable manager & semantic handoff / 强能力管家与语义交接(RFC #4330) #4339
Description
Activity
- addedenhancementNew feature or requestNew feature or requesthelp wantedExtra attention is neededExtra attention is neededdirection/operator-surface-imOperator surfaces, frontend control plane, and bounded IM integration.Operator surfaces, frontend control plane, and bounded IM integration.direction/architecture-evolutionArchitecture evolution and research-incubator work.Architecture evolution and research-incubator work.
on Sep 13, 2026 @huangruiteng 想确认一个可独立交付的 M3 切片是否已有负责人,避免与现有工程 Todo 重叠。
- Area / milestone: M3,异步 worker 结论的送达核验恢复;对应 A8/A10 的有界子场景。
- 问题与证据: 最新 main
201da973c603de7eb4d84b167bd5706819827fdf的manager_context/roundtrip.py::drain在external_write_performed=true、reply_verified=false后写入verification_required,之后的 pump 跳过该状态。已有test_ambiguous_provider_write_is_not_blindly_resent正确保护“不盲目重发”;希望在这个保证之上补只读核验原发送的路径。 - 最小范围: 先限定 provider 已返回消息定位、随后读回暂时失败的情况。由既有 Lark adapter 私有保存最小发送事实,复用 return pump 只读核验原消息,成功后关联原结果的送达回执。无法定位或证据不充分时继续明确未验证,不自动重发、不重跑 worker/model。
- 复用与入口: 复用 fix(manager): recover saved Lark replies after format failures #4314、现有 manager-context roundtrip、Lark
manager_returns/inbox_reply;覆盖 Chat return service、manager-inbox 状态及现有前端/Lark 反馈。fix(chat): scope attached completion messages to turns #4353 的 attached completion transcript 修复与此范围不同。 - 验证计划: 发送成功但读回失败、重启、重复 pump、正文/受众冲突、权限撤销、旧记录缺定位与 provider 离线;保持重复外部效果为零,并补授权测试入口的真实读回。此前固定基线
5ce8e9fcf上相关 22 项测试通过;这不是当前 head 的完整验收或真实 Lark 验收声明。 - 不包含: M2 通用请求/session 接管、Todo/lease authority、provider 晋级,以及无法定位发送结果的通用自动恢复。
这个窄范围能由我承接吗?如果已有 M3 Todo 或正在实施的重叠工作,请指向其公开关联项或建议边界;我会复用已有任务,不另建平行账本。此留言先确认分工,尚未开始实现。
English: Could I take this bounded M3 slice: read-only reconciliation of an asynchronous worker reply whose provider message locator is known but delivery readback failed? It preserves the existing no-blind-resend guarantee and reuses the current result/return owners. Please point me to the existing M3 work item or any overlapping implementation before I start.
@LIHUA919 可以承接,而且建议把它扩成一个仍可独立合并的 M3 delivery verification vertical slice,不只停在 pump 的一个分支。
建议你负责的范围
-
Provider-neutral 送达核验事实
- 在现有 result/return owner 中定义最小、可持久化的消息定位与只读核验证据;不要新建第二套 outbox 或任务状态。
- Lark adapter 负责保存 provider 返回的最小 locator,并把读回结果关联到原 result/delivery identity。
- 兼容旧记录:有可靠 locator 的可以恢复;无 locator、locator 冲突或证据不足时保持 explicit unverified,不猜测、不重发。
-
幂等恢复状态机
- 覆盖
external_write_performed=true && reply_verified=false后的后续 pump、进程重启和重复 pump。 - 核验成功后只补 verified delivery receipt;权限撤销、provider offline、正文/受众不一致、消息不存在等情况保留可行动且真实的失败状态。
- 外部重复效果必须为 0;不得重跑 worker/model,也不得用“再发一次”代替 reconciliation。
- 覆盖
-
同源用户可见状态
- 让 Chat return service、manager-inbox CLI/readback、现有 frontend 和 Lark 投影读取同一组 queued / verification_required / delivered / explicit-unverified 事实。
- 覆盖 RFC A8–A10,包含 A9 的长正文/截断 trailer 恢复;A17 只做“已发生外部发送后的 return-pump/进程重启”这一半,不承担 session takeover。
- 提供 deterministic fixtures、重启/重复投递/权限/冲突测试和 packaged frontend 状态证据。真实私有 Lark canary 可由我在你的精确 PR head 上补做,不要求你接触私人群或凭证。
仍由我负责 / 非目标
- M2 通用 request/assessment/result 事务替换、Todo/lease/Goal authority、worker takeover;
- 未知 locator 的通用恢复、其他 provider 推广与 provider promotion;
- 将 M2 producer 接到新 receipt contract、原逻辑受众的跨 session/host 集成;
- 最终真实 Lark canary、卡片交易审批链,以及 M3/M4 全入口和三 Goal 验收;
- fix(chat): scope attached completion messages to turns #4353 的 attached completion transcript、fix(lark): preserve context without granting Turn authority #4319 的历史 context-only precedence 不并入本 PR,除非先在本 issue 协调。
请基于当前 main,先贴一个简短 ownership map(现有 writer/reader、拟改文件、兼容边界)再开实现 PR;PR 描述绑定 A8/A9/A10 的具体测试。你的切片可以与 M1/M2 并行,不需要等待 #4337,但不要宣称通用 handoff schema 已交付。开 PR 后在这里贴链接,我会避开同一代码面并做 exact-head review/集成。
English: Please take an expanded but still independently mergeable M3 delivery-verification vertical slice: provider-neutral persisted locator/verification facts, Lark read-only reconciliation, idempotent restart recovery, and truthful shared CLI/frontend/Lark delivery states. Preserve zero duplicate external effects and explicit unknowns. Generic M2 transaction/authority, unknown-locator recovery, producer integration, private canary, and full M3/M4 qualification remain with me.
-
@huangruiteng Ownership map for the approved M3 delivery-verification slice, based on
origin/main@ced85a63b2e390fbbe75c5e8da6b707307b01546.Existing owners and readers
- Result / return writer:
loopx/capabilities/manager_context/roundtrip.py.report()owns the immutable audience-ready result;drain()is the sole writer of its delivery state;reply_status()is the provider-neutral read model. - Lark provider adapter:
loopx/extensions/lark/manager_returns.pyrevalidates live route and audience authority;loopx/extensions/lark/inbox_reply.pyowns the external send and provider readback. It currently has the provider message id during the failed verification attempt, but does not return a durable locator. - Readers:
manager_context/tracking.py::query()andmanager-inbox statusconsumereply_status(); the Chat transcript consumes the stablemanager_followupid; Lark consumes the adapter result. The existing frontend reads the same persisted Chat/manager projections and should receive only public-safe status/error fields.
Proposed change and placement
- Capability id: keep
manager-context; no new built-in capability. Provider: existing Lark extension. This is the nearest existing owner because the shipped caller outcome is “return the delegated result to its original audience,” while Lark-specific locators and readback commands remain adapter-private. - Add a small provider-neutral persisted attempt/verification contract beside the existing reply delivery state in
roundtrip.py; keep one delivery writer indrain(). Private locator payloads are accepted only from the provider adapter, validated/bounded, and never projected byreply_status(). - Extend
inbox_reply.pyto return the minimum Lark locator when a send occurred and readback was inconclusive, plus a read-only “verify this exact prior reply” path that performs no send. Compose that verifier inmanager_returns.pyafter the same live route/authority checks. - State transitions remain
queued/retry_pending -> verification_required -> delivered | explicit_unverified. A recoverable known locator is rechecked after restart; missing legacy locator, conflict, revocation, missing message, content/audience mismatch, or insufficient evidence stays explicit and never resends.
Expected changed files
- Runtime:
loopx/capabilities/manager_context/roundtrip.py,loopx/extensions/lark/manager_returns.py,loopx/extensions/lark/inbox_reply.py. - Shared readback/docs:
loopx/capabilities/manager_context/README.md;tracking.py/CLI only if the current projection cannot carry the truthful public-safe state without change. - Frontend: existing personal-workspace Chat/readback components and i18n only if the live schema needs a new display mapping; reuse existing primitives and packaged assets.
- Validation:
tests/test_manager_context_roundtrip.py,tests/extensions/test_lark_manager_returns.py, focused Lark reply tests, manager-inbox projection tests, and packaged frontend contract/browser smoke.
Compatibility and non-goals
- Existing
delivered,retry_pending,superseded, and legacyverification_requiredrecords remain readable. Legacy records without a trustworthy locator become explicit unverified and are not rewritten into guessed proof. - Preserve zero duplicate external effects, the immutable result identity, current return route/audience checks, and independent transcript idempotency. No model/worker rerun.
- No M2 WorkRequest/Assessment transaction, Todo/lease/Goal authority, session takeover, unknown-locator recovery, new outbox/task database, provider promotion, or private canary in this PR.
I will implement this on
codex/m3-delivery-verification, bind the PR to A8/A9/A10 tests, and post the exact head here for review.- Result / return writer:
- 也伴随可以做些ts重构 | | huangrt01 | | ***@***.*** | ---- 回复的原邮件 ---- | 发件人 | ***@***.***> | | 发送日期 | 2026年09月14日 14:56 | | 收件人 | huangruiteng/loopx ***@***.***> | | 抄送人 | huangruiteng ***@***.***>, Mention ***@***.***> | | 主题 | Re: [huangruiteng/loopx] Help wanted: capable manager & semantic handoff / 强能力管家与语义交接(RFC #4330) (Issue #4339) | LIHUA919 left a comment (loopx-project/loopx#4339) @huangruiteng Ownership map for the approved M3 delivery-verification slice, based on ***@***.*** Existing owners and readers Result / return writer:loopx/capabilities/manager_context/roundtrip.py. report() owns the immutable audience-ready result; drain() is the sole writer of its delivery state; reply_status() is the provider-neutral read model. Lark provider adapter:loopx/extensions/lark/manager_returns.py revalidates live route and audience authority; loopx/extensions/lark/inbox_reply.py owns the external send and provider readback. It currently has the provider message id during the failed verification attempt, but does not return a durable locator. Readers:manager_context/tracking.py::query() and manager-inbox status consume reply_status(); the Chat transcript consumes the stable manager_followup id; Lark consumes the adapter result. The existing frontend reads the same persisted Chat/manager projections and should receive only public-safe status/error fields. Proposed change and placement Capability id: keep manager-context; no new built-in capability. Provider: existing Lark extension. This is the nearest existing owner because the shipped caller outcome is “return the delegated result to its original audience,” while Lark-specific locators and readback commands remain adapter-private. Add a small provider-neutral persisted attempt/verification contract beside the existing reply delivery state in roundtrip.py; keep one delivery writer in drain(). Private locator payloads are accepted only from the provider adapter, validated/bounded, and never projected by reply_status(). Extend inbox_reply.py to return the minimum Lark locator when a send occurred and readback was inconclusive, plus a read-only “verify this exact prior reply” path that performs no send. Compose that verifier in manager_returns.py after the same live route/authority checks. State transitions remain queued/retry_pending -> verification_required -> delivered | explicit_unverified. A recoverable known locator is rechecked after restart; missing legacy locator, conflict, revocation, missing message, content/audience mismatch, or insufficient evidence stays explicit and never resends. Expected changed files Runtime: loopx/capabilities/manager_context/roundtrip.py, loopx/extensions/lark/manager_returns.py, loopx/extensions/lark/inbox_reply.py. Shared readback/docs: loopx/capabilities/manager_context/README.md; tracking.py/CLI only if the current projection cannot carry the truthful public-safe state without change. Frontend: existing personal-workspace Chat/readback components and i18n only if the live schema needs a new display mapping; reuse existing primitives and packaged assets. Validation: tests/test_manager_context_roundtrip.py, tests/extensions/test_lark_manager_returns.py, focused Lark reply tests, manager-inbox projection tests, and packaged frontend contract/browser smoke. Compatibility and non-goals Existing delivered, retry_pending, superseded, and legacy verification_required records remain readable. Legacy records without a trustworthy locator become explicit unverified and are not rewritten into guessed proof. Preserve zero duplicate external effects, the immutable result identity, current return route/audience checks, and independent transcript idempotency. No model/worker rerun. No M2 WorkRequest/Assessment transaction, Todo/lease/Goal authority, session takeover, unknown-locator recovery, new outbox/task database, provider promotion, or private canary in this PR. I will implement this on codex/m3-delivery-verification, bind the PR to A8/A9/A10 tests, and post the exact head here for review. — Reply to this email directly, view it on GitHub, or unsubscribe. You are receiving this because you were mentioned.Message ID: ***@***.***>
Implemented the approved M3 delivery-verification slice in #4372: #4372
- Base:
6c1a4d2cc37280a1d652b4bf67afd9c7d69ce19e - Exact head:
a883609c13826143e1d159c5a11f810171f14c49 - Runtime: provider-neutral persisted attempt facts, Lark locator capture before readback, read-only reconciliation after restart, fail-closed legacy/conflict/revocation/missing-message handling, and zero resend/model rerun.
- Shared state: Manager CLI/readback and Chat expose the same public-safe delivery statuses; provider locator remains private.
- UI: packaged Personal Workspace updates one returned message from “verification required” to “verified after recovery” on desktop/mobile.
- Validation: 326 relevant Python tests; Dashboard build; packaged
navigation-sorting,chat-recovery,typed-actions; Ruff/diff; maintainability ratchet; premerge canary 18/18, no manual holds.
The PR records the two unchanged base failures separately and leaves the private Lark canary with you as agreed. M2/session takeover/provider promotion remain out of scope. Requesting exact-head review and the owner-side private canary when ready; I will address actionable findings on the same branch.
- Base:
Follow-up to the TS-refactor suggestion: #4372 now includes a bounded typed collaboration contract with a real Python caller.
- New exact head:
544915263e43d248ca50db9a77b1067b838ec68f control_plane/collaboration/return_delivery.tsnow owns exact provider-neutral attempt validation and verification classification.- Python retains result/delivery persistence, locking and retry orchestration; Lark retains provider locator interpretation/readback. No second writer or M2 framework was introduced.
- The PR body now includes TS migration economics and removal gates.
- New TS contract + runtime handler tests: 12/12; typecheck passed; final premerge canary 18/18; 326 Chat/manager/Lark tests passed.
The owner-side private Lark canary and exact-head review remain the next acceptance steps.
- New exact head:
Latest exact head after CI installation repair:
f15b6eb4b77da7ee287839a115c0e5313e5e00fe.The first head exposed a real wheel packaging gap for the new TS collaboration subpackage. The fix adds the package-data boundary; an isolated wheel build/install successfully invoked both new runtime methods. On the new CI run, Release Artifacts and
stage2c (installed 0)now pass. Remaining jobs are still running; the LoopXgithub-pr-4372monitor will handle CI/review/private-canary changes and repair actionable findings on the same branch.Repair receipt: #4372 (comment)
M3 delivery-verification slice update: #4372 now merges current main and rebuilds the packaged chat assets on the merged sources.
- New exact head:
be8c25a6f58048fdabb837d6dcaef689d231357c - Root cause of the red frontstage-pages build was the stale PR base (main refreshed dashboard sources and packaged assets after
6c1a4d2c); CI's pull_request build compiles the merge, so neither side's committed assets matched. Merging main and rebuilding resolves it; the new generation reproduces the exact CI clean-build bytes. - Repair receipt with full validation: feat(manager): reconcile uncertain delegated result delivery #4372 (comment)
- Note:
test-shard (3)red is a date-parameterizedtest_delivery_responsecase that reproduces identically on exactorigin/main@a1c2516f4; it predates this slice.
The reviewed projection-completeness fix is unchanged. Re-review and the owner-side private Lark canary remain the next acceptance steps.
- New exact head:
@huangruiteng 想申请承接 #4311 在 #4330 下的一个 bounded M2 vertical slice。先贴 ownership / compatibility / acceptance map,确认没有和现有工程 Todo 重叠后再开实现 PR。
当前证据
- 基线:
origin/main@f2873c72a15b1a25d79e053da9a9d561d74d977e。 - fix(scheduler): dispatch executor-excluded handoffs instead of repolling origin monitor #4311 仍 open、无 assignee;原实现 fix(scheduler): dispatch stranded peer handoffs #4312 已 closed/unmerged,关闭说明明确写了 defect remains real,并要求把
dispatchable/no_eligible_peer、legacy alias、restart-safedispatched -> read -> claimed、stale claim、canonical Todo claim readback、duplicate ingress 与 effect-after-crash 迁入 M2。 - 当前
manager_context仍由 Python 分散拥有入口和回执:__init__.py::deliver/normalize_request、tracking.py的 read/decision/link、roundtrip.py的 decision/conclusion/return。请求 identity 仍是现有 context-handoff 形态。 - 当前
control_plane/collaboration只有已经交付的return_delivery.ts;尚无 WorkRequest / SemanticContext / DispatchAttempt / Assessment / Observation 的 Core transaction,也没有 fix(scheduler): dispatch stranded peer handoffs #4312 的agent_handoff.py。 coordination/todo_continuation.ts/ feat(todos): explicit revision-guarded cross-agent session continuation (Stage A) #4094 继续拥有显式 Todo continuation 与 transfer grant;agents/directory.py只提供受限的peer_agent_directory_v0发现读模型。新 slice 不复制 Todo/lease/Goal authority。- 已做 numeric、semantic 和 open-PR file-overlap 检查:没有开放 PR 修改
manager_context/**、control_plane/collaboration/**、todo_continuation.ts或agents/agent_scope.py。feat(handoff): complete context and strict receiver independent of review packets #4444 是 review-packet 超预算分片,不修改 request/assessment/result 或 work authority。
建议负责的范围
-
一个真实的 M2 request transaction
- 在 TypeScript Core collaboration boundary 定义并执行一个版本化 request/context/dispatch/assessment/observation transaction;同一 event replay 返回原 receipt,identity 相同而 payload 不同则 conflict。
- Python
manager_context只保留现有 runtime/channel I/O 与兼容读取,不再拥有新的通用状态转换;不新建 task database、daemon 或 provider。
-
fix(scheduler): dispatch executor-excluded handoffs instead of repolling origin monitor #4311 作为第一个真实 adapter/caller
- material monitor 产生 executor-excluded
independent_handoffsuccessor 时,创建一个delegated_workrequest/dispatch obligation。 - 通过现有 peer directory 选择一个不同的已注册 eligible peer;无 eligible peer 时投影 typed blocker,不创建 user gate。
- receiver assessment 与 Todo claim/adopt 分开;真正 ownership 变化仍调用现有 claim/continuation owner并回读 canonical Todo。
- material monitor 产生 executor-excluded
-
兼容与第二消费者边界
- 把 fix(scheduler): dispatch stranded peer handoffs #4312 的 tuple-derived dispatch identity 只保留为 legacy alias;新 round 使用 request id/revision,不把 same-Goal/unclaimed/source-excluded 写成通用 admission 规则。
- 用现有 feat(todos): explicit revision-guarded cross-agent session continuation (Stage A) #4094 continuation adapter 做 worker→worker 对照 consumer,证明两个入口共享 request/assessment invariants;cross-Goal consultation 和 pre-Todo request 在本 slice 中至少进入 contract/negative fixtures,但不扩大 effectful takeover 权限。
验收与验证
- 绑定 RFC A4–A7 中与本 slice 相关的子集:缺 convenience profile 仍能发现正确 receiver;correction/rejected approach 不丢;manager→worker 与 worker→worker 使用同一 contract;duplicate ingress、late receipt、concurrent claim、crash-after-effect 不产生重复 effect。
- 保留 fix(scheduler): dispatch executor-excluded handoffs instead of repolling origin monitor #4311 的 focused cases:
dispatchable_by_peer/ executor-excluded visibility、typedno_eligible_peer、restart、stale claim、canonical claim readback、duplicate dispatch/read/claim、effect-after-crash reconciliation。 - 先做 base/head differential characterization,再跑 TypeScript contract/effect-runtime tests、Python adapter/CLI quota and Turn tests、process-restart fixture、
git diff --check、public-boundary scan 与 risk-based premerge canary。
非目标
- 不改 Goal/Todo/Vision/lease 的权威归属,不推广 SQLite/PostgreSQL/NoKV,不做 cross-host effectful takeover。
- 不改 M3 return-delivery、frontend/Lark transport 或 feat(handoff): complete context and strict receiver independent of review packets #4444 handoff-fragmentation。
- 不从关闭的 fix(scheduler): dispatch stranded peer handoffs #4312 直接 cherry-pick production owner;只复用其已记录的兼容义务与 fixtures。
请确认这个 first M2 transaction + #4311 adapter slice 是否可以由我负责;若已有 canonical Todo、正在实施的 owner,或希望先进一步缩小 transaction boundary,请指向对应公开项。我会在确认后从最新
origin/main建独立codex/worktree,并在实现前贴最终 file/owner map。English summary: Requesting ownership of a bounded first M2 transaction with #4311 as the first real adapter. The slice moves new generic request/assessment transitions into the typed Core collaboration boundary, keeps Python as runtime/channel compatibility I/O, preserves existing Todo/lease/Goal authority, migrates #4312 semantics only as legacy compatibility, and qualifies manager-to-worker plus the #4094 worker-to-worker adapter without claiming full cross-Goal or effectful takeover support. No open PR currently overlaps the actual owner files.
- 基线:
Summary / 目标
Help build a capable LoopX manager that can investigate with ordinary host tools, hand work to another agent without losing its meaning, and return the result automatically.
The bilingual RFC landed in #4330. This is an implementation invitation, not a feature-release announcement: the document records proposed behavior and acceptance requirements, not completed M1–M4 qualification.
main; check existing work before choosing a slice.Inherited acceptance: PR investigation and responsible routing (#4305)
Product triage on 2026-09-24 consolidates #4305 here; this is not a claim that its original Lark journey is fixed. The superseded dedicated-provider/automatic-evidence-handoff proposal is retired. Keep the user outcome under existing M1/M2/M3 owners:
Existing foundations: #4303 (configuration source), #4337 (private-owner host profile), #4675 (semantic collaboration), #4372 (uncertain-return verification). These do not prove the exact missing-PR/two-role/Lark scenario. Reuse that scenario in a bounded existing collaboration slice; no separate PR-only capability, scheduler, mandatory profile allowlist or new task ledger is requested. RFC §6's live investigation/routing obligation is retained here, not waived by closing #4305.
中文:#4305 的独立旧方案停止推进,真实价值保留为本 issue 的验收场景:授权内直接查证、必要时按实际职责交接、结果回原会话。双角色/唯一不适任接收方和飞书真实旅程尚未验收,不能标为已修复。
Contribution areas
Best starting points: one real M1 investigation journey, an independent M3 reply-recovery bug, or a reproducible acceptance fixture. M2 is a cohesive transaction replacement, not a collection of disconnected new fields or an unused framework.
How to participate
Comment with:
Use the thread to coordinate overlap before starting a large refactor. Contributions can be implementation, design review, real-runtime validation or public/synthetic reproductions. A whole milestone is not a prerequisite for participation.
Completion and boundaries
中文:一起把这条协作链做完整
我希望管家能在现有授权内充分调查,合适的短工作自己完成,持续或专业的工作交给对应 Agent;交出去的不只是一句话,还包括背景、纠正、已排除方案、证据和期望回报。接收方判断如何接续自己的计划,做完后,用户不用再追问“结果呢”。
欢迎参与四个方向:M1 本机调查能力、M2 通用语义交接、M3 自动回传与前端/飞书可见性、M4 真实旅程验收与旧路径退役。 M1 和独立的正文恢复可以先做,不必等完整 handoff 重构;M2 需要围绕一个完整事务交付,保留已有工作状态权威。
不必一次认领整份 RFC。可以选一个具体问题,贴出拟做范围、关联的已有任务/PR、运行时、验收 ID 和验证方法,再协调实现。真实失败案例、反例和设计审阅,同样是有价值的贡献。
本 Issue 只汇总参与入口和公开结果,不复制另一份工程任务权威;RFC 已合入,不等于其中能力已上线。