diff --git a/AGENTS.md b/AGENTS.md index 889c0be..2653b4a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -15,6 +15,15 @@ Derived from `config/harness.json`. Policy lives in config, not in vibes. - Live review / Action dogfood: **at most one** Anthropic key. Never require two keys to run the stack. Never put real keys in git. - CI self-review (`.github/workflows/self-review.yml`) uses a **key-gate** job (GitHub forbids `secrets` in `jobs..if`) so an empty secret skips the review job; keeps `fail-on: never`. Action soft-skips auth/billing/rate-limit as `SKIPPED` (no ERROR sticky) when advisory. `anthropic-api-key` input is `required: false`. +## Learnings (encode + reuse) + +See [`LEARNINGS.md`](LEARNINGS.md). Short form: + +- **SPEC → PLAN → OK → smallest unit → prove with command → merge → clean tree.** +- Gate (agent-action-gate) = auth thesis; lazycoder = trailer; LLM ≠ auth engine. +- Never leave a half PR. Offline proves without keys. No fake precision. +- Adversarial audit after every "fix"; close gaps with regression fixtures. + ## Hard rules (non-negotiable) 1. Never modify code outside the reviewed diff. @@ -28,9 +37,9 @@ Derived from `config/harness.json`. Policy lives in config, not in vibes. 1. **Specify** — load diff, harness, guardrails, review_rules; scope the review. 2. **Plan** — which files/blocks, which subagents; human OK if scope exceeds limits. -3. **Execute** — run rubric via subagents; findings cite rule_id + location. -4. **Verify** — real linter/typecheck/test output in sandbox; no self-reported green. -5. **Decide** — aggregate verdict; human confirms consequential changes. +3. **Execute (smallest unit)** — one concern; findings cite rule_id + location. +4. **Prove** — real pytest / lint / mypy / `corpus_cli.py prove` (no-key path); no self-reported green. +5. **Merge → clean tree** — finish the PR; leave no half-open branch. Verdict for consequential product changes still needs human OK. ## Build order (inside → out) @@ -65,6 +74,6 @@ mypy src ## Agent process in this repo -**Plan → you approve → small change → verify (pytest/lint/mypy) → you decide.** +**SPEC → PLAN → OK → smallest unit → prove (pytest/lint/mypy/prove) → merge → clean tree.** No 200-line unreviewed dumps. One concern per change. diff --git a/LEARNINGS.md b/LEARNINGS.md new file mode 100644 index 0000000..523f646 --- /dev/null +++ b/LEARNINGS.md @@ -0,0 +1,52 @@ +# lazycoder — delivery learnings + +Short, hard-won rules from the aisona-lab harness loop. Do not invent numbers +to make a stage look done. + +## Delivery loop (non-negotiable) + +**SPEC → PLAN → OK → smallest unit → prove with a command → merge → clean tree.** + +1. Write the spec / claim that must become true. +2. Plan the smallest change that makes that claim true. +3. Get human OK when scope is unclear or consequential. +4. Ship one concern per PR — no half-open branches left behind. +5. Prove offline with a real command (see below); paste exit 0, not vibes. +6. Merge only when CI is green and the working tree is clean. + +## Architecture thesis + +| Role | Repo | Keys | +|------|------|------| +| **Auth engine** | [agent-action-gate](https://github.com/aisona-lab/agent-action-gate) | **Zero** LLM keys. Deterministic allow / deny / approval. | +| **Trailer** | this repo (lazycoder) | Optional **one** Anthropic key for live review / Action. | +| Deterministic half | verdict, replay, corpus `prove`, pytest | **No** key. | + +LLM is never the authorization engine. Never claim the stack needs two keys. +Never put real keys in git. + +## Offline proof (definition of done without credits) + +```bash +env -u ANTHROPIC_API_KEY uv run pytest -q +env -u ANTHROPIC_API_KEY uv run python scripts/corpus_cli.py prove corpus/seed.jsonl +``` + +Live corpus scoring / Action dogfood with a real model stays **BLOCKED** until +Anthropic credits exist. Do not invent precision; do not flip `fail-on` without +Stage 2 live numbers (`docs/hardening-plan.md`). + +## After every "fix" + +Run an adversarial audit: missing key, billing soft-skip, injection-in-args as +data, sticky ERROR vs SKIPPED. Close gaps with **regression fixtures / tests**, +not prose. Soft-skip auth/billing **before** posting an ERROR sticky when +advisory (`fail-on: never`). + +## Honesty + +- No fake precision. Publish measured numbers even when bad. +- Offline proves without keys. Live paths soft-skip cleanly when the key is + absent or billing fails. +- Prefer harness (docs, fixtures, FEATURE_MAP, CI) over semantic refactors + unless the task explicitly asks for a semantic change. diff --git a/docs/FEATURE_MAP.md b/docs/FEATURE_MAP.md index 4e80ab3..49a91d3 100644 --- a/docs/FEATURE_MAP.md +++ b/docs/FEATURE_MAP.md @@ -9,6 +9,8 @@ Stage 2 corpus. key. [agent-action-gate](https://github.com/aisona-lab/agent-action-gate) is the auth engine (zero LLM); lazycoder is optional analysis. Never claim two keys. +Process lessons: [`LEARNINGS.md`](../LEARNINGS.md) (SPEC→PLAN→OK→prove→merge; gate=auth, lazycoder=trailer). + ## Deterministic core | Feature | Unit tests | Fixtures | Notes |