Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 13 additions & 4 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,15 @@ Derived from `config/harness.json`. Policy lives in config, not in vibes.
- Live review / Action dogfood: **at most one** Anthropic key. Never require two keys to run the stack. Never put real keys in git.
- CI self-review (`.github/workflows/self-review.yml`) uses a **key-gate** job (GitHub forbids `secrets` in `jobs.<id>.if`) so an empty secret skips the review job; keeps `fail-on: never`. Action soft-skips auth/billing/rate-limit as `SKIPPED` (no ERROR sticky) when advisory. `anthropic-api-key` input is `required: false`.

## Learnings (encode + reuse)

See [`LEARNINGS.md`](LEARNINGS.md). Short form:

- **SPEC → PLAN → OK → smallest unit → prove with command → merge → clean tree.**
- Gate (agent-action-gate) = auth thesis; lazycoder = trailer; LLM ≠ auth engine.
- Never leave a half PR. Offline proves without keys. No fake precision.
- Adversarial audit after every "fix"; close gaps with regression fixtures.

## Hard rules (non-negotiable)

1. Never modify code outside the reviewed diff.
Expand All @@ -28,9 +37,9 @@ Derived from `config/harness.json`. Policy lives in config, not in vibes.

1. **Specify** — load diff, harness, guardrails, review_rules; scope the review.
2. **Plan** — which files/blocks, which subagents; human OK if scope exceeds limits.
3. **Execute** — run rubric via subagents; findings cite rule_id + location.
4. **Verify** — real linter/typecheck/test output in sandbox; no self-reported green.
5. **Decide** — aggregate verdict; human confirms consequential changes.
3. **Execute (smallest unit)** — one concern; findings cite rule_id + location.
4. **Prove** — real pytest / lint / mypy / `corpus_cli.py prove` (no-key path); no self-reported green.
5. **Merge → clean tree** — finish the PR; leave no half-open branch. Verdict for consequential product changes still needs human OK.

## Build order (inside → out)

Expand Down Expand Up @@ -65,6 +74,6 @@ mypy src

## Agent process in this repo

**Plan → you approve → small change → verify (pytest/lint/mypy) → you decide.**
**SPEC → PLAN → OK → smallest unit → prove (pytest/lint/mypy/prove) → merge → clean tree.**

No 200-line unreviewed dumps. One concern per change.
52 changes: 52 additions & 0 deletions LEARNINGS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# lazycoder — delivery learnings

Short, hard-won rules from the aisona-lab harness loop. Do not invent numbers
to make a stage look done.

## Delivery loop (non-negotiable)

**SPEC → PLAN → OK → smallest unit → prove with a command → merge → clean tree.**

1. Write the spec / claim that must become true.
2. Plan the smallest change that makes that claim true.
3. Get human OK when scope is unclear or consequential.
4. Ship one concern per PR — no half-open branches left behind.
5. Prove offline with a real command (see below); paste exit 0, not vibes.
6. Merge only when CI is green and the working tree is clean.

## Architecture thesis

| Role | Repo | Keys |
|------|------|------|
| **Auth engine** | [agent-action-gate](https://github.com/aisona-lab/agent-action-gate) | **Zero** LLM keys. Deterministic allow / deny / approval. |
| **Trailer** | this repo (lazycoder) | Optional **one** Anthropic key for live review / Action. |
| Deterministic half | verdict, replay, corpus `prove`, pytest | **No** key. |

LLM is never the authorization engine. Never claim the stack needs two keys.
Never put real keys in git.

## Offline proof (definition of done without credits)

```bash
env -u ANTHROPIC_API_KEY uv run pytest -q
env -u ANTHROPIC_API_KEY uv run python scripts/corpus_cli.py prove corpus/seed.jsonl
```

Live corpus scoring / Action dogfood with a real model stays **BLOCKED** until
Anthropic credits exist. Do not invent precision; do not flip `fail-on` without
Stage 2 live numbers (`docs/hardening-plan.md`).

## After every "fix"

Run an adversarial audit: missing key, billing soft-skip, injection-in-args as
data, sticky ERROR vs SKIPPED. Close gaps with **regression fixtures / tests**,
not prose. Soft-skip auth/billing **before** posting an ERROR sticky when
advisory (`fail-on: never`).

## Honesty

- No fake precision. Publish measured numbers even when bad.
- Offline proves without keys. Live paths soft-skip cleanly when the key is
absent or billing fails.
- Prefer harness (docs, fixtures, FEATURE_MAP, CI) over semantic refactors
unless the task explicitly asks for a semantic change.
2 changes: 2 additions & 0 deletions docs/FEATURE_MAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,8 @@ Stage 2 corpus.
key. [agent-action-gate](https://github.com/aisona-lab/agent-action-gate) is the
auth engine (zero LLM); lazycoder is optional analysis. Never claim two keys.

Process lessons: [`LEARNINGS.md`](../LEARNINGS.md) (SPEC→PLAN→OK→prove→merge; gate=auth, lazycoder=trailer).

## Deterministic core

| Feature | Unit tests | Fixtures | Notes |
Expand Down
Loading