Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -14,3 +14,7 @@ logs/
.DS_Store
CLAUDE.md
.claude/

# Stage 2 harvest scratch (labelled seed is committed)
corpus/candidates.jsonl
corpus/defective_seed.jsonl
11 changes: 6 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -321,11 +321,12 @@ pytest -m integration
in the wheel, releases via trusted publishing on `v*` tags.~~ ✓
7. ~~**GitHub Action** wrapping the CLI, so `uses: aisona-lab/lazycoder` gates a
PR with the same rubric and exit codes.~~ ✓
8. **A real corpus** — 30–50 hunks from merged OSS PRs, half with a defect the
follow-up fix confirms, half genuinely clean. Harvester and scoring are
built (`scripts/corpus_cli.py`); the cull thresholds are pre-registered in
[`docs/hardening-plan.md`](docs/hardening-plan.md) so the result is a
measurement rather than a rationalisation. What remains is labelling.
8. **A real corpus** — labelled seed in `corpus/seed.jsonl` (22 clean / 15
defective from merged OSS PRs) meets the Stage 2 shape floor. Cull
thresholds are executable (`decide_cull` / `fail_on_gate`); offline proof:
`python scripts/corpus_cli.py prove corpus/seed.jsonl`. **Live** scoring
(`corpus_cli.py run`) still needs Anthropic credits; Action `fail-on`
stays `never` until the fail-on gate is ready.
9. **File-level context** — review the whole post-change file with the diff
marked inside it, so the system-level rules (state, compatibility,
concurrency) become answerable instead of abstaining. Cheaper too: a 40-hunk
Expand Down
35 changes: 35 additions & 0 deletions corpus/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# Stage 2 corpus

Labelled real hunks for precision / quiet-rate measurement. See
[`docs/hardening-plan.md`](../docs/hardening-plan.md) Stage 2.

| File | Role |
|------|------|
| `seed.jsonl` | Labelled seed (≥20 clean, ≥15 defective). Checked into git. |
| `candidates.jsonl` | Unlabelled harvest output — local only, not committed. |

## Labels

- **clean** — post-change side of a merged non-bugfix PR. Seed notes mark
provisional labels that have not been longitudinally checked for later fixes.
- **defective** — pre-image of a known bug-fix PR, with `expect_rules` naming
the defect the follow-up fixed. Extra findings on defective hunks are
*unlabelled*, not false positives.

`load_corpus` refuses any line with `"label": null`.

## Offline proof (no API key)

```bash
uv run python scripts/corpus_cli.py prove corpus/seed.jsonl
# artifact: docs/eval-runs/corpus-prove.txt — expect exit 0
```

## Live scoring (needs Anthropic credits)

```bash
ANTHROPIC_API_KEY=... uv run python scripts/corpus_cli.py run corpus/seed.jsonl
```

Do **not** flip Action `fail-on` away from `never` until the printed fail-on
gate is ready (`quiet_rate ≥ 80%` and no high rule above 5% clean noise).
37 changes: 37 additions & 0 deletions corpus/seed.jsonl

Large diffs are not rendered by default.

19 changes: 19 additions & 0 deletions docs/eval-runs/corpus-prove.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
lazycoder Stage 2 — corpus prove (offline)
date: 2026-09-29 14:25 CEST (PT)
corpus: corpus/seed.jsonl

Pre-registered thresholds (hardening-plan.md, 2026-08-26):
delete if clean_noise > 25% and caught == 0
demote if clean_noise > 10% and severity high
stage3a if abstention > 50% and clean_noise < 10%
fail-on ready if quiet_rate ≥ 80% and no high rule > 5% clean noise
shape: ≥20 clean, ≥15 defective

shape: 22 clean (≥20), 15 defective (≥15) — shape OK
repos: encode/httpx, pallets/click, psf/requests, urllib3/urllib3
hunks: 37
cull self-check: PASS (6/6 decision-table cases)
live scoring: BLOCKED — ANTHROPIC_API_KEY unset (no credits/key in this environment; offline half only)
fail-on: stays never until live corpus metrics satisfy the gate (no live scores in this prove run — gate not claimed ready)

OFFLINE_PROVE: PASS LIVE_SCORING: BLOCKED FAIL_ON: never
7 changes: 7 additions & 0 deletions docs/hardening-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,13 @@ decision log fully determine the verdict.

## Stage 2 — the number

**Phase 3 (offline) landed:** `corpus/seed.jsonl` meets the shape floor
(≥20 clean / ≥15 defective); `decide_cull` / `fail_on_gate` encode the
pre-registered thresholds; `scripts/corpus_cli.py prove` is the offline
proof. Live `run` + rule cull from real scores remain gated on Anthropic
credits — do not invent precision numbers and do not flip `fail-on`.


Corpus of 30–50 hunks from merged OSS PRs: half with a defect the follow-up fix
confirms, half genuinely clean. Not toy snippets.

Expand Down
198 changes: 188 additions & 10 deletions scripts/corpus_cli.py
Original file line number Diff line number Diff line change
@@ -1,39 +1,60 @@
"""Harvest and score a corpus of real diff hunks.
"""Harvest, prove, and score a corpus of real diff hunks.

# pull candidate hunks from merged PRs (needs gh)
python scripts/corpus_cli.py harvest --repo psf/requests --prs 8 \
--out corpus/candidates.jsonl

# label each line by hand: "label": "clean" | "defective" (+ expect_rules)

# score the labelled corpus against the live model
ANTHROPIC_API_KEY=... python scripts/corpus_cli.py run corpus/corpus.jsonl
# offline Stage 2 proof (no Anthropic key): shape + cull thresholds
python scripts/corpus_cli.py prove corpus/seed.jsonl \
--out docs/eval-runs/corpus-prove.txt

# score the labelled corpus against the live model (needs ANTHROPIC_API_KEY)
ANTHROPIC_API_KEY=... python scripts/corpus_cli.py run corpus/seed.jsonl

Harvest emits the post-change side of each hunk, which is what a *clean*
sample is. Defective samples come from bug-fix PRs and are labelled by hand
for now; a --side old extractor is only worth building if that becomes the
bottleneck.
sample is. Defective samples come from bug-fix PRs (pre-image) and are
labelled by hand.
"""

from __future__ import annotations

import argparse
import json
import os
import subprocess
import sys
from datetime import datetime
from pathlib import Path
from zoneinfo import ZoneInfo

sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "src"))

from lazycoder.config import load_all_configs # noqa: E402
from lazycoder.corpus import ( # noqa: E402
DELETE_NOISE_RATE,
DEMOTE_HIGH_NOISE_RATE,
FAIL_ON_MAX_HIGH_NOISE_RATE,
FAIL_ON_MIN_QUIET_RATE,
MIN_CLEAN_HUNKS,
MIN_DEFECTIVE_HUNKS,
STAGE3A_ABSTENTION_RATE,
STAGE3A_MAX_NOISE_RATE,
CorpusLoadError,
CullAction,
HunkResult,
Label,
RuleVerdict,
corpus_shape,
cull_plan,
decide_cull,
fail_on_gate,
load_corpus,
report,
score_hunk,
)
from lazycoder.domain import RuleId, Severity # noqa: E402
from lazycoder.llm.anthropic_client import AnthropicClient # noqa: E402
from lazycoder.orchestrator import parse_diff # noqa: E402
from lazycoder.reviewers import SingleRuleReviewer # noqa: E402
Expand All @@ -43,9 +64,8 @@
# Most recently merged PRs on a popular repo are dependency bumps. Scanning
# them finds nothing and makes an empty harvest look like a broken one.
BOT_AUTHORS = {"dependabot", "pre-commit-ci", "renovate", "github-actions"}
# Most recently merged PRs on a popular repo are dependency bumps. Scanning
# them finds nothing and makes an empty harvest look like a broken one.
BOT_AUTHORS = {"dependabot", "pre-commit-ci", "renovate", "github-actions"}

USER_TZ = ZoneInfo("Europe/Prague")


def _gh(*args: str) -> str:
Expand Down Expand Up @@ -159,6 +179,14 @@ def _ids(rules: frozenset) -> str:


def run(path: Path) -> int:
if not os.environ.get("ANTHROPIC_API_KEY"):
print(
"error: ANTHROPIC_API_KEY unset — live corpus scoring is blocked.\n"
" Offline half: python scripts/corpus_cli.py prove " + str(path),
file=sys.stderr,
)
return 2

config = load_all_configs()
rubric = config.review_rules
hunks = load_corpus(path)
Expand All @@ -185,13 +213,143 @@ def run(path: Path) -> int:
f"recall {summary.recall:.0%}"
f" ({summary.caught} caught, {summary.missed} missed)"
)
print(f"\n{'rule':6}{'noise%':>8}{'caught':>8}{'missed':>8}{'abstain':>9}")
print(
f"\n{'rule':6}{'noise%':>8}{'caught':>8}{'missed':>8}{'abstain':>9}{'cull':>14}"
)
plan = cull_plan(summary, rubric)
for rule_id, verdict in summary.per_rule.items():
print(
f"{rule_id.value:6}{verdict.clean_noise_rate:>7.0%}"
f"{verdict.caught:>8}{verdict.missed:>8}"
f"{verdict.abstention_rate:>8.0%}"
f"{plan[rule_id].value:>14}"
)
gate = fail_on_gate(summary, rubric)
print(f"\nfail-on gate ready: {gate.ready} ({gate.reason})")
if not gate.ready:
print("Action fail-on stays never — metrics do not justify a blocking gate.")
return 0


def _cull_self_check() -> list[str]:
"""Measurable offline proof that the encoded thresholds match the plan table."""
cases: list[tuple[str, RuleVerdict, Severity, CullAction]] = [
(
"delete: noise>25% and never right",
RuleVerdict(RuleId.R12, 20, 6, 0, 0, 0, 0, 20),
Severity.MEDIUM,
CullAction.DELETE,
),
(
"demote: high rule noise>10%",
RuleVerdict(RuleId.R4, 20, 3, 1, 0, 0, 0, 20),
Severity.HIGH,
CullAction.DEMOTE,
),
(
"stage3a: abstains a lot, quiet on clean",
RuleVerdict(RuleId.R14, 20, 1, 0, 0, 0, 12, 20),
Severity.MEDIUM,
CullAction.KEEP_STAGE_3A,
),
(
"keep: modest noise, sometimes right",
RuleVerdict(RuleId.R7, 20, 2, 3, 1, 0, 2, 20),
Severity.HIGH,
CullAction.KEEP,
),
(
"delete beats demote when never right",
RuleVerdict(RuleId.R4, 20, 6, 0, 0, 0, 0, 20),
Severity.HIGH,
CullAction.DELETE,
),
(
"keep: noisy medium with catches (not delete, not demote)",
RuleVerdict(RuleId.R3, 20, 6, 2, 0, 0, 0, 20),
Severity.MEDIUM,
CullAction.KEEP,
),
]
failures: list[str] = []
for name, verdict, severity, expected in cases:
got = decide_cull(verdict, severity)
if got is not expected:
failures.append(f"{name}: expected {expected.value}, got {got.value}")
return failures


def prove(path: Path, out: Path | None) -> int:
"""Offline Stage 2 proof: labelled corpus shape + cull threshold self-check.

Does not call the model. Live scoring remains a separate `run` step.
"""
now = datetime.now(tz=USER_TZ)
lines: list[str] = [
"lazycoder Stage 2 — corpus prove (offline)",
f"date: {now.strftime('%Y-%m-%d %H:%M %Z')} (PT)",
f"corpus: {path}",
"",
"Pre-registered thresholds (hardening-plan.md, 2026-08-26):",
f" delete if clean_noise > {DELETE_NOISE_RATE:.0%} and caught == 0",
f" demote if clean_noise > {DEMOTE_HIGH_NOISE_RATE:.0%} and severity high",
f" stage3a if abstention > {STAGE3A_ABSTENTION_RATE:.0%}"
f" and clean_noise < {STAGE3A_MAX_NOISE_RATE:.0%}",
f" fail-on ready if quiet_rate ≥ {FAIL_ON_MIN_QUIET_RATE:.0%}"
f" and no high rule > {FAIL_ON_MAX_HIGH_NOISE_RATE:.0%} clean noise",
f" shape: ≥{MIN_CLEAN_HUNKS} clean, ≥{MIN_DEFECTIVE_HUNKS} defective",
"",
]

hunks = load_corpus(path)
shape = corpus_shape(hunks)
lines.append(f"shape: {shape.summary}")
repos = sorted({h.source.repo for h in hunks})
lines.append(f"repos: {', '.join(repos)}")
lines.append(f"hunks: {len(hunks)}")

cull_failures = _cull_self_check()
if cull_failures:
lines.append("cull self-check: FAIL")
lines.extend(f" - {f}" for f in cull_failures)
else:
lines.append("cull self-check: PASS (6/6 decision-table cases)")

key_present = bool(os.environ.get("ANTHROPIC_API_KEY"))
if key_present:
lines.append(
"live scoring: API key present — run `corpus_cli.py run` separately"
)
live_status = "KEY_PRESENT"
else:
lines.append(
"live scoring: BLOCKED — ANTHROPIC_API_KEY unset "
"(no credits/key in this environment; offline half only)"
)
live_status = "BLOCKED"

lines.append(
"fail-on: stays never until live corpus metrics satisfy the gate "
"(no live scores in this prove run — gate not claimed ready)"
)
lines.append("")

offline_ok = shape.ok and not cull_failures
lines.append(
f"OFFLINE_PROVE: {'PASS' if offline_ok else 'FAIL'} "
f"LIVE_SCORING: {live_status} FAIL_ON: never"
)

text = "\n".join(lines) + "\n"
sys.stdout.write(text)

if out is not None:
out.parent.mkdir(parents=True, exist_ok=True)
out.write_text(text, encoding="utf-8")
print(f"artifact: {out}", file=sys.stderr)

if not offline_ok:
return 1
return 0


Expand All @@ -209,10 +367,30 @@ def main() -> int:
r = sub.add_parser("run", help="score a labelled corpus against the live model")
r.add_argument("path", type=Path)

p = sub.add_parser(
"prove",
help="offline Stage 2 proof: corpus shape + cull threshold self-check",
)
p.add_argument("path", type=Path, help="labelled corpus JSONL")
p.add_argument(
"--out",
type=Path,
default=Path("docs/eval-runs/corpus-prove.txt"),
help="where to write the prove artifact",
)
p.add_argument(
"--no-artifact",
action="store_true",
help="print only; do not write --out",
)

args = parser.parse_args()
try:
if args.command == "harvest":
return harvest(args.repo, args.prs, args.per_pr, args.out, bots=args.bots)
if args.command == "prove":
out = None if args.no_artifact else args.out
return prove(args.path, out)
return run(args.path)
except (CorpusLoadError, RuntimeError, OSError) as exc:
print(f"error: {exc}", file=sys.stderr)
Expand Down
Loading
Loading