Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
59 commits
Select commit Hold shift + click to select a range
71c099d
feat: improve Jev workflows and streamline docs
Coding-Dev-Tools Oct 3, 2026
9165f78
fix: pin README assets for PR preview
Coding-Dev-Tools Oct 3, 2026
3bec5bc
fix: bound Jev review data and disclose remote recall
Coding-Dev-Tools Oct 3, 2026
10cb5ff
fix: preserve Jev support probabilities
Coding-Dev-Tools Oct 3, 2026
b227d25
fix: refresh registered offline benchmark evidence
Coding-Dev-Tools Oct 3, 2026
8f4dcbc
docs: pin README links to refreshed evidence
Coding-Dev-Tools Oct 3, 2026
cd5f3b5
fix: require explicit consent for remote Jev eval
Coding-Dev-Tools Oct 4, 2026
fee9d0c
docs: refresh offline benchmark evidence
Coding-Dev-Tools Oct 4, 2026
24f0fd4
docs: pin README links to refreshed chart
Coding-Dev-Tools Oct 4, 2026
cf6b426
test: align contracts with pinned README links
Coding-Dev-Tools Oct 4, 2026
d7058a8
fix: type Jev advisory planner path
Coding-Dev-Tools Oct 4, 2026
d299855
docs: pin README to refreshed benchmark
Coding-Dev-Tools Oct 4, 2026
a1bdb58
test: accept refreshed benchmark image pin
Coding-Dev-Tools Oct 4, 2026
27ab1b3
fix: bind retrieval preview to grounded Jev plan
Coding-Dev-Tools Oct 4, 2026
e2acf47
docs: pin README benchmark snapshot
Coding-Dev-Tools Oct 4, 2026
a0d3e57
docs: clarify managed Jev allowance terms
Coding-Dev-Tools Oct 4, 2026
6ab0475
docs: align Jev plan details with Cloud policy
Coding-Dev-Tools Oct 4, 2026
c3bfe86
fix: screen Jev questions and correct recall fixture
Coding-Dev-Tools Oct 4, 2026
0f6d292
docs: pin README chart to corrected evidence
Coding-Dev-Tools Oct 4, 2026
52f31d6
fix: surface safe Jev fallback reasons
Coding-Dev-Tools Oct 4, 2026
b08b4e2
docs: pin Jev evidence to the updated snapshot
Coding-Dev-Tools Oct 4, 2026
7516f9c
docs: clarify pooled Team Jev allowance
Coding-Dev-Tools Oct 4, 2026
e440bf6
docs: pin pooled Jev plan guidance
Coding-Dev-Tools Oct 4, 2026
3711dc3
docs: pin README to current Jev and benchmark evidence
Coding-Dev-Tools Oct 4, 2026
4eb16d7
fix(jev): preserve review evidence and recall compatibility
Coding-Dev-Tools Oct 4, 2026
cec7326
docs: pin README to refreshed benchmark evidence
Coding-Dev-Tools Oct 4, 2026
a78a09e
Fix Jev ancestor review and positional API compatibility
Coding-Dev-Tools Oct 4, 2026
5f09d1f
fix: validate visible ancestor memories in Jev review
Coding-Dev-Tools Oct 4, 2026
2ebdb8c
Pin corrected offline benchmark snapshot in every README consumer
Coding-Dev-Tools Oct 4, 2026
4802ecc
docs: pin README to refreshed Jev review evidence
Coding-Dev-Tools Oct 4, 2026
d229892
docs: preserve pre-merge Jev evidence snapshot
Coding-Dev-Tools Oct 4, 2026
fe5d5a6
Enforce current Jev evidence visibility and serialized input bounds
Coding-Dev-Tools Oct 4, 2026
fab7441
Pin current temporal and byte-bound benchmark snapshot
Coding-Dev-Tools Oct 4, 2026
9e40e1a
Merge current Jev review fixes and evidence updates
Coding-Dev-Tools Oct 4, 2026
1d2b32c
fix: honor legacy secret classifications before Jev review
Coding-Dev-Tools Oct 4, 2026
815edc9
docs: pin Jev privacy evidence consumers to source snapshot
Coding-Dev-Tools Oct 4, 2026
cc4f894
Merge latest Jev review fixes and evidence
Coding-Dev-Tools Oct 4, 2026
3c64da4
docs: pin README to benchmark v146
Coding-Dev-Tools Oct 4, 2026
bd9b779
fix: normalize Jev advisory fallback reasons for planning off
Coding-Dev-Tools Oct 4, 2026
135bb72
docs: pin normalized Jev advisory evidence snapshot
Coding-Dev-Tools Oct 4, 2026
0c3482d
Merge current Jev privacy review and benchmark evidence
Coding-Dev-Tools Oct 4, 2026
f2f5701
docs: pin README to benchmark v147
Coding-Dev-Tools Oct 4, 2026
0bd6a8e
fix: normalize Jev fallback and refresh benchmark evidence
Coding-Dev-Tools Oct 4, 2026
e0f6ed3
docs: pin README to benchmark v148
Coding-Dev-Tools Oct 4, 2026
58082c2
fix: update vulnerable benchmark PyJWT pin
Coding-Dev-Tools Oct 4, 2026
94b8d24
docs: publish individual managed Jev limits
Coding-Dev-Tools Oct 4, 2026
fd5b060
fix: address Jev review and docs CI failures
Coding-Dev-Tools Oct 4, 2026
062bff7
chore: export fresh offline evidence for Jev review
Coding-Dev-Tools Oct 4, 2026
5834d39
chore: capture immutable offline evidence outputs
Coding-Dev-Tools Oct 4, 2026
dde3344
chore: regenerate benchmark image exports
Coding-Dev-Tools Oct 4, 2026
4a30d30
docs: register fresh offline evidence for Jev review
Coding-Dev-Tools Oct 4, 2026
7fec59d
docs: checksum fresh offline evidence
Coding-Dev-Tools Oct 4, 2026
e99d2de
docs: bind benchmark guide to refreshed artifact
Coding-Dev-Tools Oct 4, 2026
98757ff
test: bind evidence contract to refreshed artifact
Coding-Dev-Tools Oct 4, 2026
47b3362
docs: cite refreshed offline evidence
Coding-Dev-Tools Oct 4, 2026
8f8c279
docs: refresh registered context evidence chart
Coding-Dev-Tools Oct 4, 2026
f576d8b
docs: refresh registered grounded examples
Coding-Dev-Tools Oct 4, 2026
35037b6
chore: remove temporary offline evidence exporter
Coding-Dev-Tools Oct 4, 2026
c76873d
docs: pin benchmark links to refreshed evidence
Coding-Dev-Tools Oct 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/skill-assets.sha256
Original file line number Diff line number Diff line change
Expand Up @@ -3,4 +3,4 @@ c5d0c26f28c9ee14092f9deaf24c98dd8bef49d971fef2b7a537ffb1ab9f2887 .claude-plugin
aeee7a94671ceb306fe2d24c5acc9f2d96ad8a8e7410536566799eea6265f080 skills/engraphis-memory/SKILL.md
055655db84af07561d002f0c69744313d8413c39f3e873f941f0fa0b1e76dc66 skills/engraphis-memory/references/CONVENTIONS.md
9d090a03f5b3f36a34d91f66b72c3844591f6915755ac3a6c6ba5f1b16977de5 skills/engraphis-memory/references/SCOPING.md
07de31349fc135895377cfcf852c490f732c19e83a627c311d724c5d2146dbb5 skills/engraphis-memory/references/TOOLS.md
46f3ae300558cb327e6b8f5a51c500d299ce7fb61bd3caf7006e5fbbacdcff59 skills/engraphis-memory/references/TOOLS.md
50 changes: 39 additions & 11 deletions BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,15 +93,15 @@ interpretation and do not count as additional benchmark-quality gains.

### Public numeric evidence registry

Every exact public aggregate retained below comes from the checked-in, public-safe
[`offline-fixtures-v128.json`](docs/benchmark-evidence/offline-fixtures-v128.json) artifact. Its
SHA-256 is
`73f2d1a8cd6e2db070577582a2266f6605efc052800da17db9938f7a274bf755`, also recorded in the
Every exact public aggregate retained below comes from the checked-in, public-safe
[`offline-fixtures-v149.json`](docs/benchmark-evidence/offline-fixtures-v149.json) artifact. Its
SHA-256 is
`d5d36c55c4303d77b161137521dd31f77f39b7a0c9e2fed3ddb63e303b12cc6d`, also recorded in the
adjacent `.sha256` file. The artifact contains no raw questions, answers, prompts, customer data,
or per-record content fingerprints.

The fixture-suite digest is
`6e1be135db67964cfb48a2125ac846a9a420662b69ed4d7f61c21d1c87d971b2`. The artifact defines
The fixture-suite digest is
`b0caf9b03349d20da8e30fc9d75b415ed498f0670e9973fc0196d1327c08ba6e`. The artifact defines
the digest algorithm and records the SHA-256 of every suite and dataset file. Each evidence ID
also binds its exact command through `sha256(UTF-8 exact command)`:

Expand All @@ -121,12 +121,19 @@ means no number is claimed in this offline registry.
The context-efficiency chart is generated from the registry values and the selected report schema.
Historical LoCoMo, graph, handoff, consolidation, and security figures remain preserved in their
source artifacts but are omitted from the current chart until each has a matching immutable,
public-safe artifact. The chart labels coding outcomes, external datasets, and operational
capacity as pending evaluation tracks rather than implying scores. Regenerate it with
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v128.json --output docs/images/context-efficiency.svg` after selecting the report to publish.
public-safe artifact. The chart reports context size, candidate and packed retrieval quality, and
grounded checks in separate panels; provider billing and MCP transport are not measured. Regenerate
the SVG and matching PNG with:

```bash
python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v149.json --output docs/images/context-efficiency.svg --png-output docs/images/context-efficiency.png
```

The companion examples are also generated from that artifact with
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v128.json --output docs/images/evidence-backed-agent-examples.svg`.
The companion examples are also generated from that artifact with:

```bash
python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v149.json --output docs/images/evidence-backed-agent-examples.svg --png-output docs/images/evidence-backed-agent-examples.png
```
The historical-to-executable mapping is in
[`docs/BENCHMARK_CHANGE_COVERAGE.md`](docs/BENCHMARK_CHANGE_COVERAGE.md).

Expand Down Expand Up @@ -290,6 +297,27 @@ LoCoMo and LongMemEval retrieval diagnostics are retained as separate public-saf
and checksum boundaries. Those values remain evidence-retrieval metrics, not end-to-end QA accuracy
or an official LoCoMo leaderboard score.

## Jev-assisted route-selection probe (provider-backed, exploratory)

On 2026-10-03, the configured BYOK client ran `eval/jev_recall_quality.py` against 40 generated
public synthetic tasks (10 each for lexical identifiers, graph relationships, temporal changes,
and procedural context). Each call sent only the task query and at most two locally generated
routes; no recalled memory bodies or stored user data were sent. Jev selected a route on all 40
calls, but the baseline and Jev-assisted means were identical: nDCG@5 **0.740773**, Recall@5
**0.75**, and answer-token coverage **0.75**. Every paired 95% interval was zero, so the
improvement/non-inferiority gate did not pass. This synthetic probe does not establish a benefit
on held-out user workloads, and no retrieval-quality improvement is claimed.

The harness caps a run at 40 provider requests, makes at most one request per task, and requires
explicit `--allow-remote` consent. Repeating it can consume BYOK provider usage:

```bash
python -m eval.jev_recall_quality --allow-remote --max-requests 40 --timeout 8
```

This exploratory provider run is separate from the registered offline artifact and the benchmark
graphic below; it is not a general benchmark or customer-workload evaluation.

## What we do NOT yet claim

- **No official end-to-end LLM QA accuracy.** The deterministic productivity agent measures the
Expand Down
Loading
Loading