Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
33 commits
Select commit Hold shift + click to select a range
cfbccec
Align managed Jev clients with individual access and fixed advisory o…
Coding-Dev-Tools Oct 4, 2026
e584e09
Use supported managed context in credential deadline regressions
Coding-Dev-Tools Oct 4, 2026
66920e4
Refresh source-bound offline evidence after current Jev workflow inte…
Coding-Dev-Tools Oct 4, 2026
ded0675
Pin Jev documentation and charts to the current reviewed snapshot
Coding-Dev-Tools Oct 4, 2026
9ede220
Reconcile current parent Jev policy documentation and preserve stack …
Coding-Dev-Tools Oct 4, 2026
2cd9d75
Pin public Jev documentation to the reconciled parent snapshot
Coding-Dev-Tools Oct 4, 2026
12d2e40
Preserve current Jev question screening and refresh combined offline …
Coding-Dev-Tools Oct 4, 2026
4646aa0
Pin current combined Jev implementation and evidence in public docume…
Coding-Dev-Tools Oct 4, 2026
0b25802
Preserve parent chart-pinning history with the combined evidence snap…
Coding-Dev-Tools Oct 4, 2026
54eae9f
Preserve safe Jev fallback updates and exact individual allowances
Coding-Dev-Tools Oct 4, 2026
2b86007
Pin public Jev requirements and charts to the combined snapshot
Coding-Dev-Tools Oct 4, 2026
b00096c
Align all public documentation checks with the combined snapshot
Coding-Dev-Tools Oct 4, 2026
07a5ced
Reconcile current parent while preserving agreed per-seat Jev allowances
Coding-Dev-Tools Oct 4, 2026
c582797
Reconcile managed Jev allowance and request bounds
Coding-Dev-Tools Oct 4, 2026
92fb451
Fix pinned README links for current release evidence
Coding-Dev-Tools Oct 4, 2026
3e593bb
Merge Jev integration and align pooled Pro Team access
Coding-Dev-Tools Oct 4, 2026
6a93da1
Update pinned hosted plan links in pricing tests
Coding-Dev-Tools Oct 4, 2026
719e171
Merge Jev review fixes and refresh offline evidence
Coding-Dev-Tools Oct 4, 2026
616aa27
docs: pin README links to the refreshed Jev snapshot
Coding-Dev-Tools Oct 4, 2026
c707673
Merge latest hosted plan link checks
Coding-Dev-Tools Oct 4, 2026
667ebea
docs: publish individual managed Jev limits
Coding-Dev-Tools Oct 4, 2026
df8d043
Merge updated managed Jev contract docs
Coding-Dev-Tools Oct 4, 2026
56625f4
Merge reviewed Jev fixes and documentation into request-bound stack
Coding-Dev-Tools Oct 4, 2026
6ab9161
chore: export fresh managed Jev benchmark evidence
Coding-Dev-Tools Oct 4, 2026
e69af55
chore: regenerate managed Jev benchmark image exports
Coding-Dev-Tools Oct 4, 2026
2a19578
chore: run exact-head benchmark export on branch pushes
Coding-Dev-Tools Oct 4, 2026
96a9024
chore: remove superseded draft benchmark artifact
Coding-Dev-Tools Oct 4, 2026
7df48fd
chore: remove superseded draft benchmark artifact
Coding-Dev-Tools Oct 4, 2026
6332018
docs: sync benchmark contract with parent stack
Coding-Dev-Tools Oct 4, 2026
8338a4c
docs: sync benchmark contract with parent stack
Coding-Dev-Tools Oct 4, 2026
4622281
docs: sync benchmark contract with parent stack
Coding-Dev-Tools Oct 4, 2026
34982ea
docs: sync benchmark contract with parent stack
Coding-Dev-Tools Oct 4, 2026
6defd55
docs: sync benchmark contract with parent stack
Coding-Dev-Tools Oct 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/skill-assets.sha256
Original file line number Diff line number Diff line change
Expand Up @@ -3,4 +3,4 @@ c5d0c26f28c9ee14092f9deaf24c98dd8bef49d971fef2b7a537ffb1ab9f2887 .claude-plugin
aeee7a94671ceb306fe2d24c5acc9f2d96ad8a8e7410536566799eea6265f080 skills/engraphis-memory/SKILL.md
055655db84af07561d002f0c69744313d8413c39f3e873f941f0fa0b1e76dc66 skills/engraphis-memory/references/CONVENTIONS.md
9d090a03f5b3f36a34d91f66b72c3844591f6915755ac3a6c6ba5f1b16977de5 skills/engraphis-memory/references/SCOPING.md
46f3ae300558cb327e6b8f5a51c500d299ce7fb61bd3caf7006e5fbbacdcff59 skills/engraphis-memory/references/TOOLS.md
582bae24a791d059f3af9c2c13ef952965bfc9861fb891b53687999a6a43e2fc skills/engraphis-memory/references/TOOLS.md
64 changes: 64 additions & 0 deletions .github/workflows/export-offline-evidence.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
name: Temporary offline evidence export
on:
push:
branches:
- codex/jev-individual-access-20261004
paths:
- ".github/workflows/export-offline-evidence.yml"
- "engraphis/**"
- "eval/**"
- "scripts/export_offline_evidence.py"
pull_request:
types: [opened, synchronize, reopened]
paths:
- ".github/workflows/export-offline-evidence.yml"
- "engraphis/**"
- "eval/**"
- "scripts/export_offline_evidence.py"
permissions:
contents: read
jobs:
export:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
ref: ${{ github.event.pull_request.head.sha }}
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
with:
python-version: "3.11"
- run: python -m pip install numpy cairosvg
- run: python -m scripts.export_offline_evidence --output docs/benchmark-evidence/offline-fixtures-v152.json
- run: python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v152.json --output /tmp/context-efficiency.svg --png-output /tmp/context-efficiency.png
- run: python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v152.json --output /tmp/evidence-backed-agent-examples.svg --png-output /tmp/evidence-backed-agent-examples.png
- name: Show generated public artifacts
run: |
echo "BEGIN_OFFLINE_FIXTURES_JSON"
cat docs/benchmark-evidence/offline-fixtures-v152.json
echo "END_OFFLINE_FIXTURES_JSON"
echo "BEGIN_OFFLINE_FIXTURES_SHA"
cat docs/benchmark-evidence/offline-fixtures-v152.json.sha256
echo "END_OFFLINE_FIXTURES_SHA"
echo "BEGIN_CONTEXT_EFFICIENCY_SVG"
cat /tmp/context-efficiency.svg
echo "END_CONTEXT_EFFICIENCY_SVG"
echo "BEGIN_EVIDENCE_BACKED_EXAMPLES_SVG"
cat /tmp/evidence-backed-agent-examples.svg
echo "END_EVIDENCE_BACKED_EXAMPLES_SVG"
echo "BEGIN_CONTEXT_EFFICIENCY_PNG_BASE64"
base64 -w 76 /tmp/context-efficiency.png
echo "END_CONTEXT_EFFICIENCY_PNG_BASE64"
echo "BEGIN_EVIDENCE_BACKED_EXAMPLES_PNG_BASE64"
base64 -w 76 /tmp/evidence-backed-agent-examples.png
echo "END_EVIDENCE_BACKED_EXAMPLES_PNG_BASE64"
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7
with:
name: offline-fixtures-v152
path: |
docs/benchmark-evidence/offline-fixtures-v152.json
docs/benchmark-evidence/offline-fixtures-v152.json.sha256
/tmp/context-efficiency.svg
/tmp/evidence-backed-agent-examples.svg
/tmp/context-efficiency.png
/tmp/evidence-backed-agent-examples.png
retention-days: 1
10 changes: 5 additions & 5 deletions BENCHMARKS.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,14 +94,14 @@ interpretation and do not count as additional benchmark-quality gains.
### Public numeric evidence registry

Every exact public aggregate retained below comes from the checked-in, public-safe
[`offline-fixtures-v148.json`](docs/benchmark-evidence/offline-fixtures-v148.json) artifact. Its
[`offline-fixtures-v149.json`](docs/benchmark-evidence/offline-fixtures-v149.json) artifact. Its
SHA-256 is
`4cf3c335b713d5b6e479b80a2d6e314c8904250478062124eef7abf0d3d8a378`, also recorded in the
`d5d36c55c4303d77b161137521dd31f77f39b7a0c9e2fed3ddb63e303b12cc6d`, also recorded in the
adjacent `.sha256` file. The artifact contains no raw questions, answers, prompts, customer data,
or per-record content fingerprints.

The fixture-suite digest is
`6b3ad6d9fdba747908ed6b77bff31a996fe28cfce84c764448f3cb59f0c7692b`. The artifact defines
`b0caf9b03349d20da8e30fc9d75b415ed498f0670e9973fc0196d1327c08ba6e`. The artifact defines
the digest algorithm and records the SHA-256 of every suite and dataset file. Each evidence ID
also binds its exact command through `sha256(UTF-8 exact command)`:

Expand All @@ -126,13 +126,13 @@ grounded checks in separate panels; provider billing and MCP transport are not m
the SVG and matching PNG with:

```bash
python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v148.json --output docs/images/context-efficiency.svg --png-output docs/images/context-efficiency.png
python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v149.json --output docs/images/context-efficiency.svg --png-output docs/images/context-efficiency.png
```

The companion examples are also generated from that artifact with:

```bash
python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v148.json --output docs/images/evidence-backed-agent-examples.svg --png-output docs/images/evidence-backed-agent-examples.png
python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v149.json --output docs/images/evidence-backed-agent-examples.svg --png-output docs/images/evidence-backed-agent-examples.png
```
The historical-to-executable mapping is in
[`docs/BENCHMARK_CHANGE_COVERAGE.md`](docs/BENCHMARK_CHANGE_COVERAGE.md).
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,7 @@ Connect a coding agent over MCP with the [agent setup guide](https://github.com/
The current registered artifact contains three deterministic offline fixture runs. It separates context size, retrieval quality, and grounded decision checks; these small fixtures do not establish general task performance.

<p align="center">
<img src="https://raw.githubusercontent.com/Coding-Dev-Tools/engraphis/0bd6a8e8f5804a664d9850c57b159d73012ce59b/docs/images/context-efficiency.svg" alt="Three registered offline fixtures: structure-aware chunking reduced mean retrieved context from 740.3 to 214.3 tokens per question (71.1%, 18 questions); the serialized JSON-shape proxy fell from 24,590 to 11,138 tokens across 26 payload samples and 260 recalls. Candidate and packed retrieval quality (Recall@5, Hit@5, and answer-token recall) are shown separately (each 1.000). Grounded checks show 5/5 answerable queries grounded and 6/6 abstention queries rejected, including a 1/1 quarantined-evidence probe; 11/11 decisions were correct. MCP transport and provider billing were not measured. Artifact SHA-256 prefix 4cf3c335b713; see the benchmark guide for the full checksum." width="100%">
<img src="https://raw.githubusercontent.com/Coding-Dev-Tools/engraphis/35037b63f372486f66c28257e29146f74c5654f3/docs/images/context-efficiency.svg" alt="Three registered offline fixtures: structure-aware chunking reduced mean retrieved context from 740.3 to 214.3 tokens per question (71.1%, 18 questions); the serialized JSON-shape proxy fell from 24,590 to 11,138 tokens across 26 payload samples and 260 recalls. Candidate and packed retrieval quality (Recall@5, Hit@5, and answer-token recall) are shown separately (each 1.000). Grounded checks show 5/5 answerable queries grounded and 6/6 abstention queries rejected, including a 1/1 quarantined-evidence probe; 11/11 decisions were correct. MCP transport and provider billing were not measured. Artifact SHA-256 prefix d5d36c55c430; see the benchmark guide for the full checksum." width="100%">
<br>
<sup>Three offline fixtures separate context reduction, candidate and packed retrieval quality, and grounded behavior. The chart shows the artifact checksum prefix; see the benchmark guide for the full checksum and reproduction steps.</sup>
</p>
Expand All @@ -60,7 +60,7 @@ The current registered artifact contains three deterministic offline fixture run
| Recall payload proxy | JSON-shape proxy: 24,590 → 11,138 tokens (54.71% lower, 26 samples; 260 timed recalls) | Candidate and packed Recall@5, Hit@5, and answer-token recall are each 1.000 |
| Grounded decisions | 5/5 answerable queries grounded; 6/6 abstention queries rejected, including 1/1 quarantined-evidence check | 11/11 decisions correct |

The payload figure is a serialized JSON-shape estimate, not an MCP transport measurement or provider billing total. See the [Benchmark methodology](https://github.com/Coding-Dev-Tools/engraphis/blob/0bd6a8e8f5804a664d9850c57b159d73012ce59b/BENCHMARKS.md) for artifact identity, counting methods, reproduction commands, external-evaluation boundaries, and limitations.
The payload figure is a serialized JSON-shape estimate, not an MCP transport measurement or provider billing total. See the [Benchmark methodology](https://github.com/Coding-Dev-Tools/engraphis/blob/35037b63f372486f66c28257e29146f74c5654f3/BENCHMARKS.md) for artifact identity, counting methods, reproduction commands, external-evaluation boundaries, and limitations.

## Optional Jev assistance

Expand Down
14 changes: 14 additions & 0 deletions docs/CONFIGURATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,6 +84,20 @@ service-enforced, not client configuration overrides. Selecting a backend does n
enable the service or satisfy release, provider-terms, capacity, or quality gates. See
[hosted plans](HOSTED_PLANS.md#included-system-1-decision-engine-jev).

The four managed workflows require a command, a pair of facts, query/evidence, or
goal/output context and their fixed question schemas. Arbitrary `custom` and
`query_planning` payloads are rejected before credential refresh or network requests.
Experimental recall route planning requires explicit BYOK; managed failure never
silently selects a personal key. Remote consent and `public` or `internal`
classification remain mandatory, and `offline_mode=true` prevents remote requests.
Viewers can use direct Classic advisory decisions; generic Smart stateful execution
still requires admin and the private Team tool catalog is unchanged.

Managed transport callers can supply `request_key` for explicit retry deduplication. A
duplicate admitted key returns 409 without an additional provider call or usage
increment; there is no stored answer replay or automatic retry. Admitted errors remain
counted. This parameter is not an environment setting or an MCP/dashboard field.

The optional cross-encoder reranker is model- and hardware-dependent. Treat its quality and
latency as deployment-specific until a versioned model identity, exact configuration, and
reproducible evaluation artifact are available for the comparison being reported.
Expand Down
34 changes: 26 additions & 8 deletions docs/MCP_TOOLS.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,14 +39,18 @@ namesakes; advanced controls are discoverable rather than routine:
`full`. It preserves the same selected text and whitespace; it does not apply an
additional summary or promise extra token savings. Source IDs remain in `sources`.

Jev-assisted recall planning is opt-in. On the Smart context tool, set `allow_remote=true` and
Jev-assisted recall planning requires explicit BYOK and is opt-in. On the Smart context tool,
set `allow_remote=true` and
`data_classification="public"` or `"internal"`; that call both enables bounded route selection
and gives per-call consent for remote processing. The default stays deterministic and local.
Classic recall tools additionally require `planning="auto"` and `jev_assisted=true`.
The provider receives the query and bounded local routes, not recalled memory bodies. Jev cannot
change scope, time, type, or trust filters; uncertainty and failures retain deterministic route
order and appear in `planning_advisory`. Leaving the Jev controls at their defaults keeps local
deterministic behavior.
Managed access does not admit `query_planning`: it fails closed before credential refresh or
network requests, retaining deterministic route order with a visible fallback. Neither `managed`
nor `auto` silently switches to BYOK.


No user profile choice or tool switching is required. The dashboard `/mcp` endpoint and
Expand Down Expand Up @@ -201,9 +205,10 @@ Decision inputs must be nonblank for the selected kind: `guard_command` needs `s
and `query`; `verify_completion` needs `state` and `goal`, with optional `recent_actions`.
`custom` accepts either `state` or `question`. Missing required input returns `invalid_request`
before backend lookup, with unknown/null conclusions and no remote allowance consumed.
Command decisions always return `allow_auto=false` and `escalate_to_user=true`, including
successful remote answers. Provider probability and category are advice, not shell authorization.

Managed access admits only the four concrete workflows above with their fixed questions and
required context. A purpose label cannot authorize arbitrary question schemas. Managed `custom`
requests return `managed_operation_unsupported` without a credential refresh or network request;
custom remote questions require explicit BYOK. Local custom fallback remains available.
Managed Jev is currently `not_yet_available` pending release acceptance and
service-capacity qualification; client configuration does not enable it. After
enablement, every legitimate Pro user and each eligible Team named seat, including paid
Expand All @@ -221,12 +226,25 @@ questions leave each rolling window.
The existing production fleet guard remains 100 questions per day across the service. It
conflicts with the per-person rolling caps and may pause or reject requests earlier, so
resolve capacity and the fleet guard before launch. No latency, accuracy, or cost-saving
guarantee is established by configuration or a successful health check. Managed access
admits only `guard_command`, `classify_contradiction`, `verify_support`, and
`verify_completion`, with fixed question schemas and required context. Managed `custom`
questions and `query_planning` fail closed before credential refresh or network
guarantee is established by configuration or a successful health check.

Managed access admits only `guard_command`, `classify_contradiction`, `verify_support`,
and `verify_completion`, with fixed question schemas and required context. Managed
`custom` questions and `query_planning` fail closed before credential refresh or network
requests; explicit BYOK and local custom heuristics remain separate choices.

Admitted errors remain counted. Managed transport callers can preserve an explicit `request_key`
for a retry; duplicate admitted keys return 409 without a second provider call or usage increment.
No answer is stored for replay, and no automatic retry occurs. The MCP tool does not expose a
`request_key` parameter.

The direct Classic `engraphis_decide` tool permits authorized viewers and remains advisory.
It does not grant memory writes or administration. Smart discovery still classifies the action
as stateful because it can consume allowance, so `engraphis_execute_action` retains its admin
requirement. The private hosted Team tool catalog is unchanged. Managed service release
acceptance, service capacity, and live quality evaluation remain required.
Command decisions always return `allow_auto=false` and `escalate_to_user=true`, including
successful remote answers. Provider probability and category are advice, not shell authorization.
Local command labels are coarse: only one simple inspection command, without chaining, pipes,
substitution or file redirection, is `read_only`. Recognized destructive, history-rewriting,
exfiltrating or credential-file commands are `destructive_or_leak`; anything else, including a
Expand Down
Loading
Loading