✨ feat: Add the Classification Port - #561
danny-avila wants to merge 12 commits into
Conversation
A typed question in, a calibrated answer out: `src/classification/` carries the port that
LibreChat PR #16180 introduced under `packages/api` and that codegraph mirrors in ESM, so the
product, the graph and any other consumer share one implementation of the contract a System One
host (TypeSafe's Jev, directly or through a gateway) answers.
- types: boolean / choice / score questions, answers with a probability or a calibrated
confidence and distribution, `Classifier`, `ClassificationError` with typed failures,
`ClassificationDialect`, `ClassificationProviderSettings`
- dialect: boolean ↔ `noul`; a string yes-criterion becomes the `{true}` pair a System One host wants
- transport: one deadline for the whole call, bounded retries on 429/5xx/network honouring
retry-after, an `onAnswered` hook instead of a logger dependency
- http: the host over HTTP, with request/response wrapping for hosts that nest the envelope
- presets: typesafe, openrouter, cloudflare, http; `createClassifier(settings, apiKey)`
- questions: `booleanQuestion`, `choiceQuestion`, `scoreQuestion`
- seven jest tests with a fake fetch; `tsc --noEmit` clean
No LibreChat type is imported: the SDK holds the port, consumers hold their configuration.
…assifier A host whose bearer expires (the ClickHouse inference gateway mints an hourly Okta token) can be given a function instead of a key. The transport calls it before each request and once more with refresh: true after a 401, then retries that request; a 403 is a scope refusal and is not retried. The `clickhouse` preset points at the gateway's System One route.
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e695493310
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Self-review handoff for PR #561 at exact pushed head This head resolves all four inline findings on the earlier head: immutable presets across tenants, cached and refreshed per-call credentials across retries, nested classifier-prompt tool-output redaction before Langfuse export, and complete normalized measured distributions. It also moves structured-chat validation inside the abortable deadline. Tests cover each previously failing case. Local checks passed: 27 classification tests, 204 tracing tests, full workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, and package build. CI for this exact head failed before creating jobs because the current main reusable workflow has duplicate YAML keys. No CI checks ran on this head. A maintainer can trigger a new Codex review for this SHA if desired; a Lia GitHub App comment cannot trigger one. |
|
Review handoff for draft agents PR #561 at exact remote head Further invariant review found and fixed two issues missed at the preceding head: private tool results copied into free-form classifier state or question text escaped selective Langfuse redaction, and a monotonic timeout could settle before its timer callback fired without aborting the fetch signal. Marked classifier prompts now fail closed in traces under any active tool-output redaction policy. Deadline checks now abort the in-flight signal and observe late rejected tasks, including when synchronous preparation crosses the deadline. No request content or provider behavior is changed by trace redaction. Local verification on this head: 234 passed tests across 13 focused classification and Langfuse suites, workspace TypeScript typecheck, touched-file lint/import order/formatting, circular-dependency check, ESM/CJS exports, package build, and diff checks. CI for this head failed before creating any jobs; the current main workflow has duplicate YAML keys. A maintainer can trigger a new Codex review for this exact SHA. A Lia GitHub App comment does not initiate that review. |
|
@codex review the latest head |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0ac0cf9d52
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review the latest head |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5fccb734ed
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Review handoff for exact remote head This head incorporates main Verified local checks:
CI for this exact head passed all 13 validation jobs, including the Anthropic summarization lane. Independent review of this exact head remains in progress. Earlier reviews do not cover this head. Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published. |
|
Final verification for exact remote and independently reviewed head Ready to merge. CI for this head passed all 13 jobs. Fresh independent review completed with no new findings and confirmed all ledger fixes.
The prior eight inline fixes were rechecked and retained. All ten GitHub threads are resolved. No findings were rejected. Local checks on this exact head passed: 291 tests across 14 focused classification/tracing suites; workspace Strict chat support is explicitly bounded to verified OpenAI modes (including Azure's inherited implementation) and Anthropic strict tool calling. Bedrock and other unverified adapters fail locally instead of silently downgrading. The HTTP Jev/Laya adapter is unchanged. Not run: live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests. Nothing merged or published. Downstream port migration and consumers remain separate PRs. |
Summary
Keep
Classifieras a small SDK contract rather than a chat-model subclass. Jev and self-hosted Laya use the same System One HTTP adapter; an already configured chat model can be injected through a separate strict structured-output adapter. Laya is a data-only preset with an operator-supplied endpoint, optional bearer authentication, and no forced checkpoint.Semantics and safety
probability: number. A chat-only boolean has{ decision: boolean, probability: null }; choice distributions and token usage arenullwhen unmeasured. Missing HTTP answers are typed asundefined, while malformed or unexpected answers fail explicitly.confidencedefinitions, so thresholds cannot be transferred without evaluating the actual checkpoint and use case.Verification and review
Current pushed head:
f38061e8e1de5eab8c51b9d90ab1b647946d8770, incorporating main64177c706f69d08a38565525437f896e8cfd0b14and package version4.0.0.Verified locally on this head:
npx jest src/classification langfuse deterministic-trace-id activity-label-trace-masking --runInBand --silentnpx tsc --noEmit -p tsconfig.json --pretty falsegit diff --checkCI for this exact head passed all 13 validation jobs. Independent review of
f38061e8e1de5eab8c51b9d90ab1b647946d8770completed with no new findings. The reviewer verified 57 native ledger assertions and focused transport invariants; it did not rerun Jest or live provider pipelines in its isolated lane. The parent task's 291 focused tests exercised the real locked LangChain pipelines with controlled provider responses. Earlier CI/reviews are not used as coverage of this head.Final finding ledger:
7411317e; final review confirmed7411317e; final review confirmedf38061e8; final review confirmedf38061e8; final review confirmedThe earlier eight inline fixes remain present and were independently rechecked. All ten GitHub review threads are resolved. No finding was rejected. This PR is ready to merge; no merge or publication has been performed.
The last regression subset reproduced 11 failures before the current fixes. Independent review found two P2 defects at
7411317e: Bedrock's locked adapter silently ignored strict mode, and exported score parsers accepted nonnumeric keys without expected question metadata. Both are fixed inf38061e8with regressions. The verified strict adapters are OpenAI (including Azure's inherited implementation) in either mode and Anthropic in strict tool-calling mode; unverified adapters fail locally withunsupported_mode. Score distributions always require canonical consecutive numeric levels starting at zero, independent of an optional expected question.Self-review follow-up
Four inline review findings from
e69549331079ccdd0416bd7060110cc8221ebce4were reproduced and fixed ind46ae34835e4c73d05aafe85fdc93c14a4088938: presets and their registry are immutable across tenants; per-call bearer minting and the one 401 refresh persist over retries; marked structured-chat prompts keep the original provider request but preserve Langfuse tool-output redaction before trace export; measured choice/score distributions are complete and normalized within a bounded tolerance. Independent review also moved strict-chat question preparation inside the request deadline, so pre-aborted calls never inspect dynamic questions. New regressions cover all five issues.Further invariant review found two gaps fixed in
f88615413542161cd8db44dd7bf406bff3f47b0d. A private tool result copied into free-form classifier state or instructions has no tool identity after stringification, so selective field redaction could export it. The trace processor now drops the whole marked classifier prompt under any active tool-output redaction policy (provider requests are unchanged); a regression first reproduced the leak. A monotonic deadline can also expire before the timer callback runs, returning while fetch stays active. The deadline now aborts its controller on expiry, honors expiry during synchronous preparation errors, and observes late promise rejections. Regressions reproduced both paths before the fix.The latest Codex review on
0ac0cf9didentified four further correctness gaps. Commits953216cband5fccb734address them: a measured score must match its rounded distribution; structured-chat failures preserve sanitized HTTP categories and Retry-After; Bedrock cache token counts are included exactly once; and HTTP/chat validation use per-call snapshots of the questions sent. New regressions exercise real OpenAI and Anthropic failure paths, Bedrock usage metadata, and delayed responses during caller mutation.The two remaining Codex findings on
5fccb734are addressed in7411317e: measured choices must select a maximum-probability option (ties are valid), and malformed boolean criteria fail locally before credential minting or HTTP/chat provider invocation. Regressions cover both wire dialects, missing choices, ties, unmeasured choices, rejected criteria, and supported string/one-sided/structured-text criteria. Choice consistency is checked during the existing probability validation pass.Rollout
Merge and publish this shared SDK separately. Migrate the duplicate LibreChat port in #16180, then adjust the probability consumers and fallbacks in #16181 separately. Live Jev/Laya comparisons, calibration, latency/cost evaluation, live Langfuse verification, and LibreChat consumer tests have not been run in this agents PR. No merge or package publication has been performed.