Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/crisp-triage.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ name: Crisp triage

on:
schedule:
- cron: "0 */4 * * *"
- cron: "0 */3 * * *"
workflow_dispatch:
inputs:
skip_dedupe_check:
Expand Down
28 changes: 24 additions & 4 deletions PHASE2-SETUP.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,14 +65,18 @@ Phase 1's PAT has **Contents: Read-only**. Stage 1 needs to commit the advanced
**Policy: a resolved conversation is trusted as fully handled by support.** Resolving it — first time or the tenth time — never triggers investigation on its own; the resolved-conversation loop in `crisp-classify.mjs` only records "session X seen resolved at time T" into `state/resolved-seen.json`, nothing more. Three things can still trigger a full investigation:

- **Manual**: a support agent adds a private note containing `!tg-autopilot investigate`. Skips the cheap classifier entirely — a human already made the call — and goes straight to full investigation, on the *full* transcript regardless of any of the below. Works any number of times; each *new* note re-triggers it (tracked by counting matching notes per conversation, not by trying to identify "which" note, since two notes can have identical text). Detection has two layers: the cheap path only fetches full messages for conversations already touched since `cursor.last_checked` (adding a note is itself an update); separately, `searchConversationsForManualTrigger()` searches Crisp directly for the trigger phrase and appends anything found there too, regardless of `fetchActiveConversations`' page cap. The second layer exists because the first one alone isn't reliable on a high-volume account — confirmed for real on `THEMEGRILL`: a manually-noted conversation didn't rank in the top 200 most-recently-updated active conversations because 200+ *other* conversations were touched in the same window, so it never got its message history fetched at all. A manual trigger is an explicit human action and shouldn't ever silently fail to reach the pipeline just because the account is busy.
- **Reopen after resolve**: a conversation `state/resolved-seen.json` already has a record for shows up in the *active* list again (the customer replied after it was marked resolved). Investigated using only the messages timestamped after that recorded resolve time — whatever was there before is assumed already covered by support's original resolve. If it gets resolved again before a `crisp-triage` run ever catches it in this in-between "active again" state, it's never investigated for that reopen at all — the resolved-conversation loop just records the newer resolve time on its next pass, and there's nothing left to catch. This is intentional, not a bug: by the same policy, a second resolve is trusted as handled too.
- **Automatic, stale-and-never-resolved**: a conversation open longer than **12 hours** (measured from `active.last`, not `created_at`) that has *never* been resolved at all gets checked by the same cheap classifier. If it agrees this looks like a real bug/feature, it's escalated the same way. This fires **at most once per conversation** automatically, and never for a conversation older than **30 days** (`AUTO_ESCALATE_MAX_HOURS` in `crisp-classify.mjs`) -- a backlog that's sat untouched that long is treated as intentionally left open, not a scan miss. This path is unrelated to the reopen path above — it exists for tickets that were *never* marked resolved, which the "trust the resolve" policy says nothing about.
- **Reopen after resolve**: a conversation `state/resolved-seen.json` already has a record for shows up in the *active* list again (the customer replied after it was marked resolved). Investigated using only the messages timestamped after that recorded resolve time, paired with the conversation's opening messages for context (capped to a few, not the whole pre-resolve history) — otherwise the delta alone reads as routine internal follow-up with no idea what bug it's even about. If the delta is empty (nothing new since the resolve) *and* the conversation is independently stale-eligible (see below), it falls back to a full-history stale check instead of silently skipping — added after a real case (`session_bd0acc7b`) sat stale for days with a never-addressed original report, because a resolve/reopen flag alone was masking the independent staleness check. If it gets resolved again before a `crisp-triage` run ever catches it in this in-between "active again" state, it's never investigated for that reopen at all — the resolved-conversation loop just records the newer resolve time on its next pass, and there's nothing left to catch. This is intentional, not a bug: by the same policy, a second resolve is trusted as handled too.
- **Automatic, stale-and-never-resolved**: a conversation open longer than **12 hours** (measured from `active.last`, not `created_at`) that has *never* been resolved at all gets checked by the same cheap classifier. If it agrees this looks like a real bug/feature, it's escalated the same way. Re-fires whenever genuinely new message content arrives after the last check — tracked via `checkedThroughAt` in `state/escalated.json`, a timestamp taken from the actual fetched messages, **not** conversation-level metadata (see the warning below on why that distinction matters). It's never eligible in the first place for a conversation older than **30 days** (`AUTO_ESCALATE_MAX_HOURS` in `crisp-classify.mjs`) -- a backlog that's sat untouched that long is treated as intentionally left open, not a scan miss. This path is unrelated to the reopen path above — it exists for tickets that were *never* marked resolved, which the "trust the resolve" policy says nothing about.

**Onboarding a new account: run `Seed escalated state for a new Crisp account` (`.github/workflows/seed-escalated.yml`, `workflow_dispatch`, input = the account key) once, before that account's first scheduled `crisp-triage` run.** `escalated.json` has no history for a brand-new account, so every one of its currently-active conversations already past 12 hours old looks auto-escalation-eligible on day one -- confirmed for real onboarding User Registration: ~83% of its ~400-conversation active backlog auto-escalated within the first few minutes of a single run. The seed workflow marks the current backlog as already-escalated so none of it fires; only conversations that go stale *after* seeding will auto-escalate from then on. It doesn't touch the manual-note path -- `!tg-autopilot investigate` still works on any of those backlog conversations if someone wants one looked at anyway.
**Onboarding a new account: run `Seed escalated state for a new Crisp account` (`.github/workflows/seed-escalated.yml`, `workflow_dispatch`, input = the account key) once, before that account's first scheduled `crisp-triage` run.** `escalated.json` has no history for a brand-new account, so every one of its currently-active conversations already past 12 hours old looks auto-escalation-eligible on day one -- confirmed for real onboarding User Registration: ~83% of its ~400-conversation active backlog auto-escalated within the first few minutes of a single run. The seed workflow marks the current backlog as already-checked (`checkedThroughAt = now`, from real message timestamps — see below) so none of it fires; only conversations that get new messages *after* seeding will auto-escalate from then on. It doesn't touch the manual-note path -- `!tg-autopilot investigate` still works on any of those backlog conversations if someone wants one looked at anyway. Pass `force: true` to recompute and overwrite every entry instead of only unset ones — needed to correct a prior bad seed, not for normal onboarding.

All three paths write to `state/escalated.json` and/or `state/resolved-seen.json` (both committed) so none of them repeats itself needlessly.

**Separately**, `state/investigated.json` (also committed) tracks every session_id that has ever completed a full Stage 2 investigation, from *any* path. The reopen and stale-auto-escalation branches both skip a session already in this set for the *same* unresolved content — a conversation reopening with a genuinely new, unrelated problem after already being investigated once still gets picked up, since the reopen path is driven by `resolved-seen.json`'s timestamp, not this set. The manual-note path deliberately does **not** check this set — a human explicitly asking to (re-)investigate should always go through.
**⚠️ Conversation-level metadata (`active.last`, `updated_at`) can lag real message activity.** Confirmed for real on an email-origin conversation with live replies well after both fields had stopped moving. This is why every "has anything new happened?" check in `crisp-classify.mjs` and `seed-escalated.mjs` is keyed off the actual fetched messages' own timestamps (`checkedThroughAt`), never off the conversation object's own `active.last`/`updated_at`/`created_at` — those fields are only trusted for the cheap, non-final pre-filter (is it even worth fetching this conversation's messages at all). **Known residual gap, not yet fixed:** the 12h–720h staleness *window* itself (`eligibleForAutoEscalate`'s bounds) still computes `staleHours` from that same unreliable metadata — a conversation with genuinely fresh real activity but a frozen `active.last` could still misjudge its own eligibility window (e.g. look falsely past the 30-day cutoff). Fixing this properly means fetching real messages before deciding eligibility at all, not just after, which is a bigger change than today's fix — left as a follow-up.

**Separately**, `state/investigated.json` (also committed) is a write-only audit trail of every session_id that has ever completed a full Stage 2 investigation, from *any* path — it no longer gates anything (see the warning below). The manual-note path deliberately does **not** check `checkedThroughAt` either — a human explicitly asking to (re-)investigate should always go through.

**⚠️ Past bug, fixed 2026-09-29 — do not reintroduce:** `investigated.has(session_id)` and an `escalated[session_id].autoEscalated` boolean used to permanently block the stale-auto-escalate path once a session was ever investigated, with no way to clear — confirmed for real (`session_2fc63232`) where repeated resolve/reopen cycles with brand-new problems stopped producing issues entirely, forever, after the first one. Replaced with `checkedThroughAt`, which re-arms automatically the moment real new content exists. If you're tempted to add a permanent "already handled" flag anywhere in this pipeline again, re-read this paragraph first.

**Rollout note**: `state/resolved-seen.json` starts empty when this policy first ships. A conversation that was resolved *before* that point and reopens shortly after has no prior "seen resolved" record yet, so that first reopen looks like a first-time resolve and gets skipped rather than recognized as a reopen — it self-corrects from its next resolve onward. Accepted as a one-time gap rather than backfilling the whole history.

Expand All @@ -96,3 +100,19 @@ ThemeIsle's own numbers: $0.0003 per conversation when Stage 1 (or a quick Stage
2. **Dedupe**: resolve a second conversation describing the *same* bug → confirm a comment on the existing issue, not a duplicate.
3. **Client-side note-back**: resolve a conversation that's clearly a client-side misunderstanding → confirm a note lands in that Crisp conversation and no GitHub issue is filed.
4. **Cursor correctness**: run again with no new conversations → confirm nothing gets re-processed.

## Changelog

**2026-09-29 — permanent-block bug fix, real-timestamp freshness, cron/debug tooling**

Two real chat sessions (`session_2fc63232`, `session_bd0acc7b`) stopped producing issues despite genuine, never-addressed problems. Root cause and fix:

- `investigated.has(session_id)` / `escalated[session_id].autoEscalated` were permanent flags that never cleared, so a session auto-escalated once could never be auto-escalated again, ever — even after resolving and reopening with a brand new problem weeks later. Replaced with `checkedThroughAt` (timestamp from real fetched messages, not conversation metadata) — see § 4c above.
- A reopen with nothing new since a (possibly spurious) resolve took the reopen branch, found an empty delta, and skipped forever instead of falling through to the independently-stale check. Now falls back to a full-history stale check.
- Along the way, confirmed `active.last`/`updated_at` can lag real message activity for at least some conversations — the reason `checkedThroughAt` is deliberately message-based, not metadata-based. The 12h–720h staleness *window* itself still uses that metadata and is a known, unfixed residual gap (see the ⚠️ warning in § 4c).
- `seed-escalated.mjs` had the same metadata-vs-real-timestamp bug on first attempt (seeded from `updated_at`, which undercounted real activity for a large fraction of a backlog and caused it to immediately re-fire) — fixed to fetch real messages instead, and given a `--force` flag to recompute existing entries.
- All 3 accounts' current backlog was re-seeded with real-timestamp `checkedThroughAt` on 2026-09-29, so today's fix applies to new activity going forward rather than re-litigating months of history in one burst.
- Added `skip_dedupe_check` (workflow_dispatch input on `crisp-triage.yml`) to skip the slow "check active conversations for duplicates" step (~15–20 min, unrelated to classify logic) for fast manual debugging. **Caveat learned the hard way: if `matrix` isn't empty, the downstream `investigate` job still auto-fires immediately after classify finishes** — a "fast test" run isn't safe to leave unattended once real conversations match, and must be watched/cancelled if you don't want real issues filed.
- Cron cadence changed a few times this session while investigating a low signal-to-noise cadence: 5h (original) → 3h → 4h → 3h (temporary, as of this writing, to watch results more closely for a day).

PRs: themegrill/.github#91, #92 (superseded, see below), #93 (accidentally a no-op — opened from a stale branch, merge commit had zero file changes, learned to always verify a branch's actual pushed content before opening a PR from it), #94, #95 (the real fix), #96.
Loading