Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -133,4 +133,5 @@ help:
.PHONY: dashboard-preview
dashboard-preview:
mkdir -p bin
DASHBOARD_PREVIEW_DIR=$(CURDIR)/bin go test ./internal/dashboard -run TestBuildHTML_KeeperIncidentPreview -count=1 >/dev/null && echo "bin/keeper_incident_preview.html"
DASHBOARD_PREVIEW_DIR=$(CURDIR)/bin go test ./internal/dashboard -run 'TestBuildHTML_(KeeperIncidentPreview|AlertsPreview)' -count=1 >/dev/null \
&& echo "bin/keeper_incident_preview.html" && echo "bin/alerts_preview.html"
6 changes: 3 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -896,7 +896,7 @@ When `-skip-dashboard` is not set, the tool generates a single self-contained `d

| # | Section | What it shows |
|---|---|---|
| 1 | 🚨 **Alert Summary** | Fired alerts grouped by severity (critical / warning / info), with the row-level message template expanded per instance. Rules that **could not run** appear with a ⚠ marker and a separate "Could not run" chip — they are never counted in the severity badge. Rules that are **not applicable** here are listed in a muted footnote. A green "no issues" banner appears when nothing fired. |
| 1 | 🚨 **Alert Summary** | Fired alerts grouped by severity (critical / warning / info), with the row-level message template expanded per instance. Long content collapses: a message is shown up to the ClickHouse stack trace (a real one runs 1700+ characters over 15 lines, of which ~200 are the error) and capped at 260 characters, with the complete text one click away under *full message and stack trace*; only the first 5 instances are listed, the rest behind *N more instances*; and a rule description shows its first paragraph, the rest under *more about this rule*. Nothing is dropped from the page — `DATA.alerts` still carries every row in full. Rules that **could not run** appear with a ⚠ marker and a separate "Could not run" chip — they are never counted in the severity badge. Rules that are **not applicable** here are listed in a muted footnote. A green "no issues" banner appears when nothing fired. |
| 2 | 📈 **Overview** | Top-level counters: server version, uptime, total databases, total tables, active parts, total size |
| 3 | 📦 **Storage** | Size by database (horizontal bar), table-engine distribution (doughnut), and a top-20-by-size table list |
| 4 | 📋 **Tables Explorer** | Searchable / paginated table of every user table with engine, parts, rows, size, partition / sorting keys, and storage policy |
Expand All @@ -921,11 +921,11 @@ A sticky top nav at the page header lets you jump straight to any section. Secti

### Previewing the Keeper Health panel without an outage

`make dashboard-preview` renders `bin/keeper_incident_preview.html` from an anonymised fixture shaped like a real Keeper outage on a shared-storage cluster: 48 hours of Keeper counters (a blip on day one, quorum lost for eight hours on day two), the `keeper_health`, `keeper_connection_blips`, `merges_stalled`, `background_operation_failures`, `high_exception_rate` and `too_many_parts` alerts as they would fire, the error codes per hour and a re-established Keeper session. Use it to see what the panel and the alerts look like before an incident, or to review a theme or wording change.
`make dashboard-preview` renders two pages from anonymised fixtures. `bin/alerts_preview.html` exercises every Alert Summary collapse case — a message carrying a stack trace, more instances than the inline cap, a short message that needs no disclosure, a rule whose query failed, a rule that was not applicable. `bin/keeper_incident_preview.html` is shaped like a real Keeper outage on a shared-storage cluster: 48 hours of Keeper counters (a blip on day one, quorum lost for eight hours on day two), the `keeper_health`, `keeper_connection_blips`, `merges_stalled`, `background_operation_failures`, `high_exception_rate` and `too_many_parts` alerts as they would fire, the error codes per hour and a re-established Keeper session. Use it to see what the panel and the alerts look like before an incident, or to review a theme or wording change.

### What's interactive vs static

- **Interactive**: Tables Explorer (full text search, database/engine filters, pagination); all charts (hover tooltips, legend toggling).
- **Interactive**: Tables Explorer (full text search, database/engine filters, pagination); all charts (hover tooltips, legend toggling); the Alert Summary disclosures (native `<details>`, so they work without JavaScript and survive `Ctrl-F` only when open).
- **Static**: every other table — they render in a fixed order, but their underlying JSON is embedded in the page so you can `grep DATA dashboard.html | head` if you want raw values.

## Configuration Collection
Expand Down
88 changes: 88 additions & 0 deletions internal/dashboard/alerts_panel_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
package dashboard

import (
"strings"
"testing"
)

// An alert message is the rule's template with one row's values substituted,
// and those values routinely carry a full ClickHouse exception: one measured
// in a real bundle was 1764 characters over 15 lines, of which the first 216
// were the error itself. replication_queue_errors returns up to 50 such rows
// and keeper_health up to 168, so rendering every message in full turned the
// panel into hundreds of lines of stack frames. These are the bounds that
// keep the panel readable; a refactor must not drop them.
func TestTemplate_AlertMessagesCollapseVerboseParts(t *testing.T) {
for _, want := range []string{
"const ALERT_HEAD_CHARS=260;", // inline length of one instance
"const ALERT_ROWS_SHOWN=5;", // instances before the rest collapse
"const ALERT_STACK_RE=", // the stack trace is the cut point
`Stack trace \(when copying this message`, // ClickHouse's own marker, matched literally
"function alertMessageParts(msg){", // head / full split
"function alertDisclosure(parts){", // the <details> toggle
"function alertRowLine(a,row){", // one instance line
`'<details class="alert-more">`, // native disclosure, no JS needed
`'<pre class="alert-full">'+esc(parts.full)+`, // full text is kept, and escaped
"more instance", // the extra-rows toggle
"more about this rule", // the description's second half

// Two off-by-one traps found in review. The cap is INCLUSIVE of the
// ellipsis, so the slice has to be one short — otherwise the no-space
// fallback emits 261 characters. And the toggle compares the inline
// head against the whole flat message, not a version-stripped copy:
// with the stripped copy, a message that is NOTHING BUT a version
// suffix strips to empty, the fallback restores it, and the mismatch
// renders a disclosure whose body repeats the line above it.
"head.slice(0,ALERT_HEAD_CHARS-1)",
"truncated:head!==flat,",
} {
if !strings.Contains(htmlTemplate, want) {
t.Errorf("alert panel lost %q", want)
}
}

// The head must never be emitted raw: every branch that prints customer
// text goes through esc().
for _, banned := range []string{
"+parts.head+", "+p.head+", "+ep.head+", "+msg+'</li>", "+a.error+'",
// The pre-fix forms of the two traps above.
"head.slice(0,ALERT_HEAD_CHARS)+", "truncated:head!==flatTrimmed",
} {
if strings.Contains(htmlTemplate, banned) {
t.Errorf("alert panel interpolates customer text unescaped: %q", banned)
}
}

// A disclosure that is always open helps nobody; the toggle must be
// closed by default (no `open` attribute on the element we emit).
if strings.Contains(htmlTemplate, `<details class="alert-more" open`) {
t.Error("alert disclosures must start closed")
}
}

// The collapse only pays off if the full text is still in the page — support
// needs the stack trace, just not on screen by default.
func TestBuildHTML_AlertStackTraceIsKeptButCollapsed(t *testing.T) {
trace := "Code: 221. DB::Exception: No interserver IO endpoint named SharedMergeTreePartsUpdate:/clickhouse/tables/<uuid>/default/virtual_parts/node-07. " +
"(NO_SUCH_INTERSERVER_IO_ENDPOINT) (version 26.2.1.390 (official build)), Stack trace (when copying this message, always include the lines below):\n\n" +
"0. ./ci/tmp/build/./src/Common/Exception.cpp:141:1: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x000000001373469f\n" +
"1. DB::Exception::Exception(String&&, int, String, bool) @ 0x000000000cbeeb0e\n"
data := map[string]interface{}{
"generated_at": "2026-09-25 10:00:00 UTC", "mode": "onprem", "version": "26.2.1.390",
"alerts": []map[string]interface{}{{
"name": "replication_queue_errors", "title": "Replication queue entries have exceptions",
"severity": "critical", "file": "replication_queue_errors.yaml",
"message": "{database}.{table} (replica {replica_name}): {type} failed after {num_tries} tries — {last_exception}",
"rows": []map[string]interface{}{
{"database": "demo", "table": "events", "replica_name": "r-07", "type": "GET_PART", "num_tries": 128, "last_exception": trace},
},
}},
}
html := buildHTML(data)
if !strings.Contains(html, "NO_SUCH_INTERSERVER_IO_ENDPOINT") {
t.Error("the full exception must still be embedded — support reads the trace after expanding")
}
if !strings.Contains(html, "Exception.cpp:141") {
t.Error("stack frames must survive into the payload")
}
}
79 changes: 79 additions & 0 deletions internal/dashboard/alerts_preview_test.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
package dashboard

import (
"os"
"path/filepath"
"testing"

"clickhouse-diagnostic/internal/alert"
)

// TestBuildHTML_AlertsPreview renders an alerts page covering every collapse
// case the panel has to handle — a message with a stack trace, more instances
// than the inline cap, a short message that needs no disclosure at all, a rule
// whose own query failed, and a rule that was not applicable — so the
// interaction can be reviewed by eye rather than by reading the template.
// `make dashboard-preview` writes it to bin/alerts_preview.html.
func TestBuildHTML_AlertsPreview(t *testing.T) {
dir := os.Getenv("DASHBOARD_PREVIEW_DIR")
if dir == "" {
t.Skip("set DASHBOARD_PREVIEW_DIR to render")
}
trace := "Code: 999. Coordination::Exception: Session expired. (KEEPER_EXCEPTION) (version 26.2.1.390 (official build)), " +
"Stack trace (when copying this message, always include the lines below):\n\n" +
"0. ./ci/tmp/build/./src/Common/Exception.cpp:141:1: DB::Exception::Exception(DB::Exception::MessageMasked&&, int, bool) @ 0x000000001373469f\n" +
"1. ./src/Common/ZooKeeper/ZooKeeperImpl.cpp:1071: Coordination::ZooKeeper::pushRequest(Coordination::ZooKeeper::RequestInfo&&) @ 0x000000001a2b3c4d\n" +
"2. ./src/Common/ZooKeeper/ZooKeeper.cpp:412: zkutil::ZooKeeper::multiImpl(std::vector<Coordination::RequestPtr> const&, ...) @ 0x000000001a2b9f10\n" +
"3. ./src/Storages/MergeTree/ReplicatedMergeTreeQueue.cpp:1188: DB::ReplicatedMergeTreeQueue::processEntry(...) @ 0x000000001c0d4a22\n" +
"4. ./src/Storages/StorageReplicatedMergeTree.cpp:3702: DB::StorageReplicatedMergeTree::processQueueEntry(...) @ 0x000000001bf51188\n" +
"5. ./src/Storages/MergeTree/MergeTreeBackgroundExecutor.cpp:288: DB::MergeTreeBackgroundExecutor<...>::threadFunction() @ 0x000000001c113d90\n" +
"6. ./base/poco/Foundation/src/ThreadPool.cpp:205:14: Poco::PooledThread::run() @ 0x000000002334b505\n" +
"7. ./base/poco/Foundation/src/Thread_POSIX.cpp:335:5: Poco::ThreadImpl::runnableEntry(void*) @ 0x0000000023348f01\n"

queue := []map[string]interface{}{}
for i, tbl := range []string{"events_local", "device_log", "sensor_raw", "job_history", "audit_trail", "shift_report", "line_state"} {
queue = append(queue, map[string]interface{}{
"database": "demo_app", "table": tbl, "replica_name": "r-0" + string(rune('1'+i)),
"type": "GET_PART", "num_tries": 40 + i*17, "last_exception": trace,
})
}
keeper := []map[string]interface{}{}
for h := 13; h <= 20; h++ {
keeper = append(keeper, map[string]interface{}{
"hour": "2026-09-10 " + string(rune('0'+h/10)) + string(rune('0'+h%10)) + ":00:00", "hostname": "node-07",
"hw_exceptions": 10647545 - h*1000, "transactions": 11033023, "usual_transactions": 28010220, "pct_of_usual": 39 - h,
})
}

data := map[string]interface{}{
"generated_at": "2026-09-25 10:00:00 UTC", "mode": "onprem", "version": "26.2.1.390",
"uptime": "1 hours 51 minutes", "total_databases": 410, "total_tables": 33882, "active_parts": 50000, "total_size": "28.40 TiB",
"alerts": []alert.Result{
// 1. long message (stack trace) AND more instances than the cap
{Name: "replication_queue_errors", Title: "Replication queue entries have exceptions", Severity: "critical", File: "replication_queue_errors.yaml",
Description: "A replication queue entry has a non-empty last_exception: the replica tried to execute it and the server refused.\n\nCheck: the exception text and num_tries — a high count with the same error is a stuck entry, not a slow one. Then system.replicas for the same table, and the source replica named in the message.",
Message: "{database}.{table} (replica {replica_name}): {type} failed after {num_tries} tries — {last_exception}",
Rows: queue},
// 2. short messages, many instances → only the row cap applies
{Name: "keeper_health", Title: "Keeper unavailable: session loss storm while Keeper traffic collapsed", Severity: "critical", File: "keeper_health.yaml",
Description: "The two-signal Keeper health test over the last 7 days, per hour.\n\nCheck next: system.zookeeper_connection (session age), part_log errors in the same hours, then the Keeper nodes themselves.",
Message: "hour starting {hour} on {hostname}: {hw_exceptions} Keeper hardware exceptions, {transactions} Keeper transactions = {pct_of_usual}% of usual ({usual_transactions}/h median) — Keeper effectively unavailable to this server",
Rows: keeper},
// 3. the control case: one short instance, single-paragraph description → no toggles at all
{Name: "too_many_parts", Title: "Table partition approaching too-many-parts limit", Severity: "warning", File: "too_many_parts.yaml",
Description: "A table partition has more than 300 active parts.",
Message: "{database}.{table} partition '{partition_id}' has {parts_count} active parts — inserts are delayed from 1000 parts (parts_to_delay_insert) and rejected with code 252 TOO_MANY_PARTS at 3000 (parts_to_throw_insert)",
Rows: []map[string]interface{}{{"database": "demo_app", "table": "events_summary", "partition_id": "202609", "parts_count": 3469}}},
// 4. a rule whose own query failed, with a trace in the error text
{Name: "detached_parts_exist", Title: "Detached parts present", Severity: "info", File: "detached_parts_exist.yaml",
Description: "Parts exist in the detached/ folder.",
Error: "error executing query: non-OK status: 500, body: " + trace},
// 5. not applicable here
{Name: "crash_log_entries", Title: "Server crash detected", Severity: "critical", File: "crash_log_entries.yaml", Skipped: true, Reason: "table not present"},
},
}
if err := os.WriteFile(filepath.Join(dir, "alerts_preview.html"), []byte(buildHTML(data)), 0o644); err != nil {
t.Fatal(err)
}
t.Logf("written: %s", filepath.Join(dir, "alerts_preview.html"))
}
Loading
Loading