Skip to content

dashboard: table-dependency graph (ported from ClickHouse /schema), DDL credential redaction - #34

Merged
CamiloSierraH merged 9 commits into
mainfrom
cami/schema-dependency-graph
Sep 28, 2026
Merged

CamiloSierraH merged 9 commits into
mainfrom
cami/schema-dependency-graph

Conversation

@CamiloSierraH

@CamiloSierraH CamiloSierraH commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

What

Schema graph — one database section, nodes coloured by engine, MV and dictionary edges, columns per node

A Schema Graph tab in dashboard.html that opens schema_graph.html — the table-dependency graph from ClickHouse's built-in /schema page (programs/server/schema.html, Apache-2.0, attributed), rendered entirely from data embedded at collection time. Tables, materialized views, refreshable MVs, dictionaries, Distributed tables and views, coloured by engine, with the edges data flows along; a sidebar with keys, columns, neighbours and the CREATE statement; search over table and column names; drag, zoom, database filter. Light and dark follow the dashboard.

Three commits, each reviewable alone:

  1. Redact credentials from DDL columns — a pre-existing leak, wider than this feature.
  2. Collectors — the columns the graph needs, plus system.view_refreshes.
  3. The graph page and the tab.

1. The leak this fixes on the way

create_table_query, engine_full, as_select and system.dictionaries.source carry engine arguments verbatim, and the config sanitizer was wired to the XML/YAML collector only — query results went to disk as returned. Verified on 22.8.21.38: the S3 secret key and the MySQL password ship raw in today's system.tables JSONL. Servers from 23.x mask the positional secrets as '[HIDDEN]' themselves but leave AWS key ids and SETTINGS … password = … visible.

collection.RedactSQLText parses every known engine / table-function call (S3-family and the *Cluster variants, MySQL, PostgreSQL, MongoDB, remote(), Redis, Azure, ExternalDistributed; nesting and quoting respected, so a function inside an MV's SELECT is found), masks the credential positions, then runs the sanitizer's byte-shape heuristics. It writes the server's own [HIDDEN] token so old and new bundles read alike, and is idempotent over already-masked text. The executor applies it to the named string fields of the JSONL only where a value changed — every other byte, including quoted 64-bit integers and key order, is preserved. execution_log.txt records the count. Native/TSV can't be redacted field by field and say so.

Before/after on 22.8:

S3('https://…/data.csv', 'AKIAIOSFODNN7EXAMPLE', 'wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY', 'CSV')
S3('https://…/data.csv', '[HIDDEN]', '[HIDDEN]', 'CSV')

2. Collectors

system.columns + is_in_primary_key, is_in_sorting_key (all modes)
system.tables + total_rows, total_bytes on every rung; new 26.6.1.0/ rung with target_database, target_table
system.view_refreshes new, 23.12.1.0/ rung, no root file (table newer than every floor)
gov system.tables dependency arrays now collected, hashed element by element with arrayMap

target_table is the only source that names the implicit .inner_id.<uuid> target of an MV declared with an ENGINE. Pinned to 26.6 by reading StorageSystemTables.cpp at v26.5.7.64-stable (absent) and v26.6.1.1193-stable (present).

⚠️ Reviewer attention — a gov guard is loosened. dependencies_* / loading_dependencies_* move from gov_leak_test's forbidden list to mustBeHashed. An array of hashed table names exposes no more than the hashed database/name beside it, and it is what lets a gov analyst count MVs per source or follow a chain by hash. loading_dependent_* stay forbidden. If you disagree, dropping the four lines from the gov files reverts it without touching anything else.

3. Why a second file, and how the tab works

The graph needs every column of every user table embedded: ~750 KiB on my near-empty server, tens of MB for a service with thousands of tables — that would slow dashboard.html for every reader who never opens the graph. So it's a sibling file, loaded on demand.

Tested the alternatives in headless Chrome from file:// rather than assuming:

mechanism result
fetch() / XHR of a sibling file blocked (opaque origin)
<iframe src> renders
iframe.src assigned on click allowed — and the file is not requested until then
parent scripting the frame blocked
localStorage between the two documents shared — the theme toggle propagates live via the storage event

The theme tokens move out of htmlTemplate into themeTokensCSS; both pages are assembled from it. schema.html's palette is remapped onto those tokens; only the engine colours keep their values.

Payload: five safeQuery calls behind hasColumn("tables","target_table") / hasTable("view_refreshes"), plain system.* in every mode (tables and columns are replica-shared; dictionaries would duplicate under clusterAllReplicas), DDL redacted before embedding. json.Marshal HTML-escapes <, so a </script> inside a customer's DDL can't terminate the inline script — a test embeds one.

Deferred: the INSERT-pipeline heat map from system.query_views_log. It's a consumer of #33's hourly collector and lands once that merges. Gov never generates a dashboard, so gov needs no HTML work.

From human review

Four suggestions, all in: the RMV refresh schedule on the node ribbon (parsed from the DDL — view_refreshes has no interval column — and underlined when there is no RANDOMIZE FOR, so views that fire at the same instant stand out), the dictionary LIFETIME (columns, DDL fallback for an unloaded dictionary), a pretty-printed and highlighted CREATE (server-side formatQueryOrNull where available, a small span-emitting tokenizer on the page), and key roles differentiated — primary / sorting / partition / sampling as four independent flags with a legend and sidebar tags — plus skip indices, which gain a system.data_skipping_indices collector in all three modes and appear on the ribbon and in the sidebar. A follow-up from the same reviewer: refreshable MVs, plain views and Distributed tables have no dependency rows in system.tables, so their source edges are now inferred from the DDL (FROM/JOIN, Distributed engine arguments) and drawn dashed, with the sidebar saying so; metadata edges are never demoted.

Verification

  • make test, go vet ./..., gofmt clean. New tests: 18 DDL shapes for the redactor, JSONL byte-preservation, the executor hook, template pins (offline, attributed, no --link shadowing, innerHTML only static), on-demand loading, fixture embedding with </script> escaping, file-index exclusion.
  • Every new/changed SQL file executed against 22.8.21.38 and 26.7.5.10.
  • End to end on 26.7 with a schema holding every node kind (MergeTree source, MV with TO, MV with implicit target, view, dictionary, Distributed, refreshable MV): tab appears, frame src is null before the click, 22 nodes with correct kinds, sidebar shows CREATE and neighbours, theme follows the dashboard, standalone page shows the back link, search dims non-matches. Network blocked except the Chart.js CDN: zero other requests, zero page errors. dashboard.html 232 KB, schema_graph.html 85 KB.
  • End to end on 22.8: JSONL reads [HIDDEN], zero credential strings remain, execution_log.txt says 7 credential(s) replaced.
  • make dashboard-preview writes bin/schema_graph_preview.html (11 tables, all node kinds, a redacted S3 DDL), reachable from the preview dashboard's Schema tab.

Merge note

This branch adds HC-7.6 (refreshable MV stopped refreshing) to health-checks.md; #33 adds HC-7.6–7.8 for query_views_log. Whichever merges second renumbers its row — one line.

Review follow-up

Six findings on the first pass, all addressed in the fourth commit:

# Finding Verdict Fix
1 Keyword redaction stopped at an escaped quote — 'pa\'ss' → '[HIDDEN]'ss' real, the worst one value end found by scanning the literal with SQL escaping (\', '', \"); regressions added
2 Dashboard's own Dictionaries panel embedded source / last_exception raw real, same leak class, missed in collect() both fields through the redactor
3 WHERE status IN (…) dropped FAILED dictionaries right in spirit; the "disappear entirely" claim was wrong (nodes come from system.tables) every status read; sidebar shows status + error. The local server's dict_users had gone FAILED between runs and would have been hidden
4 view_refreshes read the local table while hasTable probed all replicas real g.sysTable(...) + fold to one row per view (argMax(status, status != 'RunningOnAnotherReplica')); verified on the cluster form
5 Dictionaries have no nodes unless in system.tables partly: DDL dictionaries always had nodes; config (XML) dictionaries didn't — same gap as the live page synthetic node under a (config) section, source in sidebar, no edges
6 Nodes not keyboard-accessible real role=button, tabIndex=0, aria-label, Enter/Space, focus ring, focusable sidebar links

The README gained a Schema graph section (before Dry-run mode) with this screenshot and a plain-language guide to reading the page; the design record stays under Dashboard → Schema graph.

Second pass (three more, all consistency gaps, fixed in the sixth commit): the S3 rule now masks the key id structurally as well as the secret (non-AWS-shaped ids — MinIO, GCS HMAC, R2, OSS — escaped the heuristic); system.dictionaries.last_exception is redacted in the persisted JSONL, not only in the dashboard and graph; and a server whose only user objects are config-declared dictionaries now gets a graph — the empty check runs after both system.tables and system.dictionaries are read.

Re-verified in headless Chrome against a fresh collection: Tab + Enter opens a node's sidebar, Escape closes it, dictionary nodes show their status, zero non-CDN requests, zero page errors.

🤖 Generated with Claude Code

CamiloSierraH and others added 3 commits September 25, 2026 18:28
system.tables.create_table_query / engine_full / as_select and
system.dictionaries.source carry engine arguments verbatim: an S3 secret
key, a MySQL password, kafka_sasl_password in a SETTINGS clause. The config
sanitizer never saw them — it is wired to the XML/YAML collector only, and
query results were written to disk as the server returned them. Verified on
22.8.21.38: both the S3 secret and the MySQL password ship raw. Servers from
23.x mask the positional secrets as '[HIDDEN]' themselves, but leave AWS key
ids and SETTINGS-clause passwords visible.

Adds collection.RedactSQLText: a small parser that walks every known engine
and table-function call (S3-family incl. the *Cluster variants, MySQL,
PostgreSQL, MongoDB, remote(), Redis, Azure, ExternalDistributed; nesting and
quoting respected, so a function inside an MV's SELECT is found too), masks
the argument positions that carry secrets, then runs the byte-shape
heuristics the config sanitizer already has — URL basic-auth, AWS key ids,
JWTs, PEM blocks, keyword = 'value'. The replacement is '[HIDDEN]', the
server's own token, so a 22.x bundle and a 26.x bundle read alike and every
downstream reader needs one convention. Idempotent over already-masked text.

The executor applies it to the JSONL of the collectors named in
SensitiveCollectorFields, rewriting only the named string fields and only
where a value changed — every other byte, including quoted 64-bit integers
and the server's key order, is preserved. Native/TSV cannot be redacted
field by field and are written as-is with a note in execution_log.txt; the
ok entry records how many values were replaced.

redactHeuristicsWith parameterises the sentinel so the config sanitizer keeps
writing REMOVED; a test pins that.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…freshes

system.columns gains is_in_primary_key / is_in_sorting_key in every mode
(booleans, safe in gov). system.tables gains total_rows / total_bytes on
every rung, and a new 26.6.1.0 rung adds target_database / target_table —
the table an MV writes to, including the implicit .inner_id.<uuid> table of
an MV declared with an ENGINE, which dependencies_* never names. Pinned to
26.6 by reading StorageSystemTables.cpp at v26.5.7.64-stable (absent) and
v26.6.1.1193-stable (present).

system.view_refreshes is new, at a 23.12.1.0 rung with no root file — the
table arrived with refreshable MVs in 23.12, newer than every floor, so it is
skipped rather than failing below that. A REFRESH EVERY view never appears in
query_views_log; this is the only place its schedule, last success and last
error live. Cloud reads every replica with hostName(); gov hashes database
and view and ships has_exception instead of the text.

Gov edges: dependencies_* / loading_dependencies_* used to be forbidden
outright in queries.gov. They are the edges of the graph, and an array of
hashed table names exposes no more than the hashed database / name next to
it, so they are now collected as arrayMap(x -> hex(SHA256(concat(x, salt))))
and the guard in gov_leak_test moved from "never mentioned" to "must be
hashed" — with target_database / target_table and view added to the
must-hash list. loading_dependent_* stay forbidden; one direction of every
edge is enough.

Every file was run against 22.8.21.38 (roots) and 26.7.5.10 (roots and
rungs) before committing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A new page, schema_graph.html, written next to dashboard.html: the graph of
tables, materialized views, dictionaries and Distributed tables with the
edges data flows along, nodes coloured by engine, a sidebar with keys,
columns, neighbours and the CREATE statement, search over table and column
names, drag, zoom and a database filter. Adapted from ClickHouse
programs/server/schema.html (Apache-2.0, attributed in the file header and
the page); every live query is replaced by JSON embedded at generation time
from the same five result sets — system.tables, columns, dictionaries,
view_refreshes — with system databases excluded. The heat map from
system.query_views_log is deferred to a follow-up.

Why a second file. The graph needs every column of every user table
embedded: ~750 KiB on a near-empty server, tens of MB on a service with
thousands of tables, which would slow dashboard.html for readers who never
open the graph. Tested the alternatives in headless Chrome from file://:
fetch() and XHR of a sibling file are blocked (opaque origin), an <iframe>
renders it, and assigning iframe.src on click is a navigation the browser
allows and does not request the file until then. So the dashboard gains a
Schema Graph tab whose frame gets its src on the first click. The parent
cannot script the frame across the file:// boundary and does not need to;
the two pages share the theme through localStorage, which file:// documents
in one folder do share (verified, and the toggle propagates live).

The Click UI token layer moves out of htmlTemplate into themeTokensCSS and
both pages are assembled from it, so they cannot drift apart in theme.
schema.html's own palette is remapped onto those tokens; only the engine
colours keep their values, in both dark scopes. The edge colour is --edge,
because --link is the dashboard's hyperlink token.

Payload: five safeQuery calls behind hasColumn("tables","target_table") and
hasTable("view_refreshes"), plain system.* in every mode (tables and columns
are shared across replicas; dictionaries would duplicate under
clusterAllReplicas). create_table_query, engine_full and dictionary source
pass through collection.RedactSQLText before embedding, so the sidebar can
never show more than the JSONL does. json.Marshal HTML-escapes <, > and &,
so a "</script>" inside a customer's DDL cannot terminate the inline script
— a test embeds one to prove it.

Verified against 26.7.5.10 with a schema holding every node kind (MergeTree
source, MV with TO target, MV with implicit .inner_id target, plain view,
dictionary, Distributed, refreshable MV): the tab appears, the frame has no
src until clicked, 22 nodes render with the right kinds, the sidebar opens,
the theme follows the dashboard, the standalone page shows a back link,
search dims non-matches; no non-file requests, no page errors. make
dashboard-preview now also writes bin/schema_graph_preview.html from a
fixture pipeline, reachable from the preview dashboard's Schema tab.

Docs: README (section 18, a Schema graph subsection, the credentials note,
version-table rows for view_refreshes and target_*), bundle-layout (tree,
§8 key, new §8b, the tables/columns rows, gov notes), file-guide (a
system.view_refreshes entry, system.tables and dashboard.html addenda) and
HC-7.6 for a refreshable MV that stopped refreshing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical credential-redaction issues and multiple graph correctness and accessibility issues remain unresolved.

Review effort: Lite
Findings: 2 High severity · 4 Medium severity

Open (6)
What changed in this PR

Adds an offline schema-dependency graph, expanded schema collectors, dashboard integration, and credential redaction for collected DDL.

Changes:

  • Adds on-demand graph rendering, previews, shared theming, and metadata collection.
  • Adds SQL/JSONL credential redaction and governance hashing updates.
  • Updates documentation, queries, tests, and dashboard generation.
File Changes Final review status
skills/​clickhouse-diagnostic/​references/​health-checks.md Documents refreshable-MV health checks. —
skills/​clickhouse-diagnostic/​references/​file-guide.md Documents schema and refresh metadata. —
skills/​clickhouse-diagnostic/​references/​bundle-layout.md Documents graph payloads and files. —
README.md Documents collectors, graph, and redaction. —
queries.onprem/​system.tables.sql Extends table metadata collection. —
queries.onprem/​system.columns.sql Adds key metadata collection. —
queries.onprem/​26.6.1.0/​system.tables.sql Adds materialized-view target fields. —
queries.onprem/​25.4.1.0/​system.tables.sql Updates versioned table collection. —
queries.onprem/​24.2.1.0/​system.tables.sql Updates versioned table collection. —
queries.onprem/​23.12.1.0/​system.view_refreshes.sql Adds refreshable-view collection. —
queries.gov/​system.tables.sql Adds hashed dependency metadata. —
queries.gov/​system.columns.sql Adds hashed key metadata. —
queries.gov/​26.6.1.0/​system.tables.sql Adds hashed MV target fields. —
queries.gov/​24.2.1.0/​system.tables.sql Updates versioned governance collection. —
queries.gov/​23.12.1.0/​system.view_refreshes.sql Adds governance refresh collection. —
queries.cloud/​system.tables.sql Extends cloud table metadata collection. —
queries.cloud/​system.columns.sql Adds cloud key metadata collection. —
queries.cloud/​26.6.1.0/​system.tables.sql Adds cloud MV target fields. —
queries.cloud/​25.4.1.0/​system.tables.sql Updates versioned cloud collection. —
queries.cloud/​24.2.1.0/​system.tables.sql Updates versioned cloud collection. —
queries.cloud/​23.12.1.0/​system.view_refreshes.sql Adds cloud refresh collection. —
Makefile Extends dashboard preview generation. —
internal/​query/​versioned_test.go Tests versioned collector coverage. —
internal/​query/​redact_hook_test.go Tests executor redaction wiring. —
internal/​query/​gov_leak_test.go Updates governance hashing checks. —
internal/​query/​executor.go Applies redaction before persistence. —
internal/​dashboard/​schema_graph.go Implements graph collection and rendering. Moderate (4 votes): failed dictionaries are omitted; Moderate (1 vote): empty schemas are treated as failures; Moderate (1 vote): exceptions lack the collector’s length bound; Moderate (2 votes): cloud existence checks precede local-only queries; Moderate (2 votes): graph nodes lack keyboard interaction; Moderate (2 votes): dictionary nodes and source edges are not created.
internal/​dashboard/​schema_graph_test.go Tests graph generation and embedding. —
internal/​dashboard/​keeper_preview_test.go Links graph data into previews. —
internal/​dashboard/​generator.go Generates and integrates the graph page. Critical (1 vote): raw dictionary sources can still leak credentials in the dashboard payload.
internal/​dashboard/​files_panel_test.go Verifies graph file exclusion. —
internal/​collection/​sqlredact.go Implements SQL and JSONL redaction. Moderate (1 vote): unchanged function calls are unnecessarily reconstructed, breaking byte preservation.
internal/​collection/​sqlredact_test.go Tests redaction behavior. —
internal/​collection/​heuristics.go Generalizes redaction heuristics. Moderate (1 vote): [HIDDEN] is not idempotently recognized; Critical (1 vote): escaped SQL string literals can leave password content unredacted.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread internal/collection/heuristics.go Outdated
Comment thread internal/dashboard/generator.go
Comment thread internal/dashboard/schema_graph.go Outdated
Comment thread internal/dashboard/schema_graph.go Outdated
Comment thread internal/dashboard/schema_graph.go Outdated
Comment thread internal/dashboard/schema_graph.go
…s, cloud refreshes, config dictionaries, keyboard

Six findings on #34, each verified before changing anything.

Keyword redaction stopped at an escaped quote. The rule's value pattern is
[^'"]+, so password = 'pa\'ss' matched 'pa\' and produced '[HIDDEN]'ss' —
the tail of the password stayed on disk. The value's END is now found by
scanning the literal with SQL escaping (\' and '' inside '…', \" inside
"…") instead of by the regex; an unterminated literal falls back to the
regex span. The config sanitizer's REMOVED path is unchanged and pinned.

The dashboard's own Dictionaries panel embedded system.dictionaries.source
and last_exception raw. Same class of leak as the JSONL one commit 1 fixed,
missed because it lives in collect(), not in the graph. Both fields now go
through the same redactor.

The dictionaries query kept schema.html's WHERE status IN (LOADED,
NOT_LOADED, LOADING). In a diagnostic bundle a FAILED dictionary is the
node to look at first, so every status is read and the sidebar shows
"Dictionary status" and "Dictionary error". The local test server proved
the point unprompted: its dict_users had gone FAILED (auth) between runs,
and the old filter would have hidden exactly that. Note the finding's
claim that filtered dictionaries "disappear entirely" was not accurate —
nodes come from system.tables, only the source text was missing — but the
fix is right regardless.

view_refreshes read the local table while hasTable probed every replica.
Now read through g.sysTable (clusterAllReplicas in cloud) and folded to one
row per view: argMax(status, status != 'RunningOnAnotherReplica') picks the
replica that actually runs the refresh, max() of the timestamps, argMax of
a non-empty exception. Verified on both the plain and the cluster form.

Dictionaries declared in server config (XML) have database = '' and no
system.tables row, so they had no node — true of the live page too. The
page now adds a node for any dictionary absent from tables, under a
"(config)" section, with its source in the sidebar; it has no dependency
arrays, so it draws without edges. DDL dictionaries were never affected:
they are in system.tables and rendered before.

Nodes were clickable <div>s with no keyboard path. Each is now
role=button, tabIndex=0, labelled, and Enter/Space select it; sidebar
cross-references are real links; a focus ring is drawn for keyboard users.
Driven in headless Chrome: Tab to a node, Enter opens the sidebar, Escape
closes it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread internal/collection/sqlredact.go Outdated
Comment thread internal/collection/sqlredact.go
Comment thread internal/dashboard/schema_graph.go Outdated
CamiloSierraH and others added 2 commits September 28, 2026 09:42
What the page shows and how to read it — node colours by engine, edges
along the data flow (MV between its source and its target, dictionary to
its source, Distributed to its shard table), what the sidebar holds, search
and keyboard use — with a real capture from a 26.7 run. The existing
subsection under Dashboard stays as the design record (why a second file,
theme sharing, credential redaction); the two cross-reference each other.

The capture is downscaled to 1600 px (456 KB): the README renders narrower
than that anyway, and it is the repo's first tracked image.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…action, dictionaries-only graphs

Three findings on the second pass, all consistency gaps rather than new
classes of problem.

The S3 rule masked the secret (argument 3) and left the key id (argument 2)
to the AWS-shape heuristic. AKIA…/ASIA… ids were caught; an S3-compatible
service — MinIO, GCS HMAC "GOOG1…", R2, OSS "LTAI…" — issues ids of any
shape, and the id is the credential's identifier, not a public value. Both
positions are now masked structurally, session-token and format handling
unchanged; a test uses a GOOG1… id.

system.dictionaries.last_exception was redacted for the dashboard panel and
the graph in the previous round but not for the persisted JSONL — the very
field the graph fixture shows quoting a password. Added to
SensitiveCollectorFields, with a test that the JSONL path scrubs it.

collectSchemaGraph returned before reading system.dictionaries when
system.tables had no user rows, so a server whose only user objects are
config-declared (XML) dictionaries — which have no system.tables row and
whose nodes the page builds from the dictionaries result — got no page at
all. The decision now comes after both reads: tables alone, dictionaries
alone, or both draw; nothing draws only when both are empty.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@ugosan

ugosan commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

This is amazing @CamiloSierraH ! I like how all the table types are color-coded, especially.

May I suggest some things:

  1. To make the RMV's parameters like REFRESH [EVERY|AFTER interval [OFFSET interval]] and [RANDOMIZE FOR interval] explicit in the diagram, it can be shown on the gray ribbon below the node (the one that shows the number of rows on tables), in MercadoLibre there is a db with 34 RMV's all firing every 2 minutes, a visualization like this would help to quickly see they are all firing at the same time, this is where I am going to immediately test this dashboard.

  2. To make LIFETIME({MIN min_val MAX max_val | max_val}) also explicit for the Dictionaries.

  3. the CREATE STATEMENT in the flyout to be pretty printed and syntax highlighted :)

Would indices and primary keys also be displayed/differentiated like the red color on the column names that denote ORDER BY ?

Amazing work!

… roles, skip indices, highlighted CREATE

Four suggestions from review on #34, from someone about to point this at a
database with 34 refreshable views all on the same two-minute schedule.

RMV schedule. system.view_refreshes has no interval column; the schedule
lives only in the DDL. The page parses REFRESH EVERY|AFTER <interval>
[OFFSET …] [RANDOMIZE FOR …] from create_table_query — the server's
single-line form and formatQuery's one-clause-per-line form alike — and puts
it on the grey ribbon under the node. A schedule with no RANDOMIZE FOR is
underlined and titled: thirty views firing at the same instant is the thing
to see at a glance. The sidebar gets schedule, offset and jitter rows.

Dictionary LIFETIME. system.dictionaries.lifetime_min/max, with the DDL's
LIFETIME(MIN a MAX b) / LIFETIME(n) as the fallback: an unloaded dictionary
reports 0/0 (the FAILED one on the test server did), and its declared
lifetime is still worth reading. Ribbon and sidebar.

Key roles. The four system.columns flags are independent — a column can be
in the sorting key AND the partition key — so all four travel instead of one
OR, and the node shows them apart: primary key bold red, sorting-only red,
partition key dotted amber underline, sampling key dashed teal, a legend for
each, and PRIMARY KEY / ORDER BY / PARTITION BY / SAMPLE BY tags per column
in the sidebar. Skip indices are new: system.data_skipping_indices becomes a
collector in all three modes (definitions only; gov hashes db/table/name and
drops expr) and a payload query; the ribbon shows the count, the sidebar
lists name, type(expr) and granularity, and search matches index names.

CREATE statement. Pretty-printed server-side with formatQueryOrNull when the
version has it (probed by name in system.functions; 1 of 201 system DDLs on
26.7 formats to NULL, which is why the plain formatQuery is not used — one
such row would fail the whole tables query) and syntax-highlighted on the
page by a small tokenizer that emits spans through textContent: keywords,
types, strings, numbers, comments, and '[HIDDEN]' in its own colour. Below
the function's version the single-line original is shown as before.

Verified against 26.7.5.10: the ribbon reads EVERY 1 HOUR (flagged, no
jitter) on the refreshable view, LIFETIME 60s on the unloaded dictionary via
the DDL fallback, "2 skip indices" on a table whose ts column carries both
the sorting and partition classes; the sidebar lists both indices, tags the
columns, and renders a 13-line highlighted CREATE. The collector runs on
22.8.21.38 and 26.7.5.10 in every mode.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@CamiloSierraH

CamiloSierraH commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator Author

Thanks @ugosan — all four are in as of 6c8e19d. What each became:

1. RMV schedule on the ribbon. system.view_refreshes has no interval column, so it is parsed from the DDL: REFRESH EVERY|AFTER <interval> [OFFSET …] [RANDOMIZE FOR …] lands on the grey ribbon under the node. One detail aimed at your 34-view case: a schedule with no RANDOMIZE FOR is underlined (and titled), so views that all fire at the same instant read as a row of underlines rather than something you have to open each node to find. The sidebar gets Refresh schedule / offset / jitter rows next to the live status and next-refresh time.

2. Dictionary LIFETIME. From lifetime_min/max, with the DDL's LIFETIME(MIN a MAX b) as the fallback — an unloaded dictionary reports 0/0, and its declared lifetime is still worth reading. Ribbon shows LIFETIME 60s or LIFETIME 60–120s.

3. CREATE statement. Pretty-printed by the server with formatQueryOrNull where the version has it (one clause per line; formatQuery itself is 23.11, the OrNull form is what makes it safe over every DDL), and syntax-highlighted on the page — keywords, types, strings, numbers, and [HIDDEN] in its own colour. Still offline, still no dependency: a ~40-line tokenizer emitting spans.

Your question — keys and indices. The four system.columns flags are independent (a column can be in the sorting key and the partition key), so the graph now shows them apart: primary key bold red, sorting-only red, partition key dotted amber underline, sampling key dashed teal, each in the legend; the sidebar tags every column PRIMARY KEY / ORDER BY / PARTITION BY / SAMPLE BY. Skip indices were not collected anywhere before — system.data_skipping_indices is now a collector in all three modes and on the graph: the ribbon shows 2 skip indices, the sidebar lists name, type(expr) and granularity, and search matches index names.

Verified against a 26.7 server with a refreshable view, an unloaded dictionary and a table carrying two skip indices plus a partition key inside its sorting key. If the run against that 34-view database shows something the ribbon should say differently, that is the feedback I would most like.

@ugosan

ugosan commented Sep 28, 2026

Copy link
Copy Markdown
Contributor
image Excellent, since we have the FROM in the DDL, can we also render it with an arrow (with a dashed line?) from the table?

CamiloSierraH and others added 2 commits September 28, 2026 18:09
…ema graph branch

Three additive conflicts — the dashboard-preview target, the README preview
paragraph and bundle-layout §8 — resolved as the union of both sides: the
preview target now writes all three pages, and §8 keeps main's note on the
collapsed Alert Summary plus the schema_graph key and §8b.
generator.go auto-merged; the suite, gofmt and vet pass on the result.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Review follow-up on #34. A refreshable MV, a plain View and a Distributed
table have empty dependency arrays in system.tables — the RMV does not fire
on insert so it is not registered as its source's dependent, the View is
never registered, and the Distributed table's local table is only an engine
argument — so all three floated free of the table they read.

The DDL says what they read. A second pass, after every metadata edge is in,
parses the FROM / JOIN clauses of a view's SELECT (string literals removed,
ARRAY JOIN and table functions skipped, an unqualified name resolved to the
view's own database) and the ('cluster', 'db', 'table') arguments of a
Distributed engine, and adds an edge where none exists. addEdge keeps the
first of a pair, so a metadata edge is never demoted.

Inferred edges are drawn dashed, carry a title saying where they came from,
have their own legend entry, and the sidebar's Reads-from / Writes-to lists
mark them "inferred from the DDL" — the text is evidence, not metadata, and
the picture should say which is which.

Verified against 26.7.5.10: exactly three dashed edges appear — the table to
its refreshable MV, to its view, and to its Distributed table — and the seven
metadata edges stay solid.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@CamiloSierraH

Copy link
Copy Markdown
Collaborator Author

Done in 1042621 — and yes, dashed.

You'd found a real gap: for a refreshable MV, a plain View and a Distributed table, system.tables records no dependency at all (the RMV doesn't fire on insert so it is never registered as the source's dependent; the Distributed table's local table is only an engine argument), so those nodes floated free of the table they read. The page now runs a second pass after the metadata edges: it parses FROM / JOIN from the view's SELECT (string literals removed, ARRAY JOIN and table functions skipped, unqualified names resolved to the view's own database) and the ('cluster', 'db', 'table') arguments of a Distributed engine, and adds an edge where none exists.

Those edges are dashed, titled with where they came from, have their own legend entry, and the sidebar's Reads from / Writes to lists mark them "inferred from the DDL" — the text is evidence rather than metadata, and the picture should say which is which. A metadata edge is never demoted: the first edge for a pair wins, and metadata goes first.

On a 26.7 server with your three cases side by side: exactly three dashed edges appear (table → RMV, table → view, table → Distributed) and the seven metadata edges stay solid.

@ugosan

ugosan commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Just look at this beauty!
image

@ugosan ugosan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Amazing work, the graph displays dependencies clearly

@CamiloSierraH
CamiloSierraH merged commit 86037da into main Sep 28, 2026
2 checks passed
CamiloSierraH added a commit that referenced this pull request Sep 29, 2026
…ry_views_log branch

Two additive conflicts: the README collector table (this branch's
query_views_log row beside main's view_refreshes and data_skipping_indices
rows) and HC-7, where main's 7.6 (view_refreshes) landed first — this
branch's rows are renumbered 7.7 failures, 7.8 stopped writing, 7.9 chain
divergence, and inspect_bundle's finding tags follow. The merge note in the
PR said whichever landed second would renumber; this is that.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants