Repository navigation
render: a Deploy to Render button for a metrics pipeline - #335
Conversation
… them once The template's schema. metrics_1m holds five values per minute per series, the five every coarser grain merges from exactly. series turns a dimensions_key back into dimensions and is kept by statement triggers in the writer's transaction. migrate.sh is the Bluesky demo's at 6ec4966, under its own lock name. Applied twice against Postgres 18: the second run skips both. An upsert of an older minute moves first_bucket back and leaves last_bucket. A type that is neither count nor gauge is refused by the table. Part of #331.
rollups.yml declares the ladder over metrics_1m. metrics keeps every dimension and all five values. metrics_total drops dimensions_key for a name with a series per user, and has no last. 0003_rollups.sql is sqlflow rollup ddl's output from v2026.09.18, byte for byte the migration verified locally against a build of #334. Applied to Postgres 18: 0.25 and 0.5 reach the day as 0.75 with last 0.5, and three minutes of two series total 4.75. Part of #331.
metric filters by containment: the series whose dimensions include every pair sent. It joins series in DuckDB, after name and the range have pushed down to Postgres, so a filtered request scans what an unfiltered one does. A filter that is not JSON matches nothing rather than everything. metric_total reads one row per bucket for a name with many series. The grains are written by hand because sqlflow rollup generates a dataset of at most one dimension, folded. The next commit's test queries every one. Part of #331.
… end to end pipeline.yml reads one metric from a request's top-level keys or several from metrics, casts every field in SQL so one sender's bad value cannot fail a batch, builds the series key, and merges each closed minute into five values. A timestamp more than a minute ahead is dropped: the window closes against the newest event time, and one metric dated 2030 would close every minute until then. The entrypoint exits 2 on a blank secret, an auth mode that is neither hmac nor none, and a name prefix that could leave its quotes. test/e2e.sh builds the image from v2026.09.18, posts signed metrics and reads them back at all six grains of both datasets, 59 assertions. It asks each grain for a range that holds one of its buckets: serve snaps a range to bucket boundaries, and a pinned 1d grain over the default hour is an empty range, which the test asserts on purpose. Seven malformed bodies answer 200 and store nothing. Unsigned mode with a name prefix, the collector's configuration, accepts install.deployed and drops a name outside the prefix. Part of #331.
… was hiding render.yaml declares a Postgres and two web services built from render/. The HMAC secret is sync: false, so the deploy asks for it. autoDeployTrigger is off, so a copy in someone's workspace does not redeploy when this repository's main moves. CI validates the configs, the generated migration and render.yaml, holds that an invalid plan is refused with status 1, and runs the end-to-end test. The repository ignores bin/, which also matched render/bin/. The four commits before this one therefore held no migrate.sh, entrypoint.sh, serve.sh or send.sh: the end-to-end test passed from the files on disk, and a clone could not have built the image, because the Dockerfile copies bin. .gitignore now excepts /render/bin/, and this commit adds the scripts, executable. Part of #331.
render/README.md is what a deployer reads: the one prompt, the first signed request, the rules a metric must meet, and the three datasets. It says plainly that the pipeline answers 200 to anything signed and stores only what meets the rules, and that a pinned grain needs a range at least as wide as its bucket. Every command in it was run against the local compose: both forms of the signed request answered received, and the metric read back with value_sum 2. Part of #331.
The template is built and its CI is green in #335. Two things in the plan were wrong, and both would have cost the next reader. The repository ignores bin/, which also matches render/bin/. Following the plan as written committed none of the scripts, and the end-to-end test still passed, from the files on disk. The plan now adds the exception to .gitignore in Task 1, checks it with git check-ignore, and ends with a run from a clean git archive, which is the check that catches it. The restart check grepped the ingest log for a skipped migration. psql prints two spaces where the pattern had one, and the line proves nothing anyway: both services migrate, so whichever lost the first race had already printed it. The check now reads schema_migrations and asks that the service came back. The files carry v2026.09.18 rather than a placeholder.
The template is built and its CI is green in #335. Two things in the plan were wrong, and both would have cost the next reader. The repository ignores bin/, which also matches render/bin/. Following the plan as written committed none of the scripts, and the end-to-end test still passed, from the files on disk. The plan now adds the exception to .gitignore in Task 1, checks it with git check-ignore, and ends with a run from a clean git archive, which is the check that catches it. The restart check grepped the ingest log for a skipped migration. psql prints two spaces where the pattern had one, and the line proves nothing anyway: both services migrate, so whichever lost the first race had already printed it. The check now reads schema_migrations and asks that the service came back. The files carry v2026.09.18 rather than a placeholder.
…wer for minutes Found on the first deploy to Render. The client id was generateValue. Render minted aA+SZt88...+mIlHiM=, and pasted into a URL as it is, it answered 401: '+' reads as a space. It worked only through --data-urlencode, and the deployer had to find it in the dashboard first. It is now sync: false, asked for beside the HMAC secret, so the deployer has it when the deploy finishes. serve.sh refuses a blank one and one that holds anything a URL would encode, and names `openssl rand -hex 16`. There is no default: in this template the id is the only thing between a reader and the deployer's metrics, and the README now says so. The coarse grains held an answer for 2, 5 and 10 minutes, the Bluesky demo's times. Asked at 1d before its first event had landed, the live collector answered empty with "cache": "hit" for ten minutes while the row sat in the table. A new deploy sends one metric and asks for it, so every grain now holds for the same 30 seconds. The end-to-end test gains four assertions, 63 in all: the API exits 2 on a blank id and on one with '+', '/' and '='; the id works pasted into a URL unencoded; a wrong id answers 401. Part of #331.
…invariant holds it A pipeline process holds its own window and publishes a closed minute by replacing the row for its key. Keyed on the series alone, a second process holding part of the same minute replaced the first. Run as two real instances and sent 7 and 5 of one minute, the template stored 7, and every rollup re-merged the 7. Two instances is what a scaled service runs, and what every Render deploy runs for a moment, the old beside the new. The pipeline now writes metrics_1m_writers, keyed on the series and on a writer id that entrypoint.sh makes new on every start and never reads from the environment: a value set once would be shared by every instance. A trigger merges every writer's row into metrics_1m in the same transaction. Sums add, min and max nest, and value_last is the reading with the latest event time across processes. series, the rollups, serve.yml and the API are untouched: they read metrics_1m and never see a writer. The first version of the trigger locked per minute, and four concurrent writers deadlocked within one round. Postgres's log named a rollup advisory lock on one side and a series row lock on the other: an upsert fires the trigger chain twice, once for the rows it inserted and once for the rows it updated, so a transaction can hold a rollup lock from its first pass and want a series row in its second. The trigger now takes one lock before it touches anything another writer wants, and writers run the chain one at a time. A publish is milliseconds, once per poll. pipeline.writers.merge_exactly is registered in the invariant matrix, proven for pipeline.stateless and empty for pipeline.stateful, which is accurate: nothing in the engine makes it true, a destination's schema does. internal/rendertemplate proves it against the template's own migrations. Two controls must fail and do, each storing 5|1|5|5|5 where 12|2|5|7|5 is right: the lock removed, and two processes sharing a writer id. The merge refuses REPEATABLE READ. Four writers publish and republish 36 overlapping keys through the Postgres sink for forty rounds, and metrics_1m and all ten rollup tables equal what the writers remember, worked out in Go and never from the database. Deleting the writers' rows afterwards changes nothing, which the README tells a deployer they may do. On failure the test prints the server's deadlock report, which is how the second defect was diagnosed. make test runs two instances and sends each a share of one minute: 76 assertions, the split minute and its gauge's value_last at all six grains. Part of #331.
Evidence: a clean deploy from this branch, verified from outsideDeployed on 2026-09-18 from Auth, both services
The empty-key row is the one that matters most: it is what a request looks like if a pipeline were ever running with a blank secret. What was sent, at 17:48:36Z
What came backReadable 78 seconds after the send, inside the README's "one to two minutes". At each of
Also: the quoted dimension value read back as 26 of 26 checks held. The earlier deploy from this branch, at 9cd81e9, and what it changedThat one was deployed with the HMAC prompt left blank, on purpose.
Settled by these deploys
Not covered by a live deployTwo pipeline instances at once. This deploy runs one. d50f024's writer merge is proven in |
…cannot open a signed one
Two defects in how the auth mode reaches the pipeline.
render.yaml set SQLFLOW_WEBHOOK_AUTH to hmac. Render's Blueprint
reference says it "preserves existing environment variables, even if you
omit them from the Blueprint file", and that a resource "retains any
existing environment variable values that aren't overwritten by the
Blueprint". So a value in the file is rewritten on every sync. A pipeline
set to none in the dashboard, which is what the telemetry collector is,
would go back to hmac on the next sync and start refusing its senders
without anyone having changed it. The variable is no longer in the file.
Unset, the entrypoint defaults to hmac, and a value added in the
dashboard survives a sync. No third prompt.
Checking that default found the second. entrypoint.sh and pipeline.yml
each decided the mode, and disagreed about an empty string. The script
read it as hmac and was satisfied by the secret. The template's
`default('hmac') == 'hmac'` read it as not-hmac and rendered no signature
block. Run with SQLFLOW_WEBHOOK_AUTH set to empty and a secret set, the
pipeline answered an unsigned request 200. The entrypoint now settles
the mode and exports it, so the template reads only a checked value, and
the template signs unless the mode is exactly none, so it fails closed
even without the script.
The end-to-end test starts a pipeline with the mode empty and with it
hmac and holds that both refuse an unsigned request: 78 assertions.
Compose passes the variable through rather than defaulting it, so the
main flow runs with it absent, as Render leaves it.
Part of #331.
… one word The template tells the sqlflow maintainers that it was deployed and that it received its first metric. Two events per install, ever, and nothing else. render/README.md, "What this sends", prints both payloads in full, and SQLFLOW_TELEMETRY is a third deploy prompt so it is seen before anything is sent. Blank is on; off, false, 0 and no send nothing, and the log says so. The sqlflow binary sends nothing: bin/telemetry.sh, 91 lines of shell, is the only code that does. An event is an ordinary metric posted to https://telemetry.turbolytics.io, which is this same template with signatures off and SQLFLOW_METRIC_NAME_PREFIX=install. It carries install_id, a uuid the database makes for itself from nothing about the deployer; source: render, a constant, so other vendors' templates can report to one collector and be told apart; template; and the image tag. install.first_request asks Postgres whether metrics_1m has a row whose name is not install.*, and reads nothing about it. The hostname is one we own and not the collector's onrender.com URL: a deployed copy never updates itself, so whatever ships is what it calls for as long as it runs. A send is claimed in Postgres first, a lease that expires after two minutes, so two pipeline instances starting together report one install and a process that dies mid-send does not lose the event. It is recorded as sent only on a 2xx: the collector's custom domain answered 404 while it propagated, and an install recorded as sent on a 404 is never counted. A send is bounded at five seconds, runs in the background, and cannot stop the pipeline. Verified against the live collector by hand first: ten POSTs through the hostname answered 200, a 5000-byte body 413, and an install.deployed with source render read back through its API, filtered by source. The end-to-end test runs a local collector, the same template, and no service in compose can reach the real one: after the run the real collector held only the two rows sent by hand. 87 assertions. Nine are new: one series per event from two instances starting together, the dimensions exactly, each sum 1, still 1 after both restart, off claims and sends nothing, and a dead collector leaves the event unsent and the pipeline answering. The test's first attempt failed and the client was right. The collector logged "dropped late rows": under the test's five-second idle bound a minute closes while it is still current, and the second event landed in the minute the first had closed. A deploy's bound is 60, where a closed minute has always ended. The test now starts its burst in a fresh minute. The same run showed install.deployed itself, in the shared test database, counting as the first metric, which is why install.* names are excluded. Part of #331.
A webhook pipeline deployed as a Render web service had no health check. The webhook source answered only POST /events, and the pipeline's /healthz is on the metrics listener, :8000. A platform routes one port to a service and checks health on that port, so Render marked the service live because its port was open, could not tell a closing instance from a serving one, and an uptime monitor pointed at https://telemetry.turbolytics.io/healthz got 404. Found on the first deploy of the Deploy to Render template (#331, #335). The source now answers GET and HEAD /healthz with {"status":"ok"} while it admits deliveries, and 503 once it is closing, which is what it says to a delivery then. HEAD because that is what probing services send. Any other method is 405 with Allow: the route owns its path for every method, because beside the catch-all a GET-only pattern handed POST /healthz to the delivery mux, which answered 404 for a path that exists. It reads no body and checks no signature, since a health check has neither, and admits nothing to the pipeline. It does not wait on the queue: a full queue is backpressure, and a platform that reads busy as dead restarts an instance while it holds a sender's event. It sits outside the metrics middleware: webhook_requests_total is how an operator counts deliveries, and a check every few seconds would bury them under 200s that delivered nothing. Four tests, each watched failing first with 404 or a count of 6: no signature needed while an unsigned delivery is still refused; 503 after Close; 200 while a delivery waits on a full queue; five checks and one delivery count as one request. Run from this branch with HMAC on: GET and HEAD answered 200 unsigned, POST /healthz 405 with Allow: GET, HEAD, an unsigned delivery 400, a signed one 200. make test-go passes. Part of #331.
…eline The ingest service had no healthCheckPath, because the webhook source answered only POST /events. Render called it live when its port opened, could not tell an instance that was shutting down from one that was serving, and could not hold a deploy until the new instance was ready. https://telemetry.turbolytics.io/healthz answered 404. v2026.09.18.1 adds GET and HEAD /healthz to the webhook source (#336): 200 while it admits deliveries, 503 once it is closing, no signature, not counted as a request. render.yaml sets healthCheckPath: /healthz on the ingest service, and the Dockerfile, the Makefile and compose pin the new tag. sqlflow rollup ddl from the new image generates the committed migration byte for byte: make validate's drift check passes without regenerating. The end-to-end test, with signatures on, holds that GET and HEAD /healthz answer 200 unsigned and POST is refused: 88 assertions. The install event's sqlflow_version follows the pin, and the test reads it from the image tag rather than a literal. Part of #331.
The end-to-end test failed once from a clean export, with install.first_request present and install.deployed missing, after passing in the working tree. The auth-mode probes start real pipelines, which run telemetry.sh, and kill them as soon as they are probed. One killed between claiming install.deployed and giving the claim back left the two-minute lease held. The pipelines started next could not claim the event until it expired, and the test waits 60 seconds. That is the lease doing its job: in a deploy, a process killed mid-send delays the event two minutes and does not lose it. In a test it is a flake. The probes now run with SQLFLOW_TELEMETRY=off, since they test auth and not telemetry, and the test clears any claim before it starts the pipelines it measures. Three runs in a row pass, 88 assertions each. Part of #331.
…emplate # Conflicts: # CHANGELOG.md
Evidence: a real install reporting to the real collector, on v2026.09.18.1Two separate deploys from this branch on 2026-09-18, both at the head that pins The collector:
|
| Check | Result |
|---|---|
GET /healthz on the ingest service and through the hostname |
200 {"status":"ok"}, x-render-origin-server: Render |
HEAD /healthz |
200 |
POST /healthz |
405 |
Render's own check, healthCheckPath: /healthz, new in this head |
the service went live on it |
Unsigned POST /events |
200 |
| A 5000-byte body | 413 |
| After pressing Manual sync on the Blueprint | unsigned still 200, the 5000-byte body still 413 |
The last row is what f12d143 was for. render.yaml no longer names SQLFLOW_WEBHOOK_AUTH, Render's reference says it "preserves existing environment variables, even if you omit them from the Blueprint file", and a sync left all three dashboard-set variables in effect.
The install: a second deploy, all three prompts, SQLFLOW_TELEMETRY left blank
Render created it as sqlflow-metrics-ingest-nypg and sqlflow-metrics-api-nypg: a second Blueprint in a workspace that already holds these names gets suffixed services, and does not adopt the first one's.
Its ingest log, in order:
starting webhook server {"addr": "[::]:10000"}
telemetry: sent install.deployed to https://telemetry.turbolytics.io: {"name":"install.deployed","type":"count","value":1,"dimensions":{"install_id":"304e6c70-53ae-41a6-ab67-09ea3b677d5d","source":"render","template":"render-metrics","sqlflow_version":"v2026.09.18.1"}}
telemetry: this is one of two events this install ever sends. SQLFLOW_TELEMETRY=off stops them. See render/README.md.
==> Your service is live
One signed metric was sent to it at 20:50:55Z. What the collector's API then held, read with the collector's client id:
| Event | install_id |
source |
sqlflow_version |
Bucket | value_sum |
|---|---|---|---|---|---|
install.deployed |
304e6c70-…-09ea3b677d5d |
render |
v2026.09.18.1 |
20:50 | 1 |
install.first_request |
the same | render |
v2026.09.18.1 |
20:52 | 1 |
The uuid the install logged is the uuid the collector stored. Each event arrived once. install.first_request followed the first metric by under two minutes: the install's minute closing, a 30-second poll, then the collector's minute closing.
On the install itself: unsigned POST /events 400, the signed metric 200 and read back from its API as hello, value_sum 1, dimensions {"from":"test-deploy"}. Its API answered 200 to its own client id pasted into the URL, 401 to none, and 401 to the collector's id.
What this does not cover
SQLFLOW_TELEMETRY=offon Render. It is proven in the end-to-end test and on the collector's own deploy, which logged nothing and sent nothing, but no install was deployed withoffand then watched for silence.- Two pipeline instances at once on Render.
- The collector holds two junk series,
install.checkandinstall.after-sync, from probing it by hand. An open endpoint accepts anyinstall.*name from anyone, which is the known cost of it.
Part of #331. Spec and plan are in #330. Draft: do not merge yet. The telemetry client lands on this branch first, so the button never reaches
mainwithout its disclosure.What
A Deploy to Render button.
render.yamlat the root points intorender/, which deploys:sqlflow-metrics-ingest: a webhook pipeline that aggregates a generic metric by minute,sqlflow-metrics-api:sqlflow serveover the minute table and a generated six-grain rollup ladder.Render prompts for the HMAC secret. A blank one fails the deploy on purpose: unsigned mode is
SQLFLOW_WEBHOOK_AUTH=none, chosen, never reached by omission.Pinned to
turbolytics/sql-flow:v2026.09.18, which carries the three engine changes this needs (#332, #333, #334).The metric
{name, type: count|gauge, value, dimensions, timestamp}, one per request or several undermetrics. A minute storesvalue_sum,value_count,value_min,value_maxandvalue_lastfor both types, the five values every coarser grain merges from exactly.typeonly says which to read.Three datasets:
series,metric(per series, filtered by containment:dimensions={"region":"us-east"}returns every series that includes that pair), andmetric_total(summed across every series of a name, for a name with a series per user).Evidence
make -C render test: builds the image, posts signed metrics, reads them back at all six grains of both datasets. 59 assertions. It passes from a cleangit archiveof this branch, not only from the working tree.SQLFLOW_METRIC_NAME_PREFIX=install., the collector's configuration, acceptsinstall.deployedand drops a name outside the prefix.make -C render validate: the three configs,sqlflow rollup checkon the generated migration, andrender.yamlagainst Render's published schema. CI also holds that an invalid plan is refused.metrics_1m, the day equaled sums worked out by hand,value_lastfollowed event time when a later reading was sent first, and a republished minute replaced itself at every grain.render/README.mdwas run against the local compose.Two things found on the way
bin/, which also matchedrender/bin/. The first four commits here held none of the scripts and passed only from the files on disk..gitignorenow excepts/render/bin/, and the fifth commit says so.sqlflow servesnaps a range to the grain's bucket boundaries. A pinned1dgrain over the default hour is an empty range, not an error. The test asks each grain for a range that holds one of its buckets and asserts the empty case on purpose, and the README tells a reader not to pin a grain narrower than its range.Not done here, on purpose
POST /events; the pipeline's/healthzis on:8000. That deploy becomes the install-telemetry collector.