Repository navigation
webhook: GET and HEAD /healthz on the source's own address - #336
Merged
Merged
Conversation
A webhook pipeline deployed as a Render web service had no health check. The webhook source answered only POST /events, and the pipeline's /healthz is on the metrics listener, :8000. A platform routes one port to a service and checks health on that port, so Render marked the service live because its port was open, could not tell a closing instance from a serving one, and an uptime monitor pointed at https://telemetry.turbolytics.io/healthz got 404. Found on the first deploy of the Deploy to Render template (#331, #335). The source now answers GET and HEAD /healthz with {"status":"ok"} while it admits deliveries, and 503 once it is closing, which is what it says to a delivery then. HEAD because that is what probing services send. Any other method is 405 with Allow: the route owns its path for every method, because beside the catch-all a GET-only pattern handed POST /healthz to the delivery mux, which answered 404 for a path that exists. It reads no body and checks no signature, since a health check has neither, and admits nothing to the pipeline. It does not wait on the queue: a full queue is backpressure, and a platform that reads busy as dead restarts an instance while it holds a sender's event. It sits outside the metrics middleware: webhook_requests_total is how an operator counts deliveries, and a check every few seconds would bury them under 200s that delivered nothing. Four tests, each watched failing first with 404 or a count of 6: no signature needed while an unsigned delivery is still refused; 503 after Close; 200 while a delivery waits on a full queue; five checks and one delivery count as one request. Run from this branch with HMAC on: GET and HEAD answered 200 unsigned, POST /healthz 405 with Allow: GET, HEAD, an unsigned delivery 400, a signed one 200. make test-go passes. Part of #331.
turbolytics
added a commit
that referenced
this pull request
Sep 18, 2026
…eline The ingest service had no healthCheckPath, because the webhook source answered only POST /events. Render called it live when its port opened, could not tell an instance that was shutting down from one that was serving, and could not hold a deploy until the new instance was ready. https://telemetry.turbolytics.io/healthz answered 404. v2026.09.18.1 adds GET and HEAD /healthz to the webhook source (#336): 200 while it admits deliveries, 503 once it is closing, no signature, not counted as a request. render.yaml sets healthCheckPath: /healthz on the ingest service, and the Dockerfile, the Makefile and compose pin the new tag. sqlflow rollup ddl from the new image generates the committed migration byte for byte: make validate's drift check passes without regenerating. The end-to-end test, with signatures on, holds that GET and HEAD /healthz answer 200 unsigned and POST is refused: 88 assertions. The install event's sqlflow_version follows the pin, and the test reads it from the image tag rather than a literal. Part of #331.
turbolytics
added a commit
that referenced
this pull request
Sep 18, 2026
* render: the minute table, the series table, and a runner that applies them once The template's schema. metrics_1m holds five values per minute per series, the five every coarser grain merges from exactly. series turns a dimensions_key back into dimensions and is kept by statement triggers in the writer's transaction. migrate.sh is the Bluesky demo's at 6ec4966, under its own lock name. Applied twice against Postgres 18: the second run skips both. An upsert of an older minute moves first_bucket back and leaves last_bucket. A type that is neither count nor gauge is refused by the table. Part of #331. * render: five coarser grains, generated, and a total across dimensions rollups.yml declares the ladder over metrics_1m. metrics keeps every dimension and all five values. metrics_total drops dimensions_key for a name with a series per user, and has no last. 0003_rollups.sql is sqlflow rollup ddl's output from v2026.09.18, byte for byte the migration verified locally against a build of #334. Applied to Postgres 18: 0.25 and 0.5 reach the day as 0.75 with last 0.5, and three minutes of two series total 4.75. Part of #331. * render: series, metric and metric_total over HTTP metric filters by containment: the series whose dimensions include every pair sent. It joins series in DuckDB, after name and the range have pushed down to Postgres, so a filtered request scans what an unfiltered one does. A filter that is not JSON matches nothing rather than everything. metric_total reads one row per bucket for a name with many series. The grains are written by hand because sqlflow rollup generates a dataset of at most one dimension, folded. The next commit's test queries every one. Part of #331. * render: a webhook pipeline that aggregates a metric by minute, tested end to end pipeline.yml reads one metric from a request's top-level keys or several from metrics, casts every field in SQL so one sender's bad value cannot fail a batch, builds the series key, and merges each closed minute into five values. A timestamp more than a minute ahead is dropped: the window closes against the newest event time, and one metric dated 2030 would close every minute until then. The entrypoint exits 2 on a blank secret, an auth mode that is neither hmac nor none, and a name prefix that could leave its quotes. test/e2e.sh builds the image from v2026.09.18, posts signed metrics and reads them back at all six grains of both datasets, 59 assertions. It asks each grain for a range that holds one of its buckets: serve snaps a range to bucket boundaries, and a pinned 1d grain over the default hour is an empty range, which the test asserts on purpose. Seven malformed bodies answer 200 and store nothing. Unsigned mode with a name prefix, the collector's configuration, accepts install.deployed and drops a name outside the prefix. Part of #331. * render: the Blueprint, its check, CI, and the scripts the ignore file was hiding render.yaml declares a Postgres and two web services built from render/. The HMAC secret is sync: false, so the deploy asks for it. autoDeployTrigger is off, so a copy in someone's workspace does not redeploy when this repository's main moves. CI validates the configs, the generated migration and render.yaml, holds that an invalid plan is refused with status 1, and runs the end-to-end test. The repository ignores bin/, which also matched render/bin/. The four commits before this one therefore held no migrate.sh, entrypoint.sh, serve.sh or send.sh: the end-to-end test passed from the files on disk, and a clone could not have built the image, because the Dockerfile copies bin. .gitignore now excepts /render/bin/, and this commit adds the scripts, executable. Part of #331. * docs: the Deploy to Render button, and the template's contract render/README.md is what a deployer reads: the one prompt, the first signed request, the rules a metric must meet, and the three datasets. It says plainly that the pipeline answers 200 to anything signed and stores only what meets the rules, and that a pinned grain needs a range at least as wide as its bucket. Every command in it was run against the local compose: both forms of the signed request answered received, and the metric read back with value_sum 2. Part of #331. * render: the deployer chooses the client id, and no grain holds an answer for minutes Found on the first deploy to Render. The client id was generateValue. Render minted aA+SZt88...+mIlHiM=, and pasted into a URL as it is, it answered 401: '+' reads as a space. It worked only through --data-urlencode, and the deployer had to find it in the dashboard first. It is now sync: false, asked for beside the HMAC secret, so the deployer has it when the deploy finishes. serve.sh refuses a blank one and one that holds anything a URL would encode, and names `openssl rand -hex 16`. There is no default: in this template the id is the only thing between a reader and the deployer's metrics, and the README now says so. The coarse grains held an answer for 2, 5 and 10 minutes, the Bluesky demo's times. Asked at 1d before its first event had landed, the live collector answered empty with "cache": "hit" for ten minutes while the row sat in the table. A new deploy sends one metric and asks for it, so every grain now holds for the same 30 seconds. The end-to-end test gains four assertions, 63 in all: the API exits 2 on a blank id and on one with '+', '/' and '='; the id works pasted into a URL unencoded; a wrong id answers 401. Part of #331. * render: two pipeline processes writing one minute are summed, and an invariant holds it A pipeline process holds its own window and publishes a closed minute by replacing the row for its key. Keyed on the series alone, a second process holding part of the same minute replaced the first. Run as two real instances and sent 7 and 5 of one minute, the template stored 7, and every rollup re-merged the 7. Two instances is what a scaled service runs, and what every Render deploy runs for a moment, the old beside the new. The pipeline now writes metrics_1m_writers, keyed on the series and on a writer id that entrypoint.sh makes new on every start and never reads from the environment: a value set once would be shared by every instance. A trigger merges every writer's row into metrics_1m in the same transaction. Sums add, min and max nest, and value_last is the reading with the latest event time across processes. series, the rollups, serve.yml and the API are untouched: they read metrics_1m and never see a writer. The first version of the trigger locked per minute, and four concurrent writers deadlocked within one round. Postgres's log named a rollup advisory lock on one side and a series row lock on the other: an upsert fires the trigger chain twice, once for the rows it inserted and once for the rows it updated, so a transaction can hold a rollup lock from its first pass and want a series row in its second. The trigger now takes one lock before it touches anything another writer wants, and writers run the chain one at a time. A publish is milliseconds, once per poll. pipeline.writers.merge_exactly is registered in the invariant matrix, proven for pipeline.stateless and empty for pipeline.stateful, which is accurate: nothing in the engine makes it true, a destination's schema does. internal/rendertemplate proves it against the template's own migrations. Two controls must fail and do, each storing 5|1|5|5|5 where 12|2|5|7|5 is right: the lock removed, and two processes sharing a writer id. The merge refuses REPEATABLE READ. Four writers publish and republish 36 overlapping keys through the Postgres sink for forty rounds, and metrics_1m and all ten rollup tables equal what the writers remember, worked out in Go and never from the database. Deleting the writers' rows afterwards changes nothing, which the README tells a deployer they may do. On failure the test prints the server's deadlock report, which is how the second defect was diagnosed. make test runs two instances and sends each a share of one minute: 76 assertions, the split minute and its gauge's value_last at all six grains. Part of #331. * render: a sync cannot close an open pipeline, and an empty auth mode cannot open a signed one Two defects in how the auth mode reaches the pipeline. render.yaml set SQLFLOW_WEBHOOK_AUTH to hmac. Render's Blueprint reference says it "preserves existing environment variables, even if you omit them from the Blueprint file", and that a resource "retains any existing environment variable values that aren't overwritten by the Blueprint". So a value in the file is rewritten on every sync. A pipeline set to none in the dashboard, which is what the telemetry collector is, would go back to hmac on the next sync and start refusing its senders without anyone having changed it. The variable is no longer in the file. Unset, the entrypoint defaults to hmac, and a value added in the dashboard survives a sync. No third prompt. Checking that default found the second. entrypoint.sh and pipeline.yml each decided the mode, and disagreed about an empty string. The script read it as hmac and was satisfied by the secret. The template's `default('hmac') == 'hmac'` read it as not-hmac and rendered no signature block. Run with SQLFLOW_WEBHOOK_AUTH set to empty and a secret set, the pipeline answered an unsigned request 200. The entrypoint now settles the mode and exports it, so the template reads only a checked value, and the template signs unless the mode is exactly none, so it fails closed even without the script. The end-to-end test starts a pipeline with the mode empty and with it hmac and holds that both refuse an unsigned request: 78 assertions. Compose passes the variable through rather than defaulting it, so the main flow runs with it absent, as Render leaves it. Part of #331. * render: two install events, once each, disclosed in full and off with one word The template tells the sqlflow maintainers that it was deployed and that it received its first metric. Two events per install, ever, and nothing else. render/README.md, "What this sends", prints both payloads in full, and SQLFLOW_TELEMETRY is a third deploy prompt so it is seen before anything is sent. Blank is on; off, false, 0 and no send nothing, and the log says so. The sqlflow binary sends nothing: bin/telemetry.sh, 91 lines of shell, is the only code that does. An event is an ordinary metric posted to https://telemetry.turbolytics.io, which is this same template with signatures off and SQLFLOW_METRIC_NAME_PREFIX=install. It carries install_id, a uuid the database makes for itself from nothing about the deployer; source: render, a constant, so other vendors' templates can report to one collector and be told apart; template; and the image tag. install.first_request asks Postgres whether metrics_1m has a row whose name is not install.*, and reads nothing about it. The hostname is one we own and not the collector's onrender.com URL: a deployed copy never updates itself, so whatever ships is what it calls for as long as it runs. A send is claimed in Postgres first, a lease that expires after two minutes, so two pipeline instances starting together report one install and a process that dies mid-send does not lose the event. It is recorded as sent only on a 2xx: the collector's custom domain answered 404 while it propagated, and an install recorded as sent on a 404 is never counted. A send is bounded at five seconds, runs in the background, and cannot stop the pipeline. Verified against the live collector by hand first: ten POSTs through the hostname answered 200, a 5000-byte body 413, and an install.deployed with source render read back through its API, filtered by source. The end-to-end test runs a local collector, the same template, and no service in compose can reach the real one: after the run the real collector held only the two rows sent by hand. 87 assertions. Nine are new: one series per event from two instances starting together, the dimensions exactly, each sum 1, still 1 after both restart, off claims and sends nothing, and a dead collector leaves the event unsent and the pipeline answering. The test's first attempt failed and the client was right. The collector logged "dropped late rows": under the test's five-second idle bound a minute closes while it is still current, and the second event landed in the minute the first had closed. A deploy's bound is 60, where a closed minute has always ended. The test now starts its burst in a fresh minute. The same run showed install.deployed itself, in the shared test database, counting as the first metric, which is why install.* names are excluded. Part of #331. * render: pin v2026.09.18.1, and give Render a health check for the pipeline The ingest service had no healthCheckPath, because the webhook source answered only POST /events. Render called it live when its port opened, could not tell an instance that was shutting down from one that was serving, and could not hold a deploy until the new instance was ready. https://telemetry.turbolytics.io/healthz answered 404. v2026.09.18.1 adds GET and HEAD /healthz to the webhook source (#336): 200 while it admits deliveries, 503 once it is closing, no signature, not counted as a request. render.yaml sets healthCheckPath: /healthz on the ingest service, and the Dockerfile, the Makefile and compose pin the new tag. sqlflow rollup ddl from the new image generates the committed migration byte for byte: make validate's drift check passes without regenerating. The end-to-end test, with signatures on, holds that GET and HEAD /healthz answer 200 unsigned and POST is refused: 88 assertions. The install event's sqlflow_version follows the pin, and the test reads it from the image tag rather than a literal. Part of #331. * render: the test's probes no longer leave a telemetry lease behind The end-to-end test failed once from a clean export, with install.first_request present and install.deployed missing, after passing in the working tree. The auth-mode probes start real pipelines, which run telemetry.sh, and kill them as soon as they are probed. One killed between claiming install.deployed and giving the claim back left the two-minute lease held. The pipelines started next could not claim the event until it expired, and the test waits 60 seconds. That is the lease doing its job: in a deploy, a process killed mid-send delays the event two minutes and does not lose it. In a test it is a flake. The probes now run with SQLFLOW_TELEMETRY=off, since they test auth and not telemetry, and the test clears any claim before it starts the pipelines it measures. Three runs in a row pass, 88 assertions each. Part of #331.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #331. Needed by #335.
Why
A webhook pipeline deployed as a Render web service has no health check. The webhook source answers only
POST /events, and the pipeline's/healthzis on the metrics listener,:8000. A platform routes one port to a service and checks health on that port.Found on the first deploy of the Deploy to Render template: Render marked the ingest service live because its port was open, and
https://telemetry.turbolytics.io/healthzanswers 404. Without a health path Render cannot tell a closing instance from a serving one, cannot hold a zero-downtime deploy until the new instance is ready, and an uptime monitor has nothing to call but a delivery.What
GETandHEAD /healthzon the webhook source's address.GET /healthzwhile the source admits deliveries{"status":"ok"}HEAD /healthz{"detail":"Source is closed"}, what a delivery is told thenAllow: GET, HEADDecisions:
POST /eventsis still 400.serve's/healthzdraws the same line.webhook_requests_total. That counter is how an operator counts deliveries, and a check every few seconds would bury them under 200s that delivered nothing. The route sits outside the metrics middleware.GET /healthzpattern handedPOST /healthzto the delivery mux, which answered 404 for a path that exists.POST /eventsis. The pipeline's own/healthzon the metrics listener is unchanged.Not in scope: a check that the pipeline is making progress. This says the listener is up and admitting deliveries, which is what a platform's routing decision needs.
Evidence
Close; 200 while a delivery waits on a full queue; five checks and one delivery count as one request.TestSourceWebhook_RoutesOnlyPostEventsis renamedRefusesOtherMethodsAndPaths, since its old name is no longer true. Its assertions are unchanged.GETandHEAD /healthzanswered 200 unsigned,POST /healthz405 withAllow: GET, HEAD, an unsigned delivery 400, a signed one 200.-race.make test-gopasses.After it merges
A point release, then #335 pins it and sets
healthCheckPath: /healthzon the ingest service.