Skip to content

webhook: GET and HEAD /healthz on the source's own address - #336

Merged
turbolytics merged 1 commit into
mainfrom
feat/webhook-healthz
Sep 18, 2026
Merged

turbolytics merged 1 commit into
mainfrom
feat/webhook-healthz

Conversation

@turbolytics

Copy link
Copy Markdown
Owner

Part of #331. Needed by #335.

Why

A webhook pipeline deployed as a Render web service has no health check. The webhook source answers only POST /events, and the pipeline's /healthz is on the metrics listener, :8000. A platform routes one port to a service and checks health on that port.

Found on the first deploy of the Deploy to Render template: Render marked the ingest service live because its port was open, and https://telemetry.turbolytics.io/healthz answers 404. Without a health path Render cannot tell a closing instance from a serving one, cannot hold a zero-downtime deploy until the new instance is ready, and an uptime monitor has nothing to call but a delivery.

What

GET and HEAD /healthz on the webhook source's address.

Request Answer
GET /healthz while the source admits deliveries 200 {"status":"ok"}
HEAD /healthz 200, no body. It is what probing services send.
Either, once the source is closing 503 {"detail":"Source is closed"}, what a delivery is told then
Any other method 405 with Allow: GET, HEAD

Decisions:

  • No signature, no body read. A health check has neither, and the route admits nothing to the pipeline. An unsigned POST /events is still 400.
  • It does not wait on the queue. A full queue is backpressure, not failure. A platform that reads busy as dead restarts an instance while it holds a sender's event. serve's /healthz draws the same line.
  • Not counted in webhook_requests_total. That counter is how an operator counts deliveries, and a check every few seconds would bury them under 200s that delivered nothing. The route sits outside the metrics middleware.
  • The route owns its path for every method. Beside the catch-all, a GET /healthz pattern handed POST /healthz to the delivery mux, which answered 404 for a path that exists.
  • Always on, fixed path, as POST /events is. The pipeline's own /healthz on the metrics listener is unchanged.

Not in scope: a check that the pipeline is making progress. This says the listener is up and admitting deliveries, which is what a platform's routing decision needs.

Evidence

  • Four tests, each watched failing first, with 404 or a count of 6: no signature needed while an unsigned delivery is still refused; 503 after Close; 200 while a delivery waits on a full queue; five checks and one delivery count as one request.
  • TestSourceWebhook_RoutesOnlyPostEvents is renamed RefusesOtherMethodsAndPaths, since its old name is no longer true. Its assertions are unchanged.
  • Run from this branch with HMAC on: GET and HEAD /healthz answered 200 unsigned, POST /healthz 405 with Allow: GET, HEAD, an unsigned delivery 400, a signed one 200.
  • The webhook package passes under -race. make test-go passes.

After it merges

A point release, then #335 pins it and sets healthCheckPath: /healthz on the ingest service.

A webhook pipeline deployed as a Render web service had no health check.
The webhook source answered only POST /events, and the pipeline's
/healthz is on the metrics listener, :8000. A platform routes one port
to a service and checks health on that port, so Render marked the
service live because its port was open, could not tell a closing
instance from a serving one, and an uptime monitor pointed at
https://telemetry.turbolytics.io/healthz got 404. Found on the first
deploy of the Deploy to Render template (#331, #335).

The source now answers GET and HEAD /healthz with {"status":"ok"} while
it admits deliveries, and 503 once it is closing, which is what it says
to a delivery then. HEAD because that is what probing services send.
Any other method is 405 with Allow: the route owns its path for every
method, because beside the catch-all a GET-only pattern handed
POST /healthz to the delivery mux, which answered 404 for a path that
exists.

It reads no body and checks no signature, since a health check has
neither, and admits nothing to the pipeline. It does not wait on the
queue: a full queue is backpressure, and a platform that reads busy as
dead restarts an instance while it holds a sender's event. It sits
outside the metrics middleware: webhook_requests_total is how an
operator counts deliveries, and a check every few seconds would bury
them under 200s that delivered nothing.

Four tests, each watched failing first with 404 or a count of 6: no
signature needed while an unsigned delivery is still refused; 503 after
Close; 200 while a delivery waits on a full queue; five checks and one
delivery count as one request. Run from this branch with HMAC on: GET
and HEAD answered 200 unsigned, POST /healthz 405 with Allow: GET, HEAD,
an unsigned delivery 400, a signed one 200. make test-go passes.

Part of #331.
@turbolytics
turbolytics merged commit 3d3a7b3 into main Sep 18, 2026
5 checks passed
turbolytics added a commit that referenced this pull request Sep 18, 2026
…eline

The ingest service had no healthCheckPath, because the webhook source
answered only POST /events. Render called it live when its port opened,
could not tell an instance that was shutting down from one that was
serving, and could not hold a deploy until the new instance was ready.
https://telemetry.turbolytics.io/healthz answered 404.

v2026.09.18.1 adds GET and HEAD /healthz to the webhook source (#336):
200 while it admits deliveries, 503 once it is closing, no signature,
not counted as a request. render.yaml sets healthCheckPath: /healthz on
the ingest service, and the Dockerfile, the Makefile and compose pin the
new tag.

sqlflow rollup ddl from the new image generates the committed migration
byte for byte: make validate's drift check passes without regenerating.

The end-to-end test, with signatures on, holds that GET and HEAD
/healthz answer 200 unsigned and POST is refused: 88 assertions. The
install event's sqlflow_version follows the pin, and the test reads it
from the image tag rather than a literal.

Part of #331.
turbolytics added a commit that referenced this pull request Sep 18, 2026
* render: the minute table, the series table, and a runner that applies them once

The template's schema. metrics_1m holds five values per minute per
series, the five every coarser grain merges from exactly. series turns
a dimensions_key back into dimensions and is kept by statement triggers
in the writer's transaction.

migrate.sh is the Bluesky demo's at 6ec4966, under its own lock name.
Applied twice against Postgres 18: the second run skips both. An upsert
of an older minute moves first_bucket back and leaves last_bucket. A
type that is neither count nor gauge is refused by the table.

Part of #331.

* render: five coarser grains, generated, and a total across dimensions

rollups.yml declares the ladder over metrics_1m. metrics keeps every
dimension and all five values. metrics_total drops dimensions_key for a
name with a series per user, and has no last.

0003_rollups.sql is sqlflow rollup ddl's output from v2026.09.18, byte
for byte the migration verified locally against a build of #334. Applied
to Postgres 18: 0.25 and 0.5 reach the day as 0.75 with last 0.5, and
three minutes of two series total 4.75.

Part of #331.

* render: series, metric and metric_total over HTTP

metric filters by containment: the series whose dimensions include
every pair sent. It joins series in DuckDB, after name and the range
have pushed down to Postgres, so a filtered request scans what an
unfiltered one does. A filter that is not JSON matches nothing rather
than everything. metric_total reads one row per bucket for a name with
many series.

The grains are written by hand because sqlflow rollup generates a
dataset of at most one dimension, folded. The next commit's test
queries every one.

Part of #331.

* render: a webhook pipeline that aggregates a metric by minute, tested end to end

pipeline.yml reads one metric from a request's top-level keys or
several from metrics, casts every field in SQL so one sender's bad
value cannot fail a batch, builds the series key, and merges each
closed minute into five values. A timestamp more than a minute ahead is
dropped: the window closes against the newest event time, and one
metric dated 2030 would close every minute until then.

The entrypoint exits 2 on a blank secret, an auth mode that is neither
hmac nor none, and a name prefix that could leave its quotes.

test/e2e.sh builds the image from v2026.09.18, posts signed metrics and
reads them back at all six grains of both datasets, 59 assertions. It
asks each grain for a range that holds one of its buckets: serve snaps a
range to bucket boundaries, and a pinned 1d grain over the default hour
is an empty range, which the test asserts on purpose. Seven malformed
bodies answer 200 and store nothing. Unsigned mode with a name prefix,
the collector's configuration, accepts install.deployed and drops a
name outside the prefix.

Part of #331.

* render: the Blueprint, its check, CI, and the scripts the ignore file was hiding

render.yaml declares a Postgres and two web services built from
render/. The HMAC secret is sync: false, so the deploy asks for it.
autoDeployTrigger is off, so a copy in someone's workspace does not
redeploy when this repository's main moves. CI validates the configs,
the generated migration and render.yaml, holds that an invalid plan is
refused with status 1, and runs the end-to-end test.

The repository ignores bin/, which also matched render/bin/. The four
commits before this one therefore held no migrate.sh, entrypoint.sh,
serve.sh or send.sh: the end-to-end test passed from the files on disk,
and a clone could not have built the image, because the Dockerfile
copies bin. .gitignore now excepts /render/bin/, and this commit adds
the scripts, executable.

Part of #331.

* docs: the Deploy to Render button, and the template's contract

render/README.md is what a deployer reads: the one prompt, the first
signed request, the rules a metric must meet, and the three datasets.
It says plainly that the pipeline answers 200 to anything signed and
stores only what meets the rules, and that a pinned grain needs a range
at least as wide as its bucket. Every command in it was run against the
local compose: both forms of the signed request answered received, and
the metric read back with value_sum 2.

Part of #331.

* render: the deployer chooses the client id, and no grain holds an answer for minutes

Found on the first deploy to Render.

The client id was generateValue. Render minted aA+SZt88...+mIlHiM=, and
pasted into a URL as it is, it answered 401: '+' reads as a space. It
worked only through --data-urlencode, and the deployer had to find it in
the dashboard first. It is now sync: false, asked for beside the HMAC
secret, so the deployer has it when the deploy finishes. serve.sh
refuses a blank one and one that holds anything a URL would encode, and
names `openssl rand -hex 16`. There is no default: in this template the
id is the only thing between a reader and the deployer's metrics, and
the README now says so.

The coarse grains held an answer for 2, 5 and 10 minutes, the Bluesky
demo's times. Asked at 1d before its first event had landed, the live
collector answered empty with "cache": "hit" for ten minutes while the
row sat in the table. A new deploy sends one metric and asks for it, so
every grain now holds for the same 30 seconds.

The end-to-end test gains four assertions, 63 in all: the API exits 2
on a blank id and on one with '+', '/' and '='; the id works pasted into
a URL unencoded; a wrong id answers 401.

Part of #331.

* render: two pipeline processes writing one minute are summed, and an invariant holds it

A pipeline process holds its own window and publishes a closed minute by
replacing the row for its key. Keyed on the series alone, a second
process holding part of the same minute replaced the first. Run as two
real instances and sent 7 and 5 of one minute, the template stored 7,
and every rollup re-merged the 7. Two instances is what a scaled service
runs, and what every Render deploy runs for a moment, the old beside the
new.

The pipeline now writes metrics_1m_writers, keyed on the series and on a
writer id that entrypoint.sh makes new on every start and never reads
from the environment: a value set once would be shared by every
instance. A trigger merges every writer's row into metrics_1m in the
same transaction. Sums add, min and max nest, and value_last is the
reading with the latest event time across processes. series, the
rollups, serve.yml and the API are untouched: they read metrics_1m and
never see a writer.

The first version of the trigger locked per minute, and four concurrent
writers deadlocked within one round. Postgres's log named a rollup
advisory lock on one side and a series row lock on the other: an upsert
fires the trigger chain twice, once for the rows it inserted and once
for the rows it updated, so a transaction can hold a rollup lock from
its first pass and want a series row in its second. The trigger now
takes one lock before it touches anything another writer wants, and
writers run the chain one at a time. A publish is milliseconds, once
per poll.

pipeline.writers.merge_exactly is registered in the invariant matrix,
proven for pipeline.stateless and empty for pipeline.stateful, which is
accurate: nothing in the engine makes it true, a destination's schema
does. internal/rendertemplate proves it against the template's own
migrations. Two controls must fail and do, each storing 5|1|5|5|5 where
12|2|5|7|5 is right: the lock removed, and two processes sharing a
writer id. The merge refuses REPEATABLE READ. Four writers publish and
republish 36 overlapping keys through the Postgres sink for forty
rounds, and metrics_1m and all ten rollup tables equal what the writers
remember, worked out in Go and never from the database. Deleting the
writers' rows afterwards changes nothing, which the README tells a
deployer they may do. On failure the test prints the server's deadlock
report, which is how the second defect was diagnosed.

make test runs two instances and sends each a share of one minute: 76
assertions, the split minute and its gauge's value_last at all six
grains.

Part of #331.

* render: a sync cannot close an open pipeline, and an empty auth mode cannot open a signed one

Two defects in how the auth mode reaches the pipeline.

render.yaml set SQLFLOW_WEBHOOK_AUTH to hmac. Render's Blueprint
reference says it "preserves existing environment variables, even if you
omit them from the Blueprint file", and that a resource "retains any
existing environment variable values that aren't overwritten by the
Blueprint". So a value in the file is rewritten on every sync. A pipeline
set to none in the dashboard, which is what the telemetry collector is,
would go back to hmac on the next sync and start refusing its senders
without anyone having changed it. The variable is no longer in the file.
Unset, the entrypoint defaults to hmac, and a value added in the
dashboard survives a sync. No third prompt.

Checking that default found the second. entrypoint.sh and pipeline.yml
each decided the mode, and disagreed about an empty string. The script
read it as hmac and was satisfied by the secret. The template's
`default('hmac') == 'hmac'` read it as not-hmac and rendered no signature
block. Run with SQLFLOW_WEBHOOK_AUTH set to empty and a secret set, the
pipeline answered an unsigned request 200. The entrypoint now settles
the mode and exports it, so the template reads only a checked value, and
the template signs unless the mode is exactly none, so it fails closed
even without the script.

The end-to-end test starts a pipeline with the mode empty and with it
hmac and holds that both refuse an unsigned request: 78 assertions.
Compose passes the variable through rather than defaulting it, so the
main flow runs with it absent, as Render leaves it.

Part of #331.

* render: two install events, once each, disclosed in full and off with one word

The template tells the sqlflow maintainers that it was deployed and that
it received its first metric. Two events per install, ever, and nothing
else. render/README.md, "What this sends", prints both payloads in full,
and SQLFLOW_TELEMETRY is a third deploy prompt so it is seen before
anything is sent. Blank is on; off, false, 0 and no send nothing, and
the log says so. The sqlflow binary sends nothing: bin/telemetry.sh, 91
lines of shell, is the only code that does.

An event is an ordinary metric posted to
https://telemetry.turbolytics.io, which is this same template with
signatures off and SQLFLOW_METRIC_NAME_PREFIX=install. It carries
install_id, a uuid the database makes for itself from nothing about the
deployer; source: render, a constant, so other vendors' templates can
report to one collector and be told apart; template; and the image tag.
install.first_request asks Postgres whether metrics_1m has a row whose
name is not install.*, and reads nothing about it.

The hostname is one we own and not the collector's onrender.com URL: a
deployed copy never updates itself, so whatever ships is what it calls
for as long as it runs.

A send is claimed in Postgres first, a lease that expires after two
minutes, so two pipeline instances starting together report one install
and a process that dies mid-send does not lose the event. It is recorded
as sent only on a 2xx: the collector's custom domain answered 404 while
it propagated, and an install recorded as sent on a 404 is never
counted. A send is bounded at five seconds, runs in the background, and
cannot stop the pipeline.

Verified against the live collector by hand first: ten POSTs through the
hostname answered 200, a 5000-byte body 413, and an install.deployed
with source render read back through its API, filtered by source.

The end-to-end test runs a local collector, the same template, and no
service in compose can reach the real one: after the run the real
collector held only the two rows sent by hand. 87 assertions. Nine are
new: one series per event from two instances starting together, the
dimensions exactly, each sum 1, still 1 after both restart, off claims
and sends nothing, and a dead collector leaves the event unsent and the
pipeline answering.

The test's first attempt failed and the client was right. The collector
logged "dropped late rows": under the test's five-second idle bound a
minute closes while it is still current, and the second event landed in
the minute the first had closed. A deploy's bound is 60, where a closed
minute has always ended. The test now starts its burst in a fresh
minute. The same run showed install.deployed itself, in the shared test
database, counting as the first metric, which is why install.* names
are excluded.

Part of #331.

* render: pin v2026.09.18.1, and give Render a health check for the pipeline

The ingest service had no healthCheckPath, because the webhook source
answered only POST /events. Render called it live when its port opened,
could not tell an instance that was shutting down from one that was
serving, and could not hold a deploy until the new instance was ready.
https://telemetry.turbolytics.io/healthz answered 404.

v2026.09.18.1 adds GET and HEAD /healthz to the webhook source (#336):
200 while it admits deliveries, 503 once it is closing, no signature,
not counted as a request. render.yaml sets healthCheckPath: /healthz on
the ingest service, and the Dockerfile, the Makefile and compose pin the
new tag.

sqlflow rollup ddl from the new image generates the committed migration
byte for byte: make validate's drift check passes without regenerating.

The end-to-end test, with signatures on, holds that GET and HEAD
/healthz answer 200 unsigned and POST is refused: 88 assertions. The
install event's sqlflow_version follows the pin, and the test reads it
from the image tag rather than a literal.

Part of #331.

* render: the test's probes no longer leave a telemetry lease behind

The end-to-end test failed once from a clean export, with
install.first_request present and install.deployed missing, after
passing in the working tree.

The auth-mode probes start real pipelines, which run telemetry.sh, and
kill them as soon as they are probed. One killed between claiming
install.deployed and giving the claim back left the two-minute lease
held. The pipelines started next could not claim the event until it
expired, and the test waits 60 seconds.

That is the lease doing its job: in a deploy, a process killed mid-send
delays the event two minutes and does not lose it. In a test it is a
flake. The probes now run with SQLFLOW_TELEMETRY=off, since they test
auth and not telemetry, and the test clears any claim before it starts
the pipelines it measures. Three runs in a row pass, 88 assertions each.

Part of #331.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant