diff --git a/.claude/skills/slow-soak/slow.yml b/.claude/skills/slow-soak/slow.yml index f40042e7..e68ce7c0 100644 --- a/.claude/skills/slow-soak/slow.yml +++ b/.claude/skills/slow-soak/slow.yml @@ -33,8 +33,6 @@ tables: size_seconds: 60 grace_seconds: 60 idle_close_seconds: 60 - late_rows: drop - poll_interval_seconds: 10 emit_sql: | SELECT bucket, sum(count)::BIGINT AS count, max(max_latency_s) AS max_latency_s FROM closed @@ -64,13 +62,16 @@ pipeline: auto_offset_reset: earliest topics: - "{{ SQLFLOW_TOPIC|default('slow-soak') }}" + event_time: + path: sent_at + format: rfc3339 handler: type: 'handlers.InferredMemBatch' sql: | INSERT INTO agg_slow BY NAME SELECT - date_trunc('minute', CAST(sent_at AS TIMESTAMPTZ)) AS bucket, + time_bucket(INTERVAL '1 minute', event_time) AS bucket, count(*) AS count, max(epoch(now() - CAST(sent_at AS TIMESTAMPTZ))) AS max_latency_s FROM batch diff --git a/CHANGELOG.md b/CHANGELOG.md index cd5c651f..ac48e01c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,34 @@ ### Changed +- **The window's manager no longer polls, and lateness is decided when a + record arrives.** The engine tells the manager when a window's watermark + moved; the manager publishes what closed, republishes what a late row + landed in, and deletes what is past its lateness. It has no clock. + `poll_interval_seconds` is gone: nothing is left for it to pace, and a + config that sets it fails validation naming the change. + + `late_rows` is gone too, replaced by `allowed_lateness_seconds` (default 0): + Flink's `allowedLateness`. With 0, a record for a bucket that closed is + refused before the handler and counted in + `window_late_rows_total{outcome="refused"}` -- it never reaches the window + table, which is what `drop` did after the fact. With a positive value a + closed bucket's rows are kept that long, a late record within it is written, + and the bucket is republished as a **whole value** -- `emit_sql` over every + row it has -- so the sink must replace by key; validate refuses an + append-only sink. This replaces `reemit`, which published the late rows + alone as a delta that only an additive sink could use and that a replacing + sink turned into loss. The sink now always receives the exact value, and + `window_recomputes_total` counts each republication. + + `time_column` must be `time_bucket(INTERVAL '', event_time)`, which + validate now refuses rather than warns: the engine decides lateness from + that bucket, and the six shipped windowed examples are updated to it. The + Render template follows when its image pin moves to a release that carries + this. The wire bundle's `late_rows_reemitted` is gone and + `late_rows_recomputed` is new; `late_rows_dropped` now counts refusals. + See `docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md`. + - **Windowing is now decided by a watermark the engine asserts.** A window manager needs one thing, how far the stream has got, and nothing used to tell it: it inferred the answer from the newest bucket in its own table and diff --git a/README.md b/README.md index 7e7e05d9..39e5e83b 100644 --- a/README.md +++ b/README.md @@ -1203,8 +1203,7 @@ tables: size_seconds: 3600 grace_seconds: 0 idle_close_seconds: 60 - late_rows: drop # or reemit; required - poll_interval_seconds: 10 # optional + allowed_lateness_seconds: 0 # optional; see Late rows emit_sql: | # optional; default SELECT * FROM closed SELECT bucket, city, sum(count)::INT AS count FROM closed @@ -1216,25 +1215,38 @@ tables: topic: output-tumbling-window-1 pipeline: + source: + type: kafka + kafka: + brokers: [localhost:9092] + group_id: tumbling + topics: [tumbling-window] + event_time: {path: timestamp, format: rfc3339} handler: type: handlers.InferredMemBatch sql: | INSERT INTO agg_cities_count BY NAME SELECT - date_trunc('hour', CAST(timestamp AS TIMESTAMPTZ)) AS bucket, + time_bucket(INTERVAL '1 hour', event_time) AS bucket, properties.city AS city, count(*) AS count FROM batch GROUP BY bucket, city ``` +The handler must cut `time_column` as `time_bucket(INTERVAL '', +event_time)`: the engine decides each record's lateness from that bucket, and +`sqlflow validate` refuses any other expression. `event_time` is the column +the source fills from the record's own time; see **Event time** under each +source. + **The watermark.** The engine asserts it, the window follows it. One instant per window, in event time, written to `sqlflow_watermarks` in the same commit as the rows it describes, never moving backwards. Its meaning is a promise: as of this commit, every row this pipeline will ever write for this window has `time_column` at or after the watermark. A bucket is closed when its end, `time_column + size_seconds`, is at or before it; `sqlflow_windows` holds what -each window has closed up to, and a poll moves that to the assertion and +each window has closed up to, and a pass moves that to the assertion and publishes the buckets between them. That is the whole decision -- no clock, no liveness reading, one comparison in one clock. @@ -1265,33 +1277,38 @@ websocket reconnecting, are the first case: the window holds through the outage, and the clock restarts on the return. A source that can never deliver holds every bucket open until it does, which the log says. -`poll_interval_seconds` only bounds how soon a closed bucket is noticed; on -shutdown the drain asserts first, so the final poll sees the watermark up to -the signal. A pipeline's idle ticks write nothing unless a watermark moved -- -no statement, no WAL append, no fsync -- so a quiet stream costs one write, -when its last partition goes idle. `sqlflow_progress` is liveness for `/stats` -and `/healthz` and for SQL to read; no window decision rests on it. +Nothing polls. The engine signals the window's manager after the commit that +moved its watermark, and the manager passes: publishes what closed, deletes +what is past its lateness, records where it closed up to. It passes once when +it starts and once on the drain, which asserts first so the final pass sees +the watermark up to the signal, and otherwise only when told. A pipeline's +idle ticks write nothing unless a watermark moved -- no statement, no WAL +append, no fsync -- so a quiet stream costs one write, when its last partition +goes idle. `sqlflow_progress` is liveness for `/stats` and `/healthz` and for +SQL to read; no window decision rests on it. Every decision a window makes is a row in one of two tables, rendered from the code to [docs/windows/decisions.md](docs/windows/decisions.md). -**Late rows.** A row for a bucket below the watermark arrived after that -bucket was published. `late_rows` is required, because the two policies are -different promises to the sink. `drop` deletes the row and counts it in -`window_late_rows_total`, so the sink sees each bucket once, as it closed. -`reemit` runs `emit_sql` over the late rows alone and publishes the result, -because the bucket's other rows were deleted when it closed. The sink has to -add that result to the bucket it holds. An upsert that replaces on the -bucket's key overwrites the bucket's count with the late rows' count, so pair -a replacing upsert with `drop`. An upsert that adds counts a republished close -twice; see **Guarantees**. `sqlflow validate` warns when `reemit` is paired -with the Iceberg or Kafka sink. A rising drop count means the grace is too -short for the stream. - -The watermark is one value for the whole table, which makes it the fastest -partition's clock. A topic whose partitions run at uneven rates has a slow -partition whose rows arrive late, and the grace is the allowance for them. -Size it from `window_late_rows_total`. +**Late rows.** A record whose bucket ended at or before the watermark is +late. The engine decides that when the record arrives, from its `event_time` +and the window's size, before the handler runs -- the way Flink's window +operator does. With `allowed_lateness_seconds` absent or `0` a late record is +refused and counted in `window_late_rows_total{outcome="refused"}`; the +window table never holds it. With a positive value a closed bucket's rows are +kept that long, a late record within it is written, and the bucket is +republished as a **whole value** -- `emit_sql` over every row it has -- so the +sink must replace by key; `sqlflow validate` refuses an append-only sink and +warns for one it cannot classify. Past `end + allowed_lateness_seconds` the +rows are deleted and later records are refused. Each republication counts in +`window_recomputes_total`, and the rows that caused it in +`window_late_rows_total{outcome="recomputed"}`. A rising refused count means +the grace is too short for the stream, or the lateness is. + +The watermark is the minimum over the partitions, so a slow partition holds +the buckets it shares open rather than arriving late; the grace is the +allowance for reordering within a partition. Size it from +`window_late_rows_total`. **emit_sql** shapes the rows before the sink. It reads one relation, `closed`, holding every row of every bucket that just closed. Cast a `sum` back to the @@ -1418,14 +1435,17 @@ joins, with or without a state path: |---|---|---|---| | `window_watermark_seconds` | `window_watermark_seconds` | gauge | `window` | | `window_closed` | `window_closed_total` | counter | `window` | -| `window_late_rows` | `window_late_rows_total` | counter | `window`, `policy` | +| `window_late_rows` | `window_late_rows_total` | counter | `window`, `outcome` | +| `window_recomputes` | `window_recomputes_total` | counter | `window` | | `window_close_lag_seconds` | `window_close_lag_seconds` | gauge | `window` | | `window_newest_bucket_start_seconds` | `window_newest_bucket_start_seconds` | gauge | `window` | `window_closed_total` is present from startup at zero, so a windowed pipeline is never mistaken for one without windows before its first close. -`window_late_rows_total` counts only after the close that dropped or reemitted -the rows commits. +`window_late_rows_total` is the engine's: `outcome="refused"` counts records +refused before the handler, `outcome="recomputed"` counts records admitted +within the lateness, each counted once the commit that admitted it lands. +`window_recomputes_total` counts the buckets republished whole for them. `window_close_lag_seconds` is how far the window's closes trail the data it holds, in event time, and it is the number to alert on: zero while closes keep @@ -1434,8 +1454,8 @@ uses no clock, so a sparse stream does not read as stalled and a host whose clock is wrong reports it correctly; a stream going quiet shows in `pipeline_last_message_timestamp` instead. `window_newest_bucket_start_seconds` ahead of a trusted clock means rows are stamped in the future, which moves the -watermark past every correctly-timed row and, under `late_rows: drop`, deletes -them. It is updated on every poll that finds rows, even one whose close fails. +watermark past every correctly-timed row and refuses them as late. It is +updated on every pass that finds rows, even one whose close fails. `window_watermark_seconds` trails the newest bucket by `grace_seconds` by design, which makes its age the number to size `grace_seconds` from rather than one to alert on. diff --git a/dev/bench/bluesky/demo-10x-noop-sink.yml b/dev/bench/bluesky/demo-10x-noop-sink.yml index 52e529a1..1f870b4e 100644 --- a/dev/bench/bluesky/demo-10x-noop-sink.yml +++ b/dev/bench/bluesky/demo-10x-noop-sink.yml @@ -6,7 +6,6 @@ # shrinks by the same factor, and the ratios between messages, batches, # polls and flushes stay what they are live: # -# poll_interval_seconds 10 -> 1 # flush_interval_seconds 30 (the default) -> 3 # # grace_seconds is stream time and stays. idle_close_seconds only fires when @@ -17,6 +16,7 @@ commands: sql: | CREATE TABLE IF NOT EXISTS posts ( time_us BIGINT, + event_time TIMESTAMPTZ, commit STRUCT( operation TEXT, record STRUCT(langs TEXT[]) @@ -50,18 +50,13 @@ tables: # land. After idle_close_seconds with nothing arriving, every open # bucket closes. # - # late_rows: drop, because the sink replaces. A row for a bucket that - # already closed is discarded and counted in window_late_rows_total. - # reemit would run emit_sql over the late rows alone, because the - # bucket's other rows were deleted when it closed, and the upsert would - # replace the bucket's count with theirs. validate refuses that pairing. + # No lateness: a post for a bucket that already closed is refused + # before the handler and counted in window_late_rows_total. window: time_column: bucket size_seconds: 60 grace_seconds: 60 idle_close_seconds: 60 - late_rows: drop - poll_interval_seconds: 1 emit_sql: | SELECT bucket, lang, sum(posts)::INTEGER AS posts FROM closed @@ -80,6 +75,9 @@ pipeline: type: websocket websocket: uri: "{{ SQLFLOW_JETSTREAM_URI|default('wss://jetstream2.us-east.bsky.network/subscribe?wantedCollections=app.bsky.feed.post') }}" + event_time: + path: time_us + format: unix_us handler: type: handlers.StructuredBatch @@ -87,7 +85,7 @@ pipeline: sql: | INSERT INTO posts_per_minute_by_lang SELECT - time_bucket(INTERVAL '1 minute', to_timestamp(time_us / 1000000)) AS bucket, + time_bucket(INTERVAL '1 minute', event_time) AS bucket, coalesce(commit.record.langs[1], 'unknown') AS lang, count(*) AS posts FROM posts diff --git a/dev/bench/bluesky/demo-10x-sqlcommand.yml b/dev/bench/bluesky/demo-10x-sqlcommand.yml index b117dd12..30c08938 100644 --- a/dev/bench/bluesky/demo-10x-sqlcommand.yml +++ b/dev/bench/bluesky/demo-10x-sqlcommand.yml @@ -7,7 +7,6 @@ # shrinks by the same factor, and the ratios between messages, batches, # polls and flushes stay what they are live: # -# poll_interval_seconds 10 -> 1 # flush_interval_seconds 30 (the default) -> 3 # # grace_seconds is stream time and stays. idle_close_seconds only fires when @@ -27,6 +26,7 @@ commands: sql: | CREATE TABLE IF NOT EXISTS posts ( time_us BIGINT, + event_time TIMESTAMPTZ, commit STRUCT( operation TEXT, record STRUCT(langs TEXT[]) @@ -60,18 +60,13 @@ tables: # land. After idle_close_seconds with nothing arriving, every open # bucket closes. # - # late_rows: drop, because the sink replaces. A row for a bucket that - # already closed is discarded and counted in window_late_rows_total. - # reemit would run emit_sql over the late rows alone, because the - # bucket's other rows were deleted when it closed, and the upsert would - # replace the bucket's count with theirs. validate refuses that pairing. + # No lateness: a post for a bucket that already closed is refused + # before the handler and counted in window_late_rows_total. window: time_column: bucket size_seconds: 60 grace_seconds: 60 idle_close_seconds: 60 - late_rows: drop - poll_interval_seconds: 1 emit_sql: | SELECT bucket, lang, sum(posts)::INTEGER AS posts FROM closed @@ -104,6 +99,9 @@ pipeline: type: websocket websocket: uri: "{{ SQLFLOW_JETSTREAM_URI|default('wss://jetstream2.us-east.bsky.network/subscribe?wantedCollections=app.bsky.feed.post') }}" + event_time: + path: time_us + format: unix_us handler: type: handlers.StructuredBatch @@ -111,7 +109,7 @@ pipeline: sql: | INSERT INTO posts_per_minute_by_lang SELECT - time_bucket(INTERVAL '1 minute', to_timestamp(time_us / 1000000)) AS bucket, + time_bucket(INTERVAL '1 minute', event_time) AS bucket, coalesce(commit.record.langs[1], 'unknown') AS lang, count(*) AS posts FROM posts diff --git a/dev/bench/bluesky/demo-10x.yml b/dev/bench/bluesky/demo-10x.yml index bd5ee91f..acd9795a 100644 --- a/dev/bench/bluesky/demo-10x.yml +++ b/dev/bench/bluesky/demo-10x.yml @@ -7,7 +7,6 @@ # shrinks by the same factor, and the ratios between messages, batches, # polls and flushes stay what they are live: # -# poll_interval_seconds 10 -> 1 # flush_interval_seconds 30 (the default) -> 3 # # grace_seconds is stream time and stays. idle_close_seconds only fires when @@ -18,6 +17,7 @@ commands: sql: | CREATE TABLE IF NOT EXISTS posts ( time_us BIGINT, + event_time TIMESTAMPTZ, commit STRUCT( operation TEXT, record STRUCT(langs TEXT[]) @@ -51,18 +51,13 @@ tables: # land. After idle_close_seconds with nothing arriving, every open # bucket closes. # - # late_rows: drop, because the sink replaces. A row for a bucket that - # already closed is discarded and counted in window_late_rows_total. - # reemit would run emit_sql over the late rows alone, because the - # bucket's other rows were deleted when it closed, and the upsert would - # replace the bucket's count with theirs. validate refuses that pairing. + # No lateness: a post for a bucket that already closed is refused + # before the handler and counted in window_late_rows_total. window: time_column: bucket size_seconds: 60 grace_seconds: 60 idle_close_seconds: 60 - late_rows: drop - poll_interval_seconds: 1 emit_sql: | SELECT bucket, lang, sum(posts)::INTEGER AS posts FROM closed @@ -70,7 +65,7 @@ tables: sink: # A keyed sink replaces the row a (bucket, lang) already identifies. # The engine deletes a window only after the sink accepts it, so a - # crash between the two republishes that window on the next poll, + # crash between the two republishes that window on the next pass, # and the republish lands harmlessly. Each flush costs the batch: a # COPY into a staging table and a server-side merge, in one # transaction. The target needs PRIMARY KEY (bucket, lang) or a @@ -91,6 +86,9 @@ pipeline: type: websocket websocket: uri: "{{ SQLFLOW_JETSTREAM_URI|default('wss://jetstream2.us-east.bsky.network/subscribe?wantedCollections=app.bsky.feed.post') }}" + event_time: + path: time_us + format: unix_us handler: type: handlers.StructuredBatch @@ -98,7 +96,7 @@ pipeline: sql: | INSERT INTO posts_per_minute_by_lang SELECT - time_bucket(INTERVAL '1 minute', to_timestamp(time_us / 1000000)) AS bucket, + time_bucket(INTERVAL '1 minute', event_time) AS bucket, coalesce(commit.record.langs[1], 'unknown') AS lang, count(*) AS posts FROM posts diff --git a/dev/config/examples/bluesky/bluesky.kafka.windowed.yml b/dev/config/examples/bluesky/bluesky.kafka.windowed.yml index 852fce2d..1e511fe3 100644 --- a/dev/config/examples/bluesky/bluesky.kafka.windowed.yml +++ b/dev/config/examples/bluesky/bluesky.kafka.windowed.yml @@ -20,16 +20,14 @@ tables: # whether the data arrives live or as a replay. After idle_close_seconds # with nothing arriving, every open bucket closes. # - # late_rows: drop discards a row for a bucket that already closed and - # counts it in window_late_rows_total, so each bucket reaches the topic - # once. reemit would publish a second message for the bucket, counted - # over the late rows alone, for a consumer that adds it to the first. + # No lateness, because the sink is a topic, which appends: a post for a + # bucket that already closed is refused before the handler and counted + # in window_late_rows_total, so each bucket reaches the topic once. window: time_column: bucket size_seconds: 60 grace_seconds: 60 idle_close_seconds: 60 - late_rows: drop emit_sql: | SELECT strftime(bucket, '%Y-%m-%dT%H:%M:%S') AS iso_string, @@ -49,13 +47,17 @@ pipeline: type: websocket websocket: uri: 'wss://jetstream2.us-east.bsky.network/subscribe?wantedCollections=app.bsky.feed.post' + # The post's own time: time_us, microseconds since the epoch. + event_time: + path: time_us + format: unix_us handler: type: 'handlers.InferredMemBatch' sql: | INSERT INTO agg_post_kind_count_1_minute SELECT - time_bucket(INTERVAL '1 minute', to_timestamp(time_us / 1000000)) as bucket, + time_bucket(INTERVAL '1 minute', event_time) AS bucket, kind, count(*) as count FROM batch diff --git a/dev/config/examples/bluesky/bluesky.postgres.windowed.yml b/dev/config/examples/bluesky/bluesky.postgres.windowed.yml index 9e2f8615..f9cab51a 100644 --- a/dev/config/examples/bluesky/bluesky.postgres.windowed.yml +++ b/dev/config/examples/bluesky/bluesky.postgres.windowed.yml @@ -5,6 +5,7 @@ commands: sql: | CREATE TABLE IF NOT EXISTS posts ( time_us BIGINT, + event_time TIMESTAMPTZ, commit STRUCT( operation TEXT, record STRUCT(langs TEXT[]) @@ -38,18 +39,19 @@ tables: # land. After idle_close_seconds with nothing arriving, every open # bucket closes. # - # late_rows: drop, because the sink replaces. A row for a bucket that - # already closed is discarded and counted in window_late_rows_total. - # reemit would run emit_sql over the late rows alone, because the - # bucket's other rows were deleted when it closed, and the upsert would - # replace the bucket's count with theirs. validate refuses that pairing. + # allowed_lateness_seconds keeps a closed minute's rows for five more + # minutes. A post that arrives for it in that time is written, and the + # minute is republished whole -- emit_sql over every row it has -- so + # the upsert below replaces the row for (bucket, lang) with the exact + # count. Past that a late post is refused before the handler and + # counted in window_late_rows_total. This is the one shipped example + # that admits late rows, because its sink replaces by key. window: time_column: bucket size_seconds: 60 grace_seconds: 60 idle_close_seconds: 60 - late_rows: drop - poll_interval_seconds: 10 + allowed_lateness_seconds: 300 emit_sql: | SELECT bucket, lang, sum(posts)::INTEGER AS posts FROM closed @@ -57,8 +59,9 @@ tables: sink: # A keyed sink replaces the row a (bucket, lang) already identifies. # The engine deletes a window only after the sink accepts it, so a - # crash between the two republishes that window on the next poll, - # and the republish lands harmlessly. Each flush costs the batch: a + # crash between the two republishes that window on the next pass, + # and the republish lands harmlessly; so does a late post's + # recompute. Each flush costs the batch: a # COPY into a staging table and a server-side merge, in one # transaction. The target needs PRIMARY KEY (bucket, lang) or a # unique index on the pair; the sink checks at startup. @@ -77,6 +80,11 @@ pipeline: type: websocket websocket: uri: 'wss://jetstream2.us-east.bsky.network/subscribe?wantedCollections=app.bsky.feed.post' + # The post's own time: time_us, microseconds since the epoch. The + # engine fills the batch table's event_time column from it. + event_time: + path: time_us + format: unix_us handler: type: handlers.StructuredBatch @@ -84,7 +92,7 @@ pipeline: sql: | INSERT INTO posts_per_minute_by_lang SELECT - time_bucket(INTERVAL '1 minute', to_timestamp(time_us / 1000000)) AS bucket, + time_bucket(INTERVAL '1 minute', event_time) AS bucket, coalesce(commit.record.langs[1], 'unknown') AS lang, count(*) AS posts FROM posts diff --git a/dev/config/examples/kafka.stateful.window.yml b/dev/config/examples/kafka.stateful.window.yml index a1a0cb6d..e080d485 100644 --- a/dev/config/examples/kafka.stateful.window.yml +++ b/dev/config/examples/kafka.stateful.window.yml @@ -34,8 +34,6 @@ tables: size_seconds: 3600 grace_seconds: 3600 idle_close_seconds: 3600 - late_rows: drop - poll_interval_seconds: 3600 # GROUP BY ALL groups by the output columns. GROUP BY bucket would # group by the table's bucket column, which the output renames. emit_sql: | @@ -63,6 +61,10 @@ pipeline: auto_offset_reset: earliest topics: - "{{ SQLFLOW_TOPIC|default('input-stateful-window') }}" + # The record's own time, from its payload, which the handler buckets on. + event_time: + path: timestamp + format: rfc3339 handler: type: 'handlers.InferredMemBatch' @@ -70,7 +72,7 @@ pipeline: INSERT INTO agg_city_count BY NAME SELECT - date_trunc('hour', CAST(timestamp AS TIMESTAMPTZ)) AS bucket, + time_bucket(INTERVAL '1 hour', event_time) AS bucket, properties.city AS city, count(*) AS count FROM batch diff --git a/dev/config/examples/logs.rollup.clickhouse.yml b/dev/config/examples/logs.rollup.clickhouse.yml index 85bb71a7..47bafd81 100644 --- a/dev/config/examples/logs.rollup.clickhouse.yml +++ b/dev/config/examples/logs.rollup.clickhouse.yml @@ -22,7 +22,7 @@ # bytes UInt64 # ) ENGINE = MergeTree() ORDER BY (bucket, route_class, status); # -# Read it with sum(requests), not requests -- see late_rows below. +# Each bucket is published once, so a row is the hour's count. tables: sql: - name: agg_logs_rollup @@ -49,25 +49,14 @@ tables: # default of thirty the last buckets would close 10 to 40 seconds # after the corpus ends. idle_close_seconds: 10 - # reemit, not drop, and this is the setting to understand. - # - # With `drop` this pipeline silently lost 1,060 of 3,461,542 requests - # (0.031%), all in one bucket, while logging only "dropped late rows: - # 48" -- 48 is a count of aggregate rows, each carrying many requests. - # - # The source was in perfect event-time order, so the lateness was - # manufactured here: one 50,000-row batch spans many hours of traffic, - # and that batch's own newest bucket closes earlier buckets in the same - # batch. Raising grace_seconds does NOT fix it -- 7200 dropped exactly - # the same 1,060 -- and smaller batches only scale it down (batch 1000 - # still lost 603). - # - # reemit publishes the late rows as their own row instead of discarding - # them. The count then balances exactly: 3,461,542 requests in - # ClickHouse against 3,461,613 published less 71 unparseable. The cost - # is that a bucket can appear more than once, so every reader must - # SUM(requests) rather than read a single row. - late_rows: reemit + # No lateness. A record is late when its bucket ended at or before + # the watermark, and the watermark is asserted in the commit that + # makes a batch's rows visible: nothing in a batch is late against + # the batch's own newest record, so an in-order corpus is never late + # here, whatever the batch size. A record that is late is refused + # before the handler and counted in window_late_rows_total. + # allowed_lateness_seconds above 0 would republish a bucket whole, + # and the ClickHouse sink inserts rows rather than replacing them. emit_sql: | SELECT strftime(bucket, '%Y-%m-%d %H:%M:%S') AS bucket, @@ -118,25 +107,27 @@ pipeline: # segment has no directory left once the filename is dropped, which would # otherwise render as //*.html. # - # Rows whose timestamp did not parse are excluded here and kept in full by - # the archive pipeline. A NULL bucket would corrupt the watermark. + # The bucket is cut on event_time, which the source fills from the log + # line's own timestamp; a line whose timestamp does not parse is refused + # by the source and kept in full by the archive pipeline. sql: | WITH m AS ( SELECT regexp_extract(message, '^(\S+) \S+ \S+ \[([^\]]+)\] "(\S+) (\S+)[^"]*" (\d{3}) (\S+)$', - ['host','ts','method','path','status','bytes']) AS f + ['host','ts','method','path','status','bytes']) AS f, + event_time FROM batch ), s AS ( SELECT f.*, string_split(trim(f.path,'/'),'/') AS segs, lower(regexp_extract(f.path,'\.([a-zA-Z0-9]+)$',1)) AS ext, - try_strptime(f.ts,'%d/%b/%Y:%H:%M:%S %z') AS ets + event_time FROM m ) INSERT INTO agg_logs_rollup BY NAME SELECT - date_trunc('hour', ets) AS bucket, + time_bucket(INTERVAL '1 hour', event_time) AS bucket, regexp_replace( regexp_replace( '/'||array_to_string(list_slice( @@ -147,7 +138,6 @@ pipeline: count(*) AS requests, sum(TRY_CAST(bytes AS BIGINT)) AS bytes FROM s - WHERE ets IS NOT NULL GROUP BY 1,2,3 # The handler's statement is an INSERT into the windowed table, so it returns diff --git a/dev/config/examples/mqtt.duckdb.agg.yml b/dev/config/examples/mqtt.duckdb.agg.yml index c559d3a6..2a9cdc0b 100644 --- a/dev/config/examples/mqtt.duckdb.agg.yml +++ b/dev/config/examples/mqtt.duckdb.agg.yml @@ -40,7 +40,6 @@ tables: size_seconds: 60 grace_seconds: 0 idle_close_seconds: 90 - late_rows: drop # Each batch appends its own partial aggregate. emit_sql combines the # partials of every bucket that closed. emit_sql: | @@ -81,6 +80,12 @@ pipeline: broker: {{ SQLFLOW_MQTT_BROKER|default('tcp://localhost:1883') }} client_id: {{ SQLFLOW_MQTT_CLIENT_ID|default('sqlflow-edge-agg-1') }} topics: ["sensors/#"] + # The reading's own time, from its ts field. A reading whose ts does + # not parse is refused before the handler and counted, rather than + # bucketed on the wrong clock. + event_time: + path: ts + format: rfc3339 on_error: policy: IGNORE @@ -95,7 +100,7 @@ pipeline: sql: | INSERT INTO readings_1m_partial BY NAME SELECT - time_bucket(INTERVAL 1 MINUTE, TRY_CAST(ts AS TIMESTAMPTZ)) AS bucket, + time_bucket(INTERVAL '1 minute', event_time) AS bucket, device_id, metric, count(*) AS n, @@ -104,7 +109,6 @@ pipeline: max(TRY_CAST(value AS DOUBLE)) AS max_value FROM batch WHERE TRY_CAST(value AS DOUBLE) IS NOT NULL - AND TRY_CAST(ts AS TIMESTAMPTZ) IS NOT NULL GROUP BY ALL sink: diff --git a/dev/config/examples/tumbling.window.yml b/dev/config/examples/tumbling.window.yml index 6fe9a847..01251cc9 100644 --- a/dev/config/examples/tumbling.window.yml +++ b/dev/config/examples/tumbling.window.yml @@ -30,10 +30,11 @@ tables: # first row arrives. Zero says that plainly. grace_seconds: 0 idle_close_seconds: 60 - # The sink is a topic, which appends. drop keeps each bucket to one - # message; reemit would publish a second message counted over the - # late rows alone, which the consumer has to add to the first. - late_rows: drop + # The sink is a topic, which appends, so no lateness: a row for a + # bucket that has closed is refused before the handler and counted in + # window_late_rows_total, and each bucket reaches the topic once. + # allowed_lateness_seconds above 0 would republish a bucket whole, + # which needs a sink that replaces by key. # emit_sql shapes the closed rows. closed is every row of every bucket # that just closed. GROUP BY ALL groups by the output columns, which # is what the renamed bucket needs. @@ -62,6 +63,12 @@ pipeline: auto_offset_reset: earliest topics: - "{{ SQLFLOW_TOPIC|default('tumbling-window') }}" + # The record's own time, from its payload. The handler buckets on the + # event_time column the source fills from it, and the engine decides + # each record's lateness against that same bucket. + event_time: + path: timestamp + format: rfc3339 handler: type: 'handlers.InferredMemBatch' @@ -69,7 +76,7 @@ pipeline: INSERT INTO agg_cities_count BY NAME SELECT - date_trunc('hour', CAST(timestamp as TIMESTAMPTZ)) as bucket, + time_bucket(INTERVAL '1 hour', event_time) AS bucket, properties.city as city, count(*) as count FROM batch diff --git a/docs/coverage/invariants.yml b/docs/coverage/invariants.yml index a4b66e47..fd5faf46 100644 --- a/docs/coverage/invariants.yml +++ b/docs/coverage/invariants.yml @@ -239,13 +239,26 @@ invariants: applies_to: pipeline claim: > A bucket's published value counts exactly the rows produced for it, - once. Rows the engine dropped under late_rows drop are excluded and - counted in window_late_rows_total; late_rows reemit publishes a - different contract and is out of scope. + once. Rows the engine refused as late are excluded and counted in + window_late_rows_total; a row admitted within allowed_lateness_seconds + is in the bucket's republished value. verified_by: harness requires: [] tracked_by: "#183" + - id: pipeline.window.lateness_decided_at_arrival + family: checkpoint + class: safety + applies_to: pipeline + claim: > + A record whose bucket ended at or before the watermark less + allowed_lateness_seconds is refused before the handler and counted; one + within lateness is written and its bucket is republished as a whole + value; the window table never holds a row the engine did not admit. + verified_by: harness + requires: [] + tracked_by: "the conformance harness has no windowed pipeline subject yet; internal/core, internal/managers/model.go and internal/simulate verify it" + - id: source.commit.only_processed family: checkpoint class: safety @@ -482,8 +495,9 @@ invariants: class: liveness applies_to: manager claim: > - A closed window reaches the sink without anything else happening. The - loop polls on its own, and a window that closes is published. + A closed window reaches the sink on the manager's start pass or on the + kick that follows the commit which closed it, without anything else + happening. The manager has no clock; the engine tells it. verified_by: harness requires: [] enforced: true @@ -527,37 +541,6 @@ invariants: requires: [] enforced: true - - id: manager.late.policy_holds - family: checkpoint - class: safety - applies_to: manager - claim: > - A row for a bucket below the watermark is late. Under drop it is - discarded and counted, and the bucket is never published again. Under - reemit it is published once. - verified_by: harness - requires: [] - enforced: true - - # The late-row counter is the data-loss signal the TurboStats bundle reports - # as late_rows_dropped, so it holds to what sink_rows_written holds to: - # counted when the outcome is durable, never when it is attempted. It was - # incremented before the close committed, and a close that then failed - # rolled its delete back while the count stayed, so the next close counted - # the same rows again. - - id: manager.late.counted_once - family: checkpoint - class: safety - applies_to: manager - claim: > - A late row is counted once, when the close that dropped or reemitted it - commits. A close that fails counts nothing; the close that later settles - those rows counts them. - verified_by: harness - requires: [] - enforced: true - violated_once: ["#354"] - - id: manager.drain.bounded family: lifecycle class: liveness diff --git a/docs/coverage/matrix.md b/docs/coverage/matrix.md index 1787d215..bdd1f7d2 100644 --- a/docs/coverage/matrix.md +++ b/docs/coverage/matrix.md @@ -79,7 +79,7 @@ integration behind it keeps a batch it could not deliver, or commits offsets only after a flush. Those are invariants, they are counted separately below, and the two numbers are not interchangeable. -**53 invariants declared: 45 safety and 8 liveness. Of 187 (invariant, integration) cells: 99 proven, 53 missing, 0 skipped, 0 failing, 35 exempt. 0 gap(s).** +**52 invariants declared: 44 safety and 8 liveness. Of 187 (invariant, integration) cells: 97 proven, 55 missing, 0 skipped, 0 failing, 35 exempt. 0 gap(s).** Safety says nothing bad happens. Liveness says something good eventually does, and the two are not interchangeable: a sink that @@ -146,8 +146,6 @@ drains. An invariant holds only if it holds on all four. | `manager.delete.nothing_on_failure` | A failed flush deletes nothing. Every closed window stays in the state table for the next attempt. | · | · | · | · | ✅ u | | `manager.watermark.never_regresses` | The persisted watermark never moves backwards, across polls and across a restart. A manager built over the state another one saved publishes nothing that one published. | · | · | · | · | ✅ u | | `manager.close.committed_rows_only` | A close publishes rows the pipeline has committed and no others. Rows an open batch has written are not counted, so a batch that rolls back was never published. | · | · | · | · | ✅ u | -| `manager.late.policy_holds` | A row for a bucket below the watermark is late. Under drop it is discarded and counted, and the bucket is never published again. Under reemit it is published once. | · | · | · | · | ✅ u | -| `manager.late.counted_once` | A late row is counted once, when the close that dropped or reemitted it commits. A close that fails counts nothing; the close that later settles those rows counts them. *(violated once: #354)* | · | · | · | · | ✅ u | These checkpoint invariants are properties of the consume loop rather than of anything a config file names. The columns are @@ -163,7 +161,8 @@ drains. An invariant holds only if it holds on all four. | `pipeline.shutdown.commits_only_delivered` | After the consume loop returns, clean or failed, the commits the process makes on its way out never make a position durable past the last message the sink acknowledged. Not in the state database, and not at the source. *(violated once: #279)* | ✅ u | ✅ u | | `pipeline.commit.nothing_on_failure` | A failed flush commits nothing. Not offsets, not state. | ✅ u | ✅ u | | `pipeline.state.with_offsets` | Window state and the offsets that produced it commit atomically. | ✅ u | — exempt | -| `pipeline.window.counts_every_row` | A bucket's published value counts exactly the rows produced for it, once. Rows the engine dropped under late_rows drop are excluded and counted in window_late_rows_total; late_rows reemit publishes a different contract and is out of scope. *(declared, tracked by #183)* | ❌ missing | ❌ missing | +| `pipeline.window.counts_every_row` | A bucket's published value counts exactly the rows produced for it, once. Rows the engine refused as late are excluded and counted in window_late_rows_total; a row admitted within allowed_lateness_seconds is in the bucket's republished value. *(declared, tracked by #183)* | ❌ missing | ❌ missing | +| `pipeline.window.lateness_decided_at_arrival` | A record whose bucket ended at or before the watermark less allowed_lateness_seconds is refused before the handler and counted; one within lateness is written and its bucket is republished as a whole value; the window table never holds a row the engine did not admit. *(declared, tracked by the conformance harness has no windowed pipeline subject yet; internal/core, internal/managers/model.go and internal/simulate verify it)* | ❌ missing | ❌ missing | ## Safety invariants: types @@ -223,7 +222,7 @@ drains. An invariant holds only if it holds on all four. | Invariant | Claim | `manager.watermark` | | --- | --- | --- | -| `manager.publish.eventually` | A closed window reaches the sink without anything else happening. The loop polls on its own, and a window that closes is published. | ✅ u | +| `manager.publish.eventually` | A closed window reaches the sink on the manager's start pass or on the kick that follows the commit which closed it, without anything else happening. The manager has no clock; the engine tells it. | ✅ u | | `manager.failure.exits` | A poll the sink refuses stops the manager with the sink's error, after one attempt, and the process exits with its code. The sink ran its retry ladder before the error arrived, so the manager does not retry in place, and a window the destination will not take is never collected, written and refused every tick while the process reports healthy. *(violated once: #267)* | ✅ u | | `manager.drain.bounded` | The final poll after a cancel finishes or fails inside the drain deadline. A sink that never answers cannot hold the process past it, and every closed window it did not deliver stays in the state table. | ✅ u | diff --git a/docs/coverage/status/manager.watermark.yml b/docs/coverage/status/manager.watermark.yml index 1c3bee7a..f3acdc64 100644 --- a/docs/coverage/status/manager.watermark.yml +++ b/docs/coverage/status/manager.watermark.yml @@ -5,8 +5,6 @@ manager.delete.after_flush: {unit: covered, integration: missing, release: missi manager.delete.nothing_on_failure: {unit: covered, integration: missing, release: missing} manager.drain.bounded: {unit: covered, integration: missing, release: missing} manager.failure.exits: {unit: covered, integration: missing, release: missing} -manager.late.counted_once: {unit: covered, integration: missing, release: missing} -manager.late.policy_holds: {unit: covered, integration: missing, release: missing} manager.publish.eventually: {unit: covered, integration: missing, release: missing} manager.watermark.never_regresses: {unit: covered, integration: missing, release: missing} pipeline.batch.independent_of_window_io: {unit: covered, integration: missing, release: missing} diff --git a/docs/coverage/status/pipeline.stateful.yml b/docs/coverage/status/pipeline.stateful.yml index 5eefb311..db4cfff1 100644 --- a/docs/coverage/status/pipeline.stateful.yml +++ b/docs/coverage/status/pipeline.stateful.yml @@ -13,4 +13,5 @@ pipeline.progress.no_silent_stall: {unit: missing, integration: missing, release pipeline.shutdown.commits_only_delivered: {unit: covered, integration: missing, release: missing} pipeline.state.with_offsets: {unit: covered, integration: missing, release: missing} pipeline.window.counts_every_row: {unit: missing, integration: missing, release: missing} +pipeline.window.lateness_decided_at_arrival: {unit: missing, integration: missing, release: missing} pipeline.writers.merge_exactly: {unit: missing, integration: missing, release: missing} diff --git a/docs/coverage/status/pipeline.stateless.yml b/docs/coverage/status/pipeline.stateless.yml index 9467d653..ec24888f 100644 --- a/docs/coverage/status/pipeline.stateless.yml +++ b/docs/coverage/status/pipeline.stateless.yml @@ -13,4 +13,5 @@ pipeline.progress.no_silent_stall: {unit: missing, integration: missing, release pipeline.shutdown.commits_only_delivered: {unit: covered, integration: missing, release: missing} pipeline.state.with_offsets: {unit: exempt, integration: exempt, release: exempt} pipeline.window.counts_every_row: {unit: missing, integration: missing, release: missing} +pipeline.window.lateness_decided_at_arrival: {unit: missing, integration: missing, release: missing} pipeline.writers.merge_exactly: {unit: skipped, integration: covered, release: missing} diff --git a/docs/superpowers/plans/2026-09-26-watermark-driven-close.md b/docs/superpowers/plans/2026-09-26-watermark-driven-close.md index 3d00df4b..970f7bea 100644 --- a/docs/superpowers/plans/2026-09-26-watermark-driven-close.md +++ b/docs/superpowers/plans/2026-09-26-watermark-driven-close.md @@ -515,6 +515,8 @@ For each file, remove the `late_rows:` and `poll_interval_seconds:` lines, remov | `dev/bench/bluesky/demo-10x*.yml` (three) | `time_bucket(INTERVAL '1 minute', event_time) AS bucket` | `event_time: {path: time_us, format: unix_us}`; delete the `poll_interval_seconds 10 -> 1` comment lines | | `.claude/skills/slow-soak/slow.yml` | as bluesky | as bluesky | +Ruling at execution: `render/pipeline.yml` and `render/docker-compose.yml` are **not** updated in this PR. The render workflow validates the template against the image it pins (`v2026.09.21`), whose schema still requires `late_rows`, so the template moves in the PR that bumps the pin to a release carrying this change. That PR also decides the template's bucket clock: its payload timestamp is optional and per metric, so the candidate is the request's arrival `event_time`, with the metric's own timestamp kept in `last_at`. + For the bluesky postgres example, add `allowed_lateness_seconds: 300` with a comment: the sink upserts by `(bucket, lang)`, so a late row republishes the minute whole and the row is replaced -- this is the one shipped example that exercises lateness. - [ ] **Step 5: Update the README** diff --git a/docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md b/docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md index e29ab767..108d1c61 100644 --- a/docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md +++ b/docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md @@ -1,6 +1,6 @@ # The close is driven by the watermark, and lateness is decided at arrival -**Status:** proposal, approved in discussion on 2026-09-26. +**Status:** implemented in the PR that carries this plan. **Builds on:** #393 (the engine asserts the watermark), #385 (`event_time` exposed to handler SQL), #389 (every source can be told where its event time is). **Removes:** `poll_interval_seconds`, `late_rows`, the manager's poll loop, and the manager's late-row sweep. **Adds:** `allowed_lateness_seconds`. diff --git a/docs/windows/decisions.md b/docs/windows/decisions.md index 8136bba5..f450d67c 100644 --- a/docs/windows/decisions.md +++ b/docs/windows/decisions.md @@ -1,7 +1,7 @@ # Window decisions Rendered from the truth tables in `internal/managers/decide.go` by a test, so -this page is what runs. A poll reduces what it read to one value per fact, +this page is what runs. A pass reduces what it read to one value per fact, looks the combination up, and performs the action of the one row it selects. The check that runs when the package loads proves each combination selects exactly one row and every row is reachable. @@ -12,10 +12,13 @@ ever write for the window has `time_column` at or after the watermark. How it is computed -- the newest event time seen per source partition, less the grace, combined by minimum over the partitions that could still deliver -- is `internal/core/watermarks.go`. The manager reads that one value and -nothing else: no clock, no progress row, no reading of the table's newest -bucket. See `docs/superpowers/specs/2026-09-24-window-watermark-design.md`. +nothing else: no clock, no ticker, no progress row, no reading of the +table's newest bucket. It runs a pass when it starts, when the engine kicks +it after a commit that moved the watermark or admitted a late row, and when +it drains. See `docs/superpowers/specs/2026-09-24-window-watermark-design.md` +and `docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md`. -## The watermark, once per poll +## The watermark, once per pass | Fact | Values | Computed from | |---|---|---| @@ -29,29 +32,38 @@ bucket. See `docs/superpowers/specs/2026-09-24-window-watermark-design.md`. | `hold.behind` | `behind` | `hold` | unchanged | asserted | The assertion is at or before what this window has already closed, so it has been acted on; the closed watermark never moves backwards. | | `follow` | `ahead` | `follow` | the assertion | asserted | The engine has promised that every row it will ever write ends at or after the asserted watermark, so every bucket ending at or before it is complete: the closed watermark follows it, and those buckets publish. | -## Each bucket, given the poll's watermark +## Each bucket, given the pass's watermark | Fact | Values | Computed from | |---|---|---| -| `bucket` | `late`, `due`, `open` | The bucket's end against the previous watermark and the one this poll decided. `late`: at or before the previous, so its rows arrived after it closed. `due`: after the previous and at or before the next. `open`: after the next. | -| `policy` | `drop`, `reemit` | `late_rows` in the window's config. | +| `bucket` | `open`, `due`, `retained`, `expired` | The bucket's end against the closed watermark, the asserted one, and `allowed_lateness_seconds`. `open`: ends after the asserted watermark. `due`: ends after the closed watermark and at or before the asserted one, so it closes in this pass. `retained`: closed in an earlier pass and its end plus the lateness is still past the assertion. `expired`: closed in an earlier pass and its end plus the lateness is at or before the assertion. | -6 combinations, 4 rows. +4 combinations, 4 rows. -| Rule | bucket | policy | Action | Deciding | Claim | -|---|---|---|---|---|---| -| `keep` | `open` | | `keep` | bucket | The bucket ends after the watermark, so its rows stay. | -| `close` | `due` | | `close` | bucket | The bucket ends between the previous watermark and this one: emit_sql runs over its rows and they are deleted. | -| `late.drop` | `late` | `drop` | `late.drop` | policy | The bucket closed before these rows arrived, and late_rows is drop: they are deleted and counted. | -| `late.reemit` | `late` | `reemit` | `late.reemit` | policy | The bucket closed before these rows arrived, and late_rows is reemit: emit_sql runs over the late rows alone, they are deleted and counted. | +| Rule | bucket | Action | Deciding | Claim | +|---|---|---|---|---| +| `keep` | `open` | `keep` | bucket | The bucket ends after the asserted watermark, so its rows stay. | +| `publish` | `due` | `publish` | bucket | The bucket ends between the closed watermark and the asserted one: emit_sql runs over its rows and the result is published. With allowed_lateness_seconds the rows stay for a late row to republish it whole; without, this pass purges them too. | +| `retain` | `retained` | `retain` | bucket | The bucket has been published and its lateness has not run out: its rows stay, and a late row the engine admits republishes it whole. | +| `purge` | `expired` | `purge` | bucket | The bucket ended at or before the asserted watermark less the lateness: nothing more can arrive for it, and its rows are deleted. | + +The engine decides a record's lateness at arrival by the same comparison, +before the handler sees it: a record whose bucket ended at or before the +asserted watermark less the lateness is refused and counted; one whose +bucket ended at or before the watermark but within the lateness is written, +and the bucket is republished whole on the next pass; anything else is on +time. So the window table never holds a row the engine did not admit, and +the manager never sweeps. ## What closes a bucket, by configuration Three keys decide. `grace_seconds` is how far the stream's own clock must pass a bucket's end before it closes (Flink's bounded out-of-orderness); `idle_close_seconds` is how long a partition may be silent before it stops -holding the window open (Flink's idleness); `late_rows` is what happens to a -row for a bucket that already closed. +holding the window open (Flink's idleness); `allowed_lateness_seconds` is how +long after a bucket closes a record for it is still admitted, and its rows +kept, so that the bucket can be republished whole (Flink's allowed +lateness). | | `grace_seconds` | `idle_close_seconds` | the stream this is for | what closes a bucket | the failure mode to know | |---|---|---|---|---|---| @@ -60,9 +72,9 @@ row for a bucket that already closed. | **C** | 0 | set | ordered, intermittent | event time, or every partition silent for the bound | a bound shorter than the gap *within* a burst closes mid-burst, and the rest of the burst is late | | **D** | >0 | set | bursty and reordered: the IoT default | either | both of the above | -Each crossed with `drop` (a late row is discarded and counted) or `reemit` -(a late row is republished on its own, which a replacing sink turns into -loss). +Each crossed with `allowed_lateness_seconds`: 0 refuses a late record before +the handler and counts it; a positive value admits it and republishes its +bucket as a whole value, so the sink must replace by key. ## Invariants @@ -73,4 +85,5 @@ Each is a claim with a check in the tests. | `window.never_backwards` | No reading moves the closed watermark behind the committed one, checked over ten thousand random readings here; and the engine's assertion is `max(stored, W)` by construction, checked in `internal/core`. | | `window.asserted_in_the_commit` | The engine writes the watermark in the transaction that commits the rows it describes, so no reader can see rows without the watermark that accounts for them, or a watermark without its rows. A commit that fails leaves both where they were. Checked in `internal/core`. | | `window.minimum_over_partitions` | The watermark is the minimum over the partitions that could still deliver: one that has not delivered holds it at -inf, a lagging one holds it, an idle one leaves it, a lost one holds at its last position, a revoked one is gone. Checked in `internal/core` and by the simulator. | -| `window.no_clock_in_the_manager` | The manager has no clock. Every close is `bucket end <= asserted watermark`, in event time; the engine's clock decides only which partitions are in its minimum. Checked by construction: the manager takes no clock. | +| `window.no_clock_in_the_manager` | The manager has no clock. Every close is `bucket end <= asserted watermark`, in event time; the engine's clock decides only which partitions are in its minimum. Checked by construction: the manager takes no clock and no interval. | +| `window.lateness_decided_at_arrival` | A record whose bucket ended at or before the watermark less the lateness is refused before the handler and counted; one within the lateness is written and its bucket republished whole. The window table never holds a row the engine did not admit. Checked in `internal/core`, by the model and by the simulator. | diff --git a/internal/cli/run/idle_close_test.go b/internal/cli/run/idle_close_test.go index 9a552753..7193a36f 100644 --- a/internal/cli/run/idle_close_test.go +++ b/internal/cli/run/idle_close_test.go @@ -100,7 +100,7 @@ func (f *failingAfter) Record(ctx context.Context, p core.Progress) error { // seeded, and the stream exists to keep the partition out of idleness. type idleCloseRig struct { published func() int64 - poll func() error + pass func() error turbine *core.Turbine stop func() } @@ -123,8 +123,6 @@ func newIdleCloseRig(t *testing.T, src *tickingSource, wrap func(core.ProgressSa SizeSeconds: 60, GraceSeconds: 60, IdleCloseSeconds: 1, - LateRows: "drop", - PollIntervalSecs: 3600, EmitSQL: "SELECT bucket, sum(n)::BIGINT AS n FROM closed GROUP BY ALL", Sink: config.Sink{Type: "sqlcommand", SQLCommand: &config.SQLCommandSink{ SQL: "INSERT INTO published SELECT bucket, n FROM sqlflow_sink_batch", @@ -136,7 +134,12 @@ func newIdleCloseRig(t *testing.T, src *tickingSource, wrap func(core.ProgressSa assert.NoError(t, store.Init(ctx)) assert.NoError(t, initWindowStores(ctx, conf, conn)) - managed, closeConns, err := buildManagedTables(ctx, conf, db, zap.NewNop(), + // The watermark tracker, wired and restored the way run does it. The + // pipeline has seen no rows of its own, so the close it makes when the + // stream stops is measured from what the table holds. The manager gets + // the tracker's signal, as run wires it; the rig drives Pass by hand. + watermarks, windowOpts := windowOptions(conf, conn) + managed, closeConns, err := buildManagedTables(ctx, conf, db, watermarks, zap.NewNop(), sdkmetric.NewMeterProvider(), nil, sinks.RetryEvents{}) assert.NoError(t, err) assert.Equal(t, 1, len(managed)) @@ -153,10 +156,6 @@ func newIdleCloseRig(t *testing.T, src *tickingSource, wrap func(core.ProgressSa // mode, manufactured by the rig rather than by the engine. lock := &sync.Mutex{} - // The watermark tracker, wired and restored the way run does it. The - // pipeline has seen no rows of its own, so the close it makes when the - // stream stops is measured from what the table holds. - watermarks, windowOpts := windowOptions(conf, conn) lock.Lock() assert.NoError(t, restoreWindows(ctx, conf, conn, watermarks)) lock.Unlock() @@ -189,7 +188,7 @@ func newIdleCloseRig(t *testing.T, src *tickingSource, wrap func(core.ProgressSa defer lock.Unlock() return sharedConnCount(t, conn, "published") }, - poll: func() error { return managed[0].Poll(ctx) }, + pass: func() error { return managed[0].Pass(ctx) }, turbine: tb, stop: stop, } @@ -209,7 +208,7 @@ func TestManagerWindow_AQuietStreamClosesOnTheEnginesAssertion(t *testing.T) { t.Fatal("the stream has been quiet for far longer than idle_close_seconds and the bucket " + "never closed: the engine's idle ticks are not confirming the quiet") } - assert.NoError(t, rig.poll()) + assert.NoError(t, rig.pass()) time.Sleep(25 * time.Millisecond) } assert.Equal(t, int64(1), rig.published()) @@ -234,7 +233,7 @@ func TestManagerWindow_ALiveStreamNeverClosesOnIdleness(t *testing.T) { // Three times the idle bound, polling all the way through. for end := time.Now().Add(3 * time.Second); time.Now().Before(end); { - assert.NoError(t, rig.poll()) + assert.NoError(t, rig.pass()) if n := rig.published(); n != 0 { t.Fatalf("a live stream had its open bucket closed on idleness (%d rows published): "+ "a partition that keeps delivering must never leave the minimum", n) diff --git a/internal/cli/run/managers.go b/internal/cli/run/managers.go index b7549838..b28c4126 100644 --- a/internal/cli/run/managers.go +++ b/internal/cli/run/managers.go @@ -27,7 +27,7 @@ func windowDeclaration(table config.TableSQL) managers.Declaration { Size: time.Duration(w.SizeSeconds) * time.Second, Grace: time.Duration(w.GraceSeconds) * time.Second, IdleClose: time.Duration(w.IdleCloseSeconds) * time.Second, - Late: managers.LatePolicy(w.LateRows), + Lateness: w.Lateness(), EmitSQL: w.EmitSQL, } } @@ -45,7 +45,7 @@ func windowSpecs(conf *config.Conf) []core.WindowSpec { } d := windowDeclaration(table) specs = append(specs, core.WindowSpec{ - Name: d.Table, Size: d.Size, Grace: d.Grace, IdleClose: d.IdleClose, + Name: d.Table, Size: d.Size, Grace: d.Grace, IdleClose: d.IdleClose, Lateness: d.Lateness, }) } return specs @@ -119,8 +119,18 @@ func initWindowStores(ctx context.Context, conf *config.Conf, conn adbc.Connecti return core.NewWatermarkStore(conn).Init(ctx) } +// signalFor is the window's signal from the tracker, or nil without one: a +// manager built without a signal runs its start pass and its drain pass. +func signalFor(w *core.Watermarks, name string) *core.WindowSignal { + if w == nil { + return nil + } + return w.Signal(name) +} + // buildManagedTables constructs a watermark manager per table that declares -// a window. Each manager gets two connections of its own to the pipeline's +// a window. Each gets the window's signal from the engine's tracker, which +// is how it learns the watermark moved. Each manager gets two connections of its own to the pipeline's // DuckDB: one with autocommit off for the close, so it reads committed rows // only and its delete commits with its watermark, and one for its sink, // because a transaction may write to one database only and a sink stages a @@ -130,6 +140,7 @@ func buildManagedTables( ctx context.Context, conf *config.Conf, db *duckdb.DB, + watermarks *core.Watermarks, l *zap.Logger, mp metric.MeterProvider, budget *core.DrainBudget, @@ -157,9 +168,9 @@ func buildManagedTables( if table.Window == nil { continue } - if table.Window.ReemitOverwrites() { + if table.Window.LatenessNeedsReplacingSink() { return nil, closeConns, errs.New(errs.CodeConfigInvalid, - "table %q window: %s", table.Name, config.ReemitOverwritesMessage) + "table %q window: %s", table.Name, config.LatenessNeedsReplacingSinkMessage(*table.Window)) } conn, err := db.Connect(ctx) @@ -201,8 +212,7 @@ func buildManagedTables( return nil, closeConns, fmt.Errorf("table %q window sink: %w", table.Name, err) } - m, err := managers.NewWatermark(conn, windowDeclaration(table), - time.Duration(table.Window.PollIntervalSecs)*time.Second, sink, + m, err := managers.NewWatermark(conn, windowDeclaration(table), sink, signalFor(watermarks, table.Name), managers.WithLogger(l), managers.WithDrainBudget(budget), managers.WithMeterProvider(mp)) diff --git a/internal/cli/run/managers_test.go b/internal/cli/run/managers_test.go index 2683b256..a6e679b4 100644 --- a/internal/cli/run/managers_test.go +++ b/internal/cli/run/managers_test.go @@ -140,8 +140,8 @@ func newTestWindow(t *testing.T, db *duckdb.DB, sink core.Sink, opts ...managers assert.NoError(t, po.SetOption(adbc.OptionKeyAutoCommit, adbc.OptionValueDisabled)) m, err := managers.NewWatermark(conn, managers.Declaration{ - Table: "agg", TimeColumn: "bucket", Size: time.Minute, Late: managers.LateDrop, - }, time.Hour, sink, opts...) + Table: "agg", TimeColumn: "bucket", Size: time.Minute, + }, sink, nil, opts...) assert.NoError(t, err) return m } diff --git a/internal/cli/run/metrics_test.go b/internal/cli/run/metrics_test.go index 0a1252f0..e6528a7c 100644 --- a/internal/cli/run/metrics_test.go +++ b/internal/cli/run/metrics_test.go @@ -146,6 +146,8 @@ func exportedNames(t *testing.T) []string { m.PipelineRowsAccepted.Add(ctx, 1) m.PipelineRowsWritten.Add(ctx, 1) m.PipelineLastMessage.Record(ctx, 1) + // Late rows are the engine's counter: it decides lateness at arrival. + m.WindowLateRows.Add(ctx, 1) wm, err := webhook.NewMetrics(mp) assert.NoError(t, err) @@ -157,7 +159,7 @@ func exportedNames(t *testing.T) []string { win := managers.NewWindowMetrics(mp, "w") win.Watermark.Record(ctx, 1) win.Closed.Add(ctx, 1) - win.Late.Add(ctx, 1) + win.Recomputed.Add(ctx, 1) win.NewestStart.Record(ctx, 1) win.CloseLag.Record(ctx, 1) @@ -236,6 +238,7 @@ func TestExportedSeriesNames(t *testing.T) { "window_closed_total", "window_late_rows_total", "window_newest_bucket_start_seconds", + "window_recomputes_total", "window_watermark_seconds", } assert.DeepEqual(t, want, exportedNames(t)) diff --git a/internal/cli/run/progress_wiring_test.go b/internal/cli/run/progress_wiring_test.go index 9a664909..8a32f81c 100644 --- a/internal/cli/run/progress_wiring_test.go +++ b/internal/cli/run/progress_wiring_test.go @@ -37,7 +37,7 @@ func windowedConf(idleCloseSeconds int) *config.Conf { Name: "t", Window: &config.Window{ TimeColumn: "bucket", SizeSeconds: 60, GraceSeconds: 5, - IdleCloseSeconds: idleCloseSeconds, LateRows: "drop", + IdleCloseSeconds: idleCloseSeconds, }, }}}} } diff --git a/internal/cli/run/root.go b/internal/cli/run/root.go index 7d323bb6..d770b829 100644 --- a/internal/cli/run/root.go +++ b/internal/cli/run/root.go @@ -592,7 +592,7 @@ func NewCommand() *cobra.Command { ) liveTurbine.Store(turbine) - managedTables, closeWindowConns, err := buildManagedTables(ctx, conf, db, l, + managedTables, closeWindowConns, err := buildManagedTables(ctx, conf, db, watermarks, l, meterProvider, budget, retryEvents) if err != nil { return err diff --git a/internal/cli/run/row_roles_test.go b/internal/cli/run/row_roles_test.go index 89731413..b4b3434d 100644 --- a/internal/cli/run/row_roles_test.go +++ b/internal/cli/run/row_roles_test.go @@ -164,25 +164,23 @@ func TestWindowManagerRowsAreCounted(t *testing.T) { SQL: []config.TableSQL{{ Name: "agg", Window: &config.Window{ - TimeColumn: "bucket", - SizeSeconds: 60, - LateRows: "drop", - PollIntervalSecs: 3600, - Sink: config.Sink{Type: "console"}, + TimeColumn: "bucket", + SizeSeconds: 60, + Sink: config.Sink{Type: "console"}, }, }}, }, } built, closeConns, err := buildManagedTables( - context.Background(), conf, db, zap.NewNop(), mp, nil, sinks.RetryEvents{}) + context.Background(), conf, db, nil, zap.NewNop(), mp, nil, sinks.RetryEvents{}) assert.NoError(t, err) defer closeConns() assert.Equal(t, 1, len(built)) - // One poll closes the bucket, writes the rows to the manager's sink and + // One pass closes the bucket, writes the rows to the manager's sink and // flushes them. - assert.NoError(t, built[0].Poll(context.Background())) + assert.NoError(t, built[0].Pass(context.Background())) byRole := writtenRowsByRole(t, reader) assert.Equal(t, int64(3), byRole["manager"]) diff --git a/internal/cli/run/shared_conn_test.go b/internal/cli/run/shared_conn_test.go index 2ecb05d3..c87ad2f5 100644 --- a/internal/cli/run/shared_conn_test.go +++ b/internal/cli/run/shared_conn_test.go @@ -95,11 +95,9 @@ func TestPipelineSharedConnection_EveryPartyHoldsTheLock(t *testing.T) { // One-second buckets with no grace: a bucket closes once a // later one exists. Window: &config.Window{ - TimeColumn: "bucket", - SizeSeconds: 1, - LateRows: "drop", - PollIntervalSecs: 3600, - EmitSQL: "SELECT bucket, sum(n)::BIGINT AS n FROM closed GROUP BY ALL", + TimeColumn: "bucket", + SizeSeconds: 1, + EmitSQL: "SELECT bucket, sum(n)::BIGINT AS n FROM closed GROUP BY ALL", Sink: config.Sink{Type: "sqlcommand", SQLCommand: &config.SQLCommandSink{ SQL: "INSERT INTO published SELECT bucket, n FROM sqlflow_sink_batch", }}, @@ -136,7 +134,7 @@ func TestPipelineSharedConnection_EveryPartyHoldsTheLock(t *testing.T) { assert.NoError(t, err) handler, err := handlers.New(conn, conf.Pipeline.Handler, zap.NewNop()) assert.NoError(t, err) - managed, closeConns, err := buildManagedTables(ctx, conf, db, zap.NewNop(), mp, + managed, closeConns, err := buildManagedTables(ctx, conf, db, nil, zap.NewNop(), mp, nil, sinks.RetryEvents{}) assert.NoError(t, err) defer closeConns() @@ -163,7 +161,7 @@ func TestPipelineSharedConnection_EveryPartyHoldsTheLock(t *testing.T) { _, err = tb.ConsumeLoop(ctx, 6) assert.NoError(t, err) - assert.NoError(t, managed[0].Poll(ctx)) + assert.NoError(t, managed[0].Pass(ctx)) // Every party ran at least once, or a clean result proves nothing. The // window closed the two buckets a later one exists for, and the newest diff --git a/internal/cli/run/window_refusal_test.go b/internal/cli/run/window_refusal_test.go index 85d24e61..48c7e02f 100644 --- a/internal/cli/run/window_refusal_test.go +++ b/internal/cli/run/window_refusal_test.go @@ -14,26 +14,46 @@ import ( ) // A config that skipped validate must not run the pairing validate refuses: -// a postgres sink that upserts by key, handed a reemit of the late rows -// alone, replaces a published count with theirs. run refuses it before it -// dials anything, with the exit code a config error gets. -func TestWindowUpsertWithReemitIsRefusedAtStartup(t *testing.T) { +// lateness above zero with a sink that appends, which holds a republished +// bucket twice. run refuses it before it dials anything, with the exit code +// a config error gets. +func TestWindowLatenessWithAnAppendingSinkIsRefusedAtStartup(t *testing.T) { coverage.Covers(t, "manager.window") db, _ := rowsTestDB(t) conf := &config.Conf{Tables: &config.Tables{SQL: []config.TableSQL{{ Name: "agg", Window: &config.Window{ - TimeColumn: "bucket", SizeSeconds: 60, LateRows: "reemit", + TimeColumn: "bucket", SizeSeconds: 60, AllowedLatenessSecs: 300, Sink: config.Sink{Type: "postgres", Postgres: &config.PostgresSink{ - DSN: "postgres://u:p@127.0.0.1:1/db", Table: "agg", Mode: "upsert", Key: []string{"bucket"}, + DSN: "postgres://u:p@127.0.0.1:1/db", Table: "agg", Mode: "append", }}, }, }}}} - _, closeConns, err := buildManagedTables(context.Background(), conf, db, zap.NewNop(), nil, nil, sinks.RetryEvents{}) + _, closeConns, err := buildManagedTables(context.Background(), conf, db, nil, zap.NewNop(), nil, nil, sinks.RetryEvents{}) closeConns() assert.Error(t, err) assert.Equal(t, errs.CodeConfigInvalid, errs.CodeOf(err)) assert.Equal(t, errs.ExitUserError, errs.ExitCode(err)) - assert.That(t, strings.Contains(err.Error(), "late_rows is reemit and the postgres sink upserts")) + assert.That(t, strings.Contains(err.Error(), "allowed_lateness_seconds")) + assert.That(t, strings.Contains(err.Error(), "postgres sink appends")) +} + +// The same lateness with a sink that replaces by key is accepted: a +// republished bucket overwrites its earlier value there. +func TestWindowLatenessWithAReplacingSinkIsAccepted(t *testing.T) { + coverage.Covers(t, "manager.window") + db, _ := rowsTestDB(t) + conf := &config.Conf{Tables: &config.Tables{SQL: []config.TableSQL{{ + Name: "agg", + Window: &config.Window{ + TimeColumn: "bucket", SizeSeconds: 60, AllowedLatenessSecs: 300, + Sink: config.Sink{Type: "console"}, + }, + }}}} + + built, closeConns, err := buildManagedTables(context.Background(), conf, db, nil, zap.NewNop(), nil, nil, sinks.RetryEvents{}) + defer closeConns() + assert.NoError(t, err) + assert.Equal(t, 1, len(built)) } diff --git a/internal/cli/testdata/config_example.golden b/internal/cli/testdata/config_example.golden index abb223c3..c8760c68 100644 --- a/internal/cli/testdata/config_example.golden +++ b/internal/cli/testdata/config_example.golden @@ -350,14 +350,14 @@ tables: # one flush_interval_seconds; validate warns when that interval is the # longer of the two. idle_close_seconds: - # What happens to a row for a bucket that already closed. drop discards - # it and counts it. reemit publishes emit_sql over the late rows alone, - # for a sink that adds them to the bucket it holds; a sink that replaces - # the bucket's value loses the rows published before. - # Required: the two are different promises to the sink. - late_rows: drop | reemit - # How often the engine looks for closed buckets. Absent means 10. - poll_interval_seconds: + # How long after a bucket closes its rows are kept and late rows for it + # are still accepted. A late row within this republishes the bucket as a + # whole value, so the sink must replace by key; validate refuses a sink + # that appends. Absent or 0: a row for a closed bucket is refused before + # the handler and counted in window_late_rows_total; the window table + # never holds it. Flink's allowedLateness. Measured in event time, + # against the watermark, like everything a window decides. + allowed_lateness_seconds: # Shapes the closed rows before the sink. It reads one relation, closed, # which holds every row of every bucket that just closed. Absent means # SELECT * FROM closed. diff --git a/internal/config/config.go b/internal/config/config.go index e4b84ac3..32edef3a 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -1,8 +1,10 @@ package config import ( + "fmt" "math" "net/url" + "time" "github.com/turbolytics/sql-flow/internal/errs" "github.com/turbolytics/sql-flow/internal/eventtime" @@ -154,10 +156,11 @@ type SinkRetry struct { } // Window is a tumbling window the engine closes for the user. The handler -// writes rows keyed by a bucket start into the table; the engine keeps an +// writes rows keyed by a bucket start into the table; the engine asserts an // event-time watermark, publishes every bucket the watermark has passed to -// the window's sink, and deletes it. Nothing here is SQL the user has to get -// right twice. +// the window's sink, keeps it for allowed_lateness_seconds so a late row can +// republish it whole, and deletes it. Nothing here is SQL the user has to +// get right twice, and nothing here is a clock. type Window struct { // The column holding each row's bucket start. TIMESTAMPTZ. TimeColumn string `yaml:"time_column"` @@ -175,14 +178,14 @@ type Window struct { // one flush_interval_seconds; validate warns when that interval is the // longer of the two. IdleCloseSeconds int `yaml:"idle_close_seconds,omitempty" jsonschema:"minimum=0"` - // What happens to a row for a bucket that already closed. drop discards - // it and counts it. reemit publishes emit_sql over the late rows alone, - // for a sink that adds them to the bucket it holds; a sink that replaces - // the bucket's value loses the rows published before. - // Required: the two are different promises to the sink. - LateRows string `yaml:"late_rows" jsonschema:"enum=drop,enum=reemit"` - // How often the engine looks for closed buckets. Absent means 10. - PollIntervalSecs int `yaml:"poll_interval_seconds,omitempty" jsonschema:"minimum=1"` + // How long after a bucket closes its rows are kept and late rows for it + // are still accepted. A late row within this republishes the bucket as a + // whole value, so the sink must replace by key; validate refuses a sink + // that appends. Absent or 0: a row for a closed bucket is refused before + // the handler and counted in window_late_rows_total; the window table + // never holds it. Flink's allowedLateness. Measured in event time, + // against the watermark, like everything a window decides. + AllowedLatenessSecs int `yaml:"allowed_lateness_seconds,omitempty" jsonschema:"minimum=0"` // Shapes the closed rows before the sink. It reads one relation, closed, // which holds every row of every bucket that just closed. Absent means // SELECT * FROM closed. @@ -191,20 +194,37 @@ type Window struct { Sink Sink `yaml:"sink"` } -// ReemitOverwrites reports the one pairing of late_rows and sink no pipeline -// may run: a postgres sink that upserts by key, handed a reemit. A reemit -// publishes emit_sql over the late rows alone, because the bucket's other -// rows were deleted when it closed, and the upsert replaces the bucket's -// published row with that. validate refuses it, and so does run, for a config -// that never went through validate. -func (w Window) ReemitOverwrites() bool { - return w.LateRows == "reemit" && w.Sink.Type == "postgres" && - w.Sink.Postgres != nil && w.Sink.Postgres.Mode == "upsert" +// Lateness is allowed_lateness_seconds as a duration; zero means late rows +// are refused. +func (w Window) Lateness() time.Duration { + return time.Duration(w.AllowedLatenessSecs) * time.Second } -// ReemitOverwritesMessage says why, for validate and run alike. -const ReemitOverwritesMessage = "late_rows is reemit and the postgres sink upserts. A reemit publishes " + - "emit_sql over the late rows alone, and the sink replaces the bucket's row with that. Use drop" +// LatenessNeedsReplacingSink reports the pairing no pipeline may run: +// lateness above zero with a sink known to append. A late row republishes +// its bucket as a whole value, which an appending sink then holds twice. +// validate refuses it, and so does run, for a config that never went +// through validate. A sink validate cannot classify -- sqlcommand, +// clickhouse -- is the user's call and is warned, not refused. +func (w Window) LatenessNeedsReplacingSink() bool { + if w.AllowedLatenessSecs <= 0 { + return false + } + switch w.Sink.Type { + case "iceberg", "kafka": + return true + case "postgres": + return w.Sink.Postgres != nil && w.Sink.Postgres.Mode == "append" + } + return false +} + +// LatenessNeedsReplacingSinkMessage says why, for validate and run alike. +func LatenessNeedsReplacingSinkMessage(w Window) string { + return fmt.Sprintf("allowed_lateness_seconds is %d and the %s sink appends. A late row republishes "+ + "its bucket as a whole value, which an appending sink holds twice. Use a sink that replaces by "+ + "key, or set allowed_lateness_seconds to 0 so late rows are refused", w.AllowedLatenessSecs, w.Sink.Type) +} // SQL Tables type TableSQL struct { diff --git a/internal/conformance/manager.go b/internal/conformance/manager.go index 50f5514b..661ca256 100644 --- a/internal/conformance/manager.go +++ b/internal/conformance/manager.go @@ -2,15 +2,19 @@ package conformance // The manager half of the harness. // -// A watermark manager is a second loop that reaches a sink: it collects the -// buckets its watermark has passed, flushes them, deletes them, and advances -// the watermark. The delete and the watermark are its commit. So its claims -// are the consume loop's, restated for that commit, plus what a poll loop -// owes on its own: a closed bucket leaves, a poll that cannot deliver stops -// the process, and a drain ends. And three the watermark adds: it never -// moves backwards, it publishes committed rows only, and the late-row -// policy holds. And one the pipeline is owed: a window's I/O, however long, -// never fails a batch. +// A watermark manager is a second loop that reaches a sink: on the engine's +// signal it collects the buckets the asserted watermark has passed, flushes +// them, deletes the ones past their lateness, and advances the closed +// watermark. The delete and the watermark are its commit. So its claims are +// the consume loop's, restated for that commit, plus what a loop owes on its +// own: a closed bucket leaves, a pass that cannot deliver stops the process, +// and a drain ends. And two the watermark adds: it never moves backwards, +// and it publishes committed rows only. And one the pipeline is owed: a +// window's I/O, however long, never fails a batch. +// +// The manager has no clock. It passes on start, on each kick from the +// engine, and on the drain; the harness drives Pass by hand where a check +// needs one, and Start where the check is about the loop. // // The failure-exits claim is #267. The manager logged a failed poll and // polled again, so a bucket the destination rejected was collected, written @@ -27,16 +31,13 @@ import ( "github.com/turbolytics/sql-flow/internal/core" "github.com/turbolytics/sql-flow/internal/coverage" "github.com/turbolytics/sql-flow/internal/errs" - "go.opentelemetry.io/otel/metric" "go.opentelemetry.io/otel/metric/noop" - sdkmetric "go.opentelemetry.io/otel/sdk/metric" - "go.opentelemetry.io/otel/sdk/metric/metricdata" ) -// Manager is what a subject builds: the poll loop, and one poll of it. +// Manager is what a subject builds: the loop, and one pass of it. type Manager interface { Start(ctx context.Context) error - Poll(ctx context.Context) error + Pass(ctx context.Context) error } // ManagerSubject is one manager the harness can drive. @@ -44,23 +45,29 @@ type ManagerSubject struct { // Integration is the registry id, e.g. "manager.watermark". Integration string - // New builds the manager around the sink the harness supplies, polling - // at the given interval, with its final poll bounded by the budget and - // the given late-row policy, "drop" or "reemit". The harness owns the - // sink so it can fail a flush and record what the manager asked of it, - // and owns the budget so it can run one out. Every call builds over the - // same persisted state, which is how the watermark's memory is checked. - New func(t *testing.T, sink core.Sink, poll time.Duration, budget *core.DrainBudget, late string) Manager - - // Seed replaces the table's contents with n closed buckets, the stream - // idle for longer than the window's idle bound, so every one of them - // closes on the next poll. + // New builds the manager around the sink the harness supplies, with its + // final pass bounded by the budget and no signal: the harness drives + // Pass itself. The harness owns the sink so it can fail a flush and + // record what the manager asked of it, and owns the budget so it can run + // one out. Every call builds over the same persisted state, which is how + // the watermark's memory is checked. + New func(t *testing.T, sink core.Sink, budget *core.DrainBudget) Manager + + // Seed replaces the table's contents with n buckets, the engine having + // asserted past every one of them, so every one closes on the next pass. Seed func(t *testing.T, n int) - // SeedLate adds n rows to a bucket below the watermark, which a manager - // has already closed. Called after a poll that closed what Seed wrote. + // SeedLate adds n rows to a bucket a manager has already closed: what a + // window with lateness retains, written by hand. Called after a pass that + // closed what Seed wrote. The harness uses them to see that a manager's + // memory, not the table's oldest row, decides what is published. SeedLate func(t *testing.T, n int) + // SeedNewer adds a bucket newer than every one a manager has closed, the + // engine asserted past it, so the next pass publishes it and updates the + // window's row. + SeedNewer func(t *testing.T) + // Remaining returns how many rows the table still holds. Remaining func(t *testing.T) int64 @@ -76,26 +83,10 @@ type ManagerSubject struct { // the window-I/O claim. Batch func(t *testing.T) error - // HoldCommit builds like New, polling only when asked, over a connection - // whose commit calls hold first. By then the close has written its - // delete and its watermark and committed neither. - HoldCommit func(t *testing.T, sink core.Sink, budget *core.DrainBudget, late string, hold func()) Manager - - // Metered builds like New, polling only when asked, recording its - // metrics on the given provider so the harness can read what the manager - // counted. Optional, with SeedNewer and LateInstrument: a subject without - // all three skips the counted-once claim. - Metered func(t *testing.T, sink core.Sink, budget *core.DrainBudget, late string, mp metric.MeterProvider) Manager - - // SeedNewer adds a bucket newer than every one a manager has closed, - // which the next poll closes and publishes. Beside late rows it makes a - // close that has already counted them reach the sink, so a sink that - // refuses the flush rolls the close back after the count. - SeedNewer func(t *testing.T) - - // LateInstrument is the counter the manager records late rows on. The - // harness sums it across attributes. - LateInstrument string + // HoldCommit builds like New over a connection whose commit calls hold + // first. By then the close has written its delete and its watermark and + // committed neither. + HoldCommit func(t *testing.T, sink core.Sink, budget *core.DrainBudget, hold func()) Manager } // Managers proves every manager invariant against the subject. @@ -104,8 +95,8 @@ func Managers(t *testing.T, s ManagerSubject) { if s.Integration == "" { t.Fatal("conformance: ManagerSubject.Integration is required") } - if s.New == nil || s.Seed == nil || s.SeedLate == nil || s.Remaining == nil { - t.Fatal("conformance: ManagerSubject needs New, Seed, SeedLate and Remaining") + if s.New == nil || s.Seed == nil || s.SeedLate == nil || s.SeedNewer == nil || s.Remaining == nil { + t.Fatal("conformance: ManagerSubject needs New, Seed, SeedLate, SeedNewer and Remaining") } feature, hasFeature, err := coverage.FeatureFor(s.Integration) @@ -141,8 +132,6 @@ const ( managerDrainBounded = "manager.drain.bounded" watermarkNeverRegress = "manager.watermark.never_regresses" closeCommittedRowsOnly = "manager.close.committed_rows_only" - latePolicyHolds = "manager.late.policy_holds" - lateCountedOnce = "manager.late.counted_once" batchIndependentOfIO = "pipeline.batch.independent_of_window_io" ) @@ -150,12 +139,9 @@ const ( // a manager that deletes one bucket per flush is caught. const seededWindows = 2 -// managerPoll is the interval the harness runs the loop at. -const managerPoll = 20 * time.Millisecond - -// managerWait bounds every wait on the loop. It is many times the poll -// interval on purpose: a tighter bound asserts how fast the machine is, -// which is how #245 failed on CI. +// managerWait bounds every wait on the loop. It is generous on purpose: a +// tighter bound asserts how fast the machine is, which is how #245 failed on +// CI. const managerWait = 5 * time.Second func managerVerdicts(t *testing.T, s ManagerSubject) []verdict { @@ -168,8 +154,6 @@ func managerVerdicts(t *testing.T, s ManagerSubject) []verdict { bounded := verdict{invariant: managerDrainBounded} regress := verdict{invariant: watermarkNeverRegress} committedOnly := verdict{invariant: closeCommittedRowsOnly} - late := verdict{invariant: latePolicyHolds} - countedOnce := verdict{invariant: lateCountedOnce} independent := verdict{invariant: batchIndependentOfIO} if err := checkDeleteAfterFlush(t, s); err != nil { @@ -196,15 +180,6 @@ func managerVerdicts(t *testing.T, s ManagerSubject) []verdict { } else if err := checkCommittedRowsOnly(t, s); err != nil { committedOnly.failure = err.Error() } - if err := checkLatePolicyHolds(t, s); err != nil { - late.failure = err.Error() - } - if s.Metered == nil || s.SeedNewer == nil || s.LateInstrument == "" { - countedOnce.skipped = s.Integration + " cannot report what it counted; " + - "supply Metered, SeedNewer and LateInstrument" - } else if err := checkLateCountedOnce(t, s); err != nil { - countedOnce.failure = err.Error() - } if s.Batch == nil || s.HoldCommit == nil { independent.skipped = s.Integration + " cannot run a batch beside a " + "held close; supply Batch and HoldCommit" @@ -213,7 +188,7 @@ func managerVerdicts(t *testing.T, s ManagerSubject) []verdict { } return []verdict{afterFlush, onFailure, eventually, exits, bounded, regress, - committedOnly, late, countedOnce, independent} + committedOnly, independent} } // newManagerRun seeds the table and builds the manager on a recording sink. @@ -228,11 +203,11 @@ func newManagerRun(t *testing.T, s ManagerSubject, fail bool) (Manager, *Recorde // The default deadline, so no check but the bounded one can run it out. budget := core.NewDrainBudget(core.DefaultDrainDeadline) t.Cleanup(budget.Stop) - return s.New(t, sink.counted, managerPoll, budget, "reemit"), rec, sink + return s.New(t, sink.counted, budget), rec, sink } // checkDeleteAfterFlush holds that the table still has every closed bucket -// at the moment the sink is flushed, and none once the poll returns. +// at the moment the sink is flushed, and none once the pass returns. // // Observed from inside the flush rather than inferred from the count // afterwards: a manager that deleted first and flushed second leaves the @@ -249,11 +224,11 @@ func checkDeleteAfterFlush(t *testing.T, s ManagerSubject) error { } } - if err := m.Poll(context.Background()); err != nil { - return fmt.Errorf("a poll with no fault injected failed: %v", err) + if err := m.Pass(context.Background()); err != nil { + return fmt.Errorf("a pass with no fault injected failed: %v", err) } if !sameOrder(rec.Events(), []string{"flush"}) { - return fmt.Errorf("the poll did %s; want one flush", list(rec.Events())) + return fmt.Errorf("the pass did %s; want one flush", list(rec.Events())) } if atFlush != seededWindows { return fmt.Errorf( @@ -265,7 +240,7 @@ func checkDeleteAfterFlush(t *testing.T, s ManagerSubject) error { if left := s.Remaining(t); left != 0 { return fmt.Errorf( "the sink accepted %d buckets and %d rows are still in the "+ - "table, so the next poll publishes them again", + "table, so the next pass publishes them again", seededWindows, left) } return nil @@ -276,8 +251,8 @@ func checkDeleteNothingOnFailure(t *testing.T, s ManagerSubject) error { t.Helper() m, rec, _ := newManagerRun(t, s, true) - if err := m.Poll(context.Background()); err == nil { - return fmt.Errorf("the flush failed and the poll did not") + if err := m.Pass(context.Background()); err == nil { + return fmt.Errorf("the flush failed and the pass did not") } if !sameOrder(rec.Events(), []string{"flush-failed"}) { return fmt.Errorf("after a failed flush the manager did %s; want "+ @@ -290,28 +265,28 @@ func checkDeleteNothingOnFailure(t *testing.T, s ManagerSubject) error { "nothing downstream can tell", seededWindows-left, seededWindows) } - // And the watermark did not move: a second poll publishes the same - // buckets rather than treating them as late. + // And the watermark did not move: a second pass publishes the same + // buckets rather than finding them behind it. rec2 := &Recorder{} sink2 := newRecordingSink(rec2, nil, noop.NewMeterProvider()) budget := core.NewDrainBudget(core.DefaultDrainDeadline) defer budget.Stop() - m2 := s.New(t, sink2.counted, managerPoll, budget, "drop") - if err := m2.Poll(context.Background()); err != nil { - return fmt.Errorf("the poll after a failed flush failed: %v", err) + m2 := s.New(t, sink2.counted, budget) + if err := m2.Pass(context.Background()); err != nil { + return fmt.Errorf("the pass after a failed flush failed: %v", err) } if sink2.Rows() != seededWindows { return fmt.Errorf( - "after a failed flush the next poll published %d of %d rows. The "+ - "watermark moved past buckets the sink never took, so under "+ - "drop they were discarded as late", + "after a failed flush the next pass published %d of %d rows. The "+ + "watermark moved past buckets the sink never took, so nothing "+ + "will ever publish them", sink2.Rows(), seededWindows) } return nil } // checkPublishEventually starts the loop over closed buckets and holds that -// they reach the sink without anything else happening. +// they reach the sink on the start pass, without anything else happening. func checkPublishEventually(t *testing.T, s ManagerSubject) error { t.Helper() m, rec, _ := newManagerRun(t, s, false) @@ -379,7 +354,7 @@ func checkFailureExits(t *testing.T, s ManagerSubject) error { <-done return fmt.Errorf( "the sink rejected every flush for %s and the manager kept "+ - "polling, %d attempts. Each poll collected the same buckets and "+ + "going, %d attempts. Each pass collected the same buckets and "+ "failed the same way, and the process stayed up. A supervisor "+ "never notices, and the table never drains", managerWait, sink.Flushes()) @@ -400,9 +375,11 @@ func checkFailureExits(t *testing.T, s ManagerSubject) error { return nil } -// checkManagerDrainBounded cancels Start against a sink that never answers -// and holds that it returns inside the drain deadline, with every closed -// bucket still in the table for the next start to publish. +// checkManagerDrainBounded starts the loop on a context already cancelled, +// against a sink that never answers, and holds that it returns inside the +// drain deadline, with every closed bucket still in the table for the next +// start to publish. Already cancelled, so the only pass that runs is the +// drain's: the case this check is about. func checkManagerDrainBounded(t *testing.T, s ManagerSubject) error { t.Helper() s.Seed(t, seededWindows) @@ -413,36 +390,35 @@ func checkManagerDrainBounded(t *testing.T, s ManagerSubject) error { defer watchdog.Stop() budget := core.NewDrainBudget(drainBudget) defer budget.Stop() - // An hour between polls, so the only poll that runs is the final one. - m := s.New(t, sink.counted, time.Hour, budget, "reemit") + m := s.New(t, sink.counted, budget) ctx, cancel := context.WithCancel(context.Background()) + cancel() done := make(chan error, 1) started := time.Now() go func() { done <- m.Start(ctx) }() - cancel() select { case err := <-done: if err == nil { - return fmt.Errorf("the sink never answered the final poll and Start " + + return fmt.Errorf("the sink never answered the final pass and Start " + "returned nil, so the shutdown reports buckets published that were not") } case <-time.After(2 * drainBoundedWait): - return fmt.Errorf("a %s drain deadline held the final poll past %s against "+ + return fmt.Errorf("a %s drain deadline held the final pass past %s against "+ "a sink that never answered, and the watchdog could not release it", drainBudget, 2*drainBoundedWait) } if took := time.Since(started); took >= drainBoundedWait { - return fmt.Errorf("a %s drain deadline held the final poll for %s against "+ + return fmt.Errorf("a %s drain deadline held the final pass for %s against "+ "a sink that never answered", drainBudget, took) } if sink.Flushes() != 1 { - return fmt.Errorf("the final poll flushed %d times; want the one attempt the "+ + return fmt.Errorf("the final pass flushed %d times; want the one attempt the "+ "deadline ended", sink.Flushes()) } if left := s.Remaining(t); left != seededWindows { - return fmt.Errorf("the final poll ran out of time and %d of %d closed buckets "+ + return fmt.Errorf("the final pass ran out of time and %d of %d closed buckets "+ "are gone from the table. They were never delivered, so a restart "+ "cannot publish them", seededWindows-left, seededWindows) } @@ -452,11 +428,13 @@ func checkManagerDrainBounded(t *testing.T, s ManagerSubject) error { // checkWatermarkNeverRegresses closes buckets, then builds a second manager // over the same persisted state and holds that it publishes nothing the // first one published: the watermark it loads is the one the first one -// saved, and a table that now holds older rows does not pull it back. +// saved, and a table that now holds older rows does not pull it back. Nor +// does a pass in which nothing moved touch those rows: lateness is the +// engine's to decide, and the manager deletes only what a move expired. func checkWatermarkNeverRegresses(t *testing.T, s ManagerSubject) error { t.Helper() m, _, sink := newManagerRun(t, s, false) - if err := m.Poll(context.Background()); err != nil { + if err := m.Pass(context.Background()); err != nil { return fmt.Errorf("the first close failed: %v", err) } if sink.Rows() != seededWindows { @@ -470,9 +448,9 @@ func checkWatermarkNeverRegresses(t *testing.T, s ManagerSubject) error { sink2 := newRecordingSink(rec2, nil, noop.NewMeterProvider()) budget := core.NewDrainBudget(core.DefaultDrainDeadline) defer budget.Stop() - m2 := s.New(t, sink2.counted, managerPoll, budget, "drop") - if err := m2.Poll(context.Background()); err != nil { - return fmt.Errorf("the second manager's poll failed: %v", err) + m2 := s.New(t, sink2.counted, budget) + if err := m2.Pass(context.Background()); err != nil { + return fmt.Errorf("the second manager's pass failed: %v", err) } if sink2.Rows() != 0 { return fmt.Errorf( @@ -481,8 +459,10 @@ func checkWatermarkNeverRegresses(t *testing.T, s ManagerSubject) error { "watermark than the one saved, so a restart publishes buckets "+ "twice", sink2.Rows()) } - if left := s.Remaining(t); left != 0 { - return fmt.Errorf("under drop, %d late rows are still in the table", left) + if left := s.Remaining(t); left != 1 { + return fmt.Errorf("a pass in which the watermark did not move left %d rows "+ + "of the 1 seeded behind it; it published nothing, so it had no "+ + "business deleting anything", left) } return nil } @@ -497,8 +477,8 @@ func checkCommittedRowsOnly(t *testing.T, s ManagerSubject) error { release := s.Uncommitted(t, seededWindows) defer release() - if err := m.Poll(context.Background()); err != nil { - return fmt.Errorf("a poll beside an open transaction failed: %v", err) + if err := m.Pass(context.Background()); err != nil { + return fmt.Errorf("a pass beside an open transaction failed: %v", err) } if sink.Rows() != seededWindows { return fmt.Errorf( @@ -509,139 +489,6 @@ func checkCommittedRowsOnly(t *testing.T, s ManagerSubject) error { return nil } -// checkLatePolicyHolds closes buckets, seeds rows for one of them, and holds -// that drop discards them without a flush and reemit publishes them once. -func checkLatePolicyHolds(t *testing.T, s ManagerSubject) error { - t.Helper() - for _, policy := range []string{"drop", "reemit"} { - s.Seed(t, seededWindows) - rec := &Recorder{} - sink := newRecordingSink(rec, nil, noop.NewMeterProvider()) - budget := core.NewDrainBudget(core.DefaultDrainDeadline) - m := s.New(t, sink.counted, managerPoll, budget, policy) - if err := m.Poll(context.Background()); err != nil { - budget.Stop() - return fmt.Errorf("%s: the first close failed: %v", policy, err) - } - - s.SeedLate(t, 3) - if err := m.Poll(context.Background()); err != nil { - budget.Stop() - return fmt.Errorf("%s: the poll after the late rows failed: %v", policy, err) - } - budget.Stop() - - flushes, rows, left := sink.Flushes(), sink.Rows(), s.Remaining(t) - switch policy { - case "drop": - if flushes != 1 || rows != seededWindows { - return fmt.Errorf( - "drop: late rows for a closed bucket were published: %d flushes "+ - "and %d rows, want 1 and %d. An append-only sink now holds "+ - "the bucket twice", flushes, rows, seededWindows) - } - if left != 0 { - return fmt.Errorf("drop: %d late rows are still in the table", left) - } - case "reemit": - if flushes != 2 || rows != seededWindows+3 { - return fmt.Errorf( - "reemit: late rows for a closed bucket were not published: %d "+ - "flushes and %d rows, want 2 and %d", flushes, rows, seededWindows+3) - } - if left != 0 { - return fmt.Errorf("reemit: %d late rows are still in the table after "+ - "being published", left) - } - } - } - return nil -} - -// lateRows is how many late rows the counted-once check writes. -const lateRows = 3 - -// checkLateCountedOnce proves the late-row counter counts what happened, not -// what was attempted, under both policies. -// -// A close counts late rows, then fails -- here the sink refuses the flush of a -// newer bucket -- and rolls back, restoring the rows it dropped or reemitted. -// The counter must still read zero. The close that later succeeds counts the -// same rows, once. Counted before the commit, the failed close kept its count -// and the next one added the same rows again: under drop, the data-loss -// counter read twice the rows that were lost. -func checkLateCountedOnce(t *testing.T, s ManagerSubject) error { - t.Helper() - for _, policy := range []string{"drop", "reemit"} { - reader := sdkmetric.NewManualReader() - mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) - counted := func() (int64, error) { - var rm metricdata.ResourceMetrics - if err := reader.Collect(context.Background(), &rm); err != nil { - return 0, err - } - var total int64 - for _, sm := range rm.ScopeMetrics { - for _, m := range sm.Metrics { - if m.Name != s.LateInstrument { - continue - } - if sum, ok := m.Data.(metricdata.Sum[int64]); ok { - for _, dp := range sum.DataPoints { - total += dp.Value - } - } - } - } - return total, nil - } - poll := func(fail bool) error { - budget := core.NewDrainBudget(core.DefaultDrainDeadline) - defer budget.Stop() - sink := newRecordingSink(&Recorder{}, nil, noop.NewMeterProvider()) - sink.fail = fail - return s.Metered(t, sink.counted, budget, policy, mp).Poll(context.Background()) - } - - s.Seed(t, seededWindows) - if err := poll(false); err != nil { - return fmt.Errorf("%s: the first close failed: %v", policy, err) - } - - s.SeedLate(t, lateRows) - s.SeedNewer(t) - if err := poll(true); err == nil { - return fmt.Errorf("%s: a close whose sink refused the flush reported success", policy) - } - n, err := counted() - if err != nil { - return fmt.Errorf("%s: reading %s: %v", policy, s.LateInstrument, err) - } - if n != 0 { - return fmt.Errorf( - "%s: a close that failed and rolled back counted %d late rows, want 0. "+ - "The rows are back in the table, and the next close counts them again", - policy, n) - } - - if err := poll(false); err != nil { - return fmt.Errorf("%s: the close after the failure failed: %v", policy, err) - } - n, err = counted() - if err != nil { - return fmt.Errorf("%s: reading %s: %v", policy, s.LateInstrument, err) - } - if n != lateRows { - return fmt.Errorf("%s: %d late rows were counted, want %d: each late row "+ - "counts once, when the close that settled it commits", policy, n, lateRows) - } - if left := s.Remaining(t); left != 0 { - return fmt.Errorf("%s: %d rows are still in the table after the close", policy, left) - } - } - return nil -} - // batchesDuringHold is how many batches run while a close is held. More // than one, so a handler that fails only its second batch after a refused // statement is caught. @@ -682,17 +529,17 @@ func checkBatchIndependentOfWindowIO(t *testing.T, s ManagerSubject) error { } } sink := newRecordingSink(rec, nil, noop.NewMeterProvider()) - m := s.HoldCommit(t, sink.counted, budget, "reemit", func() { hold(&commitHeld) }) + m := s.HoldCommit(t, sink.counted, budget, func() { hold(&commitHeld) }) during := func(armed *atomic.Bool, what string) error { entered, release = make(chan struct{}), make(chan struct{}) armed.Store(true) - polled := make(chan error, 1) - go func() { polled <- m.Poll(context.Background()) }() + passed := make(chan error, 1) + go func() { passed <- m.Pass(context.Background()) }() select { case <-entered: - case err := <-polled: + case err := <-passed: return fmt.Errorf("the close returned before reaching %s: %v", what, err) case <-time.After(managerWait): close(release) @@ -712,7 +559,7 @@ func checkBatchIndependentOfWindowIO(t *testing.T, s ManagerSubject) error { close(release) select { - case err := <-polled: + case err := <-passed: if batchErr != nil { return batchErr } @@ -730,9 +577,9 @@ func checkBatchIndependentOfWindowIO(t *testing.T, s ManagerSubject) error { if err := during(&flushHeld, "its sink's flush"); err != nil { return err } - // A late row under reemit makes the next close publish, delete and - // update the watermark's row, then stop before committing. - s.SeedLate(t, 1) + // A newer bucket the engine asserted past makes the next close publish, + // delete and update the watermark's row, then stop before committing. + s.SeedNewer(t) if err := during(&commitHeld, "its uncommitted delete and watermark"); err != nil { return err } diff --git a/internal/core/bucket.go b/internal/core/bucket.go new file mode 100644 index 00000000..ab78f67a --- /dev/null +++ b/internal/core/bucket.go @@ -0,0 +1,29 @@ +package core + +import "time" + +// bucketOrigin is where DuckDB's time_bucket aligns its buckets: 2000-01-03, +// a Monday, so that week-sized buckets start on Mondays. The engine buckets +// from the same origin so that the bucket it decides a record's lateness +// against is the bucket the handler's time_bucket puts the row in. +// +// Go's time.Truncate aligns to the zero time, year 1 -- also a Monday, which +// is why the two agree for any size that divides a day and for multiples of +// a week, and disagree for sizes like 25 hours or 3 days. Using DuckDB's +// origin explicitly makes them agree for every size; BucketStart's test +// proves it against DuckDB itself. +var bucketOrigin = time.Date(2000, 1, 3, 0, 0, 0, 0, time.UTC) + +// BucketStart is time_bucket(INTERVAL 'size', at): the start of the bucket +// of length size that holds at, aligned to bucketOrigin. at is after the +// origin for every event time the engine places (EventTimeFloor is 2020). +func BucketStart(at time.Time, size time.Duration) time.Time { + since := at.Sub(bucketOrigin) + return bucketOrigin.Add(since - since%size) +} + +// BucketEnd is the end of the bucket holding at: BucketStart + size. A +// bucket is closed when its end is at or before the watermark. +func BucketEnd(at time.Time, size time.Duration) time.Time { + return BucketStart(at, size).Add(size) +} diff --git a/internal/core/bucket_test.go b/internal/core/bucket_test.go new file mode 100644 index 00000000..445dbfdd --- /dev/null +++ b/internal/core/bucket_test.go @@ -0,0 +1,79 @@ +package core + +import ( + "context" + "fmt" + "testing" + "time" + + "github.com/apache/arrow-go/v18/arrow/array" + "github.com/turbolytics/sql-flow/internal/coverage" + "github.com/turbolytics/sql-flow/internal/duckdb" + "github.com/zeebo/assert" +) + +// The engine decides a record's lateness from its bucket, and the handler's +// SQL computes time_column with time_bucket. If the two disagree, lateness +// is decided against the wrong bucket. This proves BucketStart is +// time_bucket, over every size a window is likely to declare and the ones +// where a naive truncation is wrong: 25 hours and 3 days do not divide the +// span from Go's zero time to DuckDB's origin, and time.Truncate gets them +// wrong by hours. +func TestWindowBucket_AgreesWithDuckDBTimeBucket(t *testing.T) { + coverage.Covers(t, "manager.window") + ctx := context.Background() + db, err := duckdb.OpenPath(ctx, "") + assert.NoError(t, err) + defer db.Close() + conn, err := db.Connect(ctx) + assert.NoError(t, err) + defer conn.Close() + + instants := []time.Time{ + time.Date(2026, 9, 26, 12, 34, 56, 789000000, time.UTC), + time.Date(2026, 9, 26, 0, 0, 0, 0, time.UTC), + time.Date(2026, 9, 26, 23, 59, 59, 999999000, time.UTC), + time.Date(2026, 1, 1, 0, 0, 0, 0, time.UTC), + EventTimeFloor, + time.Date(2026, 9, 26, 12, 35, 0, 0, time.UTC), // exactly on a minute + time.Date(2026, 9, 26, 12, 34, 59, 999999000, time.UTC), + } + sizes := []int{1, 7, 60, 90, 300, 420, 3600, 5400, 7200, 86400, 90000, 172800, 259200, 604800} + for _, size := range sizes { + for _, at := range instants { + q := fmt.Sprintf(`SELECT epoch_us(time_bucket(INTERVAL '%d seconds', TIMESTAMPTZ '%s'))`, + size, at.Format("2006-01-02 15:04:05.999999-07:00")) + stmt, err := conn.NewStatement() + assert.NoError(t, err) + assert.NoError(t, stmt.SetSqlQuery(q)) + rdr, _, err := stmt.ExecuteQuery(ctx) + assert.NoError(t, err) + var micros int64 + for rdr.Next() { + micros = rdr.Record().Column(0).(*array.Int64).Value(0) + } + rdr.Release() + stmt.Close() + + want := time.UnixMicro(micros).UTC() + d := time.Duration(size) * time.Second + if got := BucketStart(at, d); !got.Equal(want) { + t.Fatalf("size %ds at %s: BucketStart %s, time_bucket %s", size, at, got, want) + } + assert.Equal(t, want.Add(d), BucketEnd(at, d)) + } + } +} + +// And the case the origin exists for, pinned on its own so the grid above +// cannot quietly lose it: a naive truncation from Go's zero time is wrong +// for a 25-hour bucket. +func TestWindowBucket_ANaiveTruncationWouldBeWrong(t *testing.T) { + coverage.Covers(t, "manager.window") + at := time.Date(2026, 9, 26, 12, 34, 56, 0, time.UTC) + size := 25 * time.Hour + naive := at.Truncate(size) + got := BucketStart(at, size) + assert.That(t, !got.Equal(naive)) + assert.Equal(t, time.Date(2026, 9, 25, 12, 0, 0, 0, time.UTC), got) // what DuckDB says +} diff --git a/internal/core/metrics.go b/internal/core/metrics.go index 1e8e0825..578ddd01 100644 --- a/internal/core/metrics.go +++ b/internal/core/metrics.go @@ -87,6 +87,13 @@ type Metrics struct { // WithEventTimePlacement. This is a data-loss signal, and a device with // a wrong clock is what it usually means. MessagesUnplaceable metric.Int64Counter + // WindowLateRows counts records late for a window, by outcome: refused + // (the bucket closed more than allowed_lateness_seconds before the + // watermark; never written) or recomputed (within it; written, and the + // bucket republished whole). Decided at arrival, by the engine, so it is + // counted here rather than by the manager. The data-loss half is + // refused. + WindowLateRows metric.Int64Counter // LagObserved is when consumer_lag was last recorded, as Unix seconds. // // Lag is recorded when a message is processed, so a consumer that stops @@ -395,6 +402,14 @@ func NewMetrics(mp metric.MeterProvider) (*Metrics, error) { ); err != nil { return nil, fmt.Errorf("messages_unplaceable_total: %w", err) } + // No unit: "count" is one the exporter does not know, and window_late_rows + // is the series name the bundle reads. + if m.WindowLateRows, err = meter.Int64Counter( + "window_late_rows", + metric.WithDescription("Rows late for a window: refused beyond allowed_lateness_seconds, or recomputed within it"), + ); err != nil { + return nil, fmt.Errorf("window_late_rows: %w", err) + } if m.LagObserved, err = meter.Int64Gauge( "consumer_lag_observed_timestamp", diff --git a/internal/core/turbine.go b/internal/core/turbine.go index 293b3021..114bd39b 100644 --- a/internal/core/turbine.go +++ b/internal/core/turbine.go @@ -361,6 +361,17 @@ type Turbine struct { // see notePlaced. Reused across batches, and the consume loop is its only // party, so it takes no lock. observed []Observation + // pendingLate is the buckets this batch's late-but-allowed records landed + // in, handed to the windows' signals after the commit that made the rows + // visible. Consume loop only; cleared on a rollback, because the replay + // classifies the same records again. + pendingLate []LateBucket + // lateRefusedLogged is whether a refused late record has been logged + // this run, so a device stuck in the past writes one line. + lateRefusedLogged bool + // lateAttrs caches the attribute set per window and outcome, because + // WithAttributes allocates and refusals can come per record. + lateAttrs map[string][]metric.AddOption // partitionsRelayed is whether the source reports its partitions to // windows itself. Otherwise the loop asks the source whether it can // deliver before each commit, and notDelivering is the last answer, for @@ -527,9 +538,8 @@ func WithWatermarkWriteInterval(d time.Duration) TurbineOption { // // On for a pipeline that windows. A window rests on event time, and one // record stamped in the future otherwise drags the watermark past real time -// and makes every correctly stamped row after it late, which under late_rows -// drop deletes them: one device with a fast clock empties a fleet's stream -// (#358). A pipeline with no window has nothing for such a record to damage, +// and makes every correctly stamped row after it late, which refuses them: +// one device with a fast clock empties a fleet's stream (#358). A pipeline with no window has nothing for such a record to damage, // so it keeps it. // // Refused here rather than in the window manager on purpose. The manager @@ -602,9 +612,9 @@ const progressWriteInterval = time.Second // (BenchmarkCommitStateWindowed), and on a busy pipeline event time advances // on every batch, so writing per commit is a statement per batch -- the same // cost the progress write was paced to avoid, for a value with the same -// reader. A window manager reads this once per poll_interval_seconds, ten -// seconds by default, so writing it more often than the pace below buys -// nothing at all. +// reader. The window manager reads it on the kick that follows the write, so +// the pace below is the most a close can trail its commit, and a second of +// that on a busy stream costs nothing a reader can see. const watermarkWriteInterval = time.Second // watchSource tells the tracker what the source holds. A source that @@ -1179,17 +1189,42 @@ func (t *Turbine) ConsumeLoop(ctx context.Context, maxMsgs int) (stats *Stats, e } continue } - // Placed, so it counts toward the watermark, whether or not the - // handler takes it: a record an error policy drops still says - // where the stream has got to, as Flink's assigner does. - // - // Accumulated here and handed to the tracker once, when the batch - // is processed. Calling the tracker per record takes its mutex per - // record, which measured at 52ns against a loop whose own budget - // is around 40 -- the same shape as the phase timing this loop - // already refuses. A partition is one entry, and a fetch spans - // few, so the scan below is a comparison or two. + // Placed, so the window's lateness rule applies: a record whose + // bucket closed more than allowed_lateness_seconds ago has + // nowhere to go, and is refused before the handler for the same + // reason an unplaceable one is -- a row the window's promise + // excludes must not reach the table. One whose bucket closed + // within lateness is written, and the bucket is republished + // whole. Decided here, at arrival, as Flink's window operator + // decides it; nothing sweeps the table for late rows afterwards. if t.windows != nil { + refused, late := t.windows.Classify(raw.EventAtNanos) + if refused { + t.noteRefusedLate(ctx, raw) + t.mark(raw) + totalConsumed++ + t.stats.SetNumMessagesConsumed(totalConsumed) + if maxMsgs > 0 && totalConsumed >= int64(maxMsgs) { + t.logger.Info("max messages consumed, stopping consumer loop") + hitMax = true + break + } + continue + } + if len(late) > 0 { + t.pendingLate = append(t.pendingLate, late...) + } + // And it counts toward the watermark, whether or not the + // handler takes it: a record an error policy drops still + // says where the stream has got to, as Flink's assigner does. + // + // Accumulated here and handed to the tracker once, when the + // batch is processed. Calling the tracker per record takes + // its mutex per record, which measured at 52ns against a loop + // whose own budget is around 40 -- the same shape as the + // phase timing this loop already refuses. A partition is one + // entry, and a fetch spans few, so the scan is a comparison + // or two. t.notePlaced(raw.Topic, raw.Partition, raw.EventAtNanos) } if err := t.writeMessage(raw); err != nil { @@ -1572,6 +1607,7 @@ func (t *Turbine) rollbackState(ctx context.Context) { if err := t.stateTx.Rollback(context.WithoutCancel(ctx)); err != nil { t.logger.Error("rollback failed", zap.Error(err)) } + t.dropPendingLate() } // commitState writes the processed offsets into the state database and commits @@ -1610,8 +1646,10 @@ func (t *Turbine) commitState(ctx context.Context, write progressWrite) error { case watermarkErr != nil: t.recordError(ctx, errs.Wrap(errs.CodeWatermarkWriteFailed, watermarkErr, "sqlflow_watermarks not written"), phaseStateCommit, "sqlflow_watermarks not written") + t.dropPendingLate() case t.windows != nil: t.windows.Commit(moved) + t.signalWindows(ctx, moved) } t.lock.Lock() t.commits++ @@ -1639,6 +1677,7 @@ func (t *Turbine) commitState(ctx context.Context, write progressWrite) error { if rbErr := t.stateTx.Rollback(context.WithoutCancel(ctx)); rbErr != nil { t.logger.Error("rollback after failed progress write", zap.Error(rbErr)) } + t.dropPendingLate() return errs.Wrap(errs.CodeStateCommitFailed, progressErr, "the progress write failed inside the state transaction") } @@ -1649,6 +1688,7 @@ func (t *Turbine) commitState(ctx context.Context, write progressWrite) error { if rbErr := t.stateTx.Rollback(context.WithoutCancel(ctx)); rbErr != nil { t.logger.Error("rollback after failed watermark write", zap.Error(rbErr)) } + t.dropPendingLate() return errs.Wrap(errs.CodeStateCommitFailed, watermarkErr, "asserting the watermark") } @@ -1656,6 +1696,7 @@ func (t *Turbine) commitState(ctx context.Context, write progressWrite) error { if rbErr := t.stateTx.Rollback(context.WithoutCancel(ctx)); rbErr != nil { t.logger.Error("rollback after failed offset save", zap.Error(rbErr)) } + t.dropPendingLate() return errs.Wrap(errs.CodeStateCommitFailed, err, "saving offsets") } @@ -1666,10 +1707,12 @@ func (t *Turbine) commitState(ctx context.Context, write progressWrite) error { if rbErr := t.stateTx.Rollback(context.WithoutCancel(ctx)); rbErr != nil { t.logger.Error("rollback after failed commit", zap.Error(rbErr)) } + t.dropPendingLate() return errs.Wrap(errs.CodeStateCommitFailed, err, "committing state") } if t.windows != nil { t.windows.Commit(moved) + t.signalWindows(ctx, moved) } t.commits++ @@ -1930,6 +1973,73 @@ func (t *Turbine) notePlaced(topic string, partition int32, atNanos int64) { }) } +// noteRefusedLate counts a record refused as late, once per window, and +// logs the condition once per run: a device stuck in the past writes a +// line rather than a line per record. +func (t *Turbine) noteRefusedLate(ctx context.Context, m Message) { + for _, spec := range t.windows.Specs() { + t.metrics.WindowLateRows.Add(ctx, 1, t.lateAttrsFor(spec.Name, "refused")...) + } + if t.lateRefusedLogged { + return + } + t.lateRefusedLogged = true + t.logger.Warn("refusing records late beyond allowed_lateness_seconds", + zap.Time("event_time", time.Unix(0, m.EventAtNanos).UTC())) +} + +// lateAttrsFor is the cached attribute set for one window and outcome. +func (t *Turbine) lateAttrsFor(window, outcome string) []metric.AddOption { + key := window + "\x00" + outcome + if opts, ok := t.lateAttrs[key]; ok { + return opts + } + if t.lateAttrs == nil { + t.lateAttrs = map[string][]metric.AddOption{} + } + opts := []metric.AddOption{metric.WithAttributes( + attribute.String("window", window), attribute.String("outcome", outcome))} + t.lateAttrs[key] = opts + return opts +} + +// signalWindows tells each window's manager what this commit changed: a kick +// for every window whose watermark moved, and the recompute buckets this +// batch's late rows landed in, followed by a kick. After the commit, never +// before, so a manager woken here reads what the commit wrote. Called only +// when the commit succeeded; a rollback drops pendingLate, and the replay +// classifies the same records again. The recompute counter is recorded here +// for the same reason: a batch that rolls back counts nothing. +func (t *Turbine) signalWindows(ctx context.Context, moved map[string]time.Time) { + if t.windows == nil { + return + } + kicked := map[string]bool{} + for name := range moved { + if sig := t.windows.Signal(name); sig != nil { + sig.Kick() + kicked[name] = true + } + } + for _, lb := range t.pendingLate { + sig := t.windows.Signal(lb.Window) + if sig == nil { + continue + } + sig.Recompute(lb.Bucket) + t.metrics.WindowLateRows.Add(ctx, 1, t.lateAttrsFor(lb.Window, "recomputed")...) + if !kicked[lb.Window] { + sig.Kick() + kicked[lb.Window] = true + } + } + t.pendingLate = t.pendingLate[:0] +} + +// dropPendingLate forgets this batch's late buckets: the batch rolled back, +// and its replay will find them again. +func (t *Turbine) dropPendingLate() { t.pendingLate = t.pendingLate[:0] } + // flushObservations hands the batch's observations to the tracker and keeps // the slice for the next batch, so a steady pipeline allocates nothing here. func (t *Turbine) flushObservations() { diff --git a/internal/core/watermarks.go b/internal/core/watermarks.go index b0c5e586..45920b06 100644 --- a/internal/core/watermarks.go +++ b/internal/core/watermarks.go @@ -3,7 +3,9 @@ package core import ( "context" "fmt" + "sort" "sync" + "sync/atomic" "time" "github.com/apache/arrow-adbc/go/adbc" @@ -105,13 +107,18 @@ import ( // half exists, and a stream that stops leaves its last bucket open -- shape A // and B's documented behaviour, for any stream, not a property of revocation. -// WindowSpec is what the engine needs to assert one window's watermark: the -// name it is stored under, and the three durations the config declares. +// WindowSpec is what the engine needs to assert one window's watermark and +// decide a record's lateness against it: the name it is stored under, and +// the four durations the config declares. type WindowSpec struct { Name string Size time.Duration Grace time.Duration IdleClose time.Duration + // Lateness is how long a closed bucket's rows are kept. A record whose + // bucket ended more than this before the watermark is refused; one whose + // bucket ended within it is written and the bucket republished whole. + Lateness time.Duration } // WatermarksTable holds the engine's assertion per window. Written by the @@ -164,6 +171,13 @@ type Watermarks struct { // stored is each window's asserted watermark, in nanoseconds, zero // before the first assertion. It only ever grows. stored map[string]int64 + // asserted mirrors stored, one atomic per spec in spec order, for + // Classify: it runs per record on the consume loop and must not take mu. + asserted []atomic.Int64 + specIndex map[string]int + // signals is one per spec, in spec order: what the engine tells each + // window's manager. + signals []*WindowSignal } // NewWatermarks builds the tracker for a pipeline's windows on a clock. The @@ -174,15 +188,51 @@ func NewWatermarks(specs []WindowSpec, now func() time.Time) *Watermarks { now = time.Now } w := &Watermarks{ - specs: specs, - now: now, - parts: map[partitionKey]*partitionState{}, - floor: map[string]int64{}, - stored: map[string]int64{}, + specs: specs, + now: now, + parts: map[partitionKey]*partitionState{}, + floor: map[string]int64{}, + stored: map[string]int64{}, + asserted: make([]atomic.Int64, len(specs)), + specIndex: make(map[string]int, len(specs)), + signals: make([]*WindowSignal, len(specs)), + } + for i, spec := range specs { + w.specIndex[spec.Name] = i + w.signals[i] = newWindowSignal() } return w } +// NewWatermarksWithSignals is NewWatermarks over signals that already exist, +// one per spec in spec order. A process that rebuilds its tracker -- a +// simulated restart -- keeps the managers' signals, because a manager holds +// its signal for its lifetime. +func NewWatermarksWithSignals(specs []WindowSpec, now func() time.Time, signals []*WindowSignal) *Watermarks { + w := NewWatermarks(specs, now) + for i := range specs { + if i < len(signals) && signals[i] != nil { + w.signals[i] = signals[i] + } + } + return w +} + +// Specs is the windows this tracker asserts for. +func (w *Watermarks) Specs() []WindowSpec { return w.specs } + +// setAsserted records a window's watermark under mu and mirrors it for +// Classify. A value below what is stored is ignored: monotonic here too. +func (w *Watermarks) setAsserted(name string, nanos int64) { + if nanos <= w.stored[name] { + return + } + w.stored[name] = nanos + if i, ok := w.specIndex[name]; ok { + w.asserted[i].Store(nanos) + } +} + // Restore seeds a window with what its table holds at start: the newest // bucket start, or a zero time for an empty table. And the watermark last // asserted, so the first Advance after a restart cannot move it backwards. @@ -193,7 +243,7 @@ func (w *Watermarks) Restore(name string, newestBucketStart, asserted time.Time) w.floor[name] = newestBucketStart.UnixNano() } if !asserted.IsZero() { - w.stored[name] = asserted.UnixNano() + w.setAsserted(name, asserted.UnixNano()) } } @@ -357,9 +407,7 @@ func (w *Watermarks) Commit(moved map[string]time.Time) { w.mu.Lock() defer w.mu.Unlock() for name, at := range moved { - if n := at.UnixNano(); n > w.stored[name] { - w.stored[name] = n - } + w.setAsserted(name, at.UnixNano()) } } @@ -421,6 +469,125 @@ func (w *Watermarks) compute(spec WindowSpec, now time.Time) (int64, bool) { return maxSeen + int64(spec.Size), true } +// Lateness is where one record stands against one window. Decided at +// arrival from the record's bucket and the asserted watermark, as Flink's +// window operator decides it, and nowhere else: a row below the watermark is +// in the table only because this said it could be. +type Lateness int + +const ( + // OnTime is a record whose bucket is still open. + OnTime Lateness = iota + // LateAllowed is a record whose bucket closed within allowed lateness: + // written, and the bucket republished whole. + LateAllowed + // LateRefused is a record whose bucket closed beyond allowed lateness: + // never written, counted. + LateRefused +) + +// LateBucket is a window and the bucket a late-but-allowed record landed in. +type LateBucket struct { + Window string + Bucket time.Time +} + +// Classify decides every window at once for one record. refused is true only +// when every window refuses it, so a pipeline with two windows of different +// sizes keeps a record the looser one still wants; the stricter one's purge +// deletes it from that table on its next pass. recompute is the windows that +// admitted it late and want the bucket republished. A record with no event +// time has no bucket and is never late, and before a window has asserted +// anything nothing is late for it. +// +// Lock-free and allocation-free on the common path: the asserted watermarks +// are atomics, and recompute is nil until a window admits a late record. It +// runs per record on the consume loop, where a mutex measured at 52ns. +func (w *Watermarks) Classify(atNanos int64) (refused bool, recompute []LateBucket) { + if atNanos <= 0 || len(w.specs) == 0 { + return false, nil + } + at := time.Unix(0, atNanos).UTC() + refused = true + for i, spec := range w.specs { + asserted := w.asserted[i].Load() + if asserted == 0 { + refused = false + continue + } + end := BucketEnd(at, spec.Size).UnixNano() + switch { + case end+int64(spec.Lateness) <= asserted: + // Refused by this window. + case end <= asserted: + refused = false + recompute = append(recompute, LateBucket{Window: spec.Name, Bucket: BucketStart(at, spec.Size)}) + default: + refused = false + } + } + return refused, recompute +} + +// WindowSignal is what the engine tells one window's manager: that the +// watermark moved, and which closed buckets a late row landed in. Both are +// sent after the commit that made their rows visible, so a manager woken by +// a kick reads committed state. +type WindowSignal struct { + kick chan struct{} + mu sync.Mutex + recompute map[time.Time]struct{} +} + +func newWindowSignal() *WindowSignal { + return &WindowSignal{kick: make(chan struct{}, 1), recompute: map[time.Time]struct{}{}} +} + +// Kick wakes the manager. Non-blocking, capacity one: two kicks before the +// manager wakes are one kick, because a pass reads the current state rather +// than the event. +func (s *WindowSignal) Kick() { + select { + case s.kick <- struct{}{}: + default: + } +} + +// Wait is the channel a manager selects on. +func (s *WindowSignal) Wait() <-chan struct{} { return s.kick } + +// Recompute records a bucket a late-but-allowed row landed in. A set, so a +// burst of late rows for one bucket is one republish. The caller kicks after +// the commit that made the row visible. +func (s *WindowSignal) Recompute(bucket time.Time) { + s.mu.Lock() + s.recompute[bucket] = struct{}{} + s.mu.Unlock() +} + +// TakeRecompute drains the set, oldest bucket first. +func (s *WindowSignal) TakeRecompute() []time.Time { + s.mu.Lock() + out := make([]time.Time, 0, len(s.recompute)) + for b := range s.recompute { + out = append(out, b) + } + s.recompute = map[time.Time]struct{}{} + s.mu.Unlock() + sort.Slice(out, func(i, j int) bool { return out[i].Before(out[j]) }) + return out +} + +// Signal is the window's signal, for the manager built over it; nil for a +// name this tracker does not know. +func (w *Watermarks) Signal(name string) *WindowSignal { + i, ok := w.specIndex[name] + if !ok { + return nil + } + return w.signals[i] +} + // WatermarkSaver is what the engine writes assertions through; the DuckDB // store below is the one run wires. // diff --git a/internal/managers/conformance_test.go b/internal/managers/conformance_test.go index 6e54ed61..5d6122f1 100644 --- a/internal/managers/conformance_test.go +++ b/internal/managers/conformance_test.go @@ -2,18 +2,20 @@ package managers // The watermark manager under the conformance harness. // -// The subject supplies ten things: build the manager on the sink the -// harness hands it, put closed buckets in the table, put late rows in a -// bucket that closed, count what is left, hold rows open in the pipeline's -// transaction, run a batch through the structured handler on the pipeline's -// connection, build a manager whose commit waits, and build one whose metrics -// the harness reads beside a newer bucket its sink can refuse. The harness -// owns the sink, the faults and the verdicts. +// The subject supplies what the harness drives: build the manager on the +// sink the harness hands it, put closed buckets in the table, count what is +// left, hold rows open in the pipeline's transaction, run a batch through the +// structured handler on the pipeline's connection, and build a manager whose +// commit waits. The harness owns the sink, the faults and the verdicts. +// +// The engine's assertion is written by hand where a check changes the +// table: Seed asserts exactly its buckets' end. Everything closes on that +// and on nothing else; the manager takes no clock and no signal here, +// because every check drives Pass itself. import ( "context" "testing" - "time" "github.com/apache/arrow-adbc/go/adbc" "github.com/apache/arrow-go/v18/arrow" @@ -21,11 +23,10 @@ import ( "github.com/turbolytics/sql-flow/internal/core" "github.com/turbolytics/sql-flow/internal/coverage" "github.com/turbolytics/sql-flow/internal/handlers" - "go.opentelemetry.io/otel/metric" ) // heldConn is a manager's connection whose commit calls hold first, so a -// test can stop a close with its writes made and not committed. +// test can stop a pass with its writes made and not committed. type heldConn struct { adbc.Connection hold func() @@ -46,10 +47,6 @@ func TestManagerWatermark_Conformance(t *testing.T) { d := newTestDB(t, "") createWindowTable(t, d.pipeline) - // The engine's assertion is written by hand where a check changes the - // table: Seed asserts exactly its buckets' end, SeedNewer asserts past the - // bucket it adds. Everything closes on that and on nothing else. - // The handler the bluesky demo runs, on the pipeline's connection. It // checkpoints each time it re-initialises, which is the statement a // window's uncommitted write refused. @@ -68,42 +65,32 @@ func TestManagerWatermark_Conformance(t *testing.T) { } // Each New builds on a fresh connection, so a manager whose connection - // holds an open transaction from a failed poll does not block the next. + // holds an open transaction from a failed pass does not block the next. subject := conformance.ManagerSubject{ Integration: "manager.watermark", - New: func(t *testing.T, sink core.Sink, poll time.Duration, budget *core.DrainBudget, late string) conformance.Manager { + New: func(t *testing.T, sink core.Sink, budget *core.DrainBudget) conformance.Manager { conn := managerConn(t, d.db) t.Cleanup(func() { conn.Close() }) - decl := testDecl() - decl.Late = LatePolicy(late) - w, err := NewWatermark(conn, decl, poll, sink, - WithDrainBudget(budget)) + w, err := NewWatermark(conn, testDecl(), sink, nil, WithDrainBudget(budget)) if err != nil { t.Fatal(err) } return w }, - // n buckets, one row each, and the watermark forgotten, so every - // check starts from a table that has never closed anything. + // n buckets, one row each, the watermark forgotten so every check + // starts from a table that has never closed anything, and the engine + // having asserted the last seeded bucket's end. Seed: func(t *testing.T, n int) { exec(t, d.pipeline, `DELETE FROM agg_cities_count`) exec(t, d.pipeline, `DELETE FROM sqlflow_windows`) for i := 0; i < n; i++ { insertBucket(t, d.pipeline, i, "city", i+1) } - // The engine has asserted the last seeded bucket's end. assertAt(t, d.pipeline, bucket(n)) }, - // The first bucket closed with Seed's, so rows for it are late. - SeedLate: func(t *testing.T, n int) { - for i := 0; i < n; i++ { - insertBucket(t, d.pipeline, 0, "late", 1) - } - }, - Remaining: func(t *testing.T) int64 { return countRows(t, d.pipeline, testTable) }, @@ -140,12 +127,10 @@ func TestManagerWatermark_Conformance(t *testing.T) { return handler.Init(ctx) }, - HoldCommit: func(t *testing.T, sink core.Sink, budget *core.DrainBudget, late string, hold func()) conformance.Manager { + HoldCommit: func(t *testing.T, sink core.Sink, budget *core.DrainBudget, hold func()) conformance.Manager { conn := managerConn(t, d.db) t.Cleanup(func() { conn.Close() }) - decl := testDecl() - decl.Late = LatePolicy(late) - w, err := NewWatermark(heldConn{Connection: conn, hold: hold}, decl, time.Hour, sink, + w, err := NewWatermark(heldConn{Connection: conn, hold: hold}, testDecl(), sink, nil, WithDrainBudget(budget)) if err != nil { t.Fatal(err) @@ -153,28 +138,24 @@ func TestManagerWatermark_Conformance(t *testing.T) { return w }, - Metered: func(t *testing.T, sink core.Sink, budget *core.DrainBudget, late string, mp metric.MeterProvider) conformance.Manager { - conn := managerConn(t, d.db) - t.Cleanup(func() { conn.Close() }) - decl := testDecl() - decl.Late = LatePolicy(late) - w, err := NewWatermark(conn, decl, time.Hour, sink, - WithDrainBudget(budget), WithMeterProvider(mp)) - if err != nil { - t.Fatal(err) + // A row for the first seeded bucket, which a manager has closed: what + // a window with lateness retains, written by hand. + SeedLate: func(t *testing.T, n int) { + for i := 0; i < n; i++ { + insertBucket(t, d.pipeline, 0, "late", 1) } - return w }, - // A bucket past every one Seed wrote, still under the assertion. + // A bucket newer than every seeded one, and the engine asserted past + // its end, so the next pass publishes it, deletes it, and updates the + // window's row. SeedNewer: func(t *testing.T) { - insertBucket(t, d.pipeline, 10, "newer", 1) - assertAt(t, d.pipeline, bucket(11)) + insertBucket(t, d.pipeline, seededBuckets+1, "newer", 1) + assertAt(t, d.pipeline, bucket(seededBuckets+2)) }, - - LateInstrument: "window_late_rows", } conformance.Managers(t, subject) } -var _ adbc.Connection +// seededBuckets is what the harness seeds, so SeedNewer lands past it. +const seededBuckets = 2 diff --git a/internal/managers/decide.go b/internal/managers/decide.go index 9aad98f9..c6d7b15e 100644 --- a/internal/managers/decide.go +++ b/internal/managers/decide.go @@ -8,7 +8,7 @@ import ( // Every operation a window performs is a row in one of two tables here. // -// A poll builds a State from what it read, hands it to Decide, and performs +// A pass builds a State from what it read, hands it to Decide, and performs // the Action it gets back. The tables are the only place a fact meets an // action, so the logic is read in one place, every combination is accounted // for, and checkTables proves it at load: each combination selects exactly @@ -21,7 +21,13 @@ import ( // describes visible, and the manager closes on that and on nothing else: no // clock, no progress row, no reading of the table's newest bucket. Where // the assertion stands against what this window has already closed is the -// whole of what a poll decides on. +// whole of what a pass decides on. +// +// The bucket table has one fact too: where a bucket's end stands against +// the closed watermark, the asserted one, and the window's lateness. The +// engine decides a record's lateness at arrival against the same +// comparison (core.Watermarks.Classify), so the table here is the manager's +// half of one rule. // Asserted is where the engine's watermark stands against the closed one, // in event time. @@ -55,28 +61,30 @@ const ( Follow Action = "follow" ) -// Bucket is where one bucket stands against the previous closed watermark -// and the one this poll decided. +// Bucket is where one bucket stands against the closed watermark, the +// asserted one, and the window's lateness. type Bucket string const ( - // BucketLate is a bucket that ends at or before the previous watermark: - // its rows arrived after it closed. - BucketLate Bucket = "late" - // BucketDue is a bucket that ends after the previous watermark and at or - // before the next. - BucketDue Bucket = "due" - // BucketOpen is a bucket that ends after the next watermark. + // BucketOpen ends after the asserted watermark: its rows stay. BucketOpen Bucket = "open" + // BucketDue ends after the closed watermark and at or before the asserted + // one: it closes in this pass. + BucketDue Bucket = "due" + // BucketRetained ended at or before the closed watermark and its lateness + // has not run out: published, kept for a late row to republish it whole. + BucketRetained Bucket = "retained" + // BucketExpired ended at or before the asserted watermark less the + // lateness: nothing more can arrive for it, and its rows go. + BucketExpired Bucket = "expired" ) // BucketState is what is known about one bucket before it is acted on. type BucketState struct { Bucket Bucket - Policy LatePolicy } -func (b BucketState) String() string { return fmt.Sprintf("bucket=%s policy=%s", b.Bucket, b.Policy) } +func (b BucketState) String() string { return fmt.Sprintf("bucket=%s", b.Bucket) } // BucketAction is what happens to a bucket's rows. type BucketAction string @@ -84,13 +92,12 @@ type BucketAction string const ( // Keep leaves the rows in the table. Keep BucketAction = "keep" - // Close publishes emit_sql over the bucket and deletes its rows. - Close BucketAction = "close" - // DropLate deletes the rows and counts them as late. - DropLate BucketAction = "late.drop" - // ReemitLate publishes emit_sql over the late rows alone, deletes them - // and counts them as late. - ReemitLate BucketAction = "late.reemit" + // Publish runs emit_sql over the bucket and hands the result to the sink. + Publish BucketAction = "publish" + // Retain leaves the rows for a late row to republish the bucket whole. + Retain BucketAction = "retain" + // Purge deletes the rows. + Purge BucketAction = "purge" ) // StateOf reduces a poll's readings to the fact the table decides on: @@ -107,18 +114,21 @@ func StateOf(asserted time.Time, hasAsserted bool, closed time.Time, hadClosed b } } -// BucketStateOf reduces one bucket's end to its fact. -func BucketStateOf(decl Declaration, end, previous, next time.Time) BucketState { - b := BucketState{Policy: decl.Late} +// BucketStateOf reduces one bucket's end to its fact. A bucket that closes +// in this pass is due whatever the lateness; the publish rule says what +// then happens to its rows. Retained and expired are the two fates of a +// bucket that closed in an earlier pass. +func BucketStateOf(decl Declaration, end, closed time.Time, hadClosed bool, asserted time.Time) BucketState { switch { - case !end.After(previous): - b.Bucket = BucketLate - case !end.After(next): - b.Bucket = BucketDue + case end.After(asserted): + return BucketState{BucketOpen} + case !hadClosed || end.After(closed): + return BucketState{BucketDue} + case !end.Add(decl.Lateness).After(asserted): + return BucketState{BucketExpired} default: - b.Bucket = BucketOpen + return BucketState{BucketRetained} } - return b } // watermarkRule is one row of the watermark table. An empty Asserted matches @@ -166,15 +176,13 @@ var watermarkTable = []watermarkRule{ type bucketRule struct { Name string Bucket []Bucket - Policy []LatePolicy Action BucketAction Deciding string Claim string } func (r bucketRule) matches(b BucketState) bool { - return (len(r.Bucket) == 0 || contains(r.Bucket, b.Bucket)) && - (len(r.Policy) == 0 || contains(r.Policy, b.Policy)) + return len(r.Bucket) == 0 || contains(r.Bucket, b.Bucket) } // bucketTable is the bucket's truth table. @@ -184,30 +192,28 @@ var bucketTable = []bucketRule{ Bucket: []Bucket{BucketOpen}, Action: Keep, Deciding: "bucket", - Claim: "The bucket ends after the watermark, so its rows stay.", + Claim: "The bucket ends after the asserted watermark, so its rows stay.", }, { - Name: "close", + Name: "publish", Bucket: []Bucket{BucketDue}, - Action: Close, + Action: Publish, Deciding: "bucket", - Claim: "The bucket ends between the previous watermark and this one: emit_sql runs over its rows and they are deleted.", + Claim: "The bucket ends between the closed watermark and the asserted one: emit_sql runs over its rows and the result is published. With allowed_lateness_seconds the rows stay for a late row to republish it whole; without, this pass purges them too.", }, { - Name: "late.drop", - Bucket: []Bucket{BucketLate}, - Policy: []LatePolicy{LateDrop}, - Action: DropLate, - Deciding: "policy", - Claim: "The bucket closed before these rows arrived, and late_rows is drop: they are deleted and counted.", + Name: "retain", + Bucket: []Bucket{BucketRetained}, + Action: Retain, + Deciding: "bucket", + Claim: "The bucket has been published and its lateness has not run out: its rows stay, and a late row the engine admits republishes it whole.", }, { - Name: "late.reemit", - Bucket: []Bucket{BucketLate}, - Policy: []LatePolicy{LateReemit}, - Action: ReemitLate, - Deciding: "policy", - Claim: "The bucket closed before these rows arrived, and late_rows is reemit: emit_sql runs over the late rows alone, they are deleted and counted.", + Name: "purge", + Bucket: []Bucket{BucketExpired}, + Action: Purge, + Deciding: "bucket", + Claim: "The bucket ended at or before the asserted watermark less the lateness: nothing more can arrive for it, and its rows are deleted.", }, } @@ -259,8 +265,7 @@ func contains[T comparable](vals []T, v T) bool { // Every value of every fact, for the exhaustiveness check and the rendering. var ( assertedValues = []Asserted{AssertedNone, AssertedBehind, AssertedAhead} - bucketValues = []Bucket{BucketLate, BucketDue, BucketOpen} - policyValues = []LatePolicy{LateDrop, LateReemit} + bucketValues = []Bucket{BucketOpen, BucketDue, BucketRetained, BucketExpired} ) func allStates() []State { @@ -274,9 +279,7 @@ func allStates() []State { func allBucketStates() []BucketState { var out []BucketState for _, b := range bucketValues { - for _, p := range policyValues { - out = append(out, BucketState{Bucket: b, Policy: p}) - } + out = append(out, BucketState{Bucket: b}) } return out } diff --git a/internal/managers/decide_test.go b/internal/managers/decide_test.go index 66d503ba..6e9abf61 100644 --- a/internal/managers/decide_test.go +++ b/internal/managers/decide_test.go @@ -41,6 +41,19 @@ func TestManagerWindow_TheTableCheckerCatchesEveryFault(t *testing.T) { err = checkTables() assert.Error(t, err) assert.That(t, strings.Contains(err.Error(), "row unreachable is selected by no combination")) + + // The same three faults in the bucket table. + watermarkTable = saved + savedBuckets := bucketTable + defer func() { bucketTable = savedBuckets }() + bucketTable = savedBuckets[1:] + err = checkTables() + assert.Error(t, err) + assert.That(t, strings.Contains(err.Error(), "bucket=open matches no row")) + bucketTable = append(append([]bucketRule{}, savedBuckets...), savedBuckets[0]) + err = checkTables() + assert.Error(t, err) + assert.That(t, strings.Contains(err.Error(), "matches keep and keep")) } // One example per row, each run from raw readings through StateOf, so the @@ -74,22 +87,66 @@ func TestManagerWindow_EveryRuleHasAnExample(t *testing.T) { } } +// One example per bucket row, over (end, closed, hadClosed, asserted, +// lateness): the four fates of a bucket, with and without lateness. +func TestManagerWindow_EveryBucketRuleHasAnExample(t *testing.T) { + coverage.Covers(t, "manager.window") + closed := t0.Add(10 * time.Minute) + asserted := closed.Add(5 * time.Minute) + minute := time.Minute + + examples := []struct { + name string + end time.Time + hadClosed bool + lateness time.Duration + bucket Bucket + action BucketAction + }{ + {"ends after the assertion", asserted.Add(minute), true, 0, BucketOpen, Keep}, + {"ends after the assertion, never closed", asserted.Add(minute), false, time.Hour, BucketOpen, Keep}, + {"ends between the two watermarks", closed.Add(minute), true, 0, BucketDue, Publish}, + {"ends between the two watermarks, with lateness", closed.Add(minute), true, time.Hour, BucketDue, Publish}, + {"ends at the assertion", asserted, true, 0, BucketDue, Publish}, + {"never closed, ends before the assertion", closed.Add(-minute), false, 0, BucketDue, Publish}, + {"closed earlier, lateness has not run out", closed, true, time.Hour, BucketRetained, Retain}, + {"closed earlier, lateness ran out this pass", closed, true, 5 * minute, BucketExpired, Purge}, + {"closed earlier, no lateness", closed, true, 0, BucketExpired, Purge}, + } + for _, ex := range examples { + t.Run(ex.name, func(t *testing.T) { + decl := testDecl() + decl.Lateness = ex.lateness + state := BucketStateOf(decl, ex.end, closed, ex.hadClosed, asserted) + assert.Equal(t, ex.bucket, state.Bucket) + assert.Equal(t, ex.action, DecideBucket(state)) + }) + } +} + // The boundaries compare the way the SQL always has: an assertion exactly at // the closed watermark is behind, a microsecond past it is ahead; a bucket -// that ends exactly at the watermark is closed. +// that ends exactly at the watermark is closed, and one whose end plus +// lateness is exactly the assertion has expired. func TestManagerWindow_TheBoundariesAreExact(t *testing.T) { coverage.Covers(t, "manager.window") decl := testDecl() + decl.Lateness = 5 * time.Minute closed := t0.Add(10 * time.Minute) assert.Equal(t, AssertedBehind, StateOf(closed, true, closed, true).Asserted) assert.Equal(t, AssertedAhead, StateOf(closed.Add(time.Microsecond), true, closed, true).Asserted) - end := closed.Add(time.Hour) - assert.Equal(t, BucketLate, BucketStateOf(decl, closed, closed, end).Bucket) - assert.Equal(t, BucketDue, BucketStateOf(decl, closed.Add(time.Microsecond), closed, end).Bucket) - assert.Equal(t, BucketDue, BucketStateOf(decl, end, closed, end).Bucket) - assert.Equal(t, BucketOpen, BucketStateOf(decl, end.Add(time.Microsecond), closed, end).Bucket) + asserted := closed.Add(time.Hour) + assert.Equal(t, BucketExpired, BucketStateOf(decl, closed, closed, true, asserted).Bucket) + assert.Equal(t, BucketDue, BucketStateOf(decl, closed.Add(time.Microsecond), closed, true, asserted).Bucket) + assert.Equal(t, BucketDue, BucketStateOf(decl, asserted, closed, true, asserted).Bucket) + assert.Equal(t, BucketOpen, BucketStateOf(decl, asserted.Add(time.Microsecond), closed, true, asserted).Bucket) + + // Expiry: end + lateness at the assertion is expired; a microsecond + // short of it is retained. + assert.Equal(t, BucketExpired, BucketStateOf(decl, closed, closed, true, closed.Add(decl.Lateness)).Bucket) + assert.Equal(t, BucketRetained, BucketStateOf(decl, closed, closed, true, closed.Add(decl.Lateness-time.Microsecond)).Bucket) } // window.never_backwards: no reading moves the closed watermark behind the @@ -148,7 +205,7 @@ func renderDecisions() string { b.WriteString(`# Window decisions Rendered from the truth tables in ` + "`internal/managers/decide.go`" + ` by a test, so -this page is what runs. A poll reduces what it read to one value per fact, +this page is what runs. A pass reduces what it read to one value per fact, looks the combination up, and performs the action of the one row it selects. The check that runs when the package loads proves each combination selects exactly one row and every row is reachable. @@ -159,10 +216,13 @@ ever write for the window has ` + "`time_column`" + ` at or after the watermark. is computed -- the newest event time seen per source partition, less the grace, combined by minimum over the partitions that could still deliver -- is ` + "`internal/core/watermarks.go`" + `. The manager reads that one value and -nothing else: no clock, no progress row, no reading of the table's newest -bucket. See ` + "`docs/superpowers/specs/2026-09-24-window-watermark-design.md`" + `. +nothing else: no clock, no ticker, no progress row, no reading of the +table's newest bucket. It runs a pass when it starts, when the engine kicks +it after a commit that moved the watermark or admitted a late row, and when +it drains. See ` + "`docs/superpowers/specs/2026-09-24-window-watermark-design.md`" + ` +and ` + "`docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md`" + `. -## The watermark, once per poll +## The watermark, once per pass | Fact | Values | Computed from | |---|---|---| @@ -178,29 +238,38 @@ bucket. See ` + "`docs/superpowers/specs/2026-09-24-window-watermark-design.md`" } b.WriteString(` -## Each bucket, given the poll's watermark +## Each bucket, given the pass's watermark | Fact | Values | Computed from | |---|---|---| -| ` + "`bucket`" + ` | ` + "`late`, `due`, `open`" + ` | The bucket's end against the previous watermark and the one this poll decided. ` + "`late`" + `: at or before the previous, so its rows arrived after it closed. ` + "`due`" + `: after the previous and at or before the next. ` + "`open`" + `: after the next. | -| ` + "`policy`" + ` | ` + "`drop`, `reemit`" + ` | ` + "`late_rows`" + ` in the window's config. | +| ` + "`bucket`" + ` | ` + "`open`, `due`, `retained`, `expired`" + ` | The bucket's end against the closed watermark, the asserted one, and ` + "`allowed_lateness_seconds`" + `. ` + "`open`" + `: ends after the asserted watermark. ` + "`due`" + `: ends after the closed watermark and at or before the asserted one, so it closes in this pass. ` + "`retained`" + `: closed in an earlier pass and its end plus the lateness is still past the assertion. ` + "`expired`" + `: closed in an earlier pass and its end plus the lateness is at or before the assertion. | `) fmt.Fprintf(&b, "%d combinations, %d rows.\n\n", len(allBucketStates()), len(bucketTable)) - b.WriteString("| Rule | bucket | policy | Action | Deciding | Claim |\n|---|---|---|---|---|---|\n") + b.WriteString("| Rule | bucket | Action | Deciding | Claim |\n|---|---|---|---|---|\n") for _, r := range bucketTable { - fmt.Fprintf(&b, "| `%s` | %s | %s | `%s` | %s | %s |\n", - r.Name, joinValues(r.Bucket), joinValues(r.Policy), r.Action, r.Deciding, r.Claim) + fmt.Fprintf(&b, "| `%s` | %s | `%s` | %s | %s |\n", + r.Name, joinValues(r.Bucket), r.Action, r.Deciding, r.Claim) } b.WriteString(` +The engine decides a record's lateness at arrival by the same comparison, +before the handler sees it: a record whose bucket ended at or before the +asserted watermark less the lateness is refused and counted; one whose +bucket ended at or before the watermark but within the lateness is written, +and the bucket is republished whole on the next pass; anything else is on +time. So the window table never holds a row the engine did not admit, and +the manager never sweeps. + ## What closes a bucket, by configuration Three keys decide. ` + "`grace_seconds`" + ` is how far the stream's own clock must pass a bucket's end before it closes (Flink's bounded out-of-orderness); ` + "`idle_close_seconds`" + ` is how long a partition may be silent before it stops -holding the window open (Flink's idleness); ` + "`late_rows`" + ` is what happens to a -row for a bucket that already closed. +holding the window open (Flink's idleness); ` + "`allowed_lateness_seconds`" + ` is how +long after a bucket closes a record for it is still admitted, and its rows +kept, so that the bucket can be republished whole (Flink's allowed +lateness). | | ` + "`grace_seconds`" + ` | ` + "`idle_close_seconds`" + ` | the stream this is for | what closes a bucket | the failure mode to know | |---|---|---|---|---|---| @@ -209,9 +278,9 @@ row for a bucket that already closed. | **C** | 0 | set | ordered, intermittent | event time, or every partition silent for the bound | a bound shorter than the gap *within* a burst closes mid-burst, and the rest of the burst is late | | **D** | >0 | set | bursty and reordered: the IoT default | either | both of the above | -Each crossed with ` + "`drop`" + ` (a late row is discarded and counted) or ` + "`reemit`" + ` -(a late row is republished on its own, which a replacing sink turns into -loss). +Each crossed with ` + "`allowed_lateness_seconds`" + `: 0 refuses a late record before +the handler and counts it; a positive value admits it and republishes its +bucket as a whole value, so the sink must replace by key. ## Invariants @@ -222,7 +291,8 @@ Each is a claim with a check in the tests. | ` + "`window.never_backwards`" + ` | No reading moves the closed watermark behind the committed one, checked over ten thousand random readings here; and the engine's assertion is ` + "`max(stored, W)`" + ` by construction, checked in ` + "`internal/core`" + `. | | ` + "`window.asserted_in_the_commit`" + ` | The engine writes the watermark in the transaction that commits the rows it describes, so no reader can see rows without the watermark that accounts for them, or a watermark without its rows. A commit that fails leaves both where they were. Checked in ` + "`internal/core`" + `. | | ` + "`window.minimum_over_partitions`" + ` | The watermark is the minimum over the partitions that could still deliver: one that has not delivered holds it at -inf, a lagging one holds it, an idle one leaves it, a lost one holds at its last position, a revoked one is gone. Checked in ` + "`internal/core`" + ` and by the simulator. | -| ` + "`window.no_clock_in_the_manager`" + ` | The manager has no clock. Every close is ` + "`bucket end <= asserted watermark`" + `, in event time; the engine's clock decides only which partitions are in its minimum. Checked by construction: the manager takes no clock. | +| ` + "`window.no_clock_in_the_manager`" + ` | The manager has no clock. Every close is ` + "`bucket end <= asserted watermark`" + `, in event time; the engine's clock decides only which partitions are in its minimum. Checked by construction: the manager takes no clock and no interval. | +| ` + "`window.lateness_decided_at_arrival`" + ` | A record whose bucket ended at or before the watermark less the lateness is refused before the handler and counted; one within the lateness is written and its bucket republished whole. The window table never holds a row the engine did not admit. Checked in ` + "`internal/core`" + `, by the model and by the simulator. | `) return b.String() } @@ -238,37 +308,61 @@ func joinValues[T ~string](vals []T) string { return strings.Join(parts, ", ") } -// The bucket table's boundary is the SQL's. Production splits late from -// due with closedBefore, and BucketStateOf is the same comparison in Go; -// this runs both at the boundary and a microsecond past it so a change to -// either fails here. +// The bucket table's boundary is the SQL's. Production selects due buckets +// with dueBetween and expired ones with expiredBefore, and BucketStateOf is +// the same comparison in Go; this runs both at each boundary and a +// microsecond either side so a change to either fails here. func TestManagerWindow_TheBucketBoundaryIsTheSQLs(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() d := newTestDB(t, "") createWindowTable(t, d.pipeline) decl := testDecl() + decl.Lateness = 5 * time.Minute // One bucket, starting at t0 and ending at t0 + size. insertBucket(t, d.pipeline, 0, "NYC", 1) end := bucket(0).Add(decl.Size) - next := end.Add(time.Hour) + // Due: the closed watermark sits a minute before the bucket's end, and + // the assertion moves across the end. + closed := end.Add(-time.Minute) for _, c := range []struct { - name string - watermark time.Time - want Bucket + name string + asserted time.Time + want Bucket + }{ + {"asserted a microsecond before the end", end.Add(-time.Microsecond), BucketOpen}, + {"asserted at the end", end, BucketDue}, + {"asserted a microsecond past the end", end.Add(time.Microsecond), BucketDue}, + } { + t.Run(c.name, func(t *testing.T) { + due, _, err := queryInt64(ctx, d.pipeline, decl.countSQL(decl.dueBetween(closed, true, c.asserted))) + assert.NoError(t, err) + got := BucketStateOf(decl, end, closed, true, c.asserted).Bucket + assert.Equal(t, c.want, got) + assert.Equal(t, got == BucketDue, due == 1) + }) + } + + // Expired: the bucket closed in an earlier pass, and the assertion moves + // across end + lateness. + expiry := end.Add(decl.Lateness) + for _, c := range []struct { + name string + asserted time.Time + want Bucket }{ - {"watermark a microsecond before the end", end.Add(-time.Microsecond), BucketDue}, - {"watermark at the end", end, BucketLate}, - {"watermark a microsecond past the end", end.Add(time.Microsecond), BucketLate}, + {"asserted a microsecond before the expiry", expiry.Add(-time.Microsecond), BucketRetained}, + {"asserted at the expiry", expiry, BucketExpired}, + {"asserted a microsecond past the expiry", expiry.Add(time.Microsecond), BucketExpired}, } { t.Run(c.name, func(t *testing.T) { - closed, _, err := queryInt64(ctx, d.pipeline, decl.countClosedSQL(c.watermark)) + expired, _, err := queryInt64(ctx, d.pipeline, decl.countSQL(decl.expiredBefore(c.asserted))) assert.NoError(t, err) - got := BucketStateOf(decl, end, c.watermark, next).Bucket + got := BucketStateOf(decl, end, end, true, c.asserted).Bucket assert.Equal(t, c.want, got) - assert.Equal(t, got == BucketLate, closed == 1) + assert.Equal(t, got == BucketExpired, expired == 1) }) } } diff --git a/internal/managers/helpers_test.go b/internal/managers/helpers_test.go index c3ac62fc..03c18d58 100644 --- a/internal/managers/helpers_test.go +++ b/internal/managers/helpers_test.go @@ -181,7 +181,6 @@ func testDecl() Declaration { Size: time.Minute, Grace: time.Minute, IdleClose: 5 * time.Minute, - Late: LateReemit, } } @@ -215,21 +214,36 @@ func assertAt(tb testing.TB, conn adbc.Connection, at time.Time) { } } -// newTestWatermark builds the manager on the manager connection. The -// engine's watermark table exists from here, empty until assertAt. -func newTestWatermark(tb testing.TB, d *testDB, decl Declaration, sink interface { - WriteTable(context.Context, arrow.Table) error - Flush(context.Context) error -}, opts ...Option) *Watermark { +// newTestWatermark builds the manager on the manager connection, with the +// signal an engine tracker would hand it, so a test can kick it or queue a +// recompute the way the engine does. The engine's watermark table exists +// from here, empty until assertAt. +func newTestWatermark(tb testing.TB, d *testDB, decl Declaration, sink core.Sink, opts ...Option) (*Watermark, *core.WindowSignal) { tb.Helper() if err := core.NewWatermarkStore(d.pipeline).Init(context.Background()); err != nil { tb.Fatal(err) } - w, err := NewWatermark(d.manager, decl, time.Hour, sink, opts...) + tracker := core.NewWatermarks([]core.WindowSpec{{ + Name: decl.Table, Size: decl.Size, Grace: decl.Grace, IdleClose: decl.IdleClose, Lateness: decl.Lateness, + }}, time.Now) + sig := tracker.Signal(decl.Table) + w, err := NewWatermark(d.manager, decl, sink, sig, opts...) if err != nil { tb.Fatal(err) } - return w + return w, sig +} + +// waitFor polls cond until it holds or the timeout passes. +func waitFor(tb testing.TB, what string, timeout time.Duration, cond func() bool) { + tb.Helper() + deadline := time.Now().Add(timeout) + for !cond() { + if time.Now().After(deadline) { + tb.Fatalf("%s did not happen within %s", what, timeout) + } + time.Sleep(5 * time.Millisecond) + } } var _ = array.NewInt64Builder diff --git a/internal/managers/model.go b/internal/managers/model.go index 597d1a29..5863edd5 100644 --- a/internal/managers/model.go +++ b/internal/managers/model.go @@ -8,23 +8,27 @@ import ( // The model is the window over a sequence of events, with no database under // it. Its engine half is core.Watermarks itself -- the real computation, -// driven on a clock the model owns, not a copy of it -- and its manager half -// is the table in decide.go. checkTables proves every state selects one -// rule; this proves that a run of events ends with every row published once -// or dropped under the declared policy, which is the property no invariant -// owned while a reproduction of #183 lost 12% of its records. +// driven on a clock the model owns, not a copy of it, deciding each row's +// lateness at arrival as the engine does -- and its manager half is the +// table in decide.go. checkTables proves every state selects one rule; this +// proves that a run of events ends with the sink holding, for every bucket, +// the exact count of the rows the engine admitted to it: never a delta, +// never a stale first publish, never a row the engine refused. That is the +// property no invariant owned while a reproduction of #183 lost 12% of its +// records. // // One clock. Event time is what a producer stamps, and here producers stamp // the model's clock, so a row's bucket is where the clock stood when it was // produced -- except a row produced Ahead, which is stamped beyond the grace -// and is what moves the stream on. The engine's clock, the same one, decides -// only which partitions are idle. +// and is what moves the stream on, and a row produced Late or VeryLate, +// which is stamped behind the closed watermark. The engine's clock, the same +// one, decides only which partitions are idle. // EventKind is one thing that can happen to a pipeline. type EventKind string const ( - // Produce is rows from one partition reaching the handler and committing, + // Produce is rows from one partition reaching the engine and committing, // with the watermark asserted in that commit. A partition that is lost // delivers nothing; one that was revoked delivers to another worker. Produce EventKind = "produce" @@ -38,27 +42,37 @@ const ( // Revoke is another worker holding the partition from now on. Revoke EventKind = "revoke" // Restart is the process dying and coming back with its table and both - // watermarks, and nothing in memory. + // watermarks, and nothing in memory: the recompute set is gone, and the + // start pass republishes every retained bucket for it. Restart EventKind = "restart" ) -// Event is one step of a sequence. A poll follows every step. +// Event is one step of a sequence. A pass follows every step, as the kick +// after a commit does. type Event struct { Kind EventKind Partition int32 Rows int // Ahead stamps the rows beyond the grace, so the stream moves on. Ahead bool - By time.Duration + // Late stamps the rows in the bucket just below the closed watermark: + // late, and within the lateness if the window has any. + Late bool + // VeryLate stamps the rows below the closed watermark less the lateness + // and a bucket: late beyond any lateness, so the engine refuses them. + VeryLate bool + By time.Duration } -// Publication is what a close handed the sink. +// Publication is what a close handed the sink for a bucket closing for the +// first time. Recomputes are counted, not returned: they republish a bucket +// on purpose. type Publication struct { Bucket time.Time Rows int } -// Decision is one poll: the rule the one fact selected, and the fact. +// Decision is one pass: the rule the one fact selected, and the fact. type Decision struct { Rule string Action Action @@ -76,7 +90,7 @@ const ( ) // Model is one pipeline: its partitions, its buckets, the engine's tracker -// over both, and the two watermarks. +// over both, the two watermarks, and what the sink holds. type Model struct { decl Declaration spec core.WindowSpec @@ -87,7 +101,7 @@ type Model struct { // The newest event time each partition delivered, for the property that // an in-order row is never late. newest map[int32]time.Time - // Whether each partition was idle at the last poll, for the same + // Whether each partition was idle at the last pass, for the same // property: a partition that went idle may find its bucket closed. // Idleness is measured as the engine measures it, from the later of the // partition's last row and its assignment. @@ -100,13 +114,29 @@ type Model struct { lastRow map[int32]time.Time heldSince map[int32]time.Time + // buckets is the window table: the rows the engine admitted, per bucket, + // including buckets published and retained under the lateness. buckets map[time.Time]int - // Produced counts rows the source actually delivered. Dropped counts - // rows that arrived for a bucket the window had already closed, which - // late_rows drop discards and the property excludes. Late counts the - // rows that were dropped while in order for their partition, and not - // idle: what the minimum exists to prevent. - Produced, Dropped, EarlyClosed int + // recompute is the signal's set: buckets a late row landed in since the + // last pass. In memory only, so a Restart loses it. + recompute map[time.Time]bool + // restarted marks the pass after a Restart as the start pass, which + // republishes every retained bucket. + restarted bool + + // Produced counts rows the source delivered, and ProducedPerBucket the + // same per bucket. Refused counts rows the engine refused at arrival -- + // their bucket had closed more than the lateness before the watermark -- + // and RefusedPerBucket the same per bucket. A refused row is never in + // buckets. + Produced, Refused int + ProducedPerBucket map[time.Time]int + RefusedPerBucket map[time.Time]int + // Last is what the sink holds for each bucket: the value of the last + // publication, as a sink that replaces by key holds it. + Last map[time.Time]int + // Recomputes counts buckets republished whole for a late row. + Recomputes int asserted time.Time hasAsserted bool @@ -116,7 +146,7 @@ type Model struct { // IdleCloses counts assertions made by an idle tick with nothing // arriving: the all-idle close. IdleCloses int - // Decisions is every poll's rule and the fact it read, in order. + // Decisions is every pass's rule and the fact it read, in order. Decisions []Decision } @@ -126,16 +156,22 @@ var modelEpoch = time.Date(2026, 9, 23, 12, 0, 0, 0, time.UTC) // window and nothing asserted. func NewModel(decl Declaration, partitions ...int32) *Model { m := &Model{ - decl: decl, - spec: core.WindowSpec{Name: decl.Table, Size: decl.Size, Grace: decl.Grace, IdleClose: decl.IdleClose}, - clock: modelEpoch, - parts: map[int32]held{}, - newest: map[int32]time.Time{}, - idleAtPoll: map[int32]bool{}, - behind: map[int32]bool{}, - lastRow: map[int32]time.Time{}, - heldSince: map[int32]time.Time{}, - buckets: map[time.Time]int{}, + decl: decl, + spec: core.WindowSpec{ + Name: decl.Table, Size: decl.Size, Grace: decl.Grace, IdleClose: decl.IdleClose, Lateness: decl.Lateness, + }, + clock: modelEpoch, + parts: map[int32]held{}, + newest: map[int32]time.Time{}, + idleAtPoll: map[int32]bool{}, + behind: map[int32]bool{}, + lastRow: map[int32]time.Time{}, + heldSince: map[int32]time.Time{}, + buckets: map[time.Time]int{}, + recompute: map[time.Time]bool{}, + ProducedPerBucket: map[time.Time]int{}, + RefusedPerBucket: map[time.Time]int{}, + Last: map[time.Time]int{}, } m.engine = core.NewWatermarks([]core.WindowSpec{m.spec}, m.now) for _, p := range partitions { @@ -161,8 +197,11 @@ func (m *Model) commit() (moved bool) { return moved } -// Apply advances the model by one event and returns what the poll after it -// published. +// bucketOf is the bucket a row stamped at falls in, as time_bucket cuts it. +func (m *Model) bucketOf(at time.Time) time.Time { return core.BucketStart(at, m.decl.Size) } + +// Apply advances the model by one event and returns what the pass after it +// published for the first time. func (m *Model) Apply(e Event) []Publication { // Every event takes a second, so a sequence spans time without the // caller saying so. @@ -174,11 +213,34 @@ func (m *Model) Apply(e Event) []Publication { break } at := m.clock - if e.Ahead { + switch { + case e.Ahead: at = at.Add(m.decl.Grace + 2*m.decl.Size) + case e.Late && m.hadClosed: + // The bucket before the one the closed watermark stands in: it + // ended at or before the watermark, so it is late. + at = m.bucketOf(m.closed).Add(-time.Second) + case e.VeryLate && m.hadClosed: + // A bucket that ended more than the lateness before the closed + // watermark: late beyond what any lateness admits. + at = m.bucketOf(m.closed).Add(-m.decl.Lateness - m.decl.Size - time.Second) } + b := m.bucketOf(at) m.Produced += e.Rows - m.buckets[at.Truncate(m.decl.Size)] += e.Rows + m.ProducedPerBucket[b] += e.Rows + + // The engine's decision at arrival, before the handler: refused + // rows reach neither the table nor the tracker. + refused, recompute := m.engine.Classify(at.UnixNano()) + if refused { + m.Refused += e.Rows + m.RefusedPerBucket[b] += e.Rows + break + } + m.buckets[b] += e.Rows + for _, lb := range recompute { + m.recompute[lb.Bucket] = true + } m.engine.Observe("t", e.Partition, at.UnixNano()) m.lastRow[e.Partition] = m.clock if at.After(m.newest[e.Partition]) { @@ -208,7 +270,8 @@ func (m *Model) Apply(e Event) []Publication { // Nothing in memory survives: a new tracker, restored from the table // and the row, and the group assigns every partition this worker is // still a member for. A lost session is over; a revoked partition is - // another worker's. + // another worker's. The recompute set is gone with the process, and + // the pass that follows is the start pass. m.engine = core.NewWatermarks([]core.WindowSpec{m.spec}, m.now) var newest time.Time for b := range m.buckets { @@ -218,6 +281,8 @@ func (m *Model) Apply(e Event) []Publication { } m.engine.Restore(m.spec.Name, newest, m.asserted) m.lastRow = map[int32]time.Time{} + m.recompute = map[time.Time]bool{} + m.restarted = true for p, h := range m.parts { if h != revoked { m.assign(p) @@ -229,20 +294,10 @@ func (m *Model) Apply(e Event) []Publication { return m.poll() } -// poll is one manager pass: reduce to the fact, decide, act. +// poll is one manager pass: reduce to the fact, decide, act. Publish what is +// due, republish what a late row landed in (and, on the start pass, every +// retained bucket), purge what is past its lateness. func (m *Model) poll() []Publication { - // Late rows first, against the closed watermark, as the manager does. - // Under drop they are deleted and counted; an in-order row from a - // partition that was in the minimum should never be among them. - if m.hadClosed { - for b, rows := range m.buckets { - if !b.Add(m.decl.Size).After(m.closed) { - m.Dropped += rows - delete(m.buckets, b) - } - } - } - state := StateOf(m.asserted, m.hasAsserted, m.closed, m.hadClosed) rule := watermarkRuleFor(state) m.Decisions = append(m.Decisions, Decision{ @@ -250,14 +305,56 @@ func (m *Model) poll() []Publication { }) next, moved := rule.Action.Next(m.asserted, m.closed) m.noteIdle() - if !moved { + + republish := m.restarted && m.hadClosed && m.decl.Lateness > 0 + m.restarted = false + if !moved && len(m.recompute) == 0 && !republish { return nil } - m.closed, m.hadClosed = next, true - return m.collect() + watermark := m.closed + if moved { + watermark = next + } + + var out []Publication + if moved { + // Due: ended after the closed watermark, at or before the new one. + for b, rows := range m.buckets { + end := b.Add(m.decl.Size) + if (!m.hadClosed || end.After(m.closed)) && !end.After(watermark) { + out = append(out, Publication{Bucket: b, Rows: rows}) + m.Last[b] = rows + } + } + } + // Recomputes: the whole bucket, again. + for b := range m.recompute { + if rows, ok := m.buckets[b]; ok { + m.Last[b] = rows + m.Recomputes++ + } + } + m.recompute = map[time.Time]bool{} + // The start pass: every bucket closed earlier and not yet expired. + if republish { + for b, rows := range m.buckets { + end := b.Add(m.decl.Size) + if !end.After(m.closed) && end.Add(m.decl.Lateness).After(watermark) { + m.Last[b] = rows + } + } + } + // Purge: past the lateness, nothing more can arrive. + for b := range m.buckets { + if !b.Add(m.decl.Size).Add(m.decl.Lateness).After(watermark) { + delete(m.buckets, b) + } + } + m.closed, m.hadClosed = watermark, true + return out } -// noteIdle records which partitions are idle as of this poll, for the +// noteIdle records which partitions are idle as of this pass, for the // in-order property: a partition that has been silent for the bound has // left the minimum, and a row it delivers afterwards may be late by // design. @@ -279,22 +376,10 @@ func (m *Model) noteIdle() { } } -// collect publishes and deletes every bucket the watermark has passed. -func (m *Model) collect() []Publication { - var out []Publication - for b, rows := range m.buckets { - if !b.Add(m.decl.Size).After(m.closed) { - out = append(out, Publication{Bucket: b, Rows: rows}) - delete(m.buckets, b) - } - } - return out -} - // InOrderRowWouldBeLate reports whether a row the partition would deliver // now, in order and stamped with the clock, lands in a bucket the window // has already closed -- while the partition holds, was not idle at the last -// poll, and is not still behind from an idleness it has not caught up from. +// pass, and is not still behind from an idleness it has not caught up from. // That is a close the minimum should have prevented. func (m *Model) InOrderRowWouldBeLate(p int32) bool { if m.parts[p] != holding || m.idleAtPoll[p] || m.behind[p] || !m.hadClosed { @@ -304,15 +389,20 @@ func (m *Model) InOrderRowWouldBeLate(p int32) bool { if at.Before(m.newest[p]) { return false // out of order for its own partition: may be late } - return !at.Truncate(m.decl.Size).Add(m.decl.Size).After(m.closed) + return !m.bucketOf(at).Add(m.decl.Size).After(m.closed) } -// Open is how many rows the window still holds, and the newest bucket's end -// among them. +// Open is how many rows the window holds in buckets not yet published, and +// the newest bucket's end among them. Retained buckets are published, and +// not open. func (m *Model) Open() (rows int, newestEnd time.Time) { for b, n := range m.buckets { + end := b.Add(m.decl.Size) + if m.hadClosed && !end.After(m.closed) { + continue + } rows += n - if end := b.Add(m.decl.Size); end.After(newestEnd) { + if end.After(newestEnd) { newestEnd = end } } @@ -322,7 +412,15 @@ func (m *Model) Open() (rows int, newestEnd time.Time) { // Drain publishes what is still open, as a close that finally comes would, // and reports the rows it carried. func (m *Model) Drain() int { - rows, _ := m.Open() + rows := 0 + for b, n := range m.buckets { + end := b.Add(m.decl.Size) + if m.hadClosed && !end.After(m.closed) { + continue + } + rows += n + m.Last[b] = n + } m.buckets = map[time.Time]int{} return rows } diff --git a/internal/managers/model_test.go b/internal/managers/model_test.go index ea7b93d2..e2121f5a 100644 --- a/internal/managers/model_test.go +++ b/internal/managers/model_test.go @@ -11,8 +11,9 @@ import ( // modelAlphabet is every event that has produced a defect in this package // and that the engine is meant to survive: rows from each of two partitions, -// rows that move the stream on, silence past the idle bound, a partition -// lost and assigned again, one revoked for good, and a restart. +// rows that move the stream on, a row for a bucket that closed and one for a +// bucket long expired, silence past the idle bound, a partition lost and +// assigned again, one revoked for good, and a restart. func modelAlphabet(idleClose time.Duration) []Event { elapse := 3 * time.Second if idleClose > 0 { @@ -22,6 +23,8 @@ func modelAlphabet(idleClose time.Duration) []Event { {Kind: Produce, Partition: 0, Rows: 2}, {Kind: Produce, Partition: 0, Rows: 2, Ahead: true}, {Kind: Produce, Partition: 1, Rows: 2}, + {Kind: Produce, Partition: 0, Rows: 1, Late: true}, + {Kind: Produce, Partition: 0, Rows: 1, VeryLate: true}, {Kind: Elapse, By: elapse}, {Kind: Lose, Partition: 0}, {Kind: Assign, Partition: 0}, @@ -30,19 +33,33 @@ func modelAlphabet(idleClose time.Duration) []Event { } } -// shape is one of the four configurations the spec works backwards from. +// shape is one of the four configurations the spec works backwards from, +// each with and without lateness. type shape struct { - name string - grace time.Duration - idle time.Duration + name string + grace time.Duration + idle time.Duration + lateness time.Duration } -var shapes = []shape{ - {"A: grace 0, no idle bound", 0, 0}, - {"B: grace, no idle bound", time.Minute, 0}, - {"C: grace 0, idle bound", 0, 2 * time.Second}, - {"D: grace and idle bound, the IoT default", time.Minute, 2 * time.Second}, -} +var shapes = func() []shape { + base := []shape{ + {"A: grace 0, no idle bound", 0, 0, 0}, + {"B: grace, no idle bound", time.Minute, 0, 0}, + {"C: grace 0, idle bound", 0, 2 * time.Second, 0}, + {"D: grace and idle bound, the IoT default", time.Minute, 2 * time.Second, 0}, + } + // The four first, so shapes[1] is B and shapes[3] is D for the tests + // that name them; then each again with a minute of lateness. + out := append([]shape(nil), base...) + for _, s := range base { + late := s + late.name += ", lateness" + late.lateness = time.Minute + out = append(out, late) + } + return out +}() // The idle bound is two seconds because an event takes one: a bound longer // than a sequence can run is a bound no sequence ever reaches, and the first @@ -55,14 +72,15 @@ func (s shape) decl() Declaration { Size: time.Minute, Grace: s.grace, IdleClose: s.idle, - Late: LateDrop, + Lateness: s.lateness, } } // Every sequence of events up to four long, under each configuration shape, -// ends with each row published exactly once or dropped as late, and never -// drops a row that was in order for a partition still in the minimum. Once -// the run quiesces, nothing that can close stays open. +// ends with the sink holding, for every bucket, exactly the rows the engine +// admitted to it -- produced less refused -- and never refuses a row that +// was in order for a partition still in the minimum. Once the run quiesces, +// nothing that can close stays open. func TestManagerWindow_EverySequenceCountsEveryRowOnce(t *testing.T) { coverage.Covers(t, "manager.window") for _, s := range shapes { @@ -89,8 +107,8 @@ func TestManagerWindow_EverySequenceCountsEveryRowOnce(t *testing.T) { } } walk(0) - // 8 + 64 + 512 + 4096 - assert.Equal(t, 4680, runs) + // 10 + 100 + 1000 + 10000 + assert.Equal(t, 11110, runs) // An enumeration that never reaches a rule proves nothing about // it, and says so in no way a reader would notice: the suite is @@ -186,12 +204,32 @@ func checkSequence(t *testing.T, s shape, seq []Event, fired map[string]int) int t.Fatalf("%v: %d rows still open in a bucket ending %s, under the assertion %s", kinds(seq), open, newestEnd, asserted) } - published += m.Drain() + m.Drain() - if m.Produced != published+m.Dropped { - t.Fatalf("%v: produced %d, published %d, dropped %d", - kinds(seq), m.Produced, published, m.Dropped) + // The exact-value property. For every bucket, the sink's last value is + // the rows produced for it less the rows refused for it -- never a delta, + // never a stale first publish. And a refused row is never in the table, + // by construction: the model inserts only what Classify admitted. + for b, produced := range m.ProducedPerBucket { + want := produced - m.RefusedPerBucket[b] + if got := m.Last[b]; got != want { + t.Fatalf("%v: bucket %s: sink holds %d, produced %d, refused %d", + kinds(seq), b.Format("15:04"), got, produced, m.RefusedPerBucket[b]) + } } + refused := 0 + for _, n := range m.RefusedPerBucket { + refused += n + } + if refused != m.Refused { + t.Fatalf("%v: refused %d in total and %d per bucket", kinds(seq), m.Refused, refused) + } + // Without lateness nothing is ever recomputed: a late row is refused + // or the bucket had not closed. + if s.lateness == 0 && m.Recomputes != 0 { + t.Fatalf("%v: %d recomputes under no lateness", kinds(seq), m.Recomputes) + } + _ = published for i, d := range m.Decisions { if i < body { fired[d.Rule]++ @@ -210,10 +248,55 @@ func kinds(seq []Event) []string { if e.Ahead { out[i] += "+" } + if e.Late { + out[i] += "-" + } + if e.VeryLate { + out[i] += "--" + } } return out } +// A late row within the lateness is admitted and republishes its bucket +// whole; beyond it the row is refused and the bucket stays as published. +// The exact-value property, on one sequence a reader can follow. +func TestManagerWindow_ALateRowRecomputesOrIsRefused(t *testing.T) { + coverage.Covers(t, "manager.window") + for _, lateness := range []time.Duration{0, time.Minute} { + decl := shapes[0].decl() + decl.Lateness = lateness + m := NewModel(decl, 0) + m.Apply(Event{Kind: Produce, Partition: 0, Rows: 3}) + first := m.bucketOf(m.clock) + pubs := m.Apply(Event{Kind: Produce, Partition: 0, Rows: 1, Ahead: true}) + assert.Equal(t, 1, len(pubs)) + assert.Equal(t, 3, pubs[0].Rows) + assert.Equal(t, 3, m.Last[first]) + + pubs = m.Apply(Event{Kind: Produce, Partition: 0, Rows: 2, Late: true}) + assert.Equal(t, 0, len(pubs)) + if lateness == 0 { + assert.Equal(t, 2, m.Refused) + assert.Equal(t, 0, m.Recomputes) + } else { + assert.Equal(t, 0, m.Refused) + assert.Equal(t, 1, m.Recomputes) + } + // Whichever closed bucket the late row landed in, the sink holds + // every admitted row for it. + for b, produced := range m.ProducedPerBucket { + if m.hadClosed && !b.Add(decl.Size).After(m.closed) { + assert.Equal(t, produced-m.RefusedPerBucket[b], m.Last[b]) + } + } + + before := m.Refused + m.Apply(Event{Kind: Produce, Partition: 0, Rows: 1, VeryLate: true}) + assert.Equal(t, before+1, m.Refused) + } +} + // A partition racing ahead in event time cannot close the buckets a slower // one is still filling: the minimum holds for the slow one, and its rows are // never late. This is the defect the simulator pinned before the watermark @@ -229,7 +312,7 @@ func TestManagerWindow_AFastPartitionHoldsForTheSlowOne(t *testing.T) { m.Apply(Event{Kind: Produce, Partition: 0, Rows: 4, Ahead: true}) } m.Apply(Event{Kind: Produce, Partition: 1, Rows: 5}) - assert.Equal(t, 0, m.Dropped) + assert.Equal(t, 0, m.Refused) assert.That(t, !m.InOrderRowWouldBeLate(1)) } diff --git a/internal/managers/sql.go b/internal/managers/sql.go index 7cd09084..7cb874ed 100644 --- a/internal/managers/sql.go +++ b/internal/managers/sql.go @@ -13,14 +13,13 @@ import ( // but the watermark appears in any of them. now() is absent on purpose: it // is frozen at the start of an open transaction, which is the bug #158 // fixed and the reason the user's predicates were hard to get right. The -// watermark itself is read through core.LoadWatermark. The // watermark itself is read through core.LoadWatermark. // closedView is the relation emit_sql reads: the rows of every bucket that -// has just closed. It is spliced into emit_sql as a common table expression -// rather than created as a view, because a CREATE, even of a temporary -// view, is a write to a catalog, and the close's transaction may write to -// one database only. +// has just closed, or of the one bucket being republished. It is spliced +// into emit_sql as a common table expression rather than created as a view, +// because a CREATE, even of a temporary view, is a write to a catalog, and +// the pass's transaction may write to one database only. const closedView = "closed" // defaultEmitSQL publishes the closed rows as they are. @@ -38,34 +37,67 @@ func (d Declaration) closedBefore(instant time.Time) string { quoteIdent(d.TimeColumn), int64(d.Size/time.Second), core.UTCLiteral(instant)) } -// newestSQL reads the newest bucket start the table holds, as microseconds -// since the epoch. NULL on an empty table. +// dueBetween is the predicate for buckets that close in this pass: ended +// after the closed watermark and at or before the asserted one. Before the +// first close everything at or before the assertion is due. +func (d Declaration) dueBetween(closed time.Time, hadClosed bool, asserted time.Time) string { + if !hadClosed { + return d.closedBefore(asserted) + } + return fmt.Sprintf("%s + INTERVAL '%d' SECOND > TIMESTAMPTZ '%s' AND %s", + quoteIdent(d.TimeColumn), int64(d.Size/time.Second), core.UTCLiteral(closed), d.closedBefore(asserted)) +} + +// bucketIs is the predicate for one bucket's rows, for a recompute. +func (d Declaration) bucketIs(bucket time.Time) string { + return fmt.Sprintf("%s = TIMESTAMPTZ '%s'", quoteIdent(d.TimeColumn), core.UTCLiteral(bucket)) +} + +// expiredBefore is the predicate for buckets past their lateness: ended at +// or before asserted − lateness. With no lateness that is every bucket the +// pass publishes, so the close deletes what it published, as it always did. +func (d Declaration) expiredBefore(asserted time.Time) string { + return d.closedBefore(asserted.Add(-d.Lateness)) +} + +// retainedBetween is the predicate for buckets that closed in an earlier +// pass and whose lateness has not run out: ended after asserted − lateness +// and at or before the closed watermark. The start pass republishes them, +// because the recompute set a late row landed in before a restart was in +// memory. +func (d Declaration) retainedBetween(closed, asserted time.Time) string { + return fmt.Sprintf("%s + INTERVAL '%d' SECOND > TIMESTAMPTZ '%s' AND %s", + quoteIdent(d.TimeColumn), int64(d.Size/time.Second), core.UTCLiteral(asserted.Add(-d.Lateness)), d.closedBefore(closed)) +} + // oldestSQL reads the oldest bucket the window holds, as microseconds since // the epoch. NULL when the table is empty. func (d Declaration) oldestSQL() string { return fmt.Sprintf("SELECT epoch_us(min(%s)) FROM %s", quoteIdent(d.TimeColumn), quoteIdent(d.Table)) } +// newestSQL reads the newest bucket start the table holds, as microseconds +// since the epoch. NULL on an empty table. func (d Declaration) newestSQL() string { return fmt.Sprintf("SELECT epoch_us(max(%s)) FROM %s", quoteIdent(d.TimeColumn), quoteIdent(d.Table)) } -// countClosedSQL counts the rows a close at the instant would collect. -func (d Declaration) countClosedSQL(instant time.Time) string { - return fmt.Sprintf("SELECT count(*)::BIGINT FROM %s WHERE %s", quoteIdent(d.Table), d.closedBefore(instant)) +// countSQL counts the rows where selects. +func (d Declaration) countSQL(where string) string { + return fmt.Sprintf("SELECT count(*)::BIGINT FROM %s WHERE %s", quoteIdent(d.Table), where) } -// deleteClosedSQL removes the rows a close at the instant collected. -func (d Declaration) deleteClosedSQL(instant time.Time) string { - return fmt.Sprintf("DELETE FROM %s WHERE %s", quoteIdent(d.Table), d.closedBefore(instant)) +// deleteSQL removes the rows where selects. +func (d Declaration) deleteSQL(where string) string { + return fmt.Sprintf("DELETE FROM %s WHERE %s", quoteIdent(d.Table), where) } // collectSQL is what the sink receives: emit_sql with the closed relation -// spliced in front of it as a CTE. An emit_sql that starts with its own WITH -// keeps it; the closed CTE is added to its list. -func (d Declaration) collectSQL(instant time.Time) string { - closed := fmt.Sprintf("%s AS (SELECT * FROM %s WHERE %s)", - closedView, quoteIdent(d.Table), d.closedBefore(instant)) +// spliced in front of it as a CTE, over the rows where selects. An emit_sql +// that starts with its own WITH keeps it; the closed CTE is added to its +// list. +func (d Declaration) collectSQL(where string) string { + closed := fmt.Sprintf("%s AS (SELECT * FROM %s WHERE %s)", closedView, quoteIdent(d.Table), where) emit := strings.TrimSpace(d.EmitSQL) if emit == "" { emit = defaultEmitSQL diff --git a/internal/managers/sql_test.go b/internal/managers/sql_test.go index e40740b5..36364923 100644 --- a/internal/managers/sql_test.go +++ b/internal/managers/sql_test.go @@ -9,8 +9,8 @@ import ( ) // Every statement the manager runs, byte for byte, for one declaration. -// The collect and the delete share one predicate, and no statement reads -// now(): the watermark is the only clock. +// The collect and the delete share one predicate builder, and no statement +// reads now(): the watermark is the only clock. func TestManagerWindow_GeneratedSQL(t *testing.T) { coverage.Covers(t, "manager.window") d := Declaration{ @@ -18,8 +18,9 @@ func TestManagerWindow_GeneratedSQL(t *testing.T) { TimeColumn: "bucket", Size: time.Minute, Grace: 30 * time.Second, - Late: LateDrop, + Lateness: 5 * time.Minute, } + closed := time.Date(2026, 9, 13, 10, 0, 0, 0, time.UTC) at := time.Date(2026, 9, 13, 10, 5, 0, 0, time.UTC) assert.Equal(t, @@ -29,45 +30,70 @@ func TestManagerWindow_GeneratedSQL(t *testing.T) { `SELECT epoch_us(max("bucket")) FROM "posts_per_minute"`, d.newestSQL()) assert.Equal(t, - `SELECT count(*)::BIGINT FROM "posts_per_minute" WHERE "bucket" + INTERVAL '60' SECOND <= TIMESTAMPTZ '2026-09-13 10:05:00+00:00'`, - d.countClosedSQL(at)) + `SELECT epoch_us(min("bucket")) FROM "posts_per_minute"`, + d.oldestSQL()) + + // Due: after the closed watermark, at or before the asserted one. Before + // the first close, everything at or before the assertion. + assert.Equal(t, + `"bucket" + INTERVAL '60' SECOND > TIMESTAMPTZ '2026-09-13 10:00:00+00:00' AND "bucket" + INTERVAL '60' SECOND <= TIMESTAMPTZ '2026-09-13 10:05:00+00:00'`, + d.dueBetween(closed, true, at)) + assert.Equal(t, d.closedBefore(at), d.dueBetween(time.Time{}, false, at)) + // One bucket, for a recompute. + assert.Equal(t, `"bucket" = TIMESTAMPTZ '2026-09-13 10:00:00+00:00'`, d.bucketIs(closed)) + // Expired: ended at or before asserted - lateness. With no lateness, at + // or before the assertion itself. + assert.Equal(t, + `"bucket" + INTERVAL '60' SECOND <= TIMESTAMPTZ '2026-09-13 10:00:00+00:00'`, + d.expiredBefore(at)) + none := d + none.Lateness = 0 + assert.Equal(t, d.closedBefore(at), none.expiredBefore(at)) + // Retained: closed in an earlier pass, ended after asserted - lateness. + assert.Equal(t, + `"bucket" + INTERVAL '60' SECOND > TIMESTAMPTZ '2026-09-13 10:00:00+00:00' AND "bucket" + INTERVAL '60' SECOND <= TIMESTAMPTZ '2026-09-13 10:00:00+00:00'`, + d.retainedBetween(closed, at)) + + where := d.closedBefore(at) + assert.Equal(t, + `SELECT count(*)::BIGINT FROM "posts_per_minute" WHERE `+where, + d.countSQL(where)) assert.Equal(t, - `DELETE FROM "posts_per_minute" WHERE "bucket" + INTERVAL '60' SECOND <= TIMESTAMPTZ '2026-09-13 10:05:00+00:00'`, - d.deleteClosedSQL(at)) - closed := `WITH closed AS (SELECT * FROM "posts_per_minute" WHERE "bucket" + INTERVAL '60' SECOND <= TIMESTAMPTZ '2026-09-13 10:05:00+00:00')` - assert.Equal(t, closed+` SELECT * FROM closed`, d.collectSQL(at)) + `DELETE FROM "posts_per_minute" WHERE `+where, + d.deleteSQL(where)) + cte := `WITH closed AS (SELECT * FROM "posts_per_minute" WHERE ` + where + `)` + assert.Equal(t, cte+` SELECT * FROM closed`, d.collectSQL(where)) d.EmitSQL = " SELECT bucket, sum(n) FROM closed GROUP BY ALL " - assert.Equal(t, closed+` SELECT bucket, sum(n) FROM closed GROUP BY ALL`, d.collectSQL(at)) + assert.Equal(t, cte+` SELECT bucket, sum(n) FROM closed GROUP BY ALL`, d.collectSQL(where)) // An emit_sql with its own WITH keeps it: closed joins its list. d.EmitSQL = "WITH totals AS (SELECT sum(n) AS n FROM closed)\nSELECT n FROM totals" - assert.Equal(t, closed+`, totals AS (SELECT sum(n) AS n FROM closed) -SELECT n FROM totals`, d.collectSQL(at)) + assert.Equal(t, cte+`, totals AS (SELECT sum(n) AS n FROM closed) +SELECT n FROM totals`, d.collectSQL(where)) d.EmitSQL = "with t as (select 1) select * from t, closed" - assert.Equal(t, closed+`, t as (select 1) select * from t, closed`, d.collectSQL(at)) + assert.Equal(t, cte+`, t as (select 1) select * from t, closed`, d.collectSQL(where)) } // An identifier with a quote in it is quoted, not injected. func TestManagerWindow_IdentifiersAreQuoted(t *testing.T) { coverage.Covers(t, "manager.window") - d := Declaration{Table: `odd"name`, TimeColumn: "when", Size: time.Second, Late: LateDrop} + d := Declaration{Table: `odd"name`, TimeColumn: "when", Size: time.Second} assert.Equal(t, `SELECT epoch_us(max("when")) FROM "odd""name"`, d.newestSQL()) } // A declaration the engine cannot run is refused when the manager is built, -// with a config code, not at the first poll. +// with a config code, not at the first pass. func TestManagerWindow_DeclarationIsValidated(t *testing.T) { coverage.Covers(t, "manager.window") cases := map[string]Declaration{ - "no table": {TimeColumn: "b", Size: time.Second, Late: LateDrop}, - "no time column": {Table: "t", Size: time.Second, Late: LateDrop}, - "zero size": {Table: "t", TimeColumn: "b", Late: LateDrop}, - "negative grace": {Table: "t", TimeColumn: "b", Size: time.Second, Grace: -1, Late: LateDrop}, - "negative idle": {Table: "t", TimeColumn: "b", Size: time.Second, IdleClose: -1, Late: LateDrop}, - "bad late policy": {Table: "t", TimeColumn: "b", Size: time.Second, Late: "keep"}, - "no late policy": {Table: "t", TimeColumn: "b", Size: time.Second}, - "engine table": {Table: "sqlflow_offsets", TimeColumn: "b", Size: time.Second, Late: LateDrop}, + "no table": {TimeColumn: "b", Size: time.Second}, + "no time column": {Table: "t", Size: time.Second}, + "zero size": {Table: "t", TimeColumn: "b"}, + "negative grace": {Table: "t", TimeColumn: "b", Size: time.Second, Grace: -1}, + "negative idle": {Table: "t", TimeColumn: "b", Size: time.Second, IdleClose: -1}, + "negative lateness": {Table: "t", TimeColumn: "b", Size: time.Second, Lateness: -1}, + "engine table": {Table: "sqlflow_offsets", TimeColumn: "b", Size: time.Second}, } for name, d := range cases { t.Run(name, func(t *testing.T) { @@ -75,5 +101,6 @@ func TestManagerWindow_DeclarationIsValidated(t *testing.T) { assert.Error(t, d.validate()) }) } - assert.NoError(t, Declaration{Table: "t", TimeColumn: "b", Size: time.Second, Late: LateReemit}.validate()) + assert.NoError(t, Declaration{Table: "t", TimeColumn: "b", Size: time.Second}.validate()) + assert.NoError(t, Declaration{Table: "t", TimeColumn: "b", Size: time.Second, Lateness: time.Hour}.validate()) } diff --git a/internal/managers/watermark.go b/internal/managers/watermark.go index 73e575d1..5ba0dcf1 100644 --- a/internal/managers/watermark.go +++ b/internal/managers/watermark.go @@ -20,50 +20,23 @@ import ( "go.uber.org/zap" ) -// defaultPollInterval is how often a manager looks for closed buckets when -// the declaration does not say. -const defaultPollInterval = 10 * time.Second - -// LatePolicy says what happens to a row for a bucket that already closed. -type LatePolicy string - -const ( - // LateReemit publishes emit_sql over the late rows alone: the bucket's - // other rows were deleted when it closed. For a sink that adds them to - // the bucket it holds. - LateReemit LatePolicy = "reemit" - // LateDrop discards the row and counts it. For a sink that appends or - // replaces. - LateDrop LatePolicy = "drop" -) - -// ParseLatePolicy resolves the configured name. There is no default: the -// two policies are different promises to the sink. -func ParseLatePolicy(s string) (LatePolicy, error) { - switch LatePolicy(s) { - case LateReemit: - return LateReemit, nil - case LateDrop: - return LateDrop, nil - case "": - return "", errs.New(errs.CodeConfigInvalid, "late_rows is required: drop, or reemit for a sink that adds late rows to the bucket it holds") - default: - return "", errs.New(errs.CodeConfigInvalid, "late_rows must be drop or reemit, not %q", s) - } -} - // Declaration is a window as the config declares it: which table, which -// column holds the bucket start, how long a bucket is, and how the close is -// decided. +// column holds the bucket start, how long a bucket is, and how long a closed +// bucket is kept. type Declaration struct { Table string TimeColumn string Size time.Duration Grace time.Duration - // IdleClose is how long the stream may be quiet before every open bucket - // closes. Zero means never. + // IdleClose is how long a source partition may be silent before it stops + // holding the window open. Zero means never. The engine reads it; the + // manager carries it only so one declaration describes the window. IdleClose time.Duration - Late LatePolicy + // Lateness is how long after a bucket closes its rows are kept, and a + // late row the engine admits republishes it whole. Zero: the rows are + // deleted in the pass that publishes them, and the engine refuses late + // rows before they arrive. Flink's allowedLateness. + Lateness time.Duration // EmitSQL shapes the closed rows for the sink. Empty means SELECT * FROM // closed. EmitSQL string @@ -81,44 +54,48 @@ func (d Declaration) validate() error { return errs.New(errs.CodeConfigInvalid, "window %s: grace_seconds cannot be negative", d.Table) case d.IdleClose < 0: return errs.New(errs.CodeConfigInvalid, "window %s: idle_close_seconds cannot be negative", d.Table) + case d.Lateness < 0: + return errs.New(errs.CodeConfigInvalid, "window %s: allowed_lateness_seconds cannot be negative", d.Table) case core.IsEngineTable(d.Table): return errs.New(errs.CodeConfigInvalid, "window %s: an engine table cannot carry a window", d.Table) } - if _, err := ParseLatePolicy(string(d.Late)); err != nil { - return err - } return nil } // Watermark closes one window. The engine asserts the window's watermark in // sqlflow_watermarks, in the commit that makes the rows it describes -// visible (core.Watermarks); the manager keeps what it has closed up to in -// sqlflow_windows: every bucket ending at or before it has been published to -// the sink and deleted from the table. A poll moves the second to the first -// and publishes the buckets between them. Neither moves backwards, so a -// bucket is closed once, and a row that arrives for it afterwards is late. +// visible (core.Watermarks), and tells the manager through the window's +// signal when it moved and which closed buckets a late row landed in. The +// manager keeps what it has closed up to in sqlflow_windows. A pass moves +// the second to the first and publishes the buckets between them, +// republishes the buckets the signal named, and deletes the buckets past +// their lateness. Neither watermark moves backwards. // -// The manager has no clock. Every decision is `bucket end <= watermark`, in -// event time; the engine's clock decides only which partitions are in its -// minimum, and no reading of any clock reaches here. +// The manager has no clock, no ticker and no interval. It runs a pass on +// start, on every kick, and on the drain. Every decision is `bucket end <= +// watermark`, in event time; the engine's clock decides only which +// partitions are in its minimum, and no reading of any clock reaches here. // // It runs on a connection of its own, with autocommit off, so it reads -// committed rows only and its delete commits together with its watermark. +// committed rows only and its deletes commit together with its watermark. // Nothing here shares the pipeline's lock or transaction. type Watermark struct { - conn adbc.Connection - tx transaction - decl Declaration - store *Store - sink core.Sink - poll time.Duration - - // pollTrigger replaces the poll ticker when set; see WithPollTrigger. - pollTrigger <-chan time.Time + conn adbc.Connection + tx transaction + decl Declaration + store *Store + sink core.Sink + signal *core.WindowSignal logger *zap.Logger drain *core.DrainBudget metrics WindowMetrics + attrs metric.MeasurementOption + + // republishRetained is set for the start pass: every bucket still + // retained is republished whole, because a late row that landed before a + // restart left its recompute in memory only. + republishRetained bool } // transaction is the boundary on the manager's connection. An ADBC @@ -134,7 +111,7 @@ func WithLogger(l *zap.Logger) Option { return func(w *Watermark) { w.logger = l.Named("manager.watermark") } } -// WithDrainBudget bounds the final poll after a cancel. +// WithDrainBudget bounds the final pass after a cancel. func WithDrainBudget(b *core.DrainBudget) Option { return func(w *Watermark) { w.drain = b } } @@ -145,17 +122,12 @@ func WithMeterProvider(mp metric.MeterProvider) Option { return func(w *Watermark) { w.metrics = NewWindowMetrics(mp, w.decl.Table) } } -// WithPollTrigger replaces the poll ticker, so a caller decides when the -// manager looks for closed buckets. A simulator owns the order of its events -// this way; production leaves it nil and polls on the interval. -func WithPollTrigger(c <-chan time.Time) Option { - return func(w *Watermark) { w.pollTrigger = c } -} - // NewWatermark builds the manager for one window on conn, which must be a -// connection of its own with autocommit off: every poll ends with a commit or -// a rollback on it. -func NewWatermark(conn adbc.Connection, d Declaration, poll time.Duration, sink core.Sink, opts ...Option) (*Watermark, error) { +// connection of its own with autocommit off: every pass ends with a commit +// or a rollback on it. signal is the window's, from core.Watermarks; nil is +// allowed for a caller that drives Pass itself, and Start then runs its +// start pass and waits for the drain. +func NewWatermark(conn adbc.Connection, d Declaration, sink core.Sink, signal *core.WindowSignal, opts ...Option) (*Watermark, error) { if err := d.validate(); err != nil { return nil, err } @@ -163,9 +135,6 @@ func NewWatermark(conn adbc.Connection, d Declaration, poll time.Duration, sink if !ok { return nil, errs.New(errs.CodeStateInternal, "window %s: the connection does not support transactions", d.Table) } - if poll <= 0 { - poll = defaultPollInterval - } w := &Watermark{ conn: conn, @@ -173,9 +142,10 @@ func NewWatermark(conn adbc.Connection, d Declaration, poll time.Duration, sink decl: d, store: NewStore(conn), sink: sink, - poll: poll, + signal: signal, logger: zap.NewNop(), metrics: NewWindowMetrics(nil, d.Table), + attrs: metric.WithAttributes(attribute.String("window", d.Table)), } for _, opt := range opts { opt(w) @@ -189,88 +159,127 @@ func NewWatermark(conn adbc.Connection, d Declaration, poll time.Duration, sink // Declaration is what the manager was built from. func (w *Watermark) Declaration() Declaration { return w.decl } -// Start polls until the context is cancelled, then polls once more so buckets -// that closed during the final interval are not stranded in the table. +// Start runs one pass, then one on every kick, then one on the drain. It has +// no clock: the engine kicks when the watermark moved or a late row landed, +// and the engine's own flush tick is the only timer in the design. A +// context already cancelled runs the drain pass only, which is the shape a +// shutdown that raced startup has. // -// A failed poll returns. The rows are still in the table, because Poll +// The pass on start is what a restart needs and nothing more: buckets the +// watermark had passed are published, and every retained bucket is +// republished whole, since the recompute set a late row landed in before the +// crash was in memory. Bounded by lateness over size buckets per key, and +// idempotent because the value is the whole bucket. +// +// A failed pass returns. The rows are still in the table, because a pass // deletes only after the sink accepted them, so a restart republishes the // same bucket. Start does not retry in place: the sink already ran its retry // ladder before the error reached here, so what arrives is a destination // that rejected the rows or stayed unreachable past the deadline. The one -// exception is a write conflict with the pipeline, which the next poll +// exception is a write conflict with the pipeline, which the next kick // retries. func (w *Watermark) Start(ctx context.Context) error { - w.logger.Info("starting watermark manager", - zap.String("table", w.decl.Table), - zap.Duration("poll_interval", w.poll)) - - pollC := w.pollTrigger - if pollC == nil { - ticker := time.NewTicker(w.poll) - defer ticker.Stop() - pollC = ticker.C + w.logger.Info("starting watermark manager", zap.String("table", w.decl.Table)) + if ctx.Err() != nil { + return w.finalPass() + } + if err := w.StartPass(ctx); err != nil && ctx.Err() == nil && !isConflict(err) { + w.logger.Error("start pass failed, stopping the manager", zap.Error(err)) + return fmt.Errorf("watermark manager %s: %w", w.decl.Table, err) } + var wake <-chan struct{} + if w.signal != nil { + wake = w.signal.Wait() + } for { select { - case <-pollC: - err := w.Poll(ctx) + case <-wake: + err := w.Pass(ctx) if err != nil && ctx.Err() == nil { if isConflict(err) { - w.logger.Warn("poll conflicted with the pipeline, retrying next poll", zap.Error(err)) + w.logger.Warn("pass conflicted with the pipeline, retrying on the next kick", zap.Error(err)) continue } - w.logger.Error("poll failed, stopping the manager", zap.Error(err)) + w.logger.Error("pass failed, stopping the manager", zap.Error(err)) return fmt.Errorf("watermark manager %s: %w", w.decl.Table, err) } if ctx.Err() == nil { continue } - // The cancel landed during that poll. The final poll below is + // The cancel landed during that pass. The final pass below is // the one that counts. case <-ctx.Done(): } - return w.finalPoll() + return w.finalPass() } } -// finalPoll publishes what closed during the last interval, on the drain -// budget. A poll the deadline ended is reported as the drain running out of -// time, not as the sink's own failure: the rows are still in the table, and -// the next start publishes them. -func (w *Watermark) finalPoll() error { - err := w.Poll(w.drain.Context()) +// finalPass publishes what closed since the last kick, on the drain budget. +// A pass the deadline ended is reported as the drain running out of time, +// not as the sink's own failure: the rows are still in the table, and the +// next start publishes them. +func (w *Watermark) finalPass() error { + err := w.Pass(w.drain.Context()) if err == nil { return nil } if w.drain.Exceeded() { err = errs.Wrap(errs.CodeDrainIncomplete, err, - "drain deadline %s reached before the final poll finished", w.drain.Deadline()) + "drain deadline %s reached before the final pass finished", w.drain.Deadline()) } - w.logger.Error("final poll failed", zap.Error(err)) - return fmt.Errorf("watermark manager %s: final poll: %w", w.decl.Table, err) + w.logger.Error("final pass failed", zap.Error(err)) + return fmt.Errorf("watermark manager %s: final pass: %w", w.decl.Table, err) +} + +// StartPass is the pass Start runs first: a Pass that also republishes, +// whole, every bucket still retained under the window's lateness. A late +// row admitted before a restart put its bucket in a recompute set that lived +// in memory; the rows are in the table, so the start pass republishes every +// bucket that could hold one. Bounded by lateness over size buckets per key, +// and idempotent because the value is the whole bucket. +func (w *Watermark) StartPass(ctx context.Context) error { + w.republishRetained = true + return w.Pass(ctx) } -// Poll runs one close. It ends its transaction before returning, committed -// or rolled back, so the next poll reads a fresh snapshot. -func (w *Watermark) Poll(ctx context.Context) (err error) { +// Pass is one unit of the manager's work: publish every bucket the +// assertion closed, republish every bucket a late row landed in, delete +// every bucket past its lateness, and record where the window has closed up +// to. It ends its transaction before returning, committed or rolled back, so +// the next pass reads a fresh snapshot. +// +// The order inside is the guarantee. Publishes come first and deletes last, +// so a flush that fails leaves every row for the next pass. With no lateness +// the purge is exactly the buckets this pass published, which is the close +// this manager has always made; with lateness they stay until the watermark +// passes their end plus it. +func (w *Watermark) Pass(ctx context.Context) (err error) { committed := false - // What the close lag needs, filled in as the poll learns it. Recorded in - // the defer so a poll that fails after computing the close still reports - // how far behind it is -- that is the case the gauge exists for. + // What the close lag needs, filled in as the pass learns it. Recorded in + // the defer so a pass that fails after reading the assertion still + // reports how far behind it is -- that is the case the gauge exists for. var ( candidate time.Time candidateKnown bool settled time.Time settledKnown bool + recompute []time.Time ) defer func() { if !committed { // Every read opened a transaction, and a transaction left open - // would freeze the next poll's view of the table. + // would freeze the next pass's view of the table. if rbErr := w.tx.Rollback(context.WithoutCancel(ctx)); rbErr != nil && err == nil { err = fmt.Errorf("rolling back: %w", rbErr) } + // The buckets this pass took from the signal were not published; + // hand them back so the next pass does them. + if w.signal != nil { + for _, b := range recompute { + w.signal.Recompute(b) + } + } } if candidateKnown && settledKnown { w.recordCloseLag(context.WithoutCancel(ctx), candidate, settled) @@ -285,37 +294,6 @@ func (w *Watermark) Poll(ctx context.Context) (err error) { settled, settledKnown = closed, true } - // Late rows belong to buckets that already closed. Under drop they leave - // now, before the close is computed, so they are never collected. - // lateCounted is recorded only after the commit below. It used to be - // recorded here, and a close that then lost a write conflict rolled the - // delete back while the counter kept the rows: the next poll found the - // same rows and counted them again, so 500 rows dropped once read as - // 1,000. This counter is the data-loss signal, so it counts what - // happened rather than what was attempted. - var lateToReemit, dropped, lateCounted int64 - if hadClosed { - late, _, err := queryInt64(ctx, w.conn, w.decl.countClosedSQL(closed)) - if err != nil { - return fmt.Errorf("counting late rows: %w", err) - } - if late > 0 { - lateCounted = late - switch DecideBucket(BucketState{Bucket: BucketLate, Policy: w.decl.Late}) { - case DropLate: - if _, err := execRows(ctx, w.conn, w.decl.deleteClosedSQL(closed)); err != nil { - return fmt.Errorf("dropping late rows: %w", err) - } - dropped = late - w.logger.Info("dropped late rows", zap.Int64("rows", late)) - case ReemitLate: - // They stay for the close below, which collects them with - // the buckets that are due and runs emit_sql over the lot. - lateToReemit = late - } - } - } - // The one fact: what the engine has asserted. asserted, hasAsserted, err := core.LoadWatermark(ctx, w.conn, w.decl.Table) if err != nil { @@ -332,11 +310,7 @@ func (w *Watermark) Poll(ctx context.Context) (err error) { return fmt.Errorf("reading the newest bucket: %w", err) } if hasRows { - // Every poll that sees rows, not only one that commits. Rows stamped - // in the future are reported the moment they arrive, even if the close - // that would record them fails. - w.metrics.NewestStart.Record(ctx, time.UnixMicro(newestMicros).Unix(), - metric.WithAttributes(attribute.String("window", w.decl.Table))) + w.metrics.NewestStart.Record(ctx, time.UnixMicro(newestMicros).Unix(), w.attrs) } if !hadClosed && hasRows { // Never closed: the first close is due once the watermark reaches the @@ -352,6 +326,12 @@ func (w *Watermark) Poll(ctx context.Context) (err error) { } } + // The buckets a late row landed in since the last pass. Taken before the + // decision, so a kick that carried only recomputes still does them. + if w.signal != nil { + recompute = w.signal.TakeRecompute() + } + state := StateOf(asserted, hasAsserted, closed, hadClosed) rule := watermarkRuleFor(state) watermark, moved := rule.Action.Next(asserted, closed) @@ -361,94 +341,121 @@ func (w *Watermark) Poll(ctx context.Context) (err error) { zap.Stringer("state", state), zap.Time("watermark", watermark)) } - if !moved && lateToReemit == 0 && dropped == 0 { + republish := w.republishRetained && hadClosed && w.decl.Lateness > 0 + if !moved && len(recompute) == 0 && !republish { + w.republishRetained = false return nil } if !moved { watermark = closed } - // Anything to publish? A watermark that moved over an empty stretch - // still has to be saved, or the next poll recomputes the same move. - // The rows counted here end at or before the watermark: the buckets - // that are due, and under reemit the late rows kept above. - rows, _, err := queryInt64(ctx, w.conn, w.decl.countClosedSQL(watermark)) - if err != nil { - return fmt.Errorf("counting closed rows: %w", err) - } - if rows > 0 { - switch DecideBucket(BucketState{Bucket: BucketDue, Policy: w.decl.Late}) { - case Close: - if err := w.publish(ctx, watermark); err != nil { + // Publish what is due: the buckets that ended after the closed + // watermark and at or before the asserted one. A watermark that moved + // over an empty stretch still has to be saved, or the next pass decides + // the same move. + if moved { + due := w.decl.dueBetween(closed, hadClosed, watermark) + rows, _, err := queryInt64(ctx, w.conn, w.decl.countSQL(due)) + if err != nil { + return fmt.Errorf("counting due rows: %w", err) + } + if rows > 0 { + if err := w.publish(ctx, due); err != nil { return err } - if _, err := execRows(ctx, w.conn, w.decl.deleteClosedSQL(watermark)); err != nil { - return fmt.Errorf("deleting closed rows: %w", err) + w.metrics.Closed.Add(ctx, 1, w.attrs) + w.logger.Debug("closed", zap.Time("watermark", watermark), zap.Int64("rows", rows)) + } + } + + // Republish, whole, every bucket a late row landed in. The value the + // sink receives is emit_sql over every row the bucket has, never the late + // rows alone, so a sink that replaces by key holds the exact count. + for _, b := range recompute { + if err := w.publish(ctx, w.decl.bucketIs(b)); err != nil { + return err + } + w.metrics.Recomputed.Add(ctx, 1, w.attrs) + w.logger.Debug("recomputed", zap.Time("bucket", b)) + } + // Then the transaction owns it: nothing hands these back on failure past + // this point, because the publish already happened. + recompute = nil + + // The start pass: every bucket closed earlier whose lateness has not run + // out, republished whole. Retained is measured against the watermark this + // pass settles on, so a bucket the same pass expires is not republished + // and then deleted. + if republish { + retained := w.decl.retainedBetween(closed, watermark) + rows, _, err := queryInt64(ctx, w.conn, w.decl.countSQL(retained)) + if err != nil { + return fmt.Errorf("counting retained rows: %w", err) + } + if rows > 0 { + if err := w.publish(ctx, retained); err != nil { + return err } + w.logger.Debug("republished retained buckets on start", zap.Int64("rows", rows)) } } + // Purge what is past its lateness: nothing more can arrive for those + // buckets, because the engine refuses it. With no lateness this is what + // the pass just published. + if _, err := execRows(ctx, w.conn, w.decl.deleteSQL(w.decl.expiredBefore(watermark))); err != nil { + return fmt.Errorf("deleting expired rows: %w", err) + } + // closed_at is the wall clock, for an operator reading the table; no // decision reads it back. if err := w.store.Save(ctx, w.decl.Table, watermark, time.Now().UTC()); err != nil { return err } if err := w.tx.Commit(ctx); err != nil { - return errs.Wrap(errs.CodeStateCommitFailed, err, "committing the close") + return errs.Wrap(errs.CodeStateCommitFailed, err, "committing the pass") } committed = true + w.republishRetained = false settled, settledKnown = watermark, true - - if lateCounted > 0 { - w.metrics.Late.Add(ctx, lateCounted, metric.WithAttributes( - attribute.String("window", w.decl.Table), - attribute.String("policy", string(w.decl.Late)))) - } - w.metrics.Watermark.Record(ctx, watermark.Unix(), metric.WithAttributes( - attribute.String("window", w.decl.Table))) - if rows > 0 { - w.metrics.Closed.Add(ctx, 1, metric.WithAttributes( - attribute.String("window", w.decl.Table))) - w.logger.Debug("closed", zap.Time("watermark", watermark), zap.Int64("rows", rows)) - } + w.metrics.Watermark.Record(ctx, watermark.Unix(), w.attrs) return nil } // recordCloseLag records how far the window's closes trail its own data, in // event time. // -// candidate is where the watermark should be, given the rows the window -// holds; settled is where it actually is, committed. The difference is the -// close that is overdue. It is zero whenever a close commits, zero after an -// idle close has closed everything, and grows while rows arrive and closes -// fail -- a sink that is down, a transaction that keeps conflicting. +// candidate is where the watermark should be, given what the engine has +// asserted; settled is where it actually is, committed. The difference is the +// close that is overdue. It is zero whenever a pass commits, and grows while +// the engine's assertion moves on and passes fail -- a sink that is down, a +// transaction that keeps conflicting. // // No wall clock enters it. Two earlier readings compared event time with // this host's clock: the watermark's age, which trailed by size and grace by // design, and wall time past the next close, which grew for any stream that // went quiet. Both read a sparse stream as stalled, both were wrong on a // gateway whose clock was never set, and both needed a close since startup -// to report anything. A stream going quiet is the source's to report, as -// last_message_at does; this is the window's. +// to report anything. func (w *Watermark) recordCloseLag(ctx context.Context, candidate, settled time.Time) { lag := int64(candidate.Sub(settled) / time.Second) if lag < 0 { lag = 0 } - w.metrics.CloseLag.Record(ctx, lag, - metric.WithAttributes(attribute.String("window", w.decl.Table))) + w.metrics.CloseLag.Record(ctx, lag, w.attrs) } -// publish runs emit_sql over the closed rows and hands the result to the -// sink. Flushed before the delete, so a failure leaves the rows in the table -// to be retried rather than dropping them. +// publish runs emit_sql over the rows where selects and hands the result to +// the sink. Flushed before any delete, so a failure leaves the rows in the +// table to be retried rather than dropping them. // // The sink runs on a connection of its own, not this one. This connection -// holds the close's transaction, and a transaction may write to one database +// holds the pass's transaction, and a transaction may write to one database // only: a sink that writes into an attached Postgres, or stages a batch // table, would fail it. -func (w *Watermark) publish(ctx context.Context, watermark time.Time) error { - table, err := w.collect(ctx, watermark) +func (w *Watermark) publish(ctx context.Context, where string) error { + table, err := w.collect(ctx, where) if err != nil { return err } @@ -469,13 +476,13 @@ func (w *Watermark) publish(ctx context.Context, watermark time.Time) error { return nil } -func (w *Watermark) collect(ctx context.Context, watermark time.Time) (arrow.Table, error) { +func (w *Watermark) collect(ctx context.Context, where string) (arrow.Table, error) { stmt, err := w.conn.NewStatement() if err != nil { return nil, err } defer stmt.Close() - if err := stmt.SetSqlQuery(w.decl.collectSQL(watermark)); err != nil { + if err := stmt.SetSqlQuery(w.decl.collectSQL(where)); err != nil { return nil, err } reader, _, err := stmt.ExecuteQuery(ctx) @@ -506,7 +513,7 @@ func (w *Watermark) collect(ctx context.Context, watermark time.Time) (arrow.Tab } // isConflict reports a DuckDB transaction conflict: the pipeline's -// transaction touched a row this poll deleted. The next poll retries. An +// transaction touched a row this pass deleted. The next kick retries. An // error that carries a code is never one; a sink's failure keeps its code // and stops the manager. func isConflict(err error) bool { @@ -517,13 +524,16 @@ func isConflict(err error) bool { return strings.Contains(msg, "Conflict on") || strings.Contains(msg, "write-write conflict") } -// WindowMetrics is the three instruments a window records. Exported so the +// WindowMetrics is the instruments a window records. Exported so the // series-name test drives them through the real constructor, the way it -// drives core.NewMetrics. +// drives core.NewMetrics. Late rows are the engine's to count now, in +// core.Metrics.WindowLateRows, because the engine is where they are decided. type WindowMetrics struct { Watermark metric.Int64Gauge Closed metric.Int64Counter - Late metric.Int64Counter + // Recomputed counts buckets republished whole because a late row arrived + // within allowed_lateness_seconds. + Recomputed metric.Int64Counter // CloseLag is how far the window's closes trail its own data, in event // seconds. See recordCloseLag. CloseLag metric.Int64Gauge @@ -554,8 +564,8 @@ func NewWindowMetrics(mp metric.MeterProvider, table string) WindowMetrics { metric.WithDescription("Closes that published at least one bucket")); err != nil { return NewWindowMetrics(noop.NewMeterProvider(), table) } - if m.Late, err = meter.Int64Counter("window_late_rows", - metric.WithDescription("Rows that arrived for a bucket that had already closed, by policy")); err != nil { + if m.Recomputed, err = meter.Int64Counter("window_recomputes", + metric.WithDescription("Buckets republished whole because a late row arrived within allowed_lateness_seconds")); err != nil { return NewWindowMetrics(noop.NewMeterProvider(), table) } if m.CloseLag, err = meter.Int64Gauge("window_close_lag_seconds", @@ -570,9 +580,9 @@ func NewWindowMetrics(mp metric.MeterProvider, table string) WindowMetrics { } // The window exists from here, and says so. Nothing else is recorded - // until the first close commits, which after a start or a restart can be - // a poll interval away, and a reader that sees no window series concludes - // that nothing here drops rows -- on a pipeline configured to drop them. + // until the first pass commits, and a reader that sees no window series + // concludes that nothing here drops rows -- on a pipeline configured to + // refuse them. m.Closed.Add(context.Background(), 0, metric.WithAttributes(attribute.String("window", table))) return m diff --git a/internal/managers/watermark_test.go b/internal/managers/watermark_test.go index 5065b1c7..19bbea39 100644 --- a/internal/managers/watermark_test.go +++ b/internal/managers/watermark_test.go @@ -16,7 +16,8 @@ import ( ) // The manager against the one fact it reads. assertAt is the engine's -// commit, written by hand; every test here is a poll against it. +// commit, written by hand; every test here is a pass against it, driven by +// the test the way a kick would drive it. // Buckets close up to the asserted watermark and no further: with the // engine at bucket 0's end, bucket 0 publishes and bucket 1 stays. @@ -26,17 +27,17 @@ func TestManagerWindow_ClosesUpToTheAssertedWatermark(t *testing.T) { d := newTestDB(t, "") createWindowTable(t, d.pipeline) sink := &recordingSink{} - w := newTestWatermark(t, d, testDecl(), sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 1, "NYC", 1) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, flushes := sink.counts() assert.Equal(t, int64(0), rows) assert.Equal(t, 0, flushes) assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, flushes = sink.counts() assert.Equal(t, int64(1), rows) assert.Equal(t, 1, flushes) @@ -48,7 +49,7 @@ func TestManagerWindow_ClosesUpToTheAssertedWatermark(t *testing.T) { assert.That(t, wm.Equal(bucket(1))) // Nothing new: nothing published, nothing deleted. - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, flushes = sink.counts() assert.Equal(t, int64(1), rows) assert.Equal(t, 1, flushes) @@ -62,12 +63,12 @@ func TestManagerWindow_AnAssertionPastEveryBucketClosesEverything(t *testing.T) d := newTestDB(t, "") createWindowTable(t, d.pipeline) sink := &recordingSink{} - w := newTestWatermark(t, d, testDecl(), sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 1, "SF", 1) assertAt(t, d.pipeline, bucket(2)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, flushes := sink.counts() assert.Equal(t, int64(2), rows) assert.Equal(t, 1, flushes) @@ -78,7 +79,7 @@ func TestManagerWindow_AnAssertionPastEveryBucketClosesEverything(t *testing.T) assert.That(t, wm.Equal(bucket(2))) } -// Without an assertion nothing closes, however many polls run and whatever +// Without an assertion nothing closes, however many passes run and whatever // the table holds: the manager has no clock to grow impatient on, and the // newest bucket is an observation, not a fact it decides on. func TestManagerWindow_AnUnassertedWindowNeverCloses(t *testing.T) { @@ -87,12 +88,12 @@ func TestManagerWindow_AnUnassertedWindowNeverCloses(t *testing.T) { d := newTestDB(t, "") createWindowTable(t, d.pipeline) sink := &recordingSink{} - w := newTestWatermark(t, d, testDecl(), sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 5, "SF", 1) for i := 0; i < 3; i++ { - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) if rows, _ := sink.counts(); rows != 0 { t.Fatalf("%d rows published with no watermark asserted", rows) } @@ -104,17 +105,17 @@ func TestManagerWindow_AnUnassertedWindowNeverCloses(t *testing.T) { } // A watermark that moved over an empty stretch is still saved, or the next -// poll would decide the same move again. +// pass would decide the same move again. func TestManagerWindow_AnAssertionOverAnEmptyTableIsSaved(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() d := newTestDB(t, "") createWindowTable(t, d.pipeline) sink := &recordingSink{} - w := newTestWatermark(t, d, testDecl(), sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) assertAt(t, d.pipeline, bucket(5)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) _, flushes := sink.counts() assert.Equal(t, 0, flushes) wm, ok, err := NewStore(d.pipeline).Load(ctx, testTable) @@ -124,36 +125,37 @@ func TestManagerWindow_AnAssertionOverAnEmptyTableIsSaved(t *testing.T) { } // The closed watermark never moves backwards. An assertion at or below it -// is one that has been acted on, and a poll changes nothing; the rows it -// would cover are late. +// is one that has been acted on, and a pass changes nothing: it publishes +// nothing and, with nothing moved, deletes nothing either. func TestManagerWindow_WatermarkNeverRegresses(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() d := newTestDB(t, "") createWindowTable(t, d.pipeline) sink := &recordingSink{} - decl := testDecl() - decl.Late = LateDrop - w := newTestWatermark(t, d, decl, sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 5, "NYC", 1) assertAt(t, d.pipeline, bucket(4)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) wm, _, err := NewStore(d.pipeline).Load(ctx, testTable) assert.NoError(t, err) assert.That(t, wm.Equal(bucket(4))) // The engine cannot assert lower, by construction; were the row to say - // so anyway, the manager holds. + // so anyway, the manager holds. The row for bucket 1 is one the engine + // would have refused; written by hand, it sits there, unpublished. exec(t, d.pipeline, `DELETE FROM agg_cities_count`) insertBucket(t, d.pipeline, 1, "late", 1) assertAt(t, d.pipeline, bucket(2)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) wm2, _, err := NewStore(d.pipeline).Load(ctx, testTable) assert.NoError(t, err) assert.That(t, wm2.Equal(bucket(4))) - assert.Equal(t, int64(0), countRows(t, d.pipeline, testTable)) + rows, _ := sink.counts() + assert.Equal(t, int64(1), rows) + assert.Equal(t, int64(1), countRows(t, d.pipeline, testTable)) } // A manager built over the state another one saved starts from its @@ -166,109 +168,215 @@ func TestManagerWindow_ARestartResumesFromTheWatermark(t *testing.T) { createWindowTable(t, d.pipeline) sink := &recordingSink{} - w := newTestWatermark(t, d, testDecl(), sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, _ := sink.counts() assert.Equal(t, int64(1), rows) - // A late row for bucket 0 arrives, and a fresh manager under drop reads - // only the persisted watermark. + // A row for bucket 0 is in the table -- one the engine would have + // refused, written by hand -- and a fresh manager reads only the + // persisted watermark: bucket 0 is behind it, and nothing is published. insertBucket(t, d.pipeline, 0, "late", 1) sink2 := &recordingSink{} - decl := testDecl() - decl.Late = LateDrop - w2, err := NewWatermark(managerConn(t, d.db), decl, time.Hour, sink2) + w2, err := NewWatermark(managerConn(t, d.db), testDecl(), sink2, nil) assert.NoError(t, err) - assert.NoError(t, w2.Poll(ctx)) + assert.NoError(t, w2.Pass(ctx)) rows2, flushes2 := sink2.counts() assert.Equal(t, int64(0), rows2) assert.Equal(t, 0, flushes2) - assert.Equal(t, int64(1), countRows(t, d.pipeline, testTable)) + assert.Equal(t, int64(2), countRows(t, d.pipeline, testTable)) } -// drop discards late rows without a flush and counts them; reemit publishes -// them. -func TestManagerWindow_LateRowsFollowThePolicy(t *testing.T) { +// With lateness, a closed bucket's rows stay until the watermark passes +// end + lateness, and a recompute republishes the whole bucket: the sink's +// last value for it is the exact count, not the late rows alone. +// +// step table closed published (last value for bucket 0) +// ---------------------------------------------------------------------------------------- +// buckets 0 (3 rows), 2 (1) 0:3, 2:1 - - +// assert 10:01, pass 0:3, 2:1 10:01 bucket 0 → 3 rows kept: lateness 5m +// late row for 0, recompute 0:4, 2:1 10:01 bucket 0 → 4 the whole bucket, again +// assert 10:07, pass - 10:07 bucket 2 → 1 0 expired: 10:01 + 5m <= 10:07 +func TestManagerWindow_ARecomputeRepublishesTheWholeBucket(t *testing.T) { coverage.Covers(t, "manager.window") - for _, policy := range []LatePolicy{LateDrop, LateReemit} { - t.Run(string(policy), func(t *testing.T) { - coverage.Covers(t, "manager.window") - ctx := context.Background() - d := newTestDB(t, "") - createWindowTable(t, d.pipeline) - reader := sdkmetric.NewManualReader() - mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) - sink := &recordingSink{} - decl := testDecl() - decl.Late = policy - w := newTestWatermark(t, d, decl, sink, WithMeterProvider(mp)) - - insertBucket(t, d.pipeline, 0, "NYC", 3) - insertBucket(t, d.pipeline, 2, "NYC", 1) - assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, w.Poll(ctx)) - - insertBucket(t, d.pipeline, 0, "late", 1) - insertBucket(t, d.pipeline, 0, "late", 1) - assert.NoError(t, w.Poll(ctx)) - - rows, flushes := sink.counts() - if policy == LateDrop { - assert.Equal(t, int64(1), rows) - assert.Equal(t, 1, flushes) - } else { - assert.Equal(t, int64(3), rows) - assert.Equal(t, 2, flushes) - } - assert.Equal(t, int64(1), countRows(t, d.pipeline, testTable)) - assert.Equal(t, int64(2), counterValue(t, reader, "window_late_rows")) - }) - } + ctx := context.Background() + d := newTestDB(t, "") + createWindowTable(t, d.pipeline) + sink := &recordingSink{} + decl := testDecl() + decl.Lateness = 5 * time.Minute + decl.EmitSQL = "SELECT sum(count)::INT AS total, bucket FROM closed GROUP BY ALL" + w, sig := newTestWatermark(t, d, decl, sink) + + insertBucket(t, d.pipeline, 0, "NYC", 3) + insertBucket(t, d.pipeline, 2, "NYC", 1) + assertAt(t, d.pipeline, bucket(1)) + assert.NoError(t, w.Pass(ctx)) + assert.DeepEqual(t, [][]string{{"3"}}, sink.published()) + assert.Equal(t, int64(2), countRows(t, d.pipeline, testTable)) // bucket 0 retained + + insertBucket(t, d.pipeline, 0, "late", 1) + sig.Recompute(bucket(0)) + assert.NoError(t, w.Pass(ctx)) + assert.DeepEqual(t, [][]string{{"3"}, {"4"}}, sink.published()) + assert.Equal(t, int64(3), countRows(t, d.pipeline, testTable)) + + assertAt(t, d.pipeline, bucket(7)) + assert.NoError(t, w.Pass(ctx)) + assert.DeepEqual(t, [][]string{{"3"}, {"4"}, {"1"}}, sink.published()) + // Bucket 0 expired at 10:06 and is gone; bucket 2 ended 10:03, and + // 10:03 + 5m = 10:08 is past the assertion, so it is retained. + assert.Equal(t, int64(1), countRows(t, d.pipeline, testTable)) } -// Late rows are counted once, by the close that commits their fate. -// -// The counter used to be incremented before the close committed. A close -// that then failed rolled the delete back and kept the count, so the next -// poll found the same rows and counted them again: two rows dropped once read -// as four. This counter is the data-loss signal the TurboStats bundle -// reports, so it counts what happened rather than what was attempted. -func TestManagerWindow_LateRowsAreCountedOnlyWhenTheCloseCommits(t *testing.T) { +// A recompute the pass could not publish is not lost: the buckets go back on +// the signal, and the pass that succeeds republishes them. +func TestManagerWindow_AFailedRecomputeIsRetried(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() d := newTestDB(t, "") createWindowTable(t, d.pipeline) - reader := sdkmetric.NewManualReader() - mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) decl := testDecl() - decl.Late = LateDrop - - first := newTestWatermark(t, d, decl, &recordingSink{}, WithMeterProvider(mp)) + decl.Lateness = 5 * time.Minute + decl.EmitSQL = "SELECT sum(count)::INT AS total, bucket FROM closed GROUP BY ALL" + sink := &recordingSink{} + w, sig := newTestWatermark(t, d, decl, sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, first.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) - // Two late rows for the closed bucket, and a newer bucket so the next - // close has something to publish -- to a sink that is down. insertBucket(t, d.pipeline, 0, "late", 1) + sig.Recompute(bucket(0)) + down, _ := newTestWatermark(t, d, decl, &failingSink{err: errors.New("sink down")}) + // The failing manager shares the signal only through the test: hand it + // the same one by taking and re-queueing on the same signal. + down.signal = sig + assert.Error(t, down.Pass(ctx)) + + assert.NoError(t, w.Pass(ctx)) + assert.DeepEqual(t, [][]string{{"3"}, {"4"}}, sink.published()) +} + +// The start pass republishes every retained bucket whole. A late row admitted +// just before a restart left its recompute in memory only; its rows are in +// the table, and the manager that starts over them publishes the exact value +// without being told which bucket it was. +func TestManagerWindow_TheStartPassRepublishesRetainedBuckets(t *testing.T) { + coverage.Covers(t, "manager.window") + ctx := context.Background() + d := newTestDB(t, "") + createWindowTable(t, d.pipeline) + decl := testDecl() + decl.Lateness = 5 * time.Minute + decl.EmitSQL = "SELECT sum(count)::INT AS total, bucket FROM closed GROUP BY ALL ORDER BY bucket" + first := &recordingSink{} + w, _ := newTestWatermark(t, d, decl, first) + insertBucket(t, d.pipeline, 0, "NYC", 3) + insertBucket(t, d.pipeline, 2, "NYC", 1) + assertAt(t, d.pipeline, bucket(1)) + assert.NoError(t, w.Pass(ctx)) + assert.DeepEqual(t, [][]string{{"3"}}, first.published()) + + // The late row lands, the engine commits it, and the process dies before + // the manager passes. A new manager over the same database. insertBucket(t, d.pipeline, 0, "late", 1) - insertBucket(t, d.pipeline, 4, "NYC", 1) + second := &recordingSink{} + restarted, err := NewWatermark(managerConn(t, d.db), decl, second, nil) + assert.NoError(t, err) + assert.NoError(t, restarted.StartPass(ctx)) + assert.DeepEqual(t, [][]string{{"4"}}, second.published()) + + // Only the first pass republishes; the next one has nothing to do. + assert.NoError(t, restarted.Pass(ctx)) + assert.DeepEqual(t, [][]string{{"4"}}, second.published()) + + // Without lateness there is nothing retained, and the start pass is a + // pass: a manager over a table whose watermark has not moved publishes + // nothing. + plain := &recordingSink{} + none, err := NewWatermark(managerConn(t, d.db), testDecl(), plain, nil) + assert.NoError(t, err) + assert.NoError(t, none.StartPass(ctx)) + _, flushes := plain.counts() + assert.Equal(t, 0, flushes) +} + +// With no lateness the close deletes what it publishes, as it always did. +func TestManagerWindow_WithoutLatenessACloseDeletesWhatItPublishes(t *testing.T) { + coverage.Covers(t, "manager.window") + ctx := context.Background() + d := newTestDB(t, "") + createWindowTable(t, d.pipeline) + sink := &recordingSink{} + w, _ := newTestWatermark(t, d, testDecl(), sink) + insertBucket(t, d.pipeline, 0, "NYC", 3) + insertBucket(t, d.pipeline, 2, "NYC", 1) + assertAt(t, d.pipeline, bucket(1)) + assert.NoError(t, w.Pass(ctx)) + assert.Equal(t, int64(1), countRows(t, d.pipeline, testTable)) +} + +// Start runs one pass on start, one on every kick, and one on the drain; +// nothing wakes it otherwise. An assertion nobody kicks for sees no pass; +// the kick sees one. +func TestManagerWindow_StartPassesOnStartKickAndDrain(t *testing.T) { + coverage.Covers(t, "manager.window") + d := newTestDB(t, "") + createWindowTable(t, d.pipeline) + sink := &recordingSink{} + w, sig := newTestWatermark(t, d, testDecl(), sink) + insertBucket(t, d.pipeline, 0, "NYC", 3) + insertBucket(t, d.pipeline, 2, "NYC", 1) + assertAt(t, d.pipeline, bucket(1)) + + ctx, cancel := context.WithCancel(context.Background()) + done := make(chan error, 1) + go func() { done <- w.Start(ctx) }() + + // The start pass publishes bucket 0. + waitFor(t, "the start pass", 5*time.Second, func() bool { r, _ := sink.counts(); return r == 1 }) + insertBucket(t, d.pipeline, 3, "NYC", 1) assertAt(t, d.pipeline, bucket(3)) - failing := newTestWatermark(t, d, decl, &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) - assert.Error(t, failing.Poll(ctx)) - late, _ := metricValue(t, reader, "window_late_rows") - assert.Equal(t, int64(0), late) - - retry := newTestWatermark(t, d, decl, &recordingSink{}, WithMeterProvider(mp)) - assert.NoError(t, retry.Poll(ctx)) - assert.Equal(t, int64(2), counterValue(t, reader, "window_late_rows")) + time.Sleep(100 * time.Millisecond) + rows, _ := sink.counts() + assert.Equal(t, int64(1), rows) // nothing woke it + sig.Kick() + waitFor(t, "the kicked pass", 5*time.Second, func() bool { r, _ := sink.counts(); return r == 2 }) + + insertBucket(t, d.pipeline, 5, "NYC", 1) + assertAt(t, d.pipeline, bucket(5)) + cancel() + assert.NoError(t, <-done) + rows, _ = sink.counts() + assert.Equal(t, int64(3), rows) // the drain pass + assert.Equal(t, int64(1), countRows(t, d.pipeline, testTable)) +} + +// A context already cancelled when Start is called runs the drain pass only: +// the shape a shutdown that raced startup has, and the one the conformance +// harness's drain check builds. +func TestManagerWindow_StartOnACancelledContextDrainsOnly(t *testing.T) { + coverage.Covers(t, "manager.window") + d := newTestDB(t, "") + createWindowTable(t, d.pipeline) + sink := &recordingSink{} + w, _ := newTestWatermark(t, d, testDecl(), sink) + insertBucket(t, d.pipeline, 0, "NYC", 3) + insertBucket(t, d.pipeline, 2, "NYC", 1) + assertAt(t, d.pipeline, bucket(1)) + ctx, cancel := context.WithCancel(context.Background()) + cancel() + assert.NoError(t, w.Start(ctx)) + rows, flushes := sink.counts() + assert.Equal(t, int64(1), rows) + assert.Equal(t, 1, flushes) } -// The newest bucket is reported on every poll that finds rows, including one +// The newest bucket is reported on every pass that finds rows, including one // whose close fails. Rows stamped in the future are therefore visible the // moment they arrive, even while the sink the close publishes to is down. func TestManagerWindow_TheNewestBucketIsReportedEvenWhenTheCloseFails(t *testing.T) { @@ -282,8 +390,8 @@ func TestManagerWindow_TheNewestBucketIsReportedEvenWhenTheCloseFails(t *testing insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - failing := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) - assert.Error(t, failing.Poll(ctx)) + failing, _ := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) + assert.Error(t, failing.Pass(ctx)) assert.Equal(t, bucket(2).Unix(), gaugeValue(t, reader, "window_newest_bucket_start_seconds")) } @@ -305,22 +413,22 @@ func TestManagerWindow_AStalledCloseIsLagInEventTime(t *testing.T) { insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - ok := newTestWatermark(t, d, testDecl(), &recordingSink{}, WithMeterProvider(mp)) - assert.NoError(t, ok.Poll(ctx)) + ok, _ := newTestWatermark(t, d, testDecl(), &recordingSink{}, WithMeterProvider(mp)) + assert.NoError(t, ok.Pass(ctx)) assert.Equal(t, int64(0), gaugeValue(t, reader, "window_close_lag_seconds")) insertBucket(t, d.pipeline, 3, "NYC", 1) insertBucket(t, d.pipeline, 4, "NYC", 1) assertAt(t, d.pipeline, bucket(3)) - failing := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) - assert.Error(t, failing.Poll(ctx)) + failing, _ := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) + assert.Error(t, failing.Pass(ctx)) assert.Equal(t, int64(120), gaugeValue(t, reader, "window_close_lag_seconds")) } -// A stall that began before a restart is reported by the first poll after +// A stall that began before a restart is reported by the first pass after // it, with no close needed: the stored watermark says where the window is, // and the assertion says where it should be. -func TestManagerWindow_AStallIsReportedByTheFirstPollAfterARestart(t *testing.T) { +func TestManagerWindow_AStallIsReportedByTheFirstPassAfterARestart(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() d := newTestDB(t, "") @@ -329,16 +437,16 @@ func TestManagerWindow_AStallIsReportedByTheFirstPollAfterARestart(t *testing.T) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - before := newTestWatermark(t, d, testDecl(), &recordingSink{}) - assert.NoError(t, before.Poll(ctx)) + before, _ := newTestWatermark(t, d, testDecl(), &recordingSink{}) + assert.NoError(t, before.Pass(ctx)) insertBucket(t, d.pipeline, 4, "NYC", 1) assertAt(t, d.pipeline, bucket(3)) // A new process: fresh metrics, the same database, a sink that is down. reader := sdkmetric.NewManualReader() mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) - after := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) - assert.Error(t, after.Poll(ctx)) + after, _ := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) + assert.Error(t, after.Pass(ctx)) assert.Equal(t, int64(120), gaugeValue(t, reader, "window_close_lag_seconds")) } @@ -357,8 +465,8 @@ func TestManagerWindow_ANeverClosedWindowReportsLag(t *testing.T) { insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 4, "NYC", 1) assertAt(t, d.pipeline, bucket(3)) - failing := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) - assert.Error(t, failing.Poll(ctx)) + failing, _ := newTestWatermark(t, d, testDecl(), &failingSink{err: errors.New("sink down")}, WithMeterProvider(mp)) + assert.Error(t, failing.Pass(ctx)) assert.Equal(t, int64(120), gaugeValue(t, reader, "window_close_lag_seconds")) } @@ -372,17 +480,17 @@ func TestManagerWindow_AQuietStreamAfterAnIdleCloseIsNotLag(t *testing.T) { createWindowTable(t, d.pipeline) reader := sdkmetric.NewManualReader() mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) - w := newTestWatermark(t, d, testDecl(), &recordingSink{}, WithMeterProvider(mp)) + w, _ := newTestWatermark(t, d, testDecl(), &recordingSink{}, WithMeterProvider(mp)) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 1, "SF", 1) assertAt(t, d.pipeline, bucket(2)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) assert.Equal(t, int64(0), countRows(t, d.pipeline, testTable)) assert.Equal(t, int64(0), gaugeValue(t, reader, "window_close_lag_seconds")) for i := 0; i < 3; i++ { - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) assert.Equal(t, int64(0), gaugeValue(t, reader, "window_close_lag_seconds")) } } @@ -398,32 +506,6 @@ func TestManagerWindow_ExistsInTheMetricsBeforeItsFirstClose(t *testing.T) { assert.Equal(t, int64(0), closed) } -// reemit runs emit_sql over the late rows alone. The bucket's earlier rows -// were deleted when it closed, so a sum over the bucket after a late row is -// the late rows' sum, not the bucket's total. A sink that replaces the -// bucket's value with it loses everything published before. -func TestManagerWindow_ReemitPublishesTheLateRowsAlone(t *testing.T) { - coverage.Covers(t, "manager.window") - ctx := context.Background() - d := newTestDB(t, "") - createWindowTable(t, d.pipeline) - sink := &recordingSink{} - decl := testDecl() - decl.Late = LateReemit - decl.EmitSQL = "SELECT sum(count)::INT AS total, bucket, city FROM closed GROUP BY ALL" - w := newTestWatermark(t, d, decl, sink) - - insertBucket(t, d.pipeline, 0, "NYC", 5) - insertBucket(t, d.pipeline, 2, "NYC", 1) - assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, w.Poll(ctx)) - - insertBucket(t, d.pipeline, 0, "NYC", 1) - assert.NoError(t, w.Poll(ctx)) - - assert.DeepEqual(t, sink.published(), [][]string{{"5"}, {"1"}}) -} - // A close reads committed rows only. Rows an open transaction on another // connection has written are not in the bucket the sink receives. func TestManagerWindow_UncommittedRowsAreNotPublished(t *testing.T) { @@ -432,7 +514,7 @@ func TestManagerWindow_UncommittedRowsAreNotPublished(t *testing.T) { d := newTestDB(t, "") createWindowTable(t, d.pipeline) sink := &recordingSink{} - w := newTestWatermark(t, d, testDecl(), sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) @@ -442,53 +524,53 @@ func TestManagerWindow_UncommittedRowsAreNotPublished(t *testing.T) { defer open.Close() insertBucket(t, open, 0, "uncommitted", 9) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, _ := sink.counts() assert.Equal(t, int64(1), rows) assert.NoError(t, open.(transaction).Rollback(ctx)) } // A failed flush leaves the rows and the watermark where they were, and the -// next poll publishes the same bucket rather than treating it as late. +// next pass publishes the same bucket rather than treating it as late. func TestManagerWindow_AFailedFlushLeavesEverything(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() d := newTestDB(t, "") createWindowTable(t, d.pipeline) failing := &failingSink{err: errors.New("sink down")} - w := newTestWatermark(t, d, testDecl(), failing) + w, _ := newTestWatermark(t, d, testDecl(), failing) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - assert.Error(t, w.Poll(ctx)) + assert.Error(t, w.Pass(ctx)) assert.Equal(t, int64(2), countRows(t, d.pipeline, testTable)) _, ok, err := NewStore(d.pipeline).Load(ctx, testTable) assert.NoError(t, err) assert.That(t, !ok) sink := &recordingSink{} - w2 := newTestWatermark(t, d, testDecl(), sink) - assert.NoError(t, w2.Poll(ctx)) + w2, _ := newTestWatermark(t, d, testDecl(), sink) + assert.NoError(t, w2.Pass(ctx)) rows, _ := sink.counts() assert.Equal(t, int64(1), rows) } -// A poll that finds nothing ends its transaction, so the next poll sees rows +// A pass that finds nothing ends its transaction, so the next pass sees rows // and assertions committed in between. -func TestManagerWindow_AnEmptyPollDoesNotFreezeTheView(t *testing.T) { +func TestManagerWindow_AnEmptyPassDoesNotFreezeTheView(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() d := newTestDB(t, "") createWindowTable(t, d.pipeline) sink := &recordingSink{} - w := newTestWatermark(t, d, testDecl(), sink) + w, _ := newTestWatermark(t, d, testDecl(), sink) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, _ := sink.counts() assert.Equal(t, int64(1), rows) } @@ -504,21 +586,21 @@ func TestManagerWindow_EmitSQLShapesTheClosedRows(t *testing.T) { decl := testDecl() decl.EmitSQL = `WITH totals AS (SELECT city, sum(count)::INT AS count FROM closed GROUP BY ALL) SELECT city, count FROM totals ORDER BY city` - w := newTestWatermark(t, d, decl, sink) + w, _ := newTestWatermark(t, d, decl, sink) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 0, "NYC", 4) insertBucket(t, d.pipeline, 0, "SF", 1) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) rows, _ := sink.counts() assert.Equal(t, int64(2), rows) assert.DeepEqual(t, [][]string{{"NYC", "SF"}}, sink.published()) assert.Equal(t, int64(1), countRows(t, d.pipeline, testTable)) } -// The metrics: the watermark as Unix time, closes, and late rows. +// The metrics: the watermark as Unix time, closes, and recomputes. func TestManagerWindow_MetricsReportTheWatermark(t *testing.T) { coverage.Covers(t, "manager.window") ctx := context.Background() @@ -526,27 +608,34 @@ func TestManagerWindow_MetricsReportTheWatermark(t *testing.T) { createWindowTable(t, d.pipeline) reader := sdkmetric.NewManualReader() mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) - w := newTestWatermark(t, d, testDecl(), &recordingSink{}, WithMeterProvider(mp)) + decl := testDecl() + decl.Lateness = 5 * time.Minute + w, sig := newTestWatermark(t, d, decl, &recordingSink{}, WithMeterProvider(mp)) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) - assert.NoError(t, w.Poll(ctx)) + assert.NoError(t, w.Pass(ctx)) assert.Equal(t, bucket(1).Unix(), gaugeValue(t, reader, "window_watermark_seconds")) assert.Equal(t, int64(1), counterValue(t, reader, "window_closed")) + + sig.Recompute(bucket(0)) + assert.NoError(t, w.Pass(ctx)) + assert.Equal(t, int64(1), counterValue(t, reader, "window_recomputes")) + assert.Equal(t, int64(1), counterValue(t, reader, "window_closed")) } -// Start, cancel, final poll: the shape #270 and #273 gave the loop. A sink -// that never answers the final poll is bounded by the drain deadline, the +// Start, cancel, final pass: the shape #270 and #273 gave the loop. A sink +// that never answers the final pass is bounded by the drain deadline, the // bucket stays, and the error carries the drain code. -func TestManagerWindow_FinalPollStopsAtTheDrainDeadline(t *testing.T) { +func TestManagerWindow_FinalPassStopsAtTheDrainDeadline(t *testing.T) { coverage.Covers(t, "manager.window") d := newTestDB(t, "") createWindowTable(t, d.pipeline) budget := core.NewDrainBudget(200 * time.Millisecond) defer budget.Stop() - w := newTestWatermark(t, d, testDecl(), hangingSink{}, WithDrainBudget(budget)) + w, _ := newTestWatermark(t, d, testDecl(), hangingSink{}, WithDrainBudget(budget)) insertBucket(t, d.pipeline, 0, "NYC", 3) insertBucket(t, d.pipeline, 2, "NYC", 1) assertAt(t, d.pipeline, bucket(1)) @@ -559,7 +648,7 @@ func TestManagerWindow_FinalPollStopsAtTheDrainDeadline(t *testing.T) { case err := <-done: assert.Equal(t, errs.CodeDrainIncomplete, errs.CodeOf(err)) case <-time.After(5 * time.Second): - t.Fatal("the final poll outlived the drain deadline") + t.Fatal("the final pass outlived the drain deadline") } assert.Equal(t, int64(2), countRows(t, d.pipeline, testTable)) } diff --git a/internal/managers/window_leak_test.go b/internal/managers/window_leak_test.go index 9fc01b78..37f313a5 100644 --- a/internal/managers/window_leak_test.go +++ b/internal/managers/window_leak_test.go @@ -104,7 +104,6 @@ func (s leakScenario) declaration() Declaration { TimeColumn: "bucket", Size: time.Minute, Grace: time.Duration(leakCloseAfter-1) * time.Minute, - Late: LateReemit, } if !s.upsert { d.EmitSQL = "SELECT bucket, lang, sum(posts) AS posts FROM closed GROUP BY bucket, lang" @@ -237,7 +236,7 @@ func leakLoop(tb testing.TB, sc leakScenario, batches int) (before, after leakSa // The manager on its own connection, as run builds it. sink := &recordingSink{} - m, err := NewWatermark(mconn, sc.declaration(), time.Hour, sink) + m, err := NewWatermark(mconn, sc.declaration(), sink, nil) if err != nil { tb.Fatal(err) } @@ -278,9 +277,9 @@ func leakLoop(tb testing.TB, sc leakScenario, batches int) (before, after leakSa if err := watermarks.Save(ctx, "w", newest.Add(-sc.declaration().Grace)); err != nil { tb.Fatal(err) } - // The pipeline polls on a timer, about six times a minute at the - // demo's rate. Once per batch keeps the ratio close enough. - if err := m.Poll(ctx); err != nil { + // The engine kicks the manager after every commit that moved the + // watermark; one pass per batch is that. + if err := m.Pass(ctx); err != nil { tb.Fatal(err) } // Init is where StructuredBatch truncates and checkpoints, after diff --git a/internal/managers/window_leakloop_test.go b/internal/managers/window_leakloop_test.go index 926c3a2c..e2c869cc 100644 --- a/internal/managers/window_leakloop_test.go +++ b/internal/managers/window_leakloop_test.go @@ -83,9 +83,8 @@ func runDemoWindow(t *testing.T, name string, sink core.Sink, setup func(*testDB Size: time.Minute, Grace: time.Minute, IdleClose: time.Minute, - Late: LateDrop, EmitSQL: "SELECT bucket, lang, sum(posts)::INTEGER AS posts FROM closed GROUP BY bucket, lang", - }, time.Hour, sink) + }, sink, nil) if err != nil { t.Fatal(err) } @@ -114,9 +113,8 @@ func runDemoWindow(t *testing.T, name string, sink core.Sink, setup func(*testDB if err := progress.Record(ctx, core.Progress{LastArrival: now, LastCommit: now, Messages: consumed}); err != nil { t.Fatal(err) } - // One poll per batch: on Render a batch of 500 lands every 10 to 17 - // seconds and the manager polls every 10. - if err := m.Poll(ctx); err != nil { + // One pass per batch: the kick the engine sends after each commit. + if err := m.Pass(ctx); err != nil { t.Fatal(err) } if err := h.Init(ctx); err != nil { diff --git a/internal/simulate/README.md b/internal/simulate/README.md index 2757404a..f6114794 100644 --- a/internal/simulate/README.md +++ b/internal/simulate/README.md @@ -3,15 +3,16 @@ A scripted run of the **real** consume loop and the **real** window manager over one **real** DuckDB, with a fake Kafka coordinator underneath. You write a sequence of events; it tells you how many rows were produced, published, -still open, and lost. +still open, refused, and what the sink holds for each bucket. ```go r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, }) -// r.Produced, r.Published, r.StillOpen, r.LateDropped, r.Republished +// r.Produced, r.Published, r.StillOpen, r.LateRefused, r.Recomputes, +// r.LastValue[bucket], r.Republished ``` ## Why it exists @@ -47,13 +48,13 @@ the scenarios below expressible. ## How to read the tables Every scenario carries a table in its doc comment. The columns are what the -manager could see if it polled at that moment: +manager sees when it passes at that moment: ``` step what the script does event time the row's own clock bucket event time truncated to the window size -window table what the handler has written, and what a poll would read +window table what the handler has written, and what a pass would read asserted the engine's watermark after the step's commit: the minimum over the partitions that could still deliver, each at its newest event time less the grace. A bucket closes when its @@ -66,7 +67,7 @@ clock of its own, no progress row, no reading of the table's newest bucket. Worked example — `OutOfOrderRowsInsideTheGraceAreNotLate`, under `windowDecl()`: one-minute buckets, one minute of grace, a ten-second idle -close, late rows dropped. +close, no lateness. ``` step event time bucket window table asserted @@ -76,11 +77,28 @@ Produce 2 12:01:05 12:01 12:01: 5 12:00:30 ^ older than the row before it, same bucket, not late: the newest event time is still 12:01:30, so nothing moved Produce 4 12:05:00 12:05 12:01: 5, 12:05: 4 12:04 -Poll - - 12:05: 4 12:04 +Pass - - 12:05: 4 12:04 ^ 12:01 ends at 12:02, at or before 12:04, so it publishes all 5 rows. 12:05 ends at 12:06 and stays open. ``` +And with lateness -- `ALateRowRecomputesTheBucket`, `Lateness: 10 * +time.Minute`. The sink holds the last value it was handed for each bucket, +as a sink that replaces by key does, and that value is always the whole +bucket: + +``` +step event time bucket window table asserted sink's value for 12:00 +----------------------------------------------------------------------------------------- +Produce 5 12:00:30 12:00 12:00: 5 11:59:30 - +Produce 4 12:05:00 12:05 12:00: 5, 12:05: 4 12:04 - +Pass 12:00: 5, 12:05: 4 12:04 5 kept: lateness 10m +Produce 1 12:00:45 12:00 12:00: 6, 12:05: 4 12:04 5 late, allowed: recompute queued +Pass (same) 12:04 6 the whole bucket +Produce 2 12:00:50 12:00 12:00: 8, 12:05: 4 12:04 6 +Pass (same) 12:04 8 +``` + ## The four rules that explain every outcome 1. **While partitions deliver**, the assertion is the minimum over them of @@ -98,14 +116,19 @@ Poll - - 12:05: 4 12:04 holding for them would freeze this worker's windows for good. Idleness is measured from the later of a partition's last row and its assignment, so an outage counts for nothing and the clock restarts on the return. -4. **Late collection happens on a close, not on a poll.** The manager - collects late rows only on a poll that has something to do. A run that - ends without a close reports whatever the timing gave, which is why every - scenario asserting a loss ends by forcing one. - -Rule 4 is a property of the implementation rather than of the design, and it -made three scenarios flaky before it was understood: one read six drops -plainly and four under `-race`. +4. **Lateness is decided at arrival, by the engine.** A record whose bucket + ended at or before the assertion is late. With no lateness it is refused + before the handler and `LateRefused` counts it; it never reaches the + table. With lateness, a late record within it is written and the bucket + is republished whole on the next pass -- `Recomputes` counts the + republications, `LastValue` is what the sink holds. Past + `end + lateness` the rows are deleted and later records are refused. + +Nothing polls. The manager passes when it starts, when the engine kicks it +after a commit that moved the watermark or admitted a late row, and when it +drains. A script drives those passes with `Pass` so a scenario is a +sequence; `RunWindowedDriven` runs the manager's own loop, and +`TheKickFollowsTheCommit` is the one scenario that needs it. ## The scenarios @@ -139,7 +162,12 @@ Each test's doc comment carries its own table. This is the index. | scenario | what it proves | |---|---| | `OutOfOrderRowsInsideTheGraceAreNotLate` | the case `grace_seconds` exists for: records shuffled by less than the tolerance are not late | -| `ARowForAClosedBucketIsDropped` | the contract `late_rows: drop` promises | +| `ARowForAClosedBucketIsRefused` | rule 4 with no lateness: refused before the handler, counted, and the sink's value stands | +| `ALateRowRecomputesTheBucket` | rule 4 with lateness: written, and the bucket republished as a whole value | +| `ALateRowBeyondLatenessIsRefused` | the lateness bound: past `end + lateness` a late row is refused even though the watermark only just passed | +| `ABurstOfLateRowsIsOneRecompute` | several late rows before a pass republish the bucket once | +| `ARecomputeSurvivesARestart` | the recompute set is in memory; the start pass republishes every retained bucket for it | +| `TheKickFollowsTheCommit` | with the manager's loop running, the commit's kick publishes with no `Pass` step | | `AFastPartitionHoldsForTheSlowOne` | rule 1's minimum: the lagging partition holds the bucket both were filling, and its rows are never late | | `APoisonTimestampClosesEveryBucket` | one wrong clock costs the pipeline exactly the record that carried it | | `ReplayingEveryRowPublishesTheSameTotals` | **a defect** — see below | @@ -188,14 +216,16 @@ replay safe — and is what lets a destination skip deduplicating at read time. The published and still-open counts are pinned in the windowed scenarios now, because the close no longer races the run's own stop: a bucket closes when the engine has asserted past its end, and the assertion is committed before -the commit that carries it is visible. A scenario that ends without a poll -still leaves whatever the last poll did not close. +the commit that carries it is visible. A scenario that ends without a pass +still leaves whatever the last pass did not close. ## Adding a scenario 1. Write the script. `Produce{Partition, Rows, At}`, `IdleTick`, `Elapse`, - `Revoke`, `Assign`, `Restart`, `Poll`. -2. If it asserts a loss, end with `Elapse` / `IdleTick` / `Poll` — rule 4. + `Revoke`, `Assign`, `Restart`, `Pass`, `StartPass`; in a driven run, + `AwaitPublished`. +2. If it is about lateness, set `decl.Lateness` and assert `LastValue` for + the bucket: the sink's value is the claim, not the row count. 3. Run it and *look at the numbers* before writing the assertion. Several of these behaved differently from how their author expected, and the difference was the interesting part. diff --git a/internal/simulate/event_time_test.go b/internal/simulate/event_time_test.go index 4820fd59..de4a1da1 100644 --- a/internal/simulate/event_time_test.go +++ b/internal/simulate/event_time_test.go @@ -15,7 +15,7 @@ var base = time.Date(2026, 9, 23, 12, 0, 0, 0, time.UTC) // Reading the diagrams in this file // // Every scenario runs under windowDecl(): one-minute buckets, one minute of -// grace, a ten-second idle close, and late rows dropped. +// grace, a ten-second idle close, and no lateness, so a late row is refused. // // Two clocks, and keeping them apart is the point of the whole file: // @@ -25,12 +25,12 @@ var base = time.Date(2026, 9, 23, 12, 0, 0, 0, time.UTC) // Produce{At}. This is what a bucket is cut on, and what the // watermark is computed from. // -// The columns are what the manager could see if it polled at that moment: +// The columns are what the manager sees when it passes at that moment: // // step what the script does // event time the row's own clock // bucket where that event time falls: event time truncated to a minute -// window table what the handler has written, and what a poll would read +// window table what the handler has written, and what a pass would read // asserted the engine's watermark after the step's commit: for each // partition that holds and is not idle, its newest event time // less the grace, and the minimum over them. A bucket closes @@ -39,9 +39,9 @@ var base = time.Date(2026, 9, 23, 12, 0, 0, 0, time.UTC) // Two rules explain every outcome below. While partitions deliver, the // assertion is `min over partitions of (newest event time - grace)`; when // every partition has been silent for the idle bound, it is `newest event -// time + size`, which closes the newest bucket too. And the manager collects -// late rows only on a poll that finds something to do, which is why the -// scenarios that assert a loss end by forcing a close. +// time + size`, which closes the newest bucket too. And a record's lateness +// is decided by the engine when it arrives, against that assertion: refused +// before the handler with no lateness, written and recomputed within it. // Rows out of order inside the grace all land in their own buckets, and the // stream moving on closes them. This is the case grace_seconds exists for: a @@ -54,7 +54,7 @@ var base = time.Date(2026, 9, 23, 12, 0, 0, 0, time.UTC) // ^ older than the row before it, same bucket, not late: the // newest event time is still 12:01:30, so nothing moved // Produce 4 12:05:00 12:05 12:01: 5, 12:05: 4 12:04 -// Poll - - 12:05: 4 12:04 +// Pass - - 12:05: 4 12:04 // ^ 12:01 ends at 12:02, at or before 12:04, so it publishes // all 5 rows. 12:05 ends at 12:06 and stays open. func TestSimulate_OutOfOrderRowsInsideTheGraceAreNotLate(t *testing.T) { @@ -65,56 +65,151 @@ func TestSimulate_OutOfOrderRowsInsideTheGraceAreNotLate(t *testing.T) { Produce{Partition: 0, Rows: 2, At: base.Add(65 * time.Second)}, // The stream moves on by more than the grace, which closes 12:01. Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, }) assert.Equal(t, 9, r.Produced) assert.Equal(t, int64(5), r.Published) assert.Equal(t, int64(4), r.StillOpen) - assert.Equal(t, int64(0), r.LateDropped) + assert.Equal(t, int64(0), r.LateRefused) assert.Equal(t, 0, r.Republished) } -// A row for a bucket the watermark has passed is late, and late_rows drop -// discards it. That is the declared contract, so this is what correct looks -// like rather than a defect. +// A row for a bucket the watermark has passed is late. With no lateness the +// engine refuses it before the handler: it never reaches the table, and the +// sink's value for the bucket is the one the close published. That is the +// declared contract, so this is what correct looks like rather than a defect. // -// step event time bucket window table asserted -// ----------------------------------------------------------------------- +// step event time bucket window table asserted engine +// ------------------------------------------------------------------------------- // Produce 5 12:00:30 12:00 12:00: 5 11:59:30 // Produce 4 12:05:00 12:05 12:00: 5, 12:05: 4 12:04 -// Poll - - 12:05: 4 12:04 -// ^ 12:00 publishes its 5 rows -// Produce 6 12:00:30 12:00 12:00: 6, 12:05: 4 12:04 +// Pass - - 12:05: 4 12:04 12:00 → 5 published, deleted +// Produce 6 12:00:30 (none) 12:05: 4 12:04 end 12:01 <= 12:04: refused, 6 counted // ^ back into a bucket that ended at 12:01, already at or -// before the watermark. These 6 are late on arrival, and -// do not move the assertion: 12:00:30 is not the newest. -// Elapse 30s -// IdleTick - - (same) 12:06 -// ^ the partition has been silent 31s > 10s: idle, and it is -// the only one, so newest 12:05 + 1m -// Poll - - empty 12:06 -// ^ 12:05 publishes its 4 rows, and the 6 late rows are -// collected and dropped. -func TestSimulate_ARowForAClosedBucketIsDropped(t *testing.T) { +// before the watermark. Late on arrival, refused before +// the handler, and the assertion does not move: 12:00:30 +// is not the newest. +func TestSimulate_ARowForAClosedBucketIsRefused(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, // Back in the bucket that just closed. Produce{Partition: 0, Rows: 6, At: base.Add(30 * time.Second)}, - // The manager collects late rows on a poll that has something to - // do, so the run reaches a close for the drop to have happened. - Elapse{By: 30 * time.Second}, - IdleTick{}, - Poll{}, }) assert.Equal(t, 15, r.Produced) - assert.Equal(t, int64(9), r.Published) - assert.Equal(t, int64(0), r.StillOpen) - assert.Equal(t, int64(6), r.LateDropped) + assert.Equal(t, int64(5), r.Published) + assert.Equal(t, int64(4), r.StillOpen) + assert.Equal(t, int64(6), r.LateRefused) + assert.Equal(t, int64(5), r.LastValue[base]) + assert.Equal(t, int64(0), r.Recomputes) +} + +// With lateness, a late row is written and the bucket republished whole: the +// sink's last value is the exact count, never the late rows alone. +// +// step event time bucket window table asserted sink's value for 12:00 +// ----------------------------------------------------------------------------------------- +// Produce 5 12:00:30 12:00 12:00: 5 11:59:30 - +// Produce 4 12:05:00 12:05 12:00: 5, 12:05: 4 12:04 - +// Pass 12:00: 5, 12:05: 4 12:04 5 kept: lateness 10m +// Produce 1 12:00:45 12:00 12:00: 6, 12:05: 4 12:04 5 late, allowed: recompute queued +// Pass (same) 12:04 6 the whole bucket +// Produce 2 12:00:50 12:00 12:00: 8, 12:05: 4 12:04 6 +// Pass (same) 12:04 8 +func TestSimulate_ALateRowRecomputesTheBucket(t *testing.T) { + coverage.Covers(t, "manager.window") + decl := windowDecl() + decl.Lateness = 10 * time.Minute + r := RunWindowed(t, []int32{0}, decl, []Step{ + Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, + Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, + Pass{}, + Produce{Partition: 0, Rows: 1, At: base.Add(45 * time.Second)}, + Pass{}, + Produce{Partition: 0, Rows: 2, At: base.Add(50 * time.Second)}, + Pass{}, + }) + assert.Equal(t, 12, r.Produced) + assert.Equal(t, int64(8), r.LastValue[base]) + assert.Equal(t, int64(0), r.LateRefused) + assert.Equal(t, int64(2), r.Recomputes) + // The bucket is retained: its rows are still in the table. + assert.Equal(t, int64(12), r.StillOpen) +} + +// Beyond lateness a late row is refused even though the watermark has only +// just passed: 12:00 ended at 12:01, the lateness is two minutes, and the +// assertion is 12:04, so 12:01 + 2m = 12:03 <= 12:04 had already expired it. +func TestSimulate_ALateRowBeyondLatenessIsRefused(t *testing.T) { + coverage.Covers(t, "manager.window") + decl := windowDecl() + decl.Lateness = 2 * time.Minute + r := RunWindowed(t, []int32{0}, decl, []Step{ + Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, + Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, + Pass{}, + Produce{Partition: 0, Rows: 1, At: base.Add(45 * time.Second)}, + }) + assert.Equal(t, int64(1), r.LateRefused) + assert.Equal(t, int64(5), r.LastValue[base]) + assert.Equal(t, int64(0), r.Recomputes) + // And the pass purged the expired bucket: only 12:05 is left. + assert.Equal(t, int64(4), r.StillOpen) +} + +// Three late rows for one bucket before the next pass, one recompute. +func TestSimulate_ABurstOfLateRowsIsOneRecompute(t *testing.T) { + coverage.Covers(t, "manager.window") + decl := windowDecl() + decl.Lateness = 10 * time.Minute + r := RunWindowed(t, []int32{0}, decl, []Step{ + Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, + Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, + Pass{}, + Produce{Partition: 0, Rows: 1, At: base.Add(40 * time.Second)}, + Produce{Partition: 0, Rows: 1, At: base.Add(41 * time.Second)}, + Produce{Partition: 0, Rows: 1, At: base.Add(42 * time.Second)}, + Pass{}, + }) + assert.Equal(t, int64(8), r.LastValue[base]) + assert.Equal(t, int64(1), r.Recomputes) +} + +// A late row committed, then a restart before the manager passed: the +// recompute set was in memory and is gone, but the rows are in the table, +// and the start pass republishes every retained bucket, so the exact value +// still reaches the sink. +func TestSimulate_ARecomputeSurvivesARestart(t *testing.T) { + coverage.Covers(t, "manager.window") + decl := windowDecl() + decl.Lateness = 10 * time.Minute + r := RunWindowed(t, []int32{0}, decl, []Step{ + Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, + Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, + Pass{}, + Produce{Partition: 0, Rows: 1, At: base.Add(45 * time.Second)}, + Restart{}, + StartPass{}, // what Start does first + }) + assert.Equal(t, int64(6), r.LastValue[base]) +} + +// The kick follows the commit: with the manager's loop running, a produce +// that moves the watermark publishes without any Pass step. +func TestSimulate_TheKickFollowsTheCommit(t *testing.T) { + coverage.Covers(t, "manager.window") + r := RunWindowedDriven(t, []int32{0}, windowDecl(), []Step{ + Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, + Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, + AwaitPublished{Rows: 5}, + }) + assert.Equal(t, int64(5), r.Published) + assert.Equal(t, int64(4), r.StillOpen) + assert.Equal(t, int64(5), r.LastValue[base]) } // Two partitions, one lagging in event time. @@ -133,13 +228,13 @@ func TestSimulate_ARowForAClosedBucketIsDropped(t *testing.T) { // Produce 3 p1 12:00:30 12:00 12:00: 6 11:59:30 // Produce 4 p0 12:05:00 12:05 12:00: 6, 12:05: 4 11:59:30 // ^ p0 alone races five minutes ahead; the minimum is p1's -// Poll (same) 11:59:30 nothing closes +// Pass (same) 11:59:30 nothing closes // Produce 5 p1 12:00:45 12:00 12:00: 11, 12:05: 4 11:59:45 // ^ p1's share of the bucket, and it is not late -// Poll (same) 11:59:45 still open +// Pass (same) 11:59:45 still open // Produce 1 p1 12:05:00 12:05 12:00: 11, 12:05: 5 12:04 // ^ p1 catches up; the minimum follows -// Poll 12:05: 5 12:04 12:00 publishes 11 +// Pass 12:05: 5 12:04 12:00 publishes 11 func TestSimulate_AFastPartitionHoldsForTheSlowOne(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0, 1}, windowDecl(), []Step{ @@ -148,19 +243,19 @@ func TestSimulate_AFastPartitionHoldsForTheSlowOne(t *testing.T) { Produce{Partition: 1, Rows: 3, At: base.Add(30 * time.Second)}, // Partition 0 races five minutes ahead in event time. Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, // Partition 1 is still delivering its own share of the early bucket. Produce{Partition: 1, Rows: 5, At: base.Add(45 * time.Second)}, - Poll{}, + Pass{}, // And catches up, which is what closes the bucket both filled. Produce{Partition: 1, Rows: 1, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, }) assert.Equal(t, 16, r.Produced) assert.Equal(t, int64(11), r.Published) assert.Equal(t, int64(5), r.StillOpen) - assert.Equal(t, int64(0), r.LateDropped) + assert.Equal(t, int64(0), r.LateRefused) assert.Equal(t, 0, r.Republished) } @@ -175,13 +270,13 @@ func TestSimulate_AFastPartitionHoldsForTheSlowOne(t *testing.T) { // engine cannot place an event time ahead of its own // clock, so the record never reaches the handler, never // enters the table, and never reaches the watermark. -// Poll - - 12:00: 5 11:59:30 +// Pass - - 12:00: 5 11:59:30 // Produce 7 12:01:30 12:01 12:00: 5, 12:01: 7 12:00:30 // ^ the stream carries on where it actually is, and it is // not late, because nothing moved -// Poll, Elapse 30s, IdleTick 12:02:30 +// Pass, Elapse 30s, IdleTick 12:02:30 // ^ idle: newest 12:01:30 + 1m -// Poll empty 12:02:30 +// Pass empty 12:02:30 // ^ both buckets publish, all 12 rows. The bad record cost // the pipeline exactly itself. // @@ -195,21 +290,22 @@ func TestSimulate_AFastPartitionHoldsForTheSlowOne(t *testing.T) { // The record is refused before the handler, not held in the table: it is // not late, because no earlier publication of its bucket exists to amend, // and it is not open, because no watermark this engine computes will reach -// it. Unplaceable, so discarded, and counted in messages_unplaceable_total. +// it. Unplaceable, so discarded, and counted in messages_unplaceable_total, +// not as late: it never had a bucket to be late for. func TestSimulate_APoisonTimestampClosesEveryBucket(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, // One row from a device that thinks it is 2099. Produce{Partition: 0, Rows: 1, At: time.Date(2099, 1, 1, 0, 0, 0, 0, time.UTC)}, - Poll{}, + Pass{}, // The stream carries on where it actually is. Produce{Partition: 0, Rows: 7, At: base.Add(90 * time.Second)}, - Poll{}, - // A close at the end, so late collection has certainly happened. + Pass{}, + // A close at the end, so every honest bucket has published. Elapse{By: 30 * time.Second}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 13, r.Produced) @@ -217,7 +313,8 @@ func TestSimulate_APoisonTimestampClosesEveryBucket(t *testing.T) { // dishonest one, refused before it could reach the table. assert.Equal(t, int64(12), r.Published) assert.Equal(t, int64(0), r.StillOpen) - assert.Equal(t, int64(1), r.LateDropped) + assert.Equal(t, int64(1), r.Unplaceable) + assert.Equal(t, int64(0), r.LateRefused) } // Replay: the same rows delivered twice. @@ -236,7 +333,7 @@ func TestSimulate_APoisonTimestampClosesEveryBucket(t *testing.T) { // event times // Produce 4 12:05:00 -> 12:05: 4 Produce 4 -> 12:05: 4 // Produce 4 -> 12:05: 8 -// Poll publishes 12:00, 5 rows Poll -> publishes 10 +// Pass publishes 12:00, 5 rows Poll -> publishes 10 // // The engine has no record identity, so a replayed row is simply a new row // and the bucket's value doubles. Dedupe on an observation id is what closes @@ -246,7 +343,7 @@ func TestSimulate_ReplayingEveryRowPublishesTheSameTotals(t *testing.T) { once := RunWindowed(t, []int32{0}, windowDecl(), []Step{ Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, }) twice := RunWindowed(t, []int32{0}, windowDecl(), []Step{ @@ -256,7 +353,7 @@ func TestSimulate_ReplayingEveryRowPublishesTheSameTotals(t *testing.T) { Produce{Partition: 0, Rows: 5, At: base.Add(30 * time.Second)}, Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, }) // The engine has no record identity, so a replayed row is a new row and diff --git a/internal/simulate/simulate.go b/internal/simulate/simulate.go index fb421349..b04161ff 100644 --- a/internal/simulate/simulate.go +++ b/internal/simulate/simulate.go @@ -441,6 +441,9 @@ func (r *run) sinkTurbine() { opts = append(opts, core.WithProgressStore(core.NewProgressStore(r.window.db.pipeline)), core.WithWindows(w, core.NewWatermarkStore(r.window.db.pipeline)), + // The engine's instruments, so the run can report what it refused + // from the counter rather than from a residual. + core.WithMetrics(r.window.metrics), // Every commit that moves a watermark writes it. Production paces // the write at a second, which a script's steps happen to clear // today; pinning it here keeps a scenario with finer steps from @@ -483,18 +486,24 @@ func (r *run) deliver(p int32, ids []int64) { }) } // What the loop will accept. A windowing run places event time, and a - // record the engine cannot place never reaches the handler or the table, - // so waiting for it would wait forever. The prediction uses the engine's - // own rule, against the same clock it reads, so it cannot disagree with - // what the engine does. + // record the engine cannot place, or refuses as late, never reaches the + // handler or the table, so waiting for it would wait forever. The + // prediction uses the engine's own rules -- CanPlace against the clock it + // reads, Classify against the watermark it has asserted -- so it cannot + // disagree with what the engine does. Every earlier batch has committed + // by the time this runs, so the assertion Classify reads is current. accepted := len(batch) if r.window != nil { nowNanos := time.Now().UnixNano() accepted = 0 for _, m := range batch { - if core.CanPlace(m.EventAtNanos, nowNanos) { - accepted++ + if !core.CanPlace(m.EventAtNanos, nowNanos) { + continue } + if refused, _ := r.window.tracker.Classify(m.EventAtNanos); refused { + continue + } + accepted++ } } var want, commitsBefore int diff --git a/internal/simulate/window.go b/internal/simulate/window.go index 77d7285a..9c0c5af8 100644 --- a/internal/simulate/window.go +++ b/internal/simulate/window.go @@ -17,22 +17,38 @@ import ( "github.com/turbolytics/sql-flow/internal/core" "github.com/turbolytics/sql-flow/internal/duckdb" "github.com/turbolytics/sql-flow/internal/managers" + "go.opentelemetry.io/otel/attribute" + sdkmetric "go.opentelemetry.io/otel/sdk/metric" + "go.opentelemetry.io/otel/sdk/metric/metricdata" ) // The windowed runner pairs the real consume loop with a real window manager // over one DuckDB, which is the interaction neither the model nor the manager // tests cover: the model has no engine under it, and the manager tests write -// the progress row by hand rather than having a loop write it. +// the watermark by hand rather than having a loop write it. // // The handler does what a windowed pipeline's handler does: it inserts a row // per batch into the window table, bucketed by event time, and the pipeline -// sink receives nothing. The manager polls its own connection, as it does in -// production. +// sink receives nothing. The manager runs on its own connection, as it does +// in production, and the engine decides each record's lateness before the +// handler sees it. A script drives the manager's passes by hand, so a +// scenario is a sequence; RunWindowedDriven runs the manager's own loop +// instead, so the kick that follows a commit can be seen to work. const windowTable = "sim_window" -// Poll runs the manager once, as its ticker would. -type Poll struct{} +// Pass runs the manager once, as the engine's kick after a commit would. +type Pass struct{} + +// StartPass runs the pass Start runs first: a Pass that also republishes, +// whole, every bucket still retained under the window's lateness. It is +// what a restarted manager does, so a script that Restarts and wants the +// manager's recovery says so with this step. +type StartPass struct{} + +// AwaitPublished waits, in a driven run, until the window's sink holds at +// least Rows rows: the kick has been sent and the manager has acted on it. +type AwaitPublished struct{ Rows int64 } // windowDB is the pipeline's connection and the manager's, over one database. type windowDB struct { @@ -158,7 +174,9 @@ type windowHandler struct { func (h *windowHandler) Init(context.Context) error { return nil } // Write reads the record's event time out of the payload, the way a real -// windowed pipeline's SQL reads it out of the message, and buckets on it. +// windowed pipeline's SQL reads it out of the message, and buckets on it +// the way time_bucket does: from DuckDB's origin, which is the bucket the +// engine decided the record's lateness against. func (h *windowHandler) Write(msg []byte) error { at := h.clock() if _, after, ok := strings.Cut(string(msg), ","); ok { @@ -172,7 +190,7 @@ func (h *windowHandler) Write(msg []byte) error { if h.buffered == nil { h.buffered = map[time.Time]int64{} } - h.buffered[at.UTC().Truncate(h.size)]++ + h.buffered[core.BucketStart(at.UTC(), h.size)]++ h.n++ return nil } @@ -213,14 +231,22 @@ func (h *windowHandler) RowsRead() int64 { return h.n } -// windowSink counts the rows every close published. +// windowSink counts the rows every close published, and keeps the last value +// it was handed for each bucket, the way a sink that replaces by key holds +// it: a bucket published twice is worth its second value, not the sum. type windowSink struct { mu sync.Mutex rows int64 buckets map[int64]int + // pending is this flush's rows per bucket; last is what the sink holds + // for each bucket once the flush lands. + pending map[int64]int64 + last map[int64]int64 } -func newWindowSink() *windowSink { return &windowSink{buckets: map[int64]int{}} } +func newWindowSink() *windowSink { + return &windowSink{buckets: map[int64]int{}, last: map[int64]int64{}} +} func (s *windowSink) WriteTable(_ context.Context, t arrow.Table) error { s.mu.Lock() @@ -237,31 +263,49 @@ func (s *windowSink) WriteTable(_ context.Context, t arrow.Table) error { if nCol < 0 { return fmt.Errorf("published table has no n column") } - for _, chunk := range t.Column(nCol).Data().Chunks() { + nChunks := t.Column(nCol).Data().Chunks() + var bucketChunks []arrow.Array + if bucketCol >= 0 { + bucketChunks = t.Column(bucketCol).Data().Chunks() + } + if s.pending == nil { + s.pending = map[int64]int64{} + } + // Counted once per flush: one close publishes several rows for a + // bucket, and republication is the same bucket in two closes. + inThisFlush := map[int64]bool{} + for c, chunk := range nChunks { ns := chunk.(*array.Int64) + var ts *array.Timestamp + if bucketChunks != nil { + ts, _ = bucketChunks[c].(*array.Timestamp) + } for i := 0; i < ns.Len(); i++ { s.rows += ns.Value(i) - } - } - if bucketCol >= 0 { - // Counted once per flush: one close publishes several rows for a - // bucket, and republication is the same bucket in two closes. - inThisFlush := map[int64]bool{} - for _, chunk := range t.Column(bucketCol).Data().Chunks() { - if ts, ok := chunk.(*array.Timestamp); ok { - for i := 0; i < ts.Len(); i++ { - inThisFlush[int64(ts.Value(i))] = true - } + if ts != nil { + b := int64(ts.Value(i)) + inThisFlush[b] = true + s.pending[b] += ns.Value(i) } } - for b := range inThisFlush { - s.buckets[b]++ - } + } + for b := range inThisFlush { + s.buckets[b]++ } return nil } -func (s *windowSink) Flush(context.Context) error { return nil } +// Flush is where a value lands: what this flush carried for a bucket +// replaces what the sink held for it. +func (s *windowSink) Flush(context.Context) error { + s.mu.Lock() + defer s.mu.Unlock() + for b, n := range s.pending { + s.last[b] = n + } + s.pending = nil + return nil +} func (s *windowSink) counts() (rows int64, republished int) { s.mu.Lock() @@ -274,21 +318,73 @@ func (s *windowSink) counts() (rows int64, republished int) { return s.rows, republished } +// lastValues is what the sink holds per bucket start. +func (s *windowSink) lastValues() map[time.Time]int64 { + s.mu.Lock() + defer s.mu.Unlock() + out := make(map[time.Time]int64, len(s.last)) + for b, n := range s.last { + out[time.UnixMicro(b).UTC()] = n + } + return out +} + // WindowResult is what a windowed run produced and what its window published. type WindowResult struct { - Produced int - Published int64 - StillOpen int64 + Produced int + // Published sums every row the sink was handed, across every flush; a + // republished bucket counts each time. + Published int64 + // StillOpen is what the window table holds at the end. + StillOpen int64 + // Republished is how many times a bucket reached the sink after its + // first time. Republished int - LateDropped int64 + // LateRefused is what the engine refused before the handler: records for + // a bucket that had closed more than the lateness before the watermark. + // From the engine's own counter. + LateRefused int64 + // Unplaceable is what the engine could not place at all: a record stamped + // beyond its own clock, refused before the handler and never late. + Unplaceable int64 + // Recomputes is how many buckets the manager republished whole because a + // late row landed in them. + Recomputes int64 + // LastValue is what the sink holds for each bucket start: the last value + // it was handed, as a sink that replaces by key holds it. + LastValue map[time.Time]int64 } // RunWindowed replays a script against a real consume loop and a real window -// manager over one database, and reports what the window published. +// manager over one database, and reports what the window published. The +// script drives the manager with Pass steps. func RunWindowed(t *testing.T, owned []int32, decl managers.Declaration, script []Step) WindowResult { + t.Helper() + return runWindowed(t, owned, decl, script, false) +} + +// RunWindowedDriven is RunWindowed with the manager's own loop running: +// the engine's kick after each commit is what makes it pass, and a script +// waits with AwaitPublished rather than passing by hand. +func RunWindowedDriven(t *testing.T, owned []int32, decl managers.Declaration, script []Step) WindowResult { + t.Helper() + return runWindowed(t, owned, decl, script, true) +} + +func runWindowed(t *testing.T, owned []int32, decl managers.Declaration, script []Step, driven bool) WindowResult { t.Helper() db := openWindowDB(t) + // The engine's and the manager's instruments, read at the end: what was + // refused and what was recomputed come from the counters, not from a + // residual, so a row lost some other way is not reported as refused. + reader := sdkmetric.NewManualReader() + mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) + metrics, err := core.NewMetrics(mp) + if err != nil { + t.Fatal(err) + } + r := &run{ t: t, coord: newCoordinator(), @@ -297,42 +393,95 @@ func RunWindowed(t *testing.T, owned []int32, decl managers.Declaration, script owned: owned, eventAt: map[int64]time.Time{}, } - r.window = &windowRun{db: db, decl: decl, sink: newWindowSink()} + r.window = &windowRun{db: db, decl: decl, sink: newWindowSink(), metrics: metrics, reader: reader} r.window.handler = &windowHandler{t: t, conn: db.pipeline, clock: r.clock, size: decl.Size} - wm, err := managers.NewWatermark(db.manager, decl, time.Hour, r.window.sink) + // One signal for the window's lifetime, shared by every tracker the run + // builds: a Restart rebuilds the tracker and the manager keeps its signal, + // as a manager does across the engine's own restarts. + r.window.signal = core.NewWatermarks([]core.WindowSpec{{Name: decl.Table}}, r.clock).Signal(decl.Table) + wm, err := managers.NewWatermark(db.manager, decl, r.window.sink, r.window.signal, managers.WithMeterProvider(mp)) if err != nil { t.Fatal(err) } r.window.manager = wm r.start() + var ( + managerDone chan error + cancelManager context.CancelFunc + ) + if driven { + ctx, cancel := context.WithCancel(context.Background()) + cancelManager = cancel + managerDone = make(chan error, 1) + go func() { managerDone <- wm.Start(ctx) }() + } for _, s := range script { r.elapse() s.apply(r) } r.stop() + if driven { + cancelManager() + select { + case err := <-managerDone: + if err != nil { + t.Fatalf("the manager's loop failed: %v", err) + } + case <-time.After(5 * time.Second): + t.Fatal("the manager's loop did not stop") + } + } rows, republished := r.window.sink.counts() stillOpen := queryInt(t, db.reader, fmt.Sprintf(`SELECT coalesce(sum(n), 0)::BIGINT FROM %s`, windowTable)) - // What the run cannot account for: produced, not published, not still - // held. Under late_rows drop that is exactly what the drop policy - // discarded, and it is measured from the run's own totals rather than - // taken from the manager's counter, so a row lost some other way shows - // up here too rather than being reported as zero. - lateDropped := int64(r.produced) - rows - stillOpen - if lateDropped < 0 { - lateDropped = 0 - } return WindowResult{ Produced: r.produced, Published: rows, StillOpen: stillOpen, Republished: republished, - LateDropped: lateDropped, + LateRefused: counterSum(t, reader, "window_late_rows", attribute.String("outcome", "refused")), + Unplaceable: counterSum(t, reader, "messages_unplaceable_total"), + Recomputes: counterSum(t, reader, "window_recomputes"), + LastValue: r.window.sink.lastValues(), } } +// counterSum reads a counter's total across every data point carrying the +// given attributes. +func counterSum(t *testing.T, reader *sdkmetric.ManualReader, name string, with ...attribute.KeyValue) int64 { + t.Helper() + var rm metricdata.ResourceMetrics + if err := reader.Collect(context.Background(), &rm); err != nil { + t.Fatal(err) + } + var total int64 + for _, sm := range rm.ScopeMetrics { + for _, m := range sm.Metrics { + if m.Name != name { + continue + } + sum, ok := m.Data.(metricdata.Sum[int64]) + if !ok { + continue + } + for _, dp := range sum.DataPoints { + matches := true + for _, kv := range with { + if v, ok := dp.Attributes.Value(kv.Key); !ok || v.AsString() != kv.Value.AsString() { + matches = false + } + } + if matches { + total += dp.Value + } + } + } + } + return total +} + // windowRun is the window half of a run: the database, the manager, and what // the window's sink received. type windowRun struct { @@ -341,27 +490,52 @@ type windowRun struct { handler *windowHandler manager *managers.Watermark sink *windowSink + signal *core.WindowSignal + // tracker is the engine's, for the current process: what deliver asks + // whether a record will be refused. + tracker *core.Watermarks + metrics *core.Metrics + reader *sdkmetric.ManualReader +} + +func (Pass) apply(r *run) { + if r.window == nil { + r.t.Fatal("Pass needs a windowed run") + } + if err := r.window.manager.Pass(context.Background()); err != nil { + r.t.Fatalf("the manager's pass failed: %v", err) + } } -func (Poll) apply(r *run) { +func (StartPass) apply(r *run) { if r.window == nil { - r.t.Fatal("Poll needs a windowed run") + r.t.Fatal("StartPass needs a windowed run") + } + if err := r.window.manager.StartPass(context.Background()); err != nil { + r.t.Fatalf("the manager's start pass failed: %v", err) } - if err := r.window.manager.Poll(context.Background()); err != nil { - r.t.Fatalf("the manager's poll failed: %v", err) +} + +func (s AwaitPublished) apply(r *run) { + if r.window == nil { + r.t.Fatal("AwaitPublished needs a windowed run") } + r.await(fmt.Sprintf("the window's sink to hold %d rows", s.Rows), + func() bool { rows, _ := r.window.sink.counts(); return rows >= s.Rows }) } // newWatermarks builds the engine's tracker for the window, restored from // what the table and the row hold, the way run does at start. On the first // start both are empty; after a Restart they are what the previous process -// left. +// left. The recompute set is in memory, so a Restart loses it: the start +// pass is what recovers those buckets. func (w *windowRun) newWatermarks(t *testing.T, clock func() time.Time) *core.Watermarks { t.Helper() ctx := context.Background() - wm := core.NewWatermarks([]core.WindowSpec{{ - Name: w.decl.Table, Size: w.decl.Size, Grace: w.decl.Grace, IdleClose: w.decl.IdleClose, - }}, clock) + w.signal.TakeRecompute() + wm := core.NewWatermarksWithSignals([]core.WindowSpec{{ + Name: w.decl.Table, Size: w.decl.Size, Grace: w.decl.Grace, IdleClose: w.decl.IdleClose, Lateness: w.decl.Lateness, + }}, clock, []*core.WindowSignal{w.signal}) newest, _, err := core.NewestBucketStart(ctx, w.db.reader, w.decl.Table, w.decl.TimeColumn) if err != nil { t.Fatal(err) @@ -371,6 +545,7 @@ func (w *windowRun) newWatermarks(t *testing.T, clock func() time.Time) *core.Wa t.Fatal(err) } wm.Restore(w.decl.Table, newest, asserted) + w.tracker = wm return wm } diff --git a/internal/simulate/window_test.go b/internal/simulate/window_test.go index bec78273..110ae32a 100644 --- a/internal/simulate/window_test.go +++ b/internal/simulate/window_test.go @@ -16,7 +16,6 @@ func windowDecl() managers.Declaration { Size: time.Minute, Grace: time.Minute, IdleClose: 10 * time.Second, - Late: managers.LateDrop, } } @@ -37,7 +36,7 @@ func windowDecl() managers.Declaration { // Produce 6 12:01:02 12:00: 4, 12:01: 6 12:00:02 - // Elapse 1m // Produce 5 12:02:03 12:00: 4, 12:01: 6, 12:02: 5 12:01:03 - -// Poll 12:01: 6, 12:02: 5 12:01:03 12:01:03 +// Pass 12:01: 6, 12:02: 5 12:01:03 12:01:03 // ^ 12:00 ends at 12:01, at or before the assertion: it // publishes. 12:01 ends at 12:02 and stays open. // Published + still open = 15, the whole run. @@ -49,7 +48,7 @@ func TestSimulate_AWindowedRunAccountsForEveryRow(t *testing.T) { Produce{Partition: 0, Rows: 6}, Elapse{By: time.Minute}, Produce{Partition: 0, Rows: 5}, - Poll{}, + Pass{}, }) assert.Equal(t, 15, r.Produced) @@ -66,24 +65,24 @@ func TestSimulate_AWindowedRunAccountsForEveryRow(t *testing.T) { // step engine clock window table asserted why // -------------------------------------------------------------------------- // Produce 7 12:00:01 12:00: 7 11:59:01 grace: 12:00:01 - 1m -// Poll 12:00:02 12:00: 7 11:59:01 held: 12:00 ends 12:01 +// Pass 12:00:02 12:00: 7 11:59:01 held: 12:00 ends 12:01 // Elapse 30s 12:00:33 12:00: 7 11:59:01 time passed; no commit, // so the row did not move // IdleTick 12:00:34 12:00: 7 12:01:01 the partition has been // silent 33s > 10s: idle. // Nothing in the minimum: // newest 12:00:01 + 1m -// Poll 12:00:35 empty 12:01:01 12:00 ends 12:01: closed +// Pass 12:00:35 empty 12:01:01 12:00 ends 12:01: closed func TestSimulate_AnIdleCloseNeedsTheLoopToAssertIt(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ Produce{Partition: 0, Rows: 7}, // A poll before any tick closes nothing: the assertion is a minute // behind the only bucket, and nothing has said the stream stopped. - Poll{}, + Pass{}, Elapse{By: 30 * time.Second}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 7, r.Produced) @@ -102,9 +101,9 @@ func TestSimulate_AnIdleCloseNeedsTheLoopToAssertIt(t *testing.T) { // Lose p0 p0 lost 12:00: 9 11:59:01 holds at 12:00:01 - 1m // Elapse 5m // IdleTick p0 lost 12:00: 9 11:59:01 lost is not idle -// Poll p0 lost 12:00: 9 11:59:01 held +// Pass p0 lost 12:00: 9 11:59:01 held // IdleTick p0 lost 12:00: 9 11:59:01 still -// Poll p0 lost 12:00: 9 11:59:01 held +// Pass p0 lost 12:00: 9 11:59:01 held func TestSimulate_ALostPartitionHoldsTheWindow(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ @@ -113,9 +112,9 @@ func TestSimulate_ALostPartitionHoldsTheWindow(t *testing.T) { // Far past the idle bound, with the loop committing all the while. Elapse{By: 5 * time.Minute}, IdleTick{}, - Poll{}, + Pass{}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 9, r.Produced) @@ -134,11 +133,11 @@ func TestSimulate_ALostPartitionHoldsTheWindow(t *testing.T) { // Lose p0 lost - lost: never 11:59:01 // Elapse 5m // IdleTick lost - no 11:59:01 -// Poll lost held +// Pass lost held // Assign p0 holds p0 12:05:05 0s: no 11:59:01 // Elapse 3s // IdleTick holds p0 12:05:05 5s < 10s: no 11:59:01 -// Poll held +// Pass held func TestSimulate_TheAssignmentBoundsTheIdleness(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ @@ -146,13 +145,13 @@ func TestSimulate_TheAssignmentBoundsTheIdleness(t *testing.T) { Lose{Partition: 0}, Elapse{By: 5 * time.Minute}, IdleTick{}, - Poll{}, + Pass{}, Assign{Partition: 0}, // Three seconds, plus the second each step takes: under the ten the // declaration closes on. Elapse{By: 3 * time.Second}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 9, r.Produced) @@ -167,11 +166,11 @@ func TestSimulate_TheAssignmentBoundsTheIdleness(t *testing.T) { // ---------------------------------------------------------------------- // Produce 9 holds p0 12:00:01 - 11:59:01 // Lose p0 lost lost: never 11:59:01 -// Elapse 5m, IdleTick, Poll held +// Elapse 5m, IdleTick, Pass held // Assign p0 holds p0 12:05:05 0s 11:59:01 // Elapse 1m // IdleTick holds p0 12:05:05 1m > 10s: yes 12:01:01 = 12:00:01 + 1m -// Poll 12:00 closes +// Pass 12:00 closes func TestSimulate_TheIdleCloseResumesWhenThePartitionIsBack(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ @@ -179,11 +178,11 @@ func TestSimulate_TheIdleCloseResumesWhenThePartitionIsBack(t *testing.T) { Lose{Partition: 0}, Elapse{By: 5 * time.Minute}, IdleTick{}, - Poll{}, + Pass{}, Assign{Partition: 0}, Elapse{By: time.Minute}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 9, r.Produced) @@ -208,7 +207,7 @@ func TestSimulate_TheIdleCloseResumesWhenThePartitionIsBack(t *testing.T) { // Produce 3 p1 12:00:30 12:00: 6 11:59:30 min(p0, p1) - 1m // Revoke p1 12:00: 6 11:59:30 p1 gone // Produce 4 p0 12:05:00 12:00: 6, 12:05: 4 12:04 p0 alone -// Poll 12:05: 4 closed 12:04 12:00 publishes 6 +// Pass 12:05: 4 closed 12:04 12:00 publishes 6 // Produce 5 p1 12:00:45 (not ours) another worker's rows func TestSimulate_ARevokedPartitionLeavesTheMinimum(t *testing.T) { coverage.Covers(t, "manager.window") @@ -217,7 +216,7 @@ func TestSimulate_ARevokedPartitionLeavesTheMinimum(t *testing.T) { Produce{Partition: 1, Rows: 3, At: base.Add(30 * time.Second)}, Revoke{Partition: 1}, Produce{Partition: 0, Rows: 4, At: base.Add(5 * time.Minute)}, - Poll{}, + Pass{}, // Produced against the other worker, not this one: nothing arrives // here, and it is not a loss. Produce{Partition: 1, Rows: 5, At: base.Add(45 * time.Second)}, @@ -226,7 +225,7 @@ func TestSimulate_ARevokedPartitionLeavesTheMinimum(t *testing.T) { assert.Equal(t, 10, r.Produced) assert.Equal(t, int64(6), r.Published) assert.Equal(t, int64(4), r.StillOpen) - assert.Equal(t, int64(0), r.LateDropped) + assert.Equal(t, int64(0), r.LateRefused) } // The rows a revoked partition left behind still close, even after every @@ -251,7 +250,7 @@ func TestSimulate_ARevokedPartitionLeavesTheMinimum(t *testing.T) { // minimum: newest // 12:02 + 1m, and the // newest is p1's -// Poll empty 12:03 both buckets close +// Pass empty 12:03 both buckets close func TestSimulate_ARevokedPartitionsRowsStillCloseOnIdleness(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0, 1}, windowDecl(), []Step{ @@ -260,13 +259,13 @@ func TestSimulate_ARevokedPartitionsRowsStillCloseOnIdleness(t *testing.T) { Revoke{Partition: 1}, Elapse{By: 30 * time.Second}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 7, r.Produced) assert.Equal(t, int64(7), r.Published) assert.Equal(t, int64(0), r.StillOpen) - assert.Equal(t, int64(0), r.LateDropped) + assert.Equal(t, int64(0), r.LateRefused) } // And a worker that loses its whole assignment closes what it holds. Nothing @@ -281,7 +280,7 @@ func TestSimulate_AWorkerHoldingNothingClosesWhatItHas(t *testing.T) { Revoke{Partition: 1}, Elapse{By: 30 * time.Second}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 7, r.Produced) @@ -301,7 +300,7 @@ func TestSimulate_AWorkerHoldingNothingClosesWhatItHas(t *testing.T) { // Elapse 30s // IdleTick 12:00: 7 12:01 idle with nothing seen: the // newest bucket's end -// Poll empty 12:01 12:00 closes +// Pass empty 12:01 12:00 closes func TestSimulate_ARestartStillClosesByIdleness(t *testing.T) { coverage.Covers(t, "manager.window") r := RunWindowed(t, []int32{0}, windowDecl(), []Step{ @@ -309,7 +308,7 @@ func TestSimulate_ARestartStillClosesByIdleness(t *testing.T) { Restart{}, Elapse{By: 30 * time.Second}, IdleTick{}, - Poll{}, + Pass{}, }) assert.Equal(t, 7, r.Produced) diff --git a/internal/turbostats/collect.go b/internal/turbostats/collect.go index 3cf65533..f5118e69 100644 --- a/internal/turbostats/collect.go +++ b/internal/turbostats/collect.go @@ -119,12 +119,12 @@ func pipelineSection(ctx context.Context, flat map[string]int64, floats map[stri } if dim.windowSeen { p.WindowClosedCount = &dim.windowClosed - // A policy with no field of its own leaves both out. Reporting the + // An outcome with no field of its own leaves both out. Reporting the // two known ones would state a count that omits rows, on the field // an operator reads for data loss. - if !dim.lateUnknownPolicy { - p.LateRowsDropped = &dim.lateDropped - p.LateRowsReemitted = &dim.lateReemitted + if !dim.lateUnknownOutcome { + p.LateRowsDropped = &dim.lateRefused + p.LateRowsRecomputed = &dim.lateRecomputed } } if dim.closeLagSeen { @@ -259,12 +259,12 @@ func later(a, b *time.Time) *time.Time { // - A shard attribute -- topic, partition, window, sink -- splits one // measurement across parts of one system. Collapsing it is arithmetic // that stays true. -// - An outcome attribute -- result, policy -- splits points that measure +// - An outcome attribute -- result, outcome -- splits points that measure // different things. Collapsing it reports a number true of nothing. // sink_flush_count carries result=ok and result=error, and their sum is // a count of flushes that never happened; window_late_rows carries -// policy=drop and policy=reemit, and one of those lost data while the -// other did not. +// outcome=refused and outcome=recomputed, and one of those lost data +// while the other did not. // // So flat keeps ignoring every attributed point, and dim collapses shards // only, splitting each outcome into a field of its own. @@ -406,11 +406,11 @@ type dimensional struct { lagPoints int sinkRetries int64 - windowSeen bool - lateDropped int64 - lateReemitted int64 - lateUnknownPolicy bool - windowClosed int64 + windowSeen bool + lateRefused int64 + lateRecomputed int64 + lateUnknownOutcome bool + windowClosed int64 // The most behind window's close lag, with a seen flag because zero lag // is a reading. And the newest bucket across windows, a Unix second that @@ -477,18 +477,18 @@ func (d *dimensional) add(name string, attrs attribute.Set, v int64) { d.windowSeen = true d.windowClosed += v case "window_late_rows": - // Collapses window, a shard. Splits policy, an outcome: dropped rows - // are gone and reemitted rows are not. + // Collapses window, a shard. Splits outcome: a refused row is gone + // and a recomputed row is in its bucket's republished value. d.windowSeen = true - switch policy, _ := attrs.Value(attribute.Key("policy")); policy.AsString() { - case "drop": - d.lateDropped += v - case "reemit": - d.lateReemitted += v + switch outcome, _ := attrs.Value(attribute.Key("outcome")); outcome.AsString() { + case "refused": + d.lateRefused += v + case "recomputed": + d.lateRecomputed += v default: // An outcome this contract has no field for. It cannot join // either count without making that count false. - d.lateUnknownPolicy = true + d.lateUnknownOutcome = true } case "window_close_lag_seconds": // Already a duration in event time; the most behind window wins. diff --git a/internal/turbostats/collect_leak_test.go b/internal/turbostats/collect_leak_test.go index fa54b93d..a4cc2988 100644 --- a/internal/turbostats/collect_leak_test.go +++ b/internal/turbostats/collect_leak_test.go @@ -57,10 +57,10 @@ func TestCollect_DoesNotGrowOverManyCollects(t *testing.T) { for _, name := range []string{"hourly", "daily"} { window := managers.NewWindowMetrics(mp, name) w := metric.WithAttributes(attribute.String("window", name)) - window.Late.Add(ctx, 1, metric.WithAttributes( - attribute.String("window", name), attribute.String("policy", string(managers.LateDrop)))) - window.Late.Add(ctx, 1, metric.WithAttributes( - attribute.String("window", name), attribute.String("policy", string(managers.LateReemit)))) + m.WindowLateRows.Add(ctx, 1, metric.WithAttributes( + attribute.String("window", name), attribute.String("outcome", "refused"))) + m.WindowLateRows.Add(ctx, 1, metric.WithAttributes( + attribute.String("window", name), attribute.String("outcome", "recomputed"))) window.Closed.Add(ctx, 1, w) window.NewestStart.Record(ctx, 1757570000, w) window.CloseLag.Record(ctx, 0, w) diff --git a/internal/turbostats/collect_test.go b/internal/turbostats/collect_test.go index 2881c33d..8baaa718 100644 --- a/internal/turbostats/collect_test.go +++ b/internal/turbostats/collect_test.go @@ -244,11 +244,12 @@ func TestCollect_ARealisticRunBundleStaysUnderItsCeiling(t *testing.T) { } window := managers.NewWindowMetrics(mp, "posts_by_lang") w := metric.WithAttributes(attribute.String("window", "posts_by_lang")) - window.Late.Add(ctx, 184203, metric.WithAttributes( - attribute.String("window", "posts_by_lang"), attribute.String("policy", "drop"))) - window.Late.Add(ctx, 184203, metric.WithAttributes( - attribute.String("window", "posts_by_lang"), attribute.String("policy", "reemit"))) + m.WindowLateRows.Add(ctx, 184203, metric.WithAttributes( + attribute.String("window", "posts_by_lang"), attribute.String("outcome", "refused"))) + m.WindowLateRows.Add(ctx, 184203, metric.WithAttributes( + attribute.String("window", "posts_by_lang"), attribute.String("outcome", "recomputed"))) window.Closed.Add(ctx, 1842033, w) + window.Recomputed.Add(ctx, 184203, w) window.NewestStart.Record(ctx, 1757570000, w) window.CloseLag.Record(ctx, 0, w) sinks.RetryCounter(mp, "postgres")(1, errors.New("refused")) diff --git a/internal/turbostats/dimensional_test.go b/internal/turbostats/dimensional_test.go index eaf00758..4ee4acde 100644 --- a/internal/turbostats/dimensional_test.go +++ b/internal/turbostats/dimensional_test.go @@ -11,6 +11,7 @@ import ( "time" "github.com/turbolytics/sql-flow/internal/config" + "github.com/turbolytics/sql-flow/internal/core" "github.com/turbolytics/sql-flow/internal/coverage" "github.com/turbolytics/sql-flow/internal/managers" "github.com/turbolytics/sql-flow/internal/sinks" @@ -146,30 +147,30 @@ func TestCollect_AKafkaPipelineHoldingNoPartitionsReportsZero(t *testing.T) { // Late rows are split by what happened to them, and never summed. // -// policy is an outcome: a dropped row is gone and a reemitted row is not, so -// their sum is true of neither. window is a shard, and collapses. This drives -// the real instruments, so a rename in managers fails here. +// outcome is an outcome: a refused row is gone and a recomputed row is in +// its bucket's republished value, so their sum is true of neither. window is +// a shard, and collapses. This drives the real instruments -- the engine's +// counter, since the engine decides lateness -- so a rename fails here. func TestCollect_LateRowsSplitByOutcomeAndCollapseByWindow(t *testing.T) { coverage.Covers(t, "observability.turbostats") ctx := context.Background() reader, mp := windowed() + engine, err := core.NewMetrics(mp) + assert.NoError(t, err) hourly := managers.NewWindowMetrics(mp, "hourly") daily := managers.NewWindowMetrics(mp, "daily") - drop := func(w string) metric.AddOption { - return metric.WithAttributes(attribute.String("window", w), attribute.String("policy", string(managers.LateDrop))) - } - reemit := func(w string) metric.AddOption { - return metric.WithAttributes(attribute.String("window", w), attribute.String("policy", string(managers.LateReemit))) + late := func(w, outcome string) metric.AddOption { + return metric.WithAttributes(attribute.String("window", w), attribute.String("outcome", outcome)) } - hourly.Late.Add(ctx, 3, drop("hourly")) - daily.Late.Add(ctx, 4, drop("daily")) - hourly.Late.Add(ctx, 10, reemit("hourly")) + engine.WindowLateRows.Add(ctx, 3, late("hourly", "refused")) + engine.WindowLateRows.Add(ctx, 4, late("daily", "refused")) + engine.WindowLateRows.Add(ctx, 10, late("hourly", "recomputed")) hourly.Closed.Add(ctx, 2, win("hourly")) daily.Closed.Add(ctx, 1, win("daily")) p := pipelineOf(t, reader) assert.Equal(t, int64(7), *p.LateRowsDropped) - assert.Equal(t, int64(10), *p.LateRowsReemitted) + assert.Equal(t, int64(10), *p.LateRowsRecomputed) assert.Equal(t, int64(3), *p.WindowClosedCount) } @@ -185,33 +186,35 @@ func TestCollect_WindowCountersArePresentBeforeTheFirstClose(t *testing.T) { p := pipelineOf(t, reader) assert.That(t, p.LateRowsDropped != nil) - assert.That(t, p.LateRowsReemitted != nil) + assert.That(t, p.LateRowsRecomputed != nil) assert.That(t, p.WindowClosedCount != nil) assert.Equal(t, int64(0), *p.LateRowsDropped) reader, _, _ = provider(t) plain := pipelineOf(t, reader) assert.That(t, plain.LateRowsDropped == nil) - assert.That(t, plain.LateRowsReemitted == nil) + assert.That(t, plain.LateRowsRecomputed == nil) assert.That(t, plain.WindowClosedCount == nil) assert.That(t, plain.WindowLagSeconds == nil) assert.That(t, plain.WindowNewestBucketAt == nil) } -// A policy the contract has no field for leaves both late counts out. +// An outcome the contract has no field for leaves both late counts out. // // Adding it to neither would report the known counts as complete while some // rows went uncounted, on the field an operator reads for data loss. -func TestCollect_AnUnknownLatePolicyReportsNeitherCount(t *testing.T) { +func TestCollect_AnUnknownLateOutcomeReportsNeitherCount(t *testing.T) { coverage.Covers(t, "observability.turbostats") reader, mp := windowed() - w := managers.NewWindowMetrics(mp, "hourly") - w.Late.Add(context.Background(), 5, metric.WithAttributes( - attribute.String("window", "hourly"), attribute.String("policy", "quarantine"))) + managers.NewWindowMetrics(mp, "hourly") + engine, err := core.NewMetrics(mp) + assert.NoError(t, err) + engine.WindowLateRows.Add(context.Background(), 5, metric.WithAttributes( + attribute.String("window", "hourly"), attribute.String("outcome", "quarantine"))) p := pipelineOf(t, reader) assert.That(t, p.LateRowsDropped == nil) - assert.That(t, p.LateRowsReemitted == nil) + assert.That(t, p.LateRowsRecomputed == nil) // The window is still there, and says so. assert.That(t, p.WindowClosedCount != nil) } @@ -455,7 +458,7 @@ func TestCollect_AFullBundleStaysUnderTheCeiling(t *testing.T) { StateCommitCount: big, StateDBSizeBytes: &big, LastMessageAt: &at, SinkRetryCount: &big, LagMaxMessages: &big, LagTotalMessages: &big, LagPartitions: &n, LagObservedAt: &at, LateRowsDropped: &big, - LateRowsReemitted: &big, WindowClosedCount: &big, + LateRowsRecomputed: &big, WindowClosedCount: &big, WindowLagSeconds: &big, WindowNewestBucketAt: &at, SourceErrorCount: &big, HandlerErrorCount: &big, SinkErrorCount: &big, StateErrorCount: &big, DLQRows: &big, LastErrorCode: &code, diff --git a/internal/validate/drain_test.go b/internal/validate/drain_test.go index 540010e7..17d1b73e 100644 --- a/internal/validate/drain_test.go +++ b/internal/validate/drain_test.go @@ -128,7 +128,6 @@ tables: window: time_column: bucket size_seconds: 60 - late_rows: drop sink: type: iceberg iceberg: diff --git a/internal/validate/schemas/config.json b/internal/validate/schemas/config.json index 8fd9186f..31da6a3f 100644 --- a/internal/validate/schemas/config.json +++ b/internal/validate/schemas/config.json @@ -987,18 +987,10 @@ "minimum": 0, "description": "How long a source partition may be silent before it stops holding the\nwindow open. Absent means never: a stream that stops leaves its last\nbucket open, and only event time moving on closes anything. When every\npartition has been silent this long the stream is done with what it\nhas, and every open bucket closes. Measured by the engine from its own\nmonotonic clock, at each commit, so the close can trail this by up to\none flush_interval_seconds; validate warns when that interval is the\nlonger of the two." }, - "late_rows": { - "type": "string", - "enum": [ - "drop", - "reemit" - ], - "description": "What happens to a row for a bucket that already closed. drop discards\nit and counts it. reemit publishes emit_sql over the late rows alone,\nfor a sink that adds them to the bucket it holds; a sink that replaces\nthe bucket's value loses the rows published before.\nRequired: the two are different promises to the sink." - }, - "poll_interval_seconds": { + "allowed_lateness_seconds": { "type": "integer", - "minimum": 1, - "description": "How often the engine looks for closed buckets. Absent means 10." + "minimum": 0, + "description": "How long after a bucket closes its rows are kept and late rows for it\nare still accepted. A late row within this republishes the bucket as a\nwhole value, so the sink must replace by key; validate refuses a sink\nthat appends. Absent or 0: a row for a closed bucket is refused before\nthe handler and counted in window_late_rows_total; the window table\nnever holds it. Flink's allowedLateness. Measured in event time,\nagainst the watermark, like everything a window decides." }, "emit_sql": { "type": "string", @@ -1254,7 +1246,6 @@ "required": [ "time_column", "size_seconds", - "late_rows", "sink" ], "description": "A tumbling window over this table, closed by the engine." diff --git a/internal/validate/sinks.go b/internal/validate/sinks.go index 44559c4e..ae731459 100644 --- a/internal/validate/sinks.go +++ b/internal/validate/sinks.go @@ -60,11 +60,6 @@ func checkSinks(rendered []byte, rep *Report) { } node := windowNode(&root, i) checkSink(fmt.Sprintf("tables.sql[%d] window sink", i), table.Window.Sink, mappingValue(node, "sink"), attached, fail, warn) - - if table.Window.ReemitOverwrites() { - fail(fmt.Sprintf("tables.sql[%d] window: %s", i, config.ReemitOverwritesMessage), - position(mappingKey(node, "late_rows"))) - } } } diff --git a/internal/validate/sinks_test.go b/internal/validate/sinks_test.go index a2d3225e..b1b052ca 100644 --- a/internal/validate/sinks_test.go +++ b/internal/validate/sinks_test.go @@ -10,8 +10,8 @@ import ( ) // sinksConfig has four holes: the attach command's TYPE, the window's -// late_rows, the window's sink block at ten spaces, and the pipeline's sink -// block at four. +// allowed_lateness_seconds, the window's sink block at ten spaces, and the +// pipeline's sink block at four. const sinksConfig = `commands: - name: attach sql: | @@ -23,7 +23,7 @@ tables: window: time_column: bucket size_seconds: 60 - late_rows: %s + allowed_lateness_seconds: %s sink: %s pipeline: @@ -71,10 +71,10 @@ sqlcommand: noopBlock = `type: noop` ) -func validateSinks(t *testing.T, attachType, lateRows, windowSink, pipelineSink string) Report { +func validateSinks(t *testing.T, attachType, lateness, windowSink, pipelineSink string) Report { t.Helper() cfg := sinksConfig - for _, v := range []string{attachType, lateRows, indent(windowSink, 10), indent(pipelineSink, 4)} { + for _, v := range []string{attachType, lateness, indent(windowSink, 10), indent(pipelineSink, 4)} { cfg = strings.Replace(cfg, "%s", v, 1) } rep, err := Validate(context.Background(), Request{Path: "p.yml", Config: cfg}) @@ -93,33 +93,35 @@ func sinkDiagnostics(rep Report) []Diagnostic { return out } -// upsert with reemit is refused: the sink replaces a bucket's row with what -// it is handed, and a reemit hands it emit_sql over the late rows alone. -func TestValidateSchema_PostgresUpsertRefusesReemit(t *testing.T) { +// Lateness with an upsert is the pairing that works: a republished bucket +// replaces its earlier row. Nothing is said about the sink. +func TestValidateSchema_PostgresUpsertTakesLateness(t *testing.T) { coverage.Covers(t, "validate.schema") - rep := validateSinks(t, "POSTGRES", "reemit", upsertBlock, noopBlock) - assert.That(t, !rep.OK) - diags := sinkDiagnostics(rep) - assert.Equal(t, 1, len(diags)) - assert.Equal(t, SeverityError, diags[0].Severity) - assert.That(t, strings.Contains(diags[0].Message, "late_rows is reemit and the postgres sink upserts")) - assert.Equal(t, StatusFail, checkStatus(t, rep, "sinks.postgres")) - - rep = validateSinks(t, "POSTGRES", "drop", upsertBlock, noopBlock) + rep := validateSinks(t, "POSTGRES", "300", upsertBlock, noopBlock) assert.That(t, rep.OK) assert.Equal(t, 0, len(sinkDiagnostics(rep))) assert.Equal(t, StatusPass, checkStatus(t, rep, "sinks.postgres")) + + rep = validateSinks(t, "POSTGRES", "0", upsertBlock, noopBlock) + assert.That(t, rep.OK) + assert.Equal(t, 0, len(sinkDiagnostics(rep))) } -// append with reemit warns, as kafka and iceberg do, and only once. -func TestValidateSchema_PostgresAppendWarnsOnReemit(t *testing.T) { +// append with lateness is refused, as kafka and iceberg are, and only once: +// a republished bucket lands beside its first publication. +func TestValidateSchema_PostgresAppendRefusesLateness(t *testing.T) { coverage.Covers(t, "validate.schema") - rep := validateSinks(t, "POSTGRES", "reemit", appendBlock, noopBlock) - assert.That(t, rep.OK) + rep := validateSinks(t, "POSTGRES", "300", appendBlock, noopBlock) + assert.That(t, !rep.OK) diags := sinkDiagnostics(rep) assert.Equal(t, 1, len(diags)) - assert.Equal(t, SeverityWarning, diags[0].Severity) + assert.Equal(t, SeverityError, diags[0].Severity) assert.That(t, strings.Contains(diags[0].Message, "the postgres sink appends")) + + // Without lateness an appending sink sees each bucket once, and passes. + rep = validateSinks(t, "POSTGRES", "0", appendBlock, noopBlock) + assert.That(t, rep.OK) + assert.Equal(t, 0, len(sinkDiagnostics(rep))) } // A sqlcommand upsert into an attached Postgres reads the whole target table @@ -131,7 +133,7 @@ func TestValidateSchema_SqlcommandUpsertIntoPostgresWarns(t *testing.T) { "window sink": {conflictBlock, noopBlock}, "pipeline sink": {consoleBlock, conflictBlock}, } { - rep := validateSinks(t, "POSTGRES", "drop", c.window, c.pipeline) + rep := validateSinks(t, "POSTGRES", "0", c.window, c.pipeline) assert.That(t, rep.OK) diags := sinkDiagnostics(rep) if len(diags) != 1 { @@ -144,7 +146,7 @@ func TestValidateSchema_SqlcommandUpsertIntoPostgresWarns(t *testing.T) { // Without an attached Postgres, ON CONFLICT is DuckDB's own and the // extension's cost does not apply. - rep := validateSinks(t, "DUCKDB", "drop", conflictBlock, noopBlock) + rep := validateSinks(t, "DUCKDB", "0", conflictBlock, noopBlock) assert.Equal(t, 0, len(sinkDiagnostics(rep))) } @@ -158,7 +160,7 @@ func TestValidateSchema_PostgresBlockShape(t *testing.T) { "needs key": "type: postgres\npostgres:\n dsn: x\n table: t\n mode: upsert", "takes no key": "type: postgres\npostgres:\n dsn: x\n table: t\n mode: append\n key: [a]", } { - rep := validateSinks(t, "POSTGRES", "drop", block, noopBlock) + rep := validateSinks(t, "POSTGRES", "0", block, noopBlock) assert.That(t, !rep.OK) found := false for _, d := range sinkDiagnostics(rep) { @@ -181,13 +183,13 @@ func TestValidateSchema_SqlcommandUpsertWarnsOnlyForThePostgresTarget(t *testing local := `type: sqlcommand sqlcommand: sql: INSERT INTO agg SELECT * FROM sqlflow_sink_batch ON CONFLICT (bucket) DO UPDATE SET count = EXCLUDED.count` - rep := validateSinks(t, "POSTGRES", "drop", local, noopBlock) + rep := validateSinks(t, "POSTGRES", "0", local, noopBlock) assert.Equal(t, 0, len(sinkDiagnostics(rep))) replace := `type: sqlcommand sqlcommand: sql: INSERT OR REPLACE INTO "pg".agg SELECT * FROM sqlflow_sink_batch` - rep = validateSinks(t, "POSTGRES", "drop", replace, noopBlock) + rep = validateSinks(t, "POSTGRES", "0", replace, noopBlock) diags := sinkDiagnostics(rep) assert.Equal(t, 1, len(diags)) assert.That(t, strings.Contains(diags[0].Message, "whole target table")) @@ -201,7 +203,7 @@ func TestValidateSchema_SqlcommandUpsertWarnsWhenTheAliasIsImplicit(t *testing.T local := `type: sqlcommand sqlcommand: sql: INSERT INTO agg SELECT * FROM sqlflow_sink_batch ON CONFLICT (bucket) DO UPDATE SET count = EXCLUDED.count` - for _, v := range []string{"POSTGRES", "drop", indent(local, 10), indent(noopBlock, 4)} { + for _, v := range []string{"POSTGRES", "0", indent(local, 10), indent(noopBlock, 4)} { cfg = strings.Replace(cfg, "%s", v, 1) } rep, err := Validate(context.Background(), Request{Path: "p.yml", Config: cfg}) diff --git a/internal/validate/window.go b/internal/validate/window.go index 606330f2..3079b3ce 100644 --- a/internal/validate/window.go +++ b/internal/validate/window.go @@ -3,6 +3,7 @@ package validate import ( "fmt" "regexp" + "strconv" "strings" "github.com/turbolytics/sql-flow/internal/config" @@ -12,14 +13,22 @@ import ( // checkWindows checks every window declaration without running anything. // -// Three faults it catches before the pipeline starts: +// The faults it catches before the pipeline starts: // -// - A `manager:` block. The engine closes windows now, and the two -// predicates the block carried are gone. The message names what -// replaces them, because the schema check alone says "unknown key". +// - A `manager:` block, `late_rows`, or `poll_interval_seconds`: keys +// whose meaning is gone. Each message names what replaces it, because +// the schema check alone says "unknown key". // - A time column the table's CREATE does not declare as TIMESTAMPTZ. A // TIMESTAMP bucket is read in the session's zone, and the watermark // compares instants, so it closes windows early or never. +// - A time column the handler's SQL does not compute as +// time_bucket(INTERVAL '', event_time). The engine decides a +// record's lateness from that bucket, computed the same way from the +// same clock, and a row bucketed any other way is not in the bucket the +// engine decided against. +// - allowed_lateness_seconds above zero with a sink that appends. A late +// row republishes its bucket as a whole value, which an appending sink +// then holds twice. A sink validate cannot classify is warned instead. // - An emit_sql that does not read `closed`, which is the only relation a // close supplies. // - A window table with an index and no state path. DuckDB never reclaims @@ -28,8 +37,8 @@ import ( // fifth footgun, as a warning: the pipeline runs, and the operator is // told what it costs. // -// The column check is textual: validate links no DuckDB, so it reads the -// CREATE statement rather than parsing it. +// The column checks are textual: validate links no DuckDB, so it reads the +// CREATE statement and the handler's SQL rather than parsing them. func checkWindows(rendered []byte, rep *Report) { var root yaml.Node if err := yaml.Unmarshal(rendered, &root); err != nil { @@ -56,10 +65,23 @@ func checkWindows(rendered []byte, rep *Report) { if key := mappingKey(table, "manager"); key != nil { fail(fmt.Sprintf("tables.sql[%d]: manager is gone. Declare the window instead: "+ "window with time_column, size_seconds, grace_seconds, idle_close_seconds, "+ - "late_rows, emit_sql and sink. The engine closes it; collect_closed_windows_sql "+ - "and delete_closed_windows_sql are not written by the user", i), + "allowed_lateness_seconds, emit_sql and sink. The engine closes it; "+ + "collect_closed_windows_sql and delete_closed_windows_sql are not written by the user", i), position(key)) } + win := mappingValue(table, "window") + if key := mappingKey(win, "late_rows"); key != nil { + fail(fmt.Sprintf("tables.sql[%d] window: late_rows is gone. Set allowed_lateness_seconds "+ + "instead: 0 refuses a row for a closed bucket before the handler, which is what drop did "+ + "after the fact; a positive value keeps a closed bucket that long and republishes it whole "+ + "when a late row arrives, which is what reemit tried to do without the delta. See "+ + "docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md", i), position(key)) + } + if key := mappingKey(win, "poll_interval_seconds"); key != nil { + fail(fmt.Sprintf("tables.sql[%d] window: poll_interval_seconds is gone. The engine closes a "+ + "window the moment its watermark passes a bucket's end; nothing polls. Remove the key. See "+ + "docs/superpowers/specs/2026-09-26-watermark-driven-close-design.md", i), position(key)) + } } if conf.Tables != nil { @@ -79,23 +101,21 @@ func checkWindows(rendered []byte, rep *Report) { // block, complying makes the two clocks one by construction, and // this becomes an error. if hasWindow(conf) { + // The per-window rule below requires time_column to be cut from + // event_time. A structured handler can only do that if its batch + // table declares the column, so that is refused here first, with + // the message that says what to add. This was a warning until + // every source could be told where its time is (#389) and the + // engine decided lateness from that bucket (#396); it is an error + // now, because a bucket on another clock is a bucket the engine + // decides against wrongly. h := conf.Pipeline.Handler - warn := func(msg string) { - rep.Add(diagnostic(errs.CodeConfigInvalid, SeverityWarning, msg, position(handlerNode(&root)))) - } - if !mentionsEventTime(h.SQL) { - warn("pipeline.handler: a windowing pipeline's handler sql does not read the " + - "event_time column, the time the source assigned the record. A bucket cut " + - "from another field of the payload is on a different clock from the one the " + - "engine checks, and a record the engine refuses as unplaceable is judged on " + - "a time the window never sees. Cut the window's time from event_time where " + - "the source can be told where its time is") - } if isStructuredHandler(h.Type) { if ddl, ok := tableDDL(conf, h.Table); !ok || !declaresTimestamptz(ddl, "event_time") { - warn(fmt.Sprintf("pipeline.handler: a windowing pipeline's table %q does not declare "+ + fail(fmt.Sprintf("pipeline.handler: a windowing pipeline's table %q does not declare "+ "event_time TIMESTAMPTZ, so the time the source assigned each record cannot "+ - "reach the handler's SQL. Declare it and the engine fills it from the record", h.Table)) + "reach the handler's SQL. Declare it and the engine fills it from the record", h.Table), + position(handlerNode(&root))) } } } @@ -142,16 +162,35 @@ func checkWindows(rendered []byte, rep *Report) { i, w.IdleCloseSeconds, flush, w.IdleCloseSeconds, w.IdleCloseSeconds+flush), position(mappingKey(node, "idle_close_seconds")))) } - // reemit publishes emit_sql over the late rows alone, for a bucket - // the sink already holds. A sink that appends keeps both rows, and - // its reader has to add them rather than keep the newest. - if w.LateRows == "reemit" && appendsOnly(w.Sink) { - rep.Add(diagnostic(errs.CodeConfigInvalid, SeverityWarning, fmt.Sprintf( - "tables.sql[%d] window: late_rows is reemit and the %s sink appends, so a "+ - "late row publishes a second row for a bucket the sink already holds, "+ - "computed over the late rows alone. Its reader has to add the two. "+ - "Use drop unless it does", - i, w.Sink.Type), position(mappingKey(node, "late_rows")))) + // The engine decides a record's lateness from trunc(event_time, + // size); the handler must put the row in that bucket and no + // other. This was a warning while the engine only swept the + // table for late rows and could tolerate the disagreement. It + // cannot now. + if !derivesTimeColumn(conf.Pipeline.Handler.SQL, w.TimeColumn, w.SizeSeconds) { + fail(fmt.Sprintf("tables.sql[%d] window: time_column %q must be computed as "+ + "time_bucket(INTERVAL '%d seconds', event_time) in the handler's SQL. The engine "+ + "decides a record's lateness from that bucket, and a row bucketed any other way is "+ + "not in the bucket the engine decided against. Tell the source where the record's "+ + "time is with event_time: {path, format} if it is in the payload", + i, w.TimeColumn, w.SizeSeconds), position(node)) + } + // A late row within lateness republishes its bucket whole, which + // a sink that appends then holds twice. Flink's contract is the + // same: a downstream of a window with allowed lateness must + // handle updates. + if w.AllowedLatenessSecs > 0 { + switch { + case w.LatenessNeedsReplacingSink(): + fail(fmt.Sprintf("tables.sql[%d] window: %s", i, config.LatenessNeedsReplacingSinkMessage(*w)), + position(mappingKey(node, "allowed_lateness_seconds"))) + case w.Sink.Type == "sqlcommand" || w.Sink.Type == "clickhouse": + rep.Add(diagnostic(errs.CodeConfigInvalid, SeverityWarning, fmt.Sprintf( + "tables.sql[%d] window: allowed_lateness_seconds is %d, so a late row republishes "+ + "its bucket as a whole value. The %s sink must replace the row for (bucket, key) "+ + "rather than add to it, which is its SQL's or its table engine's to guarantee", + i, w.AllowedLatenessSecs, w.Sink.Type), position(mappingKey(node, "allowed_lateness_seconds")))) + } } } } @@ -159,6 +198,47 @@ func checkWindows(rendered []byte, rep *Report) { rep.SetCheck("tables.window", status, "") } +// timeBucketOverEventTime matches `time_bucket(INTERVAL '', event_time) AS `, +// the one form a windowing pipeline's time_column may take. The engine +// buckets a record with core.BucketStart, which is time_bucket with DuckDB's +// origin, so the handler must bucket the same way or lateness is decided +// against a bucket the row is not in. +var timeBucketOverEventTime = regexp.MustCompile( + `(?is)time_bucket\s*\(\s*INTERVAL\s+'([^']+)'\s*,\s*event_time\s*\)\s+AS\s+"?([A-Za-z_][A-Za-z0-9_]*)"?`) + +// intervalSeconds reads DuckDB's quoted interval literal for the units a +// window size is written in. false for a form it does not read, which +// validate reports rather than guesses at. +func intervalSeconds(literal string) (int, bool) { + fields := strings.Fields(strings.ToLower(strings.TrimSpace(literal))) + if len(fields) != 2 { + return 0, false + } + n, err := strconv.Atoi(fields[0]) + if err != nil || n <= 0 { + return 0, false + } + per := map[string]int{"second": 1, "minute": 60, "hour": 3600, "day": 86400, "week": 604800}[strings.TrimSuffix(fields[1], "s")] + if per == 0 { + return 0, false + } + return n * per, true +} + +// derivesTimeColumn reports whether handlerSQL computes column as +// time_bucket over event_time with a size of sizeSeconds. +func derivesTimeColumn(handlerSQL, column string, sizeSeconds int) bool { + for _, m := range timeBucketOverEventTime.FindAllStringSubmatch(handlerSQL, -1) { + if !strings.EqualFold(m[2], column) { + continue + } + if secs, ok := intervalSeconds(m[1]); ok && secs == sizeSeconds { + return true + } + } + return false +} + // declaresTimestamptz reports whether a CREATE statement declares column as // TIMESTAMPTZ or TIMESTAMP WITH TIME ZONE. func declaresTimestamptz(createSQL, column string) bool { @@ -191,10 +271,6 @@ func appendsOnly(s config.Sink) bool { func mentionsClosed(sql string) bool { return closedRef.MatchString(sql) } -var eventTimeRef = regexp.MustCompile(`(?i)\bevent_time\b`) - -func mentionsEventTime(sql string) bool { return eventTimeRef.MatchString(sql) } - func hasWindow(conf config.Conf) bool { if conf.Tables == nil { return false @@ -212,19 +288,30 @@ func isStructuredHandler(typ string) bool { return typ == "handlers.StructuredBatch" || typ == "structured" } -// tableDDL is the CREATE the config declares for a named table. +// tableDDL is the CREATE the config declares for a named table: an entry +// under tables.sql, or a CREATE TABLE in a command, which is where the +// bluesky examples declare the structured handler's batch table. func tableDDL(conf config.Conf, name string) (string, bool) { - if conf.Tables == nil { - return "", false + if conf.Tables != nil { + for _, t := range conf.Tables.SQL { + if t.Name == name { + return t.SQL, true + } + } } - for _, t := range conf.Tables.SQL { - if t.Name == name { - return t.SQL, true + for _, c := range conf.Commands { + for _, stmt := range strings.Split(c.SQL, ";") { + if m := createTable.FindStringSubmatch(stmt); m != nil && strings.EqualFold(strings.Trim(m[1], `"`), name) { + return stmt, true + } } } return "", false } +// createTable matches the name a CREATE TABLE statement declares. +var createTable = regexp.MustCompile(`(?is)\bCREATE\s+(?:OR\s+REPLACE\s+)?(?:TEMP(?:ORARY)?\s+)?TABLE\s+(?:IF\s+NOT\s+EXISTS\s+)?("?[A-Za-z_][A-Za-z0-9_]*"?)`) + // handlerNode is the mapping node of pipeline.handler, for a diagnostic's // position; nil, and so no position, if the document has no such node. func handlerNode(root *yaml.Node) *yaml.Node { diff --git a/internal/validate/window_test.go b/internal/validate/window_test.go index a3892801..e270c7b5 100644 --- a/internal/validate/window_test.go +++ b/internal/validate/window_test.go @@ -17,7 +17,6 @@ const windowedConfig = `tables: window: time_column: bucket size_seconds: 60 - late_rows: drop %s sink: type: console @@ -125,63 +124,100 @@ func TestValidateSchema_IndexedWindowTableWithoutStateWarns(t *testing.T) { assert.Equal(t, 0, len(windowDiagnostics(rep))) } -// reemit against a sink that appends is a warning: the config runs, and the -// operator is told the downstream will see corrections. -func TestValidateSchema_ReemitOnAnAppendOnlySinkWarns(t *testing.T) { +// Lateness above zero republishes a bucket as a whole value, which an +// append-only sink holds twice. Refused, not warned: Flink's contract is the +// same. A sink validate cannot classify -- sqlcommand, clickhouse -- is +// warned, because whether it replaces is the SQL's or the table engine's +// call. A sink that replaces by key passes with nothing said. +func TestValidateSchema_LatenessNeedsAReplacingSink(t *testing.T) { coverage.Covers(t, "validate.schema") - rep, err := Validate(context.Background(), Request{Path: "w.yml", Config: `tables: - sql: - - name: agg - sql: CREATE TABLE agg (bucket TIMESTAMPTZ, count INT) - window: - time_column: bucket - size_seconds: 60 - late_rows: reemit - sink: + withSink := func(sink string) string { + cfg := strings.Replace(strings.Replace(windowedConfig, "%s", "bucket TIMESTAMPTZ", 1), + "%s", "allowed_lateness_seconds: 300", 1) + return strings.Replace(cfg, " sink:\n type: console\n", sink, 1) + } + upsert := ` sink: + type: postgres + postgres: + dsn: postgres://u:p@localhost:5432/db + table: agg + mode: upsert + key: [bucket] +` + append := ` sink: type: kafka kafka: brokers: ["localhost:9092"] topic: out -pipeline: - batch_size: 1 - source: - type: kafka - kafka: - brokers: ["localhost:9092"] - group_id: g - auto_offset_reset: earliest - topics: ["t"] - handler: - type: handlers.InferredMemBatch - sql: SELECT time_bucket(INTERVAL '1 minute', event_time) AS bucket, city, count(*) FROM batch GROUP BY ALL - sink: - type: noop -`}) +` + sqlcommand := ` sink: + type: sqlcommand + sqlcommand: + sql: INSERT OR REPLACE INTO out SELECT * FROM sqlflow_sink_batch +` + + rep, err := Validate(context.Background(), Request{Path: "w.yml", Config: withSink(upsert)}) assert.NoError(t, err) assert.That(t, rep.OK) + assert.Equal(t, 0, len(windowDiagnostics(rep))) + + rep, err = Validate(context.Background(), Request{Path: "w.yml", Config: withSink(append)}) + assert.NoError(t, err) + assert.That(t, !rep.OK) diags := windowDiagnostics(rep) assert.Equal(t, 1, len(diags)) + assert.Equal(t, SeverityError, diags[0].Severity) + assert.That(t, strings.Contains(diags[0].Message, "allowed_lateness_seconds is 300 and the kafka sink appends")) + // On the key that made the pairing wrong, not on the sink. + assert.Equal(t, 9, diags[0].Position.Line) + + rep, err = Validate(context.Background(), Request{Path: "w.yml", Config: withSink(sqlcommand)}) + assert.NoError(t, err) + assert.That(t, rep.OK) + diags = windowDiagnostics(rep) + assert.Equal(t, 1, len(diags)) assert.Equal(t, SeverityWarning, diags[0].Severity) - assert.That(t, strings.Contains(diags[0].Message, "the kafka sink appends")) - assert.Equal(t, 8, diags[0].Position.Line) + assert.That(t, strings.Contains(diags[0].Message, "must replace the row")) } -// A window with no late_rows does not pass the schema: the two policies are -// different promises to the sink, and a config has to say which it makes. -func TestValidateSchema_LateRowsIsRequired(t *testing.T) { +// A window with no allowed_lateness_seconds refuses late rows: the key is +// optional and its absence is 0. Nothing is said about it. +func TestValidateSchema_LatenessDefaultsToRefusing(t *testing.T) { coverage.Covers(t, "validate.schema") - rep, err := Validate(context.Background(), Request{Path: "w.yml", Config: strings.Replace( - strings.Replace(strings.Replace(windowedConfig, "%s", "bucket TIMESTAMPTZ", 1), "%s", "", 1), - " late_rows: drop\n", "", 1)}) - assert.NoError(t, err) - assert.That(t, !rep.OK) - var found bool - for _, d := range rep.Diagnostics { - if strings.Contains(d.Message, "late_rows") { - found = true - } + rep := validateWindowed(t, "bucket TIMESTAMPTZ", "") + assert.That(t, rep.OK) + assert.Equal(t, 0, len(windowDiagnostics(rep))) + rep = validateWindowed(t, "bucket TIMESTAMPTZ", "allowed_lateness_seconds: 0") + assert.That(t, rep.OK) + assert.Equal(t, 0, len(windowDiagnostics(rep))) +} + +// The removed keys fail with a message naming the replacement, on the line +// they sit on, so a config written against the old schema learns what +// changed rather than "unknown key". +func TestValidateSchema_RemovedWindowKeysNameTheirReplacement(t *testing.T) { + coverage.Covers(t, "validate.schema") + for key, want := range map[string]string{ + "late_rows: drop": "Set allowed_lateness_seconds instead", + "late_rows: reemit": "Set allowed_lateness_seconds instead", + "poll_interval_seconds: 10": "nothing polls", + } { + t.Run(key, func(t *testing.T) { + coverage.Covers(t, "validate.schema") + rep := validateWindowed(t, "bucket TIMESTAMPTZ", key) + assert.That(t, !rep.OK) + var found *Diagnostic + for i, d := range rep.Diagnostics { + if d.Severity == SeverityError && strings.Contains(d.Message, want) { + found = &rep.Diagnostics[i] + } + } + if found == nil { + t.Fatalf("no error names %q: %v", want, rep.Diagnostics) + } + assert.Equal(t, 9, found.Position.Line) + }) } - assert.That(t, found) } // The old block is refused with its replacement named, on the line it sits @@ -221,7 +257,7 @@ pipeline: if strings.Contains(d.Message, "manager is gone") { found = true assert.That(t, strings.Contains(d.Message, "time_column")) - assert.That(t, strings.Contains(d.Message, "late_rows")) + assert.That(t, strings.Contains(d.Message, "allowed_lateness_seconds")) assert.Equal(t, 5, d.Position.Line) } } @@ -272,34 +308,53 @@ const plainConfig = `pipeline: type: noop ` -// A windowing pipeline's handler should derive the window's time from the -// event_time column -- the time the source assigned -- or the buckets and -// the engine's placement check are on different clocks. A warning until -// every source can be told where its time is. A pipeline with no window is -// free to cut time from any field it likes. -func TestValidateSchema_AWindowingHandlerMustReadEventTime(t *testing.T) { +// A windowing pipeline's time_column must be time_bucket over event_time +// with the window's own size, because the engine decides a record's lateness +// from that bucket and a different expression would put the row somewhere +// else. This was a warning while the manager only swept the table; the +// engine refuses records against that bucket now, so it is an error. A +// pipeline with no window is free to cut time from any field it likes. +func TestValidateSchema_AWindowingHandlerMustBucketEventTime(t *testing.T) { coverage.Covers(t, "validate.schema") - cfg := strings.Replace(strings.Replace(windowedConfig, "%s", "bucket TIMESTAMPTZ", 1), "%s", "", 1) - cfg = strings.Replace(cfg, "time_bucket(INTERVAL '1 minute', event_time)", - "time_bucket(INTERVAL '1 minute', to_timestamp(time_us / 1000000))", 1) - rep, err := Validate(context.Background(), Request{Path: "w.yml", Config: cfg}) - assert.NoError(t, err) - // A warning, not a failure: the config still runs, and validate says - // why it might be on two clocks. - assert.That(t, rep.OK) - diags := windowDiagnostics(rep) - assert.Equal(t, 1, len(diags)) - assert.Equal(t, SeverityWarning, diags[0].Severity) - assert.That(t, strings.Contains(diags[0].Message, "does not read the event_time column")) - assert.That(t, diags[0].Position != nil) + base := strings.Replace(strings.Replace(windowedConfig, "%s", "bucket TIMESTAMPTZ", 1), "%s", "", 1) + const good = "time_bucket(INTERVAL '1 minute', event_time)" - rep, err = Validate(context.Background(), Request{Path: "p.yml", Config: plainConfig}) + // The same size spelled in seconds, and the column quoted, both pass. + for _, ok := range []string{ + "time_bucket(INTERVAL '60 seconds', event_time)", + "TIME_BUCKET(interval '1 MINUTE', event_time)", + } { + rep, err := Validate(context.Background(), Request{Path: "w.yml", Config: strings.Replace(base, good, ok, 1)}) + assert.NoError(t, err) + assert.Equal(t, 0, len(windowDiagnostics(rep))) + } + + for name, bad := range map[string]string{ + "payload field": "time_bucket(INTERVAL '1 minute', to_timestamp(time_us / 1000000))", + "wrong size": "time_bucket(INTERVAL '5 minutes', event_time)", + "date_trunc": "date_trunc('minute', event_time)", + "no derivation": "now()", + } { + t.Run(name, func(t *testing.T) { + coverage.Covers(t, "validate.schema") + rep, err := Validate(context.Background(), Request{Path: "w.yml", Config: strings.Replace(base, good, bad, 1)}) + assert.NoError(t, err) + assert.That(t, !rep.OK) + diags := windowDiagnostics(rep) + assert.Equal(t, 1, len(diags)) + assert.Equal(t, SeverityError, diags[0].Severity) + assert.That(t, strings.Contains(diags[0].Message, "time_bucket(INTERVAL '60 seconds', event_time)")) + assert.That(t, diags[0].Position != nil) + }) + } + + rep, err := Validate(context.Background(), Request{Path: "p.yml", Config: plainConfig}) assert.NoError(t, err) assert.Equal(t, 0, len(windowDiagnostics(rep))) } // A structured handler's batch table is the user's, and the engine fills its -// event_time column from the record; a windowing pipeline on one should +// event_time column from the record; a windowing pipeline on one must // declare that column, as TIMESTAMPTZ, or there is nothing to cut on. func TestValidateSchema_AStructuredWindowingHandlerMustDeclareEventTime(t *testing.T) { coverage.Covers(t, "validate.schema") @@ -320,10 +375,20 @@ func TestValidateSchema_AStructuredWindowingHandlerMustDeclareEventTime(t *testi rep, err = Validate(context.Background(), Request{Path: "s.yml", Config: structured("CREATE TABLE posts (text TEXT, time_us BIGINT)")}) assert.NoError(t, err) - assert.That(t, rep.OK) + assert.That(t, !rep.OK) diags := windowDiagnostics(rep) assert.Equal(t, 1, len(diags)) - assert.Equal(t, SeverityWarning, diags[0].Severity) + assert.Equal(t, SeverityError, diags[0].Severity) + + // The table declared in a command, as the bluesky examples do, is read + // the same way. + inCommand := strings.Replace(strings.Replace(windowedConfig, "%s", "bucket TIMESTAMPTZ", 1), "%s", "", 1) + inCommand = "commands:\n - name: posts\n sql: |\n CREATE TABLE IF NOT EXISTS posts (text TEXT, event_time TIMESTAMPTZ);\n" + inCommand + inCommand = strings.Replace(inCommand, "type: handlers.InferredMemBatch\n", + "type: handlers.StructuredBatch\n table: posts\n", 1) + rep, err = Validate(context.Background(), Request{Path: "s.yml", Config: inCommand}) + assert.NoError(t, err) + assert.Equal(t, 0, len(windowDiagnostics(rep))) assert.That(t, strings.Contains(diags[0].Message, `table "posts" does not declare event_time TIMESTAMPTZ`)) assert.That(t, diags[0].Position != nil) } diff --git a/tests/release/test_image.py b/tests/release/test_image.py index 691dfa1f..ab03e942 100644 --- a/tests/release/test_image.py +++ b/tests/release/test_image.py @@ -1523,10 +1523,10 @@ def test_turbostats_bundle_reports_consumer_lag(image, stack): assert pipeline["message_payload_bytes"] == payload_bytes, ( pipeline["message_payload_bytes"], payload_bytes) - # This config declares a window with late_rows: drop, and its counters are - # present from startup, before any close. Absent would read as "nothing - # here drops rows". - for name in ("late_rows_dropped", "late_rows_reemitted", "window_closed_count"): + # This config declares a window that refuses late rows, and its counters + # are present from startup, before any close. Absent would read as + # "nothing here drops rows". + for name in ("late_rows_dropped", "late_rows_recomputed", "window_closed_count"): assert name in pipeline, (name, pipeline) assert pipeline["late_rows_dropped"] == 0, pipeline diff --git a/turbostats/wire/bundle.go b/turbostats/wire/bundle.go index cd49864e..10e77e52 100644 --- a/turbostats/wire/bundle.go +++ b/turbostats/wire/bundle.go @@ -281,16 +281,21 @@ type Pipeline struct { // when it runs none. // // That is what makes LateRowsDropped readable. It counts rows the engine - // deleted because they arrived after the watermark, which is silent data + // refused before the handler because their bucket had closed more than + // allowed_lateness_seconds before the watermark, which is silent data // loss, and an operator has to be able to tell "no rows were dropped" // from "nothing here drops rows". Absent says the second; zero says the - // first. + // first. LateRowsRecomputed counts rows admitted within the lateness, + // each of which republished its bucket whole; nothing was lost. // - // Dropped and reemitted are separate fields because the policy that - // splits them is an outcome, not a shard: one number loses data and the - // other does not, and a sum of the two is true of neither. Both are - // absent if a window reports a policy this contract has no field for, - // because a count that leaves some rows out is worse than none. + // Dropped and recomputed are separate fields because the outcome that + // splits them is not a shard: one number loses data and the other does + // not, and a sum of the two is true of neither. Both are absent if a + // window reports an outcome this contract has no field for, because a + // count that leaves some rows out is worse than none. late_rows_reemitted + // was the field before recomputes replaced the delta; an older receiver + // sees it absent, the way it sees every window field of a pipeline with + // no window. // // The unit is rows of the window table, not source events. The handler's // SQL runs before the window sees anything, so a batch of forty late @@ -300,9 +305,9 @@ type Pipeline struct { // input count is comparing different units. A count in events would // need the window to know which column carries each row's event count, // which nothing declares today. - LateRowsDropped *int64 `json:"late_rows_dropped,omitempty"` - LateRowsReemitted *int64 `json:"late_rows_reemitted,omitempty"` - WindowClosedCount *int64 `json:"window_closed_count,omitempty"` + LateRowsDropped *int64 `json:"late_rows_dropped,omitempty"` + LateRowsRecomputed *int64 `json:"late_rows_recomputed,omitempty"` + WindowClosedCount *int64 `json:"window_closed_count,omitempty"` // WindowLagSeconds is how far the most behind window's closes trail the // data it holds, in event seconds: where its watermark should be, given