diff --git a/MEMORY.md b/MEMORY.md index f40fa307..67d70f2c 100644 --- a/MEMORY.md +++ b/MEMORY.md @@ -41,3 +41,6 @@ One install answers at several addresses at once (the loopback port, the tunnel ## Every feature lands on three fronts: levers, defaults, protection Before building anything, answer all three, and say so. **Levers**: what a person can reach and from where (a node setting, a CLI flag, a graph control, a ctx call, a `metadata.json` key, a word in the language). Name the lever and where it lives; a knob nobody can turn is not a lever, a knob nobody needs is clutter. **Defaults**: what everybody gets without asking, which is the part nobody should have to know exists. Two questions: is it tedious and useful to make somebody fill it in every time, and is there a value right for almost everybody? Yes to both, set it and keep the lever. Never invent a default to paper over a setting that honestly depends on the case: if leaving it unset is legitimate, it stays explicit and optional and the VALIDATION is what makes the bad shape impossible. A hidden default that is right half the time is worse than a refusal naming what is missing. **Protection**: as long as there is a way to build it properly, the bad shape is refused and the message points at the right one, at the earliest place that can see it (compiler, parser, a node's metadata rules, and a runtime floor where nothing earlier can). An obvious footgun (an infinite loop, a run that waits forever on somebody gone) must be impossible to configure, not warned about. The full wording: `docs/src/thinking/how-we-decide.md`. + +## Every slow command runs in the background, with a realistic cap +Anything that can take more than a few seconds runs with `run_in_background`: builds, tests, installs, and cloud work too (`gcloud ... create`, a Cloud SQL instance, `terraform apply`, project or machine creation). Never in the foreground, never "just this once"; an unknown duration counts as slow. Every wait on it is capped at how long that thing normally takes (a cargo check about a minute, an e2e test at most 5, a Cloud SQL create about 10), never a catch-all like 25 minutes: at the cap, look at what it is doing (its output, its operation's state) and tell the [user]. A foreground `gcloud sql instances create` once sat silent for many minutes until the [user] backgrounded it by hand. diff --git a/catalog/postgres/database/mod.rs b/catalog/postgres/database/mod.rs index ecfad168..b4c2a470 100644 --- a/catalog/postgres/database/mod.rs +++ b/catalog/postgres/database/mod.rs @@ -249,9 +249,15 @@ impl Node for PostgresDatabaseNode { async fn run(&self, ctx: ExecutionContext) -> WeftResult<()> { let database: String = ctx.inputs.get("database")?; - let sql = ctx.endpoint("sql").await?; + // Three asks that need nothing from each other, so they go out + // together. + let (sql, credential, published) = + tokio::try_join!(ctx.endpoint("sql"), ctx.endpoint("credential"), ctx.published_access())?; let (host, port) = sql.host_and_port()?; - let credential = ctx.endpoint("credential").await?; + let published = match published { + Some(access) => Some(ctx.open(&access).await?), + None => None, + }; // The password comes from the connection this node published // last time when there is one, and from the database itself @@ -263,9 +269,9 @@ impl Node for PostgresDatabaseNode { // Asking the database whether it still knows a password also // retires it, so a run that gets a yes here has already done // the retiring this run owes. - let (password, retired) = match ctx.published_access().await? { - Some(published) => { - let held = ctx.open(&published).await?.value("password")?.to_string(); + let (password, retired) = match &published { + Some(opened) => { + let held = opened.value("password")?.to_string(); if confirm_stored(&credential, &held).await? { (held, true) } else { @@ -285,6 +291,8 @@ impl Node for PostgresDatabaseNode { values.insert("password".to_string(), password.clone()); // Inside the project's own network; the database serves no TLS. values.insert("sslmode".to_string(), "disable".to_string()); + // Published every run, values unchanged or not: publishing also + // brings the connection's recipe and label up to this node's. let access = ctx.publish_access(values).await?; // Retire the password now that a connection holds it, unless diff --git a/crates/weft-broker-client/src/client.rs b/crates/weft-broker-client/src/client.rs index 11184ddd..5bee27b8 100644 --- a/crates/weft-broker-client/src/client.rs +++ b/crates/weft-broker-client/src/client.rs @@ -208,6 +208,15 @@ where } } +/// Whether a failed broker call never reached the broker: the connection +/// itself was refused or could not be made, so the request was never +/// sent. Only such a write is safe to send again; a write that reached it +/// and failed some other way may have landed, and sending it twice would +/// apply it twice. +fn never_sent(e: &anyhow::Error) -> bool { + e.chain().any(|cause| cause.downcast_ref::().is_some_and(reqwest::Error::is_connect)) +} + /// The first wait before asking again after the broker could not answer /// a read; each further failure doubles it, up to /// [`READ_RETRY_LONGEST`]. @@ -275,6 +284,15 @@ impl BrokerJournalClient { #[async_trait] impl JournalClient for BrokerJournalClient { + /// The broker takes a body of at most [`JOURNAL_RECORD_BODY_LIMIT`]. + fn parts<'e>(&self, events: &'e [ExecEvent]) -> Result> { + record_chunks(events) + } + + fn never_reached(&self, error: &anyhow::Error) -> bool { + never_sent(error) + } + async fn record_event( &self, event: &ExecEvent, @@ -1229,7 +1247,9 @@ mod read_retry_tests { } /// `events` cut, in order, into runs whose request bodies stay under -/// [`JOURNAL_RECORD_BODY_LIMIT`]. An event too big on its own goes alone, +/// [`JOURNAL_RECORD_BODY_LIMIT`]: each one request of +/// [`BrokerJournalClient::record_events`], and the parts a writer that +/// sends a write again sends one at a time (`JournalClient::parts`). An event too big on its own goes alone, /// and the broker refuses it (`413`): the engine refuses an output that /// big at the node first, so reaching that refusal is a broken contract. fn record_chunks(events: &[ExecEvent]) -> Result> { diff --git a/crates/weft-broker/src/handlers.rs b/crates/weft-broker/src/handlers.rs index 4b626d1e..7f65c2fd 100644 --- a/crates/weft-broker/src/handlers.rs +++ b/crates/weft-broker/src/handlers.rs @@ -646,16 +646,19 @@ pub async fn task_claim_one( require_worker(&caller)?; require_replica_matches(&caller, &req.replica)?; let filter = req.filter; - if let ClaimFilter::ExecutionId { project_id, .. } = &filter { - scope::require_project_owned_by(&state.scope_cache, &state.pool, &caller, *project_id).await?; - } else { + let ClaimFilter::ExecutionId { project_id, execution_id } = &filter else { return Err((StatusCode::FORBIDDEN, "a worker claims only the execution it was called for".into())); - } - let task = state - .tasks - .claim_one(&req.replica, filter, held(req.wait_ms)) - .await - .map_err(internal)?; + }; + scope::require_project_owned_by(&state.scope_cache, &state.pool, &caller, *project_id).await?; + // The worker's next asks (its run's journal first) are scoped by the + // execution, so its scope is read while the claim runs rather than + // after it, on that first ask. + let execution_id = execution_id.clone(); + let (task, ()) = tokio::join!( + state.tasks.claim_one(&req.replica, filter, held(req.wait_ms)), + scope::warm_execution_id_scope(&state.scope_cache, &state.pool, &execution_id), + ); + let task = task.map_err(internal)?; // Latest-claim-wins execution ownership is bound IN the claim's own // transaction by the `task_claim_binds_execution_id_owner` DB trigger: // claiming an execution-bearing task atomically stamps @@ -805,7 +808,7 @@ async fn declared_infra_under( project: uuid::Uuid, digest: String, ) -> Result>, (StatusCode, String)> { - if let Some(declared) = state.declared_infra.get(project, &digest) { + if let Some(declared) = state.declared_infra.get(&(project, digest.clone())) { return Ok(Some(declared)); } // Read with its own digest, so what is kept is filed under the @@ -820,7 +823,7 @@ async fn declared_infra_under( let definition: weft_core::project::ProjectDefinition = serde_json::from_str(&project_json) .map_err(|e| internal(anyhow::anyhow!("project {project}: definition: {e}")))?; let declared = Arc::new(weft_core::project::DeclaredInfra::of(&definition)); - state.declared_infra.put(project, digest, declared.clone()); + state.declared_infra.put((project, digest), declared.clone()); Ok(Some(declared)) } diff --git a/crates/weft-broker/src/scope.rs b/crates/weft-broker/src/scope.rs index a4dd51eb..78ad6d90 100644 --- a/crates/weft-broker/src/scope.rs +++ b/crates/weft-broker/src/scope.rs @@ -237,6 +237,16 @@ async fn lookup_project_tenant( Ok(tenant) } +/// Read `execution_id`'s scope into the cache ahead of the asks that +/// need it. Only a head start: an execution that cannot be read here is +/// read again, and refused with the reason, by the first ask that needs +/// it, so nothing is lost by not answering here. +pub async fn warm_execution_id_scope(cache: &ScopeCache, pool: &PgPool, execution_id: &str) { + if let Err((_, why)) = lookup_execution_id_scope(cache, pool, execution_id).await { + tracing::debug!(target: "weft_broker::scope", execution_id, why, "could not read an execution's scope ahead of its asks"); + } +} + async fn lookup_execution_id_scope( cache: &ScopeCache, pool: &PgPool, diff --git a/crates/weft-broker/src/state.rs b/crates/weft-broker/src/state.rs index f6210136..baa66611 100644 --- a/crates/weft-broker/src/state.rs +++ b/crates/weft-broker/src/state.rs @@ -49,7 +49,7 @@ pub struct BrokerState { /// project and a digest of its definition: a definition never changes /// under its digest, so an entry is never stale, and a program asking /// for an endpoint on every run reads its definition once. - pub declared_infra: weft_core::content_cache::ContentCache, + pub declared_infra: weft_core::content_cache::ContentCache<(uuid::Uuid, String), weft_core::project::DeclaredInfra>, /// Where runtime-file bytes live: the install's bucket. pub object_store: Arc, /// The runtime-file plane (`ctx.storage`): PG metadata + bucket bytes, diff --git a/crates/weft-cli/src/commands/ci.rs b/crates/weft-cli/src/commands/ci.rs index ab3bf4db..0a8322b6 100644 --- a/crates/weft-cli/src/commands/ci.rs +++ b/crates/weft-cli/src/commands/ci.rs @@ -117,7 +117,7 @@ mod tests { assert!(first.contains("-frontends/my-app-front:"), "{first}"); write_workflow(dir.path(), "my app", Cloud::Gcp).unwrap(); - std::fs::write(&path, first.replace("runs-on: ubuntu-latest", "runs-on: self-hosted")).unwrap(); + std::fs::write(&path, first.replace("runs-on: ubuntu-24.04", "runs-on: self-hosted")).unwrap(); let e = write_workflow(dir.path(), "my app", Cloud::Gcp).unwrap_err().to_string(); assert!(e.contains("changed since weft wrote it"), "{e}"); assert!(std::fs::read_to_string(&path).unwrap().contains("self-hosted"), "left as it is"); diff --git a/crates/weft-cli/src/commands/executions.rs b/crates/weft-cli/src/commands/executions.rs index 39ab88bd..798f0b12 100644 --- a/crates/weft-cli/src/commands/executions.rs +++ b/crates/weft-cli/src/commands/executions.rs @@ -5,7 +5,7 @@ use anyhow::Context; use weft_core::program::ExecutionPage; -use super::{local_time, Ctx}; +use super::{utc_time, Ctx}; /// A value put into a query string. A node id is the author's own /// spelling, so it can hold anything they typed; only the handful of @@ -91,7 +91,7 @@ pub async fn list(ctx: Ctx, filter: ListFilter) -> anyhow::Result<()> { return Ok(()); } println!( - "{:<36} {:<9} {:<13} {:<19} {:<36} entry_node tags", + "{:<36} {:<9} {:<13} {:<23} {:<36} entry_node tags", "execution_id", "status", "phase", "started", "project_id" ); for row in &page.executions { @@ -104,8 +104,8 @@ pub async fn list(ctx: Ctx, filter: ListFilter) -> anyhow::Result<()> { // Which instance the run is in, when it is in one. let instance = row.instance.as_ref().map(|m| format!(" (instance {m})")).unwrap_or_default(); println!( - "{execution_id:<36} {status:<9} {phase:<13} {:<19} {project:<36} {entry}{tags}{instance}", - local_time(row.started_at) + "{execution_id:<36} {status:<9} {phase:<13} {:<23} {project:<36} {entry}{tags}{instance}", + utc_time(row.started_at) ); } // The server clamps the page size, so a big --limit can come back @@ -248,8 +248,8 @@ pub fn event_line(row: &serde_json::Value, full: bool) -> String { // pulse, a corruption the replay found) have none and get a blank // of the same width, so the columns still line up. let at = match row.get("at_unix").and_then(|v| v.as_u64()) { - Some(at) => format!("[{:<19}]", local_time(at)), - None => " ".repeat(21), + Some(at) => format!("[{:<23}]", utc_time(at)), + None => " ".repeat(25), }; let node = row_node(row).unwrap_or(""); let mut line = format!("{at} {kind:<23} {node}"); diff --git a/crates/weft-cli/src/commands/frontend.rs b/crates/weft-cli/src/commands/frontend.rs index 40a2600f..b0e83ed3 100644 --- a/crates/weft-cli/src/commands/frontend.rs +++ b/crates/weft-cli/src/commands/frontend.rs @@ -10,7 +10,7 @@ //! to the terminal. use anyhow::{Context, Result}; -use weft_core::frontend::{AddFrontendRequest, Frontend, FrontendHost, FrontendWithToken, Repository}; +use weft_core::frontend::{AddFrontendRequest, AddedFrontend, Frontend, FrontendHost, FrontendWithToken, Repository}; use super::Ctx; @@ -71,9 +71,10 @@ pub async fn run(ctx: Ctx, action: FrontendAction) -> Result<()> { } else { made.await }; - let made: FrontendWithToken = + let made: AddedFrontend = serde_json::from_value(made?).context("read the frontend the install made")?; - hand_over(&ctx, &project, &install_url, made)?; + let token = made.token.zip(made.frontend.token_id); + hand_over(&ctx, &project, &install_url, &made.frontend, token)?; } FrontendAction::Token { name, done: None } => { // A hosted frontend's token is its workflow's: a new one made @@ -92,7 +93,7 @@ pub async fn run(ctx: Ctx, action: FrontendAction) -> Result<()> { serde_json::from_value(client.post_json(&format!("{base}/{name}/token"), &serde_json::json!({})).await?) .context("read the frontend's new token")?; let id = renewed.token_id; - hand_over(&ctx, &project, &install_url, renewed)?; + hand_over(&ctx, &project, &install_url, &renewed.frontend, Some((renewed.token, id)))?; if !ctx.json() { println!( "its old token keeps working until the new one is in place; then `weft frontend token {name} --done {id}` retires it" @@ -138,15 +139,14 @@ pub async fn run(ctx: Ctx, action: FrontendAction) -> Result<()> { Ok(()) } -/// Put a fresh token where only this person can read it, and say what -/// goes where. A frontend the install hosts gets no file: its service -/// gets a token of its own from the repository's workflow (`weft target -/// export` mints it, the deploy puts it in place and retires every other), -/// so one written here would only sit on disk until that deploy killed it. -fn hand_over(ctx: &Ctx, project: &str, install_url: &str, made: FrontendWithToken) -> Result<()> { - let f = &made.frontend; +/// Put a fresh `token` (with its id) where only this person can read it, +/// and say what goes where. A frontend the install hosts has none to +/// write: its service gets its token from the repository's workflow +/// (`weft target export` mints it, the deploy puts it in place and retires +/// every other). +fn hand_over(ctx: &Ctx, project: &str, install_url: &str, f: &Frontend, token: Option<(String, uuid::Uuid)>) -> Result<()> { if let (Some(repo), Some(service)) = (&f.repo, &f.service) { - if ctx.json_out(&serde_json::json!({ "frontend": f, "tokenId": made.token_id }))? { + if ctx.json_out(&serde_json::json!({ "frontend": f }))? { return Ok(()); } println!("frontend '{}': the install made its service {service}, and {} may deploy to it", f.name, repo.name); @@ -162,13 +162,15 @@ fn hand_over(ctx: &Ctx, project: &str, install_url: &str, made: FrontendWithToke } // A frontend running elsewhere reaches the install at its public // address. + let (token, token_id) = + token.with_context(|| format!("the install made frontend '{}' with no token for it to call with", f.name))?; let env = vec![ - ("WEFT_TOKEN", made.token.clone()), + ("WEFT_TOKEN", token), ("WEFT_DISPATCHER_URL", install_url.to_string()), ("WEFT_PUBLIC_URL", install_url.to_string()), ]; let file = super::target::write_secrets_file(&format!("{project}-frontend-{}", f.name), &env)?; - if ctx.json_out(&serde_json::json!({ "frontend": f, "tokenId": made.token_id, "tokenFile": file }))? { + if ctx.json_out(&serde_json::json!({ "frontend": f, "tokenId": token_id, "tokenFile": file }))? { return Ok(()); } println!("frontend '{}': its token is in {} (readable by you only; shown this once)", f.name, file.display()); diff --git a/crates/weft-cli/src/commands/logs.rs b/crates/weft-cli/src/commands/logs.rs index 90f335af..d9fa52b7 100644 --- a/crates/weft-cli/src/commands/logs.rs +++ b/crates/weft-cli/src/commands/logs.rs @@ -12,7 +12,7 @@ use anyhow::Context; use weft_core::program::{ExecutionLogs, ExecutionSummary}; -use super::{local_time, resolve_project_id, Ctx}; +use super::{utc_time, resolve_project_id, Ctx}; /// The nodes a run skipped, with why, spelled the way the program /// reads them (`keep.db` for the db of an included file) when the cwd @@ -117,7 +117,7 @@ pub async fn run(ctx: Ctx, target: Option, limit: Option) -> anyhow .inherited_from .map(|execution_id| format!(" [inherited from {}]", super::versions::short(&execution_id.to_string()))) .unwrap_or_default(); - println!("[{}] {level:>5}{node}{inherited} {msg}", local_time(entry.at_unix)); + println!("[{}] {level:>5}{node}{inherited} {msg}", utc_time(entry.at_unix)); } // The dispatcher answers the tail, so a full page means the run // may have written more than this; a cut log must never read as diff --git a/crates/weft-cli/src/commands/mod.rs b/crates/weft-cli/src/commands/mod.rs index 8a17dce3..bac14bb7 100644 --- a/crates/weft-cli/src/commands/mod.rs +++ b/crates/weft-cli/src/commands/mod.rs @@ -524,37 +524,35 @@ pub fn resolve_project( )) } -/// A journal unix stamp as the local wall-clock time a person reads -/// (`2026-09-02 21:36:47`), the one rendering every listing verb uses -/// so a run's start, its events and its log lines line up by eye. -/// Zero (a row that never carried a stamp) renders as a dash rather -/// than as 1970. -pub fn local_time(unix_secs: u64) -> String { +/// A journal unix stamp as the UTC time a person reads +/// (`2026-09-02 21:36:47 UTC`), the one rendering every listing verb +/// uses so a run's start, its events and its log lines line up by eye, +/// with the install's own logs too, wherever the install runs. Zero (a +/// row that never carried a stamp) renders as a dash rather than as +/// 1970. +pub fn utc_time(unix_secs: u64) -> String { if unix_secs == 0 { return "-".to_string(); } // A stamp too large for the calendar prints as its number rather // than as a wrapped-around date. match i64::try_from(unix_secs).ok().and_then(|s| chrono::DateTime::::from_timestamp(s, 0)) { - Some(t) => t.with_timezone(&chrono::Local).format("%Y-%m-%d %H:%M:%S").to_string(), + Some(t) => t.format("%Y-%m-%d %H:%M:%S UTC").to_string(), None => unix_secs.to_string(), } } #[cfg(test)] mod tests { - use super::{local_time, resolve_dispatcher_url, Project}; + use super::{utc_time, resolve_dispatcher_url, Project}; /// The stamp renders as a date and a time, and the two sentinels /// (zero, out of range) never render as a bogus date. #[test] - fn renders_a_readable_local_time() { - let text = local_time(1_756_838_207); - assert_eq!(text.len(), "2026-09-02 21:36:47".len(), "{text}"); - assert_eq!(&text[4..5], "-"); - assert_eq!(&text[10..11], " "); - assert_eq!(local_time(0), "-"); - assert_eq!(local_time(u64::MAX), u64::MAX.to_string()); + fn renders_a_readable_utc_time() { + assert_eq!(utc_time(1_756_838_207), "2025-09-02 18:36:47 UTC"); + assert_eq!(utc_time(0), "-"); + assert_eq!(utc_time(u64::MAX), u64::MAX.to_string()); } fn project_with_prod() -> (tempfile::TempDir, Project) { diff --git a/crates/weft-cli/src/commands/status.rs b/crates/weft-cli/src/commands/status.rs index a33c7b11..bd75414b 100644 --- a/crates/weft-cli/src/commands/status.rs +++ b/crates/weft-cli/src/commands/status.rs @@ -172,9 +172,12 @@ pub async fn run(ctx: Ctx) -> Result<()> { println!(" triggers:"); for entry in &data.activations { let (trigger, mode) = (&entry.trigger, entry.mode.as_str()); + // Which version of the source its fires run, as `weft tree` + // names it. + let version = entry.version.as_deref().map(|v| format!(", version {}", super::versions::short(v))).unwrap_or_default(); match &entry.instance { - Some(instance) => println!(" {trigger} (instance {instance}): {mode}"), - None => println!(" {trigger}: {mode}"), + Some(instance) => println!(" {trigger} (instance {instance}): {mode}{version}"), + None => println!(" {trigger}: {mode}{version}"), } // Fires parked until the instance is given a value they // need: its next change of values routes them again. diff --git a/crates/weft-cli/src/commands/target.rs b/crates/weft-cli/src/commands/target.rs index ca1103a6..2413f94d 100644 --- a/crates/weft-cli/src/commands/target.rs +++ b/crates/weft-cli/src/commands/target.rs @@ -361,7 +361,7 @@ async fn export( for id in &stale { println!("an earlier export's key {id} still works; once these are in place: weft token revoke {id} --on {name}"); } - if let Some((_, frontend, _)) = &minted.frontend { + if let Some(frontend) = deployed_before(&minted, frontend) { println!( "frontend '{frontend}' keeps its old token working until the deploy workflow puts the new one in place and retires it" ); @@ -380,7 +380,7 @@ async fn export( // terminal, its name included. println!(" secrets {} set", settings.secrets.len()); println!("run its deploy workflow from the Actions tab"); - if let Some((_, frontend, _)) = &minted.frontend { + if let Some(frontend) = deployed_before(&minted, frontend) { println!("frontend '{frontend}' keeps its old token working until that run deploys the new one and retires it"); } // The repository now holds the new keys, so an earlier export's are @@ -574,6 +574,14 @@ async fn mint_ci_keys( Ok((CiKeys { operator_key, frontend, front_env }, made)) } +/// The frontend this export made a new token for, when it already calls +/// with one: that one keeps working until the new one is in place. A +/// hosted frontend never deployed has none, so nothing is kept. +fn deployed_before<'a>(minted: &'a Minted, frontend: Option<&weft_core::frontend::Frontend>) -> Option<&'a str> { + let (_, name, _) = minted.frontend.as_ref()?; + frontend.is_some_and(|f| f.token_id.is_some()).then_some(name.as_str()) +} + /// The frontend an export hands the workflow: its name, its service, and /// its new token. struct HandedFrontend { @@ -891,7 +899,7 @@ mod tests { repo: repo.map(|r| weft_core::frontend::Repository { name: r.into(), id: if r == "me/shop" { 1 } else { 2 } }), service: Some(format!("fe-{name}")), url: None, - token_id: uuid::Uuid::nil(), + token_id: Some(uuid::Uuid::nil()), pending_token_ids: Vec::new(), }; let all = [ diff --git a/crates/weft-cli/src/commands/tree.rs b/crates/weft-cli/src/commands/tree.rs index 5bb3a285..89271d41 100644 --- a/crates/weft-cli/src/commands/tree.rs +++ b/crates/weft-cli/src/commands/tree.rs @@ -24,7 +24,7 @@ pub async fn run(ctx: Ctx) -> anyhow::Result<()> { ctx.json_out(&raw)?; return Ok(()); } - for line in render(&tree, &super::local_time) { + for line in render(&tree, &super::utc_time) { println!("{line}"); } Ok(()) @@ -126,7 +126,7 @@ fn render_version<'a>( for r in runs.get(v.id.as_str()).cloned().unwrap_or_default() { let head = if tree.head.head_run == Some(r.execution_id) { " <- HEAD run" } else { "" }; let seed = r.seed_execution_id.map(|s| format!(" seed {} ({} stale)", short(&s.to_string()), r.stale.len())).unwrap_or_default(); - let scope = r.spec.as_ref().map(|s| format!(" spec {}", s.name)).unwrap_or_default(); + let scope = r.spec.as_ref().map(started_by).unwrap_or_default(); let example = r.example.as_deref().map(|e| format!(" example {e}")).unwrap_or_default(); let ended = r.completed_at.map(|t| format!(" -> {}", when(t))).unwrap_or_default(); // A cancelled run says who ended it, because a person stopping @@ -154,6 +154,17 @@ fn render_version<'a>( } } +/// What started a run, in the words a person reads it by: the trigger +/// that fired it, or the starting parameters it was run with (`weft run +/// --save` names them). A plain run says nothing. +fn started_by(spec: &weft_core::run_spec::RunSpec) -> String { + match (&spec.fire, spec.name.as_str()) { + (Some((trigger, _)), _) => format!(" fired by {trigger}"), + (None, "") => String::new(), + (None, name) => format!(" with starting parameters '{name}'"), + } +} + #[cfg(test)] mod tests { use super::*; diff --git a/crates/weft-cli/templates/ci/gcp.yml b/crates/weft-cli/templates/ci/gcp.yml index 3e8dbb7d..971f2397 100644 --- a/crates/weft-cli/templates/ci/gcp.yml +++ b/crates/weft-cli/templates/ci/gcp.yml @@ -31,7 +31,9 @@ concurrency: jobs: deploy: - runs-on: ubuntu-latest + # A named release: `ubuntu-latest` moves to a new one on GitHub's + # schedule, under a workflow that worked the day before. + runs-on: ubuntu-24.04 env: WEFT_OPERATOR_KEY: ${{ secrets.WEFT_OPERATOR_KEY }} WEFT_REPO_ROOT: ${{ github.workspace }}/.weft diff --git a/crates/weft-core/src/content_cache.rs b/crates/weft-core/src/content_cache.rs index 66e0e822..0aefe3e3 100644 --- a/crates/weft-core/src/content_cache.rs +++ b/crates/weft-core/src/content_cache.rs @@ -1,35 +1,50 @@ -//! A bounded in-memory cache of values derived from content that never -//! changes under its name: a program recorded under its definition hash, -//! what a definition declares under its digest. Content under a name never -//! changes, so every process keeps its own and nothing invalidates it; the -//! bound keeps the memory flat, dropping the least recently used. Content -//! deleted from the store (a removed project's unused versions) may still -//! be answered by a process that read it before, which is harmless: nothing -//! still in use names it. +//! A bounded in-memory cache, dropping the least recently used to keep the +//! memory flat. Mostly for values derived from content that never changes +//! under its name: a program recorded under its definition hash, what a +//! definition declares under its digest. That content never changes, so +//! every process keeps its own and nothing invalidates it; content deleted +//! from the store (a removed project's unused versions) may still be +//! answered by a process that read it before, which is harmless: nothing +//! still in use names it. A copy of rows that do change +//! (`weft_task_store::held_copy`) keeps them here too, and drops an entry +//! (`forget`) the moment its rows change. +use std::hash::Hash; use std::sync::{Arc, Mutex}; -/// See the module doc. Keyed by the project and the content's own name -/// (its hash or digest). -pub struct ContentCache { - entries: Mutex>>, +/// See the module doc. Keyed by what names the value: a program by its +/// project and definition hash, what a definition declares by its digest, +/// a held copy's rows by whose they are. +pub struct ContentCache { + entries: Mutex>>, } -impl ContentCache { +impl ContentCache { /// A cache holding at most `capacity` entries. pub fn new(capacity: usize) -> Self { let capacity = std::num::NonZeroUsize::new(capacity).expect("a content cache holds at least one entry"); Self { entries: Mutex::new(lru::LruCache::new(capacity)) } } - /// The value `project` holds under `name`, if this process has it. - pub fn get(&self, project: uuid::Uuid, name: &str) -> Option> { - self.entries.lock().expect("content cache").get(&(project, name.to_string())).cloned() + /// The value held under `key`, if this process has it. + pub fn get(&self, key: &K) -> Option> { + self.entries.lock().expect("content cache").get(key).cloned() } - /// Keep `value` as what `project` holds under `name`. - pub fn put(&self, project: uuid::Uuid, name: String, value: Arc) { - self.entries.lock().expect("content cache").put((project, name), value); + /// Keep `value` as what is held under `key`. + pub fn put(&self, key: K, value: Arc) { + self.entries.lock().expect("content cache").put(key, value); + } + + /// Drop what is held under `key`: for a copy of rows that change, never + /// for content under its name. + pub fn forget(&self, key: &K) { + self.entries.lock().expect("content cache").pop(key); + } + + /// Drop everything held. + pub fn forget_all(&self) { + self.entries.lock().expect("content cache").clear(); } } @@ -43,12 +58,12 @@ mod tests { fn values_come_back_by_name_and_the_oldest_goes_first() { let cache = ContentCache::new(2); let (a, b) = (uuid::Uuid::from_u128(1), uuid::Uuid::from_u128(2)); - cache.put(a, "h1".into(), Arc::new(1)); - cache.put(b, "h1".into(), Arc::new(2)); - assert_eq!(cache.get(a, "h1").as_deref(), Some(&1)); - assert!(cache.get(a, "h2").is_none()); - cache.put(a, "h3".into(), Arc::new(3)); - assert!(cache.get(b, "h1").is_none(), "the least recently used went"); - assert_eq!(cache.get(a, "h1").as_deref(), Some(&1)); + cache.put((a, "h1".to_string()), Arc::new(1)); + cache.put((b, "h1".to_string()), Arc::new(2)); + assert_eq!(cache.get(&(a, "h1".to_string())).as_deref(), Some(&1)); + assert!(cache.get(&(a, "h2".to_string())).is_none()); + cache.put((a, "h3".to_string()), Arc::new(3)); + assert!(cache.get(&(b, "h1".to_string())).is_none(), "the least recently used went"); + assert_eq!(cache.get(&(a, "h1".to_string())).as_deref(), Some(&1)); } } diff --git a/crates/weft-core/src/frontend.rs b/crates/weft-core/src/frontend.rs index 1d51c98f..958ee4be 100644 --- a/crates/weft-core/src/frontend.rs +++ b/crates/weft-core/src/frontend.rs @@ -58,9 +58,11 @@ pub struct Frontend { #[serde(default, skip_serializing_if = "Option::is_none")] pub url: Option, /// The caller token it calls the install with (its id; the value is - /// shown once, when it is made). - #[serde(rename = "tokenId")] - pub token_id: uuid::Uuid, + /// shown once, when it is made), once one is in place. A frontend the + /// install hosts has none until its deploy workflow puts the first in + /// place: the workflow's is the only token it ever calls with. + #[serde(default, rename = "tokenId", skip_serializing_if = "Option::is_none")] + pub token_id: Option, /// New tokens made to replace it and not put in place yet: they all /// work until one is (`POST .../token/{id}/done`), so a running site is /// never left without a working token. @@ -87,8 +89,18 @@ pub struct AddFrontendRequest { pub repo: Option, } -/// What making a frontend, or giving it a new token, answers: the -/// frontend and its token, shown this once. +/// What making a frontend answers: the frontend, and, for one running +/// elsewhere, the token it calls with, shown this once. One the install +/// hosts gets its token from its deploy workflow (`weft target export`). +#[derive(Debug, Clone, Serialize, Deserialize)] +pub struct AddedFrontend { + pub frontend: Frontend, + #[serde(default, skip_serializing_if = "Option::is_none")] + pub token: Option, +} + +/// What giving a frontend a new token answers: the frontend and the +/// token, shown this once. #[derive(Debug, Clone, Serialize, Deserialize)] pub struct FrontendWithToken { pub frontend: Frontend, diff --git a/crates/weft-core/src/projects.rs b/crates/weft-core/src/projects.rs index 0d75ee0d..f2e393ef 100644 --- a/crates/weft-core/src/projects.rs +++ b/crates/weft-core/src/projects.rs @@ -258,6 +258,11 @@ pub struct ActivationEntry { /// gave an invalid) value they need: how many, and why. #[serde(default, skip_serializing_if = "Option::is_none")] pub waiting: Option, + /// The version of the source its fires run (`weft tree` shows where + /// it sits), once its setup finished; `None` while it is set up and + /// once it is taken down. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub version: Option, } /// What a deactivated project kept, for the reactivate-time choice. diff --git a/crates/weft-dispatcher/src/activation_store.rs b/crates/weft-dispatcher/src/activation_store.rs index 7962d2bd..65ce552b 100644 --- a/crates/weft-dispatcher/src/activation_store.rs +++ b/crates/weft-dispatcher/src/activation_store.rs @@ -286,6 +286,34 @@ pub static GROUP: weft_task_store::SchemaGroup = weft_task_store::SchemaGroup { FOR EACH ROW WHEN (NEW.status IS DISTINCT FROM OLD.status) EXECUTE FUNCTION signal_held_notify()"#, + // An activation's status is part of how its routes read (only an + // active one takes a caller), so a change tells every dispatcher + // to read its tenant's routes again (`crate::held::Held::routes`). + // SYNC: 'weft_routes' <-> crate::held::ROUTES_CHANNEL + r#"CREATE OR REPLACE FUNCTION trigger_activation_routes_notify() RETURNS trigger AS $$ + DECLARE + tenant TEXT; + BEGIN + SELECT p.tenant_id INTO tenant FROM project p + WHERE p.id = CASE WHEN TG_OP = 'DELETE' THEN OLD.project_id ELSE NEW.project_id END; + -- A project already gone told its tenant itself. + IF tenant IS NOT NULL THEN + PERFORM pg_notify('weft_routes', tenant); + END IF; + RETURN NULL; + END; + $$ LANGUAGE plpgsql"#, + r#"DROP TRIGGER IF EXISTS trigger_activation_routes_on_status ON trigger_activation"#, + r#"CREATE TRIGGER trigger_activation_routes_on_status + AFTER UPDATE OF status ON trigger_activation + FOR EACH ROW + WHEN (NEW.status IS DISTINCT FROM OLD.status) + EXECUTE FUNCTION trigger_activation_routes_notify()"#, + r#"DROP TRIGGER IF EXISTS trigger_activation_routes_on_row ON trigger_activation"#, + r#"CREATE TRIGGER trigger_activation_routes_on_row + AFTER INSERT OR DELETE ON trigger_activation + FOR EACH ROW + EXECUTE FUNCTION trigger_activation_routes_notify()"#, r#"DROP TRIGGER IF EXISTS trigger_activation_held_on_row ON trigger_activation"#, r#"CREATE TRIGGER trigger_activation_held_on_row AFTER INSERT OR DELETE ON trigger_activation diff --git a/crates/weft-dispatcher/src/api/mod.rs b/crates/weft-dispatcher/src/api/mod.rs index d4032ad1..75397749 100644 --- a/crates/weft-dispatcher/src/api/mod.rs +++ b/crates/weft-dispatcher/src/api/mod.rs @@ -101,19 +101,36 @@ impl axum::extract::FromRequestParts for CallerAddress { /// Whether a path is one of the token doors, where a refused answer /// means somebody presented a token that does not work. fn is_token_door(path: &str) -> bool { - ["/signal/", "/signal-token/", "/connect/", "/instance/"].iter().any(|p| path.starts_with(p)) + ["/signal/", "/signal-token/", LIVE_CALL_DOOR, "/instance/"].iter().any(|p| path.starts_with(p)) } +/// Where live calls come in (`signal::connect_live`). +const LIVE_CALL_DOOR: &str = "/connect/"; + +/// Marks the answer to a live call that was admitted (its address checked +/// for guessing tokens with its birth): what it says is the program's, or +/// the socket address its run is reached at. +#[derive(Clone, Copy)] +pub(crate) struct Admitted; + /// The layer over the outside-caller surface that stops token guessing: /// an address past the install's bound of refused tokens this minute is /// answered 429 on every token door until the minute ends, and each /// refusal a token door gives (401, 403, or a 404 for an unknown /// `/signal/` token) counts toward it. +/// +/// A live call (`/connect/`) is not held up by a read of its own for this: +/// a call that is let through is checked by its admission, in the one +/// call to the database that admits and births its run (its answer is +/// then marked [`Admitted`]), and any other answer, whatever its status, +/// is checked here before it leaves. Either way a blocked address hears +/// 429 and nothing else, so a right guess cannot be told from a wrong one. async fn guard_token_doors( axum::extract::State(state): axum::extract::State, request: axum::extract::Request, next: axum::middleware::Next, ) -> axum::response::Response { + use axum::http::StatusCode; use axum::response::IntoResponse; let path = request.uri().path().to_string(); if !is_token_door(&path) || state.edge.invalid_tokens_per_minute.is_none() { @@ -125,17 +142,33 @@ async fn guard_token_doors( Err(e) => return e.into_response(), }; let now = crate::lease::now_unix(); - match crate::entry_limits::token_guessing_blocked(&state.pg_pool, &state.edge, address, now).await { - Ok(Some(refused)) => return crate::entry_limits::too_many(refused), - Ok(None) => {} - Err(e) => { - return (axum::http::StatusCode::INTERNAL_SERVER_ERROR, format!("rate limit: {e:#}")).into_response() + let blocked = || async { + crate::entry_limits::token_guessing_blocked(&state.pg_pool, &state.edge, address, now).await.map_err(|e| { + (StatusCode::INTERNAL_SERVER_ERROR, format!("rate limit: {e:#}")).into_response() + }) + }; + let live_call = path.starts_with(LIVE_CALL_DOOR); + if !live_call { + match blocked().await { + Ok(Some(refused)) => return crate::entry_limits::too_many(refused), + Ok(None) => {} + Err(answer) => return answer, } } let response = next.run(axum::extract::Request::from_parts(parts, body)).await; let status = response.status(); - let refused_token = matches!(status, axum::http::StatusCode::UNAUTHORIZED | axum::http::StatusCode::FORBIDDEN) - || (status == axum::http::StatusCode::NOT_FOUND && path.starts_with("/signal/")); + let admitted = response.extensions().get::().is_some(); + if live_call && !admitted && status != StatusCode::TOO_MANY_REQUESTS { + match blocked().await { + Ok(Some(refused)) => return crate::entry_limits::too_many(refused), + Ok(None) => {} + Err(answer) => return answer, + } + } + // What the program answered is not a refused token, whatever it says. + let refused_token = !admitted + && (matches!(status, StatusCode::UNAUTHORIZED | StatusCode::FORBIDDEN) + || (status == StatusCode::NOT_FOUND && path.starts_with("/signal/"))); if refused_token { if let Err(e) = crate::entry_limits::note_invalid_token(&state.pg_pool, &state.edge, address, now).await { // The answer is already decided; a lost count only lets one diff --git a/crates/weft-dispatcher/src/api/project.rs b/crates/weft-dispatcher/src/api/project.rs index ecf6cfb6..7703039b 100644 --- a/crates/weft-dispatcher/src/api/project.rs +++ b/crates/weft-dispatcher/src/api/project.rs @@ -796,11 +796,11 @@ impl InfraSetupRun { async fn start_queued_execution( state: &DispatcherState, mut birth: Birth<'_>, - expected_activation: Option, + for_activation: bool, ) -> Result<(), (StatusCode, String)> { let source_version = super::versions::record_program_source(state, birth.project_id, birth.program).await?; birth.source_version = Some(&source_version); - start_queued_execution_with(state, birth, &[], expected_activation, None).await + start_queued_execution_with(state, birth, &[], for_activation, None).await } /// THE one way a queued execution is born: its birth rows in one @@ -813,7 +813,10 @@ pub(crate) async fn start_queued_execution_with( state: &DispatcherState, birth: Birth<'_>, extra_rows: &[weft_journal::ExecEvent], - expected_activation: Option, + // A trigger setup an activation asked for (the activation is the + // setup's own execution): born only while that activation still owns + // its rows. + for_activation: bool, live_connection: Option, ) -> Result<(), (StatusCode, String)> { let tenant = state @@ -837,7 +840,7 @@ pub(crate) async fn start_queued_execution_with( kick_events.extend_from_slice(extra_rows); state .journal - .start_execution(&start, &kick_events, task, expected_activation) + .start_execution(&start, &kick_events, task, for_activation) .await .map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, format!("start execution: {e}")))?; Ok(()) @@ -977,32 +980,56 @@ pub(crate) async fn missing_infra_nodes( // way, so a run cut to one call waits on that call's instance alone. // A per-instance node is looked up in `instance`'s copy: that is the // one the run or the instance's triggers read. - let mut missing: Vec = Vec::new(); + // Which copy of which place (`None` for a per-instance place with no + // instance named, which has no copy to look at), then the project's + // copies as this dispatcher holds them (`crate::held`). + let mut wanted: Vec<(String, Option>)> = Vec::new(); for place in weft_core::project::infra_places(project) { let spelled = weft_core::project::address_of(project, &place.id, &place.path); if within.is_some_and(|set| !set.contains(&spelled)) { continue; } let per_instance = weft_core::project::is_per_instance(project, &place.id); - let copy = if per_instance { instance } else { None }; - if per_instance && copy.is_none() { - // Nobody named, so there is no copy to look at: activate's note - // over the whole project lands here (it leaves these out), and - // a run never does ([`require_run_infra`] refuses it first). - missing.push(MissingCopy { place: spelled, instance: None, needs_instance: true }); - continue; + // Nobody named, so there is no copy to look at: activate's note + // over the whole project lands here (it leaves these out), and a + // run never does ([`require_run_infra`] refuses it first). + let copy = if per_instance { instance.map(Some) } else { Some(None) }; + wanted.push((spelled, copy)); + } + let missing_of = |copies: &[crate::infra_node::CopyStatus]| -> Vec { + let mut missing = Vec::new(); + for (place, copy) in &wanted { + let Some(copy) = copy else { + missing.push(MissingCopy { place: place.clone(), instance: None, needs_instance: true }); + continue; + }; + let running = copies.iter().any(|row| { + &row.node_id == place && row.instance.as_ref() == *copy && row.status == crate::infra_node::InfraNodeStatus::Running + }); + if !running { + missing.push(MissingCopy { place: place.clone(), instance: copy.cloned(), needs_instance: false }); + } } - let row = crate::infra_node::get(&state.pg_pool, project_id, &spelled, copy) - .await - .map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, format!("infra_node: {e}")))?; - let running = row - .map(|r| r.status == crate::infra_node::InfraNodeStatus::Running) - .unwrap_or(false); - if !running { - missing.push(MissingCopy { place: spelled, instance: copy.cloned(), needs_instance: false }); + missing + }; + if wanted.iter().all(|(_, copy)| copy.is_none()) { + return Ok(missing_of(&[])); + } + // A copy that came up a moment ago may not have been heard yet: a + // refusal is made on the rows themselves. + if let Some(held) = state.held.infra_status.held(&project_id) { + let missing = missing_of(&held); + if missing.iter().all(|m| m.needs_instance) { + return Ok(missing); } } - Ok(missing) + let fresh = state + .held + .infra_status + .load_fresh(project_id, || crate::infra_node::statuses(&state.pg_pool, project_id)) + .await + .map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, format!("infra_node: {e:#}")))?; + Ok(missing_of(&fresh)) } /// One infra copy a run or a trigger needs and that is not running. @@ -1247,7 +1274,7 @@ pub async fn start_infra_setup( at_unix: crate::lease::now_unix() as u64, }, &[], - None, + false, // An infra setup answers nobody. None, ) @@ -1755,6 +1782,7 @@ pub async fn status( status: a.lifecycle.status, mode: a.lifecycle.mode(), waiting: waiting.get(&a.key).cloned(), + version: a.source_version.clone(), }) .collect(), limited: limited_entries(&state, id) @@ -4572,7 +4600,7 @@ async fn run_trigger_setup( run_class: weft_core::run_class::RunClass::Short, at_unix: crate::lease::now_unix() as u64, }, - activation, + activation.is_some(), ) .await?; diff --git a/crates/weft-dispatcher/src/api/signal.rs b/crates/weft-dispatcher/src/api/signal.rs index fef98e55..be0ad84b 100644 --- a/crates/weft-dispatcher/src/api/signal.rs +++ b/crates/weft-dispatcher/src/api/signal.rs @@ -1,6 +1,8 @@ //! Signal-related dispatcher routes. Every endpoint here either relays a //! fire to the listener or reads/writes the durable signal table. +use std::sync::Arc; + use anyhow::Context; use axum::{ extract::{Path, RawQuery, State}, @@ -382,7 +384,7 @@ async fn fire_signal_inner( }; if !routing.is_resume { if let Err(refused) = - check_entry_limits(state, token, routing.project_id, &routing.limits, &caller.key(), None).await + check_entry_limits(state, token, routing.project_id, &routing.limits, &caller.key()).await { return refused; } @@ -685,54 +687,37 @@ pub(crate) async fn signal_gate( /// The per-minute and at-once limits of one outside call to the entry /// `token`, checked before anything is started, so a refused call costs /// nothing. `caller` is who is calling, spelled as its key (the verified -/// identity, or the address). `take_slot_for` is the execution of the run -/// this call will start and how long its slot holds unborn, when the -/// slot is taken now (a live handshake); a fire takes its slot when its -/// run is born, so here it only checks the entry is not already full. -/// A resume token answers one run already going and is not an entry: -/// its callers never come here. +/// identity, or the address). A fire takes its slot when its run is born +/// (`route_entry`), so here the entry is only checked for room; a live +/// call is admitted with its birth instead (`connect_live`). A resume +/// token answers one run already going and is not an entry: its callers +/// never come here. pub(crate) async fn check_entry_limits( state: &DispatcherState, token: &str, project_id: uuid::Uuid, limits: &weft_core::signal::ResolvedLimits, caller: &str, - take_slot_for: Option<(&str, i64)>, ) -> Result<(), axum::response::Response> { - let now = crate::lease::now_unix(); - let internal = |e: anyhow::Error| { - use axum::response::IntoResponse; - (StatusCode::INTERNAL_SERVER_ERROR, format!("rate limit: {e:#}")).into_response() - }; - let mut refused = crate::entry_limits::admit_call(&state.pg_pool, token, caller, limits, now) - .await - .map_err(internal)? - .err(); - if refused.is_none() { - if let Some(max) = limits.at_once { - refused = match take_slot_for { - Some((execution_id, unborn_until)) => { - crate::entry_limits::take_slot(&state.pg_pool, token, execution_id, max, unborn_until, now) - .await - .map_err(internal)? - .err() - } - None => crate::entry_limits::at_once_full(&state.pg_pool, token, max, now) - .await - .map_err(internal)?, - }; + let admission = crate::entry_limits::Admission::call(&state.edge, None, token, caller, limits, None, crate::lease::now_unix()); + match crate::entry_limits::admit(&state.pg_pool, &admission).await { + Ok(Ok(())) => Ok(()), + Ok(Err(refused)) => Err(refuse_call(token, project_id, refused)), + Err(e) => { + use axum::response::IntoResponse; + Err((StatusCode::INTERNAL_SERVER_ERROR, format!("rate limit: {e:#}")).into_response()) } } - let Some(refused) = refused else { return Ok(()) }; +} + +/// The answer to a call an entry's limits refused, logged. +fn refuse_call(token: &str, project_id: uuid::Uuid, refused: crate::entry_limits::Refused) -> axum::response::Response { tracing::info!( target: "weft_dispatcher::signal", token = %token, project_id = %project_id, "public call refused: {} is reached", refused.reason.describe() ); - crate::entry_limits::note_refusal(&state.pg_pool, token, refused.reason, now) - .await - .map_err(internal)?; - Err(crate::entry_limits::too_many(refused)) + crate::entry_limits::too_many(refused) } @@ -1520,7 +1505,7 @@ pub async fn fire_public_entry( }; // A bare-path fire is always an entry (a resume token has no mount). if let Err(refused) = - check_entry_limits(&state, &token, routing.project_id, &routing.limits, &caller.key(), None).await + check_entry_limits(&state, &token, routing.project_id, &routing.limits, &caller.key()).await { return refused; } @@ -1559,19 +1544,14 @@ async fn public_entry_target( // address holding a capture could never be reached through here at // all, whatever was called. let (tenant, called) = split_tenant(&normalized).map_err(|_| refuse())?; - let rows: Vec = sqlx::query_as::<_, (String, Vec, String)>( - "SELECT s.mount_path, s.mount_methods, s.token \ - FROM signal s \ - WHERE s.tenant_id = $1 AND s.surface_kind = 'public_entry' \ - AND s.mount_path IS NOT NULL", - ) - .bind(tenant) - .fetch_all(&state.pg_pool) - .await - .map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, format!("mount lookup: {e}")))? - .into_iter() - .map(|(mount_path, mount_methods, token)| RouteRow { mount_path, mount_methods, token }) - .collect(); + // Matched against the held routes, and against the rows themselves + // before refusing: a route activated a moment ago may not have been + // heard yet. + let rows_of = |routes: &[HeldRoute]| -> Vec { routes.iter().map(|route| route.row.clone()).collect() }; + let rows = match state.held.routes.held(&tenant.to_string()).map(|held| rows_of(&held)) { + Some(held) if resolve_route(&held, tenant, "POST", called).is_ok() => held, + _ => rows_of(&fresh_tenant_routes(state, tenant).await?), + }; let (matched, params) = resolve_route(&rows, tenant, "POST", called).map_err(|_| refuse())?; // An address with a capture in it is a live route's shape, and a // live route is served (and gated, and answered) at `/connect`. @@ -1719,7 +1699,7 @@ struct CallerRequestParts<'a> { } /// One public-entry row of the tenant, as the route matcher sees it. -#[derive(Debug)] +#[derive(Debug, Clone)] pub(crate) struct RouteRow { pub mount_path: String, pub mount_methods: Vec, @@ -1840,45 +1820,11 @@ pub async fn connect_live( .ok_or((StatusCode::BAD_REQUEST, format!("{API_PROJECT_HEADER} is not a project id")))?, ), }; - // Every route of the tenant with how it is armed, in one read: the - // match is made here, and the matched one's arming is already in hand. // How long each step took, in the line that says the run was born: // what to read first when a call is slow. let began = std::time::Instant::now(); - let armed_rows = sqlx::query(&format!( - "SELECT s.token, s.mount_path, s.mount_methods, {ARMED_COLUMNS} \ - FROM signal s \ - LEFT JOIN project p ON p.id = s.project_id \ - {} \ - WHERE s.tenant_id = $1 AND s.surface_kind = 'public_entry' AND s.mount_path IS NOT NULL \ - AND ($2::uuid IS NULL OR s.project_id = $2)", - weft_broker_client::protocol::SIGNAL_ACTIVATION_JOIN, - )) - .bind(&tenant_segment) - .bind(only_project) - .fetch_all(&state.pg_pool) - .await - .map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, format!("route lookup: {e}")))?; - let rows = armed_rows - .iter() - .map(|r| { - Ok(RouteRow { - token: r.try_get("token").map_err(row_err)?, - mount_path: r.try_get("mount_path").map_err(row_err)?, - mount_methods: r.try_get("mount_methods").map_err(row_err)?, - }) - }) - .collect::, (StatusCode, String)>>()?; - let (matched, params) = resolve_route(&rows, &tenant_segment, &method_name, &path)?; - let token = matched.token.clone(); - let matched_row = rows - .iter() - .position(|row| row.token == token) - .map(|i| &armed_rows[i]) - .expect("the matched route is one of the rows it was matched among"); - - let route = armed_route_of(matched_row)?; - let ArmedRoute { project_id, node_id, protocol, live_config, auth_kind, auth_config, program, .. } = &route; + let (route, token, params) = live_route(&state, &tenant_segment, only_project, &method_name, &path).await?; + let ArmedRoute { project_id, node_id, protocol, live_config, auth_kind, auth_config, program, .. } = route.as_ref(); // Every header the caller sent, repeats included, less the one the // install's door added for itself: what the gate checks, and what the @@ -1937,12 +1883,12 @@ pub async fn connect_live( // Project must be Active to accept a live connection. route.require_active()?; - // The entry's limits, before anything is started. The caller is who - // the gate established when the route has auth, else the instance the - // run is for, else the address. The slot is taken now, for the execution - // this call's run will carry; the run is born just below, and the slot - // counts for as long as it runs (the ticket's life only matters if the - // birth fails and its slot could not be freed). + // The entry's limits are checked as the run is born, in the same + // call to the database: the caller is who the gate established when + // the route has auth, else the instance the run is for, else the + // address. The slot is taken for the execution this call's run + // carries, and counts for as long as it runs; the ticket's life only + // matters for a run nobody ever claims. let execution_id = uuid::Uuid::new_v4(); let caller_key = match (&caller, &instance) { (Some(identity), _) => format!("id:{identity}"), @@ -1951,21 +1897,16 @@ pub async fn connect_live( }; let issued_at = crate::lease::now_unix(); let expires_at = issued_at + live_token_ttl_secs(); - let limits = route.spec.limits.resolve(); - let limits_from = began.elapsed(); - if let Err(refused) = check_entry_limits( - &state, + let admission = crate::entry_limits::Admission::call( + &state.edge, + Some(address.0), &token, - *project_id, - &limits, &caller_key, + &route.spec.limits.resolve(), Some((&execution_id.to_string(), expires_at)), - ) - .await - { - return Ok(refused); - } - let limited = began.elapsed(); + issued_at, + ); + let admitting = began.elapsed(); // The caller's opening request, as the trigger reads it: what they // sent here, which is what reaches the worker. @@ -1978,19 +1919,10 @@ pub async fn connect_live( headers: headers_sent, caller, }; - if let Err(refused) = - birth_live_run(&state, &route, &opening, &tenant_segment, execution_id, expires_at, instance.as_ref()) - .await - { - // The slot was taken for this run; nothing will start it now. - if let Err(e) = crate::entry_limits::release_slot(&state.pg_pool, &execution_id.to_string()).await { - tracing::warn!( - target: "weft_dispatcher::signal", - execution_id = %execution_id, error = %format!("{e:#}"), - "could not free the slot of a run that was refused at birth; the reaper frees it once it expires" - ); - } - return Err(refused); + let admitted = birth_live_run(&state, &route, &opening, &tenant_segment, execution_id, expires_at, instance.as_ref(), &admission) + .await?; + if let Err(refused) = admitted { + return Ok(refuse_call(&token, *project_id, refused)); } let born = began.elapsed(); @@ -2024,8 +1956,7 @@ pub async fn connect_live( execution_id = %execution_id, node = %node_id, route_ms = ms(routed), gate_ms = ms(gated - routed), - limits_ms = ms(limited - limits_from), - birth_ms = ms(born - limited), + birth_ms = ms(born - admitting), "live call: run born" ); @@ -2038,20 +1969,25 @@ pub async fn connect_live( if *protocol == weft_core::signal::Protocol::Websocket && !is_socket_opening(headers) { let url = crate::live_relay::live_url(&live_door(headers, &state.public_base_url), *project_id, &raw_path, &raw_query, &routing); let body = serde_json::json!({ "url": url, "protocol": "websocket" }); - return Ok(Response::builder() + let mut answer = Response::builder() .status(StatusCode::OK) .header(axum::http::header::CONTENT_TYPE, "application/json") .body(axum::body::Body::from(body.to_string())) - .expect("json response builds")); + .expect("json response builds"); + answer.extensions_mut().insert(crate::api::Admitted); + return Ok(answer); } let request = axum::extract::Request::from_parts(parts, body); - let answer = crate::live_relay::to_worker(&state, &claims, &routing, &raw_path, &raw_query, request).await; + let mut answer = crate::live_relay::to_worker(&state, &claims, &routing, &raw_path, &raw_query, request).await; tracing::info!( target: "weft_dispatcher::signal", execution_id = %execution_id, status = %answer.status(), worker_ms = ms(began.elapsed() - born), "live call: the worker answered (its body may still be streaming)" ); + // The admission checked the caller's address; the token guard leaves + // the answer alone. + answer.extensions_mut().insert(crate::api::Admitted); Ok(answer) } @@ -2064,11 +2000,101 @@ pub(crate) fn is_socket_opening(headers: &HeaderMap) -> bool { .is_some_and(|v| v.eq_ignore_ascii_case("websocket")) } +/// One public entry of a tenant as held in memory (`crate::held`): what +/// the matcher needs, and how the row arms it (or why it is half-armed, +/// answered only to a call that matches it). +pub(crate) struct HeldRoute { + project_id: uuid::Uuid, + row: RouteRow, + armed: Result, (StatusCode, String)>, +} + +/// A matched live route: the route as it is armed, its token, and the +/// path's captures. +type LiveRoute = (Arc, String, std::collections::BTreeMap); + +/// The live route of `tenant` (of `only_project`, when the call came by a +/// project's API domain) serving `method` on `path`: the route as it is +/// armed, its token and the path's captures. Matched against the routes +/// this dispatcher holds; a refusal (no such route, a half-armed or +/// inactive one) is made on the rows themselves, since a route activated a +/// moment ago may not have been heard yet. +async fn live_route( + state: &DispatcherState, + tenant: &str, + only_project: Option, + method: &str, + path: &str, +) -> Result { + if let Some(held) = state.held.routes.held(&tenant.to_string()) { + if let Ok(found) = match_live_route(&held, tenant, only_project, method, path) { + if found.0.require_active().is_ok() { + return Ok(found); + } + } + } + match_live_route(&fresh_tenant_routes(state, tenant).await?, tenant, only_project, method, path) +} + +fn match_live_route( + routes: &[HeldRoute], + tenant: &str, + only_project: Option, + method: &str, + path: &str, +) -> Result { + let candidates: Vec<&HeldRoute> = + routes.iter().filter(|route| only_project.is_none_or(|only| route.project_id == only)).collect(); + let rows: Vec = candidates.iter().map(|route| route.row.clone()).collect(); + let (matched, params) = resolve_route(&rows, tenant, method, path)?; + let route = candidates + .iter() + .find(|route| route.row.token == matched.token) + .expect("the matched route is one of the rows it was matched among") + .armed + .clone()?; + Ok((route, matched.token.clone(), params)) +} + +/// Every public entry of `tenant` as the rows say now, kept for the next +/// call (`crate::held`). +async fn fresh_tenant_routes(state: &DispatcherState, tenant: &str) -> Result>, (StatusCode, String)> { + state.held.routes.load_fresh(tenant.to_string(), || read_tenant_routes(&state.pg_pool, tenant)).await +} + +async fn read_tenant_routes(pool: &sqlx::PgPool, tenant: &str) -> Result, (StatusCode, String)> { + let rows = sqlx::query(&format!( + "SELECT s.token, s.mount_path, s.mount_methods, {ARMED_COLUMNS} \ + FROM signal s \ + LEFT JOIN project p ON p.id = s.project_id \ + {} \ + WHERE s.tenant_id = $1 AND s.surface_kind = 'public_entry' AND s.mount_path IS NOT NULL", + weft_broker_client::protocol::SIGNAL_ACTIVATION_JOIN, + )) + .bind(tenant) + .fetch_all(pool) + .await + .map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, format!("route lookup: {e}")))?; + rows.iter() + .map(|r| { + Ok(HeldRoute { + project_id: r.try_get("project_id").map_err(row_err)?, + row: RouteRow { + token: r.try_get("token").map_err(row_err)?, + mount_path: r.try_get("mount_path").map_err(row_err)?, + mount_methods: r.try_get("mount_methods").map_err(row_err)?, + }, + armed: armed_route_of(r).map(Arc::new), + }) + }) + .collect() +} + /// A public entry as its signal row arms it: the trigger, its spec, the /// gate's settings, and the program identity its runs are born under. /// Read at the handshake, which gates the caller and gives birth to the /// run. -struct ArmedRoute { +pub(crate) struct ArmedRoute { project_id: uuid::Uuid, node_id: String, spec: weft_core::primitive::SignalSpec, @@ -2152,13 +2178,14 @@ fn armed_route_of(row: &sqlx::postgres::PgRow) -> Result, -) -> Result<(), (StatusCode, String)> { + admission: &crate::entry_limits::Admission, +) -> Result, (StatusCode, String)> { let project_id = route.project_id; let definition_hash = &route.program.definition_hash; let project_def = state @@ -2187,11 +2215,16 @@ async fn birth_live_run( // instance provides filled and valid, and finds that instance's infra up: // refused here, to the caller standing at the door, rather than // mid-run. The values read are the run's. - let instance_values = crate::api::project::refuse_instance_gaps(state, project_id, &project_def, &subgraph, instance) - .await - .map_err(<(StatusCode, String)>::from)?; - // The program's own connections, as this install picked them. - let picks = crate::api::project::picks_for_run(state, project_id, &project_def, &subgraph).await?; + // The program's own connections, as this install picked them. The two + // reads go out together: neither needs the other. + let (instance_values, picks) = tokio::try_join!( + async { + crate::api::project::refuse_instance_gaps(state, project_id, &project_def, &subgraph, instance) + .await + .map_err(<(StatusCode, String)>::from) + }, + crate::api::project::picks_for_run(state, project_id, &project_def, &subgraph), + )?; let now = crate::lease::now_unix() as u64; // The fire's computed subgraph rides on ExecutionStarted: the @@ -2247,7 +2280,7 @@ async fn birth_live_run( state .journal - .start_execution(&start, &kick_events, task, None) + .admit_and_start_execution(admission, &start, &kick_events, task) .await .map_err(|e| (StatusCode::INTERNAL_SERVER_ERROR, format!("live run birth: {e:#}"))) } diff --git a/crates/weft-dispatcher/src/api/versions.rs b/crates/weft-dispatcher/src/api/versions.rs index d9ade413..6e50cf57 100644 --- a/crates/weft-dispatcher/src/api/versions.rs +++ b/crates/weft-dispatcher/src/api/versions.rs @@ -767,7 +767,7 @@ pub async fn run( at_unix: now, }, &birth_rows, - None, + false, fired_caller.clone(), ) .await?; diff --git a/crates/weft-dispatcher/src/app.rs b/crates/weft-dispatcher/src/app.rs index 9f32a863..22c3ac1b 100644 --- a/crates/weft-dispatcher/src/app.rs +++ b/crates/weft-dispatcher/src/app.rs @@ -128,6 +128,9 @@ pub const DISPATCHER_CHANNELS: &[&str] = &[ crate::display_feeds::LOOK_NOW_CHANNEL, crate::holders::HELD_SIGNALS_CHANNEL, crate::domains::DOMAINS_CHANNEL, + crate::held::ROUTES_CHANNEL, + crate::held::WORKER_SETTINGS_CHANNEL, + crate::held::INFRA_STATUS_CHANNEL, ]; /// What the dispatcher is built from: the install's config and what the @@ -188,6 +191,7 @@ pub async fn build_state(settings: DispatcherSettings<'_>, defaults: Defaults) - let versions: crate::versions::VersionStore = Arc::new(crate::versions::PostgresVersionStore::new(pool.clone())); let event_bus = crate::EventBus::with_notify(pool.clone(), &signals)?; let displays = crate::display_feeds::DisplayFeeds::with_look_now(&signals)?; + let held = Arc::new(crate::held::Held::follow(&signals)?); // The other roles as this dispatcher reaches them, from where its // own placement puts it (a local install's loopback is only its own // process's). @@ -243,6 +247,7 @@ pub async fn build_state(settings: DispatcherSettings<'_>, defaults: Defaults) - http, caller_token_secret: Arc::new(caller_token_secret), programs: Arc::new(weft_core::content_cache::ContentCache::new(64)), + held, }) } diff --git a/crates/weft-dispatcher/src/delivery.rs b/crates/weft-dispatcher/src/delivery.rs index 160cbeda..71281127 100644 --- a/crates/weft-dispatcher/src/delivery.rs +++ b/crates/weft-dispatcher/src/delivery.rs @@ -97,16 +97,18 @@ pub async fn worker_target( project: uuid::Uuid, binary_hash: &str, ) -> anyhow::Result { - let overrides = state - .projects - .worker_overrides(project) - .await? - .ok_or_else(|| anyhow::anyhow!("project {project} is not registered"))?; + // A project registered a moment ago may not have been heard yet: its + // absence is read from the rows themselves. + let overrides = match state.held.worker_overrides.held(&project).filter(|held| held.is_some()) { + Some(held) => held, + None => state.held.worker_overrides.load_fresh(project, || state.projects.worker_overrides(project)).await?, + }; + let overrides = overrides.as_ref().as_ref().ok_or_else(|| anyhow::anyhow!("project {project} is not registered"))?; Ok(weft_platform_traits::WorkerTarget { tenant: tenant.to_string(), project, image: state.builder.images.image_ref(&weft_compiler::build::worker_image_tag(binary_hash)), - settings: state.worker_defaults.with(&overrides), + settings: state.worker_defaults.with(overrides), }) } diff --git a/crates/weft-dispatcher/src/door.rs b/crates/weft-dispatcher/src/door.rs index a71abe0d..67876ba8 100644 --- a/crates/weft-dispatcher/src/door.rs +++ b/crates/weft-dispatcher/src/door.rs @@ -9,13 +9,10 @@ //! its root; the install's own domains and any other name reach the //! public API itself. //! -//! The domains are read from the database when a request needs them and -//! the last read is older than [`READ_EVERY`], never on a timer: a door -//! that looked on its own would keep the database awake for an install -//! nobody is using. +//! The domains are held in memory and read again when one comes or goes +//! (`crate::held`), so routing a request costs no trip to the database. use std::sync::Arc; -use std::time::{Duration, Instant}; use axum::extract::{Request, State}; use axum::http::header::InvalidHeaderValue; @@ -27,10 +24,6 @@ use weft_core::install::{Domain, DomainServes}; use crate::state::DispatcherState; -/// How old the door's copy of the domains may be. A domain just added -/// works at most this long after. -const READ_EVERY: Duration = Duration::from_secs(30); - /// The scheme a request that came by a stored domain was reached over: /// the platform's door in front of the domains holds their certificates /// (`weft_platform_traits::DomainHosting`). @@ -106,64 +99,31 @@ pub fn api_path(tenant: &str, path_and_query: &str) -> String { format!("/connect/{tenant}/{}", path_and_query.trim_start_matches('/')) } -/// The door's copy of the stored domains (`None` until a read succeeds), -/// and when a read of them was last started. -#[derive(Default)] -struct Read { - at: Option, - domains: Option>, -} - #[derive(Clone)] struct Door { inner: Router, state: DispatcherState, - read: Arc>, } impl Door { - /// The stored domains. The first request to find the copy older than - /// [`READ_EVERY`] reads them again, and every request arriving - /// meanwhile routes by the last copy, so none waits on another's - /// query. A read that fails keeps the last copy until the next one is - /// due. With no copy yet, a request reads for itself and is refused - /// when that fails: routing it as if there were no domains would send + /// The stored domains, from the install's held copy. Refused when + /// they cannot be read: routing as if there were no domains would send /// a frontend's visitors to the install. - async fn domains(&self) -> Result, Response> { - let (copy, due) = { - let mut read = self.read.lock().await; - let due = read.at.is_none_or(|at| at.elapsed() >= READ_EVERY); - if due { - read.at = Some(Instant::now()); - } - (read.domains.clone(), due) - }; - if let (Some(copy), false) = (©, due) { - return Ok(copy.clone()); - } - match crate::domains::list(&self.state.pg_pool).await { - Ok(fresh) => { - let fresh: Arc<[Domain]> = fresh.into(); - self.read.lock().await.domains = Some(fresh.clone()); - Ok(fresh) - } - Err(e) => { - tracing::warn!(target: "weft_dispatcher::door", error = %format!("{e:#}"), "could not read the install's domains"); - copy.ok_or_else(|| { - ( - StatusCode::SERVICE_UNAVAILABLE, - format!("the install could not read its domains to route this request ({e:#}); try again shortly"), - ) - .into_response() - }) - } - } + async fn domains(&self) -> Result>, Response> { + self.state.held.domains.get_or_load((), || crate::domains::list(&self.state.pg_pool)).await.map_err(|e| { + tracing::warn!(target: "weft_dispatcher::door", error = %format!("{e:#}"), "could not read the install's domains"); + ( + StatusCode::SERVICE_UNAVAILABLE, + format!("the install could not read its domains to route this request ({e:#}); try again shortly"), + ) + .into_response() + }) } } /// `inner` (the public API) behind the door that routes by name. pub fn router(state: DispatcherState, inner: Router) -> Router { - Router::new().fallback(route).with_state(Door { inner, state, read: Arc::default() }) + Router::new().fallback(route).with_state(Door { inner, state }) } async fn route(State(door): State, mut request: Request) -> Response { diff --git a/crates/weft-dispatcher/src/entry_limits.rs b/crates/weft-dispatcher/src/entry_limits.rs index 2af4431e..7616fa60 100644 --- a/crates/weft-dispatcher/src/entry_limits.rs +++ b/crates/weft-dispatcher/src/entry_limits.rs @@ -45,13 +45,109 @@ pub static GROUP: weft_task_store::SchemaGroup = weft_task_store::SchemaGroup { // toward the entry's at-once limit: taken at admission (keyed by // the execution the run will have), dropped when the run ends. A slot // whose run never started stops counting at `unborn_until`, and - // one whose run ended stops counting at once (`slot_stopped_counting`). + // one whose run ended stops counting at once (`weft_slot_stopped_counting`). r#"CREATE TABLE IF NOT EXISTS entry_slot ( execution_id TEXT PRIMARY KEY, signal_token TEXT NOT NULL, unborn_until BIGINT NOT NULL )"#, r#"CREATE INDEX IF NOT EXISTS idx_entry_slot_token ON entry_slot(signal_token)"#, + // Count one hit on a key in a minute, and answer its hits so far, + // this one included. The increment and the read are one statement, + // so two replicas counting at once each see the other's hit. + r#"CREATE OR REPLACE FUNCTION weft_rate_hit(p_key TEXT, p_window_start BIGINT) RETURNS INTEGER AS $$ + INSERT INTO entry_rate (key, window_start, hits) VALUES (p_key, p_window_start, 1) + ON CONFLICT (key, window_start) DO UPDATE SET hits = entry_rate.hits + 1 + RETURNING hits + $$ LANGUAGE sql"#, + // Whether a slot no longer counts toward its entry's at-once limit + // at `p_now`: its run never started (or was an unrecorded run, + // forgotten with its row) and its hold passed, or its run ended, + // whether or not the bridge freed the slot yet: a terminal event is + // in the journal, or an unrecorded run whose costs kept its row is + // stamped ended. + concat!( + r#"CREATE OR REPLACE FUNCTION weft_slot_stopped_counting(p_execution_id TEXT, p_unborn_until BIGINT, p_now BIGINT) + RETURNS BOOLEAN AS $$ + SELECT (p_unborn_until < p_now + AND NOT EXISTS (SELECT 1 FROM execution ec WHERE ec.execution_id = p_execution_id)) + OR EXISTS (SELECT 1 FROM execution ec WHERE ec.execution_id = p_execution_id + AND ec.ended_at_unix IS NOT NULL) + OR EXISTS (SELECT 1 FROM exec_event e WHERE e.execution_id = p_execution_id + AND e.kind IN "#, + weft_journal::execution_terminal_kinds_sql!(), + r#") + $$ LANGUAGE sql STABLE"# + ), + // One call's (or one run's) admission at the entry's limits, in the + // caller's transaction: `NULL` when admitted, else which limit + // refused it and when to try again. In order: an address blocked + // for guessing tokens (read only), then each per-minute count (the + // caller's before the entry's, so a caller at its limit never spends + // the entry's shared allowance), then the at-once limit. A slot is + // taken for `slot.execution_id` when one is named (kept, never + // counted twice, when that execution already holds it), else the + // entry is only checked for room. Taking is serialized per entry by + // a transaction-scoped lock, and the count runs after it, so it sees + // every slot a sibling took before letting go; two replicas + // admitting the last free slot at once cannot both take it. Every + // refusal but a blocked address is counted for `weft status`. + // SYNC: p's fields <-> Admission (below) + r#"CREATE OR REPLACE FUNCTION weft_admit(p JSONB) RETURNS JSONB AS $$ + DECLARE + v_window BIGINT := (p->>'window_start')::bigint; + v_later INTEGER := (p->>'retry_after_secs')::integer; + v_now BIGINT := (p->>'now')::bigint; + v_slot JSONB := p->'slot'; + v_count JSONB; + v_hits INTEGER; + v_refused JSONB; + v_room BOOLEAN; + BEGIN + IF jsonb_typeof(p->'blocked') = 'object' THEN + SELECT r.hits INTO v_hits FROM entry_rate r + WHERE r.key = p->'blocked'->>'key' AND r.window_start = v_window; + IF v_hits >= (p->'blocked'->>'limit')::integer THEN + RETURN jsonb_build_object('reason', 'invalid_tokens', 'retry_after_secs', v_later); + END IF; + END IF; + FOR v_count IN SELECT c FROM jsonb_array_elements(p->'counts') WITH ORDINALITY AS e(c, n) ORDER BY n LOOP + IF weft_rate_hit(v_count->>'key', v_window) > (v_count->>'limit')::integer THEN + v_refused := jsonb_build_object('reason', v_count->>'reason', 'retry_after_secs', v_later); + EXIT; + END IF; + END LOOP; + IF v_refused IS NULL AND jsonb_typeof(v_slot) = 'object' THEN + IF v_slot->>'execution_id' IS NULL THEN + SELECT COUNT(*) < (v_slot->>'max')::bigint INTO v_room FROM entry_slot s + WHERE s.signal_token = v_slot->>'token' + AND NOT weft_slot_stopped_counting(s.execution_id, s.unborn_until, v_now); + ELSE + PERFORM pg_advisory_xact_lock(hashtextextended('entry_slot:' || (v_slot->>'token'), 0)); + SELECT EXISTS (SELECT 1 FROM entry_slot s WHERE s.execution_id = v_slot->>'execution_id' + AND NOT weft_slot_stopped_counting(s.execution_id, s.unborn_until, v_now)) + OR (SELECT COUNT(*) FROM entry_slot s + WHERE s.signal_token = v_slot->>'token' AND s.execution_id <> v_slot->>'execution_id' + AND NOT weft_slot_stopped_counting(s.execution_id, s.unborn_until, v_now)) + < (v_slot->>'max')::bigint + INTO v_room; + IF v_room THEN + INSERT INTO entry_slot (execution_id, signal_token, unborn_until) + VALUES (v_slot->>'execution_id', v_slot->>'token', (v_slot->>'unborn_until')::bigint) + ON CONFLICT (execution_id) DO UPDATE + SET unborn_until = GREATEST(entry_slot.unborn_until, EXCLUDED.unborn_until); + END IF; + END IF; + IF NOT v_room THEN + v_refused := jsonb_build_object('reason', 'at_once', 'retry_after_secs', 5); + END IF; + END IF; + IF v_refused IS NOT NULL THEN + PERFORM weft_rate_hit((p->>'refusals_key') || (v_refused->>'reason'), v_window); + END IF; + RETURN v_refused; + END; + $$ LANGUAGE plpgsql"#, ], seed: &[], }; @@ -103,6 +199,7 @@ impl Limited { } } + // SYNC: the slugs <-> weft_admit's reasons ('at_once', 'invalid_tokens') fn slug(self) -> &'static str { match self { Limited::PerCaller => "per_caller", @@ -111,6 +208,10 @@ impl Limited { Limited::InvalidTokens => "invalid_tokens", } } + + fn from_slug(slug: &str) -> Option { + [Limited::PerCaller, Limited::PerEntry, Limited::AtOnce, Limited::InvalidTokens].into_iter().find(|l| l.slug() == slug) + } } /// The minute `now` falls in, and the seconds left in it. @@ -119,22 +220,16 @@ fn window(now: i64) -> (i64, u64) { (start, (start + WINDOW_SECS - now).max(1) as u64) } -/// Count one hit on `key` in the current minute: its hits so far, this -/// one included, and the seconds left in the minute. The increment and -/// the read are one statement, so two replicas counting at once each see -/// the other's hit. +/// Count one hit on `key` in the current minute (`weft_rate_hit`): its +/// hits so far, this one included, and the seconds left in the minute. async fn count<'e>(db: impl sqlx::PgExecutor<'e>, key: &str, now: i64) -> Result<(i64, u64)> { let (start, left) = window(now); - let (hits,): (i32,) = sqlx::query_as( - "INSERT INTO entry_rate (key, window_start, hits) VALUES ($1, $2, 1) \ - ON CONFLICT (key, window_start) DO UPDATE SET hits = entry_rate.hits + 1 \ - RETURNING hits", - ) - .bind(key) - .bind(start) - .fetch_one(db) - .await - .context("count a public call")?; + let hits: i32 = sqlx::query_scalar("SELECT weft_rate_hit($1, $2)") + .bind(key) + .bind(start) + .fetch_one(db) + .await + .context("count a public call")?; Ok((hits as i64, left)) } @@ -144,23 +239,148 @@ async fn hit<'e>(db: impl sqlx::PgExecutor<'e>, key: &str, limit: u32, reason: L Ok(if hits > limit as i64 { Err(Refused { retry_after_secs: left, reason }) } else { Ok(()) }) } -/// The per-minute limits of one call to the entry `token` by `caller` -/// (an identity or an address, already spelled as the caller's key). -/// Checked per caller first, so one caller at its limit never spends the -/// entry's shared allowance. -pub async fn admit_call( - pool: &PgPool, - token: &str, - caller: &str, - limits: &weft_core::signal::ResolvedLimits, +/// One admission at an entry's limits, as `weft_admit` takes it: what +/// is checked, counted and taken, in order. Built by [`Admission::call`] +/// for a caller's call and [`Admission::fire`] for a run an event +/// started, and handed to the database whole: on its own ([`admit`]), or +/// with the run's birth in the same call (`Journal::start_execution`). +// SYNC: Admission's fields <-> weft_admit (GROUP above) +#[derive(Debug, Clone, serde::Serialize)] +pub struct Admission { now: i64, -) -> Result> { - if let Some(limit) = limits.per_caller_per_minute { - if let Err(refused) = hit(pool, &format!("c:{token}:{caller}"), limit, Limited::PerCaller, now).await? { - return Ok(Err(refused)); + window_start: i64, + retry_after_secs: u64, + /// The address checked for guessing tokens, when the install bounds it. + blocked: Option, + /// The per-minute counts, in the order they are checked. + counts: Vec, + /// The at-once limit, when the entry has one. + slot: Option, + /// Where a refusal is counted: the key, less the limit's slug. + refusals_key: String, +} + +#[derive(Debug, Clone, serde::Serialize)] +struct Bound { + key: String, + limit: u32, +} + +#[derive(Debug, Clone, serde::Serialize)] +struct Count { + key: String, + limit: u32, + reason: &'static str, +} + +#[derive(Debug, Clone, serde::Serialize)] +struct Slot { + token: String, + max: u32, + /// The run the slot is taken for; `None` only checks there is room. + execution_id: Option, + unborn_until: i64, +} + +/// What a call's slot is for: the run born with it (and how long its slot +/// holds if the run never starts), or nothing, when only room is checked. +pub type TakeFor<'a> = Option<(&'a str, i64)>; + +impl Admission { + fn at(token: &str, now: i64) -> Self { + let (window_start, retry_after_secs) = window(now); + Self { + now, + window_start, + retry_after_secs, + blocked: None, + counts: Vec::new(), + slot: None, + refusals_key: format!("refused:{token}:"), + } + } + + fn with_slot(mut self, token: &str, limits: &weft_core::signal::ResolvedLimits, take_for: TakeFor<'_>) -> Self { + self.slot = limits.at_once.map(|max| Slot { + token: token.to_string(), + max, + execution_id: take_for.map(|(execution_id, _)| execution_id.to_string()), + unborn_until: take_for.map(|(_, until)| until).unwrap_or(self.now), + }); + self + } + + /// A call to the entry `token` by `caller` (an identity or an address, + /// already spelled as the caller's key): the caller's count and the + /// entry's, then its room, taking a slot for `take_for`'s run when one + /// is named. `address` is checked for guessing tokens first, for a + /// door whose token guard leaves that check to the call's admission + /// (`api::guard_token_doors`). + pub fn call( + edge: &EdgeConfig, + address: Option, + token: &str, + caller: &str, + limits: &weft_core::signal::ResolvedLimits, + take_for: TakeFor<'_>, + now: i64, + ) -> Self { + let mut admission = Self::at(token, now); + admission.blocked = address + .zip(edge.invalid_tokens_per_minute) + .map(|(address, limit)| Bound { key: invalid_tokens_key(address), limit }); + if let Some(limit) = limits.per_caller_per_minute { + admission.counts.push(Count { key: format!("c:{token}:{caller}"), limit, reason: Limited::PerCaller.slug() }); + } + if let Some(limit) = limits.per_minute { + admission.counts.push(Count { key: format!("e:{token}"), limit, reason: Limited::PerEntry.slug() }); } + admission.with_slot(token, limits, take_for) + } + + /// The at-once limit of the run `execution_id` an event started at the + /// entry `token` (its per-minute count was taken when the event came + /// in, [`admit_fire`]), whose slot holds [`UNBORN_FIRE_SLOT_SECS`] if + /// the run is never born. + pub fn fire(token: &str, limits: &weft_core::signal::ResolvedLimits, execution_id: &str, now: i64) -> Self { + Self::at(token, now).with_slot(token, limits, Some((execution_id, now + UNBORN_FIRE_SLOT_SECS))) + } + + /// Whether there is nothing to check, count or take. + pub fn is_empty(&self) -> bool { + self.blocked.is_none() && self.counts.is_empty() && self.slot.is_none() + } +} + +/// `weft_admit`'s answer: `None` admitted, else the refusal. +#[derive(Debug, serde::Deserialize)] +pub struct Answer { + reason: String, + retry_after_secs: u64, +} + +impl Answer { + pub fn refused(&self) -> Result { + let reason = Limited::from_slug(&self.reason) + .ok_or_else(|| anyhow::anyhow!("the database refused a call for a limit weft does not know: '{}'", self.reason))?; + Ok(Refused { retry_after_secs: self.retry_after_secs, reason }) + } +} + +/// Admit `admission` on its own: one round trip, nothing born with it. +pub async fn admit(pool: &PgPool, admission: &Admission) -> Result> { + if admission.is_empty() { + return Ok(Ok(())); + } + let answer: Option = sqlx::query_scalar("SELECT weft_admit($1)") + .bind(serde_json::to_value(admission)?) + .fetch_one(pool) + .await + .context("admit a public call")?; + match answer { + None => Ok(Ok(())), + Some(answer) => Ok(Err(serde_json::from_value::(answer).context("read the admission's answer")?.refused()?)), } - per_entry(pool, token, limits, now).await } /// The entry's own per-minute limit: one count per run it asks for, @@ -242,13 +462,18 @@ pub async fn admit_fire( Ok(admitted) } +/// The count of refused tokens from `address`. +fn invalid_tokens_key(address: IpAddr) -> String { + format!("bad:{address}") +} + /// Whether `address` is blocked for presenting too many refused tokens /// this minute. Read-only: [`note_invalid_token`] counts. pub async fn token_guessing_blocked(pool: &PgPool, edge: &EdgeConfig, address: IpAddr, now: i64) -> Result> { let Some(limit) = edge.invalid_tokens_per_minute else { return Ok(None) }; let (start, left) = window(now); let hits: Option<(i32,)> = sqlx::query_as("SELECT hits FROM entry_rate WHERE key = $1 AND window_start = $2") - .bind(format!("bad:{address}")) + .bind(invalid_tokens_key(address)) .bind(start) .fetch_optional(pool) .await @@ -263,7 +488,7 @@ pub async fn note_invalid_token(pool: &PgPool, edge: &EdgeConfig, address: IpAdd if edge.invalid_tokens_per_minute.is_none() { return Ok(()); } - count(pool, &format!("bad:{address}"), now).await?; + count(pool, &invalid_tokens_key(address), now).await?; Ok(()) } @@ -294,92 +519,6 @@ pub async fn recent_refusals(pool: &PgPool, token: &str, now: i64) -> Result String { - format!( - "((s.unborn_until < {now_param} \ - AND NOT EXISTS (SELECT 1 FROM execution ec WHERE ec.execution_id = s.execution_id)) \ - OR EXISTS (SELECT 1 FROM execution ec WHERE ec.execution_id = s.execution_id \ - AND ec.ended_at_unix IS NOT NULL) \ - OR EXISTS (SELECT 1 FROM exec_event e WHERE e.execution_id = s.execution_id \ - AND e.kind IN {terminal}))", - terminal = weft_journal::EXECUTION_TERMINAL_KINDS_SQL, - ) -} - -/// Take one of `token`'s at-once slots for the run `execution_id` will be, or -/// say the entry is full. A slot already held by `execution_id` (a retried -/// admission of the same run) is kept, never counted twice. -/// -/// Serialized per entry by a transaction-scoped advisory lock, so two -/// replicas admitting the last free slot at once cannot both take it. The -/// count and the take are then one statement, run after the lock so it -/// sees every slot a sibling took before letting go. Slots that stopped -/// counting (a handshake nobody followed, a run that ended) are left out -/// of the count; the reaper's [`sweep`] deletes them. -pub async fn take_slot( - pool: &PgPool, - token: &str, - execution_id: &str, - max: u32, - unborn_until: i64, - now: i64, -) -> Result> { - let mut tx = pool.begin().await.context("begin slot take")?; - sqlx::query("SELECT pg_advisory_xact_lock(hashtextextended('entry_slot:' || $1, 0))") - .bind(token) - .execute(&mut *tx) - .await - .context("lock the entry's slots")?; - let (taken,): (bool,) = sqlx::query_as(&format!( - "WITH held AS ( \ - SELECT 1 FROM entry_slot s WHERE s.execution_id = $2 AND NOT {stopped}), \ - others AS ( \ - SELECT COUNT(*) AS n FROM entry_slot s \ - WHERE s.signal_token = $1 AND s.execution_id <> $2 AND NOT {stopped}), \ - taken AS ( \ - INSERT INTO entry_slot (execution_id, signal_token, unborn_until) \ - SELECT $2, $1, $3 \ - WHERE EXISTS (SELECT 1 FROM held) OR (SELECT n FROM others) < $4 \ - ON CONFLICT (execution_id) DO UPDATE \ - SET unborn_until = GREATEST(entry_slot.unborn_until, EXCLUDED.unborn_until) \ - RETURNING 1) \ - SELECT EXISTS (SELECT 1 FROM taken)", - stopped = slot_stopped_counting("$5"), - )) - .bind(token) - .bind(execution_id) - .bind(unborn_until) - .bind(max as i64) - .bind(now) - .fetch_one(&mut *tx) - .await - .context("take the slot")?; - tx.commit().await.context("commit slot take")?; - Ok(if taken { Ok(()) } else { Err(Refused { retry_after_secs: 5, reason: Limited::AtOnce }) }) -} - -/// Whether the entry is already full, without taking anything: the -/// quick answer at the edge for a fire, whose slot is taken when its run -/// is born. -pub async fn at_once_full(pool: &PgPool, token: &str, max: u32, now: i64) -> Result> { - let (taken,): (i64,) = sqlx::query_as(&format!( - "SELECT COUNT(*) FROM entry_slot s WHERE s.signal_token = $1 AND NOT {}", - slot_stopped_counting("$2") - )) - .bind(token) - .bind(now) - .fetch_one(pool) - .await - .context("count the entry's runs")?; - Ok((taken >= max as i64).then_some(Refused { retry_after_secs: 5, reason: Limited::AtOnce })) -} - /// The run `execution_id` ended: its slot, if it held one, is free. pub async fn release_slot(pool: &PgPool, execution_id: &str) -> Result<()> { sqlx::query("DELETE FROM entry_slot WHERE execution_id = $1") @@ -401,7 +540,7 @@ pub async fn sweep(pool: &PgPool, now: i64) -> Result<()> { .execute(pool) .await .context("drop old rate windows")?; - sqlx::query(&format!("DELETE FROM entry_slot s WHERE {}", slot_stopped_counting("$1"))) + sqlx::query("DELETE FROM entry_slot s WHERE weft_slot_stopped_counting(s.execution_id, s.unborn_until, $1)") .bind(now) .execute(pool) .await diff --git a/crates/weft-dispatcher/src/frontends.rs b/crates/weft-dispatcher/src/frontends.rs index ebddde4d..34c0ed1e 100644 --- a/crates/weft-dispatcher/src/frontends.rs +++ b/crates/weft-dispatcher/src/frontends.rs @@ -19,7 +19,7 @@ use axum::extract::{Path, Query, State}; use axum::http::StatusCode; use axum::Json; use sqlx::PgPool; -use weft_core::frontend::{AddFrontendRequest, Frontend, FrontendHost, FrontendWithToken, Repository}; +use weft_core::frontend::{AddFrontendRequest, AddedFrontend, Frontend, FrontendHost, FrontendWithToken, Repository}; use weft_core::signal_token::MintTokenRequest; use weft_platform_traits::FrontendSite; @@ -45,8 +45,9 @@ pub static GROUP: weft_task_store::SchemaGroup = weft_task_store::SchemaGroup { service TEXT, url TEXT, -- The caller token it calls with (`signal_token.id`), and the - -- ones made to replace it and not put in place yet. - token_id UUID NOT NULL, + -- ones made to replace it and not put in place yet. NULL for a + -- hosted frontend until its deploy puts its first in place. + token_id UUID, pending_token_ids UUID[] NOT NULL DEFAULT '{}', added_unix BIGINT NOT NULL, PRIMARY KEY (project_id, name), @@ -65,7 +66,7 @@ struct Row { repo_id: Option, service: Option, url: Option, - token_id: uuid::Uuid, + token_id: Option, pending_token_ids: Vec, } @@ -215,16 +216,19 @@ async fn revoke_token(state: &DispatcherState, owner: &TenantId, id: uuid::Uuid) Ok(()) } -/// `POST /projects/{id}/frontends`: the frontend, its token (shown this -/// once) and, for one the install hosts, its service. A step that fails -/// takes back what the steps before it made: the service and the -/// repository's access (unless another frontend uses it), then the token. +/// `POST /projects/{id}/frontends`: the frontend and, for one the install +/// hosts, its service; for one running elsewhere, its token (shown this +/// once). A hosted one gets no token here: its deploy workflow puts its +/// first in place (`weft target export` makes it), and one made now would +/// be a working credential nobody holds. A step that fails takes back what +/// the steps before it made: the service and the repository's access +/// (unless another frontend uses it), then the token. pub async fn add( State(state): State, caller: CallerTenant, Path(project): Path, Json(req): Json, -) -> Result, (StatusCode, String)> { +) -> Result, (StatusCode, String)> { authorize_project(&state, &caller.0, project).await?; req.validate().map_err(|why| (StatusCode::BAD_REQUEST, why))?; let taken = || { @@ -237,7 +241,10 @@ pub async fn add( return Err(taken()); } let owner = &caller.0; - let minted = mint_token(&state, owner, project, &req.name).await?; + let minted = match req.host { + FrontendHost::CloudRun => None, + FrontendHost::External => Some(mint_token(&state, owner, project, &req.name).await?), + }; let mut frontend = Frontend { name: req.name.clone(), project, @@ -245,7 +252,7 @@ pub async fn add( repo: req.repo.clone(), service: None, url: None, - token_id: minted.id, + token_id: minted.as_ref().map(|minted| minted.id), pending_token_ids: Vec::new(), }; let site = site(&frontend); @@ -272,7 +279,7 @@ pub async fn add( } .await; let failure = match made { - Ok(true) => return Ok(Json(FrontendWithToken { frontend, token: minted.token, token_id: minted.id })), + Ok(true) => return Ok(Json(AddedFrontend { frontend, token: minted.map(|minted| minted.token) })), Ok(false) => taken(), Err(e) => (StatusCode::BAD_GATEWAY, format!("add frontend '{}': {e:#}", req.name)), }; @@ -286,8 +293,10 @@ pub async fn add( )); } } - if let Err(undo) = revoke_token(&state, owner, minted.id).await { - msg.push_str(&format!("\n(and its token {} could not be revoked: {undo:#}; `weft token revoke {}` does it)", minted.id, minted.id)); + if let Some(minted) = &minted { + if let Err(undo) = revoke_token(&state, owner, minted.id).await { + msg.push_str(&format!("\n(and its token {} could not be revoked: {undo:#}; `weft token revoke {}` does it)", minted.id, minted.id)); + } } Err((status, msg)) } @@ -402,7 +411,7 @@ pub async fn token_done( authorize_project(&state, &caller.0, project).await?; let mut tx = state.pg_pool.begin().await.map_err(internal)?; lock_frontend(&mut tx, project, &name).await.map_err(internal)?; - let row: Option<(uuid::Uuid, Vec)> = sqlx::query_as( + let row: Option<(Option, Vec)> = sqlx::query_as( "SELECT token_id, pending_token_ids FROM project_frontend WHERE project_id = $1 AND name = $2", ) .bind(project) @@ -411,7 +420,7 @@ pub async fn token_done( .await .map_err(internal)?; let Some((current, pending)) = row else { return Err(not_found(&name)) }; - if current == token { + if current == Some(token) { return Ok(StatusCode::NO_CONTENT); } if !pending.contains(&token) { @@ -424,7 +433,7 @@ pub async fn token_done( .execute(&mut *tx) .await .map_err(internal)?; - let retired: Vec = std::iter::once(current).chain(pending).filter(|t| *t != token).collect(); + let retired: Vec = current.into_iter().chain(pending).filter(|t| *t != token).collect(); revoke_all(&state, &caller.0, &retired).await.map_err(internal)?; tx.commit().await.map_err(internal)?; Ok(StatusCode::NO_CONTENT) @@ -531,7 +540,7 @@ async fn take_down( )); } } - let tokens: Vec = std::iter::once(frontend.token_id).chain(frontend.pending_token_ids).collect(); + let tokens: Vec = frontend.token_id.into_iter().chain(frontend.pending_token_ids).collect(); revoke_all(state, owner, &tokens).await?; forget(&mut *tx, project, name).await?; tx.commit().await?; diff --git a/crates/weft-dispatcher/src/held.rs b/crates/weft-dispatcher/src/held.rs new file mode 100644 index 00000000..06ec68bf --- /dev/null +++ b/crates/weft-dispatcher/src/held.rs @@ -0,0 +1,68 @@ +//! The rows a live call reads on every request, held in this process's +//! memory and read again the moment the database says they changed +//! (`weft_task_store::held_copy`): a tenant's routes, the install's +//! domains, a project's worker levers, which of its infra copies are up. +//! With these held, a call reaches the database once, to be admitted and +//! born. An answer that refuses on what a copy says reads the rows again +//! first (`HeldCopy::load_fresh`): a change can be committed a few +//! milliseconds before it is heard. + +use std::sync::Arc; + +use weft_core::install::Domain; +use weft_task_store::held_copy::{Changed, HeldCopy}; +use weft_task_store::pg_signal::PgSignalWatch; + +/// Announced with a tenant's id when one of its routes changes (the +/// signal, project and activation groups' triggers). +// SYNC: ROUTES_CHANNEL <-> crates/weft-dispatcher/src/journal/postgres.rs (routes_notify_tenant), crates/weft-dispatcher/src/activation_store.rs (trigger_activation_routes_notify) +pub const ROUTES_CHANNEL: &str = "weft_routes"; + +/// Announced with a project's id when its worker levers change or it +/// goes (the project group's trigger). +// SYNC: WORKER_SETTINGS_CHANNEL <-> crates/weft-dispatcher/src/project_store.rs (project_worker_settings_notify) +pub const WORKER_SETTINGS_CHANNEL: &str = "weft_worker_settings"; + +/// Announced with a project's id when one of its infra copies comes, goes +/// or changes status (the infra_node group's trigger). +// SYNC: INFRA_STATUS_CHANNEL <-> crates/weft-dispatcher/src/infra_node.rs (infra_node_status_notify) +pub const INFRA_STATUS_CHANNEL: &str = "weft_infra_status"; + +/// How many tenants' routes and projects' levers a process keeps; past +/// that the least recently called go and are read again when called. +const KEPT: usize = 4096; + +/// See the module doc. +pub struct Held { + /// Every public entry of a tenant, by tenant + /// (`crate::api::signal::read_tenant_routes`). + pub(crate) routes: Arc>>, + /// The install's domains. + pub(crate) domains: Arc>>, + /// A project's worker levers, `None` for a project that is not + /// registered. + pub(crate) worker_overrides: Arc>>, + /// Every infra copy of a project and its status. + pub(crate) infra_status: Arc>>, +} + +impl Held { + pub fn follow(signals: &PgSignalWatch) -> anyhow::Result { + Ok(Self { + // A tenant with no route is not kept: a made-up tenant name + // (a scanner's) would otherwise push real ones out. + routes: HeldCopy::follow(signals, ROUTES_CHANNEL, KEPT, |tenant| Changed::Key(tenant.to_string()), |routes: &Vec| !routes.is_empty())?, + domains: HeldCopy::follow(signals, crate::domains::DOMAINS_CHANNEL, 1, |_| Changed::Everything, |_| true)?, + worker_overrides: HeldCopy::follow(signals, WORKER_SETTINGS_CHANNEL, KEPT, project_of, |_| true)?, + infra_status: HeldCopy::follow(signals, INFRA_STATUS_CHANNEL, KEPT, project_of, |_| true)?, + }) + } +} + +/// The project a notification names, or everything when it names none. +fn project_of(payload: &str) -> Changed { + match payload.parse() { + Ok(project) => Changed::Key(project), + Err(_) => Changed::Everything, + } +} diff --git a/crates/weft-dispatcher/src/infra_node.rs b/crates/weft-dispatcher/src/infra_node.rs index 5648cc07..7d0ca26d 100644 --- a/crates/weft-dispatcher/src/infra_node.rs +++ b/crates/weft-dispatcher/src/infra_node.rs @@ -134,6 +134,31 @@ pub static GROUP: weft_task_store::SchemaGroup = weft_task_store::SchemaGroup { r#"CREATE UNIQUE INDEX IF NOT EXISTS idx_infra_node_copy ON infra_node(project_id, node_id, instance_id) NULLS NOT DISTINCT"#, r#"CREATE INDEX IF NOT EXISTS idx_infra_node_project ON infra_node(project_id)"#, + // Tell every dispatcher a project's copies came, went or changed + // status, so its copy of which are up (`crate::held::Held`) is + // read again. + // SYNC: 'weft_infra_status' <-> crate::held::INFRA_STATUS_CHANNEL + r#"CREATE OR REPLACE FUNCTION infra_node_status_notify() RETURNS trigger AS $$ + BEGIN + IF TG_OP = 'DELETE' THEN + PERFORM pg_notify('weft_infra_status', OLD.project_id::text); + ELSE + PERFORM pg_notify('weft_infra_status', NEW.project_id::text); + END IF; + RETURN NULL; + END; + $$ LANGUAGE plpgsql"#, + r#"DROP TRIGGER IF EXISTS infra_node_status_on_row ON infra_node"#, + r#"CREATE TRIGGER infra_node_status_on_row + AFTER INSERT OR DELETE ON infra_node + FOR EACH ROW + EXECUTE FUNCTION infra_node_status_notify()"#, + r#"DROP TRIGGER IF EXISTS infra_node_status_on_change ON infra_node"#, + r#"CREATE TRIGGER infra_node_status_on_change + AFTER UPDATE OF status ON infra_node + FOR EACH ROW + WHEN (NEW.status IS DISTINCT FROM OLD.status) + EXECUTE FUNCTION infra_node_status_notify()"#, ], seed: &[], }; @@ -190,6 +215,36 @@ pub async fn get( } } +/// One copy of an infra node and whether it is up: what a run checks +/// before it starts. +#[derive(Debug, Clone)] +pub struct CopyStatus { + pub node_id: String, + pub instance: Option, + pub status: InfraNodeStatus, +} + +/// Every copy of `project_id`'s infra nodes and its status. +pub async fn statuses(pool: &PgPool, project_id: uuid::Uuid) -> Result> { + let rows: Vec<(String, Option, String)> = + sqlx::query_as("SELECT node_id, instance_id, status FROM infra_node WHERE project_id = $1") + .bind(project_id) + .fetch_all(pool) + .await?; + rows.into_iter() + .map(|(node_id, instance, status)| { + let instance = instance + .map(weft_core::instance::InstanceId::new) + .transpose() + .map_err(|e| anyhow::anyhow!("infra_node.instance_id for project={project_id} node={node_id}: {e}"))?; + let status = InfraNodeStatus::parse(&status).ok_or_else(|| { + anyhow::anyhow!("infra_node.status='{status}' for project={project_id} node={node_id} is not a status this dispatcher knows") + })?; + Ok(CopyStatus { node_id, instance, status }) + }) + .collect() +} + /// List every row for a project. Drives the project status response. pub async fn list_for_project( pool: &PgPool, diff --git a/crates/weft-dispatcher/src/journal/fake.rs b/crates/weft-dispatcher/src/journal/fake.rs index 90eea859..190346d5 100644 --- a/crates/weft-dispatcher/src/journal/fake.rs +++ b/crates/weft-dispatcher/src/journal/fake.rs @@ -323,7 +323,7 @@ impl Journal for FakeJournal { start: &ExecEvent, kicks: &[ExecEvent], task: weft_task_store::tasks::NewTask, - _expected_activation: Option, + _for_activation: bool, ) -> anyhow::Result<()> { let mut g = self.inner.lock().unwrap(); let ExecEvent::ExecutionStarted { execution_id, project_id, phase, .. } = start else { @@ -344,6 +344,18 @@ impl Journal for FakeJournal { Ok(()) } + /// Every run is admitted: the limits are the database's, and the + /// database tests are where they are checked. + async fn admit_and_start_execution( + &self, + _admission: &crate::entry_limits::Admission, + start: &ExecEvent, + kicks: &[ExecEvent], + task: weft_task_store::tasks::NewTask, + ) -> anyhow::Result> { + self.start_execution(start, kicks, task, false).await.map(Ok) + } + async fn cancel_execution( &self, execution_id: ExecutionId, @@ -1125,7 +1137,7 @@ pub(crate) mod tests { unrecorded_birth: Some(&birth), }) .unwrap(); - j.start_execution(&start, &[kick], task, None).await.unwrap(); + j.start_execution(&start, &[kick], task, false).await.unwrap(); assert!(j.events_log(execution_id).await.unwrap().is_empty(), "no journal row for an unrecorded birth"); assert!(j.execution_owner(execution_id).await.unwrap().is_some(), "but the execution exists"); assert!(j.list_non_terminal_execution_ids_for_project(PROJECT).await.unwrap().is_empty()); diff --git a/crates/weft-dispatcher/src/journal/mod.rs b/crates/weft-dispatcher/src/journal/mod.rs index cbd92887..c4d82dcd 100644 --- a/crates/weft-dispatcher/src/journal/mod.rs +++ b/crates/weft-dispatcher/src/journal/mod.rs @@ -185,14 +185,30 @@ pub trait Journal: Send + Sync { /// ATOMICALLY journal an execution's birth together with its queued work /// item. Either everything commits or nothing does. `start` must be /// `ExecEvent::ExecutionStarted`; `kicks` are its `NodeKicked` events. + /// `for_activation`: a trigger setup an activation asked for (the + /// activation is the setup's own execution), born only while that + /// activation still owns its rows. async fn start_execution( &self, start: &ExecEvent, kicks: &[ExecEvent], task: weft_task_store::tasks::NewTask, - expected_activation: Option, + for_activation: bool, ) -> anyhow::Result<()>; + /// [`Self::start_execution`] behind the entry's limits, in the SAME + /// commit: the run is admitted (`admission`, counted and given its + /// slot) and born together, or refused and not born. The answer to a + /// caller waiting at the door costs one round trip this way. A run + /// already born is left as it is, with the slot it took. + async fn admit_and_start_execution( + &self, + admission: &crate::entry_limits::Admission, + start: &ExecEvent, + kicks: &[ExecEvent], + task: weft_task_store::tasks::NewTask, + ) -> anyhow::Result>; + /// THE dispatcher-side cancel of an execution, in ONE transaction: /// strip the execution's wake signals (the parked form, the timer, the /// webhook, so nothing can revive it), journal its cancel terminals diff --git a/crates/weft-dispatcher/src/journal/postgres.rs b/crates/weft-dispatcher/src/journal/postgres.rs index 738def45..725067a5 100644 --- a/crates/weft-dispatcher/src/journal/postgres.rs +++ b/crates/weft-dispatcher/src/journal/postgres.rs @@ -314,128 +314,147 @@ impl PostgresJournal { .await .map_err(|e| anyhow::anyhow!("{e}")); } - let mut tx = self.pool.begin().await?; - // Execution lock before the first write (`retain_source_version`): - // the ordering invariant on `weft_journal::write`. - weft_journal::lock_execution_ids(&mut tx, &[event.execution_id()]).await?; - Self::write_started_in(&mut tx, event, dedup_key).await?; - tx.commit().await?; + let started = StartedRow::of(event, dedup_key, crate::lease::now_unix())?; + sqlx::query("SELECT weft_execution_started($1)") + .bind(serde_json::to_value(&started)?) + .execute(&self.pool) + .await + .map_err(birth_refusal)?; Ok(()) } - /// Check the durable birth, whose lifetime extends beyond its initial - /// execute task. The caller holds the execution's lock - /// (`weft_journal::lock_execution_ids`), which is what serializes admission. - async fn execution_already_started(tx: &mut sqlx::PgConnection, start: &ExecEvent) -> anyhow::Result { - let ExecEvent::ExecutionStarted { execution_id, run_kind, .. } = start else { - anyhow::bail!("execution admission requires a birth event"); - }; - // An unrecorded run journals no `ExecutionStarted`: its birth is - // its `execution` row alone. - let query = if run_kind.journaled() { - "SELECT EXISTS (SELECT 1 FROM exec_event WHERE execution_id = $1 AND kind = 'execution_started')" - } else { - "SELECT EXISTS (SELECT 1 FROM execution WHERE execution_id = $1)" + /// A run's birth (`weft_start_execution`), behind `admission` when it + /// has one: one round trip whatever the birth writes. + async fn birth( + &self, + start: &ExecEvent, + kicks: &[ExecEvent], + task: weft_task_store::tasks::NewTask, + trigger_setup: Option, + admission: Option<&crate::entry_limits::Admission>, + ) -> anyhow::Result> { + let now = crate::lease::now_unix(); + let started = StartedRow::of(start, None, now)?; + if let Some(stray) = kicks.iter().find(|kick| kick.execution_id() != start.execution_id()) { + anyhow::bail!("a birth of execution {} kicks a node of execution {}", start.execution_id(), stray.execution_id()); + } + let kicks = kicks + .iter() + .map(|kick| Ok(EventRow { kind: kick.kind_str(), payload: serde_json::to_string(kick)? })) + .collect::>>()?; + let call = BirthCall { + started, + kicks, + task: weft_task_store::tasks::DedupRow::of(&task, uuid::Uuid::new_v4(), now)?, + trigger_setup, + admission, }; - Ok(sqlx::query_scalar(query).bind(execution_id.to_string()).fetch_one(&mut *tx).await?) + let answer: serde_json::Value = sqlx::query_scalar("SELECT weft_start_execution($1)") + .bind(serde_json::to_value(&call)?) + .fetch_one(&self.pool) + .await + .map_err(birth_refusal)?; + match answer.get("outcome").and_then(serde_json::Value::as_str) { + Some("started" | "already_started") => Ok(Ok(())), + Some("refused") => { + let refused: crate::entry_limits::Answer = + serde_json::from_value(answer["refused"].clone()).context("read the birth's refusal")?; + Ok(Err(refused.refused()?)) + } + _ => anyhow::bail!("the database answered a birth with '{answer}', which weft does not read"), + } } +} - /// Write an `ExecutionStarted` event AND its `execution` seed on the - /// caller's transaction (the two must commit together; see - /// `record_with_seed`'s doc). A missing project row fails the whole write - /// loudly instead of silently journaling an unsweepable execution. - /// The caller holds the execution's lock (`weft_journal::lock_execution_ids`) - /// from before its first write. - async fn write_started_in( - tx: &mut sqlx::PgConnection, - event: &ExecEvent, - dedup_key: Option<&str>, - ) -> anyhow::Result<()> { +/// A failure of the birth functions. One they raise themselves (`RAISE +/// EXCEPTION`, SQLSTATE P0001) is passed on in its own words, which name +/// what is missing and what to do; any other keeps the database's whole +/// error and where in the functions it happened, since one call now does +/// what several statements did. +fn birth_refusal(e: sqlx::Error) -> anyhow::Error { + let Some(db) = e.as_database_error() else { return anyhow::Error::from(e) }; + if db.code().as_deref() == Some("P0001") { + return anyhow::anyhow!("{}", db.message()); + } + let at = db.try_downcast_ref::().and_then(|pg| pg.r#where()).map(str::to_string); + let e = anyhow::Error::from(e); + match at { + Some(at) => e.context(format!("a run's birth failed in the database, at: {at}")), + None => e.context("a run's birth failed in the database"), + } +} + +/// One journal row as the birth functions take it. +#[derive(serde::Serialize)] +struct EventRow { + kind: &'static str, + payload: String, +} + +/// An `ExecutionStarted` and its seed, as `weft_execution_started` takes +/// them. +// SYNC: StartedRow's fields <-> weft_execution_started, weft_start_execution (GROUP below) +#[derive(serde::Serialize)] +struct StartedRow { + execution_id: String, + project_id: uuid::Uuid, + /// Whether the run keeps a journal; an unrecorded run is born with its + /// seed alone, its birth riding the execute task. + journaled: bool, + kind: &'static str, + payload: String, + at_unix: i64, + phase: &'static str, + run_kind: &'static str, + instance_id: Option, + fired_by: Option, + source_version: Option, + dedup_key: Option, + created_at: i64, +} + +impl StartedRow { + fn of(event: &ExecEvent, dedup_key: Option<&str>, created_at: i64) -> anyhow::Result { let ExecEvent::ExecutionStarted { execution_id, project_id, at_unix, phase, run_kind, source_version, instance, fired_trigger, .. } = event else { - anyhow::bail!("write_started_in requires an ExecutionStarted event"); + anyhow::bail!("an execution's birth row must be its ExecutionStarted"); }; - if let Some(version) = source_version { - crate::versions::retain_source_version(tx, *project_id, version).await?; - } - // An unrecorded run is born with its execution row alone: its - // `ExecutionStarted` rides the execute task instead. - if run_kind.journaled() { - weft_journal::record_event_in(&mut *tx, event, None, dedup_key) - .await - .map_err(|e| anyhow::anyhow!("{e}"))?; - } - let rows = sqlx::query( - "INSERT INTO execution (execution_id, project_id, tenant_id, started_at_unix, phase, kind, instance_id, fired_by) \ - SELECT $1, $2, p.tenant_id, $3, $4, $5, $6, $7 FROM project p WHERE p.id = $2 \ - ON CONFLICT (execution_id) DO NOTHING", - ) - .bind(execution_id.to_string()) - .bind(project_id) - .bind(*at_unix as i64) - .bind(phase.as_str()) - .bind(run_kind.as_str()) - .bind(instance.as_ref().map(|m| m.as_str())) - .bind(fired_trigger.as_deref()) - .execute(&mut *tx) - .await?; - if rows.rows_affected() == 0 { - // Zero rows = conflict (already seeded: a dedup'd retry) - // OR missing project. Only the latter is an error. - let (already_seeded,): (bool,) = sqlx::query_as( - "SELECT EXISTS(SELECT 1 FROM execution WHERE execution_id = $1)", - ) - .bind(execution_id.to_string()) - .fetch_one(&mut *tx) - .await?; - if !already_seeded { - anyhow::bail!( - "refuse to journal ExecutionStarted for execution {execution_id}: project \ - {project_id} has no row, so the execution seed (which the \ - broker scope check and the terminal sweeps depend on) cannot be \ - written; register the project first" - ); - } - } - Ok(()) + Ok(Self { + execution_id: execution_id.to_string(), + project_id: *project_id, + journaled: run_kind.journaled(), + kind: event.kind_str(), + payload: serde_json::to_string(event)?, + at_unix: *at_unix as i64, + phase: phase.as_str(), + run_kind: run_kind.as_str(), + instance_id: instance.as_ref().map(|m| m.as_str().to_string()), + fired_by: fired_trigger.clone(), + source_version: source_version.clone(), + dedup_key: dedup_key.map(str::to_string), + created_at, + }) } +} - /// The one-transaction execution BIRTH: `ExecutionStarted` + seed + the - /// entry kicks, written by `start_execution`, which takes the - /// execution's lock before its first write. - async fn write_birth_in( - tx: &mut sqlx::PgConnection, - start: &ExecEvent, - kicks: &[ExecEvent], - expected_activation: Option, - ) -> anyhow::Result<()> { - Self::write_started_in(tx, start, None).await?; - if let ExecEvent::ExecutionStarted { execution_id, project_id, phase: weft_core::context::Phase::TriggerSetup, .. } = start { - // Serialize with the activation's claim: an activation's setup - // is born only while that activation still owns its rows (a - // cancel between the claim and here wins). Ownership is born - // with the task, so a dead requester cannot leave an owner - // without work. A bake (no activation) claims nothing. - if let Some(expected) = expected_activation { - anyhow::ensure!(expected == *execution_id, "activation {expected} starts setup {execution_id}"); - let owned: bool = sqlx::query_scalar( - "SELECT EXISTS (SELECT 1 FROM trigger_activation \ - WHERE project_id = $1 AND activating_execution_id = $2 AND status = 'activating' FOR UPDATE)", - ).bind(project_id).bind(expected).fetch_one(&mut *tx).await?; - anyhow::ensure!(owned, "activation {expected} ended before trigger setup could start"); - } - sqlx::query("INSERT INTO trigger_setup (project_id, execution_id) VALUES ($1, $2)") - .bind(project_id).bind(execution_id.to_string()).execute(&mut *tx).await?; - } - let journaled = matches!(start, ExecEvent::ExecutionStarted { run_kind, .. } if run_kind.journaled()); - for kick in kicks.iter().filter(|_| journaled) { - weft_journal::record_event_in(&mut *tx, kick, None, None) - .await - .map_err(|e| anyhow::anyhow!("{e}"))?; - } - Ok(()) - } +/// A trigger setup's birth: recorded in flight, and, when an activation +/// asked for it (the activation is the setup's own execution), born only +/// while that activation still owns its rows. +#[derive(serde::Serialize)] +struct TriggerSetupRow { + for_activation: bool, +} + +/// Everything `weft_start_execution` takes. +// SYNC: BirthCall's fields <-> weft_start_execution (GROUP below) +#[derive(serde::Serialize)] +struct BirthCall<'a> { + started: StartedRow, + kicks: Vec, + task: weft_task_store::tasks::DedupRow, + trigger_setup: Option, + admission: Option<&'a crate::entry_limits::Admission>, } /// The journal's schema. First in `app::ALL_GROUPS` (other groups' @@ -706,6 +725,42 @@ pub static GROUP: weft_task_store::SchemaGroup = weft_task_store::SchemaGroup { FOR EACH ROW WHEN (NEW.holds IS DISTINCT FROM OLD.holds) EXECUTE FUNCTION signal_held_notify()"#, + // Tell every dispatcher a tenant's routes changed, so its copy of + // them (`crate::held::Held::routes`) is read again: a public entry + // coming, going, or changing anything the handshake reads. The + // project group's trigger uses the same function (a project row + // gone reads as an inactive route). + // SYNC: 'weft_routes' <-> crate::held::ROUTES_CHANNEL + r#"CREATE OR REPLACE FUNCTION routes_notify_tenant() RETURNS trigger AS $$ + BEGIN + IF TG_OP = 'DELETE' THEN + PERFORM pg_notify('weft_routes', OLD.tenant_id); + ELSE + PERFORM pg_notify('weft_routes', NEW.tenant_id); + END IF; + RETURN NULL; + END; + $$ LANGUAGE plpgsql"#, + r#"DROP TRIGGER IF EXISTS signal_routes_on_insert ON signal"#, + r#"CREATE TRIGGER signal_routes_on_insert + AFTER INSERT ON signal + FOR EACH ROW + WHEN (NEW.surface_kind = 'public_entry') + EXECUTE FUNCTION routes_notify_tenant()"#, + r#"DROP TRIGGER IF EXISTS signal_routes_on_delete ON signal"#, + r#"CREATE TRIGGER signal_routes_on_delete + AFTER DELETE ON signal + FOR EACH ROW + WHEN (OLD.surface_kind = 'public_entry') + EXECUTE FUNCTION routes_notify_tenant()"#, + r#"DROP TRIGGER IF EXISTS signal_routes_on_change ON signal"#, + r#"CREATE TRIGGER signal_routes_on_change + AFTER UPDATE OF surface_kind, mount_path, mount_methods, project_id, node_id, spec_json, + auth_kind, auth_config, port_snapshot, program_json, source_version, instance_id, + activation_trigger ON signal + FOR EACH ROW + WHEN (NEW.surface_kind = 'public_entry' OR OLD.surface_kind = 'public_entry') + EXECUTE FUNCTION routes_notify_tenant()"#, // Entry rows are keyed by (project_id, node_id), `node_id` // being the trigger's place spelled the way a person writes // it (`one.door`), so a file called from two places holds two @@ -835,6 +890,145 @@ pub static GROUP: weft_task_store::SchemaGroup = weft_task_store::SchemaGroup { UNIQUE (execution_id, tag) )"#, r#"CREATE INDEX IF NOT EXISTS idx_execution_tag_tag ON execution_tag(tag)"#, + // An execution's journal lock, held until the transaction ends: the + // ONE spelling of its key, taken by every journal write here and by + // `weft_journal::lock_execution_ids`. + r#"CREATE OR REPLACE FUNCTION weft_lock_execution(p_execution_id TEXT) RETURNS VOID AS $$ + BEGIN + PERFORM pg_advisory_xact_lock(hashtextextended('exec_event:' || p_execution_id, 0)); + END; + $$ LANGUAGE plpgsql"#, + // THE journal insert (`weft_journal::write`): the execution's lock, + // then the rows in the order given, fenced on `p_owner` when one is + // named. Returns how many rows went in. The lock comes first so + // rows of one execution are numbered and committed in the same + // order; taking it again in a transaction that already holds it is + // free. + r#"CREATE OR REPLACE FUNCTION weft_journal_append( + p_execution_id TEXT, p_kinds TEXT[], p_payloads TEXT[], p_created_at BIGINT, + p_replica TEXT, p_owner TEXT, p_dedup_key TEXT + ) RETURNS BIGINT AS $$ + DECLARE + written BIGINT; + BEGIN + PERFORM weft_lock_execution(p_execution_id); + INSERT INTO exec_event (execution_id, kind, payload_json, created_at, replica, dedup_key) + SELECT p_execution_id, e.kind, e.payload, p_created_at, p_replica, p_dedup_key + FROM unnest(p_kinds, p_payloads) WITH ORDINALITY AS e(kind, payload, n) + WHERE p_owner IS NULL + OR EXISTS (SELECT 1 FROM execution x + WHERE x.execution_id = p_execution_id AND x.owner_replica = p_owner) + ORDER BY e.n + ON CONFLICT (dedup_key) WHERE dedup_key IS NOT NULL DO NOTHING; + GET DIAGNOSTICS written = ROW_COUNT; + RETURN written; + END; + $$ LANGUAGE plpgsql"#, + // An execution's `ExecutionStarted` and its `execution` seed, + // which commit together: the seed is what the broker's scope check + // and the terminal sweeps see an execution by, so one without the + // other would be an execution nothing can ever sweep. A missing + // project row refuses the whole write. The source version the run + // was prepared from is held (FOR KEY SHARE) so it cannot be removed + // under the run. An unrecorded run is born with its seed alone. + // SYNC: p's fields <-> crate::journal::postgres::StartedRow + r#"CREATE OR REPLACE FUNCTION weft_execution_started(p JSONB) RETURNS VOID AS $$ + DECLARE + v_execution_id TEXT := p->>'execution_id'; + v_project UUID := (p->>'project_id')::uuid; + seeded BIGINT; + BEGIN + PERFORM weft_lock_execution(v_execution_id); + IF p->>'source_version' IS NOT NULL THEN + PERFORM 1 FROM project_version + WHERE project_id = v_project AND id = p->>'source_version' FOR KEY SHARE; + IF NOT FOUND THEN + RAISE EXCEPTION 'source version % was removed during preparation; run the command again', + p->>'source_version'; + END IF; + END IF; + IF (p->>'journaled')::boolean THEN + PERFORM weft_journal_append(v_execution_id, ARRAY[p->>'kind'], ARRAY[p->>'payload'], + (p->>'created_at')::bigint, NULL, NULL, p->>'dedup_key'); + END IF; + INSERT INTO execution (execution_id, project_id, tenant_id, started_at_unix, phase, kind, instance_id, fired_by) + SELECT v_execution_id, v_project, pr.tenant_id, (p->>'at_unix')::bigint, p->>'phase', + p->>'run_kind', p->>'instance_id', p->>'fired_by' + FROM project pr WHERE pr.id = v_project + ON CONFLICT (execution_id) DO NOTHING; + GET DIAGNOSTICS seeded = ROW_COUNT; + IF seeded = 0 AND NOT EXISTS (SELECT 1 FROM execution WHERE execution_id = v_execution_id) THEN + RAISE EXCEPTION 'refuse to journal ExecutionStarted for execution %: project % has no row, so the execution seed (which the broker scope check and the terminal sweeps depend on) cannot be written; register the project first', + v_execution_id, v_project; + END IF; + END; + $$ LANGUAGE plpgsql"#, + // A run's whole birth, in one call: one execution's admission (its + // entry's limits, when it has any), its first task, its + // `ExecutionStarted` and seed, and the kicks that start it. + // Answers `{"outcome": "started" | "already_started" | "refused"}`, + // a refusal carrying which limit (`weft_admit`'s answer). Nothing + // is written for a run already born (its birth row, or its seed for + // an unrecorded run) or whose task is already live: a retried birth + // collapses onto the first, which keeps the slot it took. A refused + // run writes only its counts. The execution's lock is taken before + // any write, which is the journal's ordering rule. + // SYNC: p's fields <-> crate::journal::postgres::BirthCall + r#"CREATE OR REPLACE FUNCTION weft_start_execution(p JSONB) RETURNS JSONB AS $$ + DECLARE + v_started JSONB := p->'started'; + v_task JSONB := p->'task'; + v_execution_id TEXT := v_started->>'execution_id'; + v_project UUID := (v_started->>'project_id')::uuid; + v_journaled BOOLEAN := (v_started->>'journaled')::boolean; + v_refused JSONB; + v_inserted BOOLEAN; + BEGIN + PERFORM weft_lock_execution(v_execution_id); + IF (v_journaled AND EXISTS (SELECT 1 FROM exec_event + WHERE execution_id = v_execution_id AND kind = 'execution_started')) + OR (NOT v_journaled AND EXISTS (SELECT 1 FROM execution WHERE execution_id = v_execution_id)) THEN + RETURN jsonb_build_object('outcome', 'already_started'); + END IF; + IF jsonb_typeof(p->'admission') = 'object' THEN + v_refused := weft_admit(p->'admission'); + IF v_refused IS NOT NULL THEN + RETURN jsonb_build_object('outcome', 'refused', 'refused', v_refused); + END IF; + END IF; + SELECT d.inserted INTO v_inserted FROM weft_enqueue_dedup( + (v_task->>'id')::uuid, v_task->>'kind', v_task->>'target', (v_task->>'project_id')::uuid, + v_task->>'dedup_key', v_task->>'execution_id', v_task->>'tenant_id', v_task->>'target_replica', + v_task->>'binary_hash', v_task->'payload', (v_task->>'created_at')::bigint) d; + IF NOT v_inserted THEN + RETURN jsonb_build_object('outcome', 'already_started'); + END IF; + PERFORM weft_execution_started(v_started); + -- A trigger setup is born only while the activation that asked + -- for it still owns its rows (a cancel between the claim and + -- here wins), and is recorded as in flight. + IF jsonb_typeof(p->'trigger_setup') = 'object' THEN + IF (p->'trigger_setup'->>'for_activation')::boolean THEN + PERFORM 1 FROM trigger_activation + WHERE project_id = v_project + AND activating_execution_id = v_execution_id::uuid + AND status = 'activating' + FOR UPDATE; + IF NOT FOUND THEN + RAISE EXCEPTION 'activation % ended before trigger setup could start', v_execution_id; + END IF; + END IF; + INSERT INTO trigger_setup (project_id, execution_id) VALUES (v_project, v_execution_id); + END IF; + IF v_journaled AND jsonb_array_length(p->'kicks') > 0 THEN + PERFORM weft_journal_append(v_execution_id, + ARRAY(SELECT k->>'kind' FROM jsonb_array_elements(p->'kicks') WITH ORDINALITY AS e(k, n) ORDER BY n), + ARRAY(SELECT k->>'payload' FROM jsonb_array_elements(p->'kicks') WITH ORDINALITY AS e(k, n) ORDER BY n), + (v_started->>'created_at')::bigint, NULL, NULL, NULL); + END IF; + RETURN jsonb_build_object('outcome', 'started'); + END; + $$ LANGUAGE plpgsql"#, ], seed: &[], }; @@ -919,26 +1113,28 @@ impl Journal for PostgresJournal { start: &ExecEvent, kicks: &[ExecEvent], task: weft_task_store::tasks::NewTask, - expected_activation: Option, + for_activation: bool, ) -> anyhow::Result<()> { - let mut tx = self.pool.begin().await?; - // Execution lock before the first write (the task enqueue): the - // ordering invariant on `weft_journal::write`. - weft_journal::lock_execution_ids(&mut tx, &[start.execution_id()]).await?; - if Self::execution_already_started(&mut tx, start).await? { - tx.commit().await?; - return Ok(()); - } - // Enqueue FIRST and only write the birth on a FRESH insert: "one - // birth per execution" holds by construction even if a caller ever replays an execution - // (the dedup'd task collapses, and the birth is not double-written; - // ExecutionStarted itself carries no dedup key). - let outcome = weft_task_store::tasks::enqueue_dedup_in(&mut tx, task).await?; - if matches!(outcome, weft_task_store::tasks::DedupOutcome::Inserted(_)) { - Self::write_birth_in(&mut tx, start, kicks, expected_activation).await?; - } - tx.commit().await?; - Ok(()) + let trigger_setup = match start { + ExecEvent::ExecutionStarted { phase: weft_core::context::Phase::TriggerSetup, .. } => Some(TriggerSetupRow { for_activation }), + _ => None, + }; + let born = self.birth(start, kicks, task, trigger_setup, None).await?; + born.map_err(|refused| anyhow::anyhow!("a birth with no limits to check was refused: {refused:?}")) + } + + async fn admit_and_start_execution( + &self, + admission: &crate::entry_limits::Admission, + start: &ExecEvent, + kicks: &[ExecEvent], + task: weft_task_store::tasks::NewTask, + ) -> anyhow::Result> { + anyhow::ensure!( + !matches!(start, ExecEvent::ExecutionStarted { phase: weft_core::context::Phase::TriggerSetup, .. }), + "a trigger setup is not admitted at an entry's limits" + ); + self.birth(start, kicks, task, None, Some(admission)).await } async fn cancel_execution( diff --git a/crates/weft-dispatcher/src/lib.rs b/crates/weft-dispatcher/src/lib.rs index 3146c244..e2058296 100644 --- a/crates/weft-dispatcher/src/lib.rs +++ b/crates/weft-dispatcher/src/lib.rs @@ -23,6 +23,7 @@ pub mod delivery; pub mod domains; pub mod door; pub mod frontends; +pub mod held; pub mod holders; pub mod entry_limits; pub mod display_feeds; diff --git a/crates/weft-dispatcher/src/project_store.rs b/crates/weft-dispatcher/src/project_store.rs index 91ce7e28..a4a3dbab 100644 --- a/crates/weft-dispatcher/src/project_store.rs +++ b/crates/weft-dispatcher/src/project_store.rs @@ -366,6 +366,35 @@ pub static GROUP: weft_task_store::SchemaGroup = weft_task_store::SchemaGroup { head_run UUID )"#, "CREATE INDEX IF NOT EXISTS idx_project_tenant ON project(tenant_id)", + // A project coming or going changes how its tenant's routes read + // (a route whose project is gone is inactive); the function is the + // journal group's, which applies first. + r#"DROP TRIGGER IF EXISTS project_routes_on_row ON project"#, + r#"CREATE TRIGGER project_routes_on_row + AFTER INSERT OR DELETE ON project + FOR EACH ROW + EXECUTE FUNCTION routes_notify_tenant()"#, + // Tell every dispatcher a project's worker levers changed (or the + // project went), so its copy of them (`crate::held::Held`) is read + // again. + // SYNC: 'weft_worker_settings' <-> crate::held::WORKER_SETTINGS_CHANNEL + r#"CREATE OR REPLACE FUNCTION project_worker_settings_notify() RETURNS trigger AS $$ + BEGIN + PERFORM pg_notify('weft_worker_settings', OLD.id::text); + RETURN NULL; + END; + $$ LANGUAGE plpgsql"#, + r#"DROP TRIGGER IF EXISTS project_worker_settings_on_change ON project"#, + r#"CREATE TRIGGER project_worker_settings_on_change + AFTER UPDATE OF worker_settings_json ON project + FOR EACH ROW + WHEN (NEW.worker_settings_json IS DISTINCT FROM OLD.worker_settings_json) + EXECUTE FUNCTION project_worker_settings_notify()"#, + r#"DROP TRIGGER IF EXISTS project_worker_settings_on_delete ON project"#, + r#"CREATE TRIGGER project_worker_settings_on_delete + AFTER DELETE ON project + FOR EACH ROW + EXECUTE FUNCTION project_worker_settings_notify()"#, // Append-only definition-version history. Workers fetch by // (project_id, definition_hash) so a suspended execution // can always resume on the EXACT shape it was started on, diff --git a/crates/weft-dispatcher/src/settled.rs b/crates/weft-dispatcher/src/settled.rs index a23c0d0a..200cea68 100644 --- a/crates/weft-dispatcher/src/settled.rs +++ b/crates/weft-dispatcher/src/settled.rs @@ -36,12 +36,15 @@ // if a transaction holds the execution's lock before it gets its xid (at its // first write), which is the invariant on `weft_journal::write`. Every // transaction that writes something else before its exec_event rows -// takes `weft_journal::lock_execution_ids` first: +// takes the execution's lock first (`weft_journal::lock_execution_ids`, or +// the same lock in SQL): // SYNC: execution lock before first write <-> crate::journal::postgres -// (record_with_seed, start_execution, cancel_execution). Transactions whose first write is the -// exec_event insert are covered by the lock that insert takes: -// weft_journal::tags::tag_execution_in (broker execution_tag) and -// every single-statement record_event_* call. +// (cancel_execution; weft_execution_started and weft_start_execution, the +// SQL functions behind record_with_seed and every birth), +// weft_journal::unrecorded (record_retroactively). Transactions whose +// first write is the exec_event insert are covered by the lock that insert +// takes (`weft_journal_append`): weft_journal::tags::tag_execution_in +// (broker execution_tag) and every single-statement record_event_* call. use std::time::{Duration, Instant}; @@ -327,6 +330,8 @@ mod db_tests { /// reopened-run bug `weft_journal::lock_execution_ids` exists for. #[sqlx::test] async fn locking_the_execution_id_before_any_write_keeps_its_rows_in_xid_order(pool: PgPool) { + // The execution's lock is the journal schema's (`weft_lock_execution`). + crate::app::apply_core_schema(&pool).await.unwrap(); table(&pool).await; for lock_first in [true, false] { let execution_id = uuid::Uuid::new_v4(); diff --git a/crates/weft-dispatcher/src/state.rs b/crates/weft-dispatcher/src/state.rs index 6dd5c37c..325d9fc9 100644 --- a/crates/weft-dispatcher/src/state.rs +++ b/crates/weft-dispatcher/src/state.rs @@ -112,7 +112,10 @@ pub struct DispatcherState { /// (see [`DispatcherState::program`]): a definition never changes /// under its hash, so a busy route stops reading and parsing its /// program on every call. - pub programs: Arc>, + pub programs: Arc>, + /// The rows every live call reads, held in memory and read again when + /// they change (`crate::held`). + pub held: Arc, } impl DispatcherState { @@ -133,13 +136,13 @@ impl DispatcherState { project: uuid::Uuid, hash: &str, ) -> anyhow::Result>> { - if let Some(program) = self.programs.get(project, hash) { + if let Some(program) = self.programs.get(&(project, hash.to_string())) { return Ok(Some(program)); } let Some(json) = self.projects.definition_for_hash(project, hash).await? else { return Ok(None) }; let program: Arc = Arc::new(serde_json::from_str(&json).map_err(|e| UnreadableProgram(e.to_string()))?); - self.programs.put(project, hash.to_string(), program.clone()); + self.programs.put((project, hash.to_string()), program.clone()); Ok(Some(program)) } } diff --git a/crates/weft-dispatcher/src/task_kinds/route_entry.rs b/crates/weft-dispatcher/src/task_kinds/route_entry.rs index f7bff914..1edc07f2 100644 --- a/crates/weft-dispatcher/src/task_kinds/route_entry.rs +++ b/crates/weft-dispatcher/src/task_kinds/route_entry.rs @@ -169,36 +169,11 @@ impl TaskExecutor for RouteEntryExecutor { // re-parks like every step before it. A write that committed // but failed to acknowledge re-parks too, and the drained twin // finds the execution born (above) and finishes it. - // The entry's at-once limit, taken for this run's execution just - // before it is born. A full entry parks the fire, which the - // reaper retries on its backoff, so the fire waits for a run - // to end instead of being lost. The same execution on a retry - // keeps the slot it already holds. - let limits = spec.limits.resolve(); - if let Some(max) = limits.at_once { - let now = crate::lease::now_unix(); - let slot = crate::entry_limits::take_slot( - &state.pg_pool, - &payload.token, - &execution_id.to_string(), - max, - now + crate::entry_limits::UNBORN_FIRE_SLOT_SECS, - now, - ) - .await; - match slot { - Ok(Ok(())) => {} - Ok(Err(refused)) => { - if let Err(e) = - crate::entry_limits::note_refusal(&state.pg_pool, &payload.token, refused.reason, now).await - { - tracing::warn!(target: "weft_dispatcher::route_entry", error = %e, "could not count a refusal"); - } - return park_fire(state, task, &payload, &Unrouted::Retry(format!("{} is reached", refused.reason.describe()))).await; - } - Err(e) => return park_fire(state, task, &payload, &Unrouted::Retry(format!("entry slot: {e}"))).await, - } - } + // The entry's at-once limit is taken for this run's execution + // in the same commit as its birth. A full entry parks the fire, + // which the reaper retries on its backoff, so the fire waits for + // a run to end instead of being lost. The same execution on a + // retry keeps the slot it already holds. let execution_task = crate::task_kinds::execute::execution_task_spec(crate::task_kinds::execute::ExecutionTask { kind: weft_task_store::TaskKind::Execute, project_id: signal.project_id, @@ -210,12 +185,20 @@ impl TaskExecutor for RouteEntryExecutor { live_connection: None, unrecorded_birth: None, })?; - if let Err(e) = state - .journal - .start_execution(&start, &kick_events, execution_task, None) - .await - { - return park_fire(state, task, &payload, &Unrouted::Retry(format!("ExecutionStarted write: {e}"))).await; + let admission = crate::entry_limits::Admission::fire( + &payload.token, + &spec.limits.resolve(), + &execution_id.to_string(), + crate::lease::now_unix(), + ); + match state.journal.admit_and_start_execution(&admission, &start, &kick_events, execution_task).await { + Ok(Ok(())) => {} + Ok(Err(refused)) => { + return park_fire(state, task, &payload, &Unrouted::Retry(format!("{} is reached", refused.reason.describe()))).await; + } + Err(e) => { + return park_fire(state, task, &payload, &Unrouted::Retry(format!("ExecutionStarted write: {e}"))).await; + } } }; diff --git a/crates/weft-dispatcher/tests/db_entry_limits.rs b/crates/weft-dispatcher/tests/db_entry_limits.rs index c526e233..cd6a79a1 100644 --- a/crates/weft-dispatcher/tests/db_entry_limits.rs +++ b/crates/weft-dispatcher/tests/db_entry_limits.rs @@ -8,12 +8,39 @@ use sqlx::PgPool; use weft_core::signal::{EntryLimits, ResolvedLimits}; -use weft_dispatcher::entry_limits::{self, Limited}; +use weft_dispatcher::entry_limits::{self, Admission, EdgeConfig, Limited, Refused}; async fn setup(pool: &PgPool) { weft_dispatcher::app::apply_core_schema(pool).await.expect("core schema"); } +fn edge(invalid_tokens_per_minute: Option) -> EdgeConfig { + EdgeConfig { + trusted_proxy_hops: weft_platform_traits::config::ProxyHops { public: 1, outside: 1, domains: 2 }, + invalid_tokens_per_minute, + } +} + +/// One call by `caller` to the entry `tok`, admitted on its own. +async fn call(pool: &PgPool, caller: &str, l: &ResolvedLimits, now: i64) -> Result<(), Refused> { + entry_limits::admit(pool, &Admission::call(&edge(None), None, "tok", caller, l, None, now)).await.expect("admit") +} + +/// A slot of `tok` (at most `max` at once) taken for the run `execution_id`, +/// holding until `unborn_until` if it never starts. +async fn take(pool: &PgPool, execution_id: &str, max: u32, unborn_until: i64, now: i64) -> Result<(), Refused> { + let l = limits(None, None, Some(max)); + entry_limits::admit(pool, &Admission::call(&edge(None), None, "tok", "ip:a", &l, Some((execution_id, unborn_until)), now)) + .await + .expect("take") +} + +/// Whether `tok` has no room for one more run, taking nothing. +async fn full(pool: &PgPool, max: u32, now: i64) -> bool { + let l = limits(None, None, Some(max)); + entry_limits::admit(pool, &Admission::call(&edge(None), None, "tok", "ip:a", &l, None, now)).await.expect("check").is_err() +} + fn limits(per_caller: Option, per_entry: Option, at_once: Option) -> ResolvedLimits { EntryLimits { per_caller_per_minute: per_caller.or(Some(0)), per_minute: per_entry.or(Some(0)), at_once: at_once.or(Some(0)) } .resolve() @@ -31,7 +58,7 @@ async fn concurrent_callers_across_replicas_share_one_count(pool: PgPool) { for i in 0..80 { let p = if i % 2 == 0 { pool.clone() } else { other_replica.clone() }; calls.push(tokio::spawn(async move { - entry_limits::admit_call(&p, "tok", &format!("ip:10.0.0.{i}"), &l, now).await.expect("count") + call(&p, &format!("ip:10.0.0.{i}"), &l, now).await })); } let mut admitted = 0; @@ -55,12 +82,12 @@ async fn one_caller_is_limited_alone_and_only_for_the_minute(pool: PgPool) { let l = limits(Some(2), None, None); let now = 600; for _ in 0..2 { - entry_limits::admit_call(&pool, "tok", "ip:a", &l, now).await.unwrap().unwrap(); + call(&pool, "ip:a", &l, now).await.unwrap(); } - let refused = entry_limits::admit_call(&pool, "tok", "ip:a", &l, now).await.unwrap().unwrap_err(); + let refused = call(&pool, "ip:a", &l, now).await.unwrap_err(); assert_eq!(refused.reason, Limited::PerCaller); - entry_limits::admit_call(&pool, "tok", "ip:b", &l, now).await.unwrap().expect("another caller"); - entry_limits::admit_call(&pool, "tok", "ip:a", &l, now + 60).await.unwrap().expect("the next minute"); + call(&pool, "ip:b", &l, now).await.expect("another caller"); + call(&pool, "ip:a", &l, now + 60).await.expect("the next minute"); } /// A fire the entry picked up itself counts once, however often its @@ -91,7 +118,7 @@ async fn an_entry_with_every_limit_off_admits_everything(pool: PgPool) { let l = limits(None, None, None); assert_eq!(l, ResolvedLimits { per_caller_per_minute: None, per_minute: None, at_once: None }); for _ in 0..200 { - entry_limits::admit_call(&pool, "tok", "ip:a", &l, 60).await.unwrap().unwrap(); + call(&pool, "ip:a", &l, 60).await.unwrap(); } } @@ -106,7 +133,7 @@ async fn at_once_slots_hold_under_contention_and_free_up(pool: PgPool) { for i in 0..40 { let p = pool.clone(); takes.push(tokio::spawn(async move { - entry_limits::take_slot(&p, "tok", &format!("execution_id-{i}"), 10, now + 100, now).await.expect("take") + take(&p, &format!("execution_id-{i}"), 10, now + 100, now).await })); } let mut taken = Vec::new(); @@ -117,17 +144,14 @@ async fn at_once_slots_hold_under_contention_and_free_up(pool: PgPool) { } assert_eq!(taken.len(), 10); let first = format!("execution_id-{}", taken[0]); - entry_limits::take_slot(&pool, "tok", &first, 10, now + 100, now) - .await - .unwrap() - .expect("a retry of a run holding a slot keeps it"); - assert!(entry_limits::at_once_full(&pool, "tok", 10, now).await.unwrap().is_some()); + take(&pool, &first, 10, now + 100, now).await.expect("a retry of a run holding a slot keeps it"); + assert!(full(&pool, 10, now).await); entry_limits::release_slot(&pool, &first).await.unwrap(); - assert!(entry_limits::at_once_full(&pool, "tok", 10, now).await.unwrap().is_none()); + assert!(!full(&pool, 10, now).await); // Every remaining slot is a run that never started: past its expiry // none counts, so the next take finds the entry free. - assert!(entry_limits::at_once_full(&pool, "tok", 10, now + 101).await.unwrap().is_none()); - entry_limits::take_slot(&pool, "tok", "late", 1, now + 300, now + 101).await.unwrap().expect("abandoned slots stop counting"); + assert!(!full(&pool, 10, now + 101).await); + take(&pool, "late", 1, now + 300, now + 101).await.expect("abandoned slots stop counting"); } /// A run that started keeps its slot past the unborn expiry until it ends. @@ -141,12 +165,12 @@ async fn a_started_run_keeps_its_slot_until_it_ends(pool: PgPool) { .execute(&pool) .await .unwrap(); - entry_limits::take_slot(&pool, "tok", "born", 1, 10, 0).await.unwrap().unwrap(); - assert!(entry_limits::at_once_full(&pool, "tok", 1, 1_000).await.unwrap().is_some()); + take(&pool, "born", 1, 10, 0).await.unwrap(); + assert!(full(&pool, 1, 1_000).await); entry_limits::sweep(&pool, 1_000).await.unwrap(); - assert!(entry_limits::take_slot(&pool, "tok", "other", 1, 2_000, 1_000).await.unwrap().is_err()); + assert!(take(&pool, "other", 1, 2_000, 1_000).await.is_err()); entry_limits::release_slot(&pool, "born").await.unwrap(); - entry_limits::take_slot(&pool, "tok", "other", 1, 2_000, 1_000).await.unwrap().unwrap(); + take(&pool, "other", 1, 2_000, 1_000).await.unwrap(); } /// A run that ended frees its slot even when its cleanup never released @@ -161,9 +185,9 @@ async fn the_sweep_frees_the_slot_of_a_run_that_ended(pool: PgPool) { .execute(&pool) .await .unwrap(); - entry_limits::take_slot(&pool, "tok", "ended", 1, 10, 0).await.unwrap().unwrap(); + take(&pool, "ended", 1, 10, 0).await.unwrap(); entry_limits::sweep(&pool, 1_000).await.unwrap(); - assert!(entry_limits::at_once_full(&pool, "tok", 1, 1_000).await.unwrap().is_some(), "a live run keeps it"); + assert!(full(&pool, 1, 1_000).await, "a live run keeps it"); sqlx::query( "INSERT INTO exec_event (execution_id, kind, payload_json, created_at) VALUES ('ended', 'execution_completed', '{}', 0)", ) @@ -171,7 +195,7 @@ async fn the_sweep_frees_the_slot_of_a_run_that_ended(pool: PgPool) { .await .unwrap(); entry_limits::sweep(&pool, 1_000).await.unwrap(); - entry_limits::take_slot(&pool, "tok", "next", 1, 2_000, 1_000).await.unwrap().unwrap(); + take(&pool, "next", 1, 2_000, 1_000).await.unwrap(); } /// An address past the bound of refused tokens is blocked for the rest of @@ -179,7 +203,7 @@ async fn the_sweep_frees_the_slot_of_a_run_that_ended(pool: PgPool) { #[sqlx::test] async fn token_guessing_blocks_the_address_for_the_minute(pool: PgPool) { setup(&pool).await; - let edge = entry_limits::EdgeConfig { trusted_proxy_hops: weft_platform_traits::config::ProxyHops { public: 1, outside: 1, domains: 2 }, invalid_tokens_per_minute: Some(3) }; + let edge = edge(Some(3)); let addr: std::net::IpAddr = "203.0.113.9".parse().unwrap(); for _ in 0..3 { assert!(entry_limits::token_guessing_blocked(&pool, &edge, addr, 120).await.unwrap().is_none()); @@ -188,8 +212,32 @@ async fn token_guessing_blocks_the_address_for_the_minute(pool: PgPool) { let blocked = entry_limits::token_guessing_blocked(&pool, &edge, addr, 150).await.unwrap().expect("blocked"); assert_eq!(blocked.retry_after_secs, 30); assert!(entry_limits::token_guessing_blocked(&pool, &edge, addr, 180).await.unwrap().is_none()); - let off = entry_limits::EdgeConfig { trusted_proxy_hops: weft_platform_traits::config::ProxyHops { public: 1, outside: 1, domains: 2 }, invalid_tokens_per_minute: None }; - assert!(entry_limits::token_guessing_blocked(&pool, &off, addr, 150).await.unwrap().is_none()); + assert!(entry_limits::token_guessing_blocked(&pool, &self::edge(None), addr, 150).await.unwrap().is_none()); +} + +/// A live call's admission checks the address the door's guard left to +/// it (`/connect/`), before it counts or takes anything: a blocked +/// address is refused and spends none of the entry's allowance. +#[sqlx::test] +async fn a_blocked_address_is_refused_by_the_calls_admission(pool: PgPool) { + setup(&pool).await; + let edge = edge(Some(1)); + let addr: std::net::IpAddr = "203.0.113.9".parse().unwrap(); + let l = limits(Some(5), None, Some(1)); + let admit = |execution_id: &'static str| { + let (pool, edge, l) = (pool.clone(), edge.clone(), l.clone()); + async move { + entry_limits::admit(&pool, &Admission::call(&edge, Some(addr), "tok", "ip:a", &l, Some((execution_id, 500)), 120)) + .await + .expect("admit") + } + }; + entry_limits::note_invalid_token(&pool, &edge, addr, 120).await.unwrap(); + let refused = admit("first").await.unwrap_err(); + assert_eq!(refused.reason, Limited::InvalidTokens); + assert!(!full(&pool, 1, 120).await, "a blocked call took no slot"); + let unblocked = Admission::call(&self::edge(None), Some(addr), "tok", "ip:a", &l, Some(("second", 500)), 120); + entry_limits::admit(&pool, &unblocked).await.unwrap().expect("with the bound off the address is not checked"); } /// Refusals are counted per entry and limit for `weft status`. @@ -217,7 +265,7 @@ async fn an_ended_unrecorded_run_stops_counting(pool: PgPool) { .execute(&pool) .await .unwrap(); - entry_limits::take_slot(&pool, "tok", "quiet", 1, 10_000, 0).await.unwrap().unwrap(); - assert!(entry_limits::at_once_full(&pool, "tok", 1, 6).await.unwrap().is_none(), "it ended"); - entry_limits::take_slot(&pool, "tok", "next", 1, 10_000, 6).await.unwrap().expect("its slot is free"); + take(&pool, "quiet", 1, 10_000, 0).await.unwrap(); + assert!(!full(&pool, 1, 6).await, "it ended"); + take(&pool, "next", 1, 10_000, 6).await.expect("its slot is free"); } diff --git a/crates/weft-dispatcher/tests/db_install_state.rs b/crates/weft-dispatcher/tests/db_install_state.rs index 6cd30fc2..b1e92f4c 100644 --- a/crates/weft-dispatcher/tests/db_install_state.rs +++ b/crates/weft-dispatcher/tests/db_install_state.rs @@ -52,7 +52,7 @@ fn frontend(project: Uuid, name: &str, repo: Option<&str>) -> weft_core::fronten repo: repo.map(|r| weft_core::frontend::Repository { name: r.into(), id: if r == "me/site" { 1 } else { 2 } }), service: repo.map(|_| format!("fe-{name}")), url: None, - token_id: Uuid::new_v4(), + token_id: repo.is_none().then(Uuid::new_v4), pending_token_ids: Vec::new(), } } diff --git a/crates/weft-dispatcher/tests/db_lifecycle.rs b/crates/weft-dispatcher/tests/db_lifecycle.rs index f5918f27..5fee0d11 100644 --- a/crates/weft-dispatcher/tests/db_lifecycle.rs +++ b/crates/weft-dispatcher/tests/db_lifecycle.rs @@ -153,7 +153,7 @@ async fn execution_birth_and_resume_pin_the_original_image(pool: PgPool) { }; let task = execute_task(id, execution_id, "bin-A", None); seed_project(&projects, id, "bin-B").await; - journal.start_execution(&start, &[], task, None).await.unwrap(); + journal.start_execution(&start, &[], task, false).await.unwrap(); let (kind, binary_hash): (String, Option) = sqlx::query_as( "SELECT kind, binary_hash FROM task WHERE execution_id = $1", @@ -219,7 +219,7 @@ async fn pruning_a_source_waits_for_setup_and_removes_its_unused_bake(pool: PgPo .bind(id).execute(&pool).await.unwrap(); let execution_id = Uuid::new_v4(); let (birth, task) = trigger_setup_birth(id, execution_id); - journal.start_execution(&birth, &[], task, None).await.unwrap(); + journal.start_execution(&birth, &[], task, false).await.unwrap(); assert!(versions.delete_versions(id, &["source".into()]).await.is_err()); let complete = weft_journal::ExecEvent::ExecutionCompleted { execution_id, at_unix: 2 }; journal.record_event(&complete).await.unwrap(); @@ -329,7 +329,7 @@ async fn activation_cleanup_cannot_finish_or_wipe_a_newer_activation(pool: PgPoo assert_eq!(removed[0].token, "old-entry"); assert!(journal.signal_get("old-entry").await.unwrap().is_none()); let (late_birth, late_task) = trigger_setup_birth(id, first); - assert!(journal.start_execution(&late_birth, &[], late_task, Some(first)).await.is_err(), + assert!(journal.start_execution(&late_birth, &[], late_task, true).await.is_err(), "a cancelled activation cannot later start its setup"); assert!(journal.events_log(first).await.unwrap().is_empty()); @@ -562,7 +562,7 @@ async fn trigger_bake_ownership_publication_and_project_cleanup(pool: PgPool) { .bind(id).execute(&pool).await.unwrap(); let first = Uuid::new_v4(); let (birth, task) = trigger_setup_birth(id, first); - journal.start_execution(&birth, &[], task, None).await.unwrap(); + journal.start_execution(&birth, &[], task, false).await.unwrap(); let complete = weft_journal::ExecEvent::ExecutionCompleted { execution_id: first, at_unix: 2 }; journal.record_event(&complete).await.unwrap(); let bake = weft_dispatcher::journal::TriggerBake::from_events(&[birth.clone(), complete]).unwrap().unwrap(); @@ -572,9 +572,9 @@ async fn trigger_bake_ownership_publication_and_project_cleanup(pool: PgPool) { assert!(activations.try_begin_activating(id, &[feed_key()], second, None).await.unwrap().is_ok()); let (second_birth, second_task) = trigger_setup_birth(id, second); let (other_birth, other_task) = trigger_setup_birth(id, Uuid::new_v4()); - assert!(journal.start_execution(&other_birth, &[], other_task, Some(second)).await.is_err(), - "an activation starts only its own setup"); - journal.start_execution(&second_birth, &[], second_task, Some(second)).await.unwrap(); + assert!(journal.start_execution(&other_birth, &[], other_task, true).await.is_err(), + "a setup no activation is setting up is never born as one"); + journal.start_execution(&second_birth, &[], second_task, true).await.unwrap(); let mut entry = governed_entry("baked-entry", id, first); entry.program = Some(bake.program.clone()); entry.source_version = Some(bake.source_version.clone()); @@ -642,7 +642,7 @@ async fn start_execution_birth_is_atomic(pool: PgPool) { payload: json!({}), }; let err = journal - .start_execution(&start, std::slice::from_ref(&kick), task.clone(), None) + .start_execution(&start, std::slice::from_ref(&kick), task.clone(), false) .await .expect_err("missing project must fail the birth"); assert!(format!("{err:#}").contains("has no row"), "{err:?}"); @@ -690,7 +690,7 @@ async fn start_execution_birth_is_atomic(pool: PgPool) { ..task }; journal - .start_execution(&start2, &[], task2.clone(), None) + .start_execution(&start2, &[], task2.clone(), false) .await .expect("birth for a registered project"); let (events2,): (i64,) = @@ -707,7 +707,7 @@ async fn start_execution_birth_is_atomic(pool: PgPool) { assert_eq!((events2, tasks2), (1, 1), "a successful birth commits the event AND the task"); sqlx::query("DELETE FROM task WHERE execution_id = $1").bind(execution_id2.to_string()).execute(&pool).await.unwrap(); - journal.start_execution(&start2, &[], task2.clone(), None).await.unwrap(); + journal.start_execution(&start2, &[], task2.clone(), false).await.unwrap(); let births: i64 = sqlx::query_scalar("SELECT COUNT(*) FROM exec_event WHERE execution_id = $1 AND kind = 'execution_started'") .bind(execution_id2.to_string()).fetch_one(&pool).await.unwrap(); let tasks: i64 = sqlx::query_scalar("SELECT COUNT(*) FROM task WHERE execution_id = $1") @@ -762,7 +762,7 @@ async fn an_unrecorded_run_is_born_unjournaled_and_forgotten_or_recorded(pool: P // Born: the execution row and the task, no journal row. let gone = weft_core::ExecutionId::new_v4(); let (start, kick, task) = birth(gone); - journal.start_execution(&start, std::slice::from_ref(&kick), task, None).await.unwrap(); + journal.start_execution(&start, std::slice::from_ref(&kick), task, false).await.unwrap(); assert_eq!(rows(gone).await, 0, "an unrecorded birth writes no journal row"); assert_eq!(kind(gone).await.as_deref(), Some("unrecorded")); let payload: serde_json::Value = sqlx::query_scalar("SELECT payload FROM task WHERE execution_id = $1") @@ -782,7 +782,7 @@ async fn an_unrecorded_run_is_born_unjournaled_and_forgotten_or_recorded(pool: P // A failed one, written afterwards, is an ordinary listed run. let failed = weft_core::ExecutionId::new_v4(); let (start, kick, task) = birth(failed); - journal.start_execution(&start, std::slice::from_ref(&kick), task, None).await.unwrap(); + journal.start_execution(&start, std::slice::from_ref(&kick), task, false).await.unwrap(); let mut record = vec![start, kick, weft_journal::ExecEvent::ExecutionFailed { execution_id: failed, error: "boom".into(), at_unix: 2 }]; if let weft_journal::ExecEvent::ExecutionStarted { run_kind, .. } = &mut record[0] { *run_kind = weft_core::exec::RunKind::Execution; @@ -798,7 +798,7 @@ async fn an_unrecorded_run_is_born_unjournaled_and_forgotten_or_recorded(pool: P // addressed by it. let paid = weft_core::ExecutionId::new_v4(); let (start, kick, task) = birth(paid); - journal.start_execution(&start, std::slice::from_ref(&kick), task, None).await.unwrap(); + journal.start_execution(&start, std::slice::from_ref(&kick), task, false).await.unwrap(); weft_journal::record_events(&pool, &[weft_journal::ExecEvent::CostReported { execution_id: paid, node_id: "llm".into(), frames: vec![], cost_id: "c".into(), service: "llm".into(), model: None, amount_usd: Some(0.1), billed: true, origin: weft_core::CredentialOwner::Author, @@ -831,7 +831,7 @@ async fn an_unrecorded_run_is_live_while_a_worker_holds_its_task(pool: PgPool) { run_class: weft_core::run_class::RunClass::Short, }; let task = execute_task(id, execution_id, "bin-A", Some(std::slice::from_ref(&start))); - journal.start_execution(&start, &[], task, None).await.unwrap(); + journal.start_execution(&start, &[], task, false).await.unwrap(); sqlx::query("INSERT INTO execution_tag (execution_id, tag, tagged_at_unix) VALUES ($1, 'poll', 1)") .bind(execution_id.to_string()).execute(&pool).await.unwrap(); @@ -882,7 +882,7 @@ async fn quiesce_waits_until_no_run_of_the_project_is_live(pool: PgPool) { run_class: weft_core::run_class::RunClass::Short, }; let task = execute_task(id, execution_id, "bin-A", Some(std::slice::from_ref(&start))); - journal.start_execution(&start, &[], task, None).await.unwrap(); + journal.start_execution(&start, &[], task, false).await.unwrap(); let watch = weft_task_store::pg_signal::PgSignalWatch::start(&pool.connect_options(), weft_dispatcher::take_down::RUN_ENDING_CHANNELS) .await @@ -927,7 +927,7 @@ async fn forgetting_an_unrecorded_run_announces_its_ending_at_the_commit(pool: P run_class: weft_core::run_class::RunClass::Short, }; let task = execute_task(id, execution_id, "bin-A", Some(std::slice::from_ref(&start))); - journal.start_execution(&start, &[], task, None).await.unwrap(); + journal.start_execution(&start, &[], task, false).await.unwrap(); sqlx::query("UPDATE task SET status = 'claimed', claimed_by = 'worker-a', claimed_until_unix = $2 WHERE execution_id = $1") .bind(execution_id.to_string()).bind(weft_dispatcher::lease::now_unix() + 60).execute(&pool).await.unwrap(); weft_journal::record_events(&pool, &[weft_journal::ExecEvent::CostReported { @@ -2154,6 +2154,149 @@ async fn a_held_row_coming_or_going_wakes_the_holder_sizing(pool: PgPool) { assert!(woken(&mut heard).await, "and one forgotten too"); } +/// The rows a dispatcher holds in memory (`weft_dispatcher::held`) say +/// when they change, naming whose: a tenant's public entries and the +/// activations that arm them, a project's worker levers, a project's infra +/// copies. A signal that is no public entry is silent: a run waiting on a +/// reply writes one, and a tenant's routes are not read again for it. +#[sqlx::test] +async fn the_held_rows_announce_their_changes(pool: PgPool) { + use weft_dispatcher::held::{INFRA_STATUS_CHANNEL, ROUTES_CHANNEL, WORKER_SETTINGS_CHANNEL}; + let (journal, projects) = setup(&pool).await; + let id = Uuid::new_v4(); + seed_project(&projects, id, "bin-A").await; + static CHANNELS: &[&str] = &[ROUTES_CHANNEL, WORKER_SETTINGS_CHANNEL, INFRA_STATUS_CHANNEL]; + let watch = weft_task_store::pg_signal::PgSignalWatch::start(&pool.connect_options(), CHANNELS).await.unwrap(); + let mut heard = watch.subscribe(); + async fn heard_on(heard: &mut weft_task_store::pg_signal::Subscription, channel: &'static str, key: &str) -> bool { + let deadline = tokio::time::Instant::now() + std::time::Duration::from_secs(2); + heard.woken_before(deadline, |c, p| c == channel && p == key).await.unwrap() + } + let project = id.to_string(); + + journal.signal_insert(&entry_at("callback", id)).await.unwrap(); + assert!(!heard_on(&mut heard, ROUTES_CHANNEL, TENANT).await, "a signal that is no public entry is silent"); + let route = SignalRegistration { + surface_kind: "public_entry".into(), + mount_path: Some(format!("/{TENANT}/notes")), + ..entry_at("route", id) + }; + journal.signal_insert(&route).await.unwrap(); + assert!(heard_on(&mut heard, ROUTES_CHANNEL, TENANT).await, "a public entry coming names its tenant"); + sqlx::query( + "INSERT INTO trigger_activation (project_id, trigger, status, accepting_fires, fires_visible_to_consumers, updated_at) \ + VALUES ($1, 'route', 'active', TRUE, TRUE, 0)", + ) + .bind(id) + .execute(&pool) + .await + .unwrap(); + assert!(heard_on(&mut heard, ROUTES_CHANNEL, TENANT).await, "an activation coming"); + sqlx::query("UPDATE trigger_activation SET status = 'inactive' WHERE project_id = $1").bind(id).execute(&pool).await.unwrap(); + assert!(heard_on(&mut heard, ROUTES_CHANNEL, TENANT).await, "an activation taken down"); + journal.signal_remove_many(&["route".to_string()]).await.unwrap(); + assert!(heard_on(&mut heard, ROUTES_CHANNEL, TENANT).await, "a public entry going"); + + sqlx::query("UPDATE project SET worker_settings_json = '{\"minInstances\": 1}'::jsonb WHERE id = $1") + .bind(id) + .execute(&pool) + .await + .unwrap(); + assert!(heard_on(&mut heard, WORKER_SETTINGS_CHANNEL, &project).await, "a project's worker levers changing name it"); + + sqlx::query("INSERT INTO infra_node (project_id, node_id, status) VALUES ($1, 'db', 'provisioning')") + .bind(id) + .execute(&pool) + .await + .unwrap(); + assert!(heard_on(&mut heard, INFRA_STATUS_CHANNEL, &project).await, "an infra copy coming names its project"); + sqlx::query("UPDATE infra_node SET status = 'running' WHERE project_id = $1").bind(id).execute(&pool).await.unwrap(); + assert!(heard_on(&mut heard, INFRA_STATUS_CHANNEL, &project).await, "and its status changing"); +} + +/// A live run's birth, as a caller's handshake makes it: its +/// `ExecutionStarted` and its execute task, waiting for the caller until +/// 1 000. +fn live_birth(project: Uuid, execution_id: weft_core::ExecutionId) -> (weft_journal::ExecEvent, weft_task_store::tasks::NewTask) { + let start = weft_journal::ExecEvent::ExecutionStarted { + execution_id, + project_id: project, + entry_node: "entry".into(), + phase: weft_core::context::Phase::Fire, + definition_hash: Some("def-1".into()), + program: None, source_version: None, run_kind: weft_core::exec::RunKind::Execution, + subgraph: None, + seed: None, + instance: None, fired_trigger: None, instance_values: Default::default(), picks: Default::default(), at_unix: 0, + run_class: weft_core::run_class::RunClass::Short, + }; + let task = weft_dispatcher::task_kinds::execute::execution_task_spec(weft_dispatcher::task_kinds::execute::ExecutionTask { + kind: weft_task_store::TaskKind::Execute, + project_id: project, + execution_id, + definition_hash: "def-1", + binary_hash: "bin-A", + tenant_id: TENANT, + run_class: weft_core::run_class::RunClass::Short, + live_connection: Some(weft_task_store::kinds::LiveConnectionStart { + spec: weft_core::primitive::SignalSpec::of_kind("route", json!({})), + request: Default::default(), + arrive_by: Some(1_000), + fired: None, + }), + unrecorded_birth: None, + }) + .unwrap(); + (start, task) +} + +/// A live call's admission at the entry `tok`, which takes `at_once` runs +/// at once, taking a slot for `execution_id` that holds until 1 000. +fn live_admission(at_once: u32, execution_id: weft_core::ExecutionId) -> weft_dispatcher::entry_limits::Admission { + let limits = weft_core::signal::EntryLimits { per_caller_per_minute: Some(0), per_minute: Some(0), at_once: Some(at_once) }.resolve(); + let edge = weft_dispatcher::entry_limits::EdgeConfig { + trusted_proxy_hops: weft_platform_traits::config::ProxyHops { public: 1, outside: 1, domains: 2 }, + invalid_tokens_per_minute: None, + }; + weft_dispatcher::entry_limits::Admission::call(&edge, None, "tok", "ip:a", &limits, Some((&execution_id.to_string(), 1_000)), 0) +} + +/// A live call its entry's limits refuse is never born: the refusal and +/// the birth are one call to the database, so nothing of the run (journal, +/// execution, task, slot) is written, and the refusal is counted. A run +/// already born is left as it is when its birth is asked for again. +#[sqlx::test] +async fn a_live_call_refused_at_its_limit_is_never_born(pool: PgPool) { + let (journal, projects) = setup(&pool).await; + let project = Uuid::new_v4(); + seed_project(&projects, project, "bin-A").await; + let first = weft_core::ExecutionId::new_v4(); + let (start, task) = live_birth(project, first); + journal.admit_and_start_execution(&live_admission(1, first), &start, &[], task.clone()).await.unwrap().expect("room for one"); + journal + .admit_and_start_execution(&live_admission(1, first), &start, &[], task) + .await + .unwrap() + .expect("a retry of a run already born collapses onto it, slot and all"); + + let second = weft_core::ExecutionId::new_v4(); + let (start, task) = live_birth(project, second); + let refused = journal.admit_and_start_execution(&live_admission(1, second), &start, &[], task).await.unwrap().unwrap_err(); + assert_eq!(refused.reason, weft_dispatcher::entry_limits::Limited::AtOnce); + for table in ["exec_event", "execution", "task", "entry_slot"] { + let (n,): (i64,) = sqlx::query_as(&format!("SELECT COUNT(*)::bigint FROM {table} WHERE execution_id = $1")) + .bind(second.to_string()) + .fetch_one(&pool) + .await + .unwrap(); + assert_eq!(n, 0, "{table} holds a run its entry refused"); + } + assert_eq!( + weft_dispatcher::entry_limits::recent_refusals(&pool, "tok", 0).await.unwrap(), + vec![(weft_dispatcher::entry_limits::Limited::AtOnce, 1)] + ); +} + /// A live run born at its caller's handshake whose caller never came is /// erased whole once their ticket expires: its journal, its execution, its /// task and its entry slot, so nothing of it is left. One the caller did @@ -2163,38 +2306,7 @@ async fn a_live_run_whose_caller_never_came_leaves_nothing(pool: PgPool) { let (journal, projects) = setup(&pool).await; let project = Uuid::new_v4(); seed_project(&projects, project, "bin-A").await; - let born = |execution_id: weft_core::ExecutionId| { - let start = weft_journal::ExecEvent::ExecutionStarted { - execution_id, - project_id: project, - entry_node: "entry".into(), - phase: weft_core::context::Phase::Fire, - definition_hash: Some("def-1".into()), - program: None, source_version: None, run_kind: weft_core::exec::RunKind::Execution, - subgraph: None, - seed: None, - instance: None, fired_trigger: None, instance_values: Default::default(), picks: Default::default(), at_unix: 0, - run_class: weft_core::run_class::RunClass::Short, - }; - let task = weft_dispatcher::task_kinds::execute::execution_task_spec(weft_dispatcher::task_kinds::execute::ExecutionTask { - kind: weft_task_store::TaskKind::Execute, - project_id: project, - execution_id, - definition_hash: "def-1", - binary_hash: "bin-A", - tenant_id: TENANT, - run_class: weft_core::run_class::RunClass::Short, - live_connection: Some(weft_task_store::kinds::LiveConnectionStart { - spec: weft_core::primitive::SignalSpec::of_kind("route", json!({})), - request: Default::default(), - arrive_by: Some(1_000), - fired: None, - }), - unrecorded_birth: None, - }) - .unwrap(); - (start, task) - }; + let born = |execution_id: weft_core::ExecutionId| live_birth(project, execution_id); let count = |table: &'static str, execution_id: weft_core::ExecutionId| { let pool = pool.clone(); async move { @@ -2207,14 +2319,18 @@ async fn a_live_run_whose_caller_never_came_leaves_nothing(pool: PgPool) { } }; + // Born as a live call is: admitted at its entry (a slot taken for it) + // in the same commit as its birth. + let admitted = |execution_id: weft_core::ExecutionId| live_admission(5, execution_id); + let absent = weft_core::ExecutionId::new_v4(); let (start, task) = born(absent); - journal.start_execution(&start, &[], task, None).await.unwrap(); - weft_dispatcher::entry_limits::take_slot(&pool, "tok", &absent.to_string(), 5, 1_000, 0).await.unwrap().unwrap(); + journal.admit_and_start_execution(&admitted(absent), &start, &[], task).await.unwrap().expect("admitted"); + assert_eq!(count("entry_slot", absent).await, 1, "the slot is taken with the birth"); let present = weft_core::ExecutionId::new_v4(); let (start, task) = born(present); - journal.start_execution(&start, &[], task, None).await.unwrap(); + journal.start_execution(&start, &[], task, false).await.unwrap(); let claimed = weft_task_store::tasks::claim_one( &pool, "worker-a", @@ -2246,8 +2362,7 @@ async fn a_live_run_whose_caller_never_came_leaves_nothing(pool: PgPool) { // the route's slot free; the run a caller did reach is never taken. let unreached = weft_core::ExecutionId::new_v4(); let (start, task) = born(unreached); - journal.start_execution(&start, &[], task, None).await.unwrap(); - weft_dispatcher::entry_limits::take_slot(&pool, "tok", &unreached.to_string(), 5, 1_000, 0).await.unwrap().unwrap(); + journal.admit_and_start_execution(&admitted(unreached), &start, &[], task).await.unwrap().expect("admitted"); const UNREACHED: weft_task_store::tasks::UnclaimedLiveRun = weft_task_store::tasks::UnclaimedLiveRun::NeverPassedOn; assert!(!journal.erase_unclaimed_live_run(unreached, weft_task_store::tasks::UnclaimedLiveRun::PastDeadline { now: 999 }).await.unwrap(), "its deadline has not passed"); assert!(journal.erase_unclaimed_live_run(unreached, UNREACHED).await.unwrap()); diff --git a/crates/weft-engine/src/context.rs b/crates/weft-engine/src/context.rs index ada34507..084aff2b 100644 --- a/crates/weft-engine/src/context.rs +++ b/crates/weft-engine/src/context.rs @@ -64,6 +64,12 @@ fn frames_dedup_key(frames: &weft_core::frames::LoopFrames) -> Result, pub tasks: Arc, + /// Where a node's spend is recorded (`metering::CostSink`): the task + /// store, never behind a drive's journal ([`JournalFirst`]). A cost + /// has no place in the run's order, can be recorded after its run's + /// drive ended, and money already spent is recorded even when the + /// run's journal failed. + pub costs: Arc, pub infra: Arc, /// Broker client for infra applied-state reads + apply enqueue. /// Used by the loop driver during `Phase::InfraSetup` to make the @@ -113,9 +119,11 @@ impl EngineClients { /// code. pub fn from_broker(broker_url: &str, token: weft_broker_client::TokenSource) -> Self { let url = || broker_url.to_string(); + let tasks: Arc = weft_broker_client::BrokerTaskStoreClient::new(url(), token.clone()); Self { journal: weft_broker_client::BrokerJournalClient::new(url(), token.clone()), - tasks: weft_broker_client::BrokerTaskStoreClient::new(url(), token.clone()), + costs: tasks.clone(), + tasks, infra: weft_broker_client::BrokerInfraClient::new(url(), token.clone()), infra_state: weft_broker_client::BrokerInfraStateClient::new(url(), token.clone()), project: weft_broker_client::BrokerProjectClient::new(url(), token.clone()), @@ -1192,144 +1200,271 @@ fn type_accepts(declared: &WeftType, value: &Value) -> bool { declared.accepts_runtime_value(value) } -/// Journal write for teardown paths that cannot propagate errors. -/// `DriveJournal` also marks the drive for exit on failure. +/// Journal write for teardown paths that cannot propagate errors: a +/// failure is logged. Through a `DriveJournal` the write is only handed +/// over, so a failure here means the drive was already poisoned or its +/// sending task is gone; the run stops on the poison itself. pub async fn record_from_replica(journal: &dyn JournalClient, event: ExecEvent, replica: &str) { if let Err(e) = journal.record_event(&event, Some(replica)).await { - tracing::error!( - target: "weft_engine::journal", - error = %e, - "journal write failed; drive is now poisoned and the worker will exit" - ); + tracing::error!(target: "weft_engine::journal", error = %e, "a journal write on a teardown path failed"); } } -/// The journal a drive writes through: it holds the driver's own rows -/// back and sends them together, and it latches the first failed write. +/// The journal a drive writes through: every row it is handed goes to the +/// database in the background, in the order handed, and nothing on the +/// run's path waits for it. /// -/// **Batching.** A drive writes several rows per step (a firing's -/// start, its emissions, its end, the next firing's start), and each -/// write is a round trip to the database. The driver DEFERS its rows -/// ([`Self::deferring`]): they wait here and go out together at the -/// points where they must be durable, in one request, in the order -/// written. Those points are the driver's own ([`Self::flush`] before a -/// node's body runs, so its start is on record before any of its side -/// effects, and before the driver waits for news), and every write that -/// is not deferred: a node's own writes through its ctx go straight out, -/// carrying whatever was held ahead of them, so the journal's order is -/// always the order things happened. A read through the drive journal -/// flushes first, so it never reads past rows the drive wrote; the -/// driver's own idle read does not (see [`Deferring`]). +/// **Order and batching.** One task per drive sends the rows: whatever has +/// queued while the previous write was on its way goes out as the next +/// write, so a burst of rows is one round trip and a quiet run sends each +/// row as it comes. The driver's rows and its nodes' own (through their +/// ctx) share the queue, so the journal's order is the order things +/// happened. +/// +/// **Waiting for it.** [`Self::flush`] answers once everything handed so +/// far is on record. Only three kinds of place ask: before the run ends or +/// parks (its terminal, a stall, a cancel walk), since what comes after +/// reads the journal from another worker; before a read (a read never runs +/// past rows the drive wrote); and before a task goes out, since a task +/// has another writer act on the run ([`JournalFirst`]). A live call never waits on its own bookkeeping: a run +/// with a live caller is never resumed elsewhere, and a crash fails it +/// whatever the journal holds. A run that is resumed redoes the steps whose +/// rows a crash lost, which is the at-least-once rule every step already +/// runs under (`ctx.run`). /// /// **Poison.** A failed write means the journal is now behind what the -/// live worker believes happened. Continuing to drive on that -/// divergence makes every later refold (stall refetch, crash resume) -/// rebuild a different world: a body whose `NodeStarted` was lost but -/// whose `PortEmitted` rows landed re-runs and double-spends. The drive -/// loop checks [`Self::is_poisoned`] every iteration and stops with an -/// error, which `run_one_execution` journals as the run's Failed terminal. +/// live worker believes happened. Continuing to drive on that divergence +/// makes every later refold (stall refetch, crash resume) rebuild a +/// different world: a body whose `NodeStarted` was lost but whose +/// `PortEmitted` rows landed re-runs and double-spends. So the first +/// failed write stops all sending, every later [`Self::flush`] and write +/// fails naming it, and the drive loop checks [`Self::is_poisoned`] every +/// iteration and stops with an error, which `run_one_execution` journals as +/// the run's Failed terminal; a drive idling on a long node hears it at +/// once ([`Self::poisoned`]). A part of a write that never reached the +/// broker (the connection could not be made) is sent again, part by part +/// and for a minute at most each, since it cannot have landed; any other +/// failure poisons at once, since sending a row twice would record it +/// twice. /// /// The bus pump deliberately keeps the UNwrapped client: bus rows are /// the inspector's replay trail, and their failures already degrade /// per-bus without killing the worker. pub struct DriveJournal { inner: Arc, - /// The deferred rows, and the replica they are written under. - held: std::sync::Mutex<(Vec, Option)>, - /// Held across take-and-send, so two flushes never overtake each - /// other on the wire. - sending: tokio::sync::Mutex<()>, - poisoned: std::sync::atomic::AtomicBool, + queue: tokio::sync::mpsc::UnboundedSender, + /// Why the drive is poisoned, once it is (see the type's docs). + failure: Arc, + /// The replica the drive writes under, fixed by its first row. + replica: std::sync::Mutex>>, +} + +/// Why a drive's journal can no longer be trusted, once it can't: the +/// first reason is kept, and whoever waits for it hears it at once. +struct Poison { + why: std::sync::Mutex>, + heard: tokio::sync::watch::Sender, } +impl Poison { + fn new() -> Self { + Self { why: std::sync::Mutex::new(None), heard: tokio::sync::watch::Sender::new(false) } + } + + fn set(&self, why: String) { + self.why.lock().expect("drive journal").get_or_insert(why); + self.heard.send_replace(true); + } + + fn get(&self) -> Option { + self.why.lock().expect("drive journal").clone() + } + + async fn wait(&self) { + let _ = self.heard.subscribe().wait_for(|poisoned| *poisoned).await; + } +} + +/// What the sending task is handed, in order. +enum Queued { + Rows(Vec, Option), + /// Answered once every row queued before it is on record. + Written(tokio::sync::oneshot::Sender>), +} + + impl DriveJournal { + /// The drive journal over `inner`, its sending task started. The task + /// ends once every handle to the queue is gone, after sending what was + /// left in it. pub fn wrap(inner: Arc) -> Arc { - Arc::new(Self { - inner, - held: std::sync::Mutex::new((Vec::new(), None)), - sending: tokio::sync::Mutex::new(()), - poisoned: std::sync::atomic::AtomicBool::new(false), - }) + let (queue, pending) = tokio::sync::mpsc::unbounded_channel(); + let failure = Arc::new(Poison::new()); + tokio::spawn(send_in_order(inner.clone(), pending, failure.clone())); + Arc::new(Self { inner, queue, failure, replica: std::sync::Mutex::new(None) }) } - /// The view the driver writes through: its rows are held until the - /// next flush. + /// The view the driver writes through. Its writes are this journal's; + /// its reads go straight to the journal beneath (see [`Deferring`]). pub fn deferring(self: &Arc) -> Deferring { Deferring(self.clone()) } /// A write failed since the drive began. pub fn is_poisoned(&self) -> bool { - self.poisoned.load(std::sync::atomic::Ordering::Acquire) + self.failure.get().is_some() } - /// Send every held row now. Refused once the drive is poisoned (see - /// `send`): a drive whose journal is already behind starts nothing more. - pub async fn flush(&self) -> anyhow::Result<()> { - self.send(&[], None).await + /// Answers once a write of this drive failed: what a drive waiting on + /// something else (a long node, news) also waits on, so it stops at + /// once. + pub async fn poisoned(&self) { + self.failure.wait().await } - fn hold(&self, events: &[ExecEvent], replica: Option<&str>) -> anyhow::Result<()> { - let mut held = self.held.lock().expect("drive journal"); - if !held.0.is_empty() && held.1.as_deref() != replica { - drop(held); - self.poisoned.store(true, std::sync::atomic::Ordering::Release); - anyhow::bail!("one drive wrote under two replicas; its journal can no longer be trusted"); + fn poisoned_error(&self) -> Option { + self.failure.get().map(|why| anyhow::anyhow!("an earlier journal write of this drive failed: {why}")) + } + + /// Answer once every row handed so far is on record. Refused once the + /// drive is poisoned: a drive whose journal is already behind starts + /// nothing more. + pub async fn flush(&self) -> anyhow::Result<()> { + if let Some(e) = self.poisoned_error() { + return Err(e); } - held.0.extend_from_slice(events); - held.1 = replica.map(str::to_string); - Ok(()) + let (answered, answer) = tokio::sync::oneshot::channel(); + self.queue + .send(Queued::Written(answered)) + .map_err(|_| anyhow::anyhow!("the drive journal's sending task is gone"))?; + answer + .await + .map_err(|_| anyhow::anyhow!("the drive journal's sending task ended before answering"))? + .map_err(|why| anyhow::anyhow!("an earlier journal write of this drive failed: {why}")) } - /// The held rows, then `events`, in one write. - async fn send(&self, events: &[ExecEvent], replica: Option<&str>) -> anyhow::Result<()> { - // Nothing more goes out once a write failed, a node's own write - // included: a row landing after a lost one is the gap a refold - // would rebuild a different world from. - let _sending = self.sending.lock().await; - // Asked holding the lock: a write that was queued behind the one - // that failed must not go out after it. - anyhow::ensure!(!self.is_poisoned(), "an earlier journal write of this drive failed"); - let (mut batch, held_replica) = { - let mut held = self.held.lock().expect("drive journal"); - // Checked before the held rows are taken, so a refused write - // drops nothing that was already held. - if !held.0.is_empty() && !events.is_empty() && held.1.as_deref() != replica { - drop(held); - self.poisoned.store(true, std::sync::atomic::Ordering::Release); - anyhow::bail!("one drive wrote under two replicas; its journal can no longer be trusted"); + /// Hand `events` to the sending task, at once. + fn enqueue(&self, events: &[ExecEvent], replica: Option<&str>) -> anyhow::Result<()> { + if let Some(e) = self.poisoned_error() { + return Err(e); + } + { + let mut fixed = self.replica.lock().expect("drive journal"); + match fixed.as_ref() { + Some(first) if first.as_deref() != replica => { + drop(fixed); + let why = "one drive wrote under two replicas; its journal can no longer be trusted"; + self.failure.set(why.to_string()); + anyhow::bail!(why); + } + Some(_) => {} + None => *fixed = Some(replica.map(str::to_string)), } - std::mem::take(&mut *held) - }; - let replica = if batch.is_empty() { replica.map(str::to_string) } else { held_replica }; - batch.extend_from_slice(events); - if batch.is_empty() { + } + if events.is_empty() { return Ok(()); } - self.inner.record_events(&batch, replica.as_deref()).await.map_err(|e| { - self.poisoned.store(true, std::sync::atomic::Ordering::Release); - let kinds: Vec<&str> = batch.iter().map(ExecEvent::kind_str).collect(); - e.context(format!("{} journal row(s) were not written: {}", batch.len(), kinds.join(", "))) - }) + self.queue + .send(Queued::Rows(events.to_vec(), replica.map(str::to_string))) + .map_err(|_| anyhow::anyhow!("the drive journal's sending task is gone")) } /// The journal this one writes to. The run's terminal goes straight - /// there, once everything held has gone out: it is retried on its own, - /// and its failing is the terminal's, never a lost row that would + /// there, once everything handed has gone out: it is retried on its + /// own, and its failing is the terminal's, never a lost row that would /// poison the drive and make a run that completed read as failed. pub fn beneath(&self) -> &dyn JournalClient { self.inner.as_ref() } } +/// The drive journal's sending task: everything queued while the previous +/// write was on its way goes out as one write, in order, and each +/// `Written` is answered once the rows ahead of it are on record. +async fn send_in_order( + inner: Arc, + mut pending: tokio::sync::mpsc::UnboundedReceiver, + failure: Arc, +) { + while let Some(first) = pending.recv().await { + let mut batch: Vec = Vec::new(); + let mut replica: Option = None; + let mut waiting = Vec::new(); + let mut next = Some(first); + while let Some(item) = next { + match item { + Queued::Rows(rows, r) => { + replica = r; + batch.extend(rows); + } + Queued::Written(answer) => waiting.push(answer), + } + next = pending.try_recv().ok(); + } + if failure.get().is_none() && !batch.is_empty() { + if let Err(why) = send(inner.as_ref(), &batch, replica.as_deref()).await { + tracing::error!(target: "weft_engine::journal", error = %why, "a drive's journal write failed; the run stops"); + failure.set(why); + } + } + let outcome = failure.get().map_or(Ok(()), Err); + for answer in waiting { + let _ = answer.send(outcome.clone()); + } + } +} + +/// How long a part of a write that never reached the journal is sent +/// again, at first and at most between tries, and in all for that part: +/// the broker is the install's own service, and one out of reach this +/// long is down, so the run stops on it like on any failed write. +const RESEND_FIRST: Duration = Duration::from_millis(100); +const RESEND_LONGEST: Duration = Duration::from_secs(5); +const RESEND_FOR: Duration = Duration::from_secs(60); + +/// Write `batch`, one part at a time, cut the way the journal sends it +/// (`JournalClient::parts`): a part that never reached the journal +/// (`JournalClient::never_reached`) is sent again, since it cannot have +/// landed, for up to [`RESEND_FOR`] each; any other failure is final, +/// since the part may have landed and sending it twice would record its +/// rows twice. Sending part by part is what keeps a part that landed from +/// being sent again with the rest. +async fn send(inner: &dyn JournalClient, batch: &[ExecEvent], replica: Option<&str>) -> Result<(), String> { + let not_written = |rows: &[ExecEvent], e: &anyhow::Error| { + let kinds: Vec<&str> = rows.iter().map(ExecEvent::kind_str).collect(); + format!("{} journal row(s) were not written ({}): {e:#}", rows.len(), kinds.join(", ")) + }; + let chunks = inner.parts(batch).map_err(|e| not_written(batch, &e))?; + for chunk in chunks { + let started = tokio::time::Instant::now(); + let mut wait = RESEND_FIRST; + loop { + match inner.record_events(chunk, replica).await { + Ok(()) => break, + Err(e) if inner.never_reached(&e) && started.elapsed() < RESEND_FOR => { + tracing::warn!( + target: "weft_engine::journal", + error = %format!("{e:#}"), rows = chunk.len(), retry_in_ms = wait.as_millis() as u64, + "a journal write did not reach the broker; sending it again" + ); + tokio::time::sleep(wait).await; + wait = (wait * 2).min(RESEND_LONGEST); + } + Err(e) => return Err(not_written(chunk, &e)), + } + } + } + Ok(()) +} + #[async_trait::async_trait] impl JournalClient for DriveJournal { async fn record_event(&self, event: &ExecEvent, replica: Option<&str>) -> anyhow::Result<()> { - self.send(std::slice::from_ref(event), replica).await + self.enqueue(std::slice::from_ref(event), replica) } async fn record_events(&self, events: &[ExecEvent], replica: Option<&str>) -> anyhow::Result<()> { - self.send(events, replica).await + self.enqueue(events, replica) } async fn raw_rows_after( @@ -1358,23 +1493,98 @@ impl JournalClient for DriveJournal { } } -/// [`DriveJournal`] as the driver writes through it: a write is held for -/// the next flush. Its reads go straight to the journal without a flush: -/// the driver flushes before every wait itself (`flush_or_stop`), and its -/// idle read is a future kept across turns of the loop, which must never -/// queue for the send lock, since a queued lock is handed over even while -/// nobody polls its future and the driver's next flush would then wait on -/// it forever. +/// A client a drive and its nodes reach another writer through (the task +/// store, the steering of runs), wrapped so every journal row handed to +/// the drive so far is on record before the other writer is asked to act. +/// A task hands the run to the dispatcher (a signal it registers and +/// journals, a call it makes), a tag is written beside the run's own rows, +/// and a stop by tag may cancel this very run: none of them may land a +/// row of this run ahead of ones the drive was handed first. +pub struct JournalFirst { + pub inner: Arc, + pub journal: Arc, +} + +#[async_trait::async_trait] +impl ExecutionSteeringClient for JournalFirst { + async fn tag_execution(&self, execution_id: ExecutionId, tags: Vec, replica: &str) -> anyhow::Result<()> { + self.journal.flush().await?; + self.inner.tag_execution(execution_id, tags, replica).await + } + + async fn stop_tagged( + &self, + execution_id: ExecutionId, + tag: String, + stop_self: weft_core::StopSelf, + replica: &str, + ) -> anyhow::Result { + self.journal.flush().await?; + self.inner.stop_tagged(execution_id, tag, stop_self, replica).await + } +} + +#[async_trait::async_trait] +impl TaskStoreClient for JournalFirst { + async fn enqueue_dedup(&self, spec: task_store::NewTask) -> anyhow::Result { + self.journal.flush().await?; + self.inner.enqueue_dedup(spec).await + } + + async fn wait_for_terminal(&self, task_id: uuid::Uuid, timeout: Duration) -> anyhow::Result { + self.inner.wait_for_terminal(task_id, timeout).await + } + + async fn claim_one( + &self, + replica: &str, + filter: task_store::ClaimFilter, + wait: Duration, + ) -> anyhow::Result> { + self.inner.claim_one(replica, filter, wait).await + } + + async fn heartbeat(&self, task_id: uuid::Uuid, replica: &str) -> anyhow::Result { + self.inner.heartbeat(task_id, replica).await + } + + async fn requeue(&self, task_id: uuid::Uuid, replica: &str) -> anyhow::Result { + self.inner.requeue(task_id, replica).await + } + + async fn complete(&self, task_id: uuid::Uuid, replica: &str, result: Value) -> anyhow::Result<()> { + self.inner.complete(task_id, replica, result).await + } + + async fn fail(&self, task_id: uuid::Uuid, replica: &str, error: String) -> anyhow::Result<()> { + self.inner.fail(task_id, replica, error).await + } + + async fn wait_cancels( + &self, + project_id: uuid::Uuid, + execution_ids: Vec, + wait: Duration, + ) -> anyhow::Result> { + self.inner.wait_cancels(project_id, execution_ids, wait).await + } +} + +/// [`DriveJournal`] as the driver holds it: its writes are the drive +/// journal's, and its reads go straight to the journal beneath without +/// waiting for what is still being sent. The driver's idle read follows +/// the run from its own last row, so its own rows arriving later read as +/// rows it already applied; waiting there would only hold the loop. pub struct Deferring(Arc); #[async_trait::async_trait] impl JournalClient for Deferring { async fn record_event(&self, event: &ExecEvent, replica: Option<&str>) -> anyhow::Result<()> { - self.0.hold(std::slice::from_ref(event), replica) + self.0.enqueue(std::slice::from_ref(event), replica) } async fn record_events(&self, events: &[ExecEvent], replica: Option<&str>) -> anyhow::Result<()> { - self.0.hold(events, replica) + self.0.enqueue(events, replica) } async fn raw_rows_after( @@ -2804,7 +3014,7 @@ impl ContextHandle for RunnerHandle { )) })?; let sink = Arc::new(crate::metering::CostSink { - tasks: self.clients.tasks.clone(), + tasks: self.clients.costs.clone(), pending: self.clients.pending_costs.clone(), open_charges: self.clients.open_charges.clone(), project_id: self.project_id, @@ -3485,6 +3695,7 @@ mod replay_tests { let clients = EngineClients { journal: Arc::new(NoopJournal), tasks: Arc::new(NoopTaskStore), + costs: Arc::new(NoopTaskStore), infra: Arc::new(NoopInfra), infra_state: Arc::new(NoopInfraState), project: Arc::new(NoopProject), @@ -3555,6 +3766,7 @@ mod replay_tests { let clients = EngineClients { journal: Arc::new(NoopJournal), tasks: Arc::new(NoopTaskStore), + costs: Arc::new(NoopTaskStore), infra: Arc::new(NoopInfra), infra_state: Arc::new(NoopInfraState), project: Arc::new(NoopProject), @@ -3820,6 +4032,7 @@ mod replay_tests { let worker_clients = EngineClients { journal: Arc::new(NoopJournal), tasks: Arc::new(NoopTaskStore), + costs: Arc::new(NoopTaskStore), infra: Arc::new(NoopInfra), infra_state: Arc::new(NoopInfraState), project: Arc::new(NoopProject), @@ -3859,6 +4072,7 @@ mod replay_tests { let test_clients = EngineClients { journal: Arc::new(NoopJournal), tasks: Arc::new(NoopTaskStore), + costs: Arc::new(NoopTaskStore), infra: Arc::new(NoopInfra), infra_state: Arc::new(NoopInfraState), project: Arc::new(NoopProject), @@ -4587,17 +4801,38 @@ mod test_journal { use std::sync::Mutex as StdMutex; use weft_journal::ExecEvent; + /// A write that could not reach the journal at all. + #[derive(Debug)] + pub(super) struct Unreachable; + impl std::fmt::Display for Unreachable { + fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + f.write_str("the journal could not be reached") + } + } + impl std::error::Error for Unreachable {} + /// Capturing journal client. Stores every `record_event` payload /// so context and bus tests can assert what was saved. `fail_count` - /// rejects that many following writes before accepting them again. + /// rejects that many following writes before accepting them again; + /// `unreachable` fails a row's write as if the journal could not be + /// reached. #[derive(Default)] pub(super) struct CaptureJournal { pub(super) events: StdMutex>, pub(super) fail_count: StdMutex, + /// Rows (by node) whose write still fails, as many more times, as + /// if the journal could not be reached. + pub(super) unreachable: StdMutex>, } #[async_trait] impl weft_journal::JournalClient for CaptureJournal { async fn record_event(&self, event: &ExecEvent, _instance: Option<&str>) -> anyhow::Result<()> { + if let ExecEvent::NodeStarted { node_id, .. } = event { + if let Some(left) = self.unreachable.lock().unwrap().get_mut(node_id).filter(|left| **left > 0) { + *left -= 1; + return Err(Unreachable.into()); + } + } { let mut fc = self.fail_count.lock().unwrap(); if *fc > 0 { @@ -4608,6 +4843,13 @@ mod test_journal { self.events.lock().unwrap().push(event.clone()); Ok(()) } + /// One part per row, so a write of several rows can fail part way. + fn parts<'e>(&self, events: &'e [ExecEvent]) -> anyhow::Result> { + Ok(events.chunks(1).collect()) + } + fn never_reached(&self, error: &anyhow::Error) -> bool { + error.downcast_ref::().is_some() + } /// Captures are for asserting writes; nothing reads them back. async fn raw_rows_after( &self, @@ -4628,6 +4870,7 @@ mod test_journal { mod drive_journal_tests { use super::*; use super::test_journal::CaptureJournal; + use std::sync::Mutex as StdMutex; fn started(execution_id: ExecutionId, node: &str) -> ExecEvent { ExecEvent::NodeStarted { execution_id, node_id: node.into(), frames: Vec::new(), at_unix: 0 } @@ -4646,29 +4889,26 @@ mod drive_journal_tests { .collect() } - /// The driver's rows wait for the next flush, and a node's own write - /// carries them out ahead of itself, so the journal's order is the - /// order things happened. + /// Writes answer at once and reach the journal in the order handed, + /// the driver's and its nodes' alike; a flush answers once they all + /// have. #[tokio::test] - async fn held_rows_go_out_first_and_in_order() { + async fn rows_go_out_in_the_background_in_order() { let capture = Arc::new(CaptureJournal::default()); let drive = DriveJournal::wrap(capture.clone()); let driver = drive.deferring(); let run = ExecutionId::new_v4(); driver.record_event(&started(run, "a"), Some("w")).await.unwrap(); - driver.record_event(&started(run, "b"), Some("w")).await.unwrap(); - assert!(nodes(&capture).is_empty(), "held until a flush"); drive.record_event(&started(run, "node-write"), Some("w")).await.unwrap(); - assert_eq!(nodes(&capture), ["a", "b", "node-write"]); - driver.record_event(&started(run, "c"), Some("w")).await.unwrap(); + driver.record_event(&started(run, "b"), Some("w")).await.unwrap(); drive.flush().await.unwrap(); - assert_eq!(nodes(&capture), ["a", "b", "node-write", "c"]); + assert_eq!(nodes(&capture), ["a", "node-write", "b"]); assert!(!drive.is_poisoned()); } /// A read never runs past rows the drive wrote. #[tokio::test] - async fn a_read_sends_what_is_held_first() { + async fn a_read_waits_for_what_was_handed_first() { let capture = Arc::new(CaptureJournal::default()); let drive = DriveJournal::wrap(capture.clone()); let run = ExecutionId::new_v4(); @@ -4677,31 +4917,139 @@ mod drive_journal_tests { assert_eq!(nodes(&capture), ["a"]); } - /// A failed send poisons the drive: the journal is now behind what + /// A task goes out only once every row handed before it is on record. + #[tokio::test] + async fn a_task_waits_for_the_rows_handed_before_it() { + let capture = Arc::new(CaptureJournal::default()); + let drive = DriveJournal::wrap(capture.clone()); + let run = ExecutionId::new_v4(); + drive.record_event(&started(run, "a"), Some("w")).await.unwrap(); + let seen = Arc::new(StdMutex::new(Vec::new())); + let inner: Arc = Arc::new(SeeingTasks { capture: capture.clone(), seen: seen.clone() }); + let tasks = JournalFirst { inner, journal: drive.clone() }; + tasks.enqueue_dedup(task_store::NewTask { + kind: "noop".into(), + target: task_store::TaskTarget::Dispatcher, + project_id: None, + dedup_key: None, + execution_id: None, + tenant_id: "t".into(), + target_replica: None, + binary_hash: None, + payload: Value::Null, + }) + .await + .unwrap(); + assert_eq!(*seen.lock().unwrap(), vec![1], "the task saw the row already written"); + } + + /// A failed write poisons the drive: the journal is now behind what /// the worker did, and the loop stops on it. #[tokio::test] - async fn a_failed_send_poisons_the_drive() { + async fn a_failed_write_poisons_the_drive() { let capture = Arc::new(CaptureJournal::default()); *capture.fail_count.lock().unwrap() = 1; let drive = DriveJournal::wrap(capture.clone()); drive.deferring().record_event(&started(ExecutionId::new_v4(), "a"), Some("w")).await.unwrap(); assert!(drive.flush().await.is_err()); assert!(drive.is_poisoned()); - drive.deferring().record_event(&started(ExecutionId::new_v4(), "b"), Some("w")).await.unwrap(); - assert!(drive.flush().await.is_err(), "a poisoned drive sends nothing more"); + assert!( + drive.deferring().record_event(&started(ExecutionId::new_v4(), "b"), Some("w")).await.is_err(), + "a poisoned drive takes nothing more" + ); + assert!(drive.flush().await.is_err()); + assert!(capture.events.lock().unwrap().is_empty()); + } + + /// A part of a write that never reached the broker is sent again, and + /// lands once, in order, without the parts before it going again (one + /// that reached it and failed is never sent again: + /// `a_failed_write_poisons_the_drive`). + #[tokio::test(start_paused = true)] + async fn only_a_write_that_never_reached_the_broker_is_sent_again() { + let capture = Arc::new(CaptureJournal::default()); + let drive = DriveJournal::wrap(capture.clone()); + let run = ExecutionId::new_v4(); + // One write of three parts: the first lands, the second cannot + // reach the journal twice and then does, and the first is never + // sent again. + capture.unreachable.lock().unwrap().insert("b".into(), 2); + drive.deferring().record_events(&[started(run, "a"), started(run, "b"), started(run, "c")], Some("w")).await.unwrap(); + drive.flush().await.expect("sent once the journal could be reached"); + assert_eq!(nodes(&capture), ["a", "b", "c"], "each row once, in order"); + } + + /// A broker out of reach for longer than the resend window stops the + /// run, like any failed write. + #[tokio::test(start_paused = true)] + async fn a_broker_out_of_reach_too_long_stops_the_run() { + let capture = Arc::new(CaptureJournal::default()); + capture.unreachable.lock().unwrap().insert("a".into(), usize::MAX); + let drive = DriveJournal::wrap(capture.clone()); + drive.deferring().record_event(&started(ExecutionId::new_v4(), "a"), Some("w")).await.unwrap(); + assert!(drive.flush().await.is_err()); + assert!(drive.is_poisoned()); + } + + /// A drive waiting on something else hears its journal fail at once. + #[tokio::test] + async fn a_waiting_drive_hears_a_failed_write() { + let capture = Arc::new(CaptureJournal::default()); + *capture.fail_count.lock().unwrap() = 1; + let drive = DriveJournal::wrap(capture.clone()); + let waiting = { + let drive = drive.clone(); + tokio::spawn(async move { drive.poisoned().await }) + }; + drive.deferring().record_event(&started(ExecutionId::new_v4(), "a"), Some("w")).await.unwrap(); + tokio::time::timeout(Duration::from_secs(5), waiting).await.expect("heard the failure").unwrap(); + assert!(drive.is_poisoned()); } - /// A write under another replica is refused without dropping what was - /// already held. + /// A write under another replica poisons the drive. #[tokio::test] - async fn a_mismatched_replica_drops_nothing_held() { + async fn a_mismatched_replica_poisons_the_drive() { let capture = Arc::new(CaptureJournal::default()); let drive = DriveJournal::wrap(capture.clone()); let run = ExecutionId::new_v4(); drive.deferring().record_event(&started(run, "a"), Some("w")).await.unwrap(); assert!(drive.record_event(&started(run, "b"), Some("other")).await.is_err()); assert!(drive.is_poisoned()); - assert_eq!(drive.held.lock().unwrap().0.len(), 1, "the held row is still there"); + } + + /// Counts the journal's rows at each enqueue. + struct SeeingTasks { + capture: Arc, + seen: Arc>>, + } + + #[async_trait::async_trait] + impl TaskStoreClient for SeeingTasks { + async fn enqueue_dedup(&self, _spec: task_store::NewTask) -> anyhow::Result { + self.seen.lock().unwrap().push(self.capture.events.lock().unwrap().len()); + Ok(task_store::DedupOutcome::Inserted(uuid::Uuid::nil())) + } + async fn wait_for_terminal(&self, _: uuid::Uuid, _: Duration) -> anyhow::Result { + unimplemented!("not asked") + } + async fn claim_one(&self, _: &str, _: task_store::ClaimFilter, _: Duration) -> anyhow::Result> { + unimplemented!("not asked") + } + async fn heartbeat(&self, _: uuid::Uuid, _: &str) -> anyhow::Result { + unimplemented!("not asked") + } + async fn requeue(&self, _: uuid::Uuid, _: &str) -> anyhow::Result { + unimplemented!("not asked") + } + async fn complete(&self, _: uuid::Uuid, _: &str, _: Value) -> anyhow::Result<()> { + unimplemented!("not asked") + } + async fn fail(&self, _: uuid::Uuid, _: &str, _: String) -> anyhow::Result<()> { + unimplemented!("not asked") + } + async fn wait_cancels(&self, _: uuid::Uuid, _: Vec, _: Duration) -> anyhow::Result> { + unimplemented!("not asked") + } } } diff --git a/crates/weft-engine/src/execution_driver.rs b/crates/weft-engine/src/execution_driver.rs index 793c6240..9d696588 100644 --- a/crates/weft-engine/src/execution_driver.rs +++ b/crates/weft-engine/src/execution_driver.rs @@ -209,8 +209,8 @@ pub(crate) async fn run_one_execution_observed( // primary exit signal. // // The pump takes the UNwrapped journal client: bus-row failures - // degrade per-bus without poisoning the drive (`drive_execution_id` wraps - // its own copy). Both live here, around the drive, so the shutdown + // degrade per-bus without poisoning the drive (the drive's journal is + // wrapped below). Both live here, around the drive, so the shutdown // below runs whether the drive returned an outcome or an error: a // run that bailed out must not leave its buses open or its pump // running. (A shutdown that panics on its deadline leaves the @@ -229,8 +229,8 @@ pub(crate) async fn run_one_execution_observed( // here so the run still gets its Failed terminal below, instead of // unwinding past it and reading as running forever. // The drive's journal is made out here, around the drive, so the rows - // it holds outlive a drive that fails or panics: they go on record - // before the Failed terminal below, and the run reads as far as it got. + // still queued when a drive fails or panics go on record before the + // Failed terminal below, and the run reads as far as it got. let drive_journal = crate::context::DriveJournal::wrap(clients.journal.clone()); let drove = match futures::FutureExt::catch_unwind(std::panic::AssertUnwindSafe(drive_execution_id( project, @@ -374,10 +374,13 @@ async fn drive_execution_id( // ExecutionStarted + NodeKicked; wait briefly for the rows. // // The drive writes through its journal (`run_one_execution_observed` - // made it): the driver's own rows are held and sent together, and a - // failed write poisons the drive (the loop checks every iteration and - // exits the worker; see `DriveJournal`). - let clients = EngineClients { journal: drive_journal.clone(), ..clients }; + // made it): rows go out in the background, in order, and a failed + // write poisons the drive (the loop checks every iteration and exits + // the worker; see `DriveJournal`). A task, a tag or a stop by tag goes + // out only once the rows handed before it are on record (`JournalFirst`). + let tasks = Arc::new(crate::context::JournalFirst { inner: clients.tasks.clone(), journal: drive_journal.clone() }); + let steering = Arc::new(crate::context::JournalFirst { inner: clients.steering.clone(), journal: drive_journal.clone() }); + let clients = EngineClients { journal: drive_journal.clone(), tasks, steering, ..clients }; let journal = clients.journal.clone(); // The run's log from its birth row, waited for briefly (see // `FIRST_ROWS_WAIT`). The wait yields to cancellation: a cancel @@ -672,8 +675,8 @@ async fn drive_execution_id( ); } - // Whatever the last drive held goes out now: a stalled run writes - // no terminal to carry it. + // Whatever the last drive queued is on record before this returns: a + // stalled run writes no terminal after it. drive_journal.flush().await?; // Journal the terminal event based on what the worker actually @@ -693,8 +696,8 @@ async fn drive_execution_id( outcome = ExecutionOutcome::Failed { error: weft_core::caller::NO_ANSWER.to_string() }; } // The terminal is written to the journal beneath the drive's - // (`DriveJournal::beneath`), after every held row went out (the flush - // above, and the one after a cancel walk below): a failed terminal is + // (`DriveJournal::beneath`), after every queued row went out (the + // flush above, and the one after a cancel walk below): a failed terminal is // then exactly that, retried by `journal_terminal`, and a failed drive // row has already stopped the run before it gets here. let terminal = match &outcome { @@ -742,17 +745,16 @@ async fn drive_execution_id( Ok(Drove { outcome, pulses, executions, loop_runtime, kicked }) } -/// Send the drive's held rows, or stop the drive: a write that failed -/// leaves the journal behind what the worker did (see `DriveJournal`). -async fn flush_or_stop(drive_journal: &crate::context::DriveJournal, execution_id: ExecutionId) -> anyhow::Result<()> { - // A poisoned drive refuses the flush itself, so nothing new starts on - // a journal that is already behind. - drive_journal.flush().await.map_err(|e| { - anyhow::anyhow!( - "a journal write failed mid-drive for execution {execution_id}; the journal no longer \ - holds what the worker did, so the run cannot go on: {e:#}" - ) - }) +/// Stop the drive once one of its journal writes failed: the journal is +/// then behind what the worker did (see `DriveJournal`), and nothing new +/// starts on it. +fn stop_if_poisoned(drive_journal: &crate::context::DriveJournal, execution_id: ExecutionId) -> anyhow::Result<()> { + anyhow::ensure!( + !drive_journal.is_poisoned(), + "a journal write failed mid-drive for execution {execution_id}; the journal no longer holds what \ + the worker did, so the run cannot go on" + ); + Ok(()) } /// The run's state as its fold says, refusing to resume over a row the @@ -1208,9 +1210,8 @@ async fn drive( live: &mut weft_journal::LiveFold, ) -> anyhow::Result { let project: &ProjectDefinition = project_arc; - // The driver's own rows are held and go out together at each flush - // below (`DriveJournal`); its nodes write through `clients.journal`, - // which sends at once and carries the held rows ahead of theirs. + // The driver's rows and its nodes' (through `clients.journal`) go to + // the journal in the background, in one order (`DriveJournal`). let deferring = drive_journal.deferring(); let journal: &dyn JournalClient = &deferring; // What infra the program declares, read once per drive: every @@ -1296,12 +1297,7 @@ async fn drive( // Stop here: the error becomes the run's Failed terminal // (`run_one_execution`), which is the only end this run gets, // since nothing respawns an execution whose task failed. - if drive_journal.is_poisoned() { - anyhow::bail!( - "a journal write failed mid-drive for execution {execution_id}; the journal no longer \ - holds what the worker did, so the run cannot go on" - ); - } + stop_if_poisoned(drive_journal, execution_id)?; // Cancellation checkpoint. Checked at the TOP of every // iteration regardless of whether the previous iteration @@ -2164,9 +2160,10 @@ async fn drive( let provision_clients = clients.clone(); let provision_copy = weft_core::instance::copy_owner(node_def.per_instance, instance).cloned(); - // The firing's start (and everything held before it) is on - // record before its body can do anything. - flush_or_stop(drive_journal, execution_id).await?; + // The firing's start goes to the journal in the background + // (`DriveJournal`): its body starts now, whatever is still being + // sent, unless an earlier write already failed. + stop_if_poisoned(drive_journal, execution_id)?; let abort_handle = in_flight.spawn(async move { if is_infra_setup_provision { // 1. Call the node's provision body. @@ -2532,12 +2529,16 @@ async fn drive( }; tokio::pin!(resume_poll); - // Everything this step wrote goes out before the drive waits. - flush_or_stop(drive_journal, execution_id).await?; + // The drive waits for news, not for its own rows, which keep + // going out in the background; a failed one stops it. + stop_if_poisoned(drive_journal, execution_id)?; let on_wait_change = waits.wait_notified(); tokio::pin!(on_wait_change); on_wait_change.as_mut().enable(); tokio::select! { + // A journal write failed while the drive waited: the next turn + // stops it, rather than waiting on a long node first. + () = drive_journal.poisoned() => {} fresh = resume_poll.as_mut() => { // Bus-held worker with a pending suspension, and the // journal answered. If a new row landed, SURGICALLY resume diff --git a/crates/weft-engine/src/execution_driver_tests/rig.rs b/crates/weft-engine/src/execution_driver_tests/rig.rs index ad238da3..229dbe73 100644 --- a/crates/weft-engine/src/execution_driver_tests/rig.rs +++ b/crates/weft-engine/src/execution_driver_tests/rig.rs @@ -571,6 +571,7 @@ EngineClients { journal, tasks: Arc::new(NoopTasks), + costs: Arc::new(NoopTasks), infra: Arc::new(NoopInfra), infra_state: Arc::new(NoopInfraState), project: Arc::new(NoopProject), diff --git a/crates/weft-journal/src/traits.rs b/crates/weft-journal/src/traits.rs index 6d06bf39..90ab3a53 100644 --- a/crates/weft-journal/src/traits.rs +++ b/crates/weft-journal/src/traits.rs @@ -59,6 +59,21 @@ pub trait JournalClient: Send + Sync { Ok(()) } + /// How a write of `events` goes out: the parts, in order, each sent in + /// one request, so a writer that sends a write again can send only the + /// part that failed. One part by default. + fn parts<'e>(&self, events: &'e [ExecEvent]) -> anyhow::Result> { + Ok(vec![events]) + } + + /// Whether a failed write of one part never reached the journal (the + /// connection itself could not be made), so sending it again cannot + /// record its rows twice. Never, by default: a write that may have + /// landed is not sent again. + fn never_reached(&self, _error: &anyhow::Error) -> bool { + false + } + /// The rows of `execution_id` after `after_id`, in order, as RAW payload /// strings, holding up to `wait` for at least one to exist (a zero /// `wait` answers at once; empty when none came). An execution's rows diff --git a/crates/weft-journal/src/unrecorded.rs b/crates/weft-journal/src/unrecorded.rs index eb55e8b3..070f263d 100644 --- a/crates/weft-journal/src/unrecorded.rs +++ b/crates/weft-journal/src/unrecorded.rs @@ -193,6 +193,20 @@ pub fn as_recorded(events: Vec) -> Vec { #[async_trait] impl JournalClient for UnrecordedJournal { + /// One row per part: rows are held in memory, so cutting costs + /// nothing, and a cost that fails to reach the real journal leaves + /// nothing of its part held, so sending the part again holds nothing + /// twice. + fn parts<'e>(&self, events: &'e [ExecEvent]) -> anyhow::Result> { + Ok(events.chunks(1).collect()) + } + + /// Only a cost goes to the real journal; whether it never reached it is + /// the real journal's to say. + fn never_reached(&self, error: &anyhow::Error) -> bool { + self.real.never_reached(error) + } + async fn record_event(&self, event: &ExecEvent, replica: Option<&str>) -> anyhow::Result<()> { anyhow::ensure!( event.execution_id() == self.execution_id, @@ -279,6 +293,8 @@ pub async fn record_retroactively( "the record of unrecorded run {execution_id} carries rows of another run" ); let mut tx = pool.begin().await?; + // The execution's lock before the row lock below, which is a write. + // SYNC: execution lock before first write <-> crates/weft-dispatcher/src/settled.rs (the list) crate::write::lock_execution_ids(&mut tx, &[execution_id]).await?; let row: Option<(uuid::Uuid, Option, Option)> = sqlx::query_as( "SELECT project_id, fired_by, instance_id FROM execution WHERE execution_id = $1 AND kind = $2 FOR UPDATE", @@ -412,13 +428,25 @@ mod tests { written: Mutex>, recorded: Mutex>>, forgotten: Mutex, + /// Writes still to fail as if the journal could not be reached. + unreachable: Mutex, } #[async_trait] impl JournalClient for Real { async fn record_event(&self, event: &ExecEvent, _: Option<&str>) -> anyhow::Result<()> { + { + let mut left = self.unreachable.lock().unwrap(); + if *left > 0 { + *left -= 1; + anyhow::bail!("unreachable"); + } + } self.written.lock().unwrap().push(event.clone()); Ok(()) } + fn never_reached(&self, error: &anyhow::Error) -> bool { + error.to_string() == "unreachable" + } async fn raw_rows_after(&self, _: ExecutionId, _: i64, _: Duration) -> anyhow::Result> { Ok(Vec::new()) } @@ -502,6 +530,25 @@ mod tests { assert_eq!(reader.await.unwrap().unwrap().len(), 1); } + /// A writer that sends a part again sends one row at a time here, and + /// hears from the real journal whether a cost never reached it; a cost + /// that failed holds nothing, so sending it again holds it once. + #[tokio::test] + async fn a_cost_that_never_reached_the_journal_can_be_sent_again() { + let execution_id = ExecutionId::new_v4(); + let real = Arc::new(Real::default()); + *real.unreachable.lock().unwrap() = 1; + let journal = UnrecordedJournal::seeded(execution_id, birth(execution_id), real.clone()).unwrap(); + let rows = [cost(execution_id), ExecEvent::ExecutionCompleted { execution_id, at_unix: 3 }]; + assert_eq!(journal.parts(&rows).unwrap().len(), 2, "one row per part"); + let failed = journal.record_event(&rows[0], None).await.unwrap_err(); + assert!(journal.never_reached(&failed), "the real journal says it never got there"); + assert_eq!(journal.rows_after(execution_id, 2, Duration::ZERO).await.unwrap().len(), 0, "nothing held"); + journal.record_event(&rows[0], None).await.unwrap(); + assert_eq!(real.written.lock().unwrap().len(), 1); + assert_eq!(journal.rows_after(execution_id, 2, Duration::ZERO).await.unwrap().len(), 1, "held once"); + } + #[tokio::test] async fn a_cost_goes_through_as_it_happens() { let execution_id = ExecutionId::new_v4(); diff --git a/crates/weft-journal/src/write.rs b/crates/weft-journal/src/write.rs index 21cc0c99..0f30d7f4 100644 --- a/crates/weft-journal/src/write.rs +++ b/crates/weft-journal/src/write.rs @@ -85,7 +85,9 @@ pub async fn record_event_in<'e, E: sqlx::PgExecutor<'e>>( Ok(()) } -/// THE insert every journal write makes. The per-execution lock is taken +/// THE insert every journal write makes, through the database's +/// `weft_journal_append` (the dispatcher's journal schema group), which a +/// run's birth calls too. The per-execution lock is taken /// BEFORE any row's id is drawn, and held until the writing transaction /// ends, so the rows of one execution are numbered and committed in the /// same order: once a reader sees row N of an execution, every earlier row @@ -114,36 +116,17 @@ async fn insert<'e, E: sqlx::PgExecutor<'e>>( let kinds: Vec<&str> = events.iter().map(|e| e.kind_str()).collect(); let payloads: Vec = events.iter().map(serde_json::to_string).collect::>()?; let now = unix_now()?; - let written = sqlx::query( - "WITH locked AS MATERIALIZED ( \ - SELECT pg_advisory_xact_lock(hashtextextended($6, 0)) \ - ) \ - INSERT INTO exec_event (execution_id, kind, payload_json, created_at, replica, dedup_key) \ - SELECT $1, e.kind, e.payload, $4, $5, $8 \ - FROM locked, unnest($2::text[], $3::text[]) WITH ORDINALITY AS e(kind, payload, n) \ - WHERE $7::text IS NULL \ - OR EXISTS (SELECT 1 FROM execution x WHERE x.execution_id = $1 AND x.owner_replica = $7) \ - ORDER BY e.n \ - ON CONFLICT (dedup_key) WHERE dedup_key IS NOT NULL DO NOTHING", - ) - .bind(execution_id.to_string()) - .bind(&kinds) - .bind(&payloads) - .bind(now) - .bind(replica) - .bind(execution_id_lock_key(execution_id)) - .bind(owner) - .bind(dedup_key) - .execute(executor) - .await? - .rows_affected(); - Ok(written) -} - -/// The advisory lock key of one execution's journal: the ONE definition, -/// shared by [`lock_execution_ids`] and the lock every write here takes. -fn execution_id_lock_key(execution_id: weft_core::ExecutionId) -> String { - format!("exec_event:{execution_id}") + let written: i64 = sqlx::query_scalar("SELECT weft_journal_append($1, $2, $3, $4, $5, $6, $7)") + .bind(execution_id.to_string()) + .bind(&kinds) + .bind(&payloads) + .bind(now) + .bind(replica) + .bind(owner) + .bind(dedup_key) + .fetch_one(executor) + .await?; + Ok(written as u64) } /// Take the journal lock of every execution in `execution_ids`, held until the @@ -159,10 +142,8 @@ pub async fn lock_execution_ids( sorted.sort_unstable(); sorted.dedup(); for execution_id in sorted { - sqlx::query("SELECT pg_advisory_xact_lock(hashtextextended($1, 0))") - .bind(execution_id_lock_key(execution_id)) - .execute(&mut *tx) - .await?; + // The lock's key is spelled once, in the journal's schema. + sqlx::query("SELECT weft_lock_execution($1)").bind(execution_id.to_string()).execute(&mut *tx).await?; } Ok(()) } diff --git a/crates/weft-runtime/src/role_waker.rs b/crates/weft-runtime/src/role_waker.rs index 0887f45f..89b64817 100644 --- a/crates/weft-runtime/src/role_waker.rs +++ b/crates/weft-runtime/src/role_waker.rs @@ -140,6 +140,8 @@ impl RoleWaker { } // Notifications may have been lost. Heard::Recheck => self.ring_every_loop(), + // The recheck after the reconnect rings every loop. + Heard::Lost => {} } } } diff --git a/crates/weft-task-store/migrations/entry_rate/20261004T190703_entry_rate_live_call_one_round_trip.sql b/crates/weft-task-store/migrations/entry_rate/20261004T190703_entry_rate_live_call_one_round_trip.sql new file mode 100644 index 00000000..dd35fbf5 --- /dev/null +++ b/crates/weft-task-store/migrations/entry_rate/20261004T190703_entry_rate_live_call_one_round_trip.sql @@ -0,0 +1,83 @@ +CREATE OR REPLACE FUNCTION weft_admit(p jsonb) + RETURNS jsonb + LANGUAGE plpgsql +AS $function$ + DECLARE + v_window BIGINT := (p->>'window_start')::bigint; + v_later INTEGER := (p->>'retry_after_secs')::integer; + v_now BIGINT := (p->>'now')::bigint; + v_slot JSONB := p->'slot'; + v_count JSONB; + v_hits INTEGER; + v_refused JSONB; + v_room BOOLEAN; + BEGIN + IF jsonb_typeof(p->'blocked') = 'object' THEN + SELECT r.hits INTO v_hits FROM entry_rate r + WHERE r.key = p->'blocked'->>'key' AND r.window_start = v_window; + IF v_hits >= (p->'blocked'->>'limit')::integer THEN + RETURN jsonb_build_object('reason', 'invalid_tokens', 'retry_after_secs', v_later); + END IF; + END IF; + FOR v_count IN SELECT c FROM jsonb_array_elements(p->'counts') WITH ORDINALITY AS e(c, n) ORDER BY n LOOP + IF weft_rate_hit(v_count->>'key', v_window) > (v_count->>'limit')::integer THEN + v_refused := jsonb_build_object('reason', v_count->>'reason', 'retry_after_secs', v_later); + EXIT; + END IF; + END LOOP; + IF v_refused IS NULL AND jsonb_typeof(v_slot) = 'object' THEN + IF v_slot->>'execution_id' IS NULL THEN + SELECT COUNT(*) < (v_slot->>'max')::bigint INTO v_room FROM entry_slot s + WHERE s.signal_token = v_slot->>'token' + AND NOT weft_slot_stopped_counting(s.execution_id, s.unborn_until, v_now); + ELSE + PERFORM pg_advisory_xact_lock(hashtextextended('entry_slot:' || (v_slot->>'token'), 0)); + SELECT EXISTS (SELECT 1 FROM entry_slot s WHERE s.execution_id = v_slot->>'execution_id' + AND NOT weft_slot_stopped_counting(s.execution_id, s.unborn_until, v_now)) + OR (SELECT COUNT(*) FROM entry_slot s + WHERE s.signal_token = v_slot->>'token' AND s.execution_id <> v_slot->>'execution_id' + AND NOT weft_slot_stopped_counting(s.execution_id, s.unborn_until, v_now)) + < (v_slot->>'max')::bigint + INTO v_room; + IF v_room THEN + INSERT INTO entry_slot (execution_id, signal_token, unborn_until) + VALUES (v_slot->>'execution_id', v_slot->>'token', (v_slot->>'unborn_until')::bigint) + ON CONFLICT (execution_id) DO UPDATE + SET unborn_until = GREATEST(entry_slot.unborn_until, EXCLUDED.unborn_until); + END IF; + END IF; + IF NOT v_room THEN + v_refused := jsonb_build_object('reason', 'at_once', 'retry_after_secs', 5); + END IF; + END IF; + IF v_refused IS NOT NULL THEN + PERFORM weft_rate_hit((p->>'refusals_key') || (v_refused->>'reason'), v_window); + END IF; + RETURN v_refused; + END; + $function$ +; + +CREATE OR REPLACE FUNCTION weft_rate_hit(p_key text, p_window_start bigint) + RETURNS integer + LANGUAGE sql +AS $function$ + INSERT INTO entry_rate (key, window_start, hits) VALUES (p_key, p_window_start, 1) + ON CONFLICT (key, window_start) DO UPDATE SET hits = entry_rate.hits + 1 + RETURNING hits + $function$ +; + +CREATE OR REPLACE FUNCTION weft_slot_stopped_counting(p_execution_id text, p_unborn_until bigint, p_now bigint) + RETURNS boolean + LANGUAGE sql + STABLE +AS $function$ + SELECT (p_unborn_until < p_now + AND NOT EXISTS (SELECT 1 FROM execution ec WHERE ec.execution_id = p_execution_id)) + OR EXISTS (SELECT 1 FROM execution ec WHERE ec.execution_id = p_execution_id + AND ec.ended_at_unix IS NOT NULL) + OR EXISTS (SELECT 1 FROM exec_event e WHERE e.execution_id = p_execution_id + AND e.kind IN ('execution_completed', 'execution_failed', 'execution_cancelled')) + $function$ +; diff --git a/crates/weft-task-store/migrations/exec_event/20261004T190703_exec_event_live_call_one_round_trip.sql b/crates/weft-task-store/migrations/exec_event/20261004T190703_exec_event_live_call_one_round_trip.sql new file mode 100644 index 00000000..a0ac5501 --- /dev/null +++ b/crates/weft-task-store/migrations/exec_event/20261004T190703_exec_event_live_call_one_round_trip.sql @@ -0,0 +1,148 @@ +CREATE OR REPLACE FUNCTION routes_notify_tenant() + RETURNS trigger + LANGUAGE plpgsql +AS $function$ + BEGIN + IF TG_OP = 'DELETE' THEN + PERFORM pg_notify('weft_routes', OLD.tenant_id); + ELSE + PERFORM pg_notify('weft_routes', NEW.tenant_id); + END IF; + RETURN NULL; + END; + $function$ +; + +CREATE OR REPLACE FUNCTION weft_execution_started(p jsonb) + RETURNS void + LANGUAGE plpgsql +AS $function$ + DECLARE + v_execution_id TEXT := p->>'execution_id'; + v_project UUID := (p->>'project_id')::uuid; + seeded BIGINT; + BEGIN + PERFORM weft_lock_execution(v_execution_id); + IF p->>'source_version' IS NOT NULL THEN + PERFORM 1 FROM project_version + WHERE project_id = v_project AND id = p->>'source_version' FOR KEY SHARE; + IF NOT FOUND THEN + RAISE EXCEPTION 'source version % was removed during preparation; run the command again', + p->>'source_version'; + END IF; + END IF; + IF (p->>'journaled')::boolean THEN + PERFORM weft_journal_append(v_execution_id, ARRAY[p->>'kind'], ARRAY[p->>'payload'], + (p->>'created_at')::bigint, NULL, NULL, p->>'dedup_key'); + END IF; + INSERT INTO execution (execution_id, project_id, tenant_id, started_at_unix, phase, kind, instance_id, fired_by) + SELECT v_execution_id, v_project, pr.tenant_id, (p->>'at_unix')::bigint, p->>'phase', + p->>'run_kind', p->>'instance_id', p->>'fired_by' + FROM project pr WHERE pr.id = v_project + ON CONFLICT (execution_id) DO NOTHING; + GET DIAGNOSTICS seeded = ROW_COUNT; + IF seeded = 0 AND NOT EXISTS (SELECT 1 FROM execution WHERE execution_id = v_execution_id) THEN + RAISE EXCEPTION 'refuse to journal ExecutionStarted for execution %: project % has no row, so the execution seed (which the broker scope check and the terminal sweeps depend on) cannot be written; register the project first', + v_execution_id, v_project; + END IF; + END; + $function$ +; + +CREATE OR REPLACE FUNCTION weft_journal_append(p_execution_id text, p_kinds text[], p_payloads text[], p_created_at bigint, p_replica text, p_owner text, p_dedup_key text) + RETURNS bigint + LANGUAGE plpgsql +AS $function$ + DECLARE + written BIGINT; + BEGIN + PERFORM weft_lock_execution(p_execution_id); + INSERT INTO exec_event (execution_id, kind, payload_json, created_at, replica, dedup_key) + SELECT p_execution_id, e.kind, e.payload, p_created_at, p_replica, p_dedup_key + FROM unnest(p_kinds, p_payloads) WITH ORDINALITY AS e(kind, payload, n) + WHERE p_owner IS NULL + OR EXISTS (SELECT 1 FROM execution x + WHERE x.execution_id = p_execution_id AND x.owner_replica = p_owner) + ORDER BY e.n + ON CONFLICT (dedup_key) WHERE dedup_key IS NOT NULL DO NOTHING; + GET DIAGNOSTICS written = ROW_COUNT; + RETURN written; + END; + $function$ +; + +CREATE OR REPLACE FUNCTION weft_lock_execution(p_execution_id text) + RETURNS void + LANGUAGE plpgsql +AS $function$ + BEGIN + PERFORM pg_advisory_xact_lock(hashtextextended('exec_event:' || p_execution_id, 0)); + END; + $function$ +; + +CREATE OR REPLACE FUNCTION weft_start_execution(p jsonb) + RETURNS jsonb + LANGUAGE plpgsql +AS $function$ + DECLARE + v_started JSONB := p->'started'; + v_task JSONB := p->'task'; + v_execution_id TEXT := v_started->>'execution_id'; + v_project UUID := (v_started->>'project_id')::uuid; + v_journaled BOOLEAN := (v_started->>'journaled')::boolean; + v_refused JSONB; + v_inserted BOOLEAN; + BEGIN + PERFORM weft_lock_execution(v_execution_id); + IF (v_journaled AND EXISTS (SELECT 1 FROM exec_event + WHERE execution_id = v_execution_id AND kind = 'execution_started')) + OR (NOT v_journaled AND EXISTS (SELECT 1 FROM execution WHERE execution_id = v_execution_id)) THEN + RETURN jsonb_build_object('outcome', 'already_started'); + END IF; + IF jsonb_typeof(p->'admission') = 'object' THEN + v_refused := weft_admit(p->'admission'); + IF v_refused IS NOT NULL THEN + RETURN jsonb_build_object('outcome', 'refused', 'refused', v_refused); + END IF; + END IF; + SELECT d.inserted INTO v_inserted FROM weft_enqueue_dedup( + (v_task->>'id')::uuid, v_task->>'kind', v_task->>'target', (v_task->>'project_id')::uuid, + v_task->>'dedup_key', v_task->>'execution_id', v_task->>'tenant_id', v_task->>'target_replica', + v_task->>'binary_hash', v_task->'payload', (v_task->>'created_at')::bigint) d; + IF NOT v_inserted THEN + RETURN jsonb_build_object('outcome', 'already_started'); + END IF; + PERFORM weft_execution_started(v_started); + -- A trigger setup is born only while the activation that asked + -- for it still owns its rows (a cancel between the claim and + -- here wins), and is recorded as in flight. + IF jsonb_typeof(p->'trigger_setup') = 'object' THEN + IF (p->'trigger_setup'->>'for_activation')::boolean THEN + PERFORM 1 FROM trigger_activation + WHERE project_id = v_project + AND activating_execution_id = v_execution_id::uuid + AND status = 'activating' + FOR UPDATE; + IF NOT FOUND THEN + RAISE EXCEPTION 'activation % ended before trigger setup could start', v_execution_id; + END IF; + END IF; + INSERT INTO trigger_setup (project_id, execution_id) VALUES (v_project, v_execution_id); + END IF; + IF v_journaled AND jsonb_array_length(p->'kicks') > 0 THEN + PERFORM weft_journal_append(v_execution_id, + ARRAY(SELECT k->>'kind' FROM jsonb_array_elements(p->'kicks') WITH ORDINALITY AS e(k, n) ORDER BY n), + ARRAY(SELECT k->>'payload' FROM jsonb_array_elements(p->'kicks') WITH ORDINALITY AS e(k, n) ORDER BY n), + (v_started->>'created_at')::bigint, NULL, NULL, NULL); + END IF; + RETURN jsonb_build_object('outcome', 'started'); + END; + $function$ +; + +CREATE TRIGGER signal_routes_on_change AFTER UPDATE OF surface_kind, mount_path, mount_methods, project_id, node_id, spec_json, auth_kind, auth_config, port_snapshot, program_json, source_version, instance_id, activation_trigger ON signal FOR EACH ROW WHEN (((new.surface_kind = 'public_entry'::text) OR (old.surface_kind = 'public_entry'::text))) EXECUTE FUNCTION routes_notify_tenant(); + +CREATE TRIGGER signal_routes_on_delete AFTER DELETE ON signal FOR EACH ROW WHEN ((old.surface_kind = 'public_entry'::text)) EXECUTE FUNCTION routes_notify_tenant(); + +CREATE TRIGGER signal_routes_on_insert AFTER INSERT ON signal FOR EACH ROW WHEN ((new.surface_kind = 'public_entry'::text)) EXECUTE FUNCTION routes_notify_tenant(); diff --git a/crates/weft-task-store/migrations/infra_node/20261004T190703_infra_node_live_call_one_round_trip.sql b/crates/weft-task-store/migrations/infra_node/20261004T190703_infra_node_live_call_one_round_trip.sql new file mode 100644 index 00000000..468adccf --- /dev/null +++ b/crates/weft-task-store/migrations/infra_node/20261004T190703_infra_node_live_call_one_round_trip.sql @@ -0,0 +1,18 @@ +CREATE OR REPLACE FUNCTION infra_node_status_notify() + RETURNS trigger + LANGUAGE plpgsql +AS $function$ + BEGIN + IF TG_OP = 'DELETE' THEN + PERFORM pg_notify('weft_infra_status', OLD.project_id::text); + ELSE + PERFORM pg_notify('weft_infra_status', NEW.project_id::text); + END IF; + RETURN NULL; + END; + $function$ +; + +CREATE TRIGGER infra_node_status_on_change AFTER UPDATE OF status ON infra_node FOR EACH ROW WHEN ((new.status IS DISTINCT FROM old.status)) EXECUTE FUNCTION infra_node_status_notify(); + +CREATE TRIGGER infra_node_status_on_row AFTER INSERT OR DELETE ON infra_node FOR EACH ROW EXECUTE FUNCTION infra_node_status_notify(); diff --git a/crates/weft-task-store/migrations/project/20261004T190703_project_live_call_one_round_trip.sql b/crates/weft-task-store/migrations/project/20261004T190703_project_live_call_one_round_trip.sql new file mode 100644 index 00000000..7d8924ac --- /dev/null +++ b/crates/weft-task-store/migrations/project/20261004T190703_project_live_call_one_round_trip.sql @@ -0,0 +1,16 @@ +CREATE OR REPLACE FUNCTION project_worker_settings_notify() + RETURNS trigger + LANGUAGE plpgsql +AS $function$ + BEGIN + PERFORM pg_notify('weft_worker_settings', OLD.id::text); + RETURN NULL; + END; + $function$ +; + +CREATE TRIGGER project_routes_on_row AFTER INSERT OR DELETE ON project FOR EACH ROW EXECUTE FUNCTION routes_notify_tenant(); + +CREATE TRIGGER project_worker_settings_on_change AFTER UPDATE OF worker_settings_json ON project FOR EACH ROW WHEN ((new.worker_settings_json IS DISTINCT FROM old.worker_settings_json)) EXECUTE FUNCTION project_worker_settings_notify(); + +CREATE TRIGGER project_worker_settings_on_delete AFTER DELETE ON project FOR EACH ROW EXECUTE FUNCTION project_worker_settings_notify(); diff --git a/crates/weft-task-store/migrations/project_frontend/20261004T190703_project_frontend_live_call_one_round_trip.sql b/crates/weft-task-store/migrations/project_frontend/20261004T190703_project_frontend_live_call_one_round_trip.sql new file mode 100644 index 00000000..24fb7936 --- /dev/null +++ b/crates/weft-task-store/migrations/project_frontend/20261004T190703_project_frontend_live_call_one_round_trip.sql @@ -0,0 +1,3 @@ +ALTER TABLE project_frontend ALTER COLUMN token_id DROP NOT NULL; + +ALTER TABLE project_frontend DROP CONSTRAINT IF EXISTS project_frontend_token_id_not_null; diff --git a/crates/weft-task-store/migrations/task/20261004T190703_task_live_call_one_round_trip.sql b/crates/weft-task-store/migrations/task/20261004T190703_task_live_call_one_round_trip.sql new file mode 100644 index 00000000..aec7d222 --- /dev/null +++ b/crates/weft-task-store/migrations/task/20261004T190703_task_live_call_one_round_trip.sql @@ -0,0 +1,34 @@ +CREATE OR REPLACE FUNCTION weft_enqueue_dedup(p_id uuid, p_kind text, p_target text, p_project uuid, p_dedup text, p_execution text, p_tenant text, p_target_replica text, p_binary_hash text, p_payload jsonb, p_now bigint, OUT task_id uuid, OUT inserted boolean) + RETURNS record + LANGUAGE plpgsql +AS $function$ + BEGIN + PERFORM weft_lock_dedup(p_tenant, p_kind, p_dedup); + SELECT t.id INTO task_id FROM task t + WHERE t.tenant_id = p_tenant AND t.kind = p_kind AND t.dedup_key = p_dedup + AND t.status IN ('pending', 'claimed') + LIMIT 1; + IF FOUND THEN + inserted := FALSE; + RETURN; + END IF; + INSERT INTO task ( + id, kind, status, target, project_id, dedup_key, execution_id, tenant_id, + target_replica, binary_hash, payload, attempts, created_at_unix + ) VALUES (p_id, p_kind, 'pending', p_target, p_project, p_dedup, p_execution, p_tenant, + p_target_replica, p_binary_hash, p_payload, 0, p_now); + task_id := p_id; + inserted := TRUE; + END; + $function$ +; + +CREATE OR REPLACE FUNCTION weft_lock_dedup(p_tenant text, p_kind text, p_dedup text) + RETURNS void + LANGUAGE plpgsql +AS $function$ + BEGIN + PERFORM pg_advisory_xact_lock(hashtextextended(p_tenant || '|' || p_kind || '|' || p_dedup, 0)); + END; + $function$ +; diff --git a/crates/weft-task-store/migrations/trigger_activation/20261004T190703_trigger_activation_live_call_one_round_trip.sql b/crates/weft-task-store/migrations/trigger_activation/20261004T190703_trigger_activation_live_call_one_round_trip.sql new file mode 100644 index 00000000..97b4dc69 --- /dev/null +++ b/crates/weft-task-store/migrations/trigger_activation/20261004T190703_trigger_activation_live_call_one_round_trip.sql @@ -0,0 +1,21 @@ +CREATE OR REPLACE FUNCTION trigger_activation_routes_notify() + RETURNS trigger + LANGUAGE plpgsql +AS $function$ + DECLARE + tenant TEXT; + BEGIN + SELECT p.tenant_id INTO tenant FROM project p + WHERE p.id = CASE WHEN TG_OP = 'DELETE' THEN OLD.project_id ELSE NEW.project_id END; + -- A project already gone told its tenant itself. + IF tenant IS NOT NULL THEN + PERFORM pg_notify('weft_routes', tenant); + END IF; + RETURN NULL; + END; + $function$ +; + +CREATE TRIGGER trigger_activation_routes_on_row AFTER INSERT OR DELETE ON trigger_activation FOR EACH ROW EXECUTE FUNCTION trigger_activation_routes_notify(); + +CREATE TRIGGER trigger_activation_routes_on_status AFTER UPDATE OF status ON trigger_activation FOR EACH ROW WHEN ((new.status IS DISTINCT FROM old.status)) EXECUTE FUNCTION trigger_activation_routes_notify(); diff --git a/crates/weft-task-store/src/held_copy.rs b/crates/weft-task-store/src/held_copy.rs new file mode 100644 index 00000000..0f38da09 --- /dev/null +++ b/crates/weft-task-store/src/held_copy.rs @@ -0,0 +1,374 @@ +//! A copy, in this process's memory, of rows that rarely change and are +//! read on every request: a tenant's routes, the install's domains, a +//! project's worker settings. +//! +//! Reading them from the database on every request costs a round trip +//! each, and with the database in another building that is most of what a +//! request spends. So each process keeps the last read, and drops an entry +//! the moment the database says its rows changed: every write to those +//! rows sends `pg_notify(channel, key)` from a trigger in the table's +//! schema group, in the writing transaction, and the process's one +//! [`PgSignalWatch`] hears it. A sibling replica that made the change +//! tells this one through the same notification, so every replica follows +//! the rows. +//! +//! The rows stay the truth. While the listening connection is down the +//! copy keeps nothing and every read goes to the database: a change made +//! then would not be heard. Listening again drops every entry and keeps +//! again; a watch that stopped for good turns the copy off for good. + +use std::hash::Hash; +use std::sync::atomic::{AtomicBool, Ordering}; +use std::sync::{Arc, Mutex, Weak}; + +use weft_core::content_cache::ContentCache; + +use crate::pg_signal::{Heard, PgSignalWatch, Subscription}; + +/// Which entries a notification on the copy's channel drops. +pub enum Changed { + /// The rows of this one key. + Key(K), + /// Anything: the payload names nothing this copy can key by. + Everything, +} + +/// See the module doc. +pub struct HeldCopy { + entries: ContentCache, + /// Bumped by every drop, under this lock, so a read that started + /// before a change never stores what it read once the change is + /// heard: what it read may be the rows from before the change. + generation: Mutex, + /// `false` while nothing would tell the copy about a change (the + /// listening connection is down, or the watch stopped for good), so it + /// keeps nothing. + following: AtomicBool, + /// The task that follows the channel; it ends with the copy. + follower: Mutex>>, + /// Whether a read is worth keeping: rows nobody will ask for again (a + /// tenant name a scanner made up) are answered and not kept. + keep: fn(&V) -> bool, +} + +impl Drop for HeldCopy { + fn drop(&mut self) { + if let Some(follower) = self.follower.lock().expect("held copy follower").take() { + follower.abort(); + } + } +} + +impl HeldCopy +where + K: Hash + Eq + Clone + Send + Sync + 'static, + V: Send + Sync + 'static, +{ + /// A copy of at most `capacity` keys, following `channel` on `signals` + /// from now on: `key_of` reads which key a notification's payload + /// names, and `keep` says whether a read is worth keeping. Refused when `signals` does not listen on `channel`, which + /// would leave the copy serving rows nothing follows. + pub fn follow( + signals: &PgSignalWatch, + channel: &'static str, + capacity: usize, + key_of: fn(&str) -> Changed, + keep: fn(&V) -> bool, + ) -> anyhow::Result> { + signals.require(channel)?; + Ok(Self::following(signals.subscribe(), channel, capacity, key_of, keep)) + } + + /// [`Self::follow`] on any subscription (a test sends into its other + /// end). + pub fn following( + subscription: Subscription, + channel: &'static str, + capacity: usize, + key_of: fn(&str) -> Changed, + keep: fn(&V) -> bool, + ) -> Arc { + let copy = Arc::new(Self { + entries: ContentCache::new(capacity), + generation: Mutex::new(0), + // A copy made while the watch is down keeps nothing until it + // listens again. + following: AtomicBool::new(subscription.listening()), + follower: Mutex::new(None), + keep, + }); + let follower = tokio::spawn(follow(Arc::downgrade(©), subscription, channel, key_of)); + *copy.follower.lock().expect("held copy follower") = Some(follower); + copy + } + + /// The rows under `key` as this process holds them, if it does. + pub fn held(&self, key: &K) -> Option> { + if !self.following.load(Ordering::Acquire) { + return None; + } + self.entries.get(key) + } + + /// The rows under `key`: this process's copy when it has one, else + /// what `load` reads, kept unless a change was heard while it read. + pub async fn get_or_load(&self, key: K, load: F) -> Result, E> + where + F: FnOnce() -> Fut, + Fut: std::future::Future>, + { + match self.held(&key) { + Some(held) => Ok(held), + None => self.load_fresh(key, load).await, + } + } + + /// The rows under `key` as `load` reads them now, kept like + /// [`Self::get_or_load`]'s. For an answer that refuses something on + /// what the copy says (no such route, a copy not running): a change + /// that has committed and not been heard yet is a few milliseconds + /// old, and the refusal is confirmed against the rows themselves. + pub async fn load_fresh(&self, key: K, load: F) -> Result, E> + where + F: FnOnce() -> Fut, + Fut: std::future::Future>, + { + let following = self.following.load(Ordering::Acquire); + let before = *self.generation.lock().expect("held copy generation"); + let read = Arc::new(load().await?); + if following && (self.keep)(&read) { + let generation = self.generation.lock().expect("held copy generation"); + if *generation == before && self.following.load(Ordering::Acquire) { + self.entries.put(key, read.clone()); + } + } + Ok(read) + } + + fn drop_changed(&self, changed: Changed) { + let mut generation = self.generation.lock().expect("held copy generation"); + *generation += 1; + match changed { + Changed::Key(key) => self.entries.forget(&key), + Changed::Everything => self.entries.forget_all(), + } + } + + /// Keep nothing until [`Self::follow_again`]: changes are not heard. + fn stop_following(&self) { + let mut generation = self.generation.lock().expect("held copy generation"); + self.following.store(false, Ordering::Release); + *generation += 1; + self.entries.forget_all(); + } + + /// Changes are heard again: start from nothing and keep again. + fn follow_again(&self) { + let mut generation = self.generation.lock().expect("held copy generation"); + *generation += 1; + self.entries.forget_all(); + self.following.store(true, Ordering::Release); + } +} + +async fn follow(copy: Weak>, mut subscription: Subscription, channel: &'static str, key_of: fn(&str) -> Changed) +where + K: Hash + Eq + Clone + Send + Sync + 'static, + V: Send + Sync + 'static, +{ + loop { + let heard = subscription.next().await; + let Some(copy) = copy.upgrade() else { return }; + match heard { + Ok(Heard::Signal { channel: heard_on, payload }) if heard_on == channel => copy.drop_changed(key_of(&payload)), + Ok(Heard::Signal { .. }) => {} + // A recheck can also mean this copy fell behind and missed what + // was said, a lost connection included: whether the watch + // listens now is what decides. + Ok(Heard::Recheck) if subscription.listening() => copy.follow_again(), + Ok(Heard::Recheck | Heard::Lost) => copy.stop_following(), + Err(e) => { + tracing::error!( + target: "weft_task_store::held_copy", + channel, error = %format!("{e:#}"), + "the copy of rows on this channel stopped following them; every read goes to the database from now on" + ); + copy.stop_following(); + return; + } + } + } +} + +#[cfg(test)] +mod tests { + use super::*; + use tokio::sync::broadcast; + + const CHANNEL: &str = "rows"; + + fn by_payload(payload: &str) -> Changed { + match payload { + "" => Changed::Everything, + key => Changed::Key(key.to_string()), + } + } + + fn copy() -> (broadcast::Sender, Arc>) { + let (tx, rx) = broadcast::channel(8); + (tx, HeldCopy::following(Subscription::from(rx), CHANNEL, 16, by_payload, |_| true)) + } + + async fn read(copy: &HeldCopy, key: &str, value: u32) -> u32 { + *copy.get_or_load(key.to_string(), || async move { Ok::<_, ()>(value) }).await.unwrap() + } + + fn generation(copy: &HeldCopy) -> u64 { + *copy.generation.lock().unwrap() + } + + /// Send `heard` and wait until the follower acted on it. + async fn hear(tx: &broadcast::Sender, copy: &HeldCopy, heard: Heard) { + let before = generation(copy); + tx.send(heard).unwrap(); + while generation(copy) == before { + tokio::task::yield_now().await; + } + } + + weft_core::stress_test!( + name: a_held_key_is_answered_from_memory_until_its_rows_change, + runs: 32, + worker_threads: 4, + async fn body() { + let (tx, copy) = copy(); + assert_eq!(read(©, "a", 1).await, 1); + assert_eq!(read(©, "a", 2).await, 1, "held"); + assert_eq!(read(©, "b", 7).await, 7); + hear(&tx, ©, Heard::Signal { channel: CHANNEL, payload: "a".into() }).await; + assert_eq!(read(©, "a", 2).await, 2, "its rows changed, so it is read again"); + assert_eq!(read(©, "b", 8).await, 7, "another key's change leaves this one held"); + hear(&tx, ©, Heard::Recheck).await; + assert_eq!(read(©, "b", 9).await, 9, "a recheck drops everything"); + } + ); + + weft_core::stress_test!( + name: another_channel_changes_nothing, + runs: 32, + worker_threads: 4, + async fn body() { + let (tx, copy) = copy(); + read(©, "a", 1).await; + tx.send(Heard::Signal { channel: "other", payload: "a".into() }).unwrap(); + // Then one it does act on, so the first was surely heard. + hear(&tx, ©, Heard::Signal { channel: CHANNEL, payload: "b".into() }).await; + assert_eq!(read(©, "a", 2).await, 1); + } + ); + + // A read that started before a change is answered but not kept: it + // may have read the rows from before the change. + weft_core::stress_test!( + name: a_read_overtaken_by_a_change_is_not_kept, + runs: 32, + worker_threads: 4, + async fn body() { + let (tx, copy) = copy(); + let (started_tx, started) = tokio::sync::oneshot::channel::<()>(); + let (go, wait) = tokio::sync::oneshot::channel::<()>(); + let reading = { + let copy = copy.clone(); + tokio::spawn(async move { + *copy + .get_or_load("a".to_string(), || async move { + started_tx.send(()).unwrap(); + wait.await.unwrap(); + Ok::<_, ()>(1) + }) + .await + .unwrap() + }) + }; + started.await.unwrap(); + hear(&tx, ©, Heard::Signal { channel: CHANNEL, payload: "a".into() }).await; + go.send(()).unwrap(); + assert_eq!(reading.await.unwrap(), 1, "the read is still answered"); + assert_eq!(read(©, "a", 2).await, 2, "but it was not kept"); + } + ); + + // While the listening connection is down nothing says what changed, + // so the copy keeps nothing; listening again keeps again, from + // nothing. + weft_core::stress_test!( + name: a_lost_connection_keeps_nothing_until_it_listens_again, + runs: 32, + worker_threads: 4, + async fn body() { + let (tx, copy) = copy(); + read(©, "a", 1).await; + hear(&tx, ©, Heard::Lost).await; + assert_eq!(read(©, "a", 2).await, 2, "nothing held"); + assert_eq!(read(©, "a", 3).await, 3, "and nothing kept"); + hear(&tx, ©, Heard::Recheck).await; + assert_eq!(read(©, "a", 4).await, 4); + assert_eq!(read(©, "a", 5).await, 4, "kept again"); + } + ); + + // A copy that fell behind and missed the connection being lost reads + // the watch's state on the recheck that tells it it fell behind. + weft_core::stress_test!( + name: a_copy_that_missed_the_loss_reads_whether_the_watch_listens, + runs: 32, + worker_threads: 4, + async fn body() { + let (tx, rx) = broadcast::channel(8); + let listening = Arc::new(std::sync::atomic::AtomicBool::new(true)); + let copy: Arc> = + HeldCopy::following(Subscription::with_listening(rx, listening.clone()), CHANNEL, 16, by_payload, |_| true); + read(©, "a", 1).await; + listening.store(false, Ordering::Release); + hear(&tx, ©, Heard::Recheck).await; + assert_eq!(read(©, "a", 2).await, 2); + assert_eq!(read(©, "a", 3).await, 3, "nothing kept while the watch is down"); + listening.store(true, Ordering::Release); + hear(&tx, ©, Heard::Recheck).await; + read(©, "a", 4).await; + assert_eq!(read(©, "a", 5).await, 4, "kept again"); + } + ); + + // A read nobody will ask for again is answered and not kept. + weft_core::stress_test!( + name: a_read_not_worth_keeping_is_answered_and_not_kept, + runs: 32, + worker_threads: 4, + async fn body() { + let (_tx, rx) = broadcast::channel(8); + let copy: Arc> = HeldCopy::following(Subscription::from(rx), CHANNEL, 16, by_payload, |v| *v != 0); + assert_eq!(read(©, "a", 0).await, 0); + assert_eq!(read(©, "a", 1).await, 1, "the empty answer was not kept"); + assert_eq!(read(©, "a", 2).await, 1); + } + ); + + // A watch that stopped can no longer say what changed, so the copy + // keeps nothing from then on. + weft_core::stress_test!( + name: a_stopped_watch_turns_the_copy_off, + runs: 32, + worker_threads: 4, + async fn body() { + let (tx, copy) = copy(); + read(©, "a", 1).await; + drop(tx); + while copy.following.load(Ordering::Acquire) { + tokio::task::yield_now().await; + } + assert_eq!(read(©, "a", 2).await, 2); + assert_eq!(read(©, "a", 3).await, 3, "nothing is kept any more"); + } + ); +} diff --git a/crates/weft-task-store/src/lib.rs b/crates/weft-task-store/src/lib.rs index 0dfc1b97..7fd48784 100644 --- a/crates/weft-task-store/src/lib.rs +++ b/crates/weft-task-store/src/lib.rs @@ -18,6 +18,8 @@ //! - `pg_signal`: the process's one Postgres `LISTEN` connection, //! which every wait on a row sleeps on (`terminal` is the task //! waiter built on it). +//! - `held_copy`: a process's copy of rows read on every request, +//! dropped the moment a notification says they changed. //! - `schema_guard`: the schema runner every boot routes its //! `SchemaGroup`s through. It builds a new database from the canonical //! `CREATE TABLE` text and carries an existing one forward with the @@ -27,6 +29,7 @@ pub mod alarm; pub mod db; pub mod drain; pub mod executor; +pub mod held_copy; pub mod kinds; pub mod locks; pub mod pg_signal; diff --git a/crates/weft-task-store/src/pg_signal.rs b/crates/weft-task-store/src/pg_signal.rs index 923547fe..4f554a18 100644 --- a/crates/weft-task-store/src/pg_signal.rs +++ b/crates/weft-task-store/src/pg_signal.rs @@ -10,7 +10,8 @@ //! poll. //! //! A notification sent while the listening connection is down is lost, -//! so a reconnect tells every waiter to read its row again: nothing that +//! so the watch says when it loses the connection ([`Heard::Lost`]) and a +//! reconnect tells every waiter to read its row again: nothing that //! changed during the gap is missed. The row is the truth; a //! notification only says when to look. Every waiter follows the same //! recipe: subscribe, read the row, wait for a signal that concerns it @@ -26,6 +27,7 @@ //! cannot listen is refused at once, naming the fix, rather than leaving //! every waiter to sleep to its deadline. +use std::sync::atomic::{AtomicBool, Ordering}; use std::sync::{Arc, Mutex}; use std::time::Duration; @@ -45,6 +47,11 @@ pub const MAX_HOLD: Duration = Duration::from_secs(25); pub enum Heard { Signal { channel: &'static str, payload: Arc }, Recheck, + /// The listening connection is gone: nothing is heard until a + /// [`Heard::Recheck`] says it listens again. A waiter that reads its + /// row anyway has nothing to do; a copy of rows kept in memory + /// (`crate::held_copy`) stops trusting itself until then. + Lost, } /// How many signals a slow waiter may fall behind by before it is told @@ -74,6 +81,11 @@ pub struct PgSignalWatch { /// sender, so if the pump ever stops, every waiter hears `Closed` /// and fails loudly instead of sleeping to its deadline. template: Mutex>, + /// Whether the watch listens right now: cleared before + /// [`Heard::Lost`] goes out, set again before the [`Heard::Recheck`] + /// that follows a reconnect. A subscriber that fell behind and missed + /// either reads the state here. + listening: Arc, /// The listening task. It lives exactly as long as the watch: when /// the process part that owns the watch goes, so does its connection /// (a test's database cannot be dropped while it is held). @@ -100,8 +112,9 @@ impl PgSignalWatch { .await?; let listener = listen(&own, channels).await?; let (tx, template) = broadcast::channel(FANOUT_CAPACITY); - let pump = tokio::spawn(pump(listener, own, channels, tx)); - Ok(Arc::new(Self { channels, template: Mutex::new(template), pump })) + let listening = Arc::new(AtomicBool::new(true)); + let pump = tokio::spawn(pump(listener, own, channels, tx, listening.clone())); + Ok(Arc::new(Self { channels, template: Mutex::new(template), listening, pump })) } /// Subscribe BEFORE reading the row: a change that lands between the @@ -109,6 +122,7 @@ impl PgSignalWatch { pub fn subscribe(&self) -> Subscription { Subscription { rx: self.template.lock().expect("the template receiver is never poisoned").resubscribe(), + listening: self.listening.clone(), } } @@ -129,13 +143,28 @@ impl PgSignalWatch { /// One waiter's view of the watch. pub struct Subscription { rx: broadcast::Receiver, + listening: Arc, } /// Any source of what a watch hears; a test drives a waiter by sending -/// into the other end. +/// into the other end. It counts as listening. impl From> for Subscription { fn from(rx: broadcast::Receiver) -> Self { - Self { rx } + Self::with_listening(rx, Arc::new(AtomicBool::new(true))) + } +} + +impl Subscription { + /// A subscription whose watch's listening state is `listening` (a + /// test sets it as a watch's pump would). + pub fn with_listening(rx: broadcast::Receiver, listening: Arc) -> Self { + Self { rx, listening } + } + + /// Whether the watch listens right now (see [`PgSignalWatch`]'s + /// `listening`). + pub fn listening(&self) -> bool { + self.listening.load(Ordering::Acquire) } } @@ -193,6 +222,8 @@ impl Subscription { fn wakes(heard: &Heard, concerns: &impl Fn(&str, &str) -> bool) -> bool { match heard { Heard::Recheck => true, + // The recheck that follows a reconnect is what wakes the waiter. + Heard::Lost => false, Heard::Signal { channel, payload } => concerns(channel, payload), } } @@ -237,6 +268,7 @@ async fn pump( own: PgPool, channels: &'static [&'static str], tx: broadcast::Sender, + listening: Arc, ) { loop { let e = hear(&mut listener, channels, &tx).await; @@ -249,6 +281,8 @@ async fn pump( // one could never get it while it lives. tracing::warn!(target: "weft_task_store::pg_signal", error = %e, "signal listener failed; listening again"); drop(listener); + listening.store(false, Ordering::Release); + let _ = tx.send(Heard::Lost); listener = loop { tokio::time::sleep(RETRY_DELAY).await; match listen(&own, channels).await { @@ -256,6 +290,7 @@ async fn pump( Err(e) => tracing::warn!(target: "weft_task_store::pg_signal", error = %e, "signal listener cannot listen yet"), } }; + listening.store(true, Ordering::Release); let _ = tx.send(Heard::Recheck); } } @@ -282,10 +317,12 @@ async fn hear( ), } } - // The connection dropped and the listener has already made a - // fresh one and listened again; what changed in between was - // not heard. - Ok(None) => { let _ = tx.send(Heard::Recheck); } + // The connection dropped. The listener would make a fresh one + // on its next call, unannounced and without proving it hears, + // so it is handed back as a failure: the pump says the + // connection is lost, listens again the way it first did, and + // tells everyone to look again. + Ok(None) => return sqlx::Error::Io(std::io::Error::other("the listening connection dropped")), Err(e) => return e, } } diff --git a/crates/weft-task-store/src/tasks.rs b/crates/weft-task-store/src/tasks.rs index 37176ed9..cc22d402 100644 --- a/crates/weft-task-store/src/tasks.rs +++ b/crates/weft-task-store/src/tasks.rs @@ -253,6 +253,44 @@ pub static GROUP: crate::SchemaGroup = crate::SchemaGroup { r#"CREATE INDEX IF NOT EXISTS idx_task_terminal_completed ON task(completed_at_unix) WHERE status IN ('complete', 'failed')"#, + // The lock that serializes producers of one live task, held until + // the transaction ends: the ONE spelling of its key, taken by + // `weft_enqueue_dedup` and `enqueue_or_rearm`. + r#"CREATE OR REPLACE FUNCTION weft_lock_dedup(p_tenant TEXT, p_kind TEXT, p_dedup TEXT) RETURNS VOID AS $$ + BEGIN + PERFORM pg_advisory_xact_lock(hashtextextended(p_tenant || '|' || p_kind || '|' || p_dedup, 0)); + END; + $$ LANGUAGE plpgsql"#, + // THE dedup'd enqueue ([`enqueue_dedup_in`], and a run's birth, + // `weft_start_execution`): the task already live under + // `(tenant, kind, dedup_key)`, or the new one. A transaction-scoped + // lock on that triple serializes two producers of the same task, so + // the second finds the first's row instead of tripping the unique + // index. Run inside the caller's transaction, which the lock lasts. + r#"CREATE OR REPLACE FUNCTION weft_enqueue_dedup( + p_id UUID, p_kind TEXT, p_target TEXT, p_project UUID, p_dedup TEXT, p_execution TEXT, + p_tenant TEXT, p_target_replica TEXT, p_binary_hash TEXT, p_payload JSONB, p_now BIGINT, + OUT task_id UUID, OUT inserted BOOLEAN + ) AS $$ + BEGIN + PERFORM weft_lock_dedup(p_tenant, p_kind, p_dedup); + SELECT t.id INTO task_id FROM task t + WHERE t.tenant_id = p_tenant AND t.kind = p_kind AND t.dedup_key = p_dedup + AND t.status IN ('pending', 'claimed') + LIMIT 1; + IF FOUND THEN + inserted := FALSE; + RETURN; + END IF; + INSERT INTO task ( + id, kind, status, target, project_id, dedup_key, execution_id, tenant_id, + target_replica, binary_hash, payload, attempts, created_at_unix + ) VALUES (p_id, p_kind, 'pending', p_target, p_project, p_dedup, p_execution, p_tenant, + p_target_replica, p_binary_hash, p_payload, 0, p_now); + task_id := p_id; + inserted := TRUE; + END; + $$ LANGUAGE plpgsql"#, // Announce every task that has just become claimable, so the // pickers sleep until there is work instead of asking on a // timer. From a trigger rather than from each writer, so no @@ -378,76 +416,77 @@ pub async fn enqueue(pool: &PgPool, spec: NewTask) -> Result { /// the lock, two producers could both pass the SELECT (their snapshots /// don't see each other's uncommitted INSERT) and the second would hit /// a unique-violation on the partial index instead of returning -/// AlreadyLive. +/// AlreadyLive. The whole of it is one statement (`weft_enqueue_dedup`), +/// so on its own it is its own transaction: one round trip. pub async fn enqueue_dedup(pool: &PgPool, spec: NewTask) -> Result { - let mut tx = pool.begin().await?; - let outcome = enqueue_dedup_in(&mut tx, spec).await?; - tx.commit().await?; - Ok(outcome) + enqueue_dedup_in(&mut *pool.acquire().await?, spec).await } -/// [`enqueue_dedup`] on a caller-owned connection. MUST run inside a -/// transaction: the advisory lock is xact-scoped (it releases when the -/// caller's transaction ends), and the insert's atomicity with whatever else -/// the caller writes is the whole point of taking a connection. +/// [`enqueue_dedup`] on a caller-owned connection: inside the caller's +/// transaction, the insert commits with whatever else it writes, and the +/// advisory lock lasts until it ends. pub async fn enqueue_dedup_in( conn: &mut sqlx::PgConnection, spec: NewTask, ) -> Result { - let dedup = spec - .dedup_key - .as_deref() - .ok_or_else(|| anyhow::anyhow!("enqueue_dedup requires dedup_key"))?; - - // Lock + SELECT are scoped by tenant_id to match the - // `(tenant_id, kind, dedup_key)` unique index: dedup never crosses - // a tenant boundary. (`tenant_id IS NOT DISTINCT FROM $3` so a - // NULL-tenant task dedups against other NULL-tenant tasks, matching - // how the unique index treats them.) - let tenant = spec.tenant_id.as_str(); - let lock_input = format!("{}|{}|{}", tenant, spec.kind.as_str(), dedup); - sqlx::query("SELECT pg_advisory_xact_lock(hashtextextended($1, 0))") - .bind(&lock_input) - .execute(&mut *conn) - .await?; - - let existing: Option<(Uuid,)> = sqlx::query_as( - r#"SELECT id FROM task - WHERE tenant_id IS NOT DISTINCT FROM $1 - AND kind = $2 AND dedup_key = $3 AND status IN ('pending', 'claimed') - LIMIT 1"#, + let row = DedupRow::of(&spec, Uuid::new_v4(), unix_now())?; + let (id, inserted): (Uuid, bool) = sqlx::query_as( + "SELECT task_id, inserted FROM weft_enqueue_dedup($1, $2, $3, $4, $5, $6, $7, $8, $9, $10, $11)", ) - .bind(spec.tenant_id.as_str()) - .bind(spec.kind.as_str()) - .bind(dedup) - .fetch_optional(&mut *conn) + .bind(row.id) + .bind(&row.kind) + .bind(row.target) + .bind(row.project_id) + .bind(&row.dedup_key) + .bind(&row.execution_id) + .bind(&row.tenant_id) + .bind(&row.target_replica) + .bind(&row.binary_hash) + .bind(&row.payload) + .bind(row.created_at) + .fetch_one(&mut *conn) .await?; - if let Some((id,)) = existing { - return Ok(DedupOutcome::AlreadyLive(id)); - } + Ok(if inserted { DedupOutcome::Inserted(id) } else { DedupOutcome::AlreadyLive(id) }) +} - let id = Uuid::new_v4(); - let now = unix_now(); - sqlx::query( - r#"INSERT INTO task ( - id, kind, status, target, project_id, dedup_key, execution_id, tenant_id, - target_replica, binary_hash, payload, attempts, created_at_unix - ) VALUES ($1, $2, 'pending', $3, $4, $5, $6, $7, $8, $9, $10, 0, $11)"#, - ) - .bind(id) - .bind(spec.kind.as_str()) - .bind(spec.target.as_str()) - .bind(spec.project_id) - .bind(dedup) - .bind(spec.execution_id.as_deref()) - .bind(spec.tenant_id.as_str()) - .bind(spec.target_replica.as_deref()) - .bind(spec.binary_hash.as_deref()) - .bind(&spec.payload) - .bind(now) - .execute(&mut *conn) - .await?; - Ok(DedupOutcome::Inserted(id)) +/// A dedup'd task as `weft_enqueue_dedup` takes it: what +/// [`enqueue_dedup_in`] binds, and what a run's birth hands the database +/// whole (`weft_start_execution`'s `task`). +// SYNC: DedupRow's fields <-> weft_enqueue_dedup's parameters, weft_start_execution's `task` +#[derive(Debug, Clone, Serialize)] +pub struct DedupRow { + pub id: Uuid, + pub kind: String, + pub target: &'static str, + pub project_id: Option, + pub dedup_key: String, + pub execution_id: Option, + pub tenant_id: String, + pub target_replica: Option, + pub binary_hash: Option, + pub payload: Value, + pub created_at: i64, +} + +impl DedupRow { + /// `spec` as the dedup'd row `id`, enqueued at `now`. Refused without + /// a dedup key: nothing would collapse onto it. + pub fn of(spec: &NewTask, id: Uuid, now: i64) -> Result { + let dedup_key = spec.dedup_key.clone().ok_or_else(|| anyhow::anyhow!("enqueue_dedup requires dedup_key"))?; + Ok(Self { + id, + kind: spec.kind.clone(), + target: spec.target.as_str(), + project_id: spec.project_id, + dedup_key, + execution_id: spec.execution_id.clone(), + tenant_id: spec.tenant_id.clone(), + target_replica: spec.target_replica.clone(), + binary_hash: spec.binary_hash.clone(), + payload: spec.payload.clone(), + created_at: now, + }) + } } /// [`enqueue_dedup`] for a task that means "go and look again": a pending @@ -465,9 +504,10 @@ pub async fn enqueue_or_rearm(pool: &PgPool, spec: NewTask) -> Result Vec { match next.expect("the watch is running") { Heard::Signal { channel, payload } if channel == TERMINAL_CHANNEL => ids.push(payload.to_string()), Heard::Signal { .. } => {} - Heard::Recheck => panic!("a recheck would hide which notifications were sent"), + Heard::Recheck | Heard::Lost => panic!("a lost connection would hide which notifications were sent"), } } ids diff --git a/docs/src/nodes/durable-execution.md b/docs/src/nodes/durable-execution.md index b827c87d..d854b119 100644 --- a/docs/src/nodes/durable-execution.md +++ b/docs/src/nodes/durable-execution.md @@ -138,10 +138,15 @@ checked what it did. If the node sets `catchErrors` and you wired its `error` output, this failure goes there like any other, and that branch carries on. +A step's start is written down in the background, as the step begins (go and +read [the journal](../running/the-journal.md#a-run-does-not-wait-for-its-writes)). +If the worker goes away in the moment before that write lands, nothing says the +step ever started, and the next worker runs it as a step that never did. That +moment is one write to the database long. + If the body was waiting on `ctx.await_signal` when the worker went away, it is -not failed: when its answer comes, it replays from the top. This is the only way a body runs -twice, and its `ctx.run` calls give back their saved results instead of doing -the work again. +not failed: when its answer comes, it replays from the top, and its `ctx.run` +calls give back their saved results instead of doing the work again. If a step emitted a value before the crash and the step reading it had not started yet, that value is still delivered and the reading step runs. diff --git a/docs/src/running/architecture.md b/docs/src/running/architecture.md index 995af297..4fc8257b 100644 --- a/docs/src/running/architecture.md +++ b/docs/src/running/architecture.md @@ -50,8 +50,8 @@ Every queued run also carries a key that turns a duplicate into a no-op, so work queued twice runs once. A run whose claim ran out may be half done, and the next worker may redo part of the runtime's own bookkeeping for it, which is why every step the runtime takes has to be safe to run twice. Your nodes are different: a node -that was running when its worker died is failed, never run again (go and read -[surviving a restart](../nodes/durable-execution.md#when-the-worker-dies-mid-step)). +whose start is on record when its worker died is failed, never run again (go +and read [surviving a restart](../nodes/durable-execution.md#when-the-worker-dies-mid-step)). ## What happens when an event arrives @@ -98,6 +98,17 @@ a compiled definition. The next request may land on another copy of the dispatcher, so it keeps nothing in memory that another copy would need. +What it does keep is a copy of the rows a call reads every time and that +rarely change: a tenant's routes, the install's domains, a project's worker +settings, which of its infrastructure is up. Every write to those rows makes +Postgres tell every copy of the dispatcher, which drops what it held and reads +it again on the next call, and anything it is about to refuse (no such route, +infrastructure not running) it checks against the rows first. While its +connection that hears those announcements is down, it keeps nothing and reads +every time. So a live call +reaches the database once before the worker has it: one call that checks the +route's limits and writes the run down together. + ## The listener One listener serves every project on the install. Each kind of trigger says diff --git a/docs/src/running/cli.md b/docs/src/running/cli.md index 279f7b68..dff4aed3 100644 --- a/docs/src/running/cli.md +++ b/docs/src/running/cli.md @@ -90,7 +90,7 @@ reads the file again. | `weft executions` | Past runs, newest first. `--limit` (50), `--offset`, `--project`, `--phase`, `--node`, `--since 2h`, `--status`, `--instance`, `--tag`. A run of a `Route` with `recorded: false` shows only if it failed | | `weft events ` | One run's events in order. `--node`, `--kind`, `--full` for whole values. If you want one time round a loop, `--iteration 3` keeps the fourth (they count from 0), and `3.0` the first time round a loop inside it | | `weft logs []` | What the nodes wrote, plus every failure. A run that wrote nothing lists what it skipped and why (under `skipped` with `--json`). `--limit` | -| `weft status` | The cwd project: registration, listener, every infra node the program declares (one never started says `not started`, a `@per_instance` one how many instances have a copy), each trigger (with the events an instance's trigger holds until a field is filled), recent runs, what drifted, and what you can do next. While the install builds, it names each image building and where its log is; while infra starts, how long it has been going and what it waits on | +| `weft status` | The cwd project: registration, listener, every infra node the program declares (one never started says `not started`, a `@per_instance` one how many instances have a copy), each trigger and the version of the source it fires (with the events an instance's trigger holds until a field is filled), recent runs, what drifted, and what you can do next. While the install builds, it names each image building and where its log is; while infra starts, how long it has been going and what it waits on | | `weft ps` | Every project the dispatcher knows | `--phase fire` hides the setup runs that an activate or a resync makes, so it @@ -195,7 +195,7 @@ all take `--instance `. | `weft ci add --cloud gcp` | Writes `.github/workflows/deploy.yml`. Refuses to replace one somebody edited | | `weft target export [--github] [--frontend ]` | Mints a CI operator key on the install, and gives the frontend the install hosts for this repository a new token along with its service's name (the old token works until the workflow's next run deploys the new one and retires it). With `--github` it sets them, and the workflow's other settings, on the repository through `gh`, and lists what it set; without it, it prints the variables and writes the secrets to a file only you can read, under the install's `exports/` folder | | `weft target show ` | Where a target's install lives: its address, and on a cloud the project and region it runs in, and the weft it runs | -| `weft frontend add [--repo owner/name]` | Makes one of the project's frontends and its token, written to a file only you can read. With `--repo` (read through `gh`), a cloud install makes it a Cloud Run service and lets that repository deploy to it; without, it runs wherever you run it | +| `weft frontend add [--repo owner/name]` | Makes one of the project's frontends. With `--repo` (read through `gh`), a cloud install makes it a Cloud Run service and lets that repository deploy to it, and its token comes from the deploy workflow (`weft target export`); without, it runs wherever you run it, and its token is written to a file only you can read | | `weft frontend ls` | The project's frontends, and where each runs | | `weft frontend rm [--force]` | Removes one: its tokens stop working, and a service the install made for it is deleted. `--force` forgets it even when the service cannot be removed, naming what stays | | `weft frontend token [--done ]` | A new token, beside the one it has, with its id; `--done ` once that token is in place retires every other token of the frontend | diff --git a/docs/src/running/cloud.md b/docs/src/running/cloud.md index 938d3dee..ce81448a 100644 --- a/docs/src/running/cloud.md +++ b/docs/src/running/cloud.md @@ -430,9 +430,10 @@ project, `weft new --ci gcp` writes the workflow for you. `weft target export prod --github` uses the GitHub CLI (`gh`) to set the variables and secrets the workflow reads, so run it with `gh` logged in. It mints an operator key for the workflow, and gives the frontend the install -hosts for this repository a new token, with the name of its service. The -old token keeps working until the workflow's next run has deployed the new -one, and then the workflow retires it. If you would rather paste them yourself, leave out +hosts for this repository a new token, with the name of its service (its +first one: `weft frontend add` makes none for a hosted frontend). Once it has +been deployed, its old token keeps working until the workflow's next run has +deployed the new one, and then the workflow retires it. If you would rather paste them yourself, leave out `--github` and it prints everything. If you use the printed version, copy the secrets straight away: they are never shown again. @@ -442,7 +443,7 @@ Cloud Run: | Variable | What it is | Value on your machine | |---|---|---| | `WEFT_DISPATCHER_URL` | where the server calls weft: the install's address, over HTTPS | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the token the server calls with; never send it to a browser | a token from `weft frontend add `, written to a file only you can read | +| `WEFT_TOKEN` | the token the server calls with; never send it to a browser (on Cloud Run, the one `weft target export` made) | a token from `weft frontend add `, written to a file only you can read | | `WEFT_PUBLIC_URL` | the start of any link a browser follows | `http://127.0.0.1:14111` | If the frontend's server needs more than those three (the program's diff --git a/docs/src/running/the-journal.md b/docs/src/running/the-journal.md index 50379075..bc1d54c2 100644 --- a/docs/src/running/the-journal.md +++ b/docs/src/running/the-journal.md @@ -73,14 +73,24 @@ works, what `catchErrors` does with this failure, and what happens when `ctx.run` cannot save a result, go and read [surviving a restart](../nodes/durable-execution.md). -## Why a failed write stops the worker +## A run does not wait for its writes + +A worker hands each event to a sender of its own and carries on: the rows +reach the database in the background, in the order they happened. The run +waits for them only where something else is about to read or act on its +record: before it reads its own journal, before it hands anything to the +dispatcher (a task, a tag, a stop), and when it ends or pauses, so a run is +never reported finished before its record is whole. A write that fails stops +the run at once, for the reason below. + +## Why a failed write stops the run If the journal is missing rows the live worker believes it wrote, every later rebuild would reconstruct a different world: a node whose start was lost but whose emissions landed would run again and spend twice. -So a failed write ends that worker rather than carrying on with a record -nobody can trust. +So a failed write ends that run, as failed, rather than carrying on with a +record nobody can trust. ## Who may write diff --git a/extension-vscode/package.json b/extension-vscode/package.json index ddd082fd..8f16159c 100644 --- a/extension-vscode/package.json +++ b/extension-vscode/package.json @@ -2,7 +2,7 @@ "name": "weft-vscode", "displayName": "Weft", "description": "Write Weft programs and watch them run: a live, editable graph of every .weft file, with diagnostics, runs and executions in the editor.", - "version": "0.2.257", + "version": "0.2.258", "publisher": "weavemind", "license": "SEE LICENSE IN LICENSE", "repository": { diff --git a/packages/weft-graph/src/protocol.ts b/packages/weft-graph/src/protocol.ts index 4308c699..1eb9863e 100644 --- a/packages/weft-graph/src/protocol.ts +++ b/packages/weft-graph/src/protocol.ts @@ -1049,6 +1049,8 @@ export interface ActivationEntry { /// how many, and why (the refusal naming the field). // SYNC: waiting <-> crates/weft-core/src/program.rs WaitingFires waiting?: { fires: number; reason: string }; + /// The version of the source its fires run, once its setup finished. + version?: string; } /// One instance's copy of an infra node, as the project status lists it. diff --git a/tangle/claude-code/.claude/skills/weft-deploying/SKILL.md b/tangle/claude-code/.claude/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/claude-code/.claude/skills/weft-deploying/SKILL.md +++ b/tangle/claude-code/.claude/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/claude-code/.claude/skills/weft-node-authoring/SKILL.md b/tangle/claude-code/.claude/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/claude-code/.claude/skills/weft-node-authoring/SKILL.md +++ b/tangle/claude-code/.claude/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/claude-code/.claude/skills/weft-running/SKILL.md b/tangle/claude-code/.claude/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/claude-code/.claude/skills/weft-running/SKILL.md +++ b/tangle/claude-code/.claude/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/cline/.cline/skills/weft-deploying/SKILL.md b/tangle/cline/.cline/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/cline/.cline/skills/weft-deploying/SKILL.md +++ b/tangle/cline/.cline/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/cline/.cline/skills/weft-node-authoring/SKILL.md b/tangle/cline/.cline/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/cline/.cline/skills/weft-node-authoring/SKILL.md +++ b/tangle/cline/.cline/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/cline/.cline/skills/weft-running/SKILL.md b/tangle/cline/.cline/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/cline/.cline/skills/weft-running/SKILL.md +++ b/tangle/cline/.cline/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/codex/.agents/skills/weft-deploying/SKILL.md b/tangle/codex/.agents/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/codex/.agents/skills/weft-deploying/SKILL.md +++ b/tangle/codex/.agents/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/codex/.agents/skills/weft-node-authoring/SKILL.md b/tangle/codex/.agents/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/codex/.agents/skills/weft-node-authoring/SKILL.md +++ b/tangle/codex/.agents/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/codex/.agents/skills/weft-running/SKILL.md b/tangle/codex/.agents/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/codex/.agents/skills/weft-running/SKILL.md +++ b/tangle/codex/.agents/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/cursor/.cursor/skills/weft-deploying/SKILL.md b/tangle/cursor/.cursor/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/cursor/.cursor/skills/weft-deploying/SKILL.md +++ b/tangle/cursor/.cursor/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/cursor/.cursor/skills/weft-node-authoring/SKILL.md b/tangle/cursor/.cursor/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/cursor/.cursor/skills/weft-node-authoring/SKILL.md +++ b/tangle/cursor/.cursor/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/cursor/.cursor/skills/weft-running/SKILL.md b/tangle/cursor/.cursor/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/cursor/.cursor/skills/weft-running/SKILL.md +++ b/tangle/cursor/.cursor/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/devin-desktop/.windsurf/skills/weft-deploying/SKILL.md b/tangle/devin-desktop/.windsurf/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/devin-desktop/.windsurf/skills/weft-deploying/SKILL.md +++ b/tangle/devin-desktop/.windsurf/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/devin-desktop/.windsurf/skills/weft-node-authoring/SKILL.md b/tangle/devin-desktop/.windsurf/skills/weft-node-authoring/SKILL.md index 7a9373ab..8776eadd 100644 --- a/tangle/devin-desktop/.windsurf/skills/weft-node-authoring/SKILL.md +++ b/tangle/devin-desktop/.windsurf/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/devin-desktop/.windsurf/skills/weft-running/SKILL.md b/tangle/devin-desktop/.windsurf/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/devin-desktop/.windsurf/skills/weft-running/SKILL.md +++ b/tangle/devin-desktop/.windsurf/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/fallback/.agents/skills/weft-deploying/SKILL.md b/tangle/fallback/.agents/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/fallback/.agents/skills/weft-deploying/SKILL.md +++ b/tangle/fallback/.agents/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/fallback/.agents/skills/weft-node-authoring/SKILL.md b/tangle/fallback/.agents/skills/weft-node-authoring/SKILL.md index 7a9373ab..8776eadd 100644 --- a/tangle/fallback/.agents/skills/weft-node-authoring/SKILL.md +++ b/tangle/fallback/.agents/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/fallback/.agents/skills/weft-running/SKILL.md b/tangle/fallback/.agents/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/fallback/.agents/skills/weft-running/SKILL.md +++ b/tangle/fallback/.agents/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/gemini-cli/.gemini/skills/weft-deploying/SKILL.md b/tangle/gemini-cli/.gemini/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/gemini-cli/.gemini/skills/weft-deploying/SKILL.md +++ b/tangle/gemini-cli/.gemini/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/gemini-cli/.gemini/skills/weft-node-authoring/SKILL.md b/tangle/gemini-cli/.gemini/skills/weft-node-authoring/SKILL.md index 7a9373ab..8776eadd 100644 --- a/tangle/gemini-cli/.gemini/skills/weft-node-authoring/SKILL.md +++ b/tangle/gemini-cli/.gemini/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/gemini-cli/.gemini/skills/weft-running/SKILL.md b/tangle/gemini-cli/.gemini/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/gemini-cli/.gemini/skills/weft-running/SKILL.md +++ b/tangle/gemini-cli/.gemini/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/github-copilot/.github/skills/weft-deploying/SKILL.md b/tangle/github-copilot/.github/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/github-copilot/.github/skills/weft-deploying/SKILL.md +++ b/tangle/github-copilot/.github/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/github-copilot/.github/skills/weft-node-authoring/SKILL.md b/tangle/github-copilot/.github/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/github-copilot/.github/skills/weft-node-authoring/SKILL.md +++ b/tangle/github-copilot/.github/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/github-copilot/.github/skills/weft-running/SKILL.md b/tangle/github-copilot/.github/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/github-copilot/.github/skills/weft-running/SKILL.md +++ b/tangle/github-copilot/.github/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/junie/.junie/skills/weft-deploying/SKILL.md b/tangle/junie/.junie/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/junie/.junie/skills/weft-deploying/SKILL.md +++ b/tangle/junie/.junie/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/junie/.junie/skills/weft-node-authoring/SKILL.md b/tangle/junie/.junie/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/junie/.junie/skills/weft-node-authoring/SKILL.md +++ b/tangle/junie/.junie/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/junie/.junie/skills/weft-running/SKILL.md b/tangle/junie/.junie/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/junie/.junie/skills/weft-running/SKILL.md +++ b/tangle/junie/.junie/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/kilo-code/.kilo/skills/weft-deploying/SKILL.md b/tangle/kilo-code/.kilo/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/kilo-code/.kilo/skills/weft-deploying/SKILL.md +++ b/tangle/kilo-code/.kilo/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/kilo-code/.kilo/skills/weft-node-authoring/SKILL.md b/tangle/kilo-code/.kilo/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/kilo-code/.kilo/skills/weft-node-authoring/SKILL.md +++ b/tangle/kilo-code/.kilo/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/kilo-code/.kilo/skills/weft-running/SKILL.md b/tangle/kilo-code/.kilo/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/kilo-code/.kilo/skills/weft-running/SKILL.md +++ b/tangle/kilo-code/.kilo/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the diff --git a/tangle/opencode/.opencode/skills/weft-deploying/SKILL.md b/tangle/opencode/.opencode/skills/weft-deploying/SKILL.md index da67b9f9..d860a753 100644 --- a/tangle/opencode/.opencode/skills/weft-deploying/SKILL.md +++ b/tangle/opencode/.opencode/skills/weft-deploying/SKILL.md @@ -145,9 +145,10 @@ is granted to that id, so the name changing hands later gives nobody anything), makes the frontend's Cloud Run service (empty until the first deploy) and lets that repository's workflow deploy to that service and nothing else on the install. It prints the service's name and the -address visitors will reach. The export then hands the repository the -service's name and a fresh token for the frontend, and the next run of the -deploy workflow builds `front/` and puts it there. The order matters: an +address visitors will reach. It makes no token: the export makes the +frontend's first one and hands the repository the service's name and that +token, and the next run of the deploy workflow builds `front/`, puts it there +and puts the token in place. The order matters: an export run before the frontend exists hands over no frontend, and the workflow skips it. @@ -232,7 +233,7 @@ Run and by `front/.env` on this machine: | Variable | On a cloud install | On this machine | |---|---|---| | `WEFT_DISPATCHER_URL` | the install's address | `http://127.0.0.1:14111` | -| `WEFT_TOKEN` | the frontend's own token, which `weft target export` renews and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | +| `WEFT_TOKEN` | the frontend's own token, which `weft target export` makes (the first one too) and hands the workflow | a token from `weft frontend add ` (no `--repo`), written to a private file | | `WEFT_PUBLIC_URL` | the install's public address | `http://127.0.0.1:14111` | On this machine, 14111 is the default port; if the install was started on diff --git a/tangle/opencode/.opencode/skills/weft-node-authoring/SKILL.md b/tangle/opencode/.opencode/skills/weft-node-authoring/SKILL.md index a6cc7fce..b332d921 100644 --- a/tangle/opencode/.opencode/skills/weft-node-authoring/SKILL.md +++ b/tangle/opencode/.opencode/skills/weft-node-authoring/SKILL.md @@ -316,9 +316,11 @@ stdlib's `nodes/base_catalog/ai/fal/fal.rs` (`run_queued`) for a whole worked ca You emit only through `ctx.pulse_downstream(NodeOutput::new().set(port, value))`; ports you did not emit are closed, which is the skip signal downstream. For -user-added output ports use `ctx.fan_declared(...)`. A step never runs twice -by itself: if the worker dies while a body runs, the step is failed (and -that failure goes to `error` like any other). The one body that runs again +user-added output ports use `ctx.fan_declared(...)`. A step whose start is on +record never runs twice by itself: if the worker dies while a body runs, the +step is failed (and that failure goes to `error` like any other). Its start is +written as it begins, so only a worker dying in that one write's time runs it +again as new. The one body that runs again is one parked on `ctx.await_signal`, which replays from the top when its answer comes, so the work before the wait goes through `ctx.run(...)`, which gives back the recorded result. diff --git a/tangle/opencode/.opencode/skills/weft-running/SKILL.md b/tangle/opencode/.opencode/skills/weft-running/SKILL.md index eddbd52a..e9b3635d 100644 --- a/tangle/opencode/.opencode/skills/weft-running/SKILL.md +++ b/tangle/opencode/.opencode/skills/weft-running/SKILL.md @@ -383,7 +383,7 @@ node and the wrong value. you never open the events. `(no logs: ...)` means the run wrote nothing and recorded no failure: it did not fail, so you check its status. - **If you want the values on the wires**: `weft events `. One line - per event: local time, kind, node, then everything the row carries as + per event: UTC time, kind, node, then everything the row carries as `key=value`, each cut to a screen's width (`input=` on `node_started`, `output=` on `node_completed`, `error=` on `node_failed`, `reason=` on a skip or a cancel, `token=` on a suspension, and so on). You narrow before @@ -443,8 +443,8 @@ node and the wrong value. --node --grant `) or send the user to the node's Connect button / `weft connect` in their terminal, then run again. A failure saying `the worker running '' went away while it was running` means the - worker died mid-step: weft never runs a step twice by itself, because the - step may have partly happened. Check what it did outside (the email, the + worker died mid-step: weft does not run a step again once its start is on + record, because the step may have partly happened. Check what it did outside (the email, the row, the post), then `weft run --seed`, which reuses what completed and runs that step again. 3. **A value is wrong, not failed.** Work upstream from the output: open the