From 894ddf0643677cd3b7307de13539326284cbfddb Mon Sep 17 00:00:00 2001 From: abrichr Date: Sat, 29 Aug 2026 01:13:06 -0400 Subject: [PATCH] docs: drop fail-closed slogan headings from public docs Rewrite human-authored docs voice to match the flow README: keep the gates, drop the memes. lint and certify still refuse; run still will not start without its admission checks. Headings no longer sell fail-closed as a posture. --- docs/commercial/acceptance-matrix.md | 2 +- docs/commercial/citrix-external-brief.md | 6 +- docs/commercial/deployment-boundaries.md | 2 +- docs/commercial/index.md | 2 +- docs/commercial/phi-handling.md | 18 +++--- docs/commercial/security-packet.md | 4 +- docs/commercial/subprocessors.md | 14 ++--- docs/concepts/backends.md | 4 +- docs/concepts/capability-ladder.md | 10 +-- docs/concepts/demonstration-compiler.md | 2 +- docs/concepts/durable-runtime.md | 33 +++++----- docs/concepts/effect-verification.md | 28 ++++----- docs/concepts/halt-learn-loop.md | 76 +++++++++++------------ docs/concepts/identity-gate.md | 38 ++++++------ docs/concepts/index.md | 5 +- docs/concepts/multi-trace-induction.md | 26 ++++---- docs/concepts/policy-and-certify.md | 52 +++++++--------- docs/concepts/regulated-execution.md | 67 +++++++++----------- docs/concepts/self-healing.md | 15 +++-- docs/concepts/settings-governance.md | 22 +++---- docs/concepts/substrate-model.md | 10 +-- docs/concepts/vlm-appliance.md | 4 +- docs/concepts/workflow-ir.md | 20 +++--- docs/get-started/index.md | 6 +- docs/get-started/what-works-today.md | 2 +- docs/guides/data-driven-loops.md | 2 +- docs/guides/deploy-on-prem.md | 6 +- docs/guides/hosted.md | 2 +- docs/guides/induce-a-program.md | 2 +- docs/guides/policy-and-certification.md | 4 +- docs/guides/run-reports.md | 2 +- docs/guides/security-and-data-handling.md | 10 +-- docs/guides/troubleshooting.md | 2 +- docs/llms.txt | 2 +- docs/reference/cli.md | 6 +- docs/reference/configuration.md | 16 ++--- docs/reference/glossary.md | 12 ++-- docs/reference/run-outcomes.md | 20 +++--- mkdocs.yml | 2 +- 39 files changed, 264 insertions(+), 292 deletions(-) diff --git a/docs/commercial/acceptance-matrix.md b/docs/commercial/acceptance-matrix.md index c133be3e..e1ef0ef1 100644 --- a/docs/commercial/acceptance-matrix.md +++ b/docs/commercial/acceptance-matrix.md @@ -24,7 +24,7 @@ Three outcomes are acceptable, one is not: Mechanics behind the outcomes: [effect verification](../concepts/effect-verification.md), [the identity gate](../concepts/identity-gate.md), -[fail-closed regulated execution](../concepts/regulated-execution.md). +[regulated execution](../concepts/regulated-execution.md). ## Template diff --git a/docs/commercial/citrix-external-brief.md b/docs/commercial/citrix-external-brief.md index ff386899..fab7c1bc 100644 --- a/docs/commercial/citrix-external-brief.md +++ b/docs/commercial/citrix-external-brief.md @@ -30,8 +30,8 @@ other substrate: fresh pixels, resolved target, and record identity, then a one-shot input lease refuses delivery if **anything** changed between resolution and the first input edge. -- Identity checks run on the pixel tier; ambiguity deliberately over-halts - rather than guessing. A collapsible or unreadable identifier halts the run. +- Identity checks run on the pixel tier; ambiguity over-halts + instead of guessing. A collapsible or unreadable identifier halts the run. - Business effects are verified out of band where a read path exists (API, database, report export, or a read-only second session). Screen-only confirmation is labeled as such, and high-risk workflows may not qualify on @@ -39,7 +39,7 @@ other substrate: - Halts are durable: the run pauses for a human with the violated expectation named in the report. -## Honest status +## Current qualification status Two claims, kept separate: diff --git a/docs/commercial/deployment-boundaries.md b/docs/commercial/deployment-boundaries.md index 5490e7b2..cb79f598 100644 --- a/docs/commercial/deployment-boundaries.md +++ b/docs/commercial/deployment-boundaries.md @@ -82,7 +82,7 @@ flowchart LR the runner's local sensitive-data policy. Verification prefers an independent read path; where only the screen is available, that is labeled same-surface confirmation and high-risk workflows may not qualify on it. -- **Status:** see the honest status statement in the +- **Status:** see the current qualification status in the [external Citrix brief](citrix-external-brief.md). ## Rules that hold in every lane diff --git a/docs/commercial/index.md b/docs/commercial/index.md index ed449241..48946f87 100644 --- a/docs/commercial/index.md +++ b/docs/commercial/index.md @@ -34,7 +34,7 @@ the evidence supports a "do not automate" decision. | [OpenAdapt Execute private-pilot guide](oem-brief.md) | A vendor that wants to embed verified execution. The API is available to approved private-pilot partners with scoped credentials. | | [Procurement FAQ](procurement-faq.md) | Procurement, legal, and vendor-risk questions. | -## Honesty rules for this section +## How claims are bounded These documents follow the same rules as the public site: diff --git a/docs/commercial/phi-handling.md b/docs/commercial/phi-handling.md index 900d0a45..ec6f28ce 100644 --- a/docs/commercial/phi-handling.md +++ b/docs/commercial/phi-handling.md @@ -22,8 +22,8 @@ engine: - There is no silent plaintext: in the default mode without the privacy extra, writing identity-like free text emits an explicit `PlaintextPHIWarning`. -**The honest boundary.** The recorded identity evidence and the identity audit -trail intentionally retain literal identifiers — scrubbing them would defeat +**The identity boundary.** The recorded identity evidence and the identity audit +trail intentionally retain literal identifiers; scrubbing them would defeat the wrong-record check they exist to power. Those artifacts are governed as PHI-at-rest **inside your boundary** (filesystem controls, retention, full-disk encryption, opt-in AES-256-GCM sealing), and the published privacy @@ -43,12 +43,12 @@ data and the wire: can the derivative be pushed, and the control plane verifies the manifest, review state, and exact archive SHA-256 before accepting a byte. -Content the sanitizer cannot fully handle — databases, video, audio, nested -archives, symlinks, unknown binaries — refuses the **entire** derivative +Content the sanitizer cannot fully handle (databases, video, audio, nested +archives, symlinks, unknown binaries) refuses the **entire** derivative rather than passing through. Sanitizer success is not treated as proof of de-identification: the operator review is the gate. -## 3. The receipt: an allow-list, not a redaction +## 3. The receipt is an allow-list The shareable run receipt is generated **additively from a closed allow-list, never redacted subtractively** from the rich operator report. Every field is a @@ -59,7 +59,7 @@ Structurally unrepresentable in a receipt: screenshots, OCR text, typed values, parameters, URLs, hostnames, coordinates, application name, organization name, user name, workflow name, step intents, and halt free text. The same principle governs the hosted attended-decision envelope (closed -enums, bounded integers, booleans — no string field, no image) and the hosted +enums, bounded integers, booleans; no string field, no image) and the hosted break-report descriptor (hashed, coarse, no free text). ## Where each artifact can live @@ -74,8 +74,8 @@ break-report descriptor (hashed, coarse, no free text). ## Related pages -- [Security packet](security-packet.md) — the reviewer summary. +- [Security packet](security-packet.md): the reviewer summary. - [Subprocessors and hosted data retention](subprocessors.md) -- [Fail-closed regulated execution](../concepts/regulated-execution.md) -- Engine [PRIVACY.md](https://github.com/OpenAdaptAI/openadapt-flow/blob/main/docs/PRIVACY.md) — +- [Regulated execution](../concepts/regulated-execution.md) +- Engine [PRIVACY.md](https://github.com/OpenAdaptAI/openadapt-flow/blob/main/docs/PRIVACY.md): the complete path-by-path PHI map. diff --git a/docs/commercial/security-packet.md b/docs/commercial/security-packet.md index d0ae6c6f..a4e245c3 100644 --- a/docs/commercial/security-packet.md +++ b/docs/commercial/security-packet.md @@ -1,6 +1,6 @@ # Security packet -The honest current state of OpenAdapt's security posture, written for a +Current OpenAdapt security posture, written for a security or vendor-risk reviewer. Deeper technical detail: [Security and data handling](../guides/security-and-data-handling.md) and the reviewer-oriented @@ -29,7 +29,7 @@ add an append-only, hash-chained audit log. read of the system of record. REFUTED and INDETERMINATE verdicts both halt; an unreachable verifier is never treated as success ([effect verification](../concepts/effect-verification.md)). -- **Fail-closed regulated path.** The `run` verb refuses to start without +- **Regulated `run` path.** The `run` verb refuses to start without certification, identity coverage, effect contracts or explicit approval, and encrypted, integrity-sealed bundles ([regulated execution](../concepts/regulated-execution.md)). diff --git a/docs/commercial/subprocessors.md b/docs/commercial/subprocessors.md index e97c9f99..b56de647 100644 --- a/docs/commercial/subprocessors.md +++ b/docs/commercial/subprocessors.md @@ -20,7 +20,7 @@ As read from the hosted control plane's deployment configuration: |---|---|---| | Netlify | Hosting for the `app.openadapt.ai` control plane | Application traffic to the control plane: account and session data in transit, and the metadata/digest surfaces described in the [security packet](security-packet.md). | | Supabase | Database, authentication, and object storage for the control plane | Accounts, organizations, workflow versions, run metadata, sanitized artifact derivatives, retention/erasure receipts. | -| Modal | Compute for the managed browser runner | Managed browser execution for explicitly initiated, public-HTTPS, non-regulated workloads only — not a lane for PHI/PII. | +| Modal | Compute for the managed browser runner | Managed browser execution for explicitly initiated, public-HTTPS, non-regulated workloads only; not a lane for PHI/PII. | | Stripe | Payments and billing | Payment and subscription data. Card data is entered on Stripe's surfaces, not OpenAdapt's. | | Resend | Transactional email (organization invites, purchase alerts) | Recipient email addresses and the fixed-template message content. Purchase alerts carry purchase metadata only, never workflow evidence. | | GitHub | Source hosting, CI, release distribution | Public source, build artifacts, and CI logs. No customer workload data. | @@ -35,7 +35,7 @@ not to your workflows. ## Hosted retention and deletion The hosted service applies a **versioned, explicitly configured retention -policy** — there is no implicit retention duration. The policy names its +policy**; there is no implicit retention duration. The policy names its version and sets explicit windows for recordings, reports, and run metadata (run metadata is never retained shorter than reports), plus a backup recovery window and a maximum restore-drill age. @@ -47,7 +47,7 @@ Current behavior: private object storage restore into an isolated scratch environment. - **Legal holds pause eligible deletion** for the held organization. - **Tenant erasure is organization-scoped** and produces an append-only, - PHI/PII-free receipt with identifiers, counts, and digests — never deleted + PHI/PII-free receipt with identifiers, counts, and digests, never deleted payloads. - The public [readiness endpoint](https://app.openadapt.ai/api/health/ready) reports the configured retention component separately from the @@ -86,8 +86,8 @@ flowchart TB ## Related pages -- [Security packet](security-packet.md) — posture summary for reviewers. -- [Vulnerability disclosure](vulnerability-disclosure.md) — how to report. -- [PHI handling](phi-handling.md) — the end-to-end PHI narrative. -- [Security and data handling](../guides/security-and-data-handling.md) — the +- [Security packet](security-packet.md): posture summary for reviewers. +- [Vulnerability disclosure](vulnerability-disclosure.md): how to report. +- [PHI handling](phi-handling.md): the end-to-end PHI narrative. +- [Security and data handling](../guides/security-and-data-handling.md): the full technical dossier, including hosted retention detail. diff --git a/docs/concepts/backends.md b/docs/concepts/backends.md index c3acf92f..f11d95ec 100644 --- a/docs/concepts/backends.md +++ b/docs/concepts/backends.md @@ -46,7 +46,7 @@ substrate; nothing about the safety model is specific to it. The public `WindowsBackend` now narrows the in-session boundary to typed `/input` and `/uia/*` operations, disables arbitrary legacy execution by default, screenshots the desktop, and reads the **UI Automation** tree for -identity. Crucially, an +identity. An element usually exposes `Name` / `Value` text **even when it has no stable `AutomationId`**, so UIA-based identity is viable on most native apps even where a durable selector is not. @@ -135,7 +135,7 @@ swappable `RDPTransport` protocol (so the adapter is CI-testable without a live server) and a real transport over the pure-Python async `aardwolf` client, behind the optional `rdp` extra. On a pure-pixel substrate the ladder runs on its visual floor and the identity gate falls back to its pixel/OCR tiers, which -is why a look-alike identifier can force a [halt rather than a verify](identity-gate.md) +is why a look-alike identifier can force a [halt instead of a verify](identity-gate.md) there. For every consequential remote action, the runtime uses a two-phase actuation diff --git a/docs/concepts/capability-ladder.md b/docs/concepts/capability-ladder.md index abd336f5..f88bab27 100644 --- a/docs/concepts/capability-ladder.md +++ b/docs/concepts/capability-ladder.md @@ -39,12 +39,12 @@ The rung you land on changes what the system can guarantee. Identity is the clearest example. Two different records with the same name and date of birth, distinguished only by an identifier differing by a single `O` versus `0` glyph, render to a byte-identical OCR band. On the **visual** rung, OCR -cannot separate them, so OpenAdapt refuses rather than guesses. On the -**structural** rung, the two rows are different strings in the tree, so the same -case verifies with no availability cost. +can't separate them, so OpenAdapt stops. On the **structural** rung, the two +rows are different strings in the tree, so the same case verifies with no +availability cost. -The principle: push each decision to the highest rung the app supports, and fail -safe below it. See [The identity gate](identity-gate.md) for how this plays out. +Push each decision to the highest rung the app supports, and stop below it. See +[The identity gate](identity-gate.md) for how that plays out. ## Capability-adaptive compilation diff --git a/docs/concepts/demonstration-compiler.md b/docs/concepts/demonstration-compiler.md index f153599e..993b8990 100644 --- a/docs/concepts/demonstration-compiler.md +++ b/docs/concepts/demonstration-compiler.md @@ -76,7 +76,7 @@ OpenAdapt uses that higher-fidelity signal via UIA, native macOS, native Linux AT-SPI, RDP, and Citrix/VDI [backends](backends.md) are all adapters to the same protocol, not rewrites. -## An API compiler for the API-less long tail +## When the app has no usable API Most enterprise software has no usable API for the workflow you actually run. The demonstration is the only interface that always exists: if a person can do diff --git a/docs/concepts/durable-runtime.md b/docs/concepts/durable-runtime.md index 235bee6d..26910050 100644 --- a/docs/concepts/durable-runtime.md +++ b/docs/concepts/durable-runtime.md @@ -1,10 +1,10 @@ # Durable runtime: checkpoint, attended decision, resume -A halt is the safety design working: the run stopped rather than guessing. But a -halt mid-workflow should not mean starting over, and must never re-perform a -write that already landed. The durable runtime turns a halt into a **durable -pause**. An authorized operator can make a bounded attended decision, and the -runtime resumes only from the last verified checkpoint. +A halt means the run stopped instead of guessing. A halt mid-workflow should +not mean starting over, and must never re-perform a write that already landed. +The durable runtime turns a halt into a **durable pause**. An authorized +operator can make a bounded attended decision, and the runtime resumes only +from the last verified checkpoint. !!! note "Off by default" The durable runtime is Tier-3 and opt-in. Enable it with `runtime.durable` @@ -32,9 +32,9 @@ from the system of record. **A halt is not a rollback.** What the halting step may already have done is stated in the run's terminal `transaction_outcome`, and is never inferred from the checkpoint: -- **`HALTED_BEFORE_EFFECT`** — absence was positively established for every +- **`HALTED_BEFORE_EFFECT`**: absence was positively established for every consequential step. There is nothing to reconcile. -- **`RECONCILIATION_REQUIRED`** — delivery or persistence is uncertain, +- **`RECONCILIATION_REQUIRED`**: delivery or persistence is uncertain, conflicting, or unverifiable. Reconcile the current state before resuming; the runtime will not blind-retry. @@ -91,24 +91,23 @@ notification. It projects a closed context only; screenshots, OCR, values, and free-text application data stay on the customer-controlled runner. See [Attended decisions and the halt-learn loop](halt-learn-loop.md). -## The bounded-recovery posture +## Bounded recovery -Durable resume is the third tier of a deliberately bounded runtime: +Durable resume is the third tier of a bounded runtime: 1. a **deterministic fast path** (the resolution ladder, $0); 2. a **bounded model recovery** of at most one local transition, when configured and permitted; 3. a **durable pause, approve, resume** from the last verified checkpoint. -It is explicitly **not** "hand the rest of the workflow to a free-form agent -after a halt." Recovery is scoped; the checkpoint is where a human takes over -when it cannot be. Same posture as the -[identity gate](identity-gate.md) and [effect verification](effect-verification.md): -when the right action is not determined, stop, and here, stop *resumably*. +Recovery is scoped. After a halt, OpenAdapt does not hand the rest of the +workflow to a free-form agent. The checkpoint is where a human takes over when +the next action is not determined. [Identity](identity-gate.md) and +[effect verification](effect-verification.md) stop the same way; here the stop +is resumable. -How that handover actually reaches a person — the bounded question, what an -answer does and does not authorize, and why the engine re-verifies rather than -trusting it — is the +How that handover reaches a person, what an answer authorizes, and why the +engine re-verifies live state, is the [attended decision path](halt-learn-loop.md#where-a-halt-goes-the-attended-decision). See the [Run a deployment](../guides/run-a-deployment.md) guide for a worked diff --git a/docs/concepts/effect-verification.md b/docs/concepts/effect-verification.md index 8a67d815..033a3f77 100644 --- a/docs/concepts/effect-verification.md +++ b/docs/concepts/effect-verification.md @@ -1,11 +1,10 @@ # Effect verification -The screen is not the system of record: a "Saved" banner can paint over an empty -database. Effect verification confirms a write actually landed in the real -record, exactly once, with the right values, by reading the record instead of -the pixels. +A "Saved" banner can paint over an empty database. Effect verification confirms +a write actually landed in the real record, exactly once, with the right values, +by reading the record instead of the pixels. -## The problem: five silent write faults +## Five silent write faults A vision postcondition asks a weak question: "do the pixels look like a save happened?" A fault-model study drove 90 replays through a real persistence @@ -24,7 +23,7 @@ help) and the screen shows success (so the screen oracle cannot help): None is render drift: the screen genuinely showed success, so only the record knows the truth. -## The mechanism: read the record, not the screen +## Read the record, not the screen A step can declare typed **effects** against the system of record. Given an `EffectVerifier`, the replayer snapshots the record before the action and, after @@ -38,10 +37,10 @@ verdict: non-2xx, expired token, unparseable body). **Halt.** An expired token is never mistaken for "record absent." -No "probably fine": both non-confirmed verdicts halt the run, mirroring the -[identity gate](identity-gate.md)'s refuse-rather-than-guess posture. The -verifier reads an API or a database, never the pixels, and makes **zero model -calls**, so the $0 runtime guarantee holds. +Both non-confirmed verdicts halt the run, the same way the +[identity gate](identity-gate.md) stops an unresolvable target. The verifier +reads an API or a database, never the pixels, and makes **zero model calls**, so +the $0 runtime guarantee holds. ```mermaid flowchart TD @@ -93,7 +92,7 @@ approach is not tied to one backend. For the five classes the screen silently mishandles, effect verification REFUTES and halts. -## The honest preconditions +## What has to be declared Two conditions are required, and both are real: @@ -103,7 +102,6 @@ Two conditions are required, and both are real: 2. A verifier must be **configured** for the deployment. A bundle with no declared effects, or a run with no verifier, still has only the -screen oracle for the write and is as silent as before. The one automatic -fail-safe: a step that declares effects while no verifier is configured is a -configuration error and halts, so an unverifiable consequential write is never -silently accepted. +screen oracle for the write. The one automatic fail-safe: a step that declares +effects while no verifier is configured is a configuration error and halts, so +an unverifiable consequential write isn't silently accepted. diff --git a/docs/concepts/halt-learn-loop.md b/docs/concepts/halt-learn-loop.md index 8aee9819..d8dda2ac 100644 --- a/docs/concepts/halt-learn-loop.md +++ b/docs/concepts/halt-learn-loop.md @@ -1,13 +1,13 @@ # Attended decisions and the halt-learn loop -A [halt](identity-gate.md) is honest, but a halt nobody hears is just a stopped -run, and a halt that teaches the system nothing means the same unhandled state -halts forever. This page covers both halves of the answer: **where a halt goes** -— the bounded question OpenAdapt puts in front of a person — and **what happens -when that person teaches the fix**, which is the halt-learn loop proper. +A [halt](identity-gate.md) nobody hears is just a stopped run, and a halt that +teaches the system nothing means the same unhandled state halts forever. This +page covers both halves of the answer: **where a halt goes** (the bounded +question OpenAdapt puts in front of a person) and **what happens when that +person teaches the fix**, which is the halt-learn loop proper. -Both halves refuse rather than guess. Neither hands control to a free-form -agent, and neither puts a generative-model API call on the runtime path. +Neither half hands control to a free-form agent, and neither puts a +generative-model API call on the runtime path. ## Terms used here @@ -59,7 +59,7 @@ start, what kind of target it was looking for, what the [resolution ladder](capability-ladder.md) tried on each rung, and what the engine will re-prove if they continue. -The decision client is a **responsive web app** — deliberately not a native iOS +The decision client is a **responsive web app**, not a native iOS or Android application, so there is no app store or separate update channel. The customer-controlled runner serves the full local portal. The hosted queue receives only a closed, PHI-free decision context. @@ -99,14 +99,14 @@ When an answer arrives, the engine: 2. takes a single-flight lease, so two people cannot both decide one task; 3. **re-reads the live application** and re-runs its identity, postcondition, and effect checks against a fresh observation; -4. continues only if that fresh check passes — and **refuses** when the +4. continues only if that fresh check passes, and **refuses** when the application is not actually in the state the step needs. So "I fixed it" means *I prepared the live state; now go and check it*. It never means "repeat the paused write". A run reaches `VERIFIED` only when the complete configured contract proves the intended effect, exactly as on an unattended run. -An operator's answer is an input to a verification, never a substitute for one — -which is why an operator who is mistaken produces a second halt rather than a +An operator's answer is an input to a verification, never a substitute for one, +which is why an operator who is mistaken produces a second halt instead of a silent wrong write. !!! note "The operator's own work is verified, not replayed" @@ -116,7 +116,7 @@ silent wrong write. A step whose postconditions or independent effect check do not pass is refused, not banked. -The reply is also honest about delivery. Three states stay distinct: not sent, +The reply also keeps delivery states distinct. Three states stay distinct: not sent, sent, and **may have been sent**. For uncertain delivery, use **Reconcile**. The runner reads the required postcondition and independent effect again. It does not send the action again. It reports a reconciled result only when that @@ -128,10 +128,10 @@ The retained screen, the observed values, the OCR, and the failing target never leave the customer-controlled runner. Only a typed, PHI-free envelope crosses to a hosted control plane: opaque identifiers, digests, closed enums, bounded counts, and expiry. There is no free-text field anywhere in that envelope, so -raw values and prose are **structurally unable** to travel rather than being +raw values and prose are **structurally unable** to travel; they are not stripped in transit. -The direct consequence is that a hosted surface shows **less** than the runner — +The direct consequence is that a hosted surface shows **less** than the runner: it can say the *shape* of a failure but not its content. That is the design working, not a gap: @@ -144,7 +144,7 @@ working, not a gap: A hosted surface that says less is a surface that cannot leak more. An operator who needs the full picture opens the run on the runner, where the evidence -already is — and the hosted surface says so, rather than letting absent detail +already is, and the hosted surface says so, so absent detail does not read as "OpenAdapt does not know". Sending a decision *back* through a hosted control plane requires an explicit @@ -158,7 +158,7 @@ runner keeps decisions on its local surface. There are two ways, and they trade fidelity against what your network has to do. Neither of them asks OpenAdapt to open a hole in it. -#### The hosted lane — nothing to configure +#### The hosted lane: nothing to configure **This is the default answer for a practice without an IT department.** The runner makes **outbound HTTPS requests only** to the control plane: no inbound @@ -184,16 +184,16 @@ certificate. What the phone shows on this lane is the *closed halt context*: which category of check failed, which resolution rungs were tried and what each one returned, which contracts a "Continue" will re-prove, and bounded counts. Every value is a -closed enum, a bounded integer, or a boolean — **there is no string field and no +closed enum, a bounded integer, or a boolean. **There is no string field and no image**, so the hosted service is structurally unable to hold a name, an MRN, an observed value, or a workflow label. It is not scrubbed; it has nowhere to put them. The one thing it gives up is the target control's own accessible name. The phone -says *"OpenAdapt could not find the button"* rather than *"the button labelled +says *"OpenAdapt could not find the button"* instead of *"the button labelled `Open`"*, and it tells you a name exists that it is not showing you. -#### The runner-local portal — full fidelity, on your own terms +#### The runner-local portal: full fidelity, on your own terms The portal on the runner serves everything, including the protected screenshot crops. That is why it is the path with a network requirement. @@ -202,8 +202,8 @@ crops. That is why it is the path with a network requirement. Out of the box the decision portal binds `127.0.0.1` and advertises a loopback URL. **A phone cannot reach it, and a fresh install will not serve one.** Publishing it to a phone requires *you* to terminate trusted TLS in - front of the runner — an enterprise reverse proxy, a VPN, or a ZTNA - hostname — and to record that decision in configuration. + front of the runner (an enterprise reverse proxy, a VPN, or a ZTNA + hostname) and to record that decision in configuration. Use the hosted lane above if you do not operate one. See [the portal settings](../reference/configuration.md#the-self-hosted-phone-portal) @@ -214,7 +214,7 @@ crops. That is why it is the path with a network requirement. We did not punch a hole in your network for our convenience, and there is no "bind everything" switch to make a demo easier. The boundary in front of a runner that can see protected records is yours to open, deliberately, under your -own certificate and access policy — so this path inherits the authentication, +own certificate and access policy, so this path inherits the authentication, device posture, and logging you already run, instead of asking you to trust a second one. @@ -236,8 +236,8 @@ The phone then shows a short, one-use pairing code. Type that code on the runner to approve that phone. This binds the phone to the local portal; it does not give the phone an engine or console capability. -Do it the intuitive way — derive the code from the pairing and show it on the -runner — and an attacker who photographed the QR from across the room and +Do it the intuitive way (derive the code from the pairing and show it on the +runner) and an attacker who photographed the QR from across the room and claimed it first would be shown the very code the runner's screen was already displaying. The "matching code" would then confirm the attacker. Minting per claim means a remote attacker's phone shows a code the operator cannot see, and @@ -245,7 +245,7 @@ the mismatch is visible immediately. The rest of the shape follows from the same posture: -- The QR link carries **only a pairing secret** — no console capability, no +- The QR link carries **only a pairing secret**: no console capability, no pause capability, no tenant, run, or pause identifier. - The secret rides in the URL **fragment**, which browsers never transmit, so it cannot land in a reverse-proxy access log or a referrer header. @@ -406,23 +406,19 @@ the re-run command. class; generalizing to arbitrary corrections is the job of a richer inducer behind that same seam, which inherits the same gate. -## Why this shape +## What a person supplies, what the engine keeps -Every other safety mechanism in OpenAdapt refuses rather than guesses. The -halt-learn loop is how the system *improves* without abandoning that posture: the -only thing trusted to generalize a fix is a demonstration plus a gate, biased -the way the runtime is, so a revision that might weaken safety is quarantined, -not shipped. It is the counterpart to +The only thing trusted to generalize a fix is a demonstration plus a gate. A +revision that might weaken safety is quarantined, not shipped. That pairs with [multi-trace induction](multi-trace-induction.md) (recover the program from several traces) and [policy and certify](policy-and-certify.md) (refuse a bundle -whose gaps were not closed): learn only what you can prove safe. - -The attended decision is the same posture pointed at people. It would be easier -to treat a tap as an answer, publish the portal on every interface so a phone -just works, and forward the failing screen to a dashboard where support can see -it. Each of those would move a decision, a network boundary, or a protected -record somewhere it does not belong. Instead the person supplies the one thing a -machine cannot — an observation about the world — and the engine keeps -everything it is actually good at: re-reading live state, checking identity and +whose gaps were not closed). + +It would be easier to treat a tap as an answer, publish the portal on every +interface so a phone just works, and forward the failing screen to a dashboard +where support can see it. Each of those would move a decision, a network +boundary, or a protected record somewhere it does not belong. The person +supplies the one thing a machine cannot (an observation about the world) and +the engine keeps the rest: re-reading live state, checking identity and effects, and refusing. A halt reaches a phone in seconds, and it still cannot turn into a wrong write because somebody was in a hurry. diff --git a/docs/concepts/identity-gate.md b/docs/concepts/identity-gate.md index ae14b0d2..e5d79954 100644 --- a/docs/concepts/identity-gate.md +++ b/docs/concepts/identity-gate.md @@ -1,12 +1,12 @@ # The identity gate -The most dangerous failure in desktop automation is not a crash. It is clicking -the **right-looking wrong thing**: opening the wrong patient, editing the wrong +The most dangerous failure in desktop automation is clicking the +**right-looking wrong thing**: opening the wrong patient, editing the wrong account, in a repeated structure where the target still looks plausible. The identity gate is a pre-action check that refuses to act when it cannot tell two records apart. -## The threat: right position, wrong entity +## Right position, wrong entity When data shifts between runs (a row added above the target, the target's row deleted, a look-alike sibling, a re-sorted table), the resolver can still find a @@ -14,21 +14,21 @@ pixel-identical target at a plausible position. Resolving is not enough. Before an armed click, OpenAdapt re-reads the resolved row's text and compares it to the recorded row; on a mismatch, it halts before clicking. -## An impossibility result, and an honest response +## OCR cannot tell those records apart Identity has a proven ceiling on pixels alone. Two **different** records with the same name and same date of birth, whose only distinguishing field is an identifier differing by a single glyph (`MG4408` vs `MG44O8`, `100512` vs `1OO512`), render to a **byte-identical OCR band**. That band is identical to what a re-read of the true row produces, so nothing downstream of OCR can -separate them. This is not a tuning gap but the limit of OCR-based identity, and -the honest response is to **refuse rather than guess**. +separate them. That's the limit of OCR-based identity, not a threshold you can +tune away. OpenAdapt stops. ## The identity ladder Identity is an ordered ladder of verifier tiers, highest-fidelity first: the -first that can judge the substrate wins, and its verdict is final. Every tier is -fail-safe: unsure abstains to the next, and if nothing verifies, the run halts. +first that can judge the substrate wins, and its verdict is final. Unsure +abstains to the next tier. If nothing verifies, the run halts. ```mermaid flowchart TD @@ -61,7 +61,7 @@ flowchart TD tier may only MISMATCH (a safe halt) or ABSTAIN, never grant a pass. 3. **Local-VLM veto** (optional, off by default). A local open model can **reject** a wrong record with high reliability but is not trusted to - **certify** a right one, so a "same" answer abstains rather than passes. It + **certify** a right one, so a "same" answer abstains instead of passing. It pulls no model on the default install and makes zero cloud calls. 4. **OCR name + DOB band.** The fallback matcher. It verifies same-identity only when there is provably no collapsible glyph in any identifier-position token, @@ -75,8 +75,7 @@ Driven through the real replayer, the integrated ladder measures **zero false-accept across every substrate configuration**, including the same-name/same-DOB homonym. On browser and desktop the structured tier closes the class at no availability cost. On pure pixels a collapsible identifier is -not safely verifiable and **halts** today, the honest cost of "OCR alone cannot -verify a collapsible identifier." +not safely verifiable and **halts** today, because OCR alone cannot verify it. !!! warning "Coverage is a first-class, auditable metric" Identity verification covers only **armed** steps, and real bundles arm a @@ -90,15 +89,14 @@ verify a collapsible identifier." wrong-entity click on an unarmed step is still silent. This is what [policy and certify](policy-and-certify.md) exists to gate. -## Why this posture +## Availability cost -The identity gate costs availability: on noisy pure-pixel rows it sometimes -halts a correct run rather than gamble. That is the cheap direction to be wrong. -Clicking by position is what caused wrong-record writes, so OpenAdapt takes the -halt. Deployments that cannot tolerate it can escalate each halt to a fallback -rather than proceed blindly. +On noisy pure-pixel rows the identity gate sometimes halts a correct run. That's +cheaper than a wrong-record write. Clicking by position is what caused those +writes, so OpenAdapt takes the halt. Deployments that can't tolerate it can +escalate each halt to a fallback. -A halt is not a dead end: it becomes a bounded question in front of a person, -answerable from a phone on your own network. What that person's answer does — -and, importantly, what it does *not* authorize — is the +A halt becomes a bounded question in front of a person, answerable from a phone +on your own network. What that person's answer does, and what it does not +authorize, is the [attended decision path](halt-learn-loop.md#where-a-halt-goes-the-attended-decision). diff --git a/docs/concepts/index.md b/docs/concepts/index.md index a4cfc287..3c86c801 100644 --- a/docs/concepts/index.md +++ b/docs/concepts/index.md @@ -21,7 +21,7 @@ jump to what you need. Hosted browser launch, customer-controlled execution, self-hosting, and the authoring-versus-runtime data boundary. -- [__Fail-closed regulated execution__](regulated-execution.md) +- [__Regulated execution__](regulated-execution.md) `replay` is the local $0 dev path; `run` refuses to execute unless certified, identity-covered, effect-verified, signed, and config-pinned. @@ -67,8 +67,7 @@ jump to what you need. - [__Policy and certify__](policy-and-certify.md) - Fail-closed safety: `lint` reports gaps, `certify` refuses an unsafe - bundle before it deploys. + `lint` reports gaps, `certify` refuses an unsafe bundle before it deploys. - [__Backends: where it runs__](backends.md) diff --git a/docs/concepts/multi-trace-induction.md b/docs/concepts/multi-trace-induction.md index 1c79dc21..f58e87f6 100644 --- a/docs/concepts/multi-trace-induction.md +++ b/docs/concepts/multi-trace-induction.md @@ -31,7 +31,7 @@ is resolving that ambiguity: ## The induction loop Induction turns several demonstrations into a program by proposing, questioning, -and validating, not guessing: +and validating: ```mermaid flowchart TD @@ -49,12 +49,12 @@ flowchart TD 1. **Bootstrap** one interpretation from one demonstration. 2. **Enumerate** candidate generalizations (is this value a parameter, is this step a loop body). -3. **Resolve ambiguity by asking**, not guessing: surface concrete - multiple-choice questions to the operator. +3. **Resolve ambiguity by asking**: surface concrete multiple-choice questions + to the operator. 4. **Fold in additional traces** and infer the shared control-flow graph. 5. **Validate** on held-out traces and synthetic perturbations. -6. **Quarantine** when intent stays underdetermined: refuse to emit rather than - ship a workflow that might do the wrong thing. +6. **Quarantine** when intent stays underdetermined: refuse to emit a workflow + that might do the wrong thing. ## Induce a program @@ -72,7 +72,7 @@ scores. A certified program can then loop over a data source with The worked walkthrough is in [Induce a program from multiple traces](../guides/induce-a-program.md). -## Ask, don't guess +## Disambiguate before certify The disambiguation step ships as a CLI verb. It surfaces the compile-time questions an ambiguous demonstration raises and applies the answers as guards or @@ -83,15 +83,11 @@ certified. openadapt flow disambiguate bundle --interactive --write ``` -A consequential (must-answer) ambiguity exits nonzero until it is resolved. Same -posture as the rest of the system: when the right action is not determined, stop -and ask, rather than proceed and hope. - -## The through-line +A consequential (must-answer) ambiguity exits nonzero until it's resolved. When +the right action isn't determined, stop and ask. OpenAdapt spent enormous effort making a *single trace* safe: the [identity ladder](identity-gate.md), volatility mining, postconditions, -[effect verification](effect-verification.md). That work is necessary but -insufficient, because the trace itself under-specifies intent. Induction -recovers the intended program, and quarantine keeps the system honest when it -cannot. +[effect verification](effect-verification.md). That work is necessary, and it +still leaves the trace under-specified. Induction recovers the intended +program. Quarantine withholds a program when it cannot. diff --git a/docs/concepts/policy-and-certify.md b/docs/concepts/policy-and-certify.md index 3e6cfc94..d6a593fe 100644 --- a/docs/concepts/policy-and-certify.md +++ b/docs/concepts/policy-and-certify.md @@ -1,10 +1,11 @@ # Policy and certify -Compiled is not certified safe. A runnable bundle can still have gaps: a write -with no identity check, a step that asserts nothing. Policy and certify separate -"runnable" from "certified safe", and fail closed. +Compiled is not the same as certified safe. A runnable bundle can still have +gaps: a write with no identity check, a step that asserts nothing. `lint` +reports those gaps. `certify` enforces a policy and exits nonzero, refusing +the bundle before it deploys, when it fails. -## Two commands, two jobs +## lint and certify ```mermaid flowchart LR @@ -22,8 +23,7 @@ flowchart LR unarmed or vacuous *irreversible* step); `--strict` also fails on warnings. `lint` is advice. - **`certify`** enforces a policy and **refuses** the bundle (exits nonzero) - when it fails. It is the gate: put it in CI or a deploy step and an unsafe - bundle never ships. + when it fails. Put it in CI or a deploy step so an unsafe bundle never ships. ```bash openadapt flow lint bundle @@ -42,7 +42,7 @@ a strict `clinical-write`. The strict policy asserts, for example: `certify` evaluates the bundle against the policy and reports each violated requirement before deploy. -## Risk is auto-classified, then enforceable +## Risk classification At compile time, write-shaped clicks (create, update, delete, submit, save, confirm, add, and siblings, matched on word boundaries) are auto-classified @@ -50,27 +50,21 @@ confirm, add, and siblings, matched on word boundaries) are auto-classified consequential writes, not only when a human marks the step. A `risk_overrides` map wins either direction. -!!! warning "The classifier is a heuristic, not understanding" +!!! warning "The classifier reads labels, not the app's effect" Risk classification reads the label and the intent, never the app's true - effect. It is deliberately biased toward irreversible (a false irreversible - costs availability; a false reversible costs safety), but it misses writes - with non-write labels (an icon-only "commit", a bare "OK" that saves) and - writes committed by a submitting ++enter++ key. It also over-flags benign - write words ("Apply filter", "Add to favourites"). A write behind a - non-write label stays reachable with a green report unless a human adds - `risk_overrides`. That is why `certify` with a strict policy is the gate that - refuses a bundle whose gaps stay open. + effect. It is biased toward irreversible (a false irreversible costs + availability; a false reversible costs safety), but it misses writes with + non-write labels (an icon-only "commit", a bare "OK" that saves) and writes + committed by a submitting ++enter++ key. It also over-flags benign write + words ("Apply filter", "Add to favourites"). A write behind a non-write + label stays reachable with a green report unless a human adds + `risk_overrides`. That is why `certify` with a strict policy refuses a + bundle whose gaps stay open. -## Fail-closed, everywhere - -Policy and certify are the compile-time and pre-deploy face of the runtime's -posture: [halt rather than guess](identity-gate.md), -[refuse rather than accept an unverifiable write](effect-verification.md), -[quarantine rather than emit an ambiguous program](multi-trace-induction.md). -The gate turns disclosure into enforcement: the residual gaps OpenAdapt is -honest about become a policy that refuses the bundle until they are closed. - -This is a compile-time and pre-deploy layer only; the replayer, identity ladder, -and healer are unchanged. See -[Write and enforce a policy](../guides/policy-and-certification.md) for a worked -example. +Policy and certify run at compile time and before deploy. They don't change +the replayer, identity ladder, or healer. See +[Write and enforce a policy](../guides/policy-and-certification.md) for a +worked example. At runtime, [`run`](regulated-execution.md) applies the same +kind of gate: [identity](identity-gate.md) stops an unresolvable target, +[effect verification](effect-verification.md) stops an unverifiable write, and +[induction](multi-trace-induction.md) withholds an ambiguous program. diff --git a/docs/concepts/regulated-execution.md b/docs/concepts/regulated-execution.md index 6d05f4e6..3d4c6b7a 100644 --- a/docs/concepts/regulated-execution.md +++ b/docs/concepts/regulated-execution.md @@ -1,19 +1,17 @@ -# Fail-closed regulated execution: `run` vs `replay` +# Regulated execution: `run` vs `replay` -Iterating on a bundle at your desk and executing a consequential write -in production are different acts, and OpenAdapt draws a hard line: **`replay`** -is the local, $0, developer-and-pilot path; **`run`** is the fail-closed -regulated path that refuses to execute unless its admission-gate requirements -are satisfied. +Iterating on a bundle at your desk and executing a consequential write in +production are different acts. **`replay`** is the local, $0, +developer-and-pilot path. **`run`** is the regulated path that refuses to +execute unless its admission-gate requirements are satisfied. -## Two verbs, two postures +## What `replay` and `run` do | | `replay` | `run` | |---|---|---| | **Purpose** | Develop, drift-test, pilot | Execute in a regulated / production deployment | -| **Posture** | Permissive | **Fail-closed** | +| **Missing safety preconditions** | Continues; it just replays | **Refuses to start** | | **Model calls on healthy path** | 0 | 0 | -| **Refuses on missing safety preconditions** | No, it just replays | **Yes, refuses to start** | `replay` is what every guide uses: it runs the bundle locally, deterministically, for free, and is where drift-testing and pilots happen. `run` adds a @@ -22,16 +20,16 @@ effect verification). *How* a step executes is unchanged; `run` just **refuses to begin** if required coverage, encryption, or integrity evidence is missing. !!! note "Governed admission is workflow-specific" - The fail-closed `run` verb and its admission-gate tests ship in the canonical - engine. That makes configured controls mandatory by default; it does not - make every backend or workflow production-ready. Use `run --dry-run` to - inspect the gate report before execution, and review + The `run` verb and its admission-gate tests ship in the canonical engine. + That makes configured controls mandatory by default; it does not make every + backend or workflow production-ready. Use `run --dry-run` to inspect the + gate report before execution, and review [Qualification evidence](../get-started/what-works-today.md). ## What `run` checks before it executes a step `run` refuses to start unless **all** of the following hold. Each is an existing -safety mechanism; `run` makes them jointly mandatory rather than individually +safety mechanism; `run` makes them jointly mandatory, not individually optional. ```mermaid @@ -44,7 +42,7 @@ flowchart TD E -->|no| REFUSE E -->|yes| B{Encrypted bundle +
manifest integrity?} B -->|no| REFUSE - B -->|yes| GO([Execute · fail-closed at every step]) + B -->|yes| GO([Execute · each step still gated]) ``` 1. **Certification.** The bundle must pass [`certify`](policy-and-certify.md) @@ -54,9 +52,8 @@ flowchart TD runs. 2. **Identity coverage.** The policy sets a floor on [identity](identity-gate.md)-armed coverage, and `run` refuses a bundle below - it. This confronts the honest limit that **identity verification covers only - armed steps**: `run` will not execute a bundle whose consequential clicks are - unarmed when the policy forbids it. + it. Identity verification covers only armed steps: `run` will not execute a + bundle whose consequential clicks are unarmed when the policy forbids it. 3. **Verified effects or explicit approval.** Every consequential write must declare a system-of-record effect, and the deployment must configure a matching verifier, unless an operator deliberately supplies the @@ -71,35 +68,35 @@ are warnings unless `--strict-templates` is set. Those escape hatches exist for development and migration; using them weakens the regulated posture and shows in the gate report. -**If any precondition fails, `run` refuses; it does not degrade to best-effort -execution.** Refusing is the cheap direction to be wrong. +**If any precondition fails, `run` refuses; it doesn't degrade to best-effort +execution.** -## Fail-closed at every step, not just at the door +## During the run -The pre-flight gate is the entry check; the same posture governs every step: +The pre-flight gate is the entry check. Each step still applies the same +controls: -- An unresolvable target [halts](identity-gate.md) rather than clicking by +- An unresolvable target [halts](identity-gate.md) instead of clicking by position. -- A [REFUTED or INDETERMINATE effect](effect-verification.md) halts rather than +- A [REFUTED or INDETERMINATE effect](effect-verification.md) halts instead of proceeding on a "Saved" banner. - An ambiguous identity abstains up the ladder and halts if nothing verifies. -- A missing scrubbing capability, under `OPENADAPT_FLOW_SCRUB=on`, aborts rather - than writing PHI/PII at all. +- A missing scrubbing capability, under `OPENADAPT_FLOW_SCRUB=on`, aborts + instead of writing PHI/PII at all. A halt is not a dead end. It feeds the [halt-learn loop](halt-learn-loop.md), where an operator demonstrates the fix, a regression gate proves it weakens nothing, and only a verified revision is promoted. -## Honest limits `run` does not repeal +## What `run` still cannot know -`run` makes safety mechanisms mandatory; it does not make them omniscient. The +`run` makes safety mechanisms mandatory; it doesn't make them omniscient. The [LIMITS](https://github.com/OpenAdaptAI/openadapt-flow/blob/main/docs/LIMITS.md) -still apply, and `run` is honest about them: +still apply: -- **Identity coverage is a floor, not omniscience.** `run` can require N of M - clicks armed, but an armed step's guarantee is only as strong as the substrate - allows: on pure-pixel Citrix a collapsible identifier halts rather than - verifies. +- **Identity coverage is a floor.** `run` can require N of M clicks armed, but + an armed step's guarantee is only as strong as the substrate allows: on + pure-pixel Citrix a collapsible identifier halts instead of verifying. - **On-screen read-back is not independent verification.** On a no-API desktop where the only oracle is the screen, effect "verification" reads the same surface the action wrote to: **same-surface** confirmation, not an independent @@ -114,6 +111,4 @@ still apply, and `run` is honest about them: enforces the policy you wrote; a gap the policy does not name is not caught by `run`. -The value of `run` is not removing these limits but turning them into **enforced -preconditions**: a deployment that has not closed them cannot start a regulated -run. +A deployment that has not closed these gaps cannot start a regulated run. diff --git a/docs/concepts/self-healing.md b/docs/concepts/self-healing.md index 26e5a91d..09868827 100644 --- a/docs/concepts/self-healing.md +++ b/docs/concepts/self-healing.md @@ -59,8 +59,8 @@ It does **not** cover changes that invalidate the evidence itself: content) halts. Per-tenant re-recording is the working assumption. When the screen stops matching expectations entirely, the run **halts with a -report** naming the violated postcondition rather than guessing. Irreversible -steps never act on a low-confidence match. +report** naming the violated postcondition. Irreversible steps never act on a +low-confidence match. ## Repairs are diffs you can review @@ -91,9 +91,8 @@ it is not a general adaptation guarantee. ## Healing versus verification -Self-healing recovers a target whose **appearance** drifted. It deliberately -does **not** rescue a step whose **effect** is wrong: a duplicate write or -partial save shows unchanged pixels, so there is nothing to heal. That is the -job of [effect verification](effect-verification.md); clicking the wrong record -is the job of [the identity gate](identity-gate.md). The three are separate on -purpose. +Self-healing recovers a target whose **appearance** drifted. It does **not** +rescue a step whose **effect** is wrong: a duplicate write or partial save +shows unchanged pixels, so there is nothing to heal. That is the job of +[effect verification](effect-verification.md). Clicking the wrong record is +the job of [the identity gate](identity-gate.md). They stay separate. diff --git a/docs/concepts/settings-governance.md b/docs/concepts/settings-governance.md index 48e40922..59064079 100644 --- a/docs/concepts/settings-governance.md +++ b/docs/concepts/settings-governance.md @@ -31,15 +31,13 @@ safety-relevant change is always attributable. Preferences do not need this weight; safety-critical controls do, and the audit trail is where the difference shows. -## Fail-closed - -Governed configuration is fail-closed. When a required safety-critical setting is -missing, unreadable, or ambiguous, the workspace refuses rather than falling back -to a permissive default. The same posture runs through the product: a run that -cannot establish its declared checks -[halts rather than guessing](regulated-execution.md), and configuration that -cannot be resolved safely blocks rather than quietly loosening a control. - -This is why the tiers matter. A preference stays easy to change and a safety -control hard, with every governed change audited and every missing one refused. -That keeps a convenience edit from becoming a safety incident. +## Missing settings block the run + +When a required safety-critical setting is missing, unreadable, or ambiguous, +the workspace refuses instead of falling back to a permissive default. The same +rule applies at runtime: a run that can't establish its declared checks +[halts](regulated-execution.md), and configuration that cannot be resolved +safely blocks instead of quietly loosening a control. + +A preference stays easy to change. A safety control stays hard, with every +governed change audited and every missing one refused. diff --git a/docs/concepts/substrate-model.md b/docs/concepts/substrate-model.md index 0a04a1f2..adbb9673 100644 --- a/docs/concepts/substrate-model.md +++ b/docs/concepts/substrate-model.md @@ -12,7 +12,7 @@ the resolution ladder, the [identity gate](identity-gate.md), [effect verification](effect-verification.md), and the [halt-learn loop](halt-learn-loop.md). -## Two axes, one contract +## Substrate and deployment Two orthogonal questions about any run, kept separate: @@ -59,10 +59,10 @@ The released backends cover it behind the same protocol: - **`WindowsBackend`** drives a native Windows desktop through an in-session agent. Its shipped typed RPC exposes bounded screenshot, input, and UIA operations while the legacy arbitrary-execution route stays disabled by - default. It reads the **UI Automation** tree for identity. Crucially, most + default. It reads the **UI Automation** tree for identity. Most native controls expose `Name` / `Value` text **even without a stable - `AutomationId`**, so structured identity is viable on desktop, not just the - browser. + `AutomationId`**, so structured identity is viable on desktop as well as + the browser. - The **native macOS backend** binds one exact application window and uses Accessibility metadata plus retained visual evidence. - **`LinuxBackend`** binds one exact AT-SPI application and top-level window, @@ -88,7 +88,7 @@ coordinates. identity gate uses its pixel/OCR tiers. When an identifier is genuinely ambiguous at that fidelity (a same-name/same-DOB record whose MRN differs by a single `O`/`0` glyph), the gate - **[halts rather than guesses](identity-gate.md)** by design. That is the same + **[halts](identity-gate.md)** instead of guessing. That is the same never-click-the-wrong-record guarantee every substrate enforces. Each surface verifies with the highest-fidelity signal it exposes, and structured layers (a browser DOM, Windows UIA) resolve that class outright. Per-substrate diff --git a/docs/concepts/vlm-appliance.md b/docs/concepts/vlm-appliance.md index df5f319b..3cd0eea1 100644 --- a/docs/concepts/vlm-appliance.md +++ b/docs/concepts/vlm-appliance.md @@ -10,7 +10,7 @@ by default; when on, the model and the data stay in the building. Enable the appliance by pointing the runtime at a local URL (`OPENADAPT_FLOW_VLM_URL`). Unset, no model tiers exist and the ladder has no grounder rung. Configured, three fail-safe tiers come online, each biased toward -halting rather than mis-acting: +halting instead of mis-acting: ```mermaid flowchart TD @@ -38,7 +38,7 @@ Every appliance path is fail-safe to halt: an outage or unsure answer keeps the halt, never a mis-click. Every rescue is recorded in the run report and counted as a model call, so the appliance never silently breaks the $0 story. -## Measured, and honestly bounded +## What the measurements showed The appliance was measured end-to-end against a real served local model. The state verifier correctly refused 7 of 8 should-halt screens and false-rescued 1 diff --git a/docs/concepts/workflow-ir.md b/docs/concepts/workflow-ir.md index 17aaa444..2942a6c2 100644 --- a/docs/concepts/workflow-ir.md +++ b/docs/concepts/workflow-ir.md @@ -68,14 +68,14 @@ model is trusted to do: 3. a **[durable pause, approve, resume](durable-runtime.md)** from the last verified checkpoint. -Explicitly **not** "hand the rest of the workflow to a free-form agent after a -halt." Recovery is scoped to a single transition; the checkpoint is where a -human takes over if it cannot be. When a human *does* demonstrate the fix at a -halt, the [halt-learn loop](halt-learn-loop.md) folds it back into the program -through the same governed induction and regression gate, so that state need not -halt again. - -The IR is how OpenAdapt gets from "replay this one demonstration safely" to "run -this program safely across the variation the real world throws at it." How a -program is *recovered* from more than one demonstration is +Recovery is scoped to a single transition. After a halt, OpenAdapt does not +hand the rest of the workflow to a free-form agent. The checkpoint is where a +human takes over if the next action is not determined. When a human +demonstrates the fix at a halt, the [halt-learn loop](halt-learn-loop.md) folds +it back into the program through the same governed induction and regression +gate, so that state need not halt again. + +The IR is how OpenAdapt gets from "replay this one demonstration safely" to +"run this program safely across the variation the real world throws at it." How +a program is *recovered* from more than one demonstration is [multi-trace induction](multi-trace-induction.md). diff --git a/docs/get-started/index.md b/docs/get-started/index.md index eefaa0e6..d9eae4ba 100644 --- a/docs/get-started/index.md +++ b/docs/get-started/index.md @@ -21,11 +21,11 @@ runs the program, and verifies the saved result. See it working before you install anything: -- **[Hosted demo](https://app.openadapt.ai/demo)** — recorded demonstrations, +- **[Hosted demo](https://app.openadapt.ai/demo)**: recorded demonstrations, verified replays, and fail-safe halts on real footage. -- **[Template gallery](https://openadapt.ai/templates)** — ready-to-adapt +- **[Template gallery](https://openadapt.ai/templates)**: ready-to-adapt workflow templates. -- **[Blog](https://blog.openadapt.ai)** — guides, updates, and automation +- **[Blog](https://blog.openadapt.ai)**: guides, updates, and automation recipes. ## First success: install, then run diff --git a/docs/get-started/what-works-today.md b/docs/get-started/what-works-today.md index 5fc17cf8..f94f1ee0 100644 --- a/docs/get-started/what-works-today.md +++ b/docs/get-started/what-works-today.md @@ -58,7 +58,7 @@ customer-controlled runtime connected to the same governance model. | Record -> compile -> lint -> certify -> replay -> report on a browser | **Supported** | Runs end to end in CI against the bundled app and end to end against a real third-party app. | A clean run is not automatically safe. Identity coverage, risk classification, postconditions, and effect contracts must be audited per bundle. | | Deterministic target re-resolution and saved heal diffs | **Supported** | Theme, movement, and rename drift are covered by the bundled drift matrix. | Scale/reflow and tenant-specific state can still halt. The base bundle is not silently promoted; save and review a healed bundle explicitly. | | `lint` and `certify` | **Supported** | Report coverage gaps and refuse bundles that violate the selected policy. | Certification is opt-in and only enforces what the policy names. An uncertified bundle remains runnable with `replay`. | -| Fail-closed `run` admission gate | **Supported** | The shipped gate checks certification, identity/effect coverage, approval fallback, encryption, and manifest integrity before executing. | Development escape hatches exist, and passing the gate does not validate the backend or prove a workflow safe. | +| `run` admission gate | **Supported** | The shipped gate checks certification, identity/effect coverage, approval fallback, encryption, and manifest integrity before executing. | Development escape hatches exist, and passing the gate does not validate the backend or prove a workflow safe. | | Identity verification | **Supported, opt-in by step coverage** | Structured identity works on armed browser steps; UIA and pixel/OCR tiers are implemented and adversarially tested. | Unarmed clicks have no identity check. Pure-pixel ambiguity intentionally over-halts, and coverage varies by workflow. | | System-of-record effect verification | **Supported** | REST, FHIR, and document-hash verifiers run in the live replay path and halt on refuted or indeterminate declared effects. | The compiler does not infer effects. A deployment must author effects and configure the matching verifier; otherwise screen checks remain the oracle. | | `teach` halt-to-correction loop | **Supported** | The deterministic reference inducer covers the optional-dialog correction class behind regression and canary gates. | Arbitrary UI corrections are not generally learned. Unsafe or underdetermined revisions are refused. | diff --git a/docs/guides/data-driven-loops.md b/docs/guides/data-driven-loops.md index b1b95d1e..9170b563 100644 --- a/docs/guides/data-driven-loops.md +++ b/docs/guides/data-driven-loops.md @@ -67,7 +67,7 @@ openadapt flow replay queue-bundle --url https://your.app \ ``` For a real deployment, use [`run --worklist`](../reference/cli.md#run) instead, -which adds the fail-closed admission gate, effect verification, and durable +which adds the admission gate, effect verification, and durable runtime from your deployment config. ## Every iteration keeps its gates diff --git a/docs/guides/deploy-on-prem.md b/docs/guides/deploy-on-prem.md index c39ed4cb..c7f8c82f 100644 --- a/docs/guides/deploy-on-prem.md +++ b/docs/guides/deploy-on-prem.md @@ -104,8 +104,8 @@ defence-in-depth layers: kernel drops all IP traffic for the runner unless a LAN CIDR is explicitly allow-listed); the Docker Compose alternative puts the runner on an `internal: true` network with no gateway to the internet. -3. **Fail-closed PHI/PII handling.** `OPENADAPT_FLOW_SCRUB=on` makes a missing - scrubbing capability *abort* rather than write plaintext PHI/PII. +3. **PHI/PII handling.** `OPENADAPT_FLOW_SCRUB=on` makes a missing + scrubbing capability *abort* instead of writing plaintext PHI/PII. 4. **Attestation.** `verify-airgap.sh` scans your config and environment for any off-LAN URL or cloud key; with `--probe` it actively curls a public canary and **asserts the call fails**; with `--audit` it walks the audit-log hash chain. @@ -284,7 +284,7 @@ Two boundary points worth carrying into a security review: | Decision envelope to a hosted control plane | No (structurally) | Opaque ids, digests, closed enums, bounded counts only; no free-text field exists to carry a value | | Deterministic replay path | n/a | No OpenAdapt-hosted dependency; target and verifier traffic remains | -## Compliance posture, stated honestly +## Compliance posture **Not legal advice, and not a compliance guarantee.** OpenAdapt provides the software substrate for running compiled automations on-premise. Whether a given diff --git a/docs/guides/hosted.md b/docs/guides/hosted.md index ac30525e..951887ad 100644 --- a/docs/guides/hosted.md +++ b/docs/guides/hosted.md @@ -59,7 +59,7 @@ support. ## Hosted recorder boundary -The hosted recorder is a real, bounded authoring path rather than a simulated +The hosted recorder is a real, bounded authoring path, not a simulated demo. A qualified hosted browser session produced PNG frames, accepted and retained input evidence, assembled a native recording, created one compileable workflow idempotently, enforced its resource limits, and removed the ephemeral diff --git a/docs/guides/induce-a-program.md b/docs/guides/induce-a-program.md index a48a444b..b1938646 100644 --- a/docs/guides/induce-a-program.md +++ b/docs/guides/induce-a-program.md @@ -3,7 +3,7 @@ A single demonstration is evidence, not a specification: it cannot show which values are parameters, where a branch or loop belongs, or what the failure path is. `induce` recovers a parameterized **program** from several demonstrations of -the same task, and refuses rather than guesses when the traces leave intent +the same task, and refuses when the traces leave intent underdetermined. For the model behind this, see [Multi-trace induction](../concepts/multi-trace-induction.md). diff --git a/docs/guides/policy-and-certification.md b/docs/guides/policy-and-certification.md index b5790950..40a159df 100644 --- a/docs/guides/policy-and-certification.md +++ b/docs/guides/policy-and-certification.md @@ -63,5 +63,5 @@ openadapt flow lint bundle --strict openadapt flow certify bundle --policy clinical-write ``` -Both exit nonzero on failure, so a standard CI step fails the build. The gate -turns the limits OpenAdapt discloses into requirements it enforces. +Both exit nonzero on failure, so a standard CI step fails the build. Gaps +that `lint` only reports become requirements `certify` enforces. diff --git a/docs/guides/run-reports.md b/docs/guides/run-reports.md index 3c4afddc..a929993b 100644 --- a/docs/guides/run-reports.md +++ b/docs/guides/run-reports.md @@ -49,7 +49,7 @@ When auditing a consequential run, check: 1. **Did every write verify an effect?** A write with only a screen postcondition is exactly as silent as the [five transactional faults](../concepts/effect-verification.md). Confirm the effect verdict is - CONFIRMED, not just a passing screen check. + CONFIRMED, not a passing screen check. 2. **Were consequential clicks identity armed?** Cross-check the identity-coverage line against the steps that navigate to or write a record. 3. **Did anything heal?** A heal means the UI drifted. Review the diff and diff --git a/docs/guides/security-and-data-handling.md b/docs/guides/security-and-data-handling.md index 00c5a0d3..e3019f21 100644 --- a/docs/guides/security-and-data-handling.md +++ b/docs/guides/security-and-data-handling.md @@ -84,7 +84,7 @@ your screens or records**: The same boundary governs the [attended decision path](../concepts/halt-learn-loop.md#where-a-halt-goes-the-attended-decision), -where a halted run is answered by staff — including from a phone. +where a halted run is answered by staff, including from a phone. - **The full-evidence decision surface is served by the runner, inside your boundary.** It is a responsive web app on the runner itself, not a native @@ -97,7 +97,7 @@ where a halted run is answered by staff — including from a phone. certificate, and nothing to configure on your network. That lane carries the signed PHI-free task and the closed halt context only: closed enums, bounded integers, and booleans, with **no string field and no image**. It is not - scrubbed evidence — it is an envelope that cannot represent a record. See + scrubbed evidence; it is an envelope that cannot represent a record. See [Reaching it from a phone](../concepts/halt-learn-loop.md#reaching-it-from-a-phone). - **Protected evidence does not leave the runner.** The retained screen, the observed values, the OCR, and the failing target stay local. Projections and @@ -114,7 +114,7 @@ where a halted run is answered by staff — including from a phone. The engine re-reads live state and re-runs its identity, postcondition, and effect checks before continuing, and refuses when the application is not in the state the step needs. An accepted tap never produces `VERIFIED`, and - notifications on any channel carry a fixed template and a count — never an + notifications on any channel carry a fixed template and a count, never an upstream string. - **Sending decisions back through a hosted control plane is opt-in.** It is off unless remote issuance is enabled in your deployment configuration and bound @@ -129,7 +129,7 @@ part of the local-execution path. The [data-boundary table](security-review.md#data-boundary-answers) states exactly which component can see and transmit what. -## PHI/PII posture: scrubbed where shareable, fail-closed where regulated +## PHI/PII posture: scrubbed where shareable, refused where regulated OpenAdapt treats PHI/PII handling as an engineering surface with a published map. [PRIVACY.md](https://github.com/OpenAdaptAI/openadapt-flow/blob/main/docs/PRIVACY.md) @@ -312,7 +312,7 @@ an on-prem, LAN-only appliance with no retention. Identity verification against recorded evidence before consequential clicks, typed postconditions, and effect verification against the system of record, with halt (not guess) on any non-confirmed verdict, and durable pause/resume -for human review. ([Fail-closed regulated execution](../concepts/regulated-execution.md), +for human review. ([Regulated execution](../concepts/regulated-execution.md), [The identity gate](../concepts/identity-gate.md)) **Where are credentials stored?** diff --git a/docs/guides/troubleshooting.md b/docs/guides/troubleshooting.md index 6575189e..ccd9efd2 100644 --- a/docs/guides/troubleshooting.md +++ b/docs/guides/troubleshooting.md @@ -39,7 +39,7 @@ dialog. The app looks like it is recording but captures nothing. glyph-confusable identifier** (an MRN where `O`/`0` or `l`/`1`/`I` are ambiguous) far more often than on the browser. -**Cause.** The identity gate doing its job, not a bug: it **halts rather than +**Cause.** The identity gate doing its job, not a bug: it **halts instead of guessing** a wrong-record click. On a pure-pixel substrate there is no structured accessibility text, so the identity ladder falls back to OCR, which cannot always disambiguate confusable glyphs. See diff --git a/docs/llms.txt b/docs/llms.txt index 64fda700..2736c8c0 100644 --- a/docs/llms.txt +++ b/docs/llms.txt @@ -17,7 +17,7 @@ - [The demonstration compiler](https://docs.openadapt.ai/concepts/demonstration-compiler/): Compiling a bounded demonstration into a deterministic workflow - [The substrate model](https://docs.openadapt.ai/concepts/substrate-model/): One governed runner contract across browser, desktop, and virtual-desktop surfaces - [The deployment matrix](https://docs.openadapt.ai/concepts/deployment-matrix/): Substrate, hosting, and responsibility combinations -- [Fail-closed regulated execution](https://docs.openadapt.ai/concepts/regulated-execution/): Halting instead of guessing in regulated environments +- [Regulated execution](https://docs.openadapt.ai/concepts/regulated-execution/): `run` refuses missing gates; `replay` is the local $0 path - [The capability ladder](https://docs.openadapt.ai/concepts/capability-ladder/): One compiled step, several possible implementations; the highest-fidelity viable one is used - [Effect verification](https://docs.openadapt.ai/concepts/effect-verification/): Verifying real effects against the system of record - [The identity gate](https://docs.openadapt.ai/concepts/identity-gate/): Refusing writes on low-confidence target identity diff --git a/docs/reference/cli.md b/docs/reference/cli.md index f449e8db..eaed15ce 100644 --- a/docs/reference/cli.md +++ b/docs/reference/cli.md @@ -19,7 +19,7 @@ is a subcommand of `openadapt flow`. | [`induce`](#induce) | Induce a parameterized program from **multiple** recordings | 0 if certified, 2 if underdetermined | | [`for-each`](#for-each) | Author a data-driven **loop** bundle: run one demonstration once per worklist record | 0 on success, nonzero on a mapping error | | [`replay`](#replay) | Replay a bundle, locally and deterministically | 0 on success, 1 on failure | -| [`run`](#run) | Execute a bundle through the fail-closed deployment gate | 0 success, 1 execution halt, 2 refusal | +| [`run`](#run) | Execute a bundle through the regulated admission gate | 0 success, 1 execution halt, 2 refusal | | [`resume`](#resume) | Resume a durably-paused run from its last checkpoint | 0 on success, 1/3 otherwise | | [`approve`](#approve) | Mark a durably-paused run's escalation approved | 0 on success, 1 if none | | [`teach`](#teach) | Resolve a halted run from a fix demonstration, governed | 0 if promoted, 1 if refused, 2 on bad inputs | @@ -151,7 +151,7 @@ openadapt flow compile rec --out bundle --name my-task Induce a parameterized **program** bundle from **two or more** recordings (or already-compiled bundles) of the same task: infer the shared parameters, loops, and branches. It **refuses** (writes no bundle, exits nonzero) when intent is -underdetermined, rather than guessing a branch. See +underdetermined, instead of guessing a branch. See [Induce a program from multiple traces](../guides/induce-a-program.md). ```bash @@ -248,7 +248,7 @@ VLM appliance engages only when `--allow-model-grounding` is passed **and** ## run -The same executor as [`replay`](#replay), behind a fail-closed admission gate: +The same executor as [`replay`](#replay), behind a regulated admission gate: the bundle must pass policy, identity coverage, effect coverage, approval, encryption, and manifest-integrity checks before any action executes. Backend, effect verification, API actuation, durable runtime, and policy come from diff --git a/docs/reference/configuration.md b/docs/reference/configuration.md index b9d4aee1..bbf631c5 100644 --- a/docs/reference/configuration.md +++ b/docs/reference/configuration.md @@ -128,8 +128,8 @@ that code on the runner. !!! warning "Loopback is the default, and a phone cannot reach it" Unconfigured, the portal binds `127.0.0.1` and advertises a loopback URL. - That is a complete, working configuration *on that computer* — the pairing - screen says a phone cannot reach it rather than minting a link that fails on + That is a complete, working configuration *on that computer*. The pairing + screen says a phone cannot reach it instead of minting a link that fails on your network. Publishing **this** surface to a phone is an explicit decision you make by standing up trusted TLS in front of the runner. See [Deploy on-prem](../guides/deploy-on-prem.md#reaching-the-decision-portal-from-a-phone). @@ -139,8 +139,8 @@ that code on the runner. ## Cloud phone access with no inbound ingress -The runner dials **out** to the control plane — no inbound port, no port -forward, no certificate, no reverse proxy, no static address — so a phone +The runner dials **out** to the control plane (no inbound port, no port +forward, no certificate, no reverse proxy, no static address), so a phone reaches the queue from anywhere. In Cloud **Needs attention**, scan the QR code, sign in on the same Cloud origin, and optionally enable generic Web Push alerts. The QR code carries no session, runner credential, or decision @@ -165,7 +165,7 @@ human_decisions: | `human_decisions.remote.enabled` | `false` | Must be literally `true`. A truthy string does not enable it. | | `human_decisions.remote.tenant_id` | *(unset)* | Required when enabled. | | `human_decisions.remote.runner_id` | *(unset)* | Required when enabled. | -| `human_decisions.remote.context_tier` | `remote_closed_context` | `remote_closed_context` (what broke, as closed enums and bounded integers) or `remote_identifiers` (identifiers and counts only). `local_full` is refused by name — protected evidence never leaves the runner. | +| `human_decisions.remote.context_tier` | `remote_closed_context` | `remote_closed_context` (what broke, as closed enums and bounded integers) or `remote_identifiers` (identifiers and counts only). `local_full` is refused by name; protected evidence never leaves the runner. | The execution profile applies its own ceiling, and the **weaker** of the two wins, so configuration cannot widen what a profile permits. @@ -173,16 +173,16 @@ wins, so configuration cannot widen what a profile permits. !!! note "Every misconfiguration stops the console" A missing runner credential, a deployment that did not enable remote issuance, a read-only console, or a plaintext control-plane origin each - refuse to start rather than run a console whose phone lane is silently + refuse to start instead of running a console whose phone lane is silently absent. A lane that looks on and is not is worse than one that is plainly off. Every widening step fails closed, and the portal **does not start** on an invalid -combination rather than falling back to something more exposed: +combination instead of falling back to something more exposed: - A wildcard bind address (`0.0.0.0`, `::`, empty, `*`) is refused in **every** mode. There is no "bind everything for testing" switch. -- The public origin must be a bare `https://` origin — no plaintext, no path, +- The public origin must be a bare `https://` origin: no plaintext, no path, no query, no embedded credentials, and no self-signed bypass. - `OPENADAPT_PORTAL_BIND_HOST` must be a literal IP, not a hostname, and never an unspecified or multicast address. diff --git a/docs/reference/glossary.md b/docs/reference/glossary.md index 11f0ca89..c8a9b284 100644 --- a/docs/reference/glossary.md +++ b/docs/reference/glossary.md @@ -60,7 +60,7 @@ decides whether the run can continue. ## Halt -The runtime's fail-closed refusal to act: when identity, a postcondition, an +The runtime's refusal to act: when identity, a postcondition, an effect verdict, a policy gate, or an unhandled screen state does not match the compiled expectation, the run stops and records what it observed instead of guessing. A halt is a governed outcome, not a crash; it can be answered by an @@ -71,9 +71,9 @@ operator (durable pause) or resolved permanently with `teach`. See ## Identity gate The pre-click check on consequential steps that verifies the on-screen record -identifier against the run's expected identity evidence before acting — the -wrong-record guard. A conflict or an unreadable identity band halts the run -rather than clicking into the wrong record. See +identifier against the run's expected identity evidence before acting (the +wrong-record guard). A conflict or an unreadable identity band halts the run +instead of clicking into the wrong record. See [The identity gate](../concepts/identity-gate.md). ## Reconciliation @@ -94,12 +94,12 @@ and by the run gate at execution. See ## Profile -A named runtime posture — `demo`, `standard`, or `regulated` — that selects +A named runtime posture (`demo`, `standard`, or `regulated`) that selects which requirements the run gate enforces (certification, identity coverage, effect contracts and their minimum tier, encryption, durability) and how the outcome may be described. Only Standard and Regulated runs can report `VERIFIED`. See [Run outcomes](run-outcomes.md) and -[Fail-closed regulated execution](../concepts/regulated-execution.md). +[Regulated execution](../concepts/regulated-execution.md). ## Qualification diff --git a/docs/reference/run-outcomes.md b/docs/reference/run-outcomes.md index d2ecf29d..0c9e84a9 100644 --- a/docs/reference/run-outcomes.md +++ b/docs/reference/run-outcomes.md @@ -18,8 +18,8 @@ The coarse, evidence-qualified result in `report.json` | Outcome | Meaning | What to do next | |---|---|---| | `VERIFIED` | Execution completed **and** every declared contract passed: governed authorization, identity coverage, postconditions, and every effect confirmed at or above the required tier. Only possible under the Standard or Regulated profile. The only production success. | Nothing. Archive the run directory; the report and receipt are the audit evidence. | -| `COMPLETED_UNVERIFIED` | Execution reached the end, but the run lacked the evidence to claim `VERIFIED` — typically a Demo-profile replay with no independent effect verifier. Never billable, never a production success. | Fine for development. For production, wire an effect verifier and run under the Standard or Regulated profile so the same workflow can terminate `VERIFIED`. See [Run a deployment](../guides/run-a-deployment.md). | -| `HALTED` | The runtime **refused to act** on a governed check: identity, postcondition, effect verdict, an unhandled state, or a policy gate. A halt is fail-closed behavior, not a crash. | Read the halt reason in `REPORT.md` (categories below). Usually: [`teach`](cli.md#teach) the correction, or fix the target/parameters and re-run. Do not blind-retry a run whose transaction outcome is `RECONCILIATION_REQUIRED`. | +| `COMPLETED_UNVERIFIED` | Execution reached the end, but the run lacked the evidence to claim `VERIFIED` (typically a Demo-profile replay with no independent effect verifier). Never billable, never a production success. | Fine for development. For production, wire an effect verifier and run under the Standard or Regulated profile so the same workflow can terminate `VERIFIED`. See [Run a deployment](../guides/run-a-deployment.md). | +| `HALTED` | The runtime **refused to act** on a governed check: identity, postcondition, effect verdict, an unhandled state, or a policy gate. A halt is a governed stop, not a crash. | Read the halt reason in `REPORT.md` (categories below). Usually: [`teach`](cli.md#teach) the correction, or fix the target/parameters and re-run. Do not blind-retry a run whose transaction outcome is `RECONCILIATION_REQUIRED`. | | `FAILED` | A non-governed runtime failure (for example the browser or agent connection died) rather than a safety refusal. | Check the environment (target reachable, backend agent up, permissions), then re-run. If the transaction outcome is `FAILED_PLATFORM`, no business effect occurred. | | `ROLLED_BACK` | A detected duplicate or collateral write was **compensated** and the compensation was re-verified. Non-success, but the system of record was restored. | Review the compensation entries in the effect journal, confirm the record state, then address the root cause before re-running. | @@ -31,7 +31,7 @@ evidence proves about the **business effect** when a run stops. | Outcome | Meaning | What to do next | |---|---|---| | `VERIFIED` | Every declared effect passed at or above the required tier under a production profile. | Nothing; this is the billable success. | -| `HALTED_BEFORE_EFFECT` | The run stopped and the evidence proves **no consequential write landed**: every consequential step was verified absent or stopped before delivery was attempted. | Safe to fix and re-run. This is the honest version of "nothing happened". | +| `HALTED_BEFORE_EFFECT` | The run stopped and the evidence proves **no consequential write landed**: every consequential step was verified absent or stopped before delivery was attempted. | Safe to fix and re-run. This is the evidenced version of "nothing happened". | | `RECONCILIATION_REQUIRED` | Delivery or persistence is **uncertain, conflicting, or unverifiable**. A write may have half-landed. | Do **not** re-run yet. Reconcile against the independent system of record first (find or rule out the record), then re-run. The per-step effect journal in the report shows which step is uncertain. | | `FAILED_PLATFORM` | An OpenAdapt/platform failure before any possible business effect. Never billable. | Re-run after the platform issue is resolved; report persistent cases. | | `CANCELED` | The run was canceled before any business effect could occur. | Re-run when ready. | @@ -48,7 +48,7 @@ system of record, recorded per effect in the report's evidence: |---|---|---| | `confirmed` | The verifier independently observed the intended effect in the system of record. | Nothing; this is what `VERIFIED` is built from. | | `refuted` | The verifier affirmatively observed the effect is **absent** (or wrong). The run halts. | The write did not land as intended. If the observed effect is `absent`, the run maps to `HALTED_BEFORE_EFFECT` and is safe to re-run after fixing the cause; if `conflicting`, reconcile the duplicate/wrong record first. | -| `indeterminate` | The verifier could not establish presence or absence (unreachable, ambiguous read). The run halts — an unreachable verifier is **never** treated as success. | Check verifier connectivity and configuration (`--effects-kind`, `--effects-base-url`). Treat the write as uncertain: reconcile before re-running. | +| `indeterminate` | The verifier could not establish presence or absence (unreachable, ambiguous read). The run halts; an unreachable verifier is **never** treated as success. | Check verifier connectivity and configuration (`--effects-kind`, `--effects-base-url`). Treat the write as uncertain: reconcile before re-running. | The evidence also records what the verifier **observed** about the record, independent of the verdict: `present`, `absent`, `conflicting` (a record was @@ -63,18 +63,18 @@ When a run reports `HALTED`, `REPORT.md` names the violated expectation and | Halt reason (code) | Stage | What happened | Remediation | |---|---|---|---| | `target_ambiguous` | target resolution | More than one (or zero) candidate matched the step's recorded evidence; acting would be a guess. | Re-record the step with cleaner evidence, or [`teach`](cli.md#teach) the disambiguation. Prefer a structural backend (DOM / UIA / AX / AT-SPI) over pixels where the app exposes one. | -| `identity_conflict` | identity verification | The pre-click identity check read a **different record identifier** than the run's parameters expect — the wrong-record gate firing. | First verify the parameters or worklist row are correct. If the UI legitimately changed, `teach` the correction. Do not lower the gate: see [Troubleshooting](../guides/troubleshooting.md#it-halts-too-much-over-halting-on-citrix-pixel-only). | +| `identity_conflict` | identity verification | The pre-click identity check read a **different record identifier** than the run's parameters expect (the wrong-record gate firing). | First verify the parameters or worklist row are correct. If the UI legitimately changed, `teach` the correction. Do not lower the gate: see [Troubleshooting](../guides/troubleshooting.md#it-halts-too-much-over-halting-on-citrix-pixel-only). | | `identity_unverifiable` | identity verification | The identity band could not be read confidently (common on pixel-only substrates where OCR meets confusable glyphs). | Raise substrate fidelity (structural backend), re-capture the identifier crop, or `teach` the specific case. Over-halting beats a silent wrong write. | | `actuation_observation_changed` | actuation revalidation | The screen changed between resolving the target and acting on it; the runtime refused to click a stale observation. | Usually transient (a late toast or dialog): re-run. If it recurs at the same step, `teach` the interstitial state so the program handles it. | | `api_path_unavailable` | API admission | A step bound to the API actuation tier could not use it. | Check `--api-base-url` / the deployment config's actuation section and the credentials it references. | -| `effect_strength_insufficient` | effect strength | The configured verifier cannot meet the minimum verification tier the policy or profile requires for this write. | Wire a stronger verifier in the [deployment configuration](deployment-config.md), or revisit the required tier in the policy — a deliberate governance decision, not a tweak. | +| `effect_strength_insufficient` | effect strength | The configured verifier cannot meet the minimum verification tier the policy or profile requires for this write. | Wire a stronger verifier in the [deployment configuration](deployment-config.md), or revisit the required tier in the policy (a deliberate governance decision, not a tweak). | | `effect_verifier_missing` | effect verification | The step declares effects but no verifier is configured; the profile refuses to treat the write as verified. | Configure `--effects-kind` (`rest`, `fhir`, `document-hash`) or, in Demo development only, explicitly approve the unverified write. | Beyond the typed refusal codes, a run also halts on an **unhandled state**: a resolution failure, a dead-end branch, an unmet `halt` guard, a non-confirmed effect, or an explicit `halt` terminal in the program. The report records where it stopped, what unexpected on-screen state it observed, and the steps that -succeeded before it — exactly the evidence [`teach`](cli.md#teach) consumes. +succeeded before it, which is exactly the evidence [`teach`](cli.md#teach) consumes. See [The halt-learn loop](../concepts/halt-learn-loop.md). `report.json` also carries a machine-readable `failure_category` per failed @@ -104,14 +104,14 @@ Terminal outcomes mirror the engine taxonomy in lowercase: `verified`, | Command | Exit codes | |---|---| | [`replay`](cli.md#replay) | `0` success (`VERIFIED` under Standard/Regulated); `1` on a halt or failure. | -| [`run`](cli.md#run) | Same as `replay` once admitted; `2` when a fail-closed gate (certification, policy, profile requirement) refuses the run before it starts. | +| [`run`](cli.md#run) | Same as `replay` once admitted; `2` when an admission gate (certification, policy, profile requirement) refuses the run before it starts. | | [`tutorial`](../get-started/index.md) | `0` when `VERIFIED`; `1` on any other outcome; `2` when the tutorial is refused before running. | | [`lint`](cli.md#lint) | `0` clean (or advice only); `1` once a finding reaches `error` severity (`--strict`: also on warnings). | -| [`certify`](cli.md#certify) | `0` pass; `2` when the bundle fails certification — the gate refusing an unsafe bundle. | +| [`certify`](cli.md#certify) | `0` pass; `2` when the bundle fails certification (the gate refusing an unsafe bundle). | | [`teach`](cli.md#teach) | `0` correction promoted; `1` governed refusal (nothing written); `2` unusable inputs. | A nonzero exit from `lint`, `certify`, or a halted `replay` is the safety -boundary doing its job — the fail-closed design refuses to guess. See the +boundary doing its job. The run stopped instead of guessing. See the per-reason remediation above before changing any gate. ## Still stuck? diff --git a/mkdocs.yml b/mkdocs.yml index 59ca9cd5..d0de22cf 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -166,7 +166,7 @@ nav: - Security and deployment review: guides/security-review.md - Account security and privileged access: guides/account-security.md - Qualification evidence: get-started/what-works-today.md - - Fail-closed regulated execution: concepts/regulated-execution.md + - Regulated execution: concepts/regulated-execution.md - Settings and policy governance: concepts/settings-governance.md - Security and data handling: guides/security-and-data-handling.md - PHI handling in one page: commercial/phi-handling.md