Skip to content

Repository files navigation

Agentflow

Agentflow is a local-first supervised loop runtime for agent workflows in real repositories.

You write a graph that states the intent, authority, context, tools, validation, and artifacts for the work. Agentflow turns that graph into an explicit loop: graph contract -> agent attempt -> mechanical completion -> fresh verification -> causal recovery -> delivery evidence. It validates the contract, runs substantial nodes through agent harnesses such as Codex CLI or Cursor CLI, supervises failures, and produces a durable delivery package for review.

Agentflow exists because long-running agent work needs more than an ad hoc prompt. Teams need the original intent preserved, the right context pointed to clearly, failures repaired without losing the thread, and final evidence organized so a human can review the result.

flowchart LR
  intentNode["Human intent: what should be done"] --> graphNode["Agentflow graph: contract and authority"]
  graphNode --> harnessNode["Agent harness: Codex CLI or Cursor CLI"]
  graphNode --> checksNode["Checks and criteria: hard gates and rubrics"]
  harnessNode --> artifactNode["Artifacts and logs: what happened"]
  checksNode --> supervisorNode["Supervisor: observe, repair, or request authority"]
  artifactNode --> supervisorNode
  supervisorNode --> deliveryNode["Delivery package: review-ready evidence"]
Loading

What It Is

Concept What it means Why it matters Details
Graph A file-backed execution contract for a workflow. Keeps intent, authority, context, validation, and delivery explicit before launch. Product scope, operations
Node A meaningful unit of work: agent, exec, check, checkpoint, or a container such as sequence, parallel, or repeat. Lets agents own outcomes instead of receiving tiny brittle prompt fragments. Examples
Context Pointer metadata Agentflow gives a node: selected files, globs, prior artifacts, plugin files, and supervisor repair packets. Context quality is prompt quality. Too much noise hurts runs; missing context causes avoidable failures. Context and artifacts
Artifacts Named durable outputs produced by nodes. Future nodes and reviewers consume artifacts, not hidden chat state. Runtime lifecycle
Checks Deterministic or AI gates inside the run. Provides hard evidence and semantic sensors without relying on final prose alone. Outcome verification
Supervisor The runtime recovery system for failed or misaligned node attempts. Keeps the graph progressing when context, validation, artifact, workspace, or environment issues are machine-fixable. Architecture
Managed patterns Higher-level workflow shapes that compile into normal graph execution. Gives authors simple contracts for common research, work, and discovered-list loops. Managed patterns
Plugins Packaged workflow exports and CLI tools. Lets teams expose reusable capabilities, credential boundaries, and composed CLI behavior. Plugins
Evals Offline workflow evaluations across scenarios, variants, trials, criteria, trajectory, and simulation. Gives teams confidence in graphs, plugins, prompts, tools, and supervisor behavior. Evals
Delivery Terminal run package with summaries, evidence, decisions, risks, and review order. Lets humans review the result without spelunking raw logs first. Operations

Why It Exists

Agentflow is for work that should be accountable and repeatable:

  • Implementing code changes that need clear validation and review evidence.
  • Running multi-node investigations where downstream work needs durable artifacts.
  • Splitting large repo work into planned, validated, reviewable chunks.
  • Giving agents access to local tools while keeping credential and tool boundaries explicit.
  • Evaluating workflow quality over repeated realistic trials instead of judging one lucky run.

Agentflow is not an invisible planner, remote devbox, or alternate chat memory system. The graph is the source of truth. Runtime state, context packets, supervisor interventions, and delivery artifacts are written to disk so the run can be inspected, resumed, and reviewed.

How It Works

flowchart TD
  author["Author graph\nintent, repos, profiles, nodes"] --> validate["agentflow validate\nschema, review, run-ready context"]
  validate --> compile["Compile graph\nprimitive runtime contract"]
  compile --> schedule["Scheduler\nsequence, parallel, repeat"]
  schedule --> attempt["Node attempt\ncontext, tools, output dir"]
  attempt --> harness["Harness execution\nCodex CLI or Cursor CLI"]
  harness --> verify["Outcome verification\nartifacts, response, diff, logs"]
  verify --> passed{"Passed?"}
  passed -- yes --> next["Next node or delivery"]
  passed -- no --> supervisor["Supervisor recovery\ncausal cone, target, material delta"]
  supervisor --> target{"Recovery target"}
  target -- current or upstream node --> attempt
  target -- authority boundary --> delivery["Paused delivery evidence"]
  target -- impossible runtime invariant --> delivery
  next --> delivery
Loading

Agentflow keeps three contracts separate:

flowchart LR
  authored["Authored graph\nhuman contract"] --> compiled["Compiled graph\nruntime contract"]
  compiled --> runroot["Run root\naudit contract"]

  authored --> authoredFields["intent, repos, profiles,\nconstraints, context, artifacts"]
  compiled --> compiledFields["primitive nodes, edges,\nresolved profiles and tools"]
  runroot --> runFields["attempts, context packets,\nlogs, interventions, delivery"]
Loading

That separation is why validation can explain the workflow before launch, execution can resume from durable state, and review can start from delivery/ instead of raw node directories.

Supervisor Recovery

Every executable node is a supervised checkpoint. The supervisor observes healthy attempts and stays out of the way. When a node fails or is rejected, the failed node is treated as a symptom: the supervisor builds an upstream causal cone, chooses the nearest intent-aligned recovery target, repairs within that target's existing authority, reruns the failed gate, and records the recovery chain.

flowchart TD
  symptom["Failed or rejected checkpoint"] --> casefile["Causal case file\nprompt, context, logs, artifacts, diff"]
  casefile --> cone["Upstream causal cone\nedges, artifacts, context, attempts"]
  cone --> rank["Rank recovery targets\ncurrent, upstream, artifact, context, workspace"]
  rank --> repair["Machine repair\nwithin target authority"]
  repair --> delta["Material delta\nwhat actually changed"]
  delta --> rerun["Rerun failed gate"]
  rerun -- healthy --> continue["Continue graph"]
  rerun -- new symptom --> cone
  rank -- typed authority request --> pause["Pause for human authority"]
Loading

Human pause is reserved for trusted typed AuthorityRequest records from runtime-owned producers: missing credentials, missing harness authentication, planned checkpoints, external side-effect approval, or explicit operator pause. Product ambiguity, graph-contract gaps, repo/sandbox/scope expansion, context, validation, artifact, workspace, and local environment failures recover autonomously or fail contractually with evidence.

Setup

Agentflow is a Node project. Use Node 24 LTS. With nvm, run:

nvm install
nvm use
npm install
npm run build
npm run setup:link
agentflow --help
agentflow graph-help

npm run setup:link links the built package so local usage matches how operators invoke Agentflow: agentflow .... The linked CLI uses the node executable on your current PATH, so run nvm use before invoking agentflow in a new shell, or set Node 24 as your default. The repository still has development scripts, but examples and docs should use the linked CLI for Agentflow commands.

To remove the linked CLI:

npm run setup:unlink

Core repository validation:

npm run typecheck
npm test
npm run build
npm run validate:smoke

validate:smoke runs lightweight package checks and verifies the built CLI against the repeat fixture across Codex CLI and Cursor CLI adapters with both inplace and worktree workspace backends. It does not rerun the full npm test suite.

First Run

  1. Create agentflow.graph.json from the minimal graph below or an example under docs/examples.

  2. Validate the graph before launching:

    agentflow validate --graph agentflow.graph.json
  3. Inspect the compiled shape when the graph changes:

    agentflow validate --graph agentflow.graph.json --show-compiled
    agentflow validate --graph agentflow.graph.json --diagram-output graph.mmd
  4. Run the graph:

    agentflow run --graph agentflow.graph.json
  5. Review the terminal delivery package, starting with delivery/01-review-brief.md, then delivery/02-run-learnings.md and delivery/03-audit-index.md when needed.

Minimal Graph

This is the canonical small graph shape: explicit repo, profiles, supervisor profile, node-level intent, declared artifact, and deterministic check.

{
  "version": "1",
  "graph_id": "ship-reviewable-change",
  "intent": {
    "goal": "Implement a focused change and leave it ready for review.",
    "constraints": [
      "Do not turn the graph into an implementation playbook.",
      "Do not include unrelated refactors."
    ],
    "acceptance_criteria": [
      "The change is implemented.",
      "Tests or checks provide evidence.",
      "The review brief explains risk and review order."
    ]
  },
  "repos": {
    "main": {
      "path": "."
    }
  },
  "defaults": {
    "launch_profile": "codex",
    "workspace_backend": "worktree"
  },
  "profiles": {
    "codex": {
      "harness": "codex-cli",
      "model": "gpt-5-codex",
      "reasoning_effort": "medium",
      "sandbox": "workspace-write"
    },
    "cursor": {
      "harness": "cursor-cli",
      "model": "auto",
      "sandbox": "workspace-write"
    },
    "supervisor": {
      "harness": "codex-cli",
      "model": "gpt-5-codex",
      "reasoning_effort": "medium",
      "sandbox": "workspace-write"
    }
  },
  "supervision": {
    "profile": "supervisor",
    "max_total_interventions": 3
  },
  "graph": {
    "type": "sequence",
    "id": "root",
    "steps": [
      {
        "type": "agent",
        "id": "implement_slice",
        "runtime": {
          "repo": "main",
          "profile": "codex"
        },
        "intent": {
          "goal": "Implement the scoped change and leave reviewer-ready evidence.",
          "acceptance_criteria": [
            "Targeted validation is run or clearly explained.",
            "The `change_summary` artifact is published with `af artifact write change_summary`.",
            "The handoff names changed files, validation, and residual risks."
          ],
          "constraints": [
            "Do not rely on the final response as the durable handoff."
          ]
        },
        "support": {
          "context": [
            {
              "name": "goal",
              "kind": "workspace_file",
              "path": "README.md",
              "what": "Repository overview and local workflow guidance.",
              "why": "It helps the implementation node stay aligned with the current repo contract."
            }
          ]
        },
        "artifacts": {
          "change_summary": {
            "from": "output_dir",
            "path": "change-summary.md",
            "description": "Implementation summary written by the agent."
          }
        }
      },
      {
        "type": "check",
        "id": "test",
        "runtime": {
          "repo": "main"
        },
        "intent": {
          "goal": "Run the repository test suite to validate the scoped change.",
          "acceptance_criteria": [
            "`npm test` exits successfully.",
            "The check output is usable as reviewer evidence."
          ],
          "constraints": []
        },
        "check_kind": "deterministic",
        "command": "npm",
        "args": [
          "test"
        ]
      }
    ]
  }
}

Switching from Codex CLI to Cursor CLI is a launch-profile choice, not a different graph language. Both harnesses receive the same context packet, runtime CLI, optional support metadata, artifact contract, and output directory. Normal worker profiles inherit native Codex/Cursor config by default; set profiles.*.harness_config.isolation: "isolated" when reproducibility matters more than native harness parity. Codex profiles default approval_policy to "never" so Agentflow owns recovery and authority pauses. model: "auto" means Agentflow does not pass an explicit model flag to the selected harness.

Codex runs have network access by default, including AI judges. They can reach the internet and start local test servers. The graph's sandbox still controls file access: a read-only judge cannot edit app files. See Codex permissions to turn network access off for a profile.

CLI Commands

Task Command
See graph syntax help agentflow graph-help
Resolve plugin packages and lock tools agentflow plugin resolve --graph agentflow.graph.json
Validate launch readiness agentflow validate --graph agentflow.graph.json
Fail on serious authoring findings agentflow validate --graph agentflow.graph.json --strict
Inspect compiled runtime contract agentflow validate --graph agentflow.graph.json --show-compiled
Write validation package agentflow validate --graph agentflow.graph.json --output-dir .task-runtime/validation/latest
Write Mermaid diagram agentflow validate --graph agentflow.graph.json --diagram-output graph.mmd
Render graph image agentflow validate --graph agentflow.graph.json --diagram-image-output graph.svg
Launch a run agentflow run --graph agentflow.graph.json
Inspect a run root agentflow inspect <run-root>
Resume a paused or failed run agentflow resume --run-root <run-root>
Resume the latest run for a graph agentflow resume --graph agentflow.graph.json --latest
List graph runs agentflow runs list --graph agentflow.graph.json
Validate an eval suite agentflow eval validate evals/<suite-id>
Run an eval suite agentflow eval run evals/<suite-id> --variant current --scenario all --trials 1
Generate an eval report agentflow eval report .task-runtime/evals/<eval-run>

Image export uses npx -y @mermaid-js/mermaid-cli by default. Use --diagram-image-package to choose a package spec, or --diagram-image-renderer mmdc with AGENTFLOW_MERMAID_CLI_BIN for an installed binary.

Documentation Map

Reader Start here What you get
New reader docs/README.md Full documentation map split by product, technical, and examples.
Workflow author docs/product/operations.md Validation, launch, resume, inspect, and delivery workflow.
Product reviewer docs/product/scope.md Active product boundary and release bar.
Managed-pattern author docs/product/managed-patterns.md, patterns When to use managed patterns and how their contracts behave.
Eval author docs/product/evals.md Suites, scenarios, variants, criteria, trajectory checks, simulation, reports, and comparison.
Plugin author docs/product/plugins.md Workflow exports, CLI tool exports, credentials, config, naming, and consumption.
Runtime implementer docs/technical/README.md Implementation reading order for runtime, context, tools, verification, and delivery.
Runtime debugger docs/technical/runtime-lifecycle.md Launch-to-delivery execution flow.
Context debugger docs/technical/context-and-artifacts.md Context pointer resolution, artifact refs, and downstream handoffs.
Tooling debugger docs/technical/runtime-tooling.md Generated af and plugin tool wrappers.
Example user docs/examples/README.md Runnable graph, eval, and plugin examples.
Agent authoring with skills skills Agentflow, plugin, and eval skill guidance aligned to the repo contract.

Graph Contract At A Glance

Field Purpose
version Graph schema version. Current value is "1".
graph_id Stable id used for run roots and inspection.
intent Top-level goal, constraints, and acceptance criteria.
repos Local repository aliases. Defaults to main at . when omitted.
defaults Launch profile and workspace backend defaults.
profiles Harness, model, sandbox, env, artifact repair, and harness-native config isolation.
supervision Required supervisor profile plus total recovery budget.
plugins Reusable workflow and managed tool packages resolved into the graph.
skill_sources Installable or local skill collections; only referenced skills are prompted.
tools Registry of managed plugin tool declarations; not a global grant.
capabilities Reusable bundles of skill refs, managed tool grants, and ambient CLI hints.
graph The execution shape: containers, executable nodes, or managed patterns.

Executable nodes are agent, exec, check, and checkpoint; all require intent.goal and non-empty intent.acceptance_criteria, with optional intent.constraints normalized to []. Constraint strings should start with Do not; positive requirements belong in acceptance criteria. Containers are sequence, parallel, and repeat. Managed patterns are pattern_deep_research, pattern_deep_work, and pattern_work_list.

Executable nodes choose repo/profile in runtime and receive non-authoritative help in support. support.context entries require what and why; skills and managed tools are selected directly or through support.capabilities; CLI hints are plain shell commands validated as callable and rendered in the prompt without wrappers, config, credentials, or ledgers.

Use checkpoint for authored human gates, usually inside a repeat body. Supervisor authority pauses are different: they are runtime pauses chosen only from trusted typed AuthorityRequest records. Free text from agents, verifiers, stderr, helper artifacts, or debug logs cannot pause a run.

Runtime Surfaces

Surface Who uses it Purpose
agentflow Humans and automation outside a run. Validate, run, resume, inspect, observe live runs, resolve plugins, auth, eval, and report.
af Agents inside a node attempt. Orient to the node contract, track milestone evidence, publish declared artifacts, and check completion readiness.
Run root Operators and debuggers. Durable state, events, attempts, context packets, logs, supervisor interventions, and delivery files.
delivery/ Human reviewers. High-signal terminal package with review brief, run learnings, audit index, and semantic evidence files.

Agentflow injects af into every agent node on PATH. Humans normally do not use af outside a running node.

Eval Setup

Agentflow includes local-first eval suites for workflow quality and real-world issue repair. Some suites materialize ignored local repos before running.

npm run setup:eval-repos
npm run setup:realworld-evals
agentflow eval validate evals/agentflow-validation
agentflow eval validate evals/agentflow-capability-workflows
agentflow eval validate evals/agentflow-realworld-issues

The eval architecture follows Anthropic's Demystifying evals for AI agents as the primary workflow-eval reference and adopts useful ADK mechanics for criteria, trajectory, and deterministic environment simulation. See docs/product/evals.md for suite authoring and operation.

Repository Map

Path Purpose
src/graph/ Authored schema, normalization, validation, review, Mermaid diagrams, and compilation.
src/runtime/ Scheduler, execution engine, harness calls, context resolution, supervision, resume, and delivery.
src/supervisor/ Policy, failure classification, recovery planning, and runtime overlays.
src/plugins/ Local or Git plugin workflows, tool exports, and credential metadata.
src/artifacts/ Run-root paths, event projection, reconciliation, and artifact readers.
src/cli/ Human/operator CLI commands and progress rendering.
docs/ Product docs, technical docs, and examples.
evals/ Committed eval suite definitions and templates.
skills/ Agentflow skills for graph authoring, plugins, and evals.
scripts/ Setup and validation scripts.
tests/ Unit and runtime tests.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages