If your senior engineer goes on holiday for two weeks and your agents keep shipping — do you trust what comes out the other side?
Harness Engineering is the gear list that makes the answer yes. The framing draws on a growing body of work on harness engineering — Ajey Gore's "The Solo Climb" and "The Anatomy of an AI-Native Org", OpenAI's "Harness engineering", and Martin Fowler's "Harness engineering for coding agents" — distilled to one question: shipping with agents unsupervised isn't a prompt problem, it's an equipment problem. The holiday test is Gore's; The Solo Climb names the gear you need before that holiday is safe. Here is each piece — and what harness ships for it.
-
Specs you operate from — the input the agent works from, the source the eval checks against, and the artifact the team reviews when something breaks. harness both authors and polices them:
brainstormingandspec-craftproduce specs and ADRs, andacceptance-evalrefuses a spec that lacks measurable, testable acceptance criteria. -
A test suite the team trusts enough to ship on — green means ship, red means don't. harness doesn't write your tests, but it makes them load-bearing:
test-advisorselects the tests a change actually needs and audits coverage gaps, and theverifygate runs test/lint/typecheck as a single hard pass/fail. -
An eval suite for "did it solve the right problem" — the layer above the unit tests that catches the change which passes every test and still does the wrong thing. harness ships
outcome-eval, a blocking post-execution gate that judges the diff against the spec's acceptance criteria before the change is allowed to ship. -
A sandboxed environment the agent can operate in — bounded, observable, reversible, so a mistake's blast radius never reaches the user. harness covers this partly: isolated git worktrees bound the work, session search and
insightsmake it observable, and therollbackskill proposes a full-context revert when a shipped change fails. It orchestrates your environment rather than providing the sandbox itself. -
Gates that actually gate — the build either goes to production or it doesn't; no soft gates that warn and wave you through. This is the core of harness: architectural boundaries enforced by ESLint, CI checks that fail the build, a
block-no-verifyhook that refuses--no-verify, and phase gates that won't advance on a failed checkpoint. -
Agent-of-agent review — one agent writes, another reviews, a third runs the tests. harness ships a multi-phase
code-reviewpipeline that fans work out to parallel persona reviewers (architecture, security, TypeScript-strict, frontend-races, adversarial) and supports peer review between agents.
AI coding agents are powerful, but unreliable without structure. Left unconstrained, they introduce circular dependencies, violate architectural boundaries, and generate drift that compounds across a codebase. Teams respond with code review backlogs and manual checklists — trading agent speed for human bottlenecks.
Harness Engineering takes a different approach: mechanical enforcement, not hope.
Instead of relying on prompts and conventions, harness encodes your architectural decisions as machine-checkable constraints. Agents get real-time feedback when they violate boundaries. Entropy is detected and cleaned automatically. Every rule is validated on every change.
For tech leads and architects: Scale AI-assisted development across your team with confidence. Define constraints once, enforce them everywhere — across agents, developers, and CI.
For individual developers: Stop babysitting your AI agent. Give it guardrails and let it execute. Spend your time on design decisions, not cleanup.
- Cross-Platform Support — Fully tested on Windows, macOS, and Linux with mechanical enforcement preventing platform-specific regressions
- Context Engineering — Repository-as-documentation keeps agents grounded in project reality, not stale training data
- Architectural Constraints — Layered dependency rules enforced by ESLint, not willpower
- Agent Feedback Loop — Self-correcting agents with peer review and real-time validation
- Entropy Management — Automated detection of dead code, doc drift, and structural decay
- Implementation Strategy — Depth-first execution: one feature to 100% before the next begins
- Key Performance Indicators — Measure agent autonomy, harness coverage, and context density
- Orchestrator Gateway API — Token-scoped bearer auth on a versioned
/api/v1/*surface with append-only audit log, three bridge-primitive endpoints (jobs/maintenance,interactions/{id}/resolve,eventsSSE), HMAC SHA-256-signed webhook subscriptions (X-Harness-Signature: sha256=<hex>) with event-bus fan-out, and a vendored OpenAPI artifact atdocs/api/openapi.yaml. External bridges (Slack, Discord, GitHub Apps) build against a published, versioned contract instead of coupling to internals. See ADR 0011. Seeexamples/slack-echo-bridge/for the canonical reference consumer — a standalone Node bridge that verifies HMAC signatures and posts to Slack onmaintenance.completed. - Granular Task Routing — Spec B extends
agent.routingwith per-skill and per-cognitive-mode axes, fallback chains, and a/routingdashboard panel +harness routing traceCLI for inspecting decisions. See the Per-skill and per-mode routing section. - Session Search & Insights — SQLite FTS5 index over
.harness/sessions/and.harness/archive/sessions/with BM25 ranking; LLM-generated retrospectivellm-summary.mdwritten on session archive; compositeharness insightsaggregator combining health, entropy, decay, attention, and impact. New CLI commandsharness search "<query>"andharness insights, plus MCP toolssearch_sessions,summarize_session,insights_summary. See ADR 0013. - Skill Proposals — Agents emit skill candidates (new or refinement) via the
emit_skill_proposalMCP tool; proposals queue in.harness/proposals/and route through a mechanical soundness gate before reviewer approval. Every skill carriesprovenance: community | agent-proposed | user-authored. CLI:harness proposals list|show|approve|reject. Dashboard review queue at/s/proposals. See ADR 0016. - Local Model Lifecycle Manager — Opt-in (
localModels.enabled) autonomy for the local model pool: hardware-aware ranking, disk-budget-bounded install/swap/evict through the review queue, and aLocalModelResolverthat consumes pool state. Manualharness models+ resolver-from-pool + drift reconciliation ship today; autonomous swap proposals await the Phase-2 candidate parser. See the operator guide, ADR 0061, and ADR 0062.
Using a coding agent? Point it at the autonomous setup prompt and let it install + initialize harness for you:
https://raw.githubusercontent.com/Intense-Visions/harness-engineering/main/docs/agent-setup/prompt.md— "follow the instructions at this URL".
Pick the install path that matches how you use harness:
- Claude Code users → install the
harness-claudemarketplace plugin (recommended). Skills, slash commands, persona subagents, lifecycle hooks, and MCP are wired up automatically — noharness setupstep. - Cursor users → install the
harness-cursormarketplace plugin (recommended). Same component surface as Claude, plus 4 curated project rules. - Gemini CLI users → install the
harness-geminimarketplace extension. Slash commands + GEMINI.md context + MCP. (Gemini extensions don't define a subagents or hooks field, so those surfaces live in GEMINI.md.) - Codex CLI users → install the
harness-codexmarketplace plugin. Skills + MCP (Codex's plugin spec defines no slash-command or agents surface). - OpenCode users → install the npm package and run
harness setup. OpenCode auto-discovers.claude/skills/and shares Claude's skill tree, so the only setup work is wiring the harness MCP server intoopencode.json, whichharness setupdoes automatically once it detects~/.config/opencode/or a project-localopencode.json. - Plain CLI / CI users (or any tool not yet covered by a plugin) → install the npm package.
harness setupdetects every supported AI client (Claude Code, Gemini CLI, Cursor, Codex CLI, OpenCode) and lays down skills, slash commands, agent personas, MCP, and hooks.
In a Claude Code session:
/plugin marketplace add Intense-Visions/harness-engineering
/plugin install harness-claude
This pulls every bundled skill (so trigger phrases like "scaffold a test suite" engage initialize-test-suite-project automatically), registers the /harness:* slash commands, installs the 12 persona subagents (harness-code-reviewer, harness-architecture-enforcer, …), wires the standard hook profile (block-no-verify, protect-config, quality-gate, pre-compact-state, adoption-tracker, telemetry-reporter), and starts the harness MCP server via npx @harness-engineering/cli harness-mcp. No per-repo harness setup is required.
In Cursor, open the marketplace and install harness-cursor from Intense-Visions/harness-engineering. Same skills, slash commands, subagents, hooks, and MCP server as the Claude plugin, plus 4 curated project rules (validate-before-commit, respect-architecture, use-harness-skills, respect-hooks) that fire automatically as alwaysApply rules in every Cursor session in this repo.
npm install -g @harness-engineering/cli
harness setupThis installs the CLI and runs interactive setup: generates global slash commands and agent personas for all detected AI clients (Claude Code, Gemini CLI, Cursor, Codex CLI), configures MCP servers, and sets up peer integrations. Once set up, every project on your machine has access to /harness:* slash commands, agent personas, and the harness-mcp server binary — no per-project setup needed.
Tip: Re-run
harness setupafter updating the CLI (harness update) to pick up new or changed skills. Marketplace plugin users update via/plugin update harness-claude(orharness-cursor).
The marketplace plugins are the agent-session interface. The npm package is what you need for shell-level workflows. Pick based on where you actually use harness:
| Surface | harness-claude / harness-cursor plugin (after /harness:initialize-project runs once per repo) |
npm install -g @harness-engineering/cli (after harness setup) |
|---|---|---|
Inside a Claude Code or Cursor session — skills, /harness:*, subagents, hooks, MCP tools |
✅ full parity | ✅ same |
| Cursor project rules (validate-before-commit, respect-architecture, use-harness-skills, respect-hooks) | ✅ shipped with harness-cursor |
❌ Cursor-only feature |
| Adoption tracking + anonymous telemetry (hooks fire, defaults enabled) | ✅ works | ✅ works |
Project bootstrap (harness.config.json, .harness/ scaffolding) |
✅ via /harness:initialize-project (Phase 2) |
✅ via harness init / harness setup |
Knowledge graph (.harness/graph/) |
✅ via /harness:initialize-project (Phase 5 step 1) |
✅ initial scan during harness setup |
| Architecture / performance baselines | ✅ via /harness:initialize-project (Phase 5 steps 2–3) |
✅ auto-refreshed on main via CI |
| Telemetry identity tagging (project / team / alias) | ✅ via /harness:initialize-project (Phase 5 step 4) |
✅ via interactive telemetry wizard |
Legacy layout migration warnings (docs/plans/, .harness/architecture/) |
✅ via /harness:initialize-project (Phase 5 step 5) |
✅ surfaced during harness setup |
Tier-0 MCP integrations (context7, sequential-thinking, playwright) added to project .mcp.json |
✅ via /harness:initialize-project (Phase 5 step 6) |
✅ wired during interactive setup |
| Tier-1 API-key integrations (Linear, Slack, Perplexity, …) | harness integrations list; user wires via npx … add <name> |
|
| Gemini CLI / Codex integration | ✅ via sibling marketplace plugins (harness-gemini, harness-codex) |
✅ harness setup configures all detected clients |
| OpenCode integration | ✅ harness setup writes opencode.json with the harness MCP server (and Tier-0 integrations) when ~/.config/opencode/ or a project-local opencode.json is present |
|
Terminal use — harness validate, harness init, harness check-arch |
npx @harness-engineering/cli <cmd> |
✅ binary in PATH |
| CI workflows (GitHub Actions, etc.) | npx (cold-start cost per job) |
✅ npm install -g once, fast thereafter |
Git pre-commit hooks (harness validate on commit) |
✅ direct binary, fast |
TL;DR: run /harness:initialize-project once per repo and the plugin covers ~95% of what harness setup does — the skill's Phase 5 (INSTRUMENT) closes the bootstrap gap. The remaining ~5% (multi-tool MCP wiring, fast CI/terminal access without npx cold-start) is what npm install -g fills. They coexist cleanly.
Adoption tracking and anonymous telemetry hooks ship in the standard hook profile, so they fire on plugin install with no extra setup. They default-enable but a privacy notice prints to stderr on first run. Opt out at any time:
# Per-shell
export DO_NOT_TRACK=1
# Or per-project, in harness.config.json:
# { "telemetry": { "enabled": false }, "adoption": { "enabled": false } }For identity-tagged telemetry (project/team/alias), run the interactive wizard once via npx @harness-engineering/cli telemetry-wizard — the plugin doesn't ship an interactive equivalent.
Plugin users have two update channels that can drift slightly:
- Bundled artifacts (skills, slash commands, subagents, hooks) ship from this git repo. Update via
/plugin update harness-claude. - MCP server binary is launched via
npx -y -p @harness-engineering/cli@<pinned-version> harness-mcp, where<pinned-version>is an exact version pinned in the plugin manifest (not@latest). Every adopter runs that exact published build, and updates arrive deliberately — a/plugin updatepulls a manifest whose pin has been bumped, rather than each new session silently pulling whatever is newest on npm. See docs/security/trust-model.md for the full trust and integrity model.
Both channels move together on a /plugin update: the git artifacts and the pinned MCP version are bumped in the same manifest revision. If you instead want to always track the newest publish yourself, npm install -g @harness-engineering/cli and use harness setup — the global install is the opt-in "latest" path.
If you only have the plugin installed and need a shell-level harness command, npx works without a global install:
npx @harness-engineering/cli validate
npx @harness-engineering/cli check-deps
npx @harness-engineering/cli check-archFirst call is slow (npx fetches the package); subsequent calls within the cache window are fast. For frequent terminal use, npm install -g is still the better path.
In an AI agent session (Claude Code, Gemini CLI):
/harness:initialize-project
The initialization skill walks you through project setup interactively — name, adoption level, framework overlay — and scaffolds everything including MCP server configuration.
CLI alternative (for scripts or CI):
harness init --name my-project --level intermediate
/harness:verify
Runs all mechanical checks in one pass — configuration, dependency boundaries, lint, typecheck, and tests.
CLI alternative:
harness validate && harness check-deps
git clone https://github.com/Intense-Visions/harness-engineering.git
cd harness-engineering/examples/hello-world
npm install && harness validate| Package | Description |
|---|---|
@harness-engineering/types |
Shared TypeScript types and interfaces |
@harness-engineering/core |
Validation, constraints, entropy detection, state management |
@harness-engineering/cli |
CLI: validate, check-deps, skill run, state show |
@harness-engineering/eslint-plugin |
12 rules: layer violations, circular deps, forbidden imports, boundary schemas, doc exports, no nested loops in critical paths, no sync IO in async, no unbounded array chains, no unix shell commands, no hardcoded path separators, require path normalization, no process env in spawn |
@harness-engineering/linter-gen |
Generate custom ESLint rules from YAML configuration |
@harness-engineering/graph |
Knowledge graph for codebase relationships and entropy detection |
@harness-engineering/intelligence |
Intelligence pipeline for spec enrichment, complexity modeling, and pre-execution simulation |
@harness-engineering/orchestrator |
Agent orchestration daemon for dispatching coding agents to issues |
@harness-engineering/dashboard |
Local web dashboard for project health and roadmap visualization |
import { validateFileStructure } from '@harness-engineering/core';
const result = await validateFileStructure('/path/to/project');
if (!result.ok) {
console.error('Validation failed:', result.error.message);
process.exit(1);
}# CLI — validate project constraints
harness validate
# Check architectural dependency boundaries
harness check-deps
# Run a skill
harness skill run harness-verificationSee Getting Started for a full walkthrough.
Harness enforces a strict layered dependency model. Each layer may only import from layers below it.
graph TD
A[agents] --> S[services]
S --> R[repository]
R --> C[config]
C --> T[types]
style A fill:#4a90d9,stroke:#2c5f8a,color:#fff
style S fill:#50b86c,stroke:#2d7a3e,color:#fff
style R fill:#f5a623,stroke:#c17d12,color:#fff
style C fill:#9b59b6,stroke:#6c3483,color:#fff
style T fill:#7f8c8d,stroke:#566573,color:#fff
Violations are caught at lint time via @harness-engineering/eslint-plugin — not at code review.
Install the CLI, MCP server, skills, and personas so they're available in every project:
npm install -g @harness-engineering/cli
harness setupThe single npm install -g provides both the harness CLI and the harness-mcp server binary, with all dependencies version-matched. harness setup then detects installed AI clients and writes to your global config directories:
| Platform | Slash Commands | Agent Definitions |
|---|---|---|
| Claude Code | ~/.claude/commands/ |
.claude/agents/ |
| Gemini CLI | ~/.gemini/commands/ |
.gemini/agents/ |
| Cursor | ~/.cursor/rules/ |
— |
| Codex CLI | ~/.codex/ |
— |
After this, /harness:* slash commands and harness agent personas are available in every conversation — no per-project install needed.
For real-time constraint validation, connect the MCP server to your project. The easiest way is during initialization:
/harness:initialize-project
This scaffolds your project and configures the MCP server automatically.
To add the MCP server to an existing project:
harness setup-mcpThis gives your AI agent access to 62 tools (validation, entropy detection, skill execution, state management, code review, graph queries, and more) and 9 resources (project context, skills catalog, rules, learnings, state, graph, entities, relationships, business-knowledge).
Manual MCP setup
Claude Code — add to .mcp.json in your project root:
{
"mcpServers": {
"harness": {
"command": "harness-mcp"
}
}
}Gemini CLI — add to .gemini/settings.json in your project root:
{
"mcpServers": {
"harness": {
"command": "harness-mcp"
}
}
}Then add your project directory to ~/.gemini/trustedFolders.json (Gemini ignores workspace MCP servers in untrusted folders):
{
"/path/to/your/project": "TRUST_FOLDER"
}Cursor — add to .cursor/mcp.json in your project root:
{
"mcpServers": {
"harness": {
"command": "harness",
"args": ["mcp"]
}
}
}Codex CLI — add to .codex/config.toml in your project root:
[mcp_servers.harness]
command = "harness"
args = ["mcp"]
enabled = trueOpenCode — add to opencode.json in your project root:
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"harness": {
"type": "local",
"command": ["harness", "mcp"],
"enabled": true
}
}
}Note:
harness-mcpis installed alongside the CLI bynpm install -g @harness-engineering/cli. Using the installed binary instead ofnpx @harness-engineering/mcp-serveravoids stale npx cache issues and ensures the MCP server uses the same package versions as the CLI.
| Client | MCP Config Location | Additional Setup |
|---|---|---|
| Claude Code | .mcp.json |
None |
| Gemini CLI | .gemini/settings.json |
Add project to ~/.gemini/trustedFolders.json |
| Cursor | .cursor/mcp.json |
None |
| Codex CLI | .codex/config.toml |
None |
| OpenCode | opencode.json |
None — skills auto-discovered from .claude/skills/ |
| Component | Count | Description |
|---|---|---|
| Packages | 9 | Core library, CLI, ESLint plugin, linter generator, graph, intelligence, orchestrator, dashboard, shared types |
| Skills | 773 | Agent workflows. 12 are load-bearing gear (Tier-0) — see below; the rest are an on-demand library |
| Personas | 12 | Architecture enforcer, code reviewer, planner, verifier, task executor, and 7 more |
| Templates | 19 | Language bases, framework overlays (Express, NestJS, Django, FastAPI, Gin, Axum, Spring Boot, and more) |
| Examples | 3 | Progressive tutorials from 5 minutes to 30 minutes |
The catalog runs to hundreds of skills, but a senior engineer only needs to hold a dozen in their head. These twelve carry the core workflow end to end — learn them first; everything else is a library you reach for on demand. Skills declare their standing with a first-class catalog_tier field in skill.yaml (0 = load-bearing, 1 = library, 2 = retire candidate), surfaced here, in the Skills Catalog, and in the dashboard command palette.
| Skill | Slash command | What it carries |
|---|---|---|
harness-initialize-project |
/harness:initialize-project |
Scaffold or migrate a harness-managed project |
harness-strategy |
/harness:strategy |
Set the durable product anchor (STRATEGY.md) |
harness-brainstorming |
/harness:brainstorming |
Turn intent into a spec |
harness-planning |
/harness:planning |
Decompose a spec into an ordered plan |
harness-execution |
/harness:execution |
Implement a plan task-by-task with state tracking |
harness-tdd |
/harness:tdd |
Test-driven development inside the loop |
harness-verification |
/harness:verification |
Verify built artifacts against spec and plan |
harness-code-review |
/harness:code-review |
Multi-persona review pipeline |
outcome-eval |
/harness:outcome-eval |
Ship gate — did the change satisfy its spec? |
harness-debugging |
/harness:debugging |
Systematic debugging with validation |
harness-autopilot |
/harness:autopilot |
Autonomous phase loop — plan → execute → verify → review |
harness-roadmap-pilot |
/harness:roadmap-pilot |
Pick the next highest-impact roadmap item and drive it |
Learn by doing. Each example builds on the previous:
| Example | Level | Time | What You Learn |
|---|---|---|---|
| Hello World | Basic | 5 min | Config, validation, AGENTS.md — see what a harness project looks like |
| Task API | Intermediate | 15 min | Express API with 3-layer architecture enforced by ESLint |
| Multi-Tenant API | Advanced | 30 min | Custom linter rules, Zod boundary validation, personas, full state lifecycle |
Getting Started
- Getting Started Guide — From zero to validated project
- Day-to-Day Workflow — Full lifecycle tutorial using slash commands
- Best Practices — Patterns for effective harness usage
- Agent Worktree Patterns — Running multiple agents in parallel
Core Concepts
- The Core Principles — Foundational concepts behind harness engineering
- Implementation Guide — Adoption levels and rollout strategy
- KPIs — Measuring agent effectiveness
Reference
- CLI Reference — All commands and flags (for CI/scripts)
- Configuration Reference —
harness.config.jsonschema
| Project | Key Contribution |
|---|---|
| GitHub Spec Kit | Constitution/principles, cross-artifact validation |
| BMAD Method | Scale-adaptive intelligence, workflow re-entry, party mode |
| GSD | Goal-backward verification, persistent state, codebase mapping |
| Superpowers | Rigid behavioral workflows, subagent dispatch, verification discipline |
| Ralph Loop | Fresh-context iteration, append-only learnings, task sizing |
These five projects most directly shaped harness engineering. See the full Inspirations & Acknowledgments for all 50 projects, standards, and tools analyzed — what we adopted, what we skipped, and why.
See CONTRIBUTING.md for development setup, coding standards, and pull request guidelines.
MIT License — see LICENSE for details.