🎉 100 assessments GET 🎉 #2
Opensiro Admin
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
First 100 Harness Assessments Complete
Milestone timestamp: 2026-09-19 01:21:11 EEST
UTC: 2026-09-18 22:21:11Z
The VSM Harness Index has reached its first 100 completed standalone harness assessments.
The 100th completed assessment was Henterprise, admitted through PR #139
Some of the most interesting organizational forms so far
These are not “the best harnesses.” They are examples that expose particularly useful organizational distinctions.
Headcount — almost the entire autonomous VSM surface
Current vector:
A A A C A AHeadcount is one of the clearest examples in the corpus of a harness extending beyond task execution into a broad organizational system: operations, coordination, current control, adaptation, and identity/policy are agent-owned.
Its remaining distinction is S3*: reviewer-class mechanisms exist, but the reviewed boundary exposes the independent audit path as a constructor surface rather than a fully closed autonomous reviewer loop.
This is currently the closest observed structure to autonomous coverage across all six indexed functions.
Henterprise — the constructor-heavy counterpart
Vector:
A C C C C CHenterprise is interesting for almost the opposite reason.
The executing Hermes agent closes S1, while explicit organizational machinery exists for coordination, portfolio/current control, independent review, future strategy, and executive policy. But the durable cross-profile transport, shared state, invocation, and enforcement loops are left for composition.
That makes Henterprise an unusually clean example of a Constructor harness: organizational functions are present, but their autonomous owners are intentionally not fully wired by the repository itself.
It is also a useful demonstration of why assessments are repository-relative. Organizational semantics can survive a migration while ownership and closure change.
Omnigent — an autonomous S1 → S3* organization without S4/S5
Vector:
A A A A — —Omnigent closes operations, inter-worker coordination, whole-team current regulation, and independent review.
It is a compact example of an agent organization that has a substantial inside-and-now metasystem but does not thereby acquire strategy or ultimate policy.
Incidentally, Omnigent is
catalog_position=100, while Henterprise became the 100th completed assessment. That difference illustrates why catalog ordering and assessment state remain separate.Paperclip, Squad, and LoopX — autonomous organization with parent authority
Each exposes the pattern:
A A A(P) A — PThese systems are useful because they show that human or parent authority does not need to be represented as “less autonomy.”
The autonomous organization can still own operations, coordination, current regulation, and independent audit while a legitimate parent retains an additional current-control mode and ultimate organizational authority.
This is exactly the kind of case that motivated keeping
A,C, andPas ownership arrangements rather than treating them as a maturity ladder.DeepSeek Harness — organization emerging from an optional team mode
Vector:
A A A(P) — — —The base agent loop supplies autonomous operation, while Agent Teams introduces durable shared commitments and model-owned team regulation.
A separate parent-governed task-board path adds an S3 parent mode.
It is a useful example of why the Index assesses actual operating modes rather than inferring organization from the existence of “agents,” “teams,” or a manager-labelled component.
Independent audit is becoming a distinct design family
Several otherwise very different harnesses converge on:
A — — A — —Examples include PenguinHarness, Proliferate, ClawGUI, Harness Evolver, and GSD.
Their implementations differ, but the recurring organizational pattern is notable: an autonomous operating process plus a genuinely separate reviewer, critic, evaluator, or fresh-context inspection path.
This is stronger than ordinary self-reflection. The important property is an independent evidence path capable of returning corrective information to the operating loop.
Self-improvement does not automatically mean S4
Meta-Harness, AutoAgent, and related systems are useful counterexamples.
They can autonomously inspect benchmark results, mutate a harness, evaluate the mutation, and retain or revert changes.
That can be sophisticated autonomous work while still remaining S1.
S4 requires a different organizational relation: outside/future distinctions must produce adaptation options and return them into current organizational capability. “The system improves itself” is not sufficient evidence by itself.
Symphony — orchestration is not automatically a metasystem
Symphony provides issue claiming, bounded concurrency, isolated workspaces, retries, reconciliation, and scheduling around autonomous coding agents.
Yet the reviewed organizational vector remains:
A — — — — —Why? Because deterministic scheduling and runtime machinery are not automatically agent-owned S2 or S3.
This has become one of the most useful recurring distinctions in the Index: orchestration infrastructure and organizational decision ownership are different things.
A useful negative result
The generated Full-A view is still empty.
Across the first 100 completed assessments, we have not yet found a reviewed harness with autonomous ownership of all six functions:
S1 + S2 + S3 + S3* + S4 + S5That is not a product-quality judgment, and it does not imply that every harness should try to reach such a configuration.
It does suggest something interesting about the current agent-harness ecosystem: autonomous execution is common, coordination and current regulation are increasingly visible, independent audit is emerging as a recognizable architecture, while explicit adaptation and ultimate-policy closure remain much less common.
What the first 100 changed
The corpus increasingly supports the original normalization thesis:
“Manager,” “planner,” “reviewer,” “memory,” “team,” “swarm,” “learning,” and “delegation” have repeatedly turned out to be insufficient classifications on their own.
The useful comparison only appears after asking:
That is the role of VSM in the Index: not to make projects adopt cybernetics terminology, but to provide a common functional coordinate system for comparing very different agent architectures.
Next
100 assessments is a milestone, not a stopping point.
The next phase is not just a larger corpus. It is better longitudinal evidence: reassessments as upstream projects evolve, clearer domain-specific downstream views, more contributor-driven reviews, and more evidence about which organizational forms recur across otherwise unrelated harness ecosystems.
Thanks to everyone building these systems in public.
All reactions