You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 077803a
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: apps/sim/lib/benchmarks/README.md
+4-4Lines changed: 4 additions & 4 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -10,10 +10,10 @@ Open **Settings → Admin → Open benchmarks** (or `/benchmark`). Select an org
10
10
11
11
Each case fixes the execution user when it is created. Changing the selector shows that user's saved cases; it never changes an existing case's identity. Legacy cases retain their original owner's execution identity. **New benchmark** always opens a creation form, and saved cases appear in a list with an **Open** action.
12
12
13
-
1. Select a workspace. Generate the reference from its active workflow exports, or paste/import a reference produced by another client. Review the original task brief and reference.
13
+
1. Select a workspace. Generate the reference by letting Mothership inspect its workflows and supporting resources, or paste/import a reference produced by another client. Review the original task brief and reference.
14
14
2. Generate the blanks and review their expected answers. Only select facts discoverable from the permitted enterprise sources. Replacing `[[BLANK:id]]` with each exact answer must restore the reference byte for byte. Remove answer hints left elsewhere in the reference.
15
15
3. Run the planner. It receives the brief in a fresh organization Plan conversation, with enterprise search/document reads and isolated memory. It receives neither the completed workspace inventory nor the reference, masks, or expected answers. Its final response is the generated spec.
16
-
4. Reconstruct and grade separately. Reconstruction is a fresh tool-free execution with only the generated specand redacted reference. Grading compares the recovered answers to the hidden targets and checks support in the generated spec. An exact supporting passage is required for credit. The score is correct details divided by the fixed number of blanks.
16
+
4. Reconstruct and grade separately. Reconstruction is a fresh execution with the redacted questions in context and one immutable document, `generated-spec.md`, available through a bounded read/search tool. That document contains only the planner’s generated spec; enterprise tools, other documents and history are unavailable. Grading compares the recovered answers to the hidden targets and checks support in the generated spec. An exact supporting passage is required for credit. The score is correct details divided by the fixed number of blanks.
17
17
18
18
The UI groups these actions into Reference, Plan, Reconstruction, and Grade. Each output is saved independently. Editing upstream inputs clears dependent current outputs. Every successful Grade automatically saves an immutable result in **Run history**, including the exact inputs, generated spec, answers, and assessments. Add an optional run label before grading to identify agent versions or experiments. History is retained when you edit the case or rerun a stage; deleting the benchmark deletes its history.
19
19
@@ -25,8 +25,8 @@ Reference specs can be prose or JSON text; **Import spec** accepts Markdown, tex
25
25
26
26
Every stage is limited to ten minutes, with a twelve-minute persistence lease. Concurrent starts or stale completions cannot overwrite newer results. Interrupted steps are retryable; a hard server failure becomes retryable when its lease expires. This first version runs a stage within its HTTP request, so the deployment's request timeout must accommodate the run. Page reload may interrupt a pending request; saved completed stages remain available.
27
27
28
-
Automatic distillation supports at most 50 workflows and 500 KB of sanitized exports. Credentials/passwords are removed by the canonical workflow exporter. Larger projects can use an imported reference. Stateless generation has a 16,384-token output cap. The planner uses the dev worker's model configuration; keep worker/model configuration constant when comparing agent changes.
28
+
Reference generation uses the existing Mothership agent loop and workspace CLI to inspect active workflows and supporting resources incrementally. There is no benchmark-wide workflow-count or export-size gate. Reads remain authorized as the selected user; the benchmark transport refuses source mutations, workflow execution, services and scratch writes. Existing CLI pagination and searchable stored outputs keep individual model inputs bounded. Specs can contain up to 1,000,000 characters; structured output uses the worker’s 32,768-token per-response limit. The planner uses the dev worker's model configuration; keep worker/model configuration constant when comparing agent changes.
29
29
30
30
This version uses **live enterprise context**, not a frozen historical snapshot. The completed Sim workspace is excluded from planning, but historical solution documents in connected enterprise sources can still reveal answers. Treat these cases as retrospective evaluations and review source availability. The score measures recovery of the selected requirements, not execution correctness or every claim in the plan.
31
31
32
-
Validation includes real PostgreSQL ownership/organization/access-revocation checks and concurrent stage/lease fencing in `repository.integration.ts`; redaction integrity and dependency invalidation in `artifacts.test.ts`; reconstruction evidence and denominator checks in `evaluation.test.ts`; and executable tool isolation in the Mothership tool-executor tests. The worker tests cover fresh conversations, memory isolation, restricted tools, restart continuation, and tool-free execution.
32
+
Validation includes real PostgreSQL ownership/organization/access-revocation checks and concurrent stage/lease fencing in `repository.integration.ts`; redaction integrity and dependency invalidation in `artifacts.test.ts`; reconstruction evidence and denominator checks in `evaluation.test.ts`; and executable tool isolation in the Mothership tool-executor tests. The worker tests cover fresh conversations, memory isolation, restricted tools, restart continuation, and isolated spec retrieval and tool-free grading.
'Describe the implemented business behavior of these workflow exports as a self-contained reference specification. Cover triggers, conditions, ownership, mappings, actions, destinations and failure behavior. Do not invent intent, requirements or missing values. Omit layout, internal block IDs, credential IDs and implementation trivia. Also draft a short taskBrief expressing the business goal a user would originally request, without disclosing the enterprise-specific answers. If a taskBrief was supplied, preserve it. The reference is for human review before evaluation.',
17
-
{workflows,taskBrief }
16
+
'Explore the selected workspace using the read-only workspace CLI and describe its implemented behavior as a detailed, self-contained reference specification. Begin by listing all active workflows, follow list pagination, inspect each workflow and its referenced workspace resources, and follow large-output continuations. Cover exact triggers and input shapes, conditions, ownership, field mappings, prompts and code behavior, actions, destinations, cross-workflow relationships, outputs and failure/recovery behavior. Preserve concrete names, values and business rules; do not compress them into a high-level overview. Distinguish implemented behavior from unresolved configuration or inferred intent. Do not invent missing values. Omit editor layout and credential values. Also draft a short taskBrief expressing the business goal a user would originally request, without disclosing the enterprise-specific answers. If a taskBrief was supplied, preserve it. The reference is for human review before evaluation. Use tools to inspect; return the final JSON only when inspection is complete.',
17
+
{ taskBrief }
18
18
)
19
19
}
20
20
@@ -26,13 +26,10 @@ export function redactionMessages(referenceSpec: string): ExecuteMessage[] {
26
26
}
27
27
28
28
/** This projection is the reader's complete input; reference answers and enterprise history have no path into it. */
'Fill every [[BLANK:id]] in redactedSpec using generatedSpec as your only evidence. For each distinct id return answer and support, where support is an exact contiguous quote from generatedSpec that establishes the answer. Use an empty answer and empty support when the generated spec does not establish it or contradicts itself. Do not use background knowledge, surviving reference text, or guesses to supply missing facts.',
35
-
{generatedSpec,redactedSpec }
31
+
'Fill every [[BLANK:id]] in redactedSpec using generated-spec.md as your only evidence. Read or search that file with read_spec. For each distinct id return answer and support, where support is an exact contiguous quote from the file that establishes the answer. Use an empty answer and empty support when the generated spec does not establish it or contradicts itself. Do not use background knowledge, surviving reference text, or guesses to supply missing facts.',
0 commit comments