Skip to content

Commit 077803a

Browse files
committed
Inspect benchmark workspaces incrementally and search generated specs
1 parent 7b6a509 commit 077803a

19 files changed

Lines changed: 291 additions & 97 deletions

File tree

‎apps/sim/app/o/[organizationId]/benchmark/components/benchmark-detail.tsx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -143,7 +143,7 @@ function BenchmarkEditor({
143143
<BenchmarkStep
144144
number={3}
145145
title='Fill in the blanks'
146-
description='A fresh reader receives only the generated plan and redacted reference, with no tools or memory.'
146+
description='A fresh reader fills the blanks by reading or searching the generated plan. It has no enterprise access or prior memory.'
147147
pending={stage === 'reconstruct'}
148148
action={
149149
<Chip

‎apps/sim/app/o/[organizationId]/benchmark/components/benchmark-reference.tsx‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -41,11 +41,11 @@ export function BenchmarkReference({
4141
if (!file) return
4242
try {
4343
if (file.size > BENCHMARK_SPEC_MAX_LENGTH * 4) {
44-
throw new Error('The reference is too large. Use a text file under 400 KB.')
44+
throw new Error('The reference is too large. Use a text file under 4 MB.')
4545
}
4646
const text = await file.text()
4747
if (text.length > BENCHMARK_SPEC_MAX_LENGTH) {
48-
throw new Error('The reference must be at most 100,000 characters.')
48+
throw new Error('The reference must be at most 1,000,000 characters.')
4949
}
5050
onChange({ referenceSpec: text })
5151
} catch (error) {

‎apps/sim/lib/benchmarks/README.md‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -10,10 +10,10 @@ Open **Settings → Admin → Open benchmarks** (or `/benchmark`). Select an org
1010

1111
Each case fixes the execution user when it is created. Changing the selector shows that user's saved cases; it never changes an existing case's identity. Legacy cases retain their original owner's execution identity. **New benchmark** always opens a creation form, and saved cases appear in a list with an **Open** action.
1212

13-
1. Select a workspace. Generate the reference from its active workflow exports, or paste/import a reference produced by another client. Review the original task brief and reference.
13+
1. Select a workspace. Generate the reference by letting Mothership inspect its workflows and supporting resources, or paste/import a reference produced by another client. Review the original task brief and reference.
1414
2. Generate the blanks and review their expected answers. Only select facts discoverable from the permitted enterprise sources. Replacing `[[BLANK:id]]` with each exact answer must restore the reference byte for byte. Remove answer hints left elsewhere in the reference.
1515
3. Run the planner. It receives the brief in a fresh organization Plan conversation, with enterprise search/document reads and isolated memory. It receives neither the completed workspace inventory nor the reference, masks, or expected answers. Its final response is the generated spec.
16-
4. Reconstruct and grade separately. Reconstruction is a fresh tool-free execution with only the generated spec and redacted reference. Grading compares the recovered answers to the hidden targets and checks support in the generated spec. An exact supporting passage is required for credit. The score is correct details divided by the fixed number of blanks.
16+
4. Reconstruct and grade separately. Reconstruction is a fresh execution with the redacted questions in context and one immutable document, `generated-spec.md`, available through a bounded read/search tool. That document contains only the planner’s generated spec; enterprise tools, other documents and history are unavailable. Grading compares the recovered answers to the hidden targets and checks support in the generated spec. An exact supporting passage is required for credit. The score is correct details divided by the fixed number of blanks.
1717

1818
The UI groups these actions into Reference, Plan, Reconstruction, and Grade. Each output is saved independently. Editing upstream inputs clears dependent current outputs. Every successful Grade automatically saves an immutable result in **Run history**, including the exact inputs, generated spec, answers, and assessments. Add an optional run label before grading to identify agent versions or experiments. History is retained when you edit the case or rerun a stage; deleting the benchmark deletes its history.
1919

@@ -25,8 +25,8 @@ Reference specs can be prose or JSON text; **Import spec** accepts Markdown, tex
2525

2626
Every stage is limited to ten minutes, with a twelve-minute persistence lease. Concurrent starts or stale completions cannot overwrite newer results. Interrupted steps are retryable; a hard server failure becomes retryable when its lease expires. This first version runs a stage within its HTTP request, so the deployment's request timeout must accommodate the run. Page reload may interrupt a pending request; saved completed stages remain available.
2727

28-
Automatic distillation supports at most 50 workflows and 500 KB of sanitized exports. Credentials/passwords are removed by the canonical workflow exporter. Larger projects can use an imported reference. Stateless generation has a 16,384-token output cap. The planner uses the dev worker's model configuration; keep worker/model configuration constant when comparing agent changes.
28+
Reference generation uses the existing Mothership agent loop and workspace CLI to inspect active workflows and supporting resources incrementally. There is no benchmark-wide workflow-count or export-size gate. Reads remain authorized as the selected user; the benchmark transport refuses source mutations, workflow execution, services and scratch writes. Existing CLI pagination and searchable stored outputs keep individual model inputs bounded. Specs can contain up to 1,000,000 characters; structured output uses the worker’s 32,768-token per-response limit. The planner uses the dev worker's model configuration; keep worker/model configuration constant when comparing agent changes.
2929

3030
This version uses **live enterprise context**, not a frozen historical snapshot. The completed Sim workspace is excluded from planning, but historical solution documents in connected enterprise sources can still reveal answers. Treat these cases as retrospective evaluations and review source availability. The score measures recovery of the selected requirements, not execution correctness or every claim in the plan.
3131

32-
Validation includes real PostgreSQL ownership/organization/access-revocation checks and concurrent stage/lease fencing in `repository.integration.ts`; redaction integrity and dependency invalidation in `artifacts.test.ts`; reconstruction evidence and denominator checks in `evaluation.test.ts`; and executable tool isolation in the Mothership tool-executor tests. The worker tests cover fresh conversations, memory isolation, restricted tools, restart continuation, and tool-free execution.
32+
Validation includes real PostgreSQL ownership/organization/access-revocation checks and concurrent stage/lease fencing in `repository.integration.ts`; redaction integrity and dependency invalidation in `artifacts.test.ts`; reconstruction evidence and denominator checks in `evaluation.test.ts`; and executable tool isolation in the Mothership tool-executor tests. The worker tests cover fresh conversations, memory isolation, restricted tools, restart continuation, and isolated spec retrieval and tool-free grading.

‎apps/sim/lib/benchmarks/application/operations.ts‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,12 @@ import { defineOperation } from '@/lib/core/application/operation'
22
import { defineWorkspaceOperation } from '@/lib/core/application/workspace-operation'
33

44
export const benchmarkOperations = {
5+
/** permission-group-exempt: platform superusers administer benchmarks; source access is checked for the selected user. */
6+
prepareReference: defineOperation({
7+
id: 'benchmarks.reference.prepare',
8+
principalKinds: ['session'],
9+
capability: 'none',
10+
}),
511
/** permission-group-exempt: platform superusers administer benchmarks; selected-user capabilities are checked separately. */
612
preparePlan: defineOperation({
713
id: 'benchmarks.plan.prepare',
Lines changed: 38 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,38 @@
1+
import type { Principal } from '@sim/auth/principal'
2+
import { db } from '@sim/db'
3+
import { copilotChats } from '@sim/db/schema'
4+
import { defineAuthorizedBenchmarkUseCase } from '@/lib/benchmarks/application/access'
5+
import { requireBenchmarkCaseAccess } from '@/lib/benchmarks/application/cases'
6+
import { benchmarkOperations } from '@/lib/benchmarks/application/operations'
7+
import { MOTHERSHIP_CHAT_DEFAULT_MODEL } from '@/lib/mothership/constants'
8+
9+
/** Source reads use an owned workspace chat, separate from the planner's enterprise discovery context. */
10+
export const prepareBenchmarkReference = defineAuthorizedBenchmarkUseCase({
11+
operation: benchmarkOperations.prepareReference,
12+
async execute({
13+
principal,
14+
input,
15+
}: {
16+
principal: Principal
17+
input: { organizationId: string; benchmarkId: string }
18+
}) {
19+
const benchmark = await requireBenchmarkCaseAccess(principal, input)
20+
const userId = benchmark.runAsUserId ?? benchmark.userId
21+
const [chat] = await db
22+
.insert(copilotChats)
23+
.values({
24+
userId,
25+
workspaceId: benchmark.sourceWorkspaceId,
26+
type: 'mothership',
27+
model: MOTHERSHIP_CHAT_DEFAULT_MODEL,
28+
config: {
29+
conversationMode: 'agent',
30+
benchmark: { id: benchmark.id, operatorUserId: benchmark.userId },
31+
},
32+
lastSeenAt: new Date(),
33+
})
34+
.returning({ id: copilotChats.id })
35+
if (!chat) throw new Error('Failed to create benchmark reference conversation')
36+
return { chatId: chat.id, userId }
37+
},
38+
})

‎apps/sim/lib/benchmarks/application/run-stage.ts‎

Lines changed: 10 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -2,13 +2,10 @@ import type { Principal } from '@sim/auth/principal'
22
import { createLogger } from '@sim/logger'
33
import { generateId } from '@sim/utils/id'
44
import { z } from 'zod'
5-
import {
6-
benchmarkSourcePrincipal,
7-
defineAuthorizedBenchmarkUseCase,
8-
} from '@/lib/benchmarks/application/access'
5+
import { defineAuthorizedBenchmarkUseCase } from '@/lib/benchmarks/application/access'
96
import { requireBenchmarkCaseAccess } from '@/lib/benchmarks/application/cases'
107
import { benchmarkOperations } from '@/lib/benchmarks/application/operations'
11-
import { readBenchmarkWorkspace } from '@/lib/benchmarks/application/source'
8+
import { prepareBenchmarkReference } from '@/lib/benchmarks/application/prepare-reference'
129
import { applyBenchmarkPatch, validateBenchmarkRedaction } from '@/lib/benchmarks/artifacts'
1310
import { getBenchmarkMothershipUrl } from '@/lib/benchmarks/config'
1411
import { gradeReconstruction, validateReconstruction } from '@/lib/benchmarks/evaluation'
@@ -92,18 +89,18 @@ async function performStage(
9289
const artifacts = benchmark.artifacts
9390
switch (stage) {
9491
case 'distill': {
95-
const sourcePrincipal = await benchmarkSourcePrincipal(principal, {
96-
organizationId: benchmark.organizationId,
97-
runAsUserId: benchmark.runAsUserId ?? benchmark.userId,
98-
sourceWorkspaceId: benchmark.sourceWorkspaceId,
92+
const target = await prepareBenchmarkReference.execute({
93+
principal,
94+
input: { organizationId: benchmark.organizationId, benchmarkId: benchmark.id },
9995
})
100-
const source = await readBenchmarkWorkspace(sourcePrincipal, benchmark.sourceWorkspaceId)
10196
signal.throwIfAborted()
10297
const result = await executeBenchmarkJson({
10398
benchmark,
10499
signal,
105100
schema: distillationSchema,
106-
messages: distillationMessages(source, artifacts.taskBrief),
101+
messages: distillationMessages(artifacts.taskBrief),
102+
profile: { stage: 'distill' },
103+
chatId: target.chatId,
107104
})
108105
return { artifacts: applyBenchmarkPatch(artifacts, result), plannerChatId: null }
109106
}
@@ -133,7 +130,8 @@ async function performStage(
133130
benchmark,
134131
signal,
135132
schema: reconstructionSchema,
136-
messages: reconstructionMessages(artifacts.generatedSpec!, artifacts.redactedSpec),
133+
messages: reconstructionMessages(artifacts.redactedSpec),
134+
profile: { stage: 'reconstruct', spec: artifacts.generatedSpec! },
137135
})
138136
validateReconstruction(artifacts.blanks, result.answers)
139137
return { artifacts: { ...artifacts, reconstruction: result.answers, grade: null } }

‎apps/sim/lib/benchmarks/application/source.ts‎

Lines changed: 0 additions & 51 deletions
This file was deleted.

‎apps/sim/lib/benchmarks/prompts.ts‎

Lines changed: 6 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -11,10 +11,10 @@ function messages(instruction: string, data: unknown): ExecuteMessage[] {
1111
]
1212
}
1313

14-
export function distillationMessages(workflows: string, taskBrief: string): ExecuteMessage[] {
14+
export function distillationMessages(taskBrief: string): ExecuteMessage[] {
1515
return messages(
16-
'Describe the implemented business behavior of these workflow exports as a self-contained reference specification. Cover triggers, conditions, ownership, mappings, actions, destinations and failure behavior. Do not invent intent, requirements or missing values. Omit layout, internal block IDs, credential IDs and implementation trivia. Also draft a short taskBrief expressing the business goal a user would originally request, without disclosing the enterprise-specific answers. If a taskBrief was supplied, preserve it. The reference is for human review before evaluation.',
17-
{ workflows, taskBrief }
16+
'Explore the selected workspace using the read-only workspace CLI and describe its implemented behavior as a detailed, self-contained reference specification. Begin by listing all active workflows, follow list pagination, inspect each workflow and its referenced workspace resources, and follow large-output continuations. Cover exact triggers and input shapes, conditions, ownership, field mappings, prompts and code behavior, actions, destinations, cross-workflow relationships, outputs and failure/recovery behavior. Preserve concrete names, values and business rules; do not compress them into a high-level overview. Distinguish implemented behavior from unresolved configuration or inferred intent. Do not invent missing values. Omit editor layout and credential values. Also draft a short taskBrief expressing the business goal a user would originally request, without disclosing the enterprise-specific answers. If a taskBrief was supplied, preserve it. The reference is for human review before evaluation. Use tools to inspect; return the final JSON only when inspection is complete.',
17+
{ taskBrief }
1818
)
1919
}
2020

@@ -26,13 +26,10 @@ export function redactionMessages(referenceSpec: string): ExecuteMessage[] {
2626
}
2727

2828
/** This projection is the reader's complete input; reference answers and enterprise history have no path into it. */
29-
export function reconstructionMessages(
30-
generatedSpec: string,
31-
redactedSpec: string
32-
): ExecuteMessage[] {
29+
export function reconstructionMessages(redactedSpec: string): ExecuteMessage[] {
3330
return messages(
34-
'Fill every [[BLANK:id]] in redactedSpec using generatedSpec as your only evidence. For each distinct id return answer and support, where support is an exact contiguous quote from generatedSpec that establishes the answer. Use an empty answer and empty support when the generated spec does not establish it or contradicts itself. Do not use background knowledge, surviving reference text, or guesses to supply missing facts.',
35-
{ generatedSpec, redactedSpec }
31+
'Fill every [[BLANK:id]] in redactedSpec using generated-spec.md as your only evidence. Read or search that file with read_spec. For each distinct id return answer and support, where support is an exact contiguous quote from the file that establishes the answer. Use an empty answer and empty support when the generated spec does not establish it or contradicts itself. Do not use background knowledge, surviving reference text, or guesses to supply missing facts.',
32+
{ redactedSpec }
3633
)
3734
}
3835

‎apps/sim/lib/benchmarks/repository.integration.ts‎

Lines changed: 23 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -22,6 +22,7 @@ vi.mock('@sim/db', () => ({
2222

2323
import { createBenchmark, getBenchmark, listBenchmarks } from '@/lib/benchmarks/application/cases'
2424
import { prepareBenchmarkPlan } from '@/lib/benchmarks/application/prepare-plan'
25+
import { prepareBenchmarkReference } from '@/lib/benchmarks/application/prepare-reference'
2526
import {
2627
getBenchmarkRun,
2728
listBenchmarkRuns,
@@ -231,6 +232,21 @@ describe('private benchmark persistence and attempt fencing', () => {
231232
},
232233
})
233234
expect(benchmark).toMatchObject({ userId: 'owner', runAsUserId: 'target' })
235+
const reference = await prepareBenchmarkReference.execute({
236+
principal,
237+
input: { organizationId: 'org', benchmarkId: benchmark.id },
238+
})
239+
const [referenceChat] =
240+
await connection`SELECT user_id, workspace_id, organization_id, config FROM copilot_chats WHERE id = ${reference.chatId}`
241+
expect(referenceChat).toMatchObject({
242+
user_id: 'target',
243+
workspace_id: 'workspace',
244+
organization_id: null,
245+
config: {
246+
conversationMode: 'agent',
247+
benchmark: { id: benchmark.id, operatorUserId: 'owner' },
248+
},
249+
})
234250
const target = await prepareBenchmarkPlan.execute({
235251
principal,
236252
input: { organizationId: 'org', benchmarkId: benchmark.id },
@@ -269,7 +285,13 @@ describe('private benchmark persistence and attempt fencing', () => {
269285
input: { organizationId: 'org', benchmarkId: benchmark.id },
270286
})
271287
).rejects.toMatchObject({ code: 'forbidden' })
272-
expect(await connection`SELECT id FROM copilot_chats`).toHaveLength(2)
288+
await expect(
289+
prepareBenchmarkReference.execute({
290+
principal,
291+
input: { organizationId: 'org', benchmarkId: benchmark.id },
292+
})
293+
).rejects.toMatchObject({ code: 'forbidden' })
294+
expect(await connection`SELECT id FROM copilot_chats`).toHaveLength(3)
273295
})
274296

275297
it('refuses a target outside the organization or with revoked membership, disabled account, or workspace access', async () => {

‎apps/sim/lib/benchmarks/types.ts‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
import { z } from 'zod'
22

3-
export const BENCHMARK_SPEC_MAX_LENGTH = 100_000
3+
export const BENCHMARK_SPEC_MAX_LENGTH = 1_000_000
44
export const BENCHMARK_MAX_BLANKS = 50
55
export const benchmarkStageSchema = z.enum(['distill', 'redact', 'plan', 'reconstruct', 'grade'])
66
export const benchmarkNameSchema = z.string().trim().min(1, 'A benchmark name is required').max(200)

0 commit comments

Comments
 (0)