You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 9d74b8d
Browse filesBrowse the repository at this point in the historyBrowse files
'Use up to 50 blank IDs (letters, numbers, underscores or hyphens, at most 64 characters), each mapped to a nonempty string of at most 10,000 characters.'
54
+
'Use blank IDs with letters, numbers, underscores or hyphens (at most 64 characters), each mapped to a nonempty answer string.'
Copy file name to clipboardExpand all lines: apps/sim/lib/benchmarks/README.md
+2-2Lines changed: 2 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -23,9 +23,9 @@ Grading opens the saved report for human review. On any detail, choose **Mark as
23
23
24
24
Reference specs can be prose or JSON text; **Import spec** accepts Markdown, text, and JSON files without reformatting their contents. **Generate blanks with AI** creates the redacted reference and expected answers together. **Edit JSON** also lets you paste or edit a mapping such as `{"queue":"Customer Escalations","handoff":"Engineering explicitly accepts the case"}` alongside the matching redacted reference. Keys are blank IDs and values are exact original passages as strings. Applying validates that the mapping restores the reference exactly and updates the draft; **Save changes** persists it.
25
25
26
-
Every stage is limited to ten minutes, with a twelve-minute persistence lease. Concurrent starts or stale completions cannot overwrite newer results. Interrupted steps are retryable; a hard server failure becomes retryable when its lease expires. This first version runs a stage within its HTTP request, so the deployment's request timeout must accommodate the run. Page reload may interrupt a pending request; saved completed stages remain available.
26
+
Benchmark stages have no benchmark-specific duration cutoff. The active process renews its persistence lease every 30 seconds; a two-minute lease detects an abandoned process rather than limiting a healthy run. Concurrent starts or stale completions cannot overwrite newer results. Interrupted steps are retryable; a hard server failure becomes retryable when its lease expires. This version runs a stage within its HTTP request, so the deploymentmust support long-lived requests. Page reload may interrupt a pending request; saved completed stages remain available.
27
27
28
-
Reference generation uses the existing Mothership agent loop and workspace CLI to inspect active workflows and supporting resources incrementally. There is no benchmark-wide workflow-count or export-size gate. Reads remain authorized as the selected user; the benchmark transport refuses source mutations, workflow execution, services and scratch writes. Existing CLI pagination and searchable stored outputs keep individual model inputs bounded. Specs can contain up to 1,000,000 characters; structured output uses the worker’s 32,768-token per-response limit. The planner uses the dev worker's model configuration; keep worker/model configuration constant when comparing agent changes.
28
+
Reference generation uses the existing Mothership agent loop and workspace CLI to inspect active workflows and supporting resources incrementally. There is no benchmark-wide workflow-count or export-size gate. Reads remain authorized as the selected user; the benchmark transport refuses source mutations, workflow execution, services and scratch writes. Existing CLI pagination and searchable stored outputs keep individual model inputs bounded. Benchmark-specific caps on spec length, blank count, answer length and model output are removed. The normal model/provider and transport constraints still apply; page sizes bound individual tool reads without limiting the complete workspace or document. The planner uses the dev worker's model configuration; keep worker/model configuration constant when comparing agent changes.
29
29
30
30
This version uses **live enterprise context**, not a frozen historical snapshot. The completed Sim workspace is excluded from planning, but historical solution documents in connected enterprise sources can still reveal answers. Treat these cases as retrospective evaluations and review source availability. The score measures recovery of the selected requirements, not execution correctness or every claim in the plan.
0 commit comments