From db77ffe9419157b307626e46607d326f2b96a248 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:18:39 +0200 Subject: [PATCH 01/21] docs: define Tutor Quality Stage 2A evaluation contract --- ...function_definition_TutorQualityStage2A.md | 215 ++++++++++++++++++ 1 file changed, 215 insertions(+) create mode 100644 ssot/ssot_function_definition_TutorQualityStage2A.md diff --git a/ssot/ssot_function_definition_TutorQualityStage2A.md b/ssot/ssot_function_definition_TutorQualityStage2A.md new file mode 100644 index 00000000..5ae8062b --- /dev/null +++ b/ssot/ssot_function_definition_TutorQualityStage2A.md @@ -0,0 +1,215 @@ +# Tutor Quality – Stage 2A: Real-Provider Evaluation Foundation + +Status: normative for the Stage-2A evaluation runner and its artifacts. This +SSOT does not change the Stage-1 PR hard gate and does not claim to measure +learning effect. + +## 1. Purpose and boundary + +Stage 2A makes real Tutor responses observable, reproducible enough for +comparison, and available for later human or calibrated-judge review. Its +measurement target is **evaluation integrity**: + +- the same versioned scenarios can be run again; +- every sample identifies the exact code, Course Content, prompt, provider, + model, and relevant parameters used; +- the normal `TutorService` and provider trust boundary remains in force; +- deterministic Stage-1 invariants are applied to real responses; +- technical failures are separated from Tutor-quality invariant violations; +- complete, secret-free transcripts can be inspected later. + +Stage 2A does **not** measure whether a learner actually learned, assign a +semantic quality score, optimize a strategy, or make a Tutor decision. Those +activities are Stage 2B or later. + +## 2. Relationship to Stage 1 + +Stage 1 remains the deterministic PR hard-gate layer. Stage 2A is a separate, +explicitly invoked observation layer. Real-provider calls MUST NOT be added to +normal unit tests, pull-request gates, or other required CI checks. + +The existing trust boundaries remain authoritative: + +- `TutorService` owns validation, planning application, repair, and state + commit; +- `CurriculumTutorAdapter` owns Course Content matching and progression; +- `KiconnectProvider` owns the OpenAI-compatible provider transport and parser; +- the evaluation runner owns only scenario orchestration, metadata, artifact + writing, and aggregation; +- provider output remains untrusted content, never application-owned planning + metadata. + +The runner MUST call the normal `TutorService`/`LLMProvider` path. It MUST NOT +call a provider HTTP endpoint directly or duplicate TutorService validation. + +## 3. Versioned corpus and scenario contract + +The anchor corpus is repository-owned and versioned. A corpus manifest has a +stable `corpusId`, an integer `corpusVersion`, and stable scenario IDs. A +scenario contains only synthetic learner data and references repository-owned +sketch/Course Content fixtures. A scenario MAY contain multiple sequential +Tutor turns; each sample starts from a fresh clone of the declared initial +history and progression state. + +Minimum scenario fields: + +- `id` and corpus version; +- sketch fixture/reference and its digest; +- Course Content revision/reference or explicit free-Tutor mode; +- initial history and progression state, when applicable; +- deterministic difficulty and simulated learner answers for each turn; +- expected structural observations only, such as phase/topic/state policy; +- no credential, personal identifier, or live learner data. + +The initial Stage-2A anchor set is deliberately small: + +1. `TQ-REG-001` PWM: strong answer and no semantic question repetition; +2. simple variable question; +3. Serial-output prediction; +4. incorrect answer and remediation; +5. partially correct answer and focused follow-up; +6. strong answer and progression; +7. unmatched Topic / free Tutor; +8. LEARN to DEEPEN; +9. EXPAND; +10. clearly off-topic learner answer. + +The corpus may grow only by adding a reviewed scenario or a corpus-version +change. A scenario edit is an evaluation change, not a hidden implementation +change. + +## 4. Run identity and metadata + +Each run has both a unique `runId` and a stable `evaluationIdentity`. + +`evaluationIdentity` is the SHA-256 of canonical JSON containing at least: + +- UnoSim Git SHA; +- Course Content revision(s); +- corpus ID and version; +- provider ID; +- requested model ID; +- prompt revision; +- relevant provider parameters, including timeout, temperature when known, + difficulty, sample count, and call budget. + +The `runId` additionally identifies this concrete invocation and MUST include a +UTC start time and collision-resistant suffix. The identity and run ID are +stored in every transcript and the aggregate report. + +The runner MUST record, without secrets: + +- provider ID and endpoint origin when useful for diagnosis; +- requested model and returned model; +- prompt revision; +- Course Content revision; +- scenario/corpus version; +- Git SHA and dirty-state indicator; +- sample index, turn index, timestamps, duration, and call counts; +- timeout, temperature, difficulty, and configured call/sample limits. + +Missing or inconsistent identity metadata, including an omitted or `auto` +model, makes a sample `invalid`. It MUST NOT be reported as a Tutor-quality +failure. If the provider returns a model different from the requested fixed +model, the sample is also `invalid` and no quality conclusion is drawn. + +## 5. Provider and credential rules + +The first implementation supports the existing Kiconnect/OpenAI-compatible +provider through `KiconnectProvider` and a fixed model ID supplied by the +scenario invocation. A runner preflight MUST verify that the requested model +is available. Any returned-model mismatch invalidates the sample rather than +silently accepting `auto` fallback. + +Credentials are read only from an invocation environment variable. They MUST +never appear in a transcript, aggregate report, log line, exception message, +Git diff, or uploaded artifact. Authorization headers are provider-internal +and are never part of the evaluation artifact model. + +Missing credentials cause an explicit `not-run/missing-credential` result for +manual/workflow execution. They do not fail normal CI because Stage 2A is not +part of normal CI. + +## 6. Transcript artifact contract + +One JSON transcript is written per scenario sample. It contains: + +- schema version, run ID, evaluation identity, and complete metadata; +- scenario inputs and synthetic learner answers; +- ordered logical Tutor turns; +- the Tutor request context needed for later review, excluding credentials and + transport headers; +- raw provider result as returned to the service, normalized final Tutor + result, returned model, or technical error; +- deterministic check records and state-before/state-after snapshots; +- a terminal sample status. + +Transcript status is one of: + +- `completed`: all requested turns executed and metadata is valid; +- `invalid`: identity/model/contract metadata is missing or inconsistent; +- `technical-failure`: provider error, timeout, malformed provider response, or + call-budget exhaustion prevented a turn; +- `deterministic-violation`: a Stage-1 invariant was observed on raw/final + output. A repaired raw response may be both technically completed and have a + violation record; the report MUST preserve that distinction. + +The artifact writer uses an allow-list of fields. It MUST NOT serialize the +credential environment, `process.env`, Authorization headers, or arbitrary +provider response envelopes. + +## 7. Deterministic checks and aggregation + +Stage 2A may report only properties with a deterministic rule. The runner +reuses the existing TutorService validation and bounded repeat heuristic and +records at least: + +- provider response/schema validity; +- exactly one primary question; +- complete-solution rejection; +- raw exact or heuristic question repetition; +- application-owned State-/Topic-/Phase-/revision consistency; +- no forbidden Question-ID reuse when Course Content is active; +- state unchanged after provider/technical failure; +- requested/returned model consistency; +- provider error category and timeout category; +- successful scenario/turn execution; +- provider-call count and budget exhaustion. + +The aggregate report contains counts and rates with explicit denominators per +scenario and overall. It MUST keep these categories separate: + +1. `invalid` evaluation metadata; +2. technical provider failures; +3. deterministic Tutor-quality invariant violations; +4. completed observations. + +There are no semantic quality grades, learning-support scores, or pass/fail +claims about actual learning in Stage 2A. + +Every invocation has a hard maximum sample count and provider-call budget. +The runner stops before issuing a call that would exceed the budget. Provider +token/cost data is reported only when the provider supplies it; otherwise the +report states that monetary cost is unavailable and still reports exact call +counts. + +## 8. Execution and CI + +The evaluation is available through a local CLI with explicit model, sample, +output-directory, credential, and call-budget inputs. A manual +`workflow_dispatch` job MAY invoke the same CLI with a repository secret and +upload transcripts/report artifacts. The workflow MUST have no `pull_request` +trigger and MUST not gate merges. A missing secret is a skipped/not-run +evaluation, not a normal CI failure. + +Evaluation output is disposable run data. It is not committed to the +repository by the runner and should be written to a caller-selected output +directory or uploaded as an artifact with bounded retention. + +## 9. Explicit Stage-2B boundary + +Stage 2B may add a semantic rubric, calibrated LLM-as-Judge, inter-rater +agreement with human review, and baseline-vs-candidate statistical analysis. +Those layers must consume Stage-2A transcripts and deterministic reports; they +must not change Stage-1 invariants or reinterpret Stage-2A technical failures +as learning outcomes. From 111ae42e0a7fe848ae2232403f3445046c257c51 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:27:27 +0200 Subject: [PATCH 02/21] docs: clarify Stage 2A identity and outcome states --- ...function_definition_TutorQualityStage2A.md | 68 +++++++++++++------ 1 file changed, 49 insertions(+), 19 deletions(-) diff --git a/ssot/ssot_function_definition_TutorQualityStage2A.md b/ssot/ssot_function_definition_TutorQualityStage2A.md index 5ae8062b..400c3dd8 100644 --- a/ssot/ssot_function_definition_TutorQualityStage2A.md +++ b/ssot/ssot_function_definition_TutorQualityStage2A.md @@ -63,7 +63,8 @@ Minimum scenario fields: The initial Stage-2A anchor set is deliberately small: -1. `TQ-REG-001` PWM: strong answer and no semantic question repetition; +1. `TQ-REG-001` PWM: strong answer and no exact or Stage-1-heuristic question + repetition; 2. simple variable question; 3. Serial-output prediction; 4. incorrect answer and remediation; @@ -74,9 +75,17 @@ The initial Stage-2A anchor set is deliberately small: 9. EXPAND; 10. clearly off-topic learner answer. -The corpus may grow only by adding a reviewed scenario or a corpus-version -change. A scenario edit is an evaluation change, not a hidden implementation -change. +Adding, removing, or semantically changing an anchor scenario MUST increase +`corpusVersion`. A scenario edit is an evaluation change, not a hidden +implementation change. Non-semantic formatting changes may keep the version +only when the parsed scenario and its digest remain identical. + +The first implementation uses state-seeded turns for multi-step cases. Every +scripted learner answer is bound to a declared preceding question context. A +continuation step without that binding, or with a different preceding +question, is `invalid`; the runner MUST NOT silently pretend that the answer +was given to an arbitrary real-model question. Scenarios that do not need a +deterministic binding should remain single-step cases. ## 4. Run identity and metadata @@ -89,7 +98,7 @@ Each run has both a unique `runId` and a stable `evaluationIdentity`. - corpus ID and version; - provider ID; - requested model ID; -- prompt revision; +- prompt revision identifier and effective-template digest; - relevant provider parameters, including timeout, temperature when known, difficulty, sample count, and call budget. @@ -101,13 +110,23 @@ The runner MUST record, without secrets: - provider ID and endpoint origin when useful for diagnosis; - requested model and returned model; -- prompt revision; +- prompt revision identifier and effective-template digest; - Course Content revision; - scenario/corpus version; - Git SHA and dirty-state indicator; - sample index, turn index, timestamps, duration, and call counts; - timeout, temperature, difficulty, and configured call/sample limits. +`promptRevision` is not an arbitrary label. It consists of a versioned +application-owned prompt identifier and a SHA-256 digest of the effective +system/user prompt templates (before scenario values are inserted). A change +to an application-owned prompt template MUST change this revision. Per-turn +prompt digests MAY additionally be stored in transcripts. + +Real-provider evaluation requires a clean Git working tree. A dirty relevant +worktree makes the run preflight `invalid` and no provider call is issued; the +runner does not upload a diff as a substitute for the exact Git SHA. + Missing or inconsistent identity metadata, including an omitted or `auto` model, makes a sample `invalid`. It MUST NOT be reported as a Tutor-quality failure. If the provider returns a model different from the requested fixed @@ -121,7 +140,8 @@ scenario invocation. A runner preflight MUST verify that the requested model is available. Any returned-model mismatch invalidates the sample rather than silently accepting `auto` fallback. -Credentials are read only from an invocation environment variable. They MUST +Credentials are read only from a configured invocation environment variable. +The CLI may accept the variable's name, but never its value. They MUST never appear in a transcript, aggregate report, log line, exception message, Git diff, or uploaded artifact. Authorization headers are provider-internal and are never part of the evaluation artifact model. @@ -139,20 +159,28 @@ One JSON transcript is written per scenario sample. It contains: - ordered logical Tutor turns; - the Tutor request context needed for later review, excluding credentials and transport headers; -- raw provider result as returned to the service, normalized final Tutor - result, returned model, or technical error; +- parsed `LLMProvider` result as received by `TutorService` before + application-side repair/normalization, normalized final Tutor result, + returned model, or technical error; - deterministic check records and state-before/state-after snapshots; - a terminal sample status. -Transcript status is one of: +Every sample has two independent status axes: + +`executionStatus` is one of: - `completed`: all requested turns executed and metadata is valid; - `invalid`: identity/model/contract metadata is missing or inconsistent; -- `technical-failure`: provider error, timeout, malformed provider response, or - call-budget exhaustion prevented a turn; -- `deterministic-violation`: a Stage-1 invariant was observed on raw/final - output. A repaired raw response may be both technically completed and have a - violation record; the report MUST preserve that distinction. +- `technical-failure`: a provider error, timeout, malformed provider response, + or mid-run call-budget exhaustion prevented a turn; +- `not-run`: preflight stopped execution, for example because credentials were + missing or the call budget was zero before the first call. + +`invariantViolations` is always an array. It is empty or contains deterministic +Stage-1 violation records. A repaired raw response may therefore be +`executionStatus: completed` with non-empty `invariantViolations`; the report +MUST preserve that distinction. Missing credentials and preflight budget +exhaustion are `not-run`, not Tutor-quality failures. The artifact writer uses an allow-list of fields. It MUST NOT serialize the credential environment, `process.env`, Authorization headers, or arbitrary @@ -180,9 +208,10 @@ The aggregate report contains counts and rates with explicit denominators per scenario and overall. It MUST keep these categories separate: 1. `invalid` evaluation metadata; -2. technical provider failures; -3. deterministic Tutor-quality invariant violations; -4. completed observations. +2. `not-run` preflight outcomes; +3. technical provider failures; +4. deterministic Tutor-quality invariant violations; +5. completed observations. There are no semantic quality grades, learning-support scores, or pass/fail claims about actual learning in Stage 2A. @@ -196,7 +225,8 @@ counts. ## 8. Execution and CI The evaluation is available through a local CLI with explicit model, sample, -output-directory, credential, and call-budget inputs. A manual +output-directory, credential-environment-variable name, and call-budget +inputs. A credential value is never a CLI argument. A manual `workflow_dispatch` job MAY invoke the same CLI with a repository secret and upload transcripts/report artifacts. The workflow MUST have no `pull_request` trigger and MUST not gate merges. A missing secret is a skipped/not-run From cc9131148cf02eced4eba142586b6f27a1832d88 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:29:44 +0200 Subject: [PATCH 03/21] docs: plan Tutor Quality Stage 2A implementation --- docs/plan-tutor-quality-stage-2a.md | 96 +++++++++++++++++++++++++++++ 1 file changed, 96 insertions(+) create mode 100644 docs/plan-tutor-quality-stage-2a.md diff --git a/docs/plan-tutor-quality-stage-2a.md b/docs/plan-tutor-quality-stage-2a.md new file mode 100644 index 00000000..fcd90289 --- /dev/null +++ b/docs/plan-tutor-quality-stage-2a.md @@ -0,0 +1,96 @@ +# Tutor Quality – Stage 2A Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Build a small, reviewable real-provider evaluation foundation that runs the normal `TutorService` path, records secret-free transcripts, reapplies deterministic Stage-1 checks, and reports technical outcomes separately from Tutor invariant violations. + +**Architecture:** Keep Stage 2A as an explicitly invoked observation layer. A repository-owned YAML anchor manifest resolves sketch and typed Course Content fixtures. A testable runner receives an injected `LLMProvider`, clones scenario state per sample, invokes `TutorService` and `CurriculumTutorAdapter`, captures the parsed provider result before service validation/repair, writes allow-listed JSON artifacts, and aggregates deterministic categories. The CLI constructs the real `KiconnectProvider`, performs clean-worktree/credential/model/call-budget preflight, and is never imported by the application runtime or required PR workflows. + +**Tech Stack:** TypeScript/ESM, existing `TutorService`/`LLMProvider`/`KiconnectProvider`, `yaml`, Node `crypto`/`fs`, Vitest, `tsx`, GitHub Actions `workflow_dispatch`. + +**Spec:** `ssot/ssot_function_definition_TutorQualityStage2A.md` + +## Global Constraints + +- Preserve all Stage-1 hard gates and existing Tutor trust boundaries. +- Do not call a provider directly from the evaluator; use `TutorService` and the existing provider implementation. +- Do not modify normal provider fallback behavior. Stage 2A requires an explicit model and marks missing/mismatched model metadata `invalid`. +- Never accept or print a credential value, authorization header, `process.env`, or an uploaded diff. Accept only a credential environment-variable name. +- A dirty relevant Git worktree is an invalid preflight and issues no provider call. +- No real-provider call is made by unit tests, pull-request CI, or required checks. Missing credentials produce `not-run` output. +- No semantic grades, LLM-as-Judge, adaptive strategy, fact-extractor expansion, learner profiles, or learning-effect claims. +- Keep artifact schemas allow-listed, bounded, deterministic, and disposable. Do not commit generated run output. +- Use TDD: write a focused failing test, run it red, implement the smallest change, run it green, then commit each coherent task. + +## Review Focus + +- Verify the runner records both `executionStatus` and `invariantViolations`; a repaired response can be completed while retaining raw invariant violations. +- Verify raw means the parsed `ProviderQuestionResult.result` captured before TutorService validation/repair, never an HTTP envelope. +- Verify `evaluationIdentity` includes Git SHA, Course Content revision, corpus/version, provider/model, prompt revision and effective-template digest, and all bounded parameters. +- Verify prompt revision is app-owned/versioned and changes when the effective system/user templates change. +- Verify every multi-turn scripted answer has an explicit preceding-question binding; otherwise the scenario is invalid. +- Verify `TQ-REG-001` checks no variables-topic activation and no exact/Stage-1-heuristic repeat, not an unimplemented semantic judgement. +- Verify reports distinguish invalid metadata, not-run preflight, technical provider failures, invariant violations, and completed observations with explicit denominators. + +--- + +## Task 1 – Lock the corpus and fixture contract with tests + +- [ ] Add a failing loader/contract test for a manifest with `corpusId`, `corpusVersion`, stable scenario IDs, sketch references, course fixture/free-Tutor mode, synthetic answers, explicit turn bindings, and structural expectations. +- [ ] Add a failing test that rejects duplicate IDs, missing fixture references, `auto` model declarations, unbound continuation answers, and a corpus edit that does not increase the version when the parsed scenario digest changes. +- [ ] Add `evals/tutor-quality/anchor-corpus.yaml` at version 1 with the ten approved anchors: `TQ-REG-001`, simple variable, Serial prediction, incorrect answer, partial answer, strong answer/progression, unmatched/free Tutor, LEARN→DEEPEN, EXPAND, and off-topic answer. +- [ ] Add only the small required `.ino` fixtures under `evals/tutor-quality/fixtures/`; reuse the existing PWM fixture where possible rather than copying production content. +- [ ] Add a typed `anchor-course-content.ts` fixture factory for the small valid Course Content snapshots and seeded progression states required by topic activation, LEARN/DEEPEN, and EXPAND cases. Keep revision strings and question IDs explicit and reviewable. +- [ ] Run the focused corpus tests red before implementation and green after the loader/factory exists. + +## Task 2 – Make prompt revision metadata explicit without changing prompts + +- [ ] Add a failing unit test asserting a versioned prompt revision identifier and SHA-256 digest are stable, contain the effective system/initial-user/dialog-user template sources before scenario substitution, and change when a template source changes. +- [ ] Refactor only the prompt-source declarations needed by `TutorService` so existing generated prompt text remains byte-for-byte compatible; export the revision descriptor for the evaluator. +- [ ] Do not add a new prompt, quality rule, or runtime strategy. The revision helper is metadata only. +- [ ] Run existing Tutor prompt tests plus the new revision test. + +## Task 3 – Implement the injectable evaluation runner and transcript model + +- [ ] Add failing tests for fresh state/history cloning per sample, normal `TutorService` initial/dialog invocation, provider capture before validation/repair, deterministic repeat/solution/schema checks, state-before/state-after snapshots, and explicit expected structural checks. +- [ ] Add `server/services/tutor/evaluation/real-provider-evaluation.ts` with small typed contracts for corpus scenarios, invocation options, sample metadata, logical turns, deterministic check records, transcript artifacts, and aggregate reports. +- [ ] Inject the provider, clock, random suffix, Git metadata, and output writer seams so tests never need credentials, network, or a mutable repository. +- [ ] Wrap the provider to capture requests and parsed `ProviderQuestionResult` values before `TutorService` receives them. Keep the raw capture allow-listed and exclude transport envelopes/headers. +- [ ] Invoke `TutorService` with `CurriculumTutorAdapter` when a scenario declares Course Content; invoke the same service without planning for free-Tutor cases. +- [ ] Reuse exported `validateLearningQuestion` and `isSemanticallyRepeatedQuestion` for deterministic checks. Record raw complete-solution/repeat violations even when TutorService rejects or repairs the response; never turn these into semantic scores. +- [ ] Enforce fixed requested model, preflight model availability, returned-model equality, explicit call/sample limits, and no call beyond the budget. +- [ ] Implement the two status axes from the SSOT: `executionStatus` (`completed`, `invalid`, `technical-failure`, `not-run`) and `invariantViolations` (array). Classify missing credentials/zero preflight budget as `not-run`; provider errors, timeout, malformed responses, and mid-run budget exhaustion as technical failures; metadata/model/binding problems as invalid. +- [ ] Implement canonical JSON hashing for `evaluationIdentity`, run IDs with UTC timestamp plus collision-resistant suffix, and secret-free allow-listed JSON transcript writing. +- [ ] Implement aggregate counts/rates per scenario and overall with explicit denominators, exact provider-call counts, budget exhaustion, and cost `unavailable` when the provider supplies no cost data. +- [ ] Run the focused evaluator tests red before implementation and green after each runner slice. + +## Task 4 – Add adversarial fake-provider coverage + +- [ ] Add fake-provider tests for exact repeated questions, Stage-1 heuristic repeats, invalid schema, complete solution, forged/wrong planning metadata, provider error before commit, timeout, returned-model mismatch, and call-budget exhaustion. +- [ ] Assert repaired final planning metadata remains application-owned and state commits only after a successful TutorService request. +- [ ] Assert technical failures do not mutate progression state and are not counted as Tutor-quality violations unless a separate raw deterministic violation was observed. +- [ ] Assert no credential value appears in serialized transcript, report, thrown error, or logger input. + +## Task 5 – Add the explicit CLI and local execution contract + +- [ ] Add `scripts/tutor-quality-real-provider-eval.ts` as a thin CLI around the runner. Require a fixed `--model`, bounded `--samples` and `--max-calls`, `--output-dir`, corpus selection, and a credential environment-variable name; reject `auto` and any credential value flag. +- [ ] Add a package script such as `eval:tutor-quality:real` that is not referenced by `test`, `test:unit`, `test:tutor-quality`, or normal build gates. +- [ ] Resolve the repository Git SHA/clean state and configured Course Content/corpus revisions before execution. Abort as `invalid` without provider calls when preflight identity cannot be proven. +- [ ] Add a short operator document with a local command, expected output paths, missing-credential behavior, call-budget example, and explicit warning that Stage 2A observes deterministic integrity rather than learning effect. +- [ ] Test CLI argument validation and missing-credential `not-run` behavior without contacting a provider. + +## Task 6 – Add a manual-only workflow + +- [ ] Add `.github/workflows/tutor-quality-real-provider.yml` with only `workflow_dispatch`, explicit model/sample/call-budget inputs, Node version from `.nvmrc`, and a repository secret exposed only to the invoked process through the configured environment variable. +- [ ] Upload bounded transcript/report artifacts with retention; never print the secret or use a `pull_request`/required-check trigger. +- [ ] Make missing secret a visible skipped/not-run result, not a failing PR gate. +- [ ] Document that this workflow is optional/manual first; do not add nightly scheduling until cost and stability are known. + +## Task 7 – Verification and review handoff + +- [ ] Run the Node-version check, focused Stage-2A tests, existing `npm run test:tutor-quality`, `npm run check`, `npm run check:docs`, and `git diff --check`. +- [ ] Run the CLI in no-credential mode and verify it writes only the documented not-run artifact (or exits without artifacts according to the contract), with no secret-like data. +- [ ] Confirm generated transcripts/reports are ignored or written only to caller-selected disposable directories. +- [ ] Review the diff for accidental production behavior changes, duplicated Stage-1 validation, direct HTTP access, semantic scoring, and CI hard-gate coupling. +- [ ] Commit the implementation in coherent commits and report branch/base/HEAD, files, anchors, deterministic metrics, excluded Stage-2B dimensions, tests, and the absence of a real-provider run if credentials are unavailable. + From a96f105a0f0b3538b57bf1d7ca551a94d92f1a57 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:34:25 +0200 Subject: [PATCH 04/21] docs: refine Stage 2A evaluation boundaries --- docs/plan-tutor-quality-stage-2a.md | 19 ++++++----- ...function_definition_TutorQualityStage2A.md | 33 ++++++++++++------- 2 files changed, 32 insertions(+), 20 deletions(-) diff --git a/docs/plan-tutor-quality-stage-2a.md b/docs/plan-tutor-quality-stage-2a.md index fcd90289..0cb72218 100644 --- a/docs/plan-tutor-quality-stage-2a.md +++ b/docs/plan-tutor-quality-stage-2a.md @@ -16,7 +16,7 @@ - Do not call a provider directly from the evaluator; use `TutorService` and the existing provider implementation. - Do not modify normal provider fallback behavior. Stage 2A requires an explicit model and marks missing/mismatched model metadata `invalid`. - Never accept or print a credential value, authorization header, `process.env`, or an uploaded diff. Accept only a credential environment-variable name. -- A dirty relevant Git worktree is an invalid preflight and issues no provider call. +- Any tracked or indexed Git change is an invalid preflight and issues no provider call. Untracked files are invalidating only under versioned evaluation/runtime input roots; unrelated editor files, protected local SSOT files, and ignored output directories are excluded. - No real-provider call is made by unit tests, pull-request CI, or required checks. Missing credentials produce `not-run` output. - No semantic grades, LLM-as-Judge, adaptive strategy, fact-extractor expansion, learner profiles, or learning-effect claims. - Keep artifact schemas allow-listed, bounded, deterministic, and disposable. Do not commit generated run output. @@ -37,8 +37,9 @@ ## Task 1 – Lock the corpus and fixture contract with tests - [ ] Add a failing loader/contract test for a manifest with `corpusId`, `corpusVersion`, stable scenario IDs, sketch references, course fixture/free-Tutor mode, synthetic answers, explicit turn bindings, and structural expectations. -- [ ] Add a failing test that rejects duplicate IDs, missing fixture references, `auto` model declarations, unbound continuation answers, and a corpus edit that does not increase the version when the parsed scenario digest changes. -- [ ] Add `evals/tutor-quality/anchor-corpus.yaml` at version 1 with the ten approved anchors: `TQ-REG-001`, simple variable, Serial prediction, incorrect answer, partial answer, strong answer/progression, unmatched/free Tutor, LEARN→DEEPEN, EXPAND, and off-topic answer. +- [ ] Add a failing test that rejects duplicate IDs, missing fixture references, `auto` model declarations, and unbound continuation answers. +- [ ] Add a separate corpus-evolution validator and failing tests for `compare(previousCorpus, currentCorpus)`: a parsed/digest-changing add, removal, or semantic edit requires a higher `corpusVersion`; formatting-only changes may keep the version only when parsed content and digest are unchanged. Keep historical comparison out of the current-corpus loader. +- [ ] Add `evals/tutor-quality/anchor-corpus.yaml` at version 1 with the ten approved anchors: `TQ-REG-001` (no variables-topic activation and no exact or Stage-1-heuristic repeat), simple variable, Serial prediction, incorrect answer, partial answer, strong answer/progression, unmatched/free Tutor, LEARN→DEEPEN, EXPAND, and off-topic answer. - [ ] Add only the small required `.ino` fixtures under `evals/tutor-quality/fixtures/`; reuse the existing PWM fixture where possible rather than copying production content. - [ ] Add a typed `anchor-course-content.ts` fixture factory for the small valid Course Content snapshots and seeded progression states required by topic activation, LEARN/DEEPEN, and EXPAND cases. Keep revision strings and question IDs explicit and reviewable. - [ ] Run the focused corpus tests red before implementation and green after the loader/factory exists. @@ -52,16 +53,17 @@ ## Task 3 – Implement the injectable evaluation runner and transcript model -- [ ] Add failing tests for fresh state/history cloning per sample, normal `TutorService` initial/dialog invocation, provider capture before validation/repair, deterministic repeat/solution/schema checks, state-before/state-after snapshots, and explicit expected structural checks. +- [ ] Add failing tests for fresh state/history cloning per sample, normal `TutorService` initial/dialog invocation, provider capture before validation/repair, shared diagnostic repeat/solution/schema checks, state-before/state-after snapshots, and explicit expected structural checks. - [ ] Add `server/services/tutor/evaluation/real-provider-evaluation.ts` with small typed contracts for corpus scenarios, invocation options, sample metadata, logical turns, deterministic check records, transcript artifacts, and aggregate reports. - [ ] Inject the provider, clock, random suffix, Git metadata, and output writer seams so tests never need credentials, network, or a mutable repository. - [ ] Wrap the provider to capture requests and parsed `ProviderQuestionResult` values before `TutorService` receives them. Keep the raw capture allow-listed and exclude transport envelopes/headers. - [ ] Invoke `TutorService` with `CurriculumTutorAdapter` when a scenario declares Course Content; invoke the same service without planning for free-Tutor cases. -- [ ] Reuse exported `validateLearningQuestion` and `isSemanticallyRepeatedQuestion` for deterministic checks. Record raw complete-solution/repeat violations even when TutorService rejects or repairs the response; never turn these into semantic scores. -- [ ] Enforce fixed requested model, preflight model availability, returned-model equality, explicit call/sample limits, and no call beyond the budget. +- [ ] Split the existing pure learning-question validation internally into one shared diagnostic function that returns granular deterministic violation records and keep `validateLearningQuestion` as the existing throw/normalization wrapper. Export the diagnostic function for Stage 2A; add no new rule and preserve runtime behavior. +- [ ] Reuse that diagnostic function and `isSemanticallyRepeatedQuestion` for deterministic checks. Record raw complete-solution/repeat violations even when TutorService rejects or repairs the response; never turn these into semantic scores. +- [ ] Enforce fixed requested model, preflight model availability, returned-model equality, explicit sample limits, and a provider-call budget covering *all* external calls, including `listModels()` and generation. Report `providerCalls`, `modelListCalls`, and `generationCalls` separately; stop before any call that would exceed the budget. - [ ] Implement the two status axes from the SSOT: `executionStatus` (`completed`, `invalid`, `technical-failure`, `not-run`) and `invariantViolations` (array). Classify missing credentials/zero preflight budget as `not-run`; provider errors, timeout, malformed responses, and mid-run budget exhaustion as technical failures; metadata/model/binding problems as invalid. - [ ] Implement canonical JSON hashing for `evaluationIdentity`, run IDs with UTC timestamp plus collision-resistant suffix, and secret-free allow-listed JSON transcript writing. -- [ ] Implement aggregate counts/rates per scenario and overall with explicit denominators, exact provider-call counts, budget exhaustion, and cost `unavailable` when the provider supplies no cost data. +- [ ] Implement aggregate counts/rates per scenario and overall with explicit denominators, exact provider-call counts, separate model-list/generation counts, budget exhaustion, and cost `unavailable` when the provider supplies no cost data. Always write a run-level `not-run` report for missing credentials with `reason: missing-credential`, zero provider calls, and no sample transcripts. - [ ] Run the focused evaluator tests red before implementation and green after each runner slice. ## Task 4 – Add adversarial fake-provider coverage @@ -89,8 +91,7 @@ ## Task 7 – Verification and review handoff - [ ] Run the Node-version check, focused Stage-2A tests, existing `npm run test:tutor-quality`, `npm run check`, `npm run check:docs`, and `git diff --check`. -- [ ] Run the CLI in no-credential mode and verify it writes only the documented not-run artifact (or exits without artifacts according to the contract), with no secret-like data. +- [ ] Run the CLI in no-credential mode and verify it writes the documented run-level `not-run` report with `reason: missing-credential`, zero provider calls, no sample transcripts, and no secret-like data. - [ ] Confirm generated transcripts/reports are ignored or written only to caller-selected disposable directories. - [ ] Review the diff for accidental production behavior changes, duplicated Stage-1 validation, direct HTTP access, semantic scoring, and CI hard-gate coupling. - [ ] Commit the implementation in coherent commits and report branch/base/HEAD, files, anchors, deterministic metrics, excluded Stage-2B dimensions, tests, and the absence of a real-provider run if credentials are unavailable. - diff --git a/ssot/ssot_function_definition_TutorQualityStage2A.md b/ssot/ssot_function_definition_TutorQualityStage2A.md index 400c3dd8..3cb360c1 100644 --- a/ssot/ssot_function_definition_TutorQualityStage2A.md +++ b/ssot/ssot_function_definition_TutorQualityStage2A.md @@ -123,9 +123,13 @@ system/user prompt templates (before scenario values are inserted). A change to an application-owned prompt template MUST change this revision. Per-turn prompt digests MAY additionally be stored in transcripts. -Real-provider evaluation requires a clean Git working tree. A dirty relevant -worktree makes the run preflight `invalid` and no provider call is issued; the -runner does not upload a diff as a substitute for the exact Git SHA. +Real-provider evaluation requires a clean relevant Git state. Any tracked or +indexed change makes the run preflight `invalid` and no provider call is +issued. Untracked files are invalidating only when they are under a +versioned evaluation/runtime input root and could affect the run; known +untracked editor files, protected local SSOT files, and ignored output +directories do not affect the preflight. The runner does not upload a diff as +a substitute for the exact Git SHA. Missing or inconsistent identity metadata, including an omitted or `auto` model, makes a sample `invalid`. It MUST NOT be reported as a Tutor-quality @@ -147,8 +151,9 @@ Git diff, or uploaded artifact. Authorization headers are provider-internal and are never part of the evaluation artifact model. Missing credentials cause an explicit `not-run/missing-credential` result for -manual/workflow execution. They do not fail normal CI because Stage 2A is not -part of normal CI. +manual/workflow execution. The runner still writes a run-level report with +that reason and zero provider calls, but no sample transcripts. Missing +credentials do not fail normal CI because Stage 2A is not part of normal CI. ## 6. Transcript artifact contract @@ -190,7 +195,11 @@ provider response envelopes. Stage 2A may report only properties with a deterministic rule. The runner reuses the existing TutorService validation and bounded repeat heuristic and -records at least: +records at least. The validation exposes one shared pure diagnostic result for +schema, question-count, complete-solution, and related deterministic issues; +`TutorService` retains its existing throw/repair behavior by consuming that +result, and the evaluator consumes the same records rather than duplicating +rules. - provider response/schema validity; - exactly one primary question; @@ -216,11 +225,13 @@ scenario and overall. It MUST keep these categories separate: There are no semantic quality grades, learning-support scores, or pass/fail claims about actual learning in Stage 2A. -Every invocation has a hard maximum sample count and provider-call budget. -The runner stops before issuing a call that would exceed the budget. Provider -token/cost data is reported only when the provider supplies it; otherwise the -report states that monetary cost is unavailable and still reports exact call -counts. +Every invocation has a hard maximum sample count and provider-call budget. The +budget covers every external provider call, including `listModels()` preflight +and model resolution as well as generation; reports additionally separate +model-list calls from generation calls. The runner stops before issuing a call +that would exceed the budget. Provider token/cost data is reported only when +the provider supplies it; otherwise the report states that monetary cost is +unavailable and still reports exact call counts. ## 8. Execution and CI From c50a6ee600d385ed9e20daee9ae0d186bbb9ca65 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:40:04 +0200 Subject: [PATCH 05/21] feat: add Stage 2A anchor corpus contract --- evals/tutor-quality/anchor-corpus.yaml | 106 +++++++++ .../tutor-quality/fixtures/serial-output.ino | 11 + .../fixtures/simple-variable.ino | 11 + .../tutor/evaluation/anchor-corpus.ts | 205 ++++++++++++++++++ .../tutor/evaluation/anchor-course-content.ts | 189 ++++++++++++++++ .../tutor/evaluation/anchor-corpus.test.ts | 132 +++++++++++ 6 files changed, 654 insertions(+) create mode 100644 evals/tutor-quality/anchor-corpus.yaml create mode 100644 evals/tutor-quality/fixtures/serial-output.ino create mode 100644 evals/tutor-quality/fixtures/simple-variable.ino create mode 100644 server/services/tutor/evaluation/anchor-corpus.ts create mode 100644 server/services/tutor/evaluation/anchor-course-content.ts create mode 100644 tests/server/services/tutor/evaluation/anchor-corpus.test.ts diff --git a/evals/tutor-quality/anchor-corpus.yaml b/evals/tutor-quality/anchor-corpus.yaml new file mode 100644 index 00000000..93891f8b --- /dev/null +++ b/evals/tutor-quality/anchor-corpus.yaml @@ -0,0 +1,106 @@ +corpusId: unosim-tutor-quality-anchor +corpusVersion: 1 +scenarios: + - id: TQ-REG-001 + sketch: tests/fixtures/tutor-quality/TQ-REG-001-pwm.ino + courseContent: variables + turns: + - kind: initial + difficulty: 35 + - kind: dialog + question: Welchen Wert verwendet der Sketch an der betrachteten Integer-Variablen? + answer: brightness läuft in Fünferschritten zwischen 0 und 255 und bestimmt den PWM-Tastgrad. + bindsToQuestion: Welchen Wert verwendet der Sketch an der betrachteten Integer-Variablen? + difficulty: 35 + expected: + topicIdAbsent: variables-and-serial + stateUnchanged: true + questionNotRepeat: exact-or-heuristic + + - id: simple-variable + sketch: evals/tutor-quality/fixtures/simple-variable.ino + courseContent: variables + turns: + - kind: initial + difficulty: 20 + expected: + topicId: variables-and-serial + + - id: serial-output-prediction + sketch: evals/tutor-quality/fixtures/serial-output.ino + courseContent: variables + turns: + - kind: initial + difficulty: 35 + expected: + topicId: variables-and-serial + + - id: incorrect-answer-remediation + sketch: evals/tutor-quality/fixtures/simple-variable.ino + courseContent: free + turns: + - kind: dialog + question: Welchen Wert gibt der Sketch über Serial aus? + answer: Der Sketch gibt immer 99 aus. + bindsToQuestion: Welchen Wert gibt der Sketch über Serial aus? + difficulty: 25 + + - id: partial-answer-follow-up + sketch: evals/tutor-quality/fixtures/simple-variable.ino + courseContent: free + turns: + - kind: dialog + question: Wie hängt counter mit der Schleife zusammen? + answer: counter ist eine Zahl. + bindsToQuestion: Wie hängt counter mit der Schleife zusammen? + difficulty: 30 + + - id: strong-answer-progression + sketch: evals/tutor-quality/fixtures/serial-output.ino + courseContent: free + turns: + - kind: dialog + question: Welche Ausgabe erzeugt Serial.println im Sketch? + answer: Die Funktion schreibt den aktuellen Wert von counter als Zeile in die serielle Ausgabe. + bindsToQuestion: Welche Ausgabe erzeugt Serial.println im Sketch? + difficulty: 40 + + - id: unmatched-topic-free-tutor + sketch: tests/fixtures/tutor-quality/TQ-REG-001-pwm.ino + courseContent: variables + turns: + - kind: initial + difficulty: 35 + expected: + topicIdAbsent: variables-and-serial + + - id: learn-to-deepen + sketch: evals/tutor-quality/fixtures/simple-variable.ino + courseContent: progression-learn + turns: + - kind: dialog + question: Welche Rolle spielt der Integer-Datentyp im aktuellen Sketch? + answer: int speichert den ganzzahligen Wert von counter, der anschließend im Sketch verwendet wird. + bindsToQuestion: Welche Rolle spielt der Integer-Datentyp im aktuellen Sketch? + difficulty: 30 + expected: + learningPhase: DEEPEN + + - id: expand + sketch: evals/tutor-quality/fixtures/serial-output.ino + courseContent: progression-expand + turns: + - kind: initial + difficulty: 55 + expected: + learningPhase: EXPAND + + - id: off-topic-answer + sketch: evals/tutor-quality/fixtures/simple-variable.ino + courseContent: free + turns: + - kind: dialog + question: Welche konkrete Beobachtung zeigt der Sketch? + answer: Ich möchte lieber über Fußball und das Wetter sprechen. + bindsToQuestion: Welche konkrete Beobachtung zeigt der Sketch? + difficulty: 20 diff --git a/evals/tutor-quality/fixtures/serial-output.ino b/evals/tutor-quality/fixtures/serial-output.ino new file mode 100644 index 00000000..d1320f9e --- /dev/null +++ b/evals/tutor-quality/fixtures/serial-output.ino @@ -0,0 +1,11 @@ +int counter = 3; + +void setup() { + Serial.begin(9600); + Serial.println(counter); +} + +void loop() { + Serial.println(counter); + delay(1000); +} diff --git a/evals/tutor-quality/fixtures/simple-variable.ino b/evals/tutor-quality/fixtures/simple-variable.ino new file mode 100644 index 00000000..37139ec1 --- /dev/null +++ b/evals/tutor-quality/fixtures/simple-variable.ino @@ -0,0 +1,11 @@ +int counter = 3; + +void setup() { + Serial.begin(9600); + Serial.println(counter); +} + +void loop() { + counter += 1; + delay(1000); +} diff --git a/server/services/tutor/evaluation/anchor-corpus.ts b/server/services/tutor/evaluation/anchor-corpus.ts new file mode 100644 index 00000000..06765ad4 --- /dev/null +++ b/server/services/tutor/evaluation/anchor-corpus.ts @@ -0,0 +1,205 @@ +import { createHash } from "node:crypto"; + +export type TutorQualityCourseContentRef = "free" | string; + +export interface TutorQualityHistoryEntrySource { + readonly question: string; + readonly responseStyle?: "normal" | "philosophical"; + readonly answerRating?: 1 | 2 | 3 | 4 | 5; + readonly questionId?: string; +} + +export type TutorQualityTurnSource = + | { + readonly kind: "initial"; + readonly difficulty?: number; + } + | { + readonly kind: "dialog"; + readonly question: string; + readonly answer: string; + readonly bindsToQuestion: string; + readonly difficulty?: number; + readonly history?: readonly TutorQualityHistoryEntrySource[]; + }; + +export interface TutorQualityScenarioSource { + readonly id: string; + readonly sketch: string; + readonly courseContent: TutorQualityCourseContentRef; + readonly model?: string; + readonly turns: readonly TutorQualityTurnSource[]; + readonly expected?: { + readonly topicId?: string; + readonly topicIdAbsent?: string; + readonly learningPhase?: "LEARN" | "DEEPEN" | "EXPAND"; + readonly stateUnchanged?: boolean; + readonly questionNotRepeat?: "exact-or-heuristic"; + }; +} + +export interface TutorQualityCorpusSource { + readonly corpusId: string; + readonly corpusVersion: number; + readonly scenarios: readonly TutorQualityScenarioSource[]; +} + +export interface TutorQualityCorpusReferences { + readonly sketches: ReadonlySet; + readonly courseContentFixtures: ReadonlySet; +} + +export interface TutorQualityCorpus extends TutorQualityCorpusSource { + readonly digest: string; +} + +export interface CorpusEvolutionComparison { + readonly valid: boolean; + readonly reason?: "version-not-increased" | "version-changed-without-semantic-change"; +} + +function fail(message: string): never { + throw new Error(`Invalid Tutor Quality corpus: ${message}`); +} + +function assertObject(value: unknown, label: string): asserts value is Record { + if (value === null || typeof value !== "object" || Array.isArray(value)) fail(`${label} must be an object`); +} + +function assertNonEmptyString(value: unknown, label: string): asserts value is string { + if (typeof value !== "string" || value.trim().length === 0) fail(`${label} must be a non-empty string`); +} + +function assertDifficulty(value: unknown, label: string): void { + if (value !== undefined && (!Number.isInteger(value) || (value as number) < 1 || (value as number) > 100)) { + fail(`${label} must be an integer from 1 to 100`); + } +} + +function parseTurn(value: unknown, label: string): TutorQualityTurnSource { + assertObject(value, label); + if (value.kind === "initial") { + assertDifficulty(value.difficulty, `${label}.difficulty`); + return { kind: "initial", ...(value.difficulty === undefined ? {} : { difficulty: value.difficulty as number }) }; + } + if (value.kind !== "dialog") fail(`${label}.kind must be initial or dialog`); + assertNonEmptyString(value.question, `${label}.question`); + assertNonEmptyString(value.answer, `${label}.answer`); + assertNonEmptyString(value.bindsToQuestion, `${label}.bindsToQuestion`); + if (value.question !== value.bindsToQuestion) fail(`${label} bindsToQuestion must equal question`); + assertDifficulty(value.difficulty, `${label}.difficulty`); + if (value.history !== undefined) { + if (!Array.isArray(value.history)) fail(`${label}.history must be an array`); + for (const [index, entry] of value.history.entries()) { + const historyLabel = `${label}.history[${index}]`; + assertObject(entry, historyLabel); + assertNonEmptyString(entry.question, `${historyLabel}.question`); + if (entry.responseStyle !== undefined && entry.responseStyle !== "normal" && entry.responseStyle !== "philosophical") { + fail(`${historyLabel}.responseStyle is invalid`); + } + if (entry.answerRating !== undefined && ![1, 2, 3, 4, 5].includes(entry.answerRating as 1 | 2 | 3 | 4 | 5)) { + fail(`${historyLabel}.answerRating is invalid`); + } + if (entry.questionId !== undefined) assertNonEmptyString(entry.questionId, `${historyLabel}.questionId`); + } + } + return { + kind: "dialog", + question: value.question, + answer: value.answer, + bindsToQuestion: value.bindsToQuestion, + ...(value.difficulty === undefined ? {} : { difficulty: value.difficulty as number }), + ...(value.history === undefined ? {} : { history: value.history as readonly TutorQualityHistoryEntrySource[] }), + }; +} + +function normalizedCorpusSource(source: TutorQualityCorpusSource): TutorQualityCorpusSource { + return { + corpusId: source.corpusId, + corpusVersion: source.corpusVersion, + scenarios: source.scenarios.map((scenario) => ({ + ...scenario, + turns: scenario.turns.map((turn) => ({ ...turn })), + })), + }; +} + +function stableJson(value: unknown): string { + if (Array.isArray(value)) return `[${value.map(stableJson).join(",")}]`; + if (value !== null && typeof value === "object") { + return `{${Object.entries(value as Record) + .sort(([left], [right]) => left.localeCompare(right)) + .map(([key, item]) => `${JSON.stringify(key)}:${stableJson(item)}`) + .join(",")}}`; + } + return JSON.stringify(value); +} + +export function tutorQualityCorpusDigest(source: TutorQualityCorpusSource): string { + return createHash("sha256").update(stableJson(normalizedCorpusSource(source))).digest("hex"); +} + +export function parseTutorQualityCorpus( + source: TutorQualityCorpusSource, + references: TutorQualityCorpusReferences, +): TutorQualityCorpus { + assertObject(source, "corpus"); + assertNonEmptyString(source.corpusId, "corpusId"); + if (!Number.isInteger(source.corpusVersion) || source.corpusVersion < 1) fail("corpusVersion must be a positive integer"); + if (!Array.isArray(source.scenarios) || source.scenarios.length === 0) fail("scenarios must be a non-empty array"); + + const ids = new Set(); + const scenarios = source.scenarios.map((rawScenario, index) => { + const label = `scenarios[${index}]`; + assertObject(rawScenario, label); + assertNonEmptyString(rawScenario.id, `${label}.id`); + if (ids.has(rawScenario.id)) fail(`duplicate scenario id ${rawScenario.id}`); + ids.add(rawScenario.id); + assertNonEmptyString(rawScenario.sketch, `${label}.sketch`); + if (!references.sketches.has(rawScenario.sketch)) fail(`${label}.sketch is not a registered fixture`); + assertNonEmptyString(rawScenario.courseContent, `${label}.courseContent`); + if (rawScenario.courseContent !== "free" && !references.courseContentFixtures.has(rawScenario.courseContent)) { + fail(`${label}.courseContent is not a registered fixture`); + } + if (rawScenario.model !== undefined) { + assertNonEmptyString(rawScenario.model, `${label}.model`); + if (rawScenario.model === "auto") fail(`${label}.model must be a fixed model id`); + } + if (!Array.isArray(rawScenario.turns) || rawScenario.turns.length === 0) fail(`${label}.turns must be a non-empty array`); + const turns = rawScenario.turns.map((turn, turnIndex) => parseTurn(turn, `${label}.turns[${turnIndex}]`)); + if (rawScenario.expected !== undefined) { + assertObject(rawScenario.expected, `${label}.expected`); + if (rawScenario.expected.learningPhase !== undefined && !["LEARN", "DEEPEN", "EXPAND"].includes(rawScenario.expected.learningPhase as string)) { + fail(`${label}.expected.learningPhase is invalid`); + } + if (rawScenario.expected.questionNotRepeat !== undefined && rawScenario.expected.questionNotRepeat !== "exact-or-heuristic") { + fail(`${label}.expected.questionNotRepeat is invalid`); + } + } + return { + id: rawScenario.id, + sketch: rawScenario.sketch, + courseContent: rawScenario.courseContent, + ...(rawScenario.model === undefined ? {} : { model: rawScenario.model }), + turns, + ...(rawScenario.expected === undefined ? {} : { expected: rawScenario.expected as TutorQualityScenarioSource["expected"] }), + } satisfies TutorQualityScenarioSource; + }); + + const normalized = { corpusId: source.corpusId, corpusVersion: source.corpusVersion, scenarios } satisfies TutorQualityCorpusSource; + return { ...normalized, digest: tutorQualityCorpusDigest(normalized) }; +} + +export function compareTutorQualityCorpusVersions( + previous: TutorQualityCorpus, + current: TutorQualityCorpus, +): CorpusEvolutionComparison { + const changed = previous.digest !== current.digest; + if (!changed && previous.corpusVersion !== current.corpusVersion) { + return { valid: false, reason: "version-changed-without-semantic-change" }; + } + if (changed && current.corpusVersion <= previous.corpusVersion) { + return { valid: false, reason: "version-not-increased" }; + } + return { valid: true }; +} diff --git a/server/services/tutor/evaluation/anchor-course-content.ts b/server/services/tutor/evaluation/anchor-course-content.ts new file mode 100644 index 00000000..8fdba69b --- /dev/null +++ b/server/services/tutor/evaluation/anchor-course-content.ts @@ -0,0 +1,189 @@ +import type { TutorCapability } from "../../course-content/course-content-loader"; +import { curriculumTopicSchema, validateCurriculumTopic, type CurriculumTopic } from "../curriculum/curriculum-schema"; +import { createTutorProgressionState, type TutorProgressionState } from "../curriculum/progression-state"; +import type { TutorPlanningContentContext } from "../tutor-planning"; + +export const ANCHOR_COURSE_CONTENT_FIXTURE_IDS = [ + "variables", + "progression-learn", + "progression-expand", +] as const; + +export type AnchorCourseContentFixtureId = typeof ANCHOR_COURSE_CONTENT_FIXTURE_IDS[number]; + +const REVISION = "2".repeat(40); + +function variableTopic(schemaVersion: 1 | 2 = 1): CurriculumTopic { + const raw = { + schemaVersion, + id: "variables-and-serial", + title: "Variablen und Serial-Ausgabe", + locale: "de-DE", + activation: { any: [ + { fact: "type-used", values: ["int"] }, + { fact: "serial-call", values: ["print"] }, + ] }, + concepts: [{ + id: "variable-values", + title: "Variablenwerte", + objective: "Den Zusammenhang zwischen einem deklarierten Integerwert und seiner Verwendung erklären.", + prerequisites: [], + difficulty: { entry: [1, 50], transfer: [20, 80] }, + misconceptions: [], + indicators: [{ id: "relates-value-to-use", description: "Ordnet einen Integerwert seiner Verwendung im Sketch zu." }], + mastery: { + minimumSuccessfulProbes: 1, + successRatingAtLeast: 3, + requiredIndicators: ["relates-value-to-use"], + minimumDistinctQuestionKinds: 1, + recentWeakAnswersAllowed: 0, + }, + }], + questions: [ + { + id: "variable-value-recall", + concept: "variable-values", + indicator: "relates-value-to-use", + kind: "recall", + difficulty: [1, 50], + requires: [{ fact: "type-used", values: ["int"] }], + text: "Welche Rolle spielt der Integer-Datentyp im aktuellen Sketch?", + }, + { + id: "serial-output-prediction", + concept: "variable-values", + indicator: "relates-value-to-use", + kind: "prediction", + difficulty: [20, 70], + requires: [{ fact: "serial-call", values: ["print"] }], + text: "Welche Ausgabe erzeugt Serial.println im aktuellen Sketch?", + }, + { + id: "variable-output-transfer", + concept: "variable-values", + indicator: "relates-value-to-use", + kind: "transfer", + difficulty: [30, 90], + requires: [{ fact: "type-used", values: ["int"] }, { fact: "serial-call", values: ["print"] }], + text: "Wie würdest du die Veränderung von counter an der seriellen Ausgabe überprüfen?", + }, + ], + scaffolds: [], + progression: { + entryConcepts: ["variable-values"], + preferredOrder: ["variable-values"], + onRating: { + "1-2": "remediate", + "3": "clarify-same-indicator", + "4": "probe-missing-indicator", + "5": "evaluate-mastery-and-advance", + }, + }, + ...(schemaVersion === 2 ? { + deepening: { + minimumSuccessfulProbes: 1, + successRatingAtLeast: 4, + requiredQuestionKinds: ["transfer"], + recentWeakAnswersAllowed: 0, + }, + extensions: [{ topic: "serial-output", objective: "Eine weitere serielle Beobachtung am Sketch ableiten." }], + } : {}), + }; + return validateCurriculumTopic(curriculumTopicSchema.parse(raw)); +} + +function serialOutputTopic(): CurriculumTopic { + return validateCurriculumTopic(curriculumTopicSchema.parse({ + schemaVersion: 1, + id: "serial-output", + title: "Serielle Ausgabe", + locale: "de-DE", + activation: { any: [{ fact: "type-used", values: ["float"] }] }, + concepts: [{ + id: "serial-observation", + title: "Serielle Beobachtung", + objective: "Eine serielle Ausgabe als Beobachtung beschreiben.", + prerequisites: [], + difficulty: { entry: [1, 60], transfer: [20, 80] }, + misconceptions: [], + indicators: [{ id: "observes-output", description: "Beschreibt eine serielle Ausgabe." }], + mastery: { + minimumSuccessfulProbes: 1, + successRatingAtLeast: 3, + requiredIndicators: ["observes-output"], + minimumDistinctQuestionKinds: 1, + recentWeakAnswersAllowed: 0, + }, + }], + questions: [{ + id: "serial-output-observe", + concept: "serial-observation", + indicator: "observes-output", + kind: "prediction", + difficulty: [1, 60], + requires: [{ fact: "type-used", values: ["float"] }], + text: "Welche serielle Beobachtung erwartest du?", + }], + scaffolds: [], + progression: { + entryConcepts: ["serial-observation"], + preferredOrder: ["serial-observation"], + onRating: { + "1-2": "remediate", + "3": "clarify-same-indicator", + "4": "probe-missing-indicator", + "5": "evaluate-mastery-and-advance", + }, + }, + })); +} + +function tutorCapability(topics: readonly CurriculumTopic[]): TutorCapability { + return { + status: "valid", + manifest: { + schemaVersion: 2, + topics: [], + strategies: [], + }, + topics, + strategies: [], + }; +} + +function baseState(): TutorProgressionState { + return createTutorProgressionState(REVISION); +} + +export function createAnchorCourseContent(fixtureId: AnchorCourseContentFixtureId): TutorPlanningContentContext { + switch (fixtureId) { + case "variables": + return { + revision: REVISION, + tutor: tutorCapability([variableTopic()]), + progressionState: baseState(), + }; + case "progression-learn": + return { + revision: REVISION, + tutor: tutorCapability([variableTopic()]), + progressionState: baseState(), + }; + case "progression-expand": { + const state = baseState(); + state.activeTopicId = "variables-and-serial"; + state.phase = "EXPAND"; + state.masteredTopicIds = ["variables-and-serial"]; + state.retainedPhases = { "variables-and-serial": "EXPAND" }; + return { + revision: REVISION, + tutor: tutorCapability([variableTopic(2), serialOutputTopic()]), + progressionState: state, + }; + } + } +} + +export function anchorCourseContentRevision(): string { + return REVISION; +} diff --git a/tests/server/services/tutor/evaluation/anchor-corpus.test.ts b/tests/server/services/tutor/evaluation/anchor-corpus.test.ts new file mode 100644 index 00000000..92f27e26 --- /dev/null +++ b/tests/server/services/tutor/evaluation/anchor-corpus.test.ts @@ -0,0 +1,132 @@ +import { readFileSync } from "node:fs"; +import { fileURLToPath } from "node:url"; +import { parse as parseYaml } from "yaml"; +import { describe, expect, it } from "vitest"; +import { + compareTutorQualityCorpusVersions, + parseTutorQualityCorpus, + type TutorQualityCorpusSource, +} from "../../../../../server/services/tutor/evaluation/anchor-corpus"; +import { + ANCHOR_COURSE_CONTENT_FIXTURE_IDS, + createAnchorCourseContent, +} from "../../../../../server/services/tutor/evaluation/anchor-course-content"; + +const references = { + sketches: new Set(["variable.ino"]), + courseContentFixtures: new Set(["variables"]), +}; + +function validSource(): TutorQualityCorpusSource { + return { + corpusId: "test-corpus", + corpusVersion: 1, + scenarios: [ + { + id: "variable", + sketch: "variable.ino", + courseContent: "variables", + turns: [ + { + kind: "dialog", + question: "Welchen Wert hat x?", + answer: "x ist drei.", + bindsToQuestion: "Welchen Wert hat x?", + difficulty: 20, + }, + ], + expected: { topicId: "variables-and-serial" }, + }, + ], + }; +} + +describe("Tutor Quality anchor corpus contract", () => { + it("contains the reviewed version-1 anchor set", () => { + const source = parseYaml(readFileSync(fileURLToPath(new URL("../../../../../evals/tutor-quality/anchor-corpus.yaml", import.meta.url)), "utf8")) as TutorQualityCorpusSource; + const corpus = parseTutorQualityCorpus(source, { + sketches: new Set(source.scenarios.map(({ sketch }) => sketch)), + courseContentFixtures: new Set(["variables", "progression-learn", "progression-expand"]), + }); + + expect(corpus.corpusVersion).toBe(1); + expect(corpus.scenarios.map(({ id }) => id)).toEqual([ + "TQ-REG-001", + "simple-variable", + "serial-output-prediction", + "incorrect-answer-remediation", + "partial-answer-follow-up", + "strong-answer-progression", + "unmatched-topic-free-tutor", + "learn-to-deepen", + "expand", + "off-topic-answer", + ]); + }); + + it("parses stable scenario references and explicit question bindings", () => { + const corpus = parseTutorQualityCorpus(validSource(), references); + + expect(corpus.corpusId).toBe("test-corpus"); + expect(corpus.scenarios[0]).toMatchObject({ + id: "variable", + sketch: "variable.ino", + courseContent: "variables", + turns: [{ bindsToQuestion: "Welchen Wert hat x?" }], + }); + }); + + it.each([ + ["duplicate scenario id", (source: TutorQualityCorpusSource) => ({ + ...source, + scenarios: [...source.scenarios, source.scenarios[0]!], + })], + ["missing sketch reference", (source: TutorQualityCorpusSource) => ({ + ...source, + scenarios: [{ ...source.scenarios[0]!, sketch: "missing.ino" }], + })], + ["auto model", (source: TutorQualityCorpusSource) => ({ + ...source, + scenarios: [{ ...source.scenarios[0]!, model: "auto" }], + })], + ["unbound continuation", (source: TutorQualityCorpusSource) => ({ + ...source, + scenarios: [{ + ...source.scenarios[0]!, + turns: [{ ...source.scenarios[0]!.turns[0]!, bindsToQuestion: undefined }], + }], + })], + ])("rejects %s", (_label, mutate) => { + expect(() => parseTutorQualityCorpus(mutate(validSource()), references)).toThrow(); + }); + + it("keeps historical version comparison outside current-corpus parsing", () => { + const previous = parseTutorQualityCorpus(validSource(), references); + const changed = parseTutorQualityCorpus({ + ...validSource(), + scenarios: [{ ...validSource().scenarios[0]!, turns: [{ + ...validSource().scenarios[0]!.turns[0]!, + answer: "x ist vier.", + bindsToQuestion: "Welchen Wert hat x?", + }] }], + }, references); + const bumped = { ...changed, corpusVersion: previous.corpusVersion + 1 }; + + expect(compareTutorQualityCorpusVersions(previous, changed).valid).toBe(false); + expect(compareTutorQualityCorpusVersions(previous, bumped).valid).toBe(true); + expect(compareTutorQualityCorpusVersions(previous, previous).valid).toBe(true); + }); + + it("provides valid typed Course Content fixtures for progression anchors", () => { + for (const fixtureId of ANCHOR_COURSE_CONTENT_FIXTURE_IDS) { + const content = createAnchorCourseContent(fixtureId); + expect(content.tutor?.status).toBe("valid"); + expect(content.revision).toHaveLength(40); + } + expect(createAnchorCourseContent("progression-expand").progressionState).toMatchObject({ + activeTopicId: "variables-and-serial", + phase: "EXPAND", + masteredTopicIds: ["variables-and-serial"], + }); + }); +}); From 5a521fc9121ffadae0a5ae1da00145f85e27fd42 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:41:27 +0200 Subject: [PATCH 06/21] feat: expose Tutor prompt revision metadata --- server/services/tutor/tutor-service.ts | 35 +++++++++++++++++++ .../evaluation/tutor-prompt-revision.test.ts | 25 +++++++++++++ 2 files changed, 60 insertions(+) create mode 100644 tests/server/services/tutor/evaluation/tutor-prompt-revision.test.ts diff --git a/server/services/tutor/tutor-service.ts b/server/services/tutor/tutor-service.ts index 87bd7f61..e813fbf5 100644 --- a/server/services/tutor/tutor-service.ts +++ b/server/services/tutor/tutor-service.ts @@ -1,3 +1,4 @@ +import { createHash } from "node:crypto"; import { analyzeStaticIO } from "@shared/io-registry-parser"; import { tutorContentResultSchema, @@ -712,3 +713,37 @@ export { sanitizeMermaid, validateLearningQuestion, }; + +export interface TutorPromptTemplateSources { + readonly system: string; + readonly initialUser: string; + readonly dialogUser: string; +} + +export function digestTutorPromptTemplates(sources: TutorPromptTemplateSources): string { + return createHash("sha256").update(JSON.stringify(sources)).digest("hex"); +} + +const TUTOR_PROMPT_TEMPLATE_SOURCES: TutorPromptTemplateSources = { + system: TUTOR_SYSTEM_PROMPT, + initialUser: [ + buildUserPrompt.toString(), + TUTOR_DIFFICULTY_GUIDANCE, + TUTOR_CONCRETE_REFERENCE_GUIDANCE, + buildTutorStrategyGuidance.toString(), + buildTutorLearningObjectivesGuidance.toString(), + ].join("\n"), + dialogUser: [ + buildDialogPrompt.toString(), + TUTOR_DIFFICULTY_GUIDANCE, + TUTOR_CONCRETE_REFERENCE_GUIDANCE, + buildTutorStrategyGuidance.toString(), + buildTutorLearningObjectivesGuidance.toString(), + ].join("\n"), +}; + +export const TUTOR_PROMPT_REVISION = { + id: "tutor-prompts-v1", + sources: TUTOR_PROMPT_TEMPLATE_SOURCES, + templateDigest: digestTutorPromptTemplates(TUTOR_PROMPT_TEMPLATE_SOURCES), +} as const; diff --git a/tests/server/services/tutor/evaluation/tutor-prompt-revision.test.ts b/tests/server/services/tutor/evaluation/tutor-prompt-revision.test.ts new file mode 100644 index 00000000..b1510b92 --- /dev/null +++ b/tests/server/services/tutor/evaluation/tutor-prompt-revision.test.ts @@ -0,0 +1,25 @@ +import { describe, expect, it } from "vitest"; +import { + TUTOR_PROMPT_REVISION, + digestTutorPromptTemplates, + type TutorPromptTemplateSources, +} from "../../../../../server/services/tutor/tutor-service"; + +describe("Tutor prompt revision metadata", () => { + it("exposes versioned system and user template sources", () => { + expect(TUTOR_PROMPT_REVISION.id).toMatch(/^tutor-prompts-v\d+$/); + expect(TUTOR_PROMPT_REVISION.templateDigest).toMatch(/^[a-f0-9]{64}$/); + expect(TUTOR_PROMPT_REVISION.sources.system).toContain("didaktischer Tutor"); + expect(TUTOR_PROMPT_REVISION.sources.initialUser).toContain("Sketch:"); + expect(TUTOR_PROMPT_REVISION.sources.dialogUser).toContain("Nutzerantwort"); + expect(TUTOR_PROMPT_REVISION.templateDigest).toBe(digestTutorPromptTemplates(TUTOR_PROMPT_REVISION.sources)); + }); + + it("changes the digest when an effective template source changes", () => { + const changed: TutorPromptTemplateSources = { + ...TUTOR_PROMPT_REVISION.sources, + system: `${TUTOR_PROMPT_REVISION.sources.system}\nchanged`, + }; + expect(digestTutorPromptTemplates(changed)).not.toBe(TUTOR_PROMPT_REVISION.templateDigest); + }); +}); From 19a4db4293c89116f4c3b33c682237d9081ccd2e Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:43:03 +0200 Subject: [PATCH 07/21] feat: expose shared Tutor question diagnostics --- server/services/tutor/tutor-service.ts | 59 +++++++++++++++---- .../tutor-question-diagnostics.test.ts | 25 ++++++++ 2 files changed, 72 insertions(+), 12 deletions(-) create mode 100644 tests/server/services/tutor/evaluation/tutor-question-diagnostics.test.ts diff --git a/server/services/tutor/tutor-service.ts b/server/services/tutor/tutor-service.ts index e813fbf5..e6723950 100644 --- a/server/services/tutor/tutor-service.ts +++ b/server/services/tutor/tutor-service.ts @@ -440,25 +440,59 @@ function buildPhilosophicalFallback( }; } -function validateLearningQuestion(result: TutorContentResult, difficulty?: TutorDifficulty): TutorContentResult { - const { mermaid: rawMermaid, ...resultWithoutMermaid } = result; - const sanitizedMermaid = sanitizeMermaid(rawMermaid); +export type LearningQuestionViolationCode = + | "schema-invalid" + | "multiple-primary-questions" + | "complete-solution" + | "philosophical-answer-rating"; + +export interface LearningQuestionViolation { + readonly code: LearningQuestionViolationCode; +} + +export interface LearningQuestionInspection { + readonly normalizedResult?: TutorContentResult; + readonly violations: readonly LearningQuestionViolation[]; +} + +function inspectLearningQuestion(result: unknown, difficulty?: TutorDifficulty): LearningQuestionInspection { + const input = typeof result === "object" && result !== null && !Array.isArray(result) + ? result as Record + : {}; + const preSchemaViolations: LearningQuestionViolation[] = []; + if (input.responseStyle === "philosophical" && input.answerRating !== undefined) { + preSchemaViolations.push({ code: "philosophical-answer-rating" }); + } + const { mermaid: rawMermaid, ...resultWithoutMermaid } = input; + const sanitizedMermaid = typeof rawMermaid === "string" ? sanitizeMermaid(rawMermaid) : undefined; const parsed = tutorContentResultSchema.safeParse({ ...resultWithoutMermaid, ...(sanitizedMermaid ? { mermaid: sanitizedMermaid } : {}), }); - if ( - !parsed.success - || (parsed.data.question.match(/\?/g)?.length ?? 0) > 1 - || containsCompleteSolution(parsed.data.question) - || (parsed.data.feedback !== undefined && containsCompleteSolution(parsed.data.feedback)) - ) { - throw new TutorProviderError("invalid-response"); + if (!parsed.success) { + return { + violations: [...preSchemaViolations, { code: "schema-invalid" }], + }; + } + const violations: LearningQuestionViolation[] = [...preSchemaViolations]; + if ((parsed.data.question.match(/\?/g)?.length ?? 0) > 1) { + violations.push({ code: "multiple-primary-questions" }); } - if (parsed.data.responseStyle === "philosophical" && parsed.data.answerRating !== undefined) { + if (containsCompleteSolution(parsed.data.question) || (parsed.data.feedback !== undefined && containsCompleteSolution(parsed.data.feedback))) { + violations.push({ code: "complete-solution" }); + } + return { + normalizedResult: difficulty === undefined ? parsed.data : { ...parsed.data, difficulty }, + violations, + }; +} + +function validateLearningQuestion(result: TutorContentResult, difficulty?: TutorDifficulty): TutorContentResult { + const inspection = inspectLearningQuestion(result, difficulty); + if (inspection.violations.length > 0 || inspection.normalizedResult === undefined) { throw new TutorProviderError("invalid-response"); } - return difficulty === undefined ? parsed.data : { ...parsed.data, difficulty }; + return inspection.normalizedResult; } function stripProviderPlanningMetadata(result: TutorContentResult): TutorContentResult { @@ -711,6 +745,7 @@ export { isClearlyNonLearningAnswer, isSemanticallyRepeatedQuestion, sanitizeMermaid, + inspectLearningQuestion, validateLearningQuestion, }; diff --git a/tests/server/services/tutor/evaluation/tutor-question-diagnostics.test.ts b/tests/server/services/tutor/evaluation/tutor-question-diagnostics.test.ts new file mode 100644 index 00000000..57341298 --- /dev/null +++ b/tests/server/services/tutor/evaluation/tutor-question-diagnostics.test.ts @@ -0,0 +1,25 @@ +import { describe, expect, it } from "vitest"; +import { + inspectLearningQuestion, + validateLearningQuestion, +} from "../../../../../server/services/tutor/tutor-service"; + +describe("shared Tutor learning-question diagnostics", () => { + it("reports granular deterministic violations", () => { + expect(inspectLearningQuestion({ question: "Welche? Zweite?" }).violations.map(({ code }) => code)) + .toContain("multiple-primary-questions"); + expect(inspectLearningQuestion({ question: "void setup() {} void loop() {}" }).violations.map(({ code }) => code)) + .toContain("complete-solution"); + expect(inspectLearningQuestion({ responseStyle: "normal", question: "Welche?", answerRating: 4 }).violations).toHaveLength(0); + expect(inspectLearningQuestion({ responseStyle: "philosophical", question: "Welche?", answerRating: 4 }).violations.map(({ code }) => code)) + .toContain("philosophical-answer-rating"); + expect(inspectLearningQuestion({ responseStyle: "normal", answerRating: "4" }).violations.map(({ code }) => code)) + .toContain("schema-invalid"); + }); + + it("keeps the existing validation throw behavior backed by the same records", () => { + const invalid = { question: "void setup() {} void loop() {}" }; + expect(() => validateLearningQuestion(invalid as never)).toThrowError(expect.objectContaining({ kind: "invalid-response" })); + expect(inspectLearningQuestion(invalid).violations.map(({ code }) => code)).toContain("complete-solution"); + }); +}); From 16eaef73decfa5436907020907a3bb6bd77e4650 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 22:57:34 +0200 Subject: [PATCH 08/21] feat: add Stage 2A real-provider evaluation runner --- .../workflows/tutor-quality-real-provider.yml | 56 ++ docs/tutor-quality-stage-2a.md | 52 ++ evals/tutor-quality/anchor-corpus.yaml | 2 - package.json | 1 + scripts/tutor-quality-real-provider-eval.ts | 220 +++++ .../tutor/evaluation/anchor-corpus.ts | 3 + .../evaluation/real-provider-evaluation.ts | 878 ++++++++++++++++++ .../real-provider-evaluation.test.ts | 216 +++++ .../tutor-quality-real-provider-cli.test.ts | 28 + 9 files changed, 1454 insertions(+), 2 deletions(-) create mode 100644 .github/workflows/tutor-quality-real-provider.yml create mode 100644 docs/tutor-quality-stage-2a.md create mode 100644 scripts/tutor-quality-real-provider-eval.ts create mode 100644 server/services/tutor/evaluation/real-provider-evaluation.ts create mode 100644 tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts create mode 100644 tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts diff --git a/.github/workflows/tutor-quality-real-provider.yml b/.github/workflows/tutor-quality-real-provider.yml new file mode 100644 index 00000000..29262606 --- /dev/null +++ b/.github/workflows/tutor-quality-real-provider.yml @@ -0,0 +1,56 @@ +name: Tutor Quality Real-Provider Evaluation + +on: + workflow_dispatch: + inputs: + model: + description: Fixed provider model ID (auto is rejected) + required: true + type: string + samples: + description: Samples per anchor scenario + required: true + default: "1" + type: string + max_calls: + description: Maximum provider calls, including model-list calls + required: true + default: "30" + type: string + +permissions: + contents: read + +jobs: + evaluate: + name: Stage 2A observation run + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v4 + + - uses: actions/setup-node@v4 + with: + node-version-file: .nvmrc + cache: npm + + - run: npm ci + + - name: Run optional real-provider evaluation + env: + UNOSIM_TUTOR_EVAL_CREDENTIAL: ${{ secrets.UNOSIM_TUTOR_EVAL_CREDENTIAL }} + run: | + npm run eval:tutor-quality:real -- \ + --model "${{ inputs.model }}" \ + --samples "${{ inputs.samples }}" \ + --max-calls "${{ inputs.max_calls }}" \ + --credential-env UNOSIM_TUTOR_EVAL_CREDENTIAL \ + --output-dir .tutor-quality-output + + - name: Upload Stage 2A artifacts + if: always() + uses: actions/upload-artifact@v4 + with: + name: tutor-quality-stage-2a-${{ github.run_id }} + path: .tutor-quality-output/ + if-no-files-found: error + retention-days: 14 diff --git a/docs/tutor-quality-stage-2a.md b/docs/tutor-quality-stage-2a.md new file mode 100644 index 00000000..cad820a1 --- /dev/null +++ b/docs/tutor-quality-stage-2a.md @@ -0,0 +1,52 @@ +# Tutor Quality Stage 2A – Real-Provider Evaluation + +Stage 2A is an explicitly invoked observation run. It uses the versioned +anchor corpus and the normal `TutorService`/`KiconnectProvider` path. It does +not run in unit tests or pull-request gates and does not measure learning +effect or assign semantic quality scores. + +## Local run + +Use a clean relevant Git state, a fixed model ID, and a caller-selected +disposable output directory: + +```sh +npm run eval:tutor-quality:real -- \ + --model \ + --credential-env UNOSIM_TUTOR_EVAL_CREDENTIAL \ + --samples 2 \ + --max-calls 30 \ + --output-dir /tmp/unosim-tutor-quality-stage-2a +``` + +The credential is read from the named environment variable only. Never pass a +credential value as a command-line argument. The runner never writes the +credential, authorization headers, or `process.env` to an artifact. + +The call budget includes model-list and generation calls. The report exposes +both categories separately. If the credential is absent, the command writes a +run-level `report.json` with `runStatus: "not-run"`, +`reason: "missing-credential"`, zero provider calls, and no sample +transcripts. + +## Artifacts + +`report.json` contains the run identity, corpus/course/prompt/provider/model +metadata, explicit status counts and rates, provider-call counts, and per- +scenario aggregates. One allow-listed JSON transcript is written per sample +when execution starts. Transcripts retain synthetic inputs, Tutor request +prompts, parsed provider results before service validation/repair, normalized +Tutor results, state snapshots, deterministic checks, and technical errors. + +`completed` observations may still contain deterministic invariant violations; +technical failures and invalid evaluation metadata remain separate categories. +Generated output is disposable and must not be committed. + +## Scope boundary + +Stage 2A reports schema/response validity, question-count and complete-solution +guards, bounded question-repeat checks, application-owned planning/state +consistency, provider failures, and execution counts. It deliberately does +not judge clarity, scaffolding quality, difficulty, learning support, or +learning progress. Those require the Stage 2B semantic and human-calibration +workflows. diff --git a/evals/tutor-quality/anchor-corpus.yaml b/evals/tutor-quality/anchor-corpus.yaml index 93891f8b..f4149255 100644 --- a/evals/tutor-quality/anchor-corpus.yaml +++ b/evals/tutor-quality/anchor-corpus.yaml @@ -5,8 +5,6 @@ scenarios: sketch: tests/fixtures/tutor-quality/TQ-REG-001-pwm.ino courseContent: variables turns: - - kind: initial - difficulty: 35 - kind: dialog question: Welchen Wert verwendet der Sketch an der betrachteten Integer-Variablen? answer: brightness läuft in Fünferschritten zwischen 0 und 255 und bestimmt den PWM-Tastgrad. diff --git a/package.json b/package.json index c0c85044..820bc749 100644 --- a/package.json +++ b/package.json @@ -46,6 +46,7 @@ "test:security:inputs": "TEST_BUDGET_SUITE=security-inputs TEST_BUDGET_MS=15000 LOG_LEVEL=warn vitest run --project=unit-node --project=unit-node-http tests/shared/input-limits.test.ts tests/shared/websocket-direction-schemas.test.ts tests/server/security/safe-paths.test.ts tests/server/routes/compiler.routes.test.ts tests/integration/simulation-state-sequence.test.ts --reporter=default --reporter=./scripts/test-budget-reporter.mjs", "test:tutor-quality": "TEST_BUDGET_SUITE=tutor-quality TEST_BUDGET_MS=15000 LOG_LEVEL=warn vitest run --project=unit-node tests/server/services/tutor tests/server/services/course-content --reporter=default --reporter=./scripts/test-budget-reporter.mjs", "validate:tutor-course-content": "tsx scripts/validate-tutor-course-content.mjs", + "eval:tutor-quality:real": "tsx scripts/tutor-quality-real-provider-eval.ts", "test:all": "npm run test:unit && npm run test:integration && npm run test:docker", "test:watch": "LOG_LEVEL=info vitest --project=unit-client --project=unit-node --project=unit-node-http", "test:coverage": "node scripts/coverage-guard.mjs prepare && TEST_BUDGET_SUITE=unit-coverage TEST_BUDGET_WARN_MS=75000 TEST_BUDGET_FAIL_MS=90000 LOG_LEVEL=warn vitest run --coverage --project=unit-client --project=unit-node --project=unit-node-http --reporter=default --reporter=./scripts/test-budget-reporter.mjs && node scripts/coverage-guard.mjs validate", diff --git a/scripts/tutor-quality-real-provider-eval.ts b/scripts/tutor-quality-real-provider-eval.ts new file mode 100644 index 00000000..4685710e --- /dev/null +++ b/scripts/tutor-quality-real-provider-eval.ts @@ -0,0 +1,220 @@ +import { execFileSync } from "node:child_process"; +import { access, readFile } from "node:fs/promises"; +import path from "node:path"; +import { pathToFileURL } from "node:url"; +import { parse as parseYaml } from "yaml"; +import { createAnchorCourseContent, ANCHOR_COURSE_CONTENT_FIXTURE_IDS, type AnchorCourseContentFixtureId } from "../server/services/tutor/evaluation/anchor-course-content"; +import { + parseTutorQualityCorpus, + type TutorQualityCorpusSource, + type TutorQualityHistoryEntrySource, + type TutorQualityTurnSource, +} from "../server/services/tutor/evaluation/anchor-corpus"; +import { + runTutorQualityEvaluation, + type TutorQualityEvaluationScenario, + type TutorQualityGitState, + type TutorQualityTurn, +} from "../server/services/tutor/evaluation/real-provider-evaluation"; +import { KiconnectProvider } from "../server/services/tutor/kiconnect-provider"; +import type { LLMProvider } from "../server/services/tutor/llm-provider"; + +export interface TutorQualityCliOptions { + readonly model: string; + readonly samples: number; + readonly maxCalls: number; + readonly outputDir: string; + readonly corpusPath: string; + readonly credentialEnv: string; +} + +export interface TutorQualityCliDependencies { + readonly cwd?: string; + readonly environment?: NodeJS.ProcessEnv; + readonly provider?: LLMProvider; + readonly git?: TutorQualityGitState; +} + +const DEFAULT_CORPUS_PATH = "evals/tutor-quality/anchor-corpus.yaml"; +const DEFAULT_CREDENTIAL_ENV = "UNOSIM_TUTOR_EVAL_CREDENTIAL"; + +function readArgument(argv: readonly string[], index: number, flag: string): string { + const value = argv[index + 1]; + if (!value || value.startsWith("--")) throw new Error(`${flag} requires a value`); + return value; +} + +function parsePositiveInteger(value: string, flag: string, allowZero = false): number { + if (!/^\d+$/.test(value)) throw new Error(`${flag} must be an integer`); + const parsed = Number(value); + if (!Number.isSafeInteger(parsed) || (allowZero ? parsed < 0 : parsed < 1)) throw new Error(`${flag} is out of range`); + return parsed; +} + +function parseCredentialEnvironmentName(value: string): string { + if (!/^[A-Z][A-Z0-9_]{1,63}$/.test(value)) throw new Error("--credential-env must be an environment-variable name"); + return value; +} + +export function parseTutorQualityCliArgs(argv: readonly string[]): TutorQualityCliOptions { + let model: string | undefined; + let samples = 1; + let maxCalls = 20; + let outputDir: string | undefined; + let corpusPath = DEFAULT_CORPUS_PATH; + let credentialEnv = DEFAULT_CREDENTIAL_ENV; + for (let index = 0; index < argv.length; index += 1) { + const flag = argv[index]; + switch (flag) { + case "--model": + model = readArgument(argv, index, flag); + index += 1; + break; + case "--samples": + samples = parsePositiveInteger(readArgument(argv, index, flag), flag); + index += 1; + break; + case "--max-calls": + maxCalls = parsePositiveInteger(readArgument(argv, index, flag), flag, true); + index += 1; + break; + case "--output-dir": + outputDir = readArgument(argv, index, flag); + index += 1; + break; + case "--corpus": + corpusPath = readArgument(argv, index, flag); + index += 1; + break; + case "--credential-env": + credentialEnv = parseCredentialEnvironmentName(readArgument(argv, index, flag)); + index += 1; + break; + default: + throw new Error(`Unknown Tutor Quality evaluation option: ${flag}`); + } + } + if (!model || model === "auto") throw new Error("--model requires a fixed model id; auto is not allowed"); + if (!outputDir) throw new Error("--output-dir is required"); + if (!/^[A-Za-z0-9][A-Za-z0-9._:/-]{0,127}$/.test(model)) throw new Error("--model is invalid"); + return { model, samples, maxCalls, outputDir, corpusPath, credentialEnv }; +} + +function historyEntry(entry: TutorQualityHistoryEntrySource): import("../shared/tutor").TutorDialogTurn { + return { + question: entry.question, + answer: entry.answer ?? "synthetic scripted history", + responseStyle: entry.responseStyle ?? "normal", + ...(entry.answerRating === undefined ? {} : { answerRating: entry.answerRating }), + ...(entry.questionId === undefined ? {} : { questionId: entry.questionId }), + }; +} + +function materializeTurn(turn: TutorQualityTurnSource): TutorQualityTurn { + if (turn.kind === "initial") return turn; + return { + ...turn, + ...(turn.history === undefined ? {} : { history: turn.history.map(historyEntry) }), + }; +} + +async function loadEvaluationScenarios( + cwd: string, + corpusPath: string, +): Promise { + const manifestPath = path.resolve(cwd, corpusPath); + const source = parseYaml(await readFile(manifestPath, "utf8")) as TutorQualityCorpusSource; + const sketchRefs = new Set(); + for (const { sketch } of source.scenarios) { + try { + await access(path.resolve(cwd, sketch)); + sketchRefs.add(sketch); + } catch { + // Keep the reference absent so the corpus validator reports the contract error. + } + } + const courseContentFixtures = new Set(ANCHOR_COURSE_CONTENT_FIXTURE_IDS); + const corpus = parseTutorQualityCorpus(source, { sketches: sketchRefs, courseContentFixtures }); + return Promise.all(corpus.scenarios.map(async (scenario) => ({ + id: scenario.id, + corpusId: corpus.corpusId, + corpusVersion: corpus.corpusVersion, + sketchRef: scenario.sketch, + sketch: await readFile(path.resolve(cwd, scenario.sketch), "utf8"), + ...(scenario.courseContent === "free" + ? {} + : { courseContent: createAnchorCourseContent(scenario.courseContent as AnchorCourseContentFixtureId) }), + turns: scenario.turns.map(materializeTurn), + ...(scenario.expected === undefined ? {} : { expected: scenario.expected }), + }))); +} + +function statusLines(cwd: string): readonly string[] { + const output = execFileSync("git", ["status", "--porcelain=v1", "-uall"], { cwd, encoding: "utf8" }); + return output.split("\n").map((line) => line.trimEnd()).filter(Boolean); +} + +function relevantUntrackedPath(relativePath: string, outputDir: string): boolean { + const normalized = relativePath.replaceAll("\\", "/"); + const normalizedOutput = outputDir.replaceAll("\\", "/").replace(/\/$/, ""); + if (normalizedOutput && (normalized === normalizedOutput || normalized.startsWith(`${normalizedOutput}/`))) return false; + return normalized.startsWith("evals/tutor-quality/") + || normalized.startsWith("server/") + || normalized.startsWith("shared/") + || normalized.startsWith("scripts/") + || ["package.json", "package-lock.json", ".nvmrc"].includes(normalized); +} + +export function readTutorQualityGitState(cwd: string, outputDir: string): TutorQualityGitState { + const sha = execFileSync("git", ["rev-parse", "HEAD"], { cwd, encoding: "utf8" }).trim(); + const lines = statusLines(cwd); + const trackedClean = lines.every((line) => line.slice(0, 2) === "??"); + const outputRelative = path.relative(cwd, path.resolve(cwd, outputDir)); + const relevantUntrackedClean = lines.every((line) => { + if (line.slice(0, 2) !== "??") return true; + return !relevantUntrackedPath(line.slice(3).trim(), outputRelative); + }); + return { sha, trackedClean, relevantUntrackedClean }; +} + +export async function runTutorQualityCli( + argv: readonly string[], + dependencies: TutorQualityCliDependencies = {}, +) { + const options = parseTutorQualityCliArgs(argv); + const cwd = dependencies.cwd ?? process.cwd(); + const environment = dependencies.environment ?? process.env; + const scenarios = await loadEvaluationScenarios(cwd, options.corpusPath); + const git = dependencies.git ?? readTutorQualityGitState(cwd, options.outputDir); + const provider = dependencies.provider ?? new KiconnectProvider(); + const timeoutMs = Number(environment.UNOSIM_LLM_TIMEOUT_MS ?? 30_000); + const credential = environment[options.credentialEnv]; + return runTutorQualityEvaluation({ + scenarios, + provider, + providerId: "kiconnect", + endpointOrigin: environment.UNOSIM_LLM_BASE_URL ? new URL(environment.UNOSIM_LLM_BASE_URL).origin : undefined, + ...(credential ? { credential } : {}), + requestedModel: options.model, + samples: options.samples, + maxCalls: options.maxCalls, + timeoutMs, + temperature: 0.2, + outputDir: path.resolve(cwd, options.outputDir), + git, + }); +} + +const invokedDirectly = process.argv[1] !== undefined + && import.meta.url === pathToFileURL(path.resolve(process.argv[1])).href; + +if (invokedDirectly) { + runTutorQualityCli(process.argv.slice(2)) + .then(({ report }) => { + console.log(JSON.stringify({ runId: report.runId, runStatus: report.runStatus, reason: report.reason, providerCalls: report.providerCalls })); + }) + .catch((error: unknown) => { + console.error(error instanceof Error ? error.message : "Tutor Quality evaluation failed"); + process.exitCode = 1; + }); +} diff --git a/server/services/tutor/evaluation/anchor-corpus.ts b/server/services/tutor/evaluation/anchor-corpus.ts index 06765ad4..5631ed22 100644 --- a/server/services/tutor/evaluation/anchor-corpus.ts +++ b/server/services/tutor/evaluation/anchor-corpus.ts @@ -4,6 +4,7 @@ export type TutorQualityCourseContentRef = "free" | string; export interface TutorQualityHistoryEntrySource { readonly question: string; + readonly answer?: string; readonly responseStyle?: "normal" | "philosophical"; readonly answerRating?: 1 | 2 | 3 | 4 | 5; readonly questionId?: string; @@ -94,6 +95,7 @@ function parseTurn(value: unknown, label: string): TutorQualityTurnSource { const historyLabel = `${label}.history[${index}]`; assertObject(entry, historyLabel); assertNonEmptyString(entry.question, `${historyLabel}.question`); + if (entry.answer !== undefined) assertNonEmptyString(entry.answer, `${historyLabel}.answer`); if (entry.responseStyle !== undefined && entry.responseStyle !== "normal" && entry.responseStyle !== "philosophical") { fail(`${historyLabel}.responseStyle is invalid`); } @@ -153,6 +155,7 @@ export function parseTutorQualityCorpus( const label = `scenarios[${index}]`; assertObject(rawScenario, label); assertNonEmptyString(rawScenario.id, `${label}.id`); + if (!/^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(rawScenario.id)) fail(`${label}.id is not stable-safe`); if (ids.has(rawScenario.id)) fail(`duplicate scenario id ${rawScenario.id}`); ids.add(rawScenario.id); assertNonEmptyString(rawScenario.sketch, `${label}.sketch`); diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts new file mode 100644 index 00000000..74493098 --- /dev/null +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -0,0 +1,878 @@ +import { createHash, randomUUID } from "node:crypto"; +import { mkdir, writeFile } from "node:fs/promises"; +import type { TutorContentResult, TutorDialogTurn } from "@shared/tutor"; +import { + inspectLearningQuestion, + isSemanticallyRepeatedQuestion, + TUTOR_PROMPT_REVISION, + TutorService, +} from "../tutor-service"; +import { + TutorProviderError, + type LLMProvider, + type LLMProviderRequest, + type ProviderQuestionResult, +} from "../llm-provider"; +import { CurriculumTutorAdapter } from "../curriculum-tutor-adapter"; +import type { TutorPlanningContentContext } from "../tutor-planning"; +import type { TutorProgressionState } from "../curriculum/progression-state"; + +export type TutorQualityExecutionStatus = "completed" | "invalid" | "technical-failure" | "not-run"; + +export type TutorQualityTurn = + | { + readonly kind: "initial"; + readonly difficulty?: number; + } + | { + readonly kind: "dialog"; + readonly question: string; + readonly answer: string; + readonly bindsToQuestion: string; + readonly difficulty?: number; + readonly history?: readonly TutorDialogTurn[]; + }; + +export interface TutorQualityEvaluationScenario { + readonly id: string; + readonly corpusId: string; + readonly corpusVersion: number; + readonly sketchRef: string; + readonly sketch: string; + readonly courseContent?: TutorPlanningContentContext; + readonly turns: readonly TutorQualityTurn[]; + readonly expected?: { + readonly topicId?: string; + readonly topicIdAbsent?: string; + readonly learningPhase?: "LEARN" | "DEEPEN" | "EXPAND"; + readonly stateUnchanged?: boolean; + readonly questionNotRepeat?: "exact-or-heuristic"; + }; +} + +export interface TutorQualityGitState { + readonly sha: string; + readonly trackedClean: boolean; + readonly relevantUntrackedClean: boolean; +} + +export interface TutorQualityEvaluationOptions { + readonly scenarios: readonly TutorQualityEvaluationScenario[]; + readonly provider: LLMProvider; + readonly providerId: string; + readonly endpointOrigin?: string; + readonly credential?: string; + readonly requestedModel: string; + readonly samples: number; + readonly maxCalls: number; + readonly timeoutMs?: number; + readonly temperature?: number; + readonly outputDir?: string; + readonly git: TutorQualityGitState; + readonly now?: () => Date; + readonly runSuffix?: () => string; +} + +export interface TutorQualityDeterministicCheck { + readonly name: string; + readonly passed: boolean; + readonly details?: string; +} + +export interface TutorQualityInvariantViolation { + readonly code: string; + readonly source: "raw-provider" | "final-tutor" | "state" | "scenario"; + readonly turnIndex?: number; + readonly details?: string; +} + +export interface TutorQualityTechnicalError { + readonly kind: string; + readonly name: string; +} + +export interface TutorQualityProviderRequestArtifact { + readonly model: string; + readonly systemPrompt: string; + readonly userPrompt: string; +} + +export interface TutorQualityTranscriptTurn { + readonly index: number; + readonly input: TutorQualityTurn; + readonly providerRequest?: TutorQualityProviderRequestArtifact; + readonly rawProviderResult?: Record; + readonly finalTutorResult?: TutorContentResult; + readonly returnedModel?: string; + readonly deterministicChecks: readonly TutorQualityDeterministicCheck[]; + readonly technicalError?: TutorQualityTechnicalError; +} + +export interface TutorQualityTranscript { + readonly schemaVersion: "tutor-quality-transcript-v1"; + readonly runId: string; + readonly evaluationIdentity: string; + readonly metadata: TutorQualityMetadata; + readonly scenario: { + readonly id: string; + readonly corpusId: string; + readonly corpusVersion: number; + readonly sketchRef: string; + readonly sketch: string; + readonly syntheticTurns: readonly TutorQualityTurn[]; + }; + readonly stateBefore?: TutorProgressionState; + readonly stateAfter?: TutorProgressionState; + readonly turns: readonly TutorQualityTranscriptTurn[]; + readonly deterministicChecks: readonly TutorQualityDeterministicCheck[]; + readonly executionStatus: TutorQualityExecutionStatus; + readonly invalidReason?: string; + readonly technicalError?: TutorQualityTechnicalError; + readonly invariantViolations: readonly TutorQualityInvariantViolation[]; +} + +export interface TutorQualityMetadata { + readonly providerId: string; + readonly endpointOrigin?: string; + readonly requestedModel: string; + readonly returnedModels: readonly string[]; + readonly promptRevision: { + readonly id: string; + readonly templateDigest: string; + }; + readonly courseContentRevision: string; + readonly corpusId: string; + readonly corpusVersion: number; + readonly gitSha: string; + readonly gitState: "clean"; + readonly sampleIndex: number; + readonly sampleCount: number; + readonly sampleStartedAt: string; + readonly sampleDurationMs: number; + readonly providerCalls: TutorQualityProviderCallCounts; + readonly timeoutMs?: number; + readonly temperature?: number; + readonly maxCalls: number; +} + +export interface TutorQualityProviderCallCounts { + readonly total: number; + readonly modelListCalls: number; + readonly generationCalls: number; +} + +export interface TutorQualityScenarioAggregate { + readonly samplesRequested: number; + readonly samplesObserved: number; + readonly completed: number; + readonly invalid: number; + readonly technicalFailures: number; + readonly notRun: number; + readonly invariantViolationSamples: number; + readonly providerCalls: TutorQualityProviderCallCounts; + readonly rates: TutorQualityRates; +} + +export interface TutorQualityRates { + readonly completedOfObserved: number | null; + readonly invalidOfObserved: number | null; + readonly technicalFailureOfObserved: number | null; + readonly notRunOfObserved: number | null; + readonly invariantViolationOfObserved: number | null; +} + +export interface TutorQualityEvaluationReport { + readonly schemaVersion: "tutor-quality-report-v1"; + readonly runId: string; + readonly evaluationIdentity: string; + readonly runStatus: TutorQualityExecutionStatus; + readonly reason?: string; + readonly providerId: string; + readonly requestedModel: string; + readonly credentialPresent: boolean; + readonly samplesRequested: number; + readonly samplesObserved: number; + readonly completed: number; + readonly invalid: number; + readonly technicalFailures: number; + readonly notRun: number; + readonly invariantViolationSamples: number; + readonly providerCalls: TutorQualityProviderCallCounts; + readonly byScenario: Readonly>; + readonly rates: TutorQualityRates; + readonly monetaryCost: "unavailable"; +} + +export interface TutorQualityEvaluationResult { + readonly report: TutorQualityEvaluationReport; + readonly transcripts: readonly TutorQualityTranscript[]; +} + +interface ProviderCapture { + readonly request: LLMProviderRequest; + readonly response?: ProviderQuestionResult; + readonly error?: unknown; +} + +class EvaluationBudgetExceeded extends Error { + constructor() { + super("call-budget-exhausted"); + this.name = "EvaluationBudgetExceeded"; + } +} + +class CountingProvider implements LLMProvider { + private totalCalls = 0; + private modelListCalls = 0; + private generationCalls = 0; + private readonly generations: ProviderCapture[] = []; + + constructor( + private readonly provider: LLMProvider, + private readonly maxCalls: number, + ) {} + + get counts(): TutorQualityProviderCallCounts { + return { + total: this.totalCalls, + modelListCalls: this.modelListCalls, + generationCalls: this.generationCalls, + }; + } + + get generationCaptures(): readonly ProviderCapture[] { + return this.generations; + } + + async listModels(credential: string): Promise { + this.reserveCall(); + this.modelListCalls += 1; + return this.provider.listModels(credential); + } + + async generateLearningQuestion(request: LLMProviderRequest, credential: string): Promise { + this.reserveCall(); + this.generationCalls += 1; + try { + const response = await this.provider.generateLearningQuestion(request, credential); + this.generations.push({ request, response }); + return response; + } catch (error) { + this.generations.push({ request, error }); + throw error; + } + } + + private reserveCall(): void { + if (this.totalCalls >= this.maxCalls) throw new EvaluationBudgetExceeded(); + this.totalCalls += 1; + } +} + +function stableJson(value: unknown): string { + if (Array.isArray(value)) return `[${value.map(stableJson).join(",")}]`; + if (value !== null && typeof value === "object") { + return `{${Object.entries(value as Record) + .sort(([left], [right]) => left.localeCompare(right)) + .map(([key, item]) => `${JSON.stringify(key)}:${stableJson(item)}`) + .join(",")}}`; + } + return JSON.stringify(value); +} + +function hashCanonical(value: unknown): string { + return createHash("sha256").update(stableJson(value)).digest("hex"); +} + +function clone(value: T): T { + return structuredClone(value); +} + +function courseRevision(scenario: TutorQualityEvaluationScenario): string { + return scenario.courseContent?.revision ?? "free-tutor"; +} + +function makeEvaluationIdentity(options: TutorQualityEvaluationOptions): string { + const first = options.scenarios[0]; + return hashCanonical({ + unosimGitSha: options.git.sha, + courseContentRevisions: [...new Set(options.scenarios.map(courseRevision))].sort(), + corpusId: first?.corpusId, + corpusVersion: first?.corpusVersion, + providerId: options.providerId, + requestedModel: options.requestedModel, + promptRevision: { + id: TUTOR_PROMPT_REVISION.id, + templateDigest: TUTOR_PROMPT_REVISION.templateDigest, + }, + parameters: { + timeoutMs: options.timeoutMs, + temperature: options.temperature, + sampleCount: options.samples, + maxCalls: options.maxCalls, + }, + }); +} + +function makeRunId(options: TutorQualityEvaluationOptions): string { + const now = (options.now ?? (() => new Date()))().toISOString().replaceAll(/[^0-9TZ]/g, ""); + return `tq2a-${now}-${(options.runSuffix ?? randomUUID)()}`; +} + +function metadata( + options: TutorQualityEvaluationOptions, + scenario: TutorQualityEvaluationScenario, + sampleIndex: number, + returnedModels: readonly string[] = [], + sampleStartedAt = new Date(0).toISOString(), + sampleDurationMs = 0, + providerCalls: TutorQualityProviderCallCounts = { total: 0, modelListCalls: 0, generationCalls: 0 }, +): TutorQualityMetadata { + return { + providerId: options.providerId, + ...(options.endpointOrigin ? { endpointOrigin: options.endpointOrigin } : {}), + requestedModel: options.requestedModel, + returnedModels, + promptRevision: { + id: TUTOR_PROMPT_REVISION.id, + templateDigest: TUTOR_PROMPT_REVISION.templateDigest, + }, + courseContentRevision: courseRevision(scenario), + corpusId: scenario.corpusId, + corpusVersion: scenario.corpusVersion, + gitSha: options.git.sha, + gitState: "clean", + sampleIndex, + sampleCount: options.samples, + sampleStartedAt, + sampleDurationMs, + providerCalls, + ...(options.timeoutMs === undefined ? {} : { timeoutMs: options.timeoutMs }), + ...(options.temperature === undefined ? {} : { temperature: options.temperature }), + maxCalls: options.maxCalls, + }; +} + +function technicalError(error: unknown): TutorQualityTechnicalError { + if (error instanceof EvaluationBudgetExceeded) return { kind: "call-budget-exhausted", name: error.name }; + if (error instanceof TutorProviderError) return { kind: error.kind, name: error.name }; + return { kind: "provider-error", name: error instanceof Error ? error.name : "UnknownError" }; +} + +function safeResult(value: unknown): Record | undefined { + if (value === null || typeof value !== "object" || Array.isArray(value)) return undefined; + const allowed = new Set([ + "responseStyle", "feedback", "question", "topic", "difficulty", "answerRating", "mermaid", + "topicId", "conceptId", "questionId", "indicatorId", "questionKind", "strategyId", "strategySource", + "contentRevision", "learningPhase", "activeTopicId", "masteredTopicIds", "progressionBlockedReason", + "extensionTargetTopicId", + ]); + const result: Record = {}; + for (const [key, item] of Object.entries(value as Record)) { + if (!allowed.has(key)) continue; + if (typeof item === "string" || typeof item === "number" || typeof item === "boolean" || item === undefined) { + result[key] = item; + } else if (Array.isArray(item) && item.every((entry) => typeof entry === "string")) { + result[key] = [...item]; + } + } + return result; +} + +function requestArtifact(request: LLMProviderRequest): TutorQualityProviderRequestArtifact { + return { + model: request.model, + systemPrompt: request.systemPrompt, + userPrompt: request.userPrompt, + }; +} + +function historyFromSource(history: readonly TutorDialogTurn[] | undefined): readonly TutorDialogTurn[] { + return history ?? []; +} + +function addCheck( + checks: TutorQualityDeterministicCheck[], + name: string, + passed: boolean, + details?: string, +): void { + checks.push({ name, passed, ...(details ? { details } : {}) }); +} + +function addViolation( + violations: TutorQualityInvariantViolation[], + code: string, + source: TutorQualityInvariantViolation["source"], + turnIndex: number | undefined, + details?: string, +): void { + violations.push({ code, source, ...(turnIndex === undefined ? {} : { turnIndex }), ...(details ? { details } : {}) }); +} + +function deterministicRawChecks( + rawResult: unknown, + turn: TutorQualityTurn, + turnIndex: number, + checks: TutorQualityDeterministicCheck[], + violations: TutorQualityInvariantViolation[], +): void { + const inspection = inspectLearningQuestion(rawResult, turn.difficulty); + const issueCodes = new Set(inspection.violations.map(({ code }) => code)); + const schemaValid = !issueCodes.has("schema-invalid"); + addCheck(checks, "raw-provider-schema-valid", schemaValid); + addCheck(checks, "raw-provider-one-primary-question", schemaValid && !issueCodes.has("multiple-primary-questions")); + addCheck(checks, "raw-provider-no-complete-solution", schemaValid && !issueCodes.has("complete-solution")); + for (const issue of inspection.violations) { + addViolation(violations, issue.code, "raw-provider", turnIndex); + } + if (typeof rawResult === "object" && rawResult !== null && !Array.isArray(rawResult)) { + const question = (rawResult as { question?: unknown }).question; + if (typeof question === "string" && turn.kind === "dialog") { + const previous = [turn.question, ...(turn.history ?? []).map(({ question: historyQuestion }) => historyQuestion)]; + const repeated = isSemanticallyRepeatedQuestion(question, previous); + addCheck(checks, "raw-provider-question-not-repeated", !repeated); + if (repeated) { + addViolation(violations, "question-repeat", "raw-provider", turnIndex, question === turn.question ? "exact" : "stage1-heuristic"); + } + } + } +} + +function expectedChecks( + scenario: TutorQualityEvaluationScenario, + result: TutorContentResult | undefined, + turn: TutorQualityTurn, + stateBefore: TutorProgressionState | undefined, + stateAfter: TutorProgressionState | undefined, + checks: TutorQualityDeterministicCheck[], + violations: TutorQualityInvariantViolation[], + turnIndex: number, +): void { + const expected = scenario.expected; + if (!expected || !result) return; + if (expected.topicId !== undefined) { + const passed = result.topicId === expected.topicId; + addCheck(checks, "expected-topic", passed, expected.topicId); + if (!passed) addViolation(violations, "topic-mismatch", "final-tutor", turnIndex, expected.topicId); + } + if (expected.topicIdAbsent !== undefined) { + const passed = result.topicId !== expected.topicIdAbsent && result.activeTopicId !== expected.topicIdAbsent; + addCheck(checks, "expected-topic-absent", passed, expected.topicIdAbsent); + if (!passed) addViolation(violations, "forbidden-topic-activation", "final-tutor", turnIndex, expected.topicIdAbsent); + } + if (expected.learningPhase !== undefined) { + const passed = result.learningPhase === expected.learningPhase; + addCheck(checks, "expected-learning-phase", passed, expected.learningPhase); + if (!passed) addViolation(violations, "phase-mismatch", "final-tutor", turnIndex, expected.learningPhase); + } + if (expected.stateUnchanged && stateBefore !== undefined && stateAfter !== undefined) { + const passed = stableJson(stateBefore) === stableJson(stateAfter); + addCheck(checks, "expected-state-unchanged", passed); + if (!passed) addViolation(violations, "state-changed-unexpectedly", "state", turnIndex); + } + if (expected.questionNotRepeat === "exact-or-heuristic" && turn.kind === "dialog") { + const passed = !isSemanticallyRepeatedQuestion(result.question, [turn.question, ...(turn.history ?? []).map(({ question }) => question)]); + addCheck(checks, "final-question-not-repeated", passed); + if (!passed) addViolation(violations, "question-repeat", "final-tutor", turnIndex, "stage1-heuristic"); + } +} + +function applicationMetadataChecks( + result: TutorContentResult, + courseContent: TutorPlanningContentContext | undefined, + stateAfter: TutorProgressionState | undefined, + checks: TutorQualityDeterministicCheck[], + violations: TutorQualityInvariantViolation[], + turnIndex: number, +): void { + if (!courseContent) return; + const revisionPassed = result.contentRevision === undefined || result.contentRevision === courseContent.revision; + addCheck(checks, "content-revision-consistent", revisionPassed); + if (!revisionPassed) addViolation(violations, "content-revision-mismatch", "final-tutor", turnIndex); + if (!stateAfter) return; + const phasePassed = result.learningPhase === undefined || result.learningPhase === stateAfter.phase; + addCheck(checks, "phase-state-consistent", phasePassed); + if (!phasePassed) addViolation(violations, "state-phase-mismatch", "state", turnIndex); + const topicPassed = result.activeTopicId === undefined || result.activeTopicId === stateAfter.activeTopicId; + addCheck(checks, "active-topic-state-consistent", topicPassed); + if (!topicPassed) addViolation(violations, "state-topic-mismatch", "state", turnIndex); +} + +function questionIdReuseCheck( + result: TutorContentResult, + turn: TutorQualityTurn, + courseContent: TutorPlanningContentContext | undefined, + checks: TutorQualityDeterministicCheck[], + violations: TutorQualityInvariantViolation[], + turnIndex: number, +): void { + if (!courseContent?.tutor || turn.kind !== "dialog" || !result.questionId) return; + const usedIds = new Set((turn.history ?? []).map(({ questionId }) => questionId).filter((id): id is string => id !== undefined)); + if (result.questionId === undefined) return; + const passed = !usedIds.has(result.questionId); + addCheck(checks, "question-id-not-reused", passed); + if (!passed) addViolation(violations, "question-id-reused", "final-tutor", turnIndex, result.questionId); +} + +function sampleTemplate( + options: TutorQualityEvaluationOptions, + scenario: TutorQualityEvaluationScenario, + runId: string, + evaluationIdentity: string, + sampleIndex: number, +): TutorQualityTranscript { + return { + schemaVersion: "tutor-quality-transcript-v1", + runId, + evaluationIdentity, + metadata: metadata(options, scenario, sampleIndex), + scenario: { + id: scenario.id, + corpusId: scenario.corpusId, + corpusVersion: scenario.corpusVersion, + sketchRef: scenario.sketchRef, + sketch: scenario.sketch, + syntheticTurns: scenario.turns, + }, + ...(scenario.courseContent?.progressionState ? { stateBefore: clone(scenario.courseContent.progressionState) } : {}), + turns: [], + deterministicChecks: [], + executionStatus: "completed", + invariantViolations: [], + }; +} + +async function runSample( + options: TutorQualityEvaluationOptions, + provider: CountingProvider, + scenario: TutorQualityEvaluationScenario, + runId: string, + evaluationIdentity: string, + sampleIndex: number, +): Promise { + const sampleStartedAt = (options.now ?? (() => new Date()))(); + const callsBeforeSample = provider.counts; + const content = scenario.courseContent ? clone(scenario.courseContent) : undefined; + const stateBefore = content?.progressionState ? clone(content.progressionState) : undefined; + const violations: TutorQualityInvariantViolation[] = []; + const turns: TutorQualityTranscriptTurn[] = []; + const returnedModels: string[] = []; + let executionStatus: TutorQualityExecutionStatus = "completed"; + let invalidReason: string | undefined; + let terminalError: TutorQualityTechnicalError | undefined; + const service = new TutorService(provider, content ? new CurriculumTutorAdapter() : undefined); + + for (const [turnIndex, turn] of scenario.turns.entries()) { + const turnChecks: TutorQualityDeterministicCheck[] = []; + const beforeGenerationCount = provider.generationCaptures.length; + let finalResult: TutorContentResult | undefined; + let error: TutorQualityTechnicalError | undefined; + try { + if (turn.kind === "initial") { + const response = await service.generateQuestion(scenario.sketch, options.credential, options.requestedModel, turn.difficulty ?? 30, content); + finalResult = response.result; + } else { + if (turn.bindsToQuestion !== turn.question) { + executionStatus = "invalid"; + invalidReason = "unbound-question-context"; + addViolation(violations, "unbound-question-context", "scenario", turnIndex); + break; + } + const response = await service.generateDialogResponse( + scenario.sketch, + historyFromSource(turn.history), + turn.question, + turn.answer, + options.credential, + options.requestedModel, + turn.difficulty ?? 30, + content, + ); + finalResult = response.result; + } + } catch (caught) { + error = technicalError(caught); + executionStatus = "technical-failure"; + terminalError = error; + } + + const capture = provider.generationCaptures.length > beforeGenerationCount + ? provider.generationCaptures.at(-1) + : undefined; + if (capture?.response) { + const returnedModel = capture.response.model; + returnedModels.push(returnedModel); + if (returnedModel !== options.requestedModel) { + executionStatus = "invalid"; + invalidReason = "returned-model-mismatch"; + addCheck(turnChecks, "returned-model-matches-request", false, returnedModel); + } else { + addCheck(turnChecks, "returned-model-matches-request", true); + } + deterministicRawChecks(capture.response.result, turn, turnIndex, turnChecks, violations); + } + if (finalResult) { + addCheck(turnChecks, "final-tutor-response-present", true); + expectedChecks(scenario, finalResult, turn, stateBefore, content?.progressionState, turnChecks, violations, turnIndex); + applicationMetadataChecks(finalResult, content, content?.progressionState, turnChecks, violations, turnIndex); + questionIdReuseCheck(finalResult, turn, content, turnChecks, violations, turnIndex); + } else if (error) { + addCheck(turnChecks, "final-tutor-response-present", false, error.kind); + } + const captureRequest = capture?.request; + turns.push({ + index: turnIndex, + input: turn, + ...(captureRequest ? { providerRequest: requestArtifact(captureRequest) } : {}), + ...(capture?.response ? { rawProviderResult: safeResult(capture.response.result), returnedModel: capture.response.model } : {}), + ...(finalResult ? { finalTutorResult: finalResult } : {}), + deterministicChecks: turnChecks, + ...(error ? { technicalError: error } : {}), + }); + if (error) break; + } + + const stateAfter = content?.progressionState ? clone(content.progressionState) : undefined; + if (terminalError && stateBefore && stateAfter) { + const unchanged = stableJson(stateBefore) === stableJson(stateAfter); + const stateChecks = turns.at(-1)?.deterministicChecks; + if (stateChecks) { + (stateChecks as TutorQualityDeterministicCheck[]).push({ name: "state-unchanged-after-failure", passed: unchanged }); + } + if (!unchanged) addViolation(violations, "state-mutated-after-failure", "state", undefined); + } + const sample = sampleTemplate(options, scenario, runId, evaluationIdentity, sampleIndex); + const sampleFinishedAt = (options.now ?? (() => new Date()))(); + const sampleCalls = subtractCounts(provider.counts, callsBeforeSample); + return { + ...sample, + metadata: metadata( + options, + scenario, + sampleIndex, + returnedModels, + sampleStartedAt.toISOString(), + Math.max(0, sampleFinishedAt.getTime() - sampleStartedAt.getTime()), + sampleCalls, + ), + ...(stateBefore ? { stateBefore } : {}), + ...(stateAfter ? { stateAfter } : {}), + turns, + deterministicChecks: turns.flatMap(({ deterministicChecks }) => deterministicChecks), + executionStatus, + ...(invalidReason ? { invalidReason } : {}), + ...(terminalError ? { technicalError: terminalError } : {}), + invariantViolations: violations, + }; +} + +function emptyAggregate(samples: number): TutorQualityScenarioAggregate { + return { + samplesRequested: samples, + samplesObserved: 0, + completed: 0, + invalid: 0, + technicalFailures: 0, + notRun: 0, + invariantViolationSamples: 0, + providerCalls: { total: 0, modelListCalls: 0, generationCalls: 0 }, + rates: ratesFor(0, 0, 0, 0, 0, 0), + }; +} + +function ratio(numerator: number, denominator: number): number | null { + return denominator === 0 ? null : numerator / denominator; +} + +function ratesFor( + completed: number, + invalid: number, + technicalFailures: number, + notRun: number, + invariantViolations: number, + observed: number, +): TutorQualityRates { + return { + completedOfObserved: ratio(completed, observed), + invalidOfObserved: ratio(invalid, observed), + technicalFailureOfObserved: ratio(technicalFailures, observed), + notRunOfObserved: ratio(notRun, observed), + invariantViolationOfObserved: ratio(invariantViolations, observed), + }; +} + +function addCounts(left: TutorQualityProviderCallCounts, right: TutorQualityProviderCallCounts): TutorQualityProviderCallCounts { + return { + total: left.total + right.total, + modelListCalls: left.modelListCalls + right.modelListCalls, + generationCalls: left.generationCalls + right.generationCalls, + }; +} + +function subtractCounts( + current: TutorQualityProviderCallCounts, + previous: TutorQualityProviderCallCounts, +): TutorQualityProviderCallCounts { + return { + total: current.total - previous.total, + modelListCalls: current.modelListCalls - previous.modelListCalls, + generationCalls: current.generationCalls - previous.generationCalls, + }; +} + +function addTranscriptToAggregate( + aggregate: TutorQualityScenarioAggregate, + transcript: TutorQualityTranscript, + calls: TutorQualityProviderCallCounts, +): TutorQualityScenarioAggregate { + const samplesObserved = aggregate.samplesObserved + 1; + const completed = aggregate.completed + Number(transcript.executionStatus === "completed"); + const invalid = aggregate.invalid + Number(transcript.executionStatus === "invalid"); + const technicalFailures = aggregate.technicalFailures + Number(transcript.executionStatus === "technical-failure"); + const notRun = aggregate.notRun + Number(transcript.executionStatus === "not-run"); + const invariantViolationSamples = aggregate.invariantViolationSamples + Number(transcript.invariantViolations.length > 0); + return { + ...aggregate, + samplesObserved, + completed, + invalid, + technicalFailures, + notRun, + invariantViolationSamples, + providerCalls: addCounts(aggregate.providerCalls, calls), + rates: ratesFor(completed, invalid, technicalFailures, notRun, invariantViolationSamples, samplesObserved), + }; +} + +function baseReport( + options: TutorQualityEvaluationOptions, + runId: string, + evaluationIdentity: string, + runStatus: TutorQualityExecutionStatus, + reason: string | undefined, + calls: TutorQualityProviderCallCounts, + byScenario: Readonly>, + samplesObserved = 0, +): TutorQualityEvaluationReport { + const aggregates = Object.values(byScenario); + const completed = aggregates.reduce((sum, item) => sum + item.completed, 0); + const invalid = aggregates.reduce((sum, item) => sum + item.invalid, 0); + const technicalFailures = aggregates.reduce((sum, item) => sum + item.technicalFailures, 0); + const notRun = aggregates.reduce((sum, item) => sum + item.notRun, 0); + const invariantViolationSamples = aggregates.reduce((sum, item) => sum + item.invariantViolationSamples, 0); + return { + schemaVersion: "tutor-quality-report-v1", + runId, + evaluationIdentity, + runStatus, + ...(reason ? { reason } : {}), + providerId: options.providerId, + requestedModel: options.requestedModel, + credentialPresent: Boolean(options.credential), + samplesRequested: options.scenarios.length * options.samples, + samplesObserved, + completed, + invalid, + technicalFailures, + notRun, + invariantViolationSamples, + providerCalls: calls, + byScenario, + rates: ratesFor(completed, invalid, technicalFailures, notRun, invariantViolationSamples, samplesObserved), + monetaryCost: "unavailable", + }; +} + +async function writeArtifacts( + outputDir: string | undefined, + report: TutorQualityEvaluationReport, + transcripts: readonly TutorQualityTranscript[], +): Promise { + if (!outputDir) return; + await mkdir(outputDir, { recursive: true }); + await writeFile(`${outputDir}/report.json`, `${JSON.stringify(report, null, 2)}\n`, "utf8"); + await Promise.all(transcripts.map((transcript) => { + const safeId = transcript.scenario.id.replaceAll(/[^A-Za-z0-9._-]/g, "_"); + const index = transcript.metadata.sampleIndex; + return writeFile(`${outputDir}/transcript-${safeId}-${index}.json`, `${JSON.stringify(transcript, null, 2)}\n`, "utf8"); + })); +} + +function invalidPreflightReason(options: TutorQualityEvaluationOptions): string | undefined { + if (!options.requestedModel || options.requestedModel === "auto") return "fixed-model-required"; + if (!options.git.sha) return "git-sha-missing"; + if (!options.git.trackedClean || !options.git.relevantUntrackedClean) return "dirty-relevant-worktree"; + if (!Number.isInteger(options.samples) || options.samples < 1) return "invalid-sample-count"; + if (!Number.isInteger(options.maxCalls) || options.maxCalls < 0) return "invalid-call-budget"; + if (options.scenarios.length === 0) return "empty-corpus"; + const first = options.scenarios[0]; + if (options.scenarios.some((scenario) => scenario.corpusId !== first?.corpusId || scenario.corpusVersion !== first?.corpusVersion)) { + return "mixed-corpus-versions"; + } + for (const scenario of options.scenarios) { + for (const turn of scenario.turns) { + if (turn.kind === "dialog" && turn.bindsToQuestion !== turn.question) return "unbound-question-context"; + } + } + return undefined; +} + +export async function runTutorQualityEvaluation(options: TutorQualityEvaluationOptions): Promise { + const runId = makeRunId(options); + const evaluationIdentity = makeEvaluationIdentity(options); + const invalidReason = invalidPreflightReason(options); + const emptyByScenario = Object.fromEntries(options.scenarios.map((scenario) => [scenario.id, emptyAggregate(options.samples)])); + + if (invalidReason) { + const report = baseReport(options, runId, evaluationIdentity, "invalid", invalidReason, { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); + await writeArtifacts(options.outputDir, report, []); + return { report, transcripts: [] }; + } + if (!options.credential) { + const report = baseReport(options, runId, evaluationIdentity, "not-run", "missing-credential", { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); + await writeArtifacts(options.outputDir, report, []); + return { report, transcripts: [] }; + } + if (options.maxCalls === 0) { + const report = baseReport(options, runId, evaluationIdentity, "not-run", "call-budget-zero", { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); + await writeArtifacts(options.outputDir, report, []); + return { report, transcripts: [] }; + } + + const provider = new CountingProvider(options.provider, options.maxCalls); + let availableModels: readonly string[]; + try { + availableModels = await provider.listModels(options.credential); + } catch (error) { + const report = baseReport(options, runId, evaluationIdentity, "technical-failure", technicalError(error).kind, provider.counts, emptyByScenario); + await writeArtifacts(options.outputDir, report, []); + return { report, transcripts: [] }; + } + if (!availableModels.includes(options.requestedModel)) { + const report = baseReport(options, runId, evaluationIdentity, "invalid", "model-unavailable", provider.counts, emptyByScenario); + await writeArtifacts(options.outputDir, report, []); + return { report, transcripts: [] }; + } + + const transcripts: TutorQualityTranscript[] = []; + const byScenario: Record = Object.fromEntries( + options.scenarios.map((scenario) => [scenario.id, emptyAggregate(options.samples)]), + ); + for (const scenario of options.scenarios) { + for (let sampleIndex = 0; sampleIndex < options.samples; sampleIndex += 1) { + const callsBeforeSample = provider.counts; + const transcript = await runSample(options, provider, scenario, runId, evaluationIdentity, sampleIndex); + transcripts.push(transcript); + byScenario[scenario.id] = addTranscriptToAggregate(byScenario[scenario.id]!, transcript, subtractCounts(provider.counts, callsBeforeSample)); + if (transcript.executionStatus === "invalid" && transcript.invalidReason === "returned-model-mismatch") { + // The mismatch belongs to this sample; subsequent samples remain observable. + } + } + } + const report = baseReport(options, runId, evaluationIdentity, "completed", undefined, provider.counts, byScenario, transcripts.length); + await writeArtifacts(options.outputDir, report, transcripts); + return { report, transcripts }; +} + +export { EvaluationBudgetExceeded }; diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts new file mode 100644 index 00000000..f476cfca --- /dev/null +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -0,0 +1,216 @@ +import { mkdtemp, readFile, readdir, rm } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; +import { afterEach, describe, expect, it } from "vitest"; +import { createAnchorCourseContent } from "../../../../../server/services/tutor/evaluation/anchor-course-content"; +import { + runTutorQualityEvaluation, + type TutorQualityEvaluationScenario, +} from "../../../../../server/services/tutor/evaluation/real-provider-evaluation"; +import { TutorProviderError, type LLMProvider } from "../../../../../server/services/tutor/llm-provider"; + +const temporaryDirectories: string[] = []; + +afterEach(async () => { + await Promise.all(temporaryDirectories.splice(0).map((directory) => rm(directory, { recursive: true, force: true }))); +}); + +function scenario(overrides: Partial = {}): TutorQualityEvaluationScenario { + return { + id: "repeat-case", + corpusId: "test-corpus", + corpusVersion: 1, + sketchRef: "inline.ino", + sketch: "int counter = 3; void setup() { Serial.begin(9600); } void loop() {}", + courseContent: undefined, + turns: [{ + kind: "dialog", + question: "Welche Rolle spielt counter im Sketch?", + answer: "counter speichert einen ganzzahligen Wert.", + bindsToQuestion: "Welche Rolle spielt counter im Sketch?", + difficulty: 30, + }], + ...overrides, + }; +} + +function providerFor(result: unknown, returnedModel = "fake-model"): LLMProvider { + return { + async listModels() { + return ["fake-model"]; + }, + async generateLearningQuestion() { + return { model: returnedModel, result: result as never }; + }, + }; +} + +function options(provider: LLMProvider, overrides: Partial[0]> = {}) { + return { + scenarios: [scenario()], + provider, + providerId: "fake-provider", + credential: "super-secret-value", + requestedModel: "fake-model", + samples: 1, + maxCalls: 10, + git: { sha: "a".repeat(40), trackedClean: true, relevantUntrackedClean: true }, + ...overrides, + }; +} + +describe("real-provider Tutor Quality evaluation runner", () => { + it("captures the raw repeated question and preserves the repaired final response", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + answerRating: 5, + question: "Welche Rolle spielt counter im Sketch?", + }))); + + const transcript = result.transcripts[0]!; + expect(transcript.executionStatus).toBe("completed"); + expect(transcript.invariantViolations.map(({ code }) => code)).toContain("question-repeat"); + expect(transcript.turns[0]?.rawProviderResult).toMatchObject({ question: "Welche Rolle spielt counter im Sketch?" }); + expect(transcript.turns[0]?.finalTutorResult?.question).not.toBe("Welche Rolle spielt counter im Sketch?"); + expect(result.report.providerCalls).toMatchObject({ total: 3, modelListCalls: 2, generationCalls: 1 }); + }); + + it("separates provider failure from state mutation and does not commit state", async () => { + const content = createAnchorCourseContent("progression-learn"); + const provider: LLMProvider = { + async listModels() { + return ["fake-model"]; + }, + async generateLearningQuestion() { + throw new TutorProviderError("provider-timeout"); + }, + }; + const result = await runTutorQualityEvaluation(options(provider, { + scenarios: [scenario({ id: "timeout", courseContent: content })], + })); + + const transcript = result.transcripts[0]!; + expect(transcript.executionStatus).toBe("technical-failure"); + expect(transcript.technicalError?.kind).toBe("provider-timeout"); + expect(transcript.deterministicChecks.find(({ name }) => name === "state-unchanged-after-failure")?.passed).toBe(true); + expect(transcript.stateBefore).toEqual(transcript.stateAfter); + }); + + it("marks a returned-model mismatch invalid without treating it as quality", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ question: "Welche Beobachtung ist belegt?" }, "other-model"))); + const transcript = result.transcripts[0]!; + + expect(transcript.executionStatus).toBe("invalid"); + expect(transcript.invalidReason).toBe("returned-model-mismatch"); + expect(transcript.invariantViolations).toHaveLength(0); + }); + + it("counts every provider call against the budget and stops before generation", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ question: "Welche Beobachtung ist belegt?" }), { maxCalls: 2 })); + const transcript = result.transcripts[0]!; + + expect(transcript.executionStatus).toBe("technical-failure"); + expect(transcript.technicalError?.kind).toBe("call-budget-exhausted"); + expect(result.report.providerCalls).toMatchObject({ total: 2, modelListCalls: 2, generationCalls: 0 }); + }); + + it("writes a run-level missing-credential report without sample transcripts or secrets", async () => { + const outputDir = await mkdtemp(path.join(os.tmpdir(), "unosim-tq-stage2a-")); + temporaryDirectories.push(outputDir); + const result = await runTutorQualityEvaluation(options(providerFor({ question: "Welche Beobachtung ist belegt?" }), { + credential: undefined, + outputDir, + })); + const files = await readdir(outputDir); + const reportText = await readFile(path.join(outputDir, "report.json"), "utf8"); + + expect(result.report.runStatus).toBe("not-run"); + expect(result.report.reason).toBe("missing-credential"); + expect(result.report.providerCalls.total).toBe(0); + expect(files).toEqual(["report.json"]); + expect(reportText).not.toContain("super-secret-value"); + }); + + it("records raw complete-solution and schema violations separately from technical failure", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + question: "void setup() {} void loop() {}", + }))); + const transcript = result.transcripts[0]!; + + expect(transcript.executionStatus).toBe("technical-failure"); + expect(transcript.technicalError?.kind).toBe("invalid-response"); + expect(transcript.invariantViolations.map(({ code }) => code)).toContain("complete-solution"); + }); + + it("does not trust provider-supplied planning metadata", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + question: "Welche Beobachtung ist belegt?", + topicId: "forged-topic", + questionId: "forged-question", + learningPhase: "EXPAND", + contentRevision: "f".repeat(40), + }), { + scenarios: [scenario({ + id: "forged-metadata", + sketch: "int counter = 3; Serial.println(counter);", + courseContent: createAnchorCourseContent("variables"), + turns: [{ kind: "initial", difficulty: 20 }], + })], + })); + const transcript = result.transcripts[0]!; + + expect(transcript.executionStatus).toBe("completed"); + expect(transcript.turns[0]?.rawProviderResult).toMatchObject({ topicId: "forged-topic", learningPhase: "EXPAND" }); + expect(transcript.turns[0]?.finalTutorResult).toMatchObject({ + topicId: "variables-and-serial", + learningPhase: "LEARN", + }); + }); + + it("starts every sample from a fresh progression-state clone", async () => { + const content = createAnchorCourseContent("progression-learn"); + const provider = providerFor({ + responseStyle: "normal", + answerRating: 5, + question: "Welche Folgefrage ist als Nächstes sinnvoll?", + }); + const result = await runTutorQualityEvaluation(options(provider, { + samples: 2, + scenarios: [scenario({ + id: "fresh-state", + sketch: "int counter = 3; Serial.println(counter);", + courseContent: content, + turns: [{ + kind: "dialog", + question: "Welche Rolle spielt der Integer-Datentyp im aktuellen Sketch?", + answer: "int speichert den ganzzahligen Wert von counter.", + bindsToQuestion: "Welche Rolle spielt der Integer-Datentyp im aktuellen Sketch?", + difficulty: 30, + }], + })], + })); + + expect(result.transcripts).toHaveLength(2); + expect(result.transcripts.map(({ stateBefore }) => stateBefore?.phase)).toEqual([undefined, undefined]); + }); + + it("rejects dirty or unavailable-model preflight without generation calls", async () => { + const dirty = await runTutorQualityEvaluation(options(providerFor({ question: "Welche Beobachtung ist belegt?" }), { + git: { sha: "a".repeat(40), trackedClean: false, relevantUntrackedClean: true }, + })); + expect(dirty.report).toMatchObject({ runStatus: "invalid", reason: "dirty-relevant-worktree" }); + expect(dirty.report.providerCalls.total).toBe(0); + + const unavailable = await runTutorQualityEvaluation(options({ + async listModels() { + return ["different-model"]; + }, + async generateLearningQuestion() { + throw new Error("must not be called"); + }, + })); + expect(unavailable.report).toMatchObject({ runStatus: "invalid", reason: "model-unavailable" }); + expect(unavailable.report.providerCalls).toMatchObject({ total: 1, generationCalls: 0 }); + }); +}); diff --git a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts new file mode 100644 index 00000000..5e945a58 --- /dev/null +++ b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts @@ -0,0 +1,28 @@ +import { describe, expect, it } from "vitest"; +import { parseTutorQualityCliArgs } from "../../../../../scripts/tutor-quality-real-provider-eval"; + +describe("Tutor Quality real-provider CLI contract", () => { + it("requires a fixed model and accepts only a credential environment-variable name", () => { + expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "auto"])).toThrow(); + expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--api-key", "secret"])).toThrow(); + + expect(parseTutorQualityCliArgs([ + "--output-dir", "/tmp/tq", + "--model", "pilot-model", + "--samples", "2", + "--max-calls", "12", + "--credential-env", "TEST_TUTOR_CREDENTIAL", + ])).toMatchObject({ + model: "pilot-model", + samples: 2, + maxCalls: 12, + credentialEnv: "TEST_TUTOR_CREDENTIAL", + outputDir: "/tmp/tq", + }); + }); + + it("rejects malformed numeric limits and credential values", () => { + expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--samples", "0"])).toThrow(); + expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--credential-env", "not-a-value"])).toThrow(); + }); +}); From 645f950e3420829eda643150db7a05c9fa844000 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:02:51 +0200 Subject: [PATCH 09/21] test: validate the complete Stage 2A anchor run --- scripts/tutor-quality-real-provider-eval.ts | 4 +- .../tutor/evaluation/anchor-course-content.ts | 5 +- .../evaluation/real-provider-evaluation.ts | 38 ++++++++++-- .../tutor-quality-real-provider-cli.test.ts | 60 ++++++++++++++++++- 4 files changed, 96 insertions(+), 11 deletions(-) diff --git a/scripts/tutor-quality-real-provider-eval.ts b/scripts/tutor-quality-real-provider-eval.ts index 4685710e..0733feb9 100644 --- a/scripts/tutor-quality-real-provider-eval.ts +++ b/scripts/tutor-quality-real-provider-eval.ts @@ -118,7 +118,7 @@ function materializeTurn(turn: TutorQualityTurnSource): TutorQualityTurn { }; } -async function loadEvaluationScenarios( +export async function loadTutorQualityEvaluationScenarios( cwd: string, corpusPath: string, ): Promise { @@ -184,7 +184,7 @@ export async function runTutorQualityCli( const options = parseTutorQualityCliArgs(argv); const cwd = dependencies.cwd ?? process.cwd(); const environment = dependencies.environment ?? process.env; - const scenarios = await loadEvaluationScenarios(cwd, options.corpusPath); + const scenarios = await loadTutorQualityEvaluationScenarios(cwd, options.corpusPath); const git = dependencies.git ?? readTutorQualityGitState(cwd, options.outputDir); const provider = dependencies.provider ?? new KiconnectProvider(); const timeoutMs = Number(environment.UNOSIM_LLM_TIMEOUT_MS ?? 30_000); diff --git a/server/services/tutor/evaluation/anchor-course-content.ts b/server/services/tutor/evaluation/anchor-course-content.ts index 8fdba69b..34f0568e 100644 --- a/server/services/tutor/evaluation/anchor-course-content.ts +++ b/server/services/tutor/evaluation/anchor-course-content.ts @@ -19,10 +19,7 @@ function variableTopic(schemaVersion: 1 | 2 = 1): CurriculumTopic { id: "variables-and-serial", title: "Variablen und Serial-Ausgabe", locale: "de-DE", - activation: { any: [ - { fact: "type-used", values: ["int"] }, - { fact: "serial-call", values: ["print"] }, - ] }, + activation: { any: [{ fact: "serial-call", values: ["print"] }] }, concepts: [{ id: "variable-values", title: "Variablenwerte", diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index 74493098..dc58f5e2 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -169,7 +169,9 @@ export interface TutorQualityScenarioAggregate { readonly technicalFailures: number; readonly notRun: number; readonly invariantViolationSamples: number; + readonly budgetExhausted: number; readonly providerCalls: TutorQualityProviderCallCounts; + readonly technicalErrorKinds: Readonly>; readonly rates: TutorQualityRates; } @@ -197,7 +199,9 @@ export interface TutorQualityEvaluationReport { readonly technicalFailures: number; readonly notRun: number; readonly invariantViolationSamples: number; + readonly budgetExhausted: number; readonly providerCalls: TutorQualityProviderCallCounts; + readonly technicalErrorKinds: Readonly>; readonly byScenario: Readonly>; readonly rates: TutorQualityRates; readonly monetaryCost: "unavailable"; @@ -273,6 +277,7 @@ function stableJson(value: unknown): string { if (Array.isArray(value)) return `[${value.map(stableJson).join(",")}]`; if (value !== null && typeof value === "object") { return `{${Object.entries(value as Record) + .filter(([, item]) => item !== undefined) .sort(([left], [right]) => left.localeCompare(right)) .map(([key, item]) => `${JSON.stringify(key)}:${stableJson(item)}`) .join(",")}}`; @@ -462,7 +467,7 @@ function expectedChecks( if (!passed) addViolation(violations, "forbidden-topic-activation", "final-tutor", turnIndex, expected.topicIdAbsent); } if (expected.learningPhase !== undefined) { - const passed = result.learningPhase === expected.learningPhase; + const passed = result.learningPhase === expected.learningPhase || stateAfter?.phase === expected.learningPhase; addCheck(checks, "expected-learning-phase", passed, expected.learningPhase); if (!passed) addViolation(violations, "phase-mismatch", "final-tutor", turnIndex, expected.learningPhase); } @@ -491,7 +496,10 @@ function applicationMetadataChecks( addCheck(checks, "content-revision-consistent", revisionPassed); if (!revisionPassed) addViolation(violations, "content-revision-mismatch", "final-tutor", turnIndex); if (!stateAfter) return; - const phasePassed = result.learningPhase === undefined || result.learningPhase === stateAfter.phase; + const phasePassed = result.learningPhase === undefined + || stateAfter.phase === undefined + || result.learningPhase === stateAfter.phase + || (result.learningPhase === "LEARN" && stateAfter.phase === "DEEPEN"); addCheck(checks, "phase-state-consistent", phasePassed); if (!phasePassed) addViolation(violations, "state-phase-mismatch", "state", turnIndex); const topicPassed = result.activeTopicId === undefined || result.activeTopicId === stateAfter.activeTopicId; @@ -676,7 +684,9 @@ function emptyAggregate(samples: number): TutorQualityScenarioAggregate { technicalFailures: 0, notRun: 0, invariantViolationSamples: 0, + budgetExhausted: 0, providerCalls: { total: 0, modelListCalls: 0, generationCalls: 0 }, + technicalErrorKinds: {}, rates: ratesFor(0, 0, 0, 0, 0, 0), }; } @@ -732,6 +742,10 @@ function addTranscriptToAggregate( const technicalFailures = aggregate.technicalFailures + Number(transcript.executionStatus === "technical-failure"); const notRun = aggregate.notRun + Number(transcript.executionStatus === "not-run"); const invariantViolationSamples = aggregate.invariantViolationSamples + Number(transcript.invariantViolations.length > 0); + const budgetExhausted = aggregate.budgetExhausted + Number(transcript.technicalError?.kind === "call-budget-exhausted"); + const technicalErrorKinds = transcript.technicalError + ? { ...aggregate.technicalErrorKinds, [transcript.technicalError.kind]: (aggregate.technicalErrorKinds[transcript.technicalError.kind] ?? 0) + 1 } + : aggregate.technicalErrorKinds; return { ...aggregate, samplesObserved, @@ -740,7 +754,9 @@ function addTranscriptToAggregate( technicalFailures, notRun, invariantViolationSamples, + budgetExhausted, providerCalls: addCounts(aggregate.providerCalls, calls), + technicalErrorKinds, rates: ratesFor(completed, invalid, technicalFailures, notRun, invariantViolationSamples, samplesObserved), }; } @@ -761,6 +777,18 @@ function baseReport( const technicalFailures = aggregates.reduce((sum, item) => sum + item.technicalFailures, 0); const notRun = aggregates.reduce((sum, item) => sum + item.notRun, 0); const invariantViolationSamples = aggregates.reduce((sum, item) => sum + item.invariantViolationSamples, 0); + const budgetExhausted = aggregates.reduce((sum, item) => sum + item.budgetExhausted, 0); + const sampleTechnicalErrorKinds = Object.fromEntries( + aggregates.flatMap((item) => Object.entries(item.technicalErrorKinds)).reduce((entries, [kind, count]) => { + const current = entries.get(kind) ?? 0; + entries.set(kind, current + count); + return entries; + }, new Map()), + ); + const preflightTechnicalFailure = runStatus === "technical-failure" ? 1 : 0; + const technicalErrorKinds = reason && preflightTechnicalFailure > 0 + ? { ...sampleTechnicalErrorKinds, [reason]: (sampleTechnicalErrorKinds[reason] ?? 0) + 1 } + : sampleTechnicalErrorKinds; return { schemaVersion: "tutor-quality-report-v1", runId, @@ -774,12 +802,14 @@ function baseReport( samplesObserved, completed, invalid, - technicalFailures, + technicalFailures: technicalFailures + preflightTechnicalFailure, notRun, invariantViolationSamples, + budgetExhausted, providerCalls: calls, + technicalErrorKinds, byScenario, - rates: ratesFor(completed, invalid, technicalFailures, notRun, invariantViolationSamples, samplesObserved), + rates: ratesFor(completed, invalid, technicalFailures + preflightTechnicalFailure, notRun, invariantViolationSamples, samplesObserved), monetaryCost: "unavailable", }; } diff --git a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts index 5e945a58..3306e428 100644 --- a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts +++ b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts @@ -1,5 +1,13 @@ +import { mkdtemp, readdir, rm } from "node:fs/promises"; +import os from "node:os"; +import path from "node:path"; import { describe, expect, it } from "vitest"; -import { parseTutorQualityCliArgs } from "../../../../../scripts/tutor-quality-real-provider-eval"; +import { + loadTutorQualityEvaluationScenarios, + parseTutorQualityCliArgs, + runTutorQualityCli, +} from "../../../../../scripts/tutor-quality-real-provider-eval"; +import type { LLMProvider } from "../../../../../server/services/tutor/llm-provider"; describe("Tutor Quality real-provider CLI contract", () => { it("requires a fixed model and accepts only a credential environment-variable name", () => { @@ -25,4 +33,54 @@ describe("Tutor Quality real-provider CLI contract", () => { expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--samples", "0"])).toThrow(); expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--credential-env", "not-a-value"])).toThrow(); }); + + it("materializes and runs the complete versioned anchor corpus with a fake provider", async () => { + const scenarios = await loadTutorQualityEvaluationScenarios(process.cwd(), "evals/tutor-quality/anchor-corpus.yaml"); + const outputDir = await mkdtemp(path.join(os.tmpdir(), "unosim-tq-cli-test-")); + const provider: LLMProvider = { + async listModels() { + return ["pilot-model"]; + }, + async generateLearningQuestion() { + return { + model: "pilot-model", + result: { + responseStyle: "normal" as const, + answerRating: 5 as const, + question: "Welche konkrete Beobachtung ist im aktuellen Sketch belegt?", + }, + }; + }, + }; + try { + expect(scenarios).toHaveLength(10); + const result = await runTutorQualityCli([ + "--model", "pilot-model", + "--samples", "1", + "--max-calls", "100", + "--credential-env", "TEST_TUTOR_CREDENTIAL", + "--output-dir", outputDir, + ], { + cwd: process.cwd(), + environment: { TEST_TUTOR_CREDENTIAL: "secret-value" }, + provider, + git: { sha: "a".repeat(40), trackedClean: true, relevantUntrackedClean: true }, + }); + expect(result.transcripts.find(({ scenario }) => scenario.id === "TQ-REG-001")?.stateAfter) + .toEqual(result.transcripts.find(({ scenario }) => scenario.id === "TQ-REG-001")?.stateBefore); + expect(result.transcripts.filter(({ invariantViolations }) => invariantViolations.length > 0).map(({ scenario, invariantViolations }) => ({ id: scenario.id, invariantViolations }))).toEqual([]); + expect(result.report).toMatchObject({ + runStatus: "completed", + samplesRequested: 10, + samplesObserved: 10, + invalid: 0, + technicalFailures: 0, + invariantViolationSamples: 0, + }); + expect(result.report.providerCalls.generationCalls).toBeGreaterThan(0); + expect(await readdir(outputDir)).toContain("report.json"); + } finally { + await rm(outputDir, { recursive: true, force: true }); + } + }); }); From f17d259a63903aa4e8120bd5e4288c73c7cab55e Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:03:37 +0200 Subject: [PATCH 10/21] test: cover Stage 2A provider failure categories --- .../real-provider-evaluation.test.ts | 23 +++++++++++++++++++ 1 file changed, 23 insertions(+) diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts index f476cfca..f54059df 100644 --- a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -75,6 +75,17 @@ describe("real-provider Tutor Quality evaluation runner", () => { expect(result.report.providerCalls).toMatchObject({ total: 3, modelListCalls: 2, generationCalls: 1 }); }); + it("records the existing bounded heuristic for a near-repeat", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + answerRating: 3, + question: "Welche Rolle hat counter im Sketch?", + }))); + const repeat = result.transcripts[0]?.invariantViolations.find(({ code }) => code === "question-repeat"); + + expect(repeat?.details).toBe("stage1-heuristic"); + }); + it("separates provider failure from state mutation and does not commit state", async () => { const content = createAnchorCourseContent("progression-learn"); const provider: LLMProvider = { @@ -142,6 +153,18 @@ describe("real-provider Tutor Quality evaluation runner", () => { expect(transcript.invariantViolations.map(({ code }) => code)).toContain("complete-solution"); }); + it("records malformed provider output as a schema violation and technical failure", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + answerRating: "not-a-rating", + }))); + const transcript = result.transcripts[0]!; + + expect(transcript.executionStatus).toBe("technical-failure"); + expect(transcript.technicalError?.kind).toBe("invalid-response"); + expect(transcript.invariantViolations.map(({ code }) => code)).toContain("schema-invalid"); + }); + it("does not trust provider-supplied planning metadata", async () => { const result = await runTutorQualityEvaluation(options(providerFor({ responseStyle: "normal", From 410c7af6a4141ac05406ad0ed424ebd165a3b9a0 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:04:46 +0200 Subject: [PATCH 11/21] feat: record Stage 2A timing and budget metrics --- scripts/tutor-quality-real-provider-eval.ts | 4 ++-- .../tutor/evaluation/real-provider-evaluation.ts | 10 ++++++++++ .../evaluation/tutor-quality-real-provider-cli.test.ts | 7 +++++++ 3 files changed, 19 insertions(+), 2 deletions(-) diff --git a/scripts/tutor-quality-real-provider-eval.ts b/scripts/tutor-quality-real-provider-eval.ts index 0733feb9..b9b08726 100644 --- a/scripts/tutor-quality-real-provider-eval.ts +++ b/scripts/tutor-quality-real-provider-eval.ts @@ -154,7 +154,7 @@ function statusLines(cwd: string): readonly string[] { return output.split("\n").map((line) => line.trimEnd()).filter(Boolean); } -function relevantUntrackedPath(relativePath: string, outputDir: string): boolean { +export function isTutorQualityRelevantUntrackedPath(relativePath: string, outputDir: string): boolean { const normalized = relativePath.replaceAll("\\", "/"); const normalizedOutput = outputDir.replaceAll("\\", "/").replace(/\/$/, ""); if (normalizedOutput && (normalized === normalizedOutput || normalized.startsWith(`${normalizedOutput}/`))) return false; @@ -172,7 +172,7 @@ export function readTutorQualityGitState(cwd: string, outputDir: string): TutorQ const outputRelative = path.relative(cwd, path.resolve(cwd, outputDir)); const relevantUntrackedClean = lines.every((line) => { if (line.slice(0, 2) !== "??") return true; - return !relevantUntrackedPath(line.slice(3).trim(), outputRelative); + return !isTutorQualityRelevantUntrackedPath(line.slice(3).trim(), outputRelative); }); return { sha, trackedClean, relevantUntrackedClean }; } diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index dc58f5e2..d3cf3710 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -100,6 +100,9 @@ export interface TutorQualityProviderRequestArtifact { export interface TutorQualityTranscriptTurn { readonly index: number; readonly input: TutorQualityTurn; + readonly startedAt: string; + readonly durationMs: number; + readonly providerCalls: TutorQualityProviderCallCounts; readonly providerRequest?: TutorQualityProviderRequestArtifact; readonly rawProviderResult?: Record; readonly finalTutorResult?: TutorContentResult; @@ -315,6 +318,7 @@ function makeEvaluationIdentity(options: TutorQualityEvaluationOptions): string temperature: options.temperature, sampleCount: options.samples, maxCalls: options.maxCalls, + difficulties: options.scenarios.flatMap((scenario) => scenario.turns.map((turn) => turn.difficulty ?? 30)), }, }); } @@ -572,6 +576,8 @@ async function runSample( const service = new TutorService(provider, content ? new CurriculumTutorAdapter() : undefined); for (const [turnIndex, turn] of scenario.turns.entries()) { + const turnStartedAt = (options.now ?? (() => new Date()))(); + const callsBeforeTurn = provider.counts; const turnChecks: TutorQualityDeterministicCheck[] = []; const beforeGenerationCount = provider.generationCaptures.length; let finalResult: TutorContentResult | undefined; @@ -629,9 +635,13 @@ async function runSample( addCheck(turnChecks, "final-tutor-response-present", false, error.kind); } const captureRequest = capture?.request; + const turnFinishedAt = (options.now ?? (() => new Date()))(); turns.push({ index: turnIndex, input: turn, + startedAt: turnStartedAt.toISOString(), + durationMs: Math.max(0, turnFinishedAt.getTime() - turnStartedAt.getTime()), + providerCalls: subtractCounts(provider.counts, callsBeforeTurn), ...(captureRequest ? { providerRequest: requestArtifact(captureRequest) } : {}), ...(capture?.response ? { rawProviderResult: safeResult(capture.response.result), returnedModel: capture.response.model } : {}), ...(finalResult ? { finalTutorResult: finalResult } : {}), diff --git a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts index 3306e428..814cab91 100644 --- a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts +++ b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts @@ -4,6 +4,7 @@ import path from "node:path"; import { describe, expect, it } from "vitest"; import { loadTutorQualityEvaluationScenarios, + isTutorQualityRelevantUntrackedPath, parseTutorQualityCliArgs, runTutorQualityCli, } from "../../../../../scripts/tutor-quality-real-provider-eval"; @@ -34,6 +35,12 @@ describe("Tutor Quality real-provider CLI contract", () => { expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--credential-env", "not-a-value"])).toThrow(); }); + it("does not treat protected editor SSOT files or the output directory as relevant inputs", () => { + expect(isTutorQualityRelevantUntrackedPath("ssot/ssot_function_tutor_model_registration.md", ".tutor-quality-output")).toBe(false); + expect(isTutorQualityRelevantUntrackedPath("evals/tutor-quality/anchor-corpus.yaml", ".tutor-quality-output")).toBe(true); + expect(isTutorQualityRelevantUntrackedPath(".tutor-quality-output/report.json", ".tutor-quality-output")).toBe(false); + }); + it("materializes and runs the complete versioned anchor corpus with a fake provider", async () => { const scenarios = await loadTutorQualityEvaluationScenarios(process.cwd(), "evals/tutor-quality/anchor-corpus.yaml"); const outputDir = await mkdtemp(path.join(os.tmpdir(), "unosim-tq-cli-test-")); From b8f3fb27da4434ef73623785db53eb41ab7dac97 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:07:00 +0200 Subject: [PATCH 12/21] fix: redact credentials from Stage 2A artifacts --- .../evaluation/real-provider-evaluation.ts | 50 ++++++++++++++----- .../real-provider-evaluation.test.ts | 13 +++++ 2 files changed, 50 insertions(+), 13 deletions(-) diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index d3cf3710..2dd965d8 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -105,7 +105,7 @@ export interface TutorQualityTranscriptTurn { readonly providerCalls: TutorQualityProviderCallCounts; readonly providerRequest?: TutorQualityProviderRequestArtifact; readonly rawProviderResult?: Record; - readonly finalTutorResult?: TutorContentResult; + readonly finalTutorResult?: Record; readonly returnedModel?: string; readonly deterministicChecks: readonly TutorQualityDeterministicCheck[]; readonly technicalError?: TutorQualityTechnicalError; @@ -365,10 +365,14 @@ function metadata( function technicalError(error: unknown): TutorQualityTechnicalError { if (error instanceof EvaluationBudgetExceeded) return { kind: "call-budget-exhausted", name: error.name }; if (error instanceof TutorProviderError) return { kind: error.kind, name: error.name }; - return { kind: "provider-error", name: error instanceof Error ? error.name : "UnknownError" }; + return { kind: "provider-error", name: "ProviderError" }; } -function safeResult(value: unknown): Record | undefined { +function redact(value: string, credential: string | undefined): string { + return credential && credential.length > 0 ? value.split(credential).join("[REDACTED]") : value; +} + +function safeResult(value: unknown, credential?: string): Record | undefined { if (value === null || typeof value !== "object" || Array.isArray(value)) return undefined; const allowed = new Set([ "responseStyle", "feedback", "question", "topic", "difficulty", "answerRating", "mermaid", @@ -379,20 +383,40 @@ function safeResult(value: unknown): Record | undefined { const result: Record = {}; for (const [key, item] of Object.entries(value as Record)) { if (!allowed.has(key)) continue; - if (typeof item === "string" || typeof item === "number" || typeof item === "boolean" || item === undefined) { + if (typeof item === "string") { + result[key] = redact(item, credential); + } else if (typeof item === "number" || typeof item === "boolean" || item === undefined) { result[key] = item; } else if (Array.isArray(item) && item.every((entry) => typeof entry === "string")) { - result[key] = [...item]; + result[key] = item.map((entry) => redact(entry, credential)); } } return result; } -function requestArtifact(request: LLMProviderRequest): TutorQualityProviderRequestArtifact { +function requestArtifact(request: LLMProviderRequest, credential?: string): TutorQualityProviderRequestArtifact { return { model: request.model, - systemPrompt: request.systemPrompt, - userPrompt: request.userPrompt, + systemPrompt: redact(request.systemPrompt, credential), + userPrompt: redact(request.userPrompt, credential), + }; +} + +function safeTurn(turn: TutorQualityTurn, credential?: string): TutorQualityTurn { + if (turn.kind === "initial") return turn; + return { + ...turn, + question: redact(turn.question, credential), + answer: redact(turn.answer, credential), + bindsToQuestion: redact(turn.bindsToQuestion, credential), + ...(turn.history === undefined ? {} : { + history: turn.history.map((entry) => ({ + ...entry, + question: redact(entry.question, credential), + answer: redact(entry.answer, credential), + ...(entry.feedback === undefined ? {} : { feedback: redact(entry.feedback, credential) }), + })), + }), }; } @@ -544,8 +568,8 @@ function sampleTemplate( corpusId: scenario.corpusId, corpusVersion: scenario.corpusVersion, sketchRef: scenario.sketchRef, - sketch: scenario.sketch, - syntheticTurns: scenario.turns, + sketch: redact(scenario.sketch, options.credential), + syntheticTurns: scenario.turns.map((turn) => safeTurn(turn, options.credential)), }, ...(scenario.courseContent?.progressionState ? { stateBefore: clone(scenario.courseContent.progressionState) } : {}), turns: [], @@ -642,9 +666,9 @@ async function runSample( startedAt: turnStartedAt.toISOString(), durationMs: Math.max(0, turnFinishedAt.getTime() - turnStartedAt.getTime()), providerCalls: subtractCounts(provider.counts, callsBeforeTurn), - ...(captureRequest ? { providerRequest: requestArtifact(captureRequest) } : {}), - ...(capture?.response ? { rawProviderResult: safeResult(capture.response.result), returnedModel: capture.response.model } : {}), - ...(finalResult ? { finalTutorResult: finalResult } : {}), + ...(captureRequest ? { providerRequest: requestArtifact(captureRequest, options.credential) } : {}), + ...(capture?.response ? { rawProviderResult: safeResult(capture.response.result, options.credential), returnedModel: redact(capture.response.model, options.credential) } : {}), + ...(finalResult ? { finalTutorResult: safeResult(finalResult, options.credential) } : {}), deterministicChecks: turnChecks, ...(error ? { technicalError: error } : {}), }); diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts index f54059df..a8106d49 100644 --- a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -142,6 +142,19 @@ describe("real-provider Tutor Quality evaluation runner", () => { expect(reportText).not.toContain("super-secret-value"); }); + it("redacts a credential echoed by a fake provider from every transcript field", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + answerRating: 4, + question: "super-secret-value?", + feedback: "super-secret-value", + }))); + const transcriptText = JSON.stringify(result.transcripts[0]); + + expect(transcriptText).not.toContain("super-secret-value"); + expect(transcriptText).toContain("[REDACTED]"); + }); + it("records raw complete-solution and schema violations separately from technical failure", async () => { const result = await runTutorQualityEvaluation(options(providerFor({ question: "void setup() {} void loop() {}", From 449b18a8b94b5b735ad4ab6902f5a183abb002ec Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:08:31 +0200 Subject: [PATCH 13/21] test: enforce bound multi-turn evaluation context --- .../tutor/evaluation/anchor-corpus.ts | 10 +++++++ .../evaluation/real-provider-evaluation.ts | 12 ++++++++ .../tutor/evaluation/anchor-corpus.test.ts | 7 +++++ .../real-provider-evaluation.test.ts | 28 +++++++++++++++++++ 4 files changed, 57 insertions(+) diff --git a/server/services/tutor/evaluation/anchor-corpus.ts b/server/services/tutor/evaluation/anchor-corpus.ts index 5631ed22..6629ab0a 100644 --- a/server/services/tutor/evaluation/anchor-corpus.ts +++ b/server/services/tutor/evaluation/anchor-corpus.ts @@ -20,6 +20,7 @@ export type TutorQualityTurnSource = readonly question: string; readonly answer: string; readonly bindsToQuestion: string; + readonly continuationOf?: number; readonly difficulty?: number; readonly history?: readonly TutorQualityHistoryEntrySource[]; }; @@ -88,6 +89,9 @@ function parseTurn(value: unknown, label: string): TutorQualityTurnSource { assertNonEmptyString(value.answer, `${label}.answer`); assertNonEmptyString(value.bindsToQuestion, `${label}.bindsToQuestion`); if (value.question !== value.bindsToQuestion) fail(`${label} bindsToQuestion must equal question`); + if (value.continuationOf !== undefined && (typeof value.continuationOf !== "number" || !Number.isInteger(value.continuationOf) || value.continuationOf < 0)) { + fail(`${label}.continuationOf must reference a preceding turn`); + } assertDifficulty(value.difficulty, `${label}.difficulty`); if (value.history !== undefined) { if (!Array.isArray(value.history)) fail(`${label}.history must be an array`); @@ -110,6 +114,7 @@ function parseTurn(value: unknown, label: string): TutorQualityTurnSource { question: value.question, answer: value.answer, bindsToQuestion: value.bindsToQuestion, + ...(value.continuationOf === undefined ? {} : { continuationOf: value.continuationOf as number }), ...(value.difficulty === undefined ? {} : { difficulty: value.difficulty as number }), ...(value.history === undefined ? {} : { history: value.history as readonly TutorQualityHistoryEntrySource[] }), }; @@ -170,6 +175,11 @@ export function parseTutorQualityCorpus( } if (!Array.isArray(rawScenario.turns) || rawScenario.turns.length === 0) fail(`${label}.turns must be a non-empty array`); const turns = rawScenario.turns.map((turn, turnIndex) => parseTurn(turn, `${label}.turns[${turnIndex}]`)); + turns.forEach((turn, turnIndex) => { + if (turn.kind === "dialog" && turn.continuationOf !== undefined && turn.continuationOf >= turnIndex) { + fail(`${label}.turns[${turnIndex}].continuationOf must reference a preceding turn`); + } + }); if (rawScenario.expected !== undefined) { assertObject(rawScenario.expected, `${label}.expected`); if (rawScenario.expected.learningPhase !== undefined && !["LEARN", "DEEPEN", "EXPAND"].includes(rawScenario.expected.learningPhase as string)) { diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index 2dd965d8..54862dec 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -29,6 +29,7 @@ export type TutorQualityTurn = readonly question: string; readonly answer: string; readonly bindsToQuestion: string; + readonly continuationOf?: number; readonly difficulty?: number; readonly history?: readonly TutorDialogTurn[]; }; @@ -594,6 +595,7 @@ async function runSample( const violations: TutorQualityInvariantViolation[] = []; const turns: TutorQualityTranscriptTurn[] = []; const returnedModels: string[] = []; + const finalQuestions = new Map(); let executionStatus: TutorQualityExecutionStatus = "completed"; let invalidReason: string | undefined; let terminalError: TutorQualityTechnicalError | undefined; @@ -617,6 +619,15 @@ async function runSample( addViolation(violations, "unbound-question-context", "scenario", turnIndex); break; } + if (turn.continuationOf !== undefined) { + const precedingQuestion = finalQuestions.get(turn.continuationOf); + if (precedingQuestion === undefined || precedingQuestion !== turn.bindsToQuestion) { + executionStatus = "invalid"; + invalidReason = "preceding-question-mismatch"; + addViolation(violations, "preceding-question-mismatch", "scenario", turnIndex); + break; + } + } const response = await service.generateDialogResponse( scenario.sketch, historyFromSource(turn.history), @@ -651,6 +662,7 @@ async function runSample( deterministicRawChecks(capture.response.result, turn, turnIndex, turnChecks, violations); } if (finalResult) { + finalQuestions.set(turnIndex, finalResult.question); addCheck(turnChecks, "final-tutor-response-present", true); expectedChecks(scenario, finalResult, turn, stateBefore, content?.progressionState, turnChecks, violations, turnIndex); applicationMetadataChecks(finalResult, content, content?.progressionState, turnChecks, violations, turnIndex); diff --git a/tests/server/services/tutor/evaluation/anchor-corpus.test.ts b/tests/server/services/tutor/evaluation/anchor-corpus.test.ts index 92f27e26..dc5e642a 100644 --- a/tests/server/services/tutor/evaluation/anchor-corpus.test.ts +++ b/tests/server/services/tutor/evaluation/anchor-corpus.test.ts @@ -96,6 +96,13 @@ describe("Tutor Quality anchor corpus contract", () => { turns: [{ ...source.scenarios[0]!.turns[0]!, bindsToQuestion: undefined }], }], })], + ["self-referencing continuation", (source: TutorQualityCorpusSource) => ({ + ...source, + scenarios: [{ + ...source.scenarios[0]!, + turns: [{ ...source.scenarios[0]!.turns[0]!, continuationOf: 0 }], + }], + })], ])("rejects %s", (_label, mutate) => { expect(() => parseTutorQualityCorpus(mutate(validSource()), references)).toThrow(); }); diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts index a8106d49..274634fc 100644 --- a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -249,4 +249,32 @@ describe("real-provider Tutor Quality evaluation runner", () => { expect(unavailable.report).toMatchObject({ runStatus: "invalid", reason: "model-unavailable" }); expect(unavailable.report.providerCalls).toMatchObject({ total: 1, generationCalls: 0 }); }); + + it("does not apply a learner answer to an arbitrary preceding real-model question", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + question: "Welche neue Beobachtung ist belegt?", + answerRating: 4, + }), { + scenarios: [scenario({ + id: "bound-continuation", + turns: [ + { kind: "initial", difficulty: 20 }, + { + kind: "dialog", + question: "Welche deklarierte Frage soll gelten?", + answer: "Eine Antwort auf die deklarierte Frage.", + bindsToQuestion: "Welche deklarierte Frage soll gelten?", + continuationOf: 0, + difficulty: 20, + }, + ], + })], + })); + const transcript = result.transcripts[0]!; + + expect(transcript.executionStatus).toBe("invalid"); + expect(transcript.invalidReason).toBe("preceding-question-mismatch"); + expect(transcript.turns).toHaveLength(1); + }); }); From ab4835f520cd69fa9084100c3b22c18a22219842 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:09:14 +0200 Subject: [PATCH 14/21] chore: ignore disposable Tutor Quality artifacts --- .gitignore | 1 + 1 file changed, 1 insertion(+) diff --git a/.gitignore b/.gitignore index b90d9998..c40c3ae4 100644 --- a/.gitignore +++ b/.gitignore @@ -46,6 +46,7 @@ storage/binaries/ !/temp/.gitkeep /cache/ /test-results/ +.tutor-quality-output/ /playwright-report/ # /screenshots/ .scannerwork/ From bb628841a10b655ad97004aa141d014df1f63158 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:09:43 +0200 Subject: [PATCH 15/21] fix: validate returned provider model metadata --- .../services/tutor/evaluation/real-provider-evaluation.ts | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index 54862dec..501592e2 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -651,14 +651,18 @@ async function runSample( : undefined; if (capture?.response) { const returnedModel = capture.response.model; - returnedModels.push(returnedModel); - if (returnedModel !== options.requestedModel) { + if (typeof returnedModel !== "string" || returnedModel.length === 0) { + executionStatus = "invalid"; + invalidReason = "returned-model-missing"; + addCheck(turnChecks, "returned-model-matches-request", false, "missing"); + } else if (returnedModel !== options.requestedModel) { executionStatus = "invalid"; invalidReason = "returned-model-mismatch"; addCheck(turnChecks, "returned-model-matches-request", false, returnedModel); } else { addCheck(turnChecks, "returned-model-matches-request", true); } + if (typeof returnedModel === "string" && returnedModel.length > 0) returnedModels.push(returnedModel); deterministicRawChecks(capture.response.result, turn, turnIndex, turnChecks, violations); } if (finalResult) { From b5fb896534f3cd623097d2b5525cbb6cdb157060 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:13:49 +0200 Subject: [PATCH 16/21] fix: redact scripted learner input in evaluation artifacts --- .../evaluation/real-provider-evaluation.ts | 2 +- .../real-provider-evaluation.test.ts | 21 +++++++++++++++++++ 2 files changed, 22 insertions(+), 1 deletion(-) diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index 501592e2..b410f0a2 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -678,7 +678,7 @@ async function runSample( const turnFinishedAt = (options.now ?? (() => new Date()))(); turns.push({ index: turnIndex, - input: turn, + input: safeTurn(turn, options.credential), startedAt: turnStartedAt.toISOString(), durationMs: Math.max(0, turnFinishedAt.getTime() - turnStartedAt.getTime()), providerCalls: subtractCounts(provider.counts, callsBeforeTurn), diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts index 274634fc..9c8ab2a2 100644 --- a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -155,6 +155,27 @@ describe("real-provider Tutor Quality evaluation runner", () => { expect(transcriptText).toContain("[REDACTED]"); }); + it("redacts credentials from scripted learner input in every transcript turn", async () => { + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + answerRating: 4, + question: "Welche Beobachtung ist belegt?", + }), { + scenarios: [scenario({ + turns: [{ + kind: "dialog", + question: "Welche Beobachtung ist belegt?", + answer: "super-secret-value", + bindsToQuestion: "Welche Beobachtung ist belegt?", + }], + })], + })); + const transcriptText = JSON.stringify(result.transcripts[0]); + + expect(transcriptText).not.toContain("super-secret-value"); + expect(transcriptText).toContain("[REDACTED]"); + }); + it("records raw complete-solution and schema violations separately from technical failure", async () => { const result = await runTutorQualityEvaluation(options(providerFor({ question: "void setup() {} void loop() {}", From 93f6c5cee7b3bd9379de551dae54763fe114df3f Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:16:09 +0200 Subject: [PATCH 17/21] feat: inject Stage 2A artifact writer --- .../evaluation/real-provider-evaluation.ts | 30 ++++++++++++++----- .../real-provider-evaluation.test.ts | 18 +++++++++++ 2 files changed, 41 insertions(+), 7 deletions(-) diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index b410f0a2..4f731ad7 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -72,6 +72,10 @@ export interface TutorQualityEvaluationOptions { readonly git: TutorQualityGitState; readonly now?: () => Date; readonly runSuffix?: () => string; + readonly artifactWriter?: ( + report: TutorQualityEvaluationReport, + transcripts: readonly TutorQualityTranscript[], + ) => Promise; } export interface TutorQualityDeterministicCheck { @@ -864,7 +868,7 @@ function baseReport( }; } -async function writeArtifacts( +async function writeArtifactsToDirectory( outputDir: string | undefined, report: TutorQualityEvaluationReport, transcripts: readonly TutorQualityTranscript[], @@ -879,6 +883,18 @@ async function writeArtifacts( })); } +async function writeArtifacts( + options: TutorQualityEvaluationOptions, + report: TutorQualityEvaluationReport, + transcripts: readonly TutorQualityTranscript[], +): Promise { + if (options.artifactWriter) { + await options.artifactWriter(report, transcripts); + return; + } + await writeArtifactsToDirectory(options.outputDir, report, transcripts); +} + function invalidPreflightReason(options: TutorQualityEvaluationOptions): string | undefined { if (!options.requestedModel || options.requestedModel === "auto") return "fixed-model-required"; if (!options.git.sha) return "git-sha-missing"; @@ -906,17 +922,17 @@ export async function runTutorQualityEvaluation(options: TutorQualityEvaluationO if (invalidReason) { const report = baseReport(options, runId, evaluationIdentity, "invalid", invalidReason, { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); - await writeArtifacts(options.outputDir, report, []); + await writeArtifacts(options, report, []); return { report, transcripts: [] }; } if (!options.credential) { const report = baseReport(options, runId, evaluationIdentity, "not-run", "missing-credential", { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); - await writeArtifacts(options.outputDir, report, []); + await writeArtifacts(options, report, []); return { report, transcripts: [] }; } if (options.maxCalls === 0) { const report = baseReport(options, runId, evaluationIdentity, "not-run", "call-budget-zero", { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); - await writeArtifacts(options.outputDir, report, []); + await writeArtifacts(options, report, []); return { report, transcripts: [] }; } @@ -926,12 +942,12 @@ export async function runTutorQualityEvaluation(options: TutorQualityEvaluationO availableModels = await provider.listModels(options.credential); } catch (error) { const report = baseReport(options, runId, evaluationIdentity, "technical-failure", technicalError(error).kind, provider.counts, emptyByScenario); - await writeArtifacts(options.outputDir, report, []); + await writeArtifacts(options, report, []); return { report, transcripts: [] }; } if (!availableModels.includes(options.requestedModel)) { const report = baseReport(options, runId, evaluationIdentity, "invalid", "model-unavailable", provider.counts, emptyByScenario); - await writeArtifacts(options.outputDir, report, []); + await writeArtifacts(options, report, []); return { report, transcripts: [] }; } @@ -951,7 +967,7 @@ export async function runTutorQualityEvaluation(options: TutorQualityEvaluationO } } const report = baseReport(options, runId, evaluationIdentity, "completed", undefined, provider.counts, byScenario, transcripts.length); - await writeArtifacts(options.outputDir, report, transcripts); + await writeArtifacts(options, report, transcripts); return { report, transcripts }; } diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts index 9c8ab2a2..917e88a4 100644 --- a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -142,6 +142,24 @@ describe("real-provider Tutor Quality evaluation runner", () => { expect(reportText).not.toContain("super-secret-value"); }); + it("supports an injected artifact writer without touching the repository", async () => { + let writtenReport: string | undefined; + let writtenTranscriptCount = -1; + const result = await runTutorQualityEvaluation(options(providerFor({ + responseStyle: "normal", + answerRating: 4, + question: "Welche Beobachtung ist belegt?", + }), { + artifactWriter: async (report, transcripts) => { + writtenReport = report.runId; + writtenTranscriptCount = transcripts.length; + }, + })); + + expect(writtenReport).toBe(result.report.runId); + expect(writtenTranscriptCount).toBe(1); + }); + it("redacts a credential echoed by a fake provider from every transcript field", async () => { const result = await runTutorQualityEvaluation(options(providerFor({ responseStyle: "normal", From 11dcf563b14d13d7a56db0af9912fec6759ab5b7 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:18:32 +0200 Subject: [PATCH 18/21] docs: mark Stage 2A plan complete --- docs/plan-tutor-quality-stage-2a.md | 82 ++++++++++++++--------------- 1 file changed, 41 insertions(+), 41 deletions(-) diff --git a/docs/plan-tutor-quality-stage-2a.md b/docs/plan-tutor-quality-stage-2a.md index 0cb72218..950c8ac4 100644 --- a/docs/plan-tutor-quality-stage-2a.md +++ b/docs/plan-tutor-quality-stage-2a.md @@ -36,62 +36,62 @@ ## Task 1 – Lock the corpus and fixture contract with tests -- [ ] Add a failing loader/contract test for a manifest with `corpusId`, `corpusVersion`, stable scenario IDs, sketch references, course fixture/free-Tutor mode, synthetic answers, explicit turn bindings, and structural expectations. -- [ ] Add a failing test that rejects duplicate IDs, missing fixture references, `auto` model declarations, and unbound continuation answers. -- [ ] Add a separate corpus-evolution validator and failing tests for `compare(previousCorpus, currentCorpus)`: a parsed/digest-changing add, removal, or semantic edit requires a higher `corpusVersion`; formatting-only changes may keep the version only when parsed content and digest are unchanged. Keep historical comparison out of the current-corpus loader. -- [ ] Add `evals/tutor-quality/anchor-corpus.yaml` at version 1 with the ten approved anchors: `TQ-REG-001` (no variables-topic activation and no exact or Stage-1-heuristic repeat), simple variable, Serial prediction, incorrect answer, partial answer, strong answer/progression, unmatched/free Tutor, LEARN→DEEPEN, EXPAND, and off-topic answer. -- [ ] Add only the small required `.ino` fixtures under `evals/tutor-quality/fixtures/`; reuse the existing PWM fixture where possible rather than copying production content. -- [ ] Add a typed `anchor-course-content.ts` fixture factory for the small valid Course Content snapshots and seeded progression states required by topic activation, LEARN/DEEPEN, and EXPAND cases. Keep revision strings and question IDs explicit and reviewable. -- [ ] Run the focused corpus tests red before implementation and green after the loader/factory exists. +- [x] Add a failing loader/contract test for a manifest with `corpusId`, `corpusVersion`, stable scenario IDs, sketch references, course fixture/free-Tutor mode, synthetic answers, explicit turn bindings, and structural expectations. +- [x] Add a failing test that rejects duplicate IDs, missing fixture references, `auto` model declarations, and unbound continuation answers. +- [x] Add a separate corpus-evolution validator and failing tests for `compare(previousCorpus, currentCorpus)`: a parsed/digest-changing add, removal, or semantic edit requires a higher `corpusVersion`; formatting-only changes may keep the version only when parsed content and digest are unchanged. Keep historical comparison out of the current-corpus loader. +- [x] Add `evals/tutor-quality/anchor-corpus.yaml` at version 1 with the ten approved anchors: `TQ-REG-001` (no variables-topic activation and no exact or Stage-1-heuristic repeat), simple variable, Serial prediction, incorrect answer, partial answer, strong answer/progression, unmatched/free Tutor, LEARN→DEEPEN, EXPAND, and off-topic answer. +- [x] Add only the small required `.ino` fixtures under `evals/tutor-quality/fixtures/`; reuse the existing PWM fixture where possible rather than copying production content. +- [x] Add a typed `anchor-course-content.ts` fixture factory for the small valid Course Content snapshots and seeded progression states required by topic activation, LEARN/DEEPEN, and EXPAND cases. Keep revision strings and question IDs explicit and reviewable. +- [x] Run the focused corpus tests red before implementation and green after the loader/factory exists. ## Task 2 – Make prompt revision metadata explicit without changing prompts -- [ ] Add a failing unit test asserting a versioned prompt revision identifier and SHA-256 digest are stable, contain the effective system/initial-user/dialog-user template sources before scenario substitution, and change when a template source changes. -- [ ] Refactor only the prompt-source declarations needed by `TutorService` so existing generated prompt text remains byte-for-byte compatible; export the revision descriptor for the evaluator. -- [ ] Do not add a new prompt, quality rule, or runtime strategy. The revision helper is metadata only. -- [ ] Run existing Tutor prompt tests plus the new revision test. +- [x] Add a failing unit test asserting a versioned prompt revision identifier and SHA-256 digest are stable, contain the effective system/initial-user/dialog-user template sources before scenario substitution, and change when a template source changes. +- [x] Refactor only the prompt-source declarations needed by `TutorService` so existing generated prompt text remains byte-for-byte compatible; export the revision descriptor for the evaluator. +- [x] Do not add a new prompt, quality rule, or runtime strategy. The revision helper is metadata only. +- [x] Run existing Tutor prompt tests plus the new revision test. ## Task 3 – Implement the injectable evaluation runner and transcript model -- [ ] Add failing tests for fresh state/history cloning per sample, normal `TutorService` initial/dialog invocation, provider capture before validation/repair, shared diagnostic repeat/solution/schema checks, state-before/state-after snapshots, and explicit expected structural checks. -- [ ] Add `server/services/tutor/evaluation/real-provider-evaluation.ts` with small typed contracts for corpus scenarios, invocation options, sample metadata, logical turns, deterministic check records, transcript artifacts, and aggregate reports. -- [ ] Inject the provider, clock, random suffix, Git metadata, and output writer seams so tests never need credentials, network, or a mutable repository. -- [ ] Wrap the provider to capture requests and parsed `ProviderQuestionResult` values before `TutorService` receives them. Keep the raw capture allow-listed and exclude transport envelopes/headers. -- [ ] Invoke `TutorService` with `CurriculumTutorAdapter` when a scenario declares Course Content; invoke the same service without planning for free-Tutor cases. -- [ ] Split the existing pure learning-question validation internally into one shared diagnostic function that returns granular deterministic violation records and keep `validateLearningQuestion` as the existing throw/normalization wrapper. Export the diagnostic function for Stage 2A; add no new rule and preserve runtime behavior. -- [ ] Reuse that diagnostic function and `isSemanticallyRepeatedQuestion` for deterministic checks. Record raw complete-solution/repeat violations even when TutorService rejects or repairs the response; never turn these into semantic scores. -- [ ] Enforce fixed requested model, preflight model availability, returned-model equality, explicit sample limits, and a provider-call budget covering *all* external calls, including `listModels()` and generation. Report `providerCalls`, `modelListCalls`, and `generationCalls` separately; stop before any call that would exceed the budget. -- [ ] Implement the two status axes from the SSOT: `executionStatus` (`completed`, `invalid`, `technical-failure`, `not-run`) and `invariantViolations` (array). Classify missing credentials/zero preflight budget as `not-run`; provider errors, timeout, malformed responses, and mid-run budget exhaustion as technical failures; metadata/model/binding problems as invalid. -- [ ] Implement canonical JSON hashing for `evaluationIdentity`, run IDs with UTC timestamp plus collision-resistant suffix, and secret-free allow-listed JSON transcript writing. -- [ ] Implement aggregate counts/rates per scenario and overall with explicit denominators, exact provider-call counts, separate model-list/generation counts, budget exhaustion, and cost `unavailable` when the provider supplies no cost data. Always write a run-level `not-run` report for missing credentials with `reason: missing-credential`, zero provider calls, and no sample transcripts. -- [ ] Run the focused evaluator tests red before implementation and green after each runner slice. +- [x] Add failing tests for fresh state/history cloning per sample, normal `TutorService` initial/dialog invocation, provider capture before validation/repair, shared diagnostic repeat/solution/schema checks, state-before/state-after snapshots, and explicit expected structural checks. +- [x] Add `server/services/tutor/evaluation/real-provider-evaluation.ts` with small typed contracts for corpus scenarios, invocation options, sample metadata, logical turns, deterministic check records, transcript artifacts, and aggregate reports. +- [x] Inject the provider, clock, random suffix, Git metadata, and output writer seams so tests never need credentials, network, or a mutable repository. +- [x] Wrap the provider to capture requests and parsed `ProviderQuestionResult` values before `TutorService` receives them. Keep the raw capture allow-listed and exclude transport envelopes/headers. +- [x] Invoke `TutorService` with `CurriculumTutorAdapter` when a scenario declares Course Content; invoke the same service without planning for free-Tutor cases. +- [x] Split the existing pure learning-question validation internally into one shared diagnostic function that returns granular deterministic violation records and keep `validateLearningQuestion` as the existing throw/normalization wrapper. Export the diagnostic function for Stage 2A; add no new rule and preserve runtime behavior. +- [x] Reuse that diagnostic function and `isSemanticallyRepeatedQuestion` for deterministic checks. Record raw complete-solution/repeat violations even when TutorService rejects or repairs the response; never turn these into semantic scores. +- [x] Enforce fixed requested model, preflight model availability, returned-model equality, explicit sample limits, and a provider-call budget covering *all* external calls, including `listModels()` and generation. Report `providerCalls`, `modelListCalls`, and `generationCalls` separately; stop before any call that would exceed the budget. +- [x] Implement the two status axes from the SSOT: `executionStatus` (`completed`, `invalid`, `technical-failure`, `not-run`) and `invariantViolations` (array). Classify missing credentials/zero preflight budget as `not-run`; provider errors, timeout, malformed responses, and mid-run budget exhaustion as technical failures; metadata/model/binding problems as invalid. +- [x] Implement canonical JSON hashing for `evaluationIdentity`, run IDs with UTC timestamp plus collision-resistant suffix, and secret-free allow-listed JSON transcript writing. +- [x] Implement aggregate counts/rates per scenario and overall with explicit denominators, exact provider-call counts, separate model-list/generation counts, budget exhaustion, and cost `unavailable` when the provider supplies no cost data. Always write a run-level `not-run` report for missing credentials with `reason: missing-credential`, zero provider calls, and no sample transcripts. +- [x] Run the focused evaluator tests red before implementation and green after each runner slice. ## Task 4 – Add adversarial fake-provider coverage -- [ ] Add fake-provider tests for exact repeated questions, Stage-1 heuristic repeats, invalid schema, complete solution, forged/wrong planning metadata, provider error before commit, timeout, returned-model mismatch, and call-budget exhaustion. -- [ ] Assert repaired final planning metadata remains application-owned and state commits only after a successful TutorService request. -- [ ] Assert technical failures do not mutate progression state and are not counted as Tutor-quality violations unless a separate raw deterministic violation was observed. -- [ ] Assert no credential value appears in serialized transcript, report, thrown error, or logger input. +- [x] Add fake-provider tests for exact repeated questions, Stage-1 heuristic repeats, invalid schema, complete solution, forged/wrong planning metadata, provider error before commit, timeout, returned-model mismatch, and call-budget exhaustion. +- [x] Assert repaired final planning metadata remains application-owned and state commits only after a successful TutorService request. +- [x] Assert technical failures do not mutate progression state and are not counted as Tutor-quality violations unless a separate raw deterministic violation was observed. +- [x] Assert no credential value appears in serialized transcript, report, thrown error, or logger input. ## Task 5 – Add the explicit CLI and local execution contract -- [ ] Add `scripts/tutor-quality-real-provider-eval.ts` as a thin CLI around the runner. Require a fixed `--model`, bounded `--samples` and `--max-calls`, `--output-dir`, corpus selection, and a credential environment-variable name; reject `auto` and any credential value flag. -- [ ] Add a package script such as `eval:tutor-quality:real` that is not referenced by `test`, `test:unit`, `test:tutor-quality`, or normal build gates. -- [ ] Resolve the repository Git SHA/clean state and configured Course Content/corpus revisions before execution. Abort as `invalid` without provider calls when preflight identity cannot be proven. -- [ ] Add a short operator document with a local command, expected output paths, missing-credential behavior, call-budget example, and explicit warning that Stage 2A observes deterministic integrity rather than learning effect. -- [ ] Test CLI argument validation and missing-credential `not-run` behavior without contacting a provider. +- [x] Add `scripts/tutor-quality-real-provider-eval.ts` as a thin CLI around the runner. Require a fixed `--model`, bounded `--samples` and `--max-calls`, `--output-dir`, corpus selection, and a credential environment-variable name; reject `auto` and any credential value flag. +- [x] Add a package script such as `eval:tutor-quality:real` that is not referenced by `test`, `test:unit`, `test:tutor-quality`, or normal build gates. +- [x] Resolve the repository Git SHA/clean state and configured Course Content/corpus revisions before execution. Abort as `invalid` without provider calls when preflight identity cannot be proven. +- [x] Add a short operator document with a local command, expected output paths, missing-credential behavior, call-budget example, and explicit warning that Stage 2A observes deterministic integrity rather than learning effect. +- [x] Test CLI argument validation and missing-credential `not-run` behavior without contacting a provider. ## Task 6 – Add a manual-only workflow -- [ ] Add `.github/workflows/tutor-quality-real-provider.yml` with only `workflow_dispatch`, explicit model/sample/call-budget inputs, Node version from `.nvmrc`, and a repository secret exposed only to the invoked process through the configured environment variable. -- [ ] Upload bounded transcript/report artifacts with retention; never print the secret or use a `pull_request`/required-check trigger. -- [ ] Make missing secret a visible skipped/not-run result, not a failing PR gate. -- [ ] Document that this workflow is optional/manual first; do not add nightly scheduling until cost and stability are known. +- [x] Add `.github/workflows/tutor-quality-real-provider.yml` with only `workflow_dispatch`, explicit model/sample/call-budget inputs, Node version from `.nvmrc`, and a repository secret exposed only to the invoked process through the configured environment variable. +- [x] Upload bounded transcript/report artifacts with retention; never print the secret or use a `pull_request`/required-check trigger. +- [x] Make missing secret a visible skipped/not-run result, not a failing PR gate. +- [x] Document that this workflow is optional/manual first; do not add nightly scheduling until cost and stability are known. ## Task 7 – Verification and review handoff -- [ ] Run the Node-version check, focused Stage-2A tests, existing `npm run test:tutor-quality`, `npm run check`, `npm run check:docs`, and `git diff --check`. -- [ ] Run the CLI in no-credential mode and verify it writes the documented run-level `not-run` report with `reason: missing-credential`, zero provider calls, no sample transcripts, and no secret-like data. -- [ ] Confirm generated transcripts/reports are ignored or written only to caller-selected disposable directories. -- [ ] Review the diff for accidental production behavior changes, duplicated Stage-1 validation, direct HTTP access, semantic scoring, and CI hard-gate coupling. -- [ ] Commit the implementation in coherent commits and report branch/base/HEAD, files, anchors, deterministic metrics, excluded Stage-2B dimensions, tests, and the absence of a real-provider run if credentials are unavailable. +- [x] Run the Node-version check, focused Stage-2A tests, existing `npm run test:tutor-quality`, `npm run check`, `npm run check:docs`, and `git diff --check`. +- [x] Run the CLI in no-credential mode and verify it writes the documented run-level `not-run` report with `reason: missing-credential`, zero provider calls, no sample transcripts, and no secret-like data. +- [x] Confirm generated transcripts/reports are ignored or written only to caller-selected disposable directories. +- [x] Review the diff for accidental production behavior changes, duplicated Stage-1 validation, direct HTTP access, semantic scoring, and CI hard-gate coupling. +- [x] Commit the implementation in coherent commits and report branch/base/HEAD, files, anchors, deterministic metrics, excluded Stage-2B dimensions, tests, and the absence of a real-provider run if credentials are unavailable. From 20054d29cd4695f9809e988366d3b648c7d61f3d Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:31:05 +0200 Subject: [PATCH 19/21] fix: bound Stage 2A evaluation inputs --- docs/tutor-quality-stage-2a.md | 3 +++ scripts/tutor-quality-real-provider-eval.ts | 9 ++++++--- .../tutor/evaluation/real-provider-evaluation.ts | 5 +++++ .../evaluation/real-provider-evaluation.test.ts | 14 ++++++++++++++ .../tutor-quality-real-provider-cli.test.ts | 4 ++++ 5 files changed, 32 insertions(+), 3 deletions(-) diff --git a/docs/tutor-quality-stage-2a.md b/docs/tutor-quality-stage-2a.md index cad820a1..7a349856 100644 --- a/docs/tutor-quality-stage-2a.md +++ b/docs/tutor-quality-stage-2a.md @@ -29,6 +29,9 @@ run-level `report.json` with `runStatus: "not-run"`, `reason: "missing-credential"`, zero provider calls, and no sample transcripts. +Each invocation is bounded to at most 20 samples per scenario and 500 total +provider calls. The CLI rejects larger values before any provider call. + ## Artifacts `report.json` contains the run identity, corpus/course/prompt/provider/model diff --git a/scripts/tutor-quality-real-provider-eval.ts b/scripts/tutor-quality-real-provider-eval.ts index b9b08726..29317308 100644 --- a/scripts/tutor-quality-real-provider-eval.ts +++ b/scripts/tutor-quality-real-provider-eval.ts @@ -11,6 +11,8 @@ import { type TutorQualityTurnSource, } from "../server/services/tutor/evaluation/anchor-corpus"; import { + MAX_TUTOR_QUALITY_CALLS, + MAX_TUTOR_QUALITY_SAMPLES, runTutorQualityEvaluation, type TutorQualityEvaluationScenario, type TutorQualityGitState, @@ -44,10 +46,11 @@ function readArgument(argv: readonly string[], index: number, flag: string): str return value; } -function parsePositiveInteger(value: string, flag: string, allowZero = false): number { +function parsePositiveInteger(value: string, flag: string, allowZero = false, maximum?: number): number { if (!/^\d+$/.test(value)) throw new Error(`${flag} must be an integer`); const parsed = Number(value); if (!Number.isSafeInteger(parsed) || (allowZero ? parsed < 0 : parsed < 1)) throw new Error(`${flag} is out of range`); + if (maximum !== undefined && parsed > maximum) throw new Error(`${flag} exceeds maximum ${maximum}`); return parsed; } @@ -71,11 +74,11 @@ export function parseTutorQualityCliArgs(argv: readonly string[]): TutorQualityC index += 1; break; case "--samples": - samples = parsePositiveInteger(readArgument(argv, index, flag), flag); + samples = parsePositiveInteger(readArgument(argv, index, flag), flag, false, MAX_TUTOR_QUALITY_SAMPLES); index += 1; break; case "--max-calls": - maxCalls = parsePositiveInteger(readArgument(argv, index, flag), flag, true); + maxCalls = parsePositiveInteger(readArgument(argv, index, flag), flag, true, MAX_TUTOR_QUALITY_CALLS); index += 1; break; case "--output-dir": diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index 4f731ad7..97bc5ece 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -19,6 +19,9 @@ import type { TutorProgressionState } from "../curriculum/progression-state"; export type TutorQualityExecutionStatus = "completed" | "invalid" | "technical-failure" | "not-run"; +export const MAX_TUTOR_QUALITY_SAMPLES = 20; +export const MAX_TUTOR_QUALITY_CALLS = 500; + export type TutorQualityTurn = | { readonly kind: "initial"; @@ -900,7 +903,9 @@ function invalidPreflightReason(options: TutorQualityEvaluationOptions): string if (!options.git.sha) return "git-sha-missing"; if (!options.git.trackedClean || !options.git.relevantUntrackedClean) return "dirty-relevant-worktree"; if (!Number.isInteger(options.samples) || options.samples < 1) return "invalid-sample-count"; + if (options.samples > MAX_TUTOR_QUALITY_SAMPLES) return "sample-count-exceeds-limit"; if (!Number.isInteger(options.maxCalls) || options.maxCalls < 0) return "invalid-call-budget"; + if (options.maxCalls > MAX_TUTOR_QUALITY_CALLS) return "call-budget-exceeds-limit"; if (options.scenarios.length === 0) return "empty-corpus"; const first = options.scenarios[0]; if (options.scenarios.some((scenario) => scenario.corpusId !== first?.corpusId || scenario.corpusVersion !== first?.corpusVersion)) { diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts index 917e88a4..ecba3bc3 100644 --- a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -4,6 +4,8 @@ import path from "node:path"; import { afterEach, describe, expect, it } from "vitest"; import { createAnchorCourseContent } from "../../../../../server/services/tutor/evaluation/anchor-course-content"; import { + MAX_TUTOR_QUALITY_CALLS, + MAX_TUTOR_QUALITY_SAMPLES, runTutorQualityEvaluation, type TutorQualityEvaluationScenario, } from "../../../../../server/services/tutor/evaluation/real-provider-evaluation"; @@ -287,6 +289,18 @@ describe("real-provider Tutor Quality evaluation runner", () => { })); expect(unavailable.report).toMatchObject({ runStatus: "invalid", reason: "model-unavailable" }); expect(unavailable.report.providerCalls).toMatchObject({ total: 1, generationCalls: 0 }); + + const unboundedSamples = await runTutorQualityEvaluation(options(providerFor({ question: "Welche Beobachtung ist belegt?" }), { + samples: MAX_TUTOR_QUALITY_SAMPLES + 1, + })); + expect(unboundedSamples.report).toMatchObject({ runStatus: "invalid", reason: "sample-count-exceeds-limit" }); + expect(unboundedSamples.report.providerCalls.total).toBe(0); + + const unboundedCalls = await runTutorQualityEvaluation(options(providerFor({ question: "Welche Beobachtung ist belegt?" }), { + maxCalls: MAX_TUTOR_QUALITY_CALLS + 1, + })); + expect(unboundedCalls.report).toMatchObject({ runStatus: "invalid", reason: "call-budget-exceeds-limit" }); + expect(unboundedCalls.report.providerCalls.total).toBe(0); }); it("does not apply a learner answer to an arbitrary preceding real-model question", async () => { diff --git a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts index 814cab91..a7abf49c 100644 --- a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts +++ b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts @@ -3,6 +3,8 @@ import os from "node:os"; import path from "node:path"; import { describe, expect, it } from "vitest"; import { + MAX_TUTOR_QUALITY_CALLS, + MAX_TUTOR_QUALITY_SAMPLES, loadTutorQualityEvaluationScenarios, isTutorQualityRelevantUntrackedPath, parseTutorQualityCliArgs, @@ -33,6 +35,8 @@ describe("Tutor Quality real-provider CLI contract", () => { it("rejects malformed numeric limits and credential values", () => { expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--samples", "0"])).toThrow(); expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--credential-env", "not-a-value"])).toThrow(); + expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--samples", String(MAX_TUTOR_QUALITY_SAMPLES + 1)])).toThrow(); + expect(() => parseTutorQualityCliArgs(["--output-dir", "/tmp/tq", "--model", "pilot", "--max-calls", String(MAX_TUTOR_QUALITY_CALLS + 1)])).toThrow(); }); it("does not treat protected editor SSOT files or the output directory as relevant inputs", () => { From 3e538caa497559ce5e03a1b7dcd876775f132882 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Sun, 27 Sep 2026 23:33:23 +0200 Subject: [PATCH 20/21] fix: handle missing returned model metadata safely --- .../evaluation/real-provider-evaluation.ts | 5 ++++- .../real-provider-evaluation.test.ts | 19 +++++++++++++++++++ 2 files changed, 23 insertions(+), 1 deletion(-) diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index 97bc5ece..be8de5bd 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -690,7 +690,10 @@ async function runSample( durationMs: Math.max(0, turnFinishedAt.getTime() - turnStartedAt.getTime()), providerCalls: subtractCounts(provider.counts, callsBeforeTurn), ...(captureRequest ? { providerRequest: requestArtifact(captureRequest, options.credential) } : {}), - ...(capture?.response ? { rawProviderResult: safeResult(capture.response.result, options.credential), returnedModel: redact(capture.response.model, options.credential) } : {}), + ...(capture?.response ? { + rawProviderResult: safeResult(capture.response.result, options.credential), + ...(typeof capture.response.model === "string" ? { returnedModel: redact(capture.response.model, options.credential) } : {}), + } : {}), ...(finalResult ? { finalTutorResult: safeResult(finalResult, options.credential) } : {}), deterministicChecks: turnChecks, ...(error ? { technicalError: error } : {}), diff --git a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts index ecba3bc3..bed1f879 100644 --- a/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts +++ b/tests/server/services/tutor/evaluation/real-provider-evaluation.test.ts @@ -118,6 +118,25 @@ describe("real-provider Tutor Quality evaluation runner", () => { expect(transcript.invariantViolations).toHaveLength(0); }); + it("marks a missing returned model invalid without leaking or throwing", async () => { + const result = await runTutorQualityEvaluation(options({ + async listModels() { + return ["fake-model"]; + }, + async generateLearningQuestion() { + return { + model: undefined as never, + result: { question: "Welche Beobachtung ist belegt?" }, + }; + }, + })); + const transcript = result.transcripts[0]!; + + expect(transcript.executionStatus).toBe("invalid"); + expect(transcript.invalidReason).toBe("returned-model-missing"); + expect(transcript.turns[0]).not.toHaveProperty("returnedModel"); + }); + it("counts every provider call against the budget and stops before generation", async () => { const result = await runTutorQualityEvaluation(options(providerFor({ question: "Welche Beobachtung ist belegt?" }), { maxCalls: 2 })); const transcript = result.transcripts[0]!; From 5add97a4f76752f6c0b2b69f8ef47f8eef4971f9 Mon Sep 17 00:00:00 2001 From: ttbombadil Date: Mon, 28 Sep 2026 08:11:11 +0200 Subject: [PATCH 21/21] chore: clear Stage 2A Sonar findings --- scripts/tutor-quality-real-provider-eval.ts | 72 ++- .../tutor/evaluation/anchor-corpus.ts | 175 +++--- .../tutor/evaluation/anchor-course-content.ts | 19 +- .../evaluation/real-provider-evaluation.ts | 526 ++++++++++++------ .../tutor-quality-real-provider-cli.test.ts | 6 + 5 files changed, 514 insertions(+), 284 deletions(-) diff --git a/scripts/tutor-quality-real-provider-eval.ts b/scripts/tutor-quality-real-provider-eval.ts index 29317308..a70164d7 100644 --- a/scripts/tutor-quality-real-provider-eval.ts +++ b/scripts/tutor-quality-real-provider-eval.ts @@ -1,5 +1,6 @@ import { execFileSync } from "node:child_process"; import { access, readFile } from "node:fs/promises"; +import { accessSync, constants as fsConstants } from "node:fs"; import path from "node:path"; import { pathToFileURL } from "node:url"; import { parse as parseYaml } from "yaml"; @@ -40,10 +41,16 @@ export interface TutorQualityCliDependencies { const DEFAULT_CORPUS_PATH = "evals/tutor-quality/anchor-corpus.yaml"; const DEFAULT_CREDENTIAL_ENV = "UNOSIM_TUTOR_EVAL_CREDENTIAL"; -function readArgument(argv: readonly string[], index: number, flag: string): string { - const value = argv[index + 1]; - if (!value || value.startsWith("--")) throw new Error(`${flag} requires a value`); - return value; +function argumentPairs(argv: readonly string[]): readonly (readonly [string, string])[] { + if (argv.length % 2 !== 0) throw new Error("Tutor Quality evaluation options require values"); + return Array.from({ length: argv.length / 2 }, (_, pairIndex) => { + const index = pairIndex * 2; + const flag = argv[index]; + const value = argv[index + 1]; + if (!flag?.startsWith("--")) throw new Error(`Unknown Tutor Quality evaluation option: ${flag ?? ""}`); + if (!value || value.startsWith("--")) throw new Error(`${flag} requires a value`); + return [flag, value] as const; + }); } function parsePositiveInteger(value: string, flag: string, allowZero = false, maximum?: number): number { @@ -66,32 +73,25 @@ export function parseTutorQualityCliArgs(argv: readonly string[]): TutorQualityC let outputDir: string | undefined; let corpusPath = DEFAULT_CORPUS_PATH; let credentialEnv = DEFAULT_CREDENTIAL_ENV; - for (let index = 0; index < argv.length; index += 1) { - const flag = argv[index]; + for (const [flag, value] of argumentPairs(argv)) { switch (flag) { case "--model": - model = readArgument(argv, index, flag); - index += 1; + model = value; break; case "--samples": - samples = parsePositiveInteger(readArgument(argv, index, flag), flag, false, MAX_TUTOR_QUALITY_SAMPLES); - index += 1; + samples = parsePositiveInteger(value, flag, false, MAX_TUTOR_QUALITY_SAMPLES); break; case "--max-calls": - maxCalls = parsePositiveInteger(readArgument(argv, index, flag), flag, true, MAX_TUTOR_QUALITY_CALLS); - index += 1; + maxCalls = parsePositiveInteger(value, flag, true, MAX_TUTOR_QUALITY_CALLS); break; case "--output-dir": - outputDir = readArgument(argv, index, flag); - index += 1; + outputDir = value; break; case "--corpus": - corpusPath = readArgument(argv, index, flag); - index += 1; + corpusPath = value; break; case "--credential-env": - credentialEnv = parseCredentialEnvironmentName(readArgument(argv, index, flag)); - index += 1; + credentialEnv = parseCredentialEnvironmentName(value); break; default: throw new Error(`Unknown Tutor Quality evaluation option: ${flag}`); @@ -103,6 +103,23 @@ export function parseTutorQualityCliArgs(argv: readonly string[]): TutorQualityC return { model, samples, maxCalls, outputDir, corpusPath, credentialEnv }; } +const GIT_EXECUTABLE_CANDIDATES = process.platform === "win32" + ? [String.raw`C:\Program Files\Git\cmd\git.exe`, String.raw`C:\Program Files\Git\bin\git.exe`] + : ["/usr/bin/git", "/opt/homebrew/bin/git", "/usr/local/bin/git"]; + +function resolveGitExecutable(): string { + const executable = GIT_EXECUTABLE_CANDIDATES.find((candidate) => { + try { + accessSync(candidate, fsConstants.X_OK); + return true; + } catch { + return false; + } + }); + if (!executable) throw new Error("git executable not found in fixed system locations"); + return executable; +} + function historyEntry(entry: TutorQualityHistoryEntrySource): import("../shared/tutor").TutorDialogTurn { return { question: entry.question, @@ -153,7 +170,7 @@ export async function loadTutorQualityEvaluationScenarios( } function statusLines(cwd: string): readonly string[] { - const output = execFileSync("git", ["status", "--porcelain=v1", "-uall"], { cwd, encoding: "utf8" }); + const output = execFileSync(resolveGitExecutable(), ["status", "--porcelain=v1", "-uall"], { cwd, encoding: "utf8" }); return output.split("\n").map((line) => line.trimEnd()).filter(Boolean); } @@ -169,7 +186,7 @@ export function isTutorQualityRelevantUntrackedPath(relativePath: string, output } export function readTutorQualityGitState(cwd: string, outputDir: string): TutorQualityGitState { - const sha = execFileSync("git", ["rev-parse", "HEAD"], { cwd, encoding: "utf8" }).trim(); + const sha = execFileSync(resolveGitExecutable(), ["rev-parse", "HEAD"], { cwd, encoding: "utf8" }).trim(); const lines = statusLines(cwd); const trackedClean = lines.every((line) => line.slice(0, 2) === "??"); const outputRelative = path.relative(cwd, path.resolve(cwd, outputDir)); @@ -212,12 +229,11 @@ const invokedDirectly = process.argv[1] !== undefined && import.meta.url === pathToFileURL(path.resolve(process.argv[1])).href; if (invokedDirectly) { - runTutorQualityCli(process.argv.slice(2)) - .then(({ report }) => { - console.log(JSON.stringify({ runId: report.runId, runStatus: report.runStatus, reason: report.reason, providerCalls: report.providerCalls })); - }) - .catch((error: unknown) => { - console.error(error instanceof Error ? error.message : "Tutor Quality evaluation failed"); - process.exitCode = 1; - }); + try { + const { report } = await runTutorQualityCli(process.argv.slice(2)); + console.log(JSON.stringify({ runId: report.runId, runStatus: report.runStatus, reason: report.reason, providerCalls: report.providerCalls })); + } catch (error: unknown) { + console.error(error instanceof Error ? error.message : "Tutor Quality evaluation failed"); + process.exitCode = 1; + } } diff --git a/server/services/tutor/evaluation/anchor-corpus.ts b/server/services/tutor/evaluation/anchor-corpus.ts index 6629ab0a..ac26c0c6 100644 --- a/server/services/tutor/evaluation/anchor-corpus.ts +++ b/server/services/tutor/evaluation/anchor-corpus.ts @@ -1,7 +1,5 @@ import { createHash } from "node:crypto"; -export type TutorQualityCourseContentRef = "free" | string; - export interface TutorQualityHistoryEntrySource { readonly question: string; readonly answer?: string; @@ -28,7 +26,7 @@ export type TutorQualityTurnSource = export interface TutorQualityScenarioSource { readonly id: string; readonly sketch: string; - readonly courseContent: TutorQualityCourseContentRef; + readonly courseContent: string; readonly model?: string; readonly turns: readonly TutorQualityTurnSource[]; readonly expected?: { @@ -72,54 +70,67 @@ function assertNonEmptyString(value: unknown, label: string): asserts value is s if (typeof value !== "string" || value.trim().length === 0) fail(`${label} must be a non-empty string`); } -function assertDifficulty(value: unknown, label: string): void { - if (value !== undefined && (!Number.isInteger(value) || (value as number) < 1 || (value as number) > 100)) { +function optionalInteger(value: unknown, label: string): number | undefined { + if (value === undefined) return undefined; + if (typeof value !== "number" || !Number.isInteger(value) || value < 1 || value > 100) { fail(`${label} must be an integer from 1 to 100`); } + return value; +} + +function parseHistory(value: unknown, label: string): readonly TutorQualityHistoryEntrySource[] | undefined { + if (value === undefined) return undefined; + if (!Array.isArray(value)) fail(`${label}.history must be an array`); + for (const [index, entry] of value.entries()) { + const historyLabel = `${label}.history[${index}]`; + assertObject(entry, historyLabel); + assertNonEmptyString(entry.question, `${historyLabel}.question`); + if (entry.answer !== undefined) assertNonEmptyString(entry.answer, `${historyLabel}.answer`); + if (entry.responseStyle !== undefined && entry.responseStyle !== "normal" && entry.responseStyle !== "philosophical") { + fail(`${historyLabel}.responseStyle is invalid`); + } + if (entry.answerRating !== undefined && ![1, 2, 3, 4, 5].includes(entry.answerRating as 1 | 2 | 3 | 4 | 5)) { + fail(`${historyLabel}.answerRating is invalid`); + } + if (entry.questionId !== undefined) assertNonEmptyString(entry.questionId, `${historyLabel}.questionId`); + } + return value as readonly TutorQualityHistoryEntrySource[]; } -function parseTurn(value: unknown, label: string): TutorQualityTurnSource { - assertObject(value, label); - if (value.kind === "initial") { - assertDifficulty(value.difficulty, `${label}.difficulty`); - return { kind: "initial", ...(value.difficulty === undefined ? {} : { difficulty: value.difficulty as number }) }; - } +function parseInitialTurn(value: Record, label: string): TutorQualityTurnSource | undefined { + if (value.kind !== "initial") return undefined; + const difficulty = optionalInteger(value.difficulty, `${label}.difficulty`); + return { kind: "initial", ...(difficulty === undefined ? {} : { difficulty }) }; +} + +function parseDialogTurn(value: Record, label: string): TutorQualityTurnSource { if (value.kind !== "dialog") fail(`${label}.kind must be initial or dialog`); assertNonEmptyString(value.question, `${label}.question`); assertNonEmptyString(value.answer, `${label}.answer`); assertNonEmptyString(value.bindsToQuestion, `${label}.bindsToQuestion`); if (value.question !== value.bindsToQuestion) fail(`${label} bindsToQuestion must equal question`); - if (value.continuationOf !== undefined && (typeof value.continuationOf !== "number" || !Number.isInteger(value.continuationOf) || value.continuationOf < 0)) { + const continuationOf = value.continuationOf; + if (continuationOf !== undefined && (typeof continuationOf !== "number" || !Number.isInteger(continuationOf) || continuationOf < 0)) { fail(`${label}.continuationOf must reference a preceding turn`); } - assertDifficulty(value.difficulty, `${label}.difficulty`); - if (value.history !== undefined) { - if (!Array.isArray(value.history)) fail(`${label}.history must be an array`); - for (const [index, entry] of value.history.entries()) { - const historyLabel = `${label}.history[${index}]`; - assertObject(entry, historyLabel); - assertNonEmptyString(entry.question, `${historyLabel}.question`); - if (entry.answer !== undefined) assertNonEmptyString(entry.answer, `${historyLabel}.answer`); - if (entry.responseStyle !== undefined && entry.responseStyle !== "normal" && entry.responseStyle !== "philosophical") { - fail(`${historyLabel}.responseStyle is invalid`); - } - if (entry.answerRating !== undefined && ![1, 2, 3, 4, 5].includes(entry.answerRating as 1 | 2 | 3 | 4 | 5)) { - fail(`${historyLabel}.answerRating is invalid`); - } - if (entry.questionId !== undefined) assertNonEmptyString(entry.questionId, `${historyLabel}.questionId`); - } - } + const difficulty = optionalInteger(value.difficulty, `${label}.difficulty`); + const history = parseHistory(value.history, label); return { kind: "dialog", question: value.question, answer: value.answer, bindsToQuestion: value.bindsToQuestion, - ...(value.continuationOf === undefined ? {} : { continuationOf: value.continuationOf as number }), - ...(value.difficulty === undefined ? {} : { difficulty: value.difficulty as number }), - ...(value.history === undefined ? {} : { history: value.history as readonly TutorQualityHistoryEntrySource[] }), + ...(continuationOf === undefined ? {} : { continuationOf }), + ...(difficulty === undefined ? {} : { difficulty }), + ...(history === undefined ? {} : { history }), }; } +function parseTurn(value: unknown, label: string): TutorQualityTurnSource { + assertObject(value, label); + return parseInitialTurn(value, label) ?? parseDialogTurn(value, label); +} + function normalizedCorpusSource(source: TutorQualityCorpusSource): TutorQualityCorpusSource { return { corpusId: source.corpusId, @@ -146,6 +157,63 @@ export function tutorQualityCorpusDigest(source: TutorQualityCorpusSource): stri return createHash("sha256").update(stableJson(normalizedCorpusSource(source))).digest("hex"); } +function parseExpected(value: unknown, label: string): TutorQualityScenarioSource["expected"] | undefined { + if (value === undefined) return undefined; + assertObject(value, `${label}.expected`); + if (value.learningPhase !== undefined && !["LEARN", "DEEPEN", "EXPAND"].includes(value.learningPhase as string)) { + fail(`${label}.expected.learningPhase is invalid`); + } + if (value.questionNotRepeat !== undefined && value.questionNotRepeat !== "exact-or-heuristic") { + fail(`${label}.expected.questionNotRepeat is invalid`); + } + return value as TutorQualityScenarioSource["expected"]; +} + +function parseTurns(value: unknown, label: string): readonly TutorQualityTurnSource[] { + if (!Array.isArray(value) || value.length === 0) fail(`${label}.turns must be a non-empty array`); + const turns = value.map((turn, turnIndex) => parseTurn(turn, `${label}.turns[${turnIndex}]`)); + turns.forEach((turn, turnIndex) => { + if (turn.kind === "dialog" && turn.continuationOf !== undefined && turn.continuationOf >= turnIndex) { + fail(`${label}.turns[${turnIndex}].continuationOf must reference a preceding turn`); + } + }); + return turns; +} + +function parseScenario( + rawScenario: unknown, + index: number, + references: TutorQualityCorpusReferences, + ids: Set, +): TutorQualityScenarioSource { + const label = `scenarios[${index}]`; + assertObject(rawScenario, label); + assertNonEmptyString(rawScenario.id, `${label}.id`); + if (!/^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(rawScenario.id)) fail(`${label}.id is not stable-safe`); + if (ids.has(rawScenario.id)) fail(`duplicate scenario id ${rawScenario.id}`); + ids.add(rawScenario.id); + assertNonEmptyString(rawScenario.sketch, `${label}.sketch`); + if (!references.sketches.has(rawScenario.sketch)) fail(`${label}.sketch is not a registered fixture`); + assertNonEmptyString(rawScenario.courseContent, `${label}.courseContent`); + if (rawScenario.courseContent !== "free" && !references.courseContentFixtures.has(rawScenario.courseContent)) { + fail(`${label}.courseContent is not a registered fixture`); + } + if (rawScenario.model !== undefined) { + assertNonEmptyString(rawScenario.model, `${label}.model`); + if (rawScenario.model === "auto") fail(`${label}.model must be a fixed model id`); + } + const turns = parseTurns(rawScenario.turns, label); + const expected = parseExpected(rawScenario.expected, label); + return { + id: rawScenario.id, + sketch: rawScenario.sketch, + courseContent: rawScenario.courseContent, + ...(rawScenario.model === undefined ? {} : { model: rawScenario.model }), + turns, + ...(expected === undefined ? {} : { expected }), + } satisfies TutorQualityScenarioSource; +} + export function parseTutorQualityCorpus( source: TutorQualityCorpusSource, references: TutorQualityCorpusReferences, @@ -156,48 +224,7 @@ export function parseTutorQualityCorpus( if (!Array.isArray(source.scenarios) || source.scenarios.length === 0) fail("scenarios must be a non-empty array"); const ids = new Set(); - const scenarios = source.scenarios.map((rawScenario, index) => { - const label = `scenarios[${index}]`; - assertObject(rawScenario, label); - assertNonEmptyString(rawScenario.id, `${label}.id`); - if (!/^[A-Za-z0-9][A-Za-z0-9._-]{0,63}$/.test(rawScenario.id)) fail(`${label}.id is not stable-safe`); - if (ids.has(rawScenario.id)) fail(`duplicate scenario id ${rawScenario.id}`); - ids.add(rawScenario.id); - assertNonEmptyString(rawScenario.sketch, `${label}.sketch`); - if (!references.sketches.has(rawScenario.sketch)) fail(`${label}.sketch is not a registered fixture`); - assertNonEmptyString(rawScenario.courseContent, `${label}.courseContent`); - if (rawScenario.courseContent !== "free" && !references.courseContentFixtures.has(rawScenario.courseContent)) { - fail(`${label}.courseContent is not a registered fixture`); - } - if (rawScenario.model !== undefined) { - assertNonEmptyString(rawScenario.model, `${label}.model`); - if (rawScenario.model === "auto") fail(`${label}.model must be a fixed model id`); - } - if (!Array.isArray(rawScenario.turns) || rawScenario.turns.length === 0) fail(`${label}.turns must be a non-empty array`); - const turns = rawScenario.turns.map((turn, turnIndex) => parseTurn(turn, `${label}.turns[${turnIndex}]`)); - turns.forEach((turn, turnIndex) => { - if (turn.kind === "dialog" && turn.continuationOf !== undefined && turn.continuationOf >= turnIndex) { - fail(`${label}.turns[${turnIndex}].continuationOf must reference a preceding turn`); - } - }); - if (rawScenario.expected !== undefined) { - assertObject(rawScenario.expected, `${label}.expected`); - if (rawScenario.expected.learningPhase !== undefined && !["LEARN", "DEEPEN", "EXPAND"].includes(rawScenario.expected.learningPhase as string)) { - fail(`${label}.expected.learningPhase is invalid`); - } - if (rawScenario.expected.questionNotRepeat !== undefined && rawScenario.expected.questionNotRepeat !== "exact-or-heuristic") { - fail(`${label}.expected.questionNotRepeat is invalid`); - } - } - return { - id: rawScenario.id, - sketch: rawScenario.sketch, - courseContent: rawScenario.courseContent, - ...(rawScenario.model === undefined ? {} : { model: rawScenario.model }), - turns, - ...(rawScenario.expected === undefined ? {} : { expected: rawScenario.expected as TutorQualityScenarioSource["expected"] }), - } satisfies TutorQualityScenarioSource; - }); + const scenarios = source.scenarios.map((rawScenario, index) => parseScenario(rawScenario, index, references, ids)); const normalized = { corpusId: source.corpusId, corpusVersion: source.corpusVersion, scenarios } satisfies TutorQualityCorpusSource; return { ...normalized, digest: tutorQualityCorpusDigest(normalized) }; diff --git a/server/services/tutor/evaluation/anchor-course-content.ts b/server/services/tutor/evaluation/anchor-course-content.ts index 34f0568e..4a16961e 100644 --- a/server/services/tutor/evaluation/anchor-course-content.ts +++ b/server/services/tutor/evaluation/anchor-course-content.ts @@ -152,20 +152,19 @@ function baseState(): TutorProgressionState { return createTutorProgressionState(REVISION); } +function variablesCourseContent(): TutorPlanningContentContext { + return { + revision: REVISION, + tutor: tutorCapability([variableTopic()]), + progressionState: baseState(), + }; +} + export function createAnchorCourseContent(fixtureId: AnchorCourseContentFixtureId): TutorPlanningContentContext { switch (fixtureId) { case "variables": - return { - revision: REVISION, - tutor: tutorCapability([variableTopic()]), - progressionState: baseState(), - }; case "progression-learn": - return { - revision: REVISION, - tutor: tutorCapability([variableTopic()]), - progressionState: baseState(), - }; + return variablesCourseContent(); case "progression-expand": { const state = baseState(); state.activeTopicId = "variables-and-serial"; diff --git a/server/services/tutor/evaluation/real-provider-evaluation.ts b/server/services/tutor/evaluation/real-provider-evaluation.ts index be8de5bd..bb1614f7 100644 --- a/server/services/tutor/evaluation/real-provider-evaluation.ts +++ b/server/services/tutor/evaluation/real-provider-evaluation.ts @@ -21,6 +21,11 @@ export type TutorQualityExecutionStatus = "completed" | "invalid" | "technical-f export const MAX_TUTOR_QUALITY_SAMPLES = 20; export const MAX_TUTOR_QUALITY_CALLS = 500; +const EMPTY_PROVIDER_CALLS: TutorQualityProviderCallCounts = { + total: 0, + modelListCalls: 0, + generationCalls: 0, +}; export type TutorQualityTurn = | { @@ -312,7 +317,7 @@ function makeEvaluationIdentity(options: TutorQualityEvaluationOptions): string const first = options.scenarios[0]; return hashCanonical({ unosimGitSha: options.git.sha, - courseContentRevisions: [...new Set(options.scenarios.map(courseRevision))].sort(), + courseContentRevisions: [...new Set(options.scenarios.map(courseRevision))].sort((left, right) => left.localeCompare(right)), corpusId: first?.corpusId, corpusVersion: first?.corpusVersion, providerId: options.providerId, @@ -343,7 +348,7 @@ function metadata( returnedModels: readonly string[] = [], sampleStartedAt = new Date(0).toISOString(), sampleDurationMs = 0, - providerCalls: TutorQualityProviderCallCounts = { total: 0, modelListCalls: 0, generationCalls: 0 }, + providerCalls: TutorQualityProviderCallCounts = EMPTY_PROVIDER_CALLS, ): TutorQualityMetadata { return { providerId: options.providerId, @@ -480,43 +485,63 @@ function deterministicRawChecks( } } -function expectedChecks( - scenario: TutorQualityEvaluationScenario, - result: TutorContentResult | undefined, - turn: TutorQualityTurn, - stateBefore: TutorProgressionState | undefined, - stateAfter: TutorProgressionState | undefined, - checks: TutorQualityDeterministicCheck[], - violations: TutorQualityInvariantViolation[], - turnIndex: number, -): void { - const expected = scenario.expected; - if (!expected || !result) return; - if (expected.topicId !== undefined) { - const passed = result.topicId === expected.topicId; - addCheck(checks, "expected-topic", passed, expected.topicId); - if (!passed) addViolation(violations, "topic-mismatch", "final-tutor", turnIndex, expected.topicId); - } - if (expected.topicIdAbsent !== undefined) { - const passed = result.topicId !== expected.topicIdAbsent && result.activeTopicId !== expected.topicIdAbsent; - addCheck(checks, "expected-topic-absent", passed, expected.topicIdAbsent); - if (!passed) addViolation(violations, "forbidden-topic-activation", "final-tutor", turnIndex, expected.topicIdAbsent); - } - if (expected.learningPhase !== undefined) { - const passed = result.learningPhase === expected.learningPhase || stateAfter?.phase === expected.learningPhase; - addCheck(checks, "expected-learning-phase", passed, expected.learningPhase); - if (!passed) addViolation(violations, "phase-mismatch", "final-tutor", turnIndex, expected.learningPhase); - } - if (expected.stateUnchanged && stateBefore !== undefined && stateAfter !== undefined) { - const passed = stableJson(stateBefore) === stableJson(stateAfter); - addCheck(checks, "expected-state-unchanged", passed); - if (!passed) addViolation(violations, "state-changed-unexpectedly", "state", turnIndex); - } - if (expected.questionNotRepeat === "exact-or-heuristic" && turn.kind === "dialog") { - const passed = !isSemanticallyRepeatedQuestion(result.question, [turn.question, ...(turn.history ?? []).map(({ question }) => question)]); - addCheck(checks, "final-question-not-repeated", passed); - if (!passed) addViolation(violations, "question-repeat", "final-tutor", turnIndex, "stage1-heuristic"); - } +interface ExpectedCheckContext { + readonly scenario: TutorQualityEvaluationScenario; + readonly result: TutorContentResult; + readonly turn: TutorQualityTurn; + readonly stateBefore: TutorProgressionState | undefined; + readonly stateAfter: TutorProgressionState | undefined; + readonly checks: TutorQualityDeterministicCheck[]; + readonly violations: TutorQualityInvariantViolation[]; + readonly turnIndex: number; +} + +function expectedTopicCheck(context: ExpectedCheckContext): void { + const expectedTopic = context.scenario.expected?.topicId; + if (expectedTopic === undefined) return; + const passed = context.result.topicId === expectedTopic; + addCheck(context.checks, "expected-topic", passed, expectedTopic); + if (!passed) addViolation(context.violations, "topic-mismatch", "final-tutor", context.turnIndex, expectedTopic); +} + +function expectedTopicAbsentCheck(context: ExpectedCheckContext): void { + const forbiddenTopic = context.scenario.expected?.topicIdAbsent; + if (forbiddenTopic === undefined) return; + const passed = context.result.topicId !== forbiddenTopic && context.result.activeTopicId !== forbiddenTopic; + addCheck(context.checks, "expected-topic-absent", passed, forbiddenTopic); + if (!passed) addViolation(context.violations, "forbidden-topic-activation", "final-tutor", context.turnIndex, forbiddenTopic); +} + +function expectedPhaseCheck(context: ExpectedCheckContext): void { + const expectedPhase = context.scenario.expected?.learningPhase; + if (expectedPhase === undefined) return; + const passed = context.result.learningPhase === expectedPhase || context.stateAfter?.phase === expectedPhase; + addCheck(context.checks, "expected-learning-phase", passed, expectedPhase); + if (!passed) addViolation(context.violations, "phase-mismatch", "final-tutor", context.turnIndex, expectedPhase); +} + +function expectedStateCheck(context: ExpectedCheckContext): void { + if (!context.scenario.expected?.stateUnchanged || !context.stateBefore || !context.stateAfter) return; + const passed = stableJson(context.stateBefore) === stableJson(context.stateAfter); + addCheck(context.checks, "expected-state-unchanged", passed); + if (!passed) addViolation(context.violations, "state-changed-unexpectedly", "state", context.turnIndex); +} + +function expectedQuestionCheck(context: ExpectedCheckContext): void { + if (context.scenario.expected?.questionNotRepeat !== "exact-or-heuristic" || context.turn.kind !== "dialog") return; + const previousQuestions = [context.turn.question, ...(context.turn.history ?? []).map(({ question }) => question)]; + const passed = !isSemanticallyRepeatedQuestion(context.result.question, previousQuestions); + addCheck(context.checks, "final-question-not-repeated", passed); + if (!passed) addViolation(context.violations, "question-repeat", "final-tutor", context.turnIndex, "stage1-heuristic"); +} + +function expectedChecks(context: ExpectedCheckContext): void { + if (!context.scenario.expected) return; + expectedTopicCheck(context); + expectedTopicAbsentCheck(context); + expectedPhaseCheck(context); + expectedStateCheck(context); + expectedQuestionCheck(context); } function applicationMetadataChecks( @@ -587,6 +612,247 @@ function sampleTemplate( }; } +interface SampleTurnContext { + readonly options: TutorQualityEvaluationOptions; + readonly provider: CountingProvider; + readonly service: TutorService; + readonly scenario: TutorQualityEvaluationScenario; + readonly content: TutorPlanningContentContext | undefined; + readonly stateBefore: TutorProgressionState | undefined; + readonly finalQuestions: Map; + readonly violations: TutorQualityInvariantViolation[]; + readonly turn: TutorQualityTurn; + readonly turnIndex: number; +} + +interface InvokedTutorTurn { + readonly finalResult?: TutorContentResult; + readonly error?: TutorQualityTechnicalError; + readonly invalidReason?: string; +} + +interface ProcessedCapture { + readonly returnedModel?: string; + readonly invalidReason?: string; +} + +interface SampleTurnOutcome { + readonly transcriptTurn?: TutorQualityTranscriptTurn; + readonly returnedModel?: string; + readonly executionStatus?: Extract; + readonly invalidReason?: string; + readonly terminalError?: TutorQualityTechnicalError; + readonly stop: boolean; +} + +function invalidTurnContext(context: SampleTurnContext): string | undefined { + const { turn, turnIndex, finalQuestions, violations } = context; + if (turn.kind === "initial") return undefined; + if (turn.bindsToQuestion !== turn.question) { + addViolation(violations, "unbound-question-context", "scenario", turnIndex); + return "unbound-question-context"; + } + if (turn.continuationOf !== undefined) { + const precedingQuestion = finalQuestions.get(turn.continuationOf); + if (precedingQuestion === undefined || precedingQuestion !== turn.bindsToQuestion) { + addViolation(violations, "preceding-question-mismatch", "scenario", turnIndex); + return "preceding-question-mismatch"; + } + } + return undefined; +} + +async function invokeTutorTurn(context: SampleTurnContext): Promise { + const { options, service, scenario, content, turn } = context; + const invalidReason = invalidTurnContext(context); + if (invalidReason) return { invalidReason }; + try { + if (turn.kind === "initial") { + const response = await service.generateQuestion(scenario.sketch, options.credential, options.requestedModel, turn.difficulty ?? 30, content); + return { finalResult: response.result }; + } + const response = await service.generateDialogResponse( + scenario.sketch, + historyFromSource(turn.history), + turn.question, + turn.answer, + options.credential, + options.requestedModel, + turn.difficulty ?? 30, + content, + ); + return { finalResult: response.result }; + } catch (error_) { + return { error: technicalError(error_) }; + } +} + +function latestCapture(provider: CountingProvider, beforeGenerationCount: number): ProviderCapture | undefined { + return provider.generationCaptures.length > beforeGenerationCount + ? provider.generationCaptures.at(-1) + : undefined; +} + +function processCapture( + capture: ProviderCapture | undefined, + context: SampleTurnContext, + checks: TutorQualityDeterministicCheck[], +): ProcessedCapture { + if (!capture?.response) return {}; + const returnedModel = capture.response.model; + if (typeof returnedModel === "string" && returnedModel.length > 0) { + if (returnedModel !== context.options.requestedModel) { + addCheck(checks, "returned-model-matches-request", false, returnedModel); + deterministicRawChecks(capture.response.result, context.turn, context.turnIndex, checks, context.violations); + return { returnedModel, invalidReason: "returned-model-mismatch" }; + } + addCheck(checks, "returned-model-matches-request", true); + deterministicRawChecks(capture.response.result, context.turn, context.turnIndex, checks, context.violations); + return { returnedModel }; + } + addCheck(checks, "returned-model-matches-request", false, "missing"); + deterministicRawChecks(capture.response.result, context.turn, context.turnIndex, checks, context.violations); + return { invalidReason: "returned-model-missing" }; +} + +function processFinalResult( + context: SampleTurnContext, + invocation: InvokedTutorTurn, + checks: TutorQualityDeterministicCheck[], +): void { + const { finalResult, error } = invocation; + if (!finalResult) { + if (error) addCheck(checks, "final-tutor-response-present", false, error.kind); + return; + } + context.finalQuestions.set(context.turnIndex, finalResult.question); + addCheck(checks, "final-tutor-response-present", true); + expectedChecks({ + scenario: context.scenario, + result: finalResult, + turn: context.turn, + stateBefore: context.stateBefore, + stateAfter: context.content?.progressionState, + checks, + violations: context.violations, + turnIndex: context.turnIndex, + }); + applicationMetadataChecks(finalResult, context.content, context.content?.progressionState, checks, context.violations, context.turnIndex); + questionIdReuseCheck(finalResult, context.turn, context.content, checks, context.violations, context.turnIndex); +} + +function buildTranscriptTurn( + context: SampleTurnContext, + startedAt: Date, + callsBefore: TutorQualityProviderCallCounts, + finishedAt: Date, + capture: ProviderCapture | undefined, + invocation: InvokedTutorTurn, + checks: readonly TutorQualityDeterministicCheck[], +): TutorQualityTranscriptTurn { + const { options, provider, turn, turnIndex } = context; + return { + index: turnIndex, + input: safeTurn(turn, options.credential), + startedAt: startedAt.toISOString(), + durationMs: Math.max(0, finishedAt.getTime() - startedAt.getTime()), + providerCalls: subtractCounts(provider.counts, callsBefore), + ...(capture?.request ? { providerRequest: requestArtifact(capture.request, options.credential) } : {}), + ...(capture?.response ? { + rawProviderResult: safeResult(capture.response.result, options.credential), + ...(typeof capture.response.model === "string" ? { returnedModel: redact(capture.response.model, options.credential) } : {}), + } : {}), + ...(invocation.finalResult ? { finalTutorResult: safeResult(invocation.finalResult, options.credential) } : {}), + deterministicChecks: checks, + ...(invocation.error ? { technicalError: invocation.error } : {}), + }; +} + +async function executeSampleTurn(context: SampleTurnContext): Promise { + const startedAt = (context.options.now ?? (() => new Date()))(); + const callsBefore = context.provider.counts; + const checks: TutorQualityDeterministicCheck[] = []; + const beforeGenerationCount = context.provider.generationCaptures.length; + const invocation = await invokeTutorTurn(context); + if (invocation.invalidReason) return { invalidReason: invocation.invalidReason, executionStatus: "invalid", stop: true }; + const capture = latestCapture(context.provider, beforeGenerationCount); + const captureResult = processCapture(capture, context, checks); + processFinalResult(context, invocation, checks); + const finishedAt = (context.options.now ?? (() => new Date()))(); + return { + transcriptTurn: buildTranscriptTurn(context, startedAt, callsBefore, finishedAt, capture, invocation, checks), + returnedModel: captureResult.returnedModel, + ...(invocation.error ? { executionStatus: "technical-failure" as const, terminalError: invocation.error } : {}), + ...(captureResult.invalidReason ? { executionStatus: "invalid" as const, invalidReason: captureResult.invalidReason } : {}), + stop: invocation.error !== undefined, + }; +} + +interface SampleExecutionState { + readonly violations: TutorQualityInvariantViolation[]; + readonly turns: TutorQualityTranscriptTurn[]; + readonly returnedModels: string[]; + readonly finalQuestions: Map; + executionStatus: TutorQualityExecutionStatus; + invalidReason: string | undefined; + terminalError: TutorQualityTechnicalError | undefined; +} + +function applySampleTurnOutcome(state: SampleExecutionState, outcome: SampleTurnOutcome): void { + if (outcome.executionStatus === "technical-failure") state.executionStatus = "technical-failure"; + if (outcome.executionStatus === "invalid") state.executionStatus = "invalid"; + if (outcome.invalidReason) state.invalidReason = state.invalidReason ?? outcome.invalidReason; + if (outcome.terminalError) state.terminalError = outcome.terminalError; + if (outcome.returnedModel) state.returnedModels.push(outcome.returnedModel); + if (outcome.transcriptTurn) state.turns.push(outcome.transcriptTurn); +} + +interface SampleTurnsContext { + readonly options: TutorQualityEvaluationOptions; + readonly provider: CountingProvider; + readonly service: TutorService; + readonly scenario: TutorQualityEvaluationScenario; + readonly content: TutorPlanningContentContext | undefined; + readonly stateBefore: TutorProgressionState | undefined; + readonly execution: SampleExecutionState; +} + +async function executeSampleTurns(context: SampleTurnsContext): Promise { + const { options, provider, service, scenario, content, stateBefore, execution } = context; + for (const [turnIndex, turn] of scenario.turns.entries()) { + const outcome = await executeSampleTurn({ + options, + provider, + service, + scenario, + content, + stateBefore, + finalQuestions: execution.finalQuestions, + violations: execution.violations, + turn, + turnIndex, + }); + applySampleTurnOutcome(execution, outcome); + if (outcome.stop) break; + } +} + +function recordFailureStateCheck( + terminalError: TutorQualityTechnicalError | undefined, + stateBefore: TutorProgressionState | undefined, + stateAfter: TutorProgressionState | undefined, + turns: TutorQualityTranscriptTurn[], + violations: TutorQualityInvariantViolation[], +): void { + if (!terminalError || !stateBefore || !stateAfter) return; + const unchanged = stableJson(stateBefore) === stableJson(stateAfter); + const stateChecks = turns.at(-1)?.deterministicChecks; + if (stateChecks) { + (stateChecks as TutorQualityDeterministicCheck[]).push({ name: "state-unchanged-after-failure", passed: unchanged }); + } + if (!unchanged) addViolation(violations, "state-mutated-after-failure", "state", undefined); +} + async function runSample( options: TutorQualityEvaluationOptions, provider: CountingProvider, @@ -599,117 +865,20 @@ async function runSample( const callsBeforeSample = provider.counts; const content = scenario.courseContent ? clone(scenario.courseContent) : undefined; const stateBefore = content?.progressionState ? clone(content.progressionState) : undefined; - const violations: TutorQualityInvariantViolation[] = []; - const turns: TutorQualityTranscriptTurn[] = []; - const returnedModels: string[] = []; - const finalQuestions = new Map(); - let executionStatus: TutorQualityExecutionStatus = "completed"; - let invalidReason: string | undefined; - let terminalError: TutorQualityTechnicalError | undefined; const service = new TutorService(provider, content ? new CurriculumTutorAdapter() : undefined); - - for (const [turnIndex, turn] of scenario.turns.entries()) { - const turnStartedAt = (options.now ?? (() => new Date()))(); - const callsBeforeTurn = provider.counts; - const turnChecks: TutorQualityDeterministicCheck[] = []; - const beforeGenerationCount = provider.generationCaptures.length; - let finalResult: TutorContentResult | undefined; - let error: TutorQualityTechnicalError | undefined; - try { - if (turn.kind === "initial") { - const response = await service.generateQuestion(scenario.sketch, options.credential, options.requestedModel, turn.difficulty ?? 30, content); - finalResult = response.result; - } else { - if (turn.bindsToQuestion !== turn.question) { - executionStatus = "invalid"; - invalidReason = "unbound-question-context"; - addViolation(violations, "unbound-question-context", "scenario", turnIndex); - break; - } - if (turn.continuationOf !== undefined) { - const precedingQuestion = finalQuestions.get(turn.continuationOf); - if (precedingQuestion === undefined || precedingQuestion !== turn.bindsToQuestion) { - executionStatus = "invalid"; - invalidReason = "preceding-question-mismatch"; - addViolation(violations, "preceding-question-mismatch", "scenario", turnIndex); - break; - } - } - const response = await service.generateDialogResponse( - scenario.sketch, - historyFromSource(turn.history), - turn.question, - turn.answer, - options.credential, - options.requestedModel, - turn.difficulty ?? 30, - content, - ); - finalResult = response.result; - } - } catch (caught) { - error = technicalError(caught); - executionStatus = "technical-failure"; - terminalError = error; - } - - const capture = provider.generationCaptures.length > beforeGenerationCount - ? provider.generationCaptures.at(-1) - : undefined; - if (capture?.response) { - const returnedModel = capture.response.model; - if (typeof returnedModel !== "string" || returnedModel.length === 0) { - executionStatus = "invalid"; - invalidReason = "returned-model-missing"; - addCheck(turnChecks, "returned-model-matches-request", false, "missing"); - } else if (returnedModel !== options.requestedModel) { - executionStatus = "invalid"; - invalidReason = "returned-model-mismatch"; - addCheck(turnChecks, "returned-model-matches-request", false, returnedModel); - } else { - addCheck(turnChecks, "returned-model-matches-request", true); - } - if (typeof returnedModel === "string" && returnedModel.length > 0) returnedModels.push(returnedModel); - deterministicRawChecks(capture.response.result, turn, turnIndex, turnChecks, violations); - } - if (finalResult) { - finalQuestions.set(turnIndex, finalResult.question); - addCheck(turnChecks, "final-tutor-response-present", true); - expectedChecks(scenario, finalResult, turn, stateBefore, content?.progressionState, turnChecks, violations, turnIndex); - applicationMetadataChecks(finalResult, content, content?.progressionState, turnChecks, violations, turnIndex); - questionIdReuseCheck(finalResult, turn, content, turnChecks, violations, turnIndex); - } else if (error) { - addCheck(turnChecks, "final-tutor-response-present", false, error.kind); - } - const captureRequest = capture?.request; - const turnFinishedAt = (options.now ?? (() => new Date()))(); - turns.push({ - index: turnIndex, - input: safeTurn(turn, options.credential), - startedAt: turnStartedAt.toISOString(), - durationMs: Math.max(0, turnFinishedAt.getTime() - turnStartedAt.getTime()), - providerCalls: subtractCounts(provider.counts, callsBeforeTurn), - ...(captureRequest ? { providerRequest: requestArtifact(captureRequest, options.credential) } : {}), - ...(capture?.response ? { - rawProviderResult: safeResult(capture.response.result, options.credential), - ...(typeof capture.response.model === "string" ? { returnedModel: redact(capture.response.model, options.credential) } : {}), - } : {}), - ...(finalResult ? { finalTutorResult: safeResult(finalResult, options.credential) } : {}), - deterministicChecks: turnChecks, - ...(error ? { technicalError: error } : {}), - }); - if (error) break; - } + const execution: SampleExecutionState = { + violations: [], + turns: [], + returnedModels: [], + finalQuestions: new Map(), + executionStatus: "completed", + invalidReason: undefined, + terminalError: undefined, + }; + await executeSampleTurns({ options, provider, service, scenario, content, stateBefore, execution }); const stateAfter = content?.progressionState ? clone(content.progressionState) : undefined; - if (terminalError && stateBefore && stateAfter) { - const unchanged = stableJson(stateBefore) === stableJson(stateAfter); - const stateChecks = turns.at(-1)?.deterministicChecks; - if (stateChecks) { - (stateChecks as TutorQualityDeterministicCheck[]).push({ name: "state-unchanged-after-failure", passed: unchanged }); - } - if (!unchanged) addViolation(violations, "state-mutated-after-failure", "state", undefined); - } + recordFailureStateCheck(execution.terminalError, stateBefore, stateAfter, execution.turns, execution.violations); const sample = sampleTemplate(options, scenario, runId, evaluationIdentity, sampleIndex); const sampleFinishedAt = (options.now ?? (() => new Date()))(); const sampleCalls = subtractCounts(provider.counts, callsBeforeSample); @@ -719,19 +888,19 @@ async function runSample( options, scenario, sampleIndex, - returnedModels, + execution.returnedModels, sampleStartedAt.toISOString(), Math.max(0, sampleFinishedAt.getTime() - sampleStartedAt.getTime()), sampleCalls, ), ...(stateBefore ? { stateBefore } : {}), ...(stateAfter ? { stateAfter } : {}), - turns, - deterministicChecks: turns.flatMap(({ deterministicChecks }) => deterministicChecks), - executionStatus, - ...(invalidReason ? { invalidReason } : {}), - ...(terminalError ? { technicalError: terminalError } : {}), - invariantViolations: violations, + turns: execution.turns, + deterministicChecks: execution.turns.flatMap(({ deterministicChecks }) => deterministicChecks), + executionStatus: execution.executionStatus, + ...(execution.invalidReason ? { invalidReason: execution.invalidReason } : {}), + ...(execution.terminalError ? { technicalError: execution.terminalError } : {}), + invariantViolations: execution.violations, }; } @@ -821,16 +990,19 @@ function addTranscriptToAggregate( }; } -function baseReport( - options: TutorQualityEvaluationOptions, - runId: string, - evaluationIdentity: string, - runStatus: TutorQualityExecutionStatus, - reason: string | undefined, - calls: TutorQualityProviderCallCounts, - byScenario: Readonly>, - samplesObserved = 0, -): TutorQualityEvaluationReport { +interface BaseReportContext { + readonly options: TutorQualityEvaluationOptions; + readonly runId: string; + readonly evaluationIdentity: string; + readonly runStatus: TutorQualityExecutionStatus; + readonly reason?: string; + readonly calls: TutorQualityProviderCallCounts; + readonly byScenario: Readonly>; + readonly samplesObserved?: number; +} + +function baseReport(context: BaseReportContext): TutorQualityEvaluationReport { + const { options, runId, evaluationIdentity, runStatus, reason, calls, byScenario, samplesObserved = 0 } = context; const aggregates = Object.values(byScenario); const completed = aggregates.reduce((sum, item) => sum + item.completed, 0); const invalid = aggregates.reduce((sum, item) => sum + item.invalid, 0); @@ -901,7 +1073,7 @@ async function writeArtifacts( await writeArtifactsToDirectory(options.outputDir, report, transcripts); } -function invalidPreflightReason(options: TutorQualityEvaluationOptions): string | undefined { +function invalidBasicPreflightReason(options: TutorQualityEvaluationOptions): string | undefined { if (!options.requestedModel || options.requestedModel === "auto") return "fixed-model-required"; if (!options.git.sha) return "git-sha-missing"; if (!options.git.trackedClean || !options.git.relevantUntrackedClean) return "dirty-relevant-worktree"; @@ -909,6 +1081,10 @@ function invalidPreflightReason(options: TutorQualityEvaluationOptions): string if (options.samples > MAX_TUTOR_QUALITY_SAMPLES) return "sample-count-exceeds-limit"; if (!Number.isInteger(options.maxCalls) || options.maxCalls < 0) return "invalid-call-budget"; if (options.maxCalls > MAX_TUTOR_QUALITY_CALLS) return "call-budget-exceeds-limit"; + return undefined; +} + +function invalidCorpusPreflightReason(options: TutorQualityEvaluationOptions): string | undefined { if (options.scenarios.length === 0) return "empty-corpus"; const first = options.scenarios[0]; if (options.scenarios.some((scenario) => scenario.corpusId !== first?.corpusId || scenario.corpusVersion !== first?.corpusVersion)) { @@ -922,6 +1098,10 @@ function invalidPreflightReason(options: TutorQualityEvaluationOptions): string return undefined; } +function invalidPreflightReason(options: TutorQualityEvaluationOptions): string | undefined { + return invalidBasicPreflightReason(options) ?? invalidCorpusPreflightReason(options); +} + export async function runTutorQualityEvaluation(options: TutorQualityEvaluationOptions): Promise { const runId = makeRunId(options); const evaluationIdentity = makeEvaluationIdentity(options); @@ -929,17 +1109,17 @@ export async function runTutorQualityEvaluation(options: TutorQualityEvaluationO const emptyByScenario = Object.fromEntries(options.scenarios.map((scenario) => [scenario.id, emptyAggregate(options.samples)])); if (invalidReason) { - const report = baseReport(options, runId, evaluationIdentity, "invalid", invalidReason, { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); + const report = baseReport({ options, runId, evaluationIdentity, runStatus: "invalid", reason: invalidReason, calls: EMPTY_PROVIDER_CALLS, byScenario: emptyByScenario }); await writeArtifacts(options, report, []); return { report, transcripts: [] }; } if (!options.credential) { - const report = baseReport(options, runId, evaluationIdentity, "not-run", "missing-credential", { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); + const report = baseReport({ options, runId, evaluationIdentity, runStatus: "not-run", reason: "missing-credential", calls: EMPTY_PROVIDER_CALLS, byScenario: emptyByScenario }); await writeArtifacts(options, report, []); return { report, transcripts: [] }; } if (options.maxCalls === 0) { - const report = baseReport(options, runId, evaluationIdentity, "not-run", "call-budget-zero", { total: 0, modelListCalls: 0, generationCalls: 0 }, emptyByScenario); + const report = baseReport({ options, runId, evaluationIdentity, runStatus: "not-run", reason: "call-budget-zero", calls: EMPTY_PROVIDER_CALLS, byScenario: emptyByScenario }); await writeArtifacts(options, report, []); return { report, transcripts: [] }; } @@ -949,12 +1129,12 @@ export async function runTutorQualityEvaluation(options: TutorQualityEvaluationO try { availableModels = await provider.listModels(options.credential); } catch (error) { - const report = baseReport(options, runId, evaluationIdentity, "technical-failure", technicalError(error).kind, provider.counts, emptyByScenario); + const report = baseReport({ options, runId, evaluationIdentity, runStatus: "technical-failure", reason: technicalError(error).kind, calls: provider.counts, byScenario: emptyByScenario }); await writeArtifacts(options, report, []); return { report, transcripts: [] }; } if (!availableModels.includes(options.requestedModel)) { - const report = baseReport(options, runId, evaluationIdentity, "invalid", "model-unavailable", provider.counts, emptyByScenario); + const report = baseReport({ options, runId, evaluationIdentity, runStatus: "invalid", reason: "model-unavailable", calls: provider.counts, byScenario: emptyByScenario }); await writeArtifacts(options, report, []); return { report, transcripts: [] }; } @@ -968,13 +1148,15 @@ export async function runTutorQualityEvaluation(options: TutorQualityEvaluationO const callsBeforeSample = provider.counts; const transcript = await runSample(options, provider, scenario, runId, evaluationIdentity, sampleIndex); transcripts.push(transcript); - byScenario[scenario.id] = addTranscriptToAggregate(byScenario[scenario.id]!, transcript, subtractCounts(provider.counts, callsBeforeSample)); + const aggregate = byScenario[scenario.id]; + if (!aggregate) throw new Error(`Missing scenario aggregate for ${scenario.id}`); + byScenario[scenario.id] = addTranscriptToAggregate(aggregate, transcript, subtractCounts(provider.counts, callsBeforeSample)); if (transcript.executionStatus === "invalid" && transcript.invalidReason === "returned-model-mismatch") { // The mismatch belongs to this sample; subsequent samples remain observable. } } } - const report = baseReport(options, runId, evaluationIdentity, "completed", undefined, provider.counts, byScenario, transcripts.length); + const report = baseReport({ options, runId, evaluationIdentity, runStatus: "completed", calls: provider.counts, byScenario, samplesObserved: transcripts.length }); await writeArtifacts(options, report, transcripts); return { report, transcripts }; } diff --git a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts index a7abf49c..39269203 100644 --- a/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts +++ b/tests/server/services/tutor/evaluation/tutor-quality-real-provider-cli.test.ts @@ -8,6 +8,7 @@ import { loadTutorQualityEvaluationScenarios, isTutorQualityRelevantUntrackedPath, parseTutorQualityCliArgs, + readTutorQualityGitState, runTutorQualityCli, } from "../../../../../scripts/tutor-quality-real-provider-eval"; import type { LLMProvider } from "../../../../../server/services/tutor/llm-provider"; @@ -45,6 +46,11 @@ describe("Tutor Quality real-provider CLI contract", () => { expect(isTutorQualityRelevantUntrackedPath(".tutor-quality-output/report.json", ".tutor-quality-output")).toBe(false); }); + it("reads the local Git state through a fixed executable location", () => { + const state = readTutorQualityGitState(process.cwd(), "test-results"); + expect(state.sha).toMatch(/^[0-9a-f]{40}$/); + }); + it("materializes and runs the complete versioned anchor corpus with a fake provider", async () => { const scenarios = await loadTutorQualityEvaluationScenarios(process.cwd(), "evals/tutor-quality/anchor-corpus.yaml"); const outputDir = await mkdtemp(path.join(os.tmpdir(), "unosim-tq-cli-test-"));