Skip to content

Run reports and the behaviour check on a local model - #110

Open
alanhc wants to merge 5 commits into
sysprog21:mainfrom
alanhc:shim-tools
Open

alanhc wants to merge 5 commits into
sysprog21:mainfrom
alanhc:shim-tools

Conversation

@alanhc

@alanhc alanhc commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

This now carries #99 as well, as asked there. Main had moved a long way since both branches were cut (multi-key reports, doubling backoffs, report recovery, the played-candidate check from #107), so rather than replay 16 commits over it, the work is rebuilt on current main as five commits. The old history is still readable on #99.

Summary

The report, the quiet-pause reviews and the interviewer behaviour check can now run on a model on the operator's own GPU instead of Gemini. The live interviewer is untouched and still talks to Gemini.

  • CODETRIAL_GEMINI_REST_BASE, read from the environment like INTERVIEW_ROOM_NAME, points every generateContent call at another server. Unset, the URL is what it was.
  • scripts/gemini-shim.py (standard library only) answers generateContent from llama-server's OpenAI-compatible endpoint. It converts the response schema to the JSON Schema llama.cpp compiles to a grammar, carries function declarations and calls both ways, passes upstream status codes through so a 503 still reads as a 503, and gives up on llama-server when the caller would. A request with no output limit gets one, and --thinking off turns thinking off for every request, which Gemma 4 needs for the behaviour check. tests/test_gemini_shim.py runs in the gate.
  • Longer deadlines for a local base only. Any base other than Google's gets 45 seconds a call and 250 for the report, against 20 and 125; 250 is what five calls and the doubling backoffs need. Only the server knows which is in force, so the page's wait before offering to leave and its wait on a regenerated report now come from /runtime-config.js, and the page's own values stay the hosted ones as floors.
  • Repairs that say what to fix. When a report names the published problem, the repair names the title and the scenario name to use; when the plan and the feedback disagree, it names each wrong item by index, each improvement with no item, and the swap when there is one of each. Told only that a rule failed, the local model sent back the same output byte for byte. None of this reaches the candidate's failure note.
  • The behaviour check follows the base, so pointed at the shim it sends what production sends. BEHAVIOR_LOCAL_BASE from Hold played candidates to the interview rules #107 stays as the direct route, without the shim.

Numbers

RTX 5070 Ti, gemma-4-12b Q4_K_M on llama.cpp, shim with --thinking off:

  • Reports on the Two Sum golden prompt, this branch: 4 of 4 valid, 14 to 16 seconds each. On the old branches, the plan repair took this from 0 of 6 to 10 of 10, and the title repair from 2 of 4 to 4 of 4.
  • Behaviour check through the shim: the Two Sum script passed in 9 seconds. On the old branch, 18 of 21 problem runs across seven runs; each miss was a second hint request answered without calling log_hint.

One prompt, one machine: read these as "this works end to end", not as a comparison with Gemini.

Test plan

  • The report budget test holds both deadline pairs to five calls plus backoffs; the escape-hatch and recovery-limit tests hold the page's floors to the hosted deadline and the server's values to both.
  • binary_web_gives_a_local_report_base_the_longer_wait starts the binary with the base set and checks /runtime-config.js says 260000 and 265; without it, the runtime-config test pins 135000 and 140.
  • report-recovery.test.js: the server can raise the retry wait and cannot lower it.
  • Repair guidance: the title test, and seven plan tests (reworded, repeated, several at once, counted once across sections, a missing weakness keeping indexes, a matching plan getting none, guidance staying out of the failure note).
  • a_local_model_writes_a_report is an ignored unit test that makes one real report through the base.
  • ./scripts/test.sh passes locally: 844 JS and 1228 Rust tests, none skipped, formatting clean. Every commit builds clean under clippy. cargo mutants --in-diff on this diff: 42 mutants, 38 caught, 4 unviable, none missed. cargo-audit and shellcheck were not installed here.

Not tested: a full interview end to end with the report coming from a local model, and anything against Gemini.

alanhc added 5 commits October 3, 2026 14:04
The report and interim-review calls always went to Google, so trying
a model on one's own GPU meant editing the URL by hand. They now go to
CODETRIAL_GEMINI_REST_BASE when it is set, read from the process
environment like INTERVIEW_ROOM_NAME; unset, nothing changes. The live
socket is untouched.

scripts/gemini-shim.py answers generateContent from llama-server's
OpenAI-compatible endpoint. It turns the Gemini schema into the JSON
Schema llama.cpp compiles to a grammar, carries function declarations
and calls both ways, passes upstream status codes through so the retry
rules still see a 503 as a 503, and gives up on llama-server when the
caller would. It caps a request that names no output limit, so a model
stuck repeating itself ends as MAX_TOKENS instead of out-waiting the
shim, and --thinking off turns thinking off for every request, which
Gemma 4 needs for a check that names no thinking budget.
tests/test_gemini_shim.py covers the mapping and runs in the gate.

An ignored unit test makes one real report through the base. Against
gemma-4-12b on an RTX 5070 Ti it wrote a valid report in 16 seconds.
A 9B to 14B model on one consumer GPU takes 14 to 32 seconds per
report call, against a 20-second attempt limit sized for Gemini, so a
local model timed out on most calls. When CODETRIAL_GEMINI_REST_BASE
points anywhere but Google, the attempt limit is 45 seconds and the
report deadline 250, so five attempts and their doubling backoffs
still fit; Gemini keeps 20 and 125. Only the server knows which
deadline is in force, so the page's wait before offering to leave and
its wait on a regenerated report now come from /runtime-config.js, and
the page's own values stay the hosted ones as floors.
A report whose improvement plan called the exercise "a 'Two Sum'
style problem" was sent back with only "names the published problem"
and the field's path. The model could not tell which words broke the
rule, wrote the same sentence twice more, and the report was lost.

The repair prompt now names the published title and the scenario's
title to use instead. The title stays out of the error itself: that
error is the failure note the candidate reads, and the title is what
it must not show them. The original prompt already carries the title,
so the model learns nothing new from it.

Against gemma-4-12b on the Two Sum weak-candidate prompt, 2 of 4
reports passed before this and 4 of 4 after, three of them on the
first repair.
When the improvement plan and the feedback disagreed, the repair was
told only that they did, and a local 12B model sent back the same plan
byte for byte on every repair: told that something in a list of four
was wrong, it could not find which. The usual cause is a weakness
reworded on its way into the plan, or a repeat standing where an
improvement was left out. The repair now names each wrong item by
index and each improvement with no item, and when there is one of
each, the swap. Improvements are counted once each, the way the
validator counts them, and an item without a weakness keeps its index.
The note the candidate reads is unchanged.

Against gemma-4-12b on the Two Sum report prompt, 0 of 6 reports
passed before this and 10 of 10 after; 7 of the 10 broke the plan on
their first attempt and were repaired.
The scripted check posted to Google's URL by hand, so it could not
follow CODETRIAL_GEMINI_REST_BASE the way the report does. It now
builds its URL with gemini_generate_content_url, so pointed at
scripts/gemini-shim.py it sends what production sends through the
shim the report uses, tools included. The direct route the played
candidates added, BEHAVIOR_LOCAL_BASE, stays for talking to an
OpenAI-compatible server without the shim. README and
docs/development.md describe both. Against gemma-4-12b through the
shim with --thinking off, the Two Sum script passed in 9 seconds.
@alanhc alanhc changed the title Let the behaviour check run against a local model Run reports and the behaviour check on a local model Oct 3, 2026
@alanhc
alanhc marked this pull request as ready for review October 3, 2026 06:21

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

7 issues found across 19 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="tests/web/assets.rs">

<violation number="1" location="tests/web/assets.rs:446">
P2: This branch accepts hosted waits even for a non-default local base, so it will not catch the regression this test is meant to detect. Distinguish the effective base from the default Google URL and assert the exact expected wait tuple.</violation>
</file>

<file name="scripts/gemini-shim.py">

<violation number="1" location="scripts/gemini-shim.py:37">
P2: Interim-review callers give up after 12 seconds, but this keeps their llama request running for up to 45 seconds. Make the upstream deadline per-request so abandoned pause reviews stop consuming local-model capacity.</violation>

<violation number="2" location="scripts/gemini-shim.py:196">
P2: The report and interim requests set a seed for repeatable outputs, but this conversion drops `generationConfig.seed`. Forward it to the local chat request so repeated prompts retain the configured determinism.</violation>
</file>

<file name="web/interview.js">

<violation number="1" location="web/interview.js:2410">
P2: If `/runtime-config.js` is unavailable, this falls back to 135 seconds even though local report generation can run for 250 seconds, exposing the leave/offline-report options before it can finish. Use a local-safe fallback when the server value is absent.</violation>
</file>

<file name="web/report-recovery.js">

<violation number="1" location="web/report-recovery.js:66">
P2: If `/runtime-config.js` is unavailable, the missing value leaves this at 140 seconds, so `finish()` finalizes the incomplete report before a local retry's 250-second deadline; a later report is ignored. Use a local-safe fallback only when the server value is absent.</violation>
</file>

<file name="README.md">

<violation number="1" location="README.md:243">
P2: The referenced setup command omits `--jinja`, which the shim requires for function calls; the behaviour check then cannot exercise its tools. Include `--jinja` in the `llama-server` command.</violation>
</file>

<file name="src/gemini.rs">

<violation number="1" location="src/gemini.rs:1185">
P2: This comparison flags weaknesses that `validate_report_candidate` already accepts after `snap_plan_weaknesses` normalizes case, spacing, or a final period. When another field needs repair, the prompt tells the model to rewrite a valid plan item; derive guidance from the snapped values or suppress these false mismatches.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread tests/web/assets.rs
assert!(
matches!(
(escape_wait_ms, retry_wait_s),
(135_000, 140) | (260_000, 265)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This branch accepts hosted waits even for a non-default local base, so it will not catch the regression this test is meant to detect. Distinguish the effective base from the default Google URL and assert the exact expected wait tuple.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At tests/web/assets.rs, line 446:

<comment>This branch accepts hosted waits even for a non-default local base, so it will not catch the regression this test is meant to detect. Distinguish the effective base from the default Google URL and assert the exact expected wait tuple.</comment>

<file context>
@@ -430,6 +430,29 @@ async fn static_server_serves_health_fixture_and_missing_asset() {
+        assert!(
+            matches!(
+                (escape_wait_ms, retry_wait_s),
+                (135_000, 140) | (260_000, 265)
+            ),
+            "{escape_wait_ms} {retry_wait_s}"
</file context>

Comment thread scripts/gemini-shim.py
# answer nobody would read while the retry queued behind it. Closing the
# upstream socket is what stops llama-server, within about two seconds, so the
# default matches that deadline; --upstream-timeout changes it.
UPSTREAM_TIMEOUT_S = 45

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Interim-review callers give up after 12 seconds, but this keeps their llama request running for up to 45 seconds. Make the upstream deadline per-request so abandoned pause reviews stop consuming local-model capacity.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At scripts/gemini-shim.py, line 37:

<comment>Interim-review callers give up after 12 seconds, but this keeps their llama request running for up to 45 seconds. Make the upstream deadline per-request so abandoned pause reviews stop consuming local-model capacity.</comment>

<file context>
@@ -0,0 +1,400 @@
+# answer nobody would read while the retry queued behind it. Closing the
+# upstream socket is what stops llama-server, within about two seconds, so the
+# default matches that deadline; --upstream-timeout changes it.
+UPSTREAM_TIMEOUT_S = 45
+
+# For a request that names no output limit. llama-server's own default is none,
</file context>

Comment thread web/interview.js
Comment on lines +2410 to +2414
const DEFAULT_REPORT_ESCAPE_WAIT_MS = 135000;
const REPORT_ESCAPE_WAIT_MS = Math.max(
DEFAULT_REPORT_ESCAPE_WAIT_MS,
Number(globalThis.CODETRIAL_REPORT_ESCAPE_WAIT_MS) || 0,
);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: If /runtime-config.js is unavailable, this falls back to 135 seconds even though local report generation can run for 250 seconds, exposing the leave/offline-report options before it can finish. Use a local-safe fallback when the server value is absent.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/interview.js, line 2410:

<comment>If `/runtime-config.js` is unavailable, this falls back to 135 seconds even though local report generation can run for 250 seconds, exposing the leave/offline-report options before it can finish. Use a local-safe fallback when the server value is absent.</comment>

<file context>
@@ -2401,12 +2401,17 @@ function flushPendingLanguagePublish() {
+/// is the hosted deadline, held against the agent's by
+/// the_browser_escape_hatch_outlasts_the_report_deadline, which reads it here,
+/// and it is a floor: a config that is missing or says less never shortens it.
+const DEFAULT_REPORT_ESCAPE_WAIT_MS = 135000;
+const REPORT_ESCAPE_WAIT_MS = Math.max(
+  DEFAULT_REPORT_ESCAPE_WAIT_MS,
</file context>
Suggested change
const DEFAULT_REPORT_ESCAPE_WAIT_MS = 135000;
const REPORT_ESCAPE_WAIT_MS = Math.max(
DEFAULT_REPORT_ESCAPE_WAIT_MS,
Number(globalThis.CODETRIAL_REPORT_ESCAPE_WAIT_MS) || 0,
);
const DEFAULT_REPORT_ESCAPE_WAIT_MS = 135000;
const configuredReportEscapeWait = Number(
globalThis.CODETRIAL_REPORT_ESCAPE_WAIT_MS,
);
const REPORT_ESCAPE_WAIT_MS =
Number.isFinite(configuredReportEscapeWait) && configuredReportEscapeWait > 0
? Math.max(DEFAULT_REPORT_ESCAPE_WAIT_MS, configuredReportEscapeWait)
: 260000;

Comment thread web/report-recovery.js
wait = timers.setTimeout(
finish,
reportRecoveryLimits.retryWaitSeconds * 1000,
const seconds = Math.max(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: If /runtime-config.js is unavailable, the missing value leaves this at 140 seconds, so finish() finalizes the incomplete report before a local retry's 250-second deadline; a later report is ignored. Use a local-safe fallback only when the server value is absent.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/report-recovery.js, line 66:

<comment>If `/runtime-config.js` is unavailable, the missing value leaves this at 140 seconds, so `finish()` finalizes the incomplete report before a local retry's 250-second deadline; a later report is ignored. Use a local-safe fallback only when the server value is absent.</comment>

<file context>
@@ -57,14 +57,17 @@ export function createReportRecovery({
-    wait = timers.setTimeout(
-      finish,
-      reportRecoveryLimits.retryWaitSeconds * 1000,
+    const seconds = Math.max(
+      reportRecoveryLimits.retryWaitSeconds,
+      Number(globalThis.CODETRIAL_REPORT_RETRY_WAIT_SECONDS) || 0,
</file context>

Comment thread scripts/gemini-shim.py
messages.extend(to_messages(body.get("contents", [])))

config = body.get("generationConfig", {})
request = {"messages": messages, "stream": False}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: The report and interim requests set a seed for repeatable outputs, but this conversion drops generationConfig.seed. Forward it to the local chat request so repeated prompts retain the configured determinism.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At scripts/gemini-shim.py, line 196:

<comment>The report and interim requests set a seed for repeatable outputs, but this conversion drops `generationConfig.seed`. Forward it to the local chat request so repeated prompts retain the configured determinism.</comment>

<file context>
@@ -0,0 +1,400 @@
+    messages.extend(to_messages(body.get("contents", [])))
+
+    config = body.get("generationConfig", {})
+    request = {"messages": messages, "stream": False}
+    tools = to_tools(body)
+    if tools:
</file context>

Comment thread README.md
The report and the quiet-pause reviews can be written by a model on your own
hardware instead. Point `CODETRIAL_GEMINI_REST_BASE` at a server that answers
Gemini's `generateContent`, such as `scripts/gemini-shim.py` in front of
llama.cpp's `llama-server`; the shim's docstring has the commands. The live

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: The referenced setup command omits --jinja, which the shim requires for function calls; the behaviour check then cannot exercise its tools. Include --jinja in the llama-server command.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At README.md, line 243:

<comment>The referenced setup command omits `--jinja`, which the shim requires for function calls; the behaviour check then cannot exercise its tools. Include `--jinja` in the `llama-server` command.</comment>

<file context>
@@ -236,6 +237,22 @@ save.
+The report and the quiet-pause reviews can be written by a model on your own
+hardware instead. Point `CODETRIAL_GEMINI_REST_BASE` at a server that answers
+Gemini's `generateContent`, such as `scripts/gemini-shim.py` in front of
+llama.cpp's `llama-server`; the shim's docstring has the commands. The live
+interviewer still talks to Gemini. Any base other than Google's gets longer
+report deadlines, 45 seconds a call and 250 in all instead of 20 and 125, since
</file context>
Suggested change
llama.cpp's `llama-server`; the shim's docstring has the commands. The live
llama.cpp's `llama-server`; the shim's docstring has the commands; add `--jinja` to the `llama-server` command for function calls. The live

Comment thread src/gemini.rs
// One entry per item, a missing weakness included, so each index is the
// item's own and the one the validator reports. Filtering those out first
// shifted every later index, and the repair named the wrong item.
let weaknesses = raw

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This comparison flags weaknesses that validate_report_candidate already accepts after snap_plan_weaknesses normalizes case, spacing, or a final period. When another field needs repair, the prompt tells the model to rewrite a valid plan item; derive guidance from the snapped values or suppress these false mismatches.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/gemini.rs, line 1185:

<comment>This comparison flags weaknesses that `validate_report_candidate` already accepts after `snap_plan_weaknesses` normalizes case, spacing, or a final period. When another field needs repair, the prompt tells the model to rewrite a valid plan item; derive guidance from the snapped values or suppress these false mismatches.</comment>

<file context>
@@ -1097,6 +1140,125 @@ pub(crate) fn report_regeneration_retry_after(
+    // One entry per item, a missing weakness included, so each index is the
+    // item's own and the one the validator reports. Filtering those out first
+    // shifted every later index, and the repair named the wrong item.
+    let weaknesses = raw
+        .get("improvementPlan")
+        .and_then(Value::as_array)
</file context>

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant