Skip to content

Let reports be written by a local model - #99

Closed
alanhc wants to merge 11 commits into
sysprog21:mainfrom
alanhc:local-llm-report
Closed

alanhc wants to merge 11 commits into
sysprog21:mainfrom
alanhc:local-llm-report

Conversation

@alanhc

@alanhc alanhc commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The report and the quiet-pause reviews can now be written by a model on the operator's own GPU instead of Gemini. The live interviewer is untouched and still talks to Gemini.

  • CODETRIAL_GEMINI_REST_BASE, read from the environment like INTERVIEW_ROOM_NAME, points the generateContent calls at another server. Unset, the URL is what it was.
  • scripts/gemini-shim.py (standard library only) answers generateContent from llama-server's OpenAI-compatible endpoint. It converts the Gemini response schema into the JSON Schema that llama.cpp compiles to a grammar, turns thinking off when the request asks for a zero budget, and passes upstream status codes through so the retry rules still see a 503 as a 503.
  • A local model is slower, so any base other than Google's gets longer report deadlines: 45 seconds per call and 240 in all, instead of 20 and 125. Gemini keeps its values. Only the server knows which deadline is in force, so the page's wait before it offers to leave now comes from /runtime-config.js. The page's own default stays the hosted 135 seconds and acts as a floor, so a missing config can never shorten it.
  • When a report names the published problem, the repair prompt now spells out the title and the scenario name to use instead. A smaller model told only names the published problem kept writing "a 'Two Sum' style problem" and lost the report. The title stays out of the failure note the candidate reads.
  • When the improvement plan and the feedback disagree, the repair now names each wrong item by index (a reworded weakness, or a repeat), each improvement with no item, and when there is one of each, the swap. Told only that the plan and the feedback disagreed, the local model sent back the same plan byte for byte on every repair.
  • README documents the variable and the deadlines.

Numbers

On an RTX 5070 Ti with the Two Sum report prompt from tests/golden/prompts.json:

  • gemma-4-12b Q4_K_M, with every change here: 10 of 10 reports passed validation, 28 to 50 seconds each. 7 of the 10 broke the plan rule on their first attempt and were repaired, 6 of them on the first repair.
  • The same model before the plan repair change: 0 of 6 in one run of six and 7 of 10 across an earlier ten, every failure the plan rule surviving both repairs. Before the title repair change, that rule alone cost 2 of 4.
  • Qwen3.5-9B Q6_K, before the repair change: 3 of 8. The failures were improvementPlan rules, not the transport.
  • Interim reviews through the shim: under a second each.

These are one prompt on one machine, so read them as "this works end to end" rather than as a quality comparison with Gemini.

Test plan

  • report_network_budget_covers_every_repair_and_retry_per_generation now checks that both the hosted and the local deadline pay for five attempts.
  • the_browser_escape_hatch_outlasts_the_report_deadline reads the page's default and checks it against the hosted deadline, then checks what the server sends against both.
  • The runtime-config test and the browser test stub carry the new line.
  • binary_web_gives_a_local_report_base_the_longer_wait in tests/cli.rs starts the binary with CODETRIAL_GEMINI_REST_BASE set and checks the page is told 250000. The base is read once per process and nothing else in the suite sets it, so this is the only test that reaches the local branch.
  • Four unit tests for the plan repair: a reworded weakness, a repeated one, several wrong at once (no positional pairing), and that the guidance stays out of other repairs and out of the candidate's failure note. Three of them fail with the guidance removed. The guidance is worked out from the response itself, and a plan that matches its feedback gets none.
  • tests/local_report.rs is an ignored check that makes one real generate_report call through CODETRIAL_GEMINI_REST_BASE. It is not part of the gate.
  • Manually: the built binary serves CODETRIAL_REPORT_ESCAPE_WAIT_MS = 135000 without the variable and 250000 with it.
  • ./scripts/test.sh passes locally, formatting included. cargo mutants --in-diff on this diff: 36 mutants, 33 caught, 3 unviable, none missed. cargo-audit and shellcheck were not installed here.

Not tested: a full interview end to end with the report coming from the local model.

cubic-dev-ai[bot]

This comment was marked as resolved.

cubic-dev-ai[bot]

This comment was marked as resolved.

The report and interim-review calls always went to Google, so trying
a model on one's own GPU meant editing the URL by hand. They now go to
CODETRIAL_GEMINI_REST_BASE when it is set, read from the process
environment like INTERVIEW_ROOM_NAME; unset, nothing changes. The live
socket is untouched.

scripts/gemini-shim.py answers generateContent from llama-server's
OpenAI-compatible endpoint: it turns the Gemini schema into the JSON
Schema llama.cpp compiles to a grammar, turns thinking off when the
request asks for a zero budget, and passes upstream status codes
through so the retry rules still see a 503 as a 503.

tests/local_report.rs is an ignored check that makes one real
generate_report call through that base. Against Qwen3.5-9B Q6_K on an
RTX 5070 Ti, 3 of 8 reports passed validation; the failures were the
improvementPlan rules, not the transport.
A 9B to 14B model on one consumer GPU takes 14 to 32 seconds per
report call, against a 20-second attempt limit sized for Gemini, so a
local model timed out on most calls. When CODETRIAL_GEMINI_REST_BASE
points anywhere but Google, the attempt limit is 45 seconds and the
report deadline 240, so five attempts still fit; Gemini keeps 20 and
125. The page's wait before offering to leave now comes from
/runtime-config.js, since only the server knows which deadline is in
force, and the page's own default stays the hosted one as a floor.
A report whose improvement plan called the exercise "a 'Two Sum'
style problem" was sent back with only "names the published problem"
and the field's path. The model could not tell which words broke the
rule, wrote the same sentence twice more, and the report was lost.

The repair prompt now names the published title and the scenario's
title to use instead. The title stays out of the error itself: that
error is the failure note the candidate reads, and the title is what
it must not show them. The original prompt already carries the title,
so the model learns nothing new from it.

Against gemma-4-12b on the Two Sum weak-candidate prompt, 2 of 4
reports passed before this and 4 of 4 after, three of them on the
first repair.
CODETRIAL_GEMINI_REST_BASE was described only in the code and in the
shim's docstring, so an operator reading the configuration section
would not learn that reports can be written locally, or that doing so
lengthens the report deadlines. The section now says both, and lists
the variable with the others that come from the environment.
When the improvement plan and the feedback disagreed, the repair was
told only that they did, and a local 12B model sent back the same plan
byte for byte on every repair: told that something in a list of four
was wrong, it could not find which. The usual cause is a weakness
reworded on its way into the plan, or a repeat standing where an
improvement was left out. The repair now names each wrong item by
index and each improvement with no item, and when there is one of
each, the swap. The note the candidate reads is unchanged.

Against gemma-4-12b on the Two Sum report prompt, 0 of 6 reports
passed before this and 10 of 10 after; 7 of the 10 broke the plan on
their first attempt and were repaired.
Two ways the plan repair disagreed with the validator it is meant to
satisfy, both reported in review.

The validator compares sets, so an improvement named under both
feedback sections wants one plan item. The guidance counted it twice,
asked for an item too many, and the model's compliance came back
rejected as a duplicate. Improvements are now taken once each.

Items without a weakness were filtered out before being numbered, so
every later index moved up by one and the repair named the item
before the one the validator had. Each item now keeps its own index,
and one without a weakness is named as such and can be the target of
the swap.
The shim waited up to 300 seconds on llama-server, while the Rust side
gives a local report attempt 45 and then retries. The shim never saw
the caller leave, so the GPU kept writing an answer nobody would read
while the retry queued behind it.

The wait now defaults to the same 45 seconds, and --upstream-timeout
changes it. Timing out closes the upstream socket, which is what stops
llama-server: with the wait set to 2 seconds and generation forced to
run on, the shim answered 503 at 2.0 seconds and no slot was busy half
a second later.
The test wrote the page's escape wait as 135000, the hosted value, so
running the suite with CODETRIAL_GEMINI_REST_BASE exported failed it,
although the server was right to say 250000 there. The expected wait
is now read from report_escape_wait(report_timeout()), made public for
it, and held to one of the two values so a third cannot slip through.
The test that pins 250000 for a local base is unchanged in tests/cli.rs.
tests/local_report.rs let REPORT_PROMPT_FILE swap in another prompt
but always validated the report as Two Sum, so a report for another
problem was checked against the wrong published title and could name
its own. REPORT_PROBLEM now names the problem, Two Sum by default, and
an id the bank does not have fails the check instead of falling back
to the default the way get_problem would.
Accepting either wait let a report_endpoint_is_local that always
answers true survive mutation testing: the only test that pinned
135000 was this one, and it no longer did. Without
CODETRIAL_GEMINI_REST_BASE the test now requires 135000 exactly, and
only a run with the variable set accepts either value.
The indent gate wants a blank line before a comment that opens a new
step, which the pre-commit hook here skipped without commentflow on
its path.
@jserv

jserv commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Merge this work into #110

@jserv jserv closed this Sep 26, 2026
@alanhc

alanhc commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator Author

Folded into #110: the work here is rebuilt there on current main as five commits, together with the behaviour-check changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants