Conversation
alanhc
force-pushed
the
local-llm-report
branch
from
September 24, 2026 16:41
94f15b3 to
edf42eb
Compare
The report and interim-review calls always went to Google, so trying a model on one's own GPU meant editing the URL by hand. They now go to CODETRIAL_GEMINI_REST_BASE when it is set, read from the process environment like INTERVIEW_ROOM_NAME; unset, nothing changes. The live socket is untouched. scripts/gemini-shim.py answers generateContent from llama-server's OpenAI-compatible endpoint: it turns the Gemini schema into the JSON Schema llama.cpp compiles to a grammar, turns thinking off when the request asks for a zero budget, and passes upstream status codes through so the retry rules still see a 503 as a 503. tests/local_report.rs is an ignored check that makes one real generate_report call through that base. Against Qwen3.5-9B Q6_K on an RTX 5070 Ti, 3 of 8 reports passed validation; the failures were the improvementPlan rules, not the transport.
A 9B to 14B model on one consumer GPU takes 14 to 32 seconds per report call, against a 20-second attempt limit sized for Gemini, so a local model timed out on most calls. When CODETRIAL_GEMINI_REST_BASE points anywhere but Google, the attempt limit is 45 seconds and the report deadline 240, so five attempts still fit; Gemini keeps 20 and 125. The page's wait before offering to leave now comes from /runtime-config.js, since only the server knows which deadline is in force, and the page's own default stays the hosted one as a floor.
A report whose improvement plan called the exercise "a 'Two Sum' style problem" was sent back with only "names the published problem" and the field's path. The model could not tell which words broke the rule, wrote the same sentence twice more, and the report was lost. The repair prompt now names the published title and the scenario's title to use instead. The title stays out of the error itself: that error is the failure note the candidate reads, and the title is what it must not show them. The original prompt already carries the title, so the model learns nothing new from it. Against gemma-4-12b on the Two Sum weak-candidate prompt, 2 of 4 reports passed before this and 4 of 4 after, three of them on the first repair.
CODETRIAL_GEMINI_REST_BASE was described only in the code and in the shim's docstring, so an operator reading the configuration section would not learn that reports can be written locally, or that doing so lengthens the report deadlines. The section now says both, and lists the variable with the others that come from the environment.
When the improvement plan and the feedback disagreed, the repair was told only that they did, and a local 12B model sent back the same plan byte for byte on every repair: told that something in a list of four was wrong, it could not find which. The usual cause is a weakness reworded on its way into the plan, or a repeat standing where an improvement was left out. The repair now names each wrong item by index and each improvement with no item, and when there is one of each, the swap. The note the candidate reads is unchanged. Against gemma-4-12b on the Two Sum report prompt, 0 of 6 reports passed before this and 10 of 10 after; 7 of the 10 broke the plan on their first attempt and were repaired.
alanhc
force-pushed
the
local-llm-report
branch
from
September 25, 2026 06:25
c846782 to
4bb9a20
Compare
Two ways the plan repair disagreed with the validator it is meant to satisfy, both reported in review. The validator compares sets, so an improvement named under both feedback sections wants one plan item. The guidance counted it twice, asked for an item too many, and the model's compliance came back rejected as a duplicate. Improvements are now taken once each. Items without a weakness were filtered out before being numbered, so every later index moved up by one and the repair named the item before the one the validator had. Each item now keeps its own index, and one without a weakness is named as such and can be the target of the swap.
The shim waited up to 300 seconds on llama-server, while the Rust side gives a local report attempt 45 and then retries. The shim never saw the caller leave, so the GPU kept writing an answer nobody would read while the retry queued behind it. The wait now defaults to the same 45 seconds, and --upstream-timeout changes it. Timing out closes the upstream socket, which is what stops llama-server: with the wait set to 2 seconds and generation forced to run on, the shim answered 503 at 2.0 seconds and no slot was busy half a second later.
The test wrote the page's escape wait as 135000, the hosted value, so running the suite with CODETRIAL_GEMINI_REST_BASE exported failed it, although the server was right to say 250000 there. The expected wait is now read from report_escape_wait(report_timeout()), made public for it, and held to one of the two values so a third cannot slip through. The test that pins 250000 for a local base is unchanged in tests/cli.rs.
tests/local_report.rs let REPORT_PROMPT_FILE swap in another prompt but always validated the report as Two Sum, so a report for another problem was checked against the wrong published title and could name its own. REPORT_PROBLEM now names the problem, Two Sum by default, and an id the bank does not have fails the check instead of falling back to the default the way get_problem would.
Accepting either wait let a report_endpoint_is_local that always answers true survive mutation testing: the only test that pinned 135000 was this one, and it no longer did. Without CODETRIAL_GEMINI_REST_BASE the test now requires 135000 exactly, and only a run with the variable set accepts either value.
The indent gate wants a blank line before a comment that opens a new step, which the pre-commit hook here skipped without commentflow on its path.
Contributor
|
Merge this work into #110 |
Collaborator
Author
|
Folded into #110: the work here is rebuilt there on current |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The report and the quiet-pause reviews can now be written by a model on the operator's own GPU instead of Gemini. The live interviewer is untouched and still talks to Gemini.
CODETRIAL_GEMINI_REST_BASE, read from the environment likeINTERVIEW_ROOM_NAME, points thegenerateContentcalls at another server. Unset, the URL is what it was.scripts/gemini-shim.py(standard library only) answersgenerateContentfrom llama-server's OpenAI-compatible endpoint. It converts the Gemini response schema into the JSON Schema that llama.cpp compiles to a grammar, turns thinking off when the request asks for a zero budget, and passes upstream status codes through so the retry rules still see a 503 as a 503./runtime-config.js. The page's own default stays the hosted 135 seconds and acts as a floor, so a missing config can never shorten it.names the published problemkept writing "a 'Two Sum' style problem" and lost the report. The title stays out of the failure note the candidate reads.Numbers
On an RTX 5070 Ti with the Two Sum report prompt from
tests/golden/prompts.json:improvementPlanrules, not the transport.These are one prompt on one machine, so read them as "this works end to end" rather than as a quality comparison with Gemini.
Test plan
report_network_budget_covers_every_repair_and_retry_per_generationnow checks that both the hosted and the local deadline pay for five attempts.the_browser_escape_hatch_outlasts_the_report_deadlinereads the page's default and checks it against the hosted deadline, then checks what the server sends against both.binary_web_gives_a_local_report_base_the_longer_waitintests/cli.rsstarts the binary withCODETRIAL_GEMINI_REST_BASEset and checks the page is told 250000. The base is read once per process and nothing else in the suite sets it, so this is the only test that reaches the local branch.tests/local_report.rsis an ignored check that makes one realgenerate_reportcall throughCODETRIAL_GEMINI_REST_BASE. It is not part of the gate.CODETRIAL_REPORT_ESCAPE_WAIT_MS = 135000without the variable and250000with it../scripts/test.shpasses locally, formatting included.cargo mutants --in-diffon this diff: 36 mutants, 33 caught, 3 unviable, none missed. cargo-audit and shellcheck were not installed here.Not tested: a full interview end to end with the report coming from the local model.