You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Benchmark local models and GPUs for the reports and the behaviour check #257
#110 lets the report, the pause review, the phase judge and the interviewer behaviour check run on a model you host yourself, through scripts/gemini-shim.py in front of llama.cpp's llama-server. So far it has been measured on one GPU with one model. This issue collects results from other models and other hardware, so that we can answer two questions:
Which models follow the interview rules? The behaviour check fails a model that names the source problem, volunteers a limit, answers a hint request without calling log_hint, or serves a hint past its rung.
Which hardware is fast enough? The local deadlines in Run reports and the behaviour check on a local model #110 are fixed: 45 s for each report call, 250 s for the whole report, and 12 s for each pause review or phase judgment. A slower GPU, or a model that writes longer answers, will hit them.
The live interviewer is not covered. In #110 it still talks to Gemini Live, so the steps below need neither a Gemini key nor LiveKit.
What you need
A GPU that llama.cpp supports (CUDA, ROCm, Vulkan or Metal), or a CPU if you are patient. gemma-4-12b Q4_K_M with a 32k context takes about 8.2 GB of VRAM.
Rust 1.98 or newer and Python 3.9 or newer. On Linux x86_64, make build downloads the Clang that WebRTC needs; on other hosts, see the README's "Dependencies by lifecycle".
About 20 GB of disk: the model, a llama.cpp build and CodeTrial's target/.
scripts/gemini-shim.py --listen 127.0.0.1:8090 --llama http://127.0.0.1:8080 --thinking off
--jinja is required, or the tool calls do not work. Use --thinking off for models that think by default (Gemma 4, Qwen3): with thinking on, gemma-4-12b sometimes repeated itself until the output limit and took about ten times as long. If the model does not fit, lower -c or -ngl and note what you used.
4. Run the bench from the CodeTrial checkout. Save this as local-bench.sh:
Please run it twice: results vary between runs even though the calls are seeded (see the reference below). On an RTX 5070 Ti one run takes about 5 minutes. The report test uses the Two Sum golden prompt by default, and REPORT_PROBLEM changes the problem. The shim's own log shows each call's tokens and time; please include the report calls (... in, ... out, STOP, ...s).
Results template
Post a comment with this filled in, and attach the full log if anything failed. These commands print most of the hardware fields:
# Linux
nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv # or rocm-smi / vulkaninfo --summary
lscpu | grep "Model name"; free -g | grep Mem; cat /etc/os-release | grep PRETTY
# macOS
system_profiler SPHardwareDataType SPDisplaysDataType | grep -E "Chip|Memory|Total Number of Cores"; sw_vers
GPU (model, VRAM):
CPU:
RAM:
OS:
GPU backend and driver: (e.g. CUDA 12.9 / driver 580, ROCm 6.4, Vulkan, Metal)
llama.cpp commit:
Model file: (e.g. gemma-4-12b-it-Q4_K_M.gguf)
llama-server flags:
llama-bench pp512 / pp4096 / tg256:
VRAM in use while serving:
For each of the two runs:
Report x3: valid? / seconds each / calls each (from the shim log)
Behaviour check: tests passed of 5 (the phase judge is one of them),
and the failure lines the script prints
Anything else you saw:
Reference result
RTX 5070 Ti 16 GB, i7-14700, 64 GB RAM, Ubuntu 24.04, driver 580.173, llama.cpp bd43117, gemma-4-12b-it Q4_K_M with the flags above and --thinking off:
Throughput from the server log: prompt about 3,400 tokens/s, generation about 77 tokens/s.
Run 1
Run 2
Report x3
all valid, 44.5 s each, two calls each: the first answer (3,267 tokens in, 1,846 out, 24 s) was refused by validation and the repair (5,431 in, 1,530 out, 20.5 s) was accepted
all valid, 15 to 16 s each, one call each
Behaviour check
2 of 5
3 of 5
Phase judge test
pass
pass
Hint order (live_interviewer_poses_the_variant_and_serves_hints_in_order)
fail: a second hint request answered without log_hint, on 3sum and on two-sum with the whiteboard
pass
Uncertain speech
fail: wrong indices not challenged against the input
fail, the same way
Played candidates
fail, 5 lines
fail, 8 lines: hint requests answered without log_hint, and limits volunteered (3sum 100,000, coin-change 10,000, two-sum 10^9)
The failures are model behaviour, not timeouts. Every report call stayed inside the 45 s limit, the longest at 24 s.
In a five-minute interview by hand on the same machine, with the interviewer also local, a report took one call and 17 s.
Worth trying
Models: Qwen3-14B, Qwen3.5-9B, gpt-oss-20b, Mistral Small 3.x, Llama 3.x 8B, and larger quantizations of gemma-4-12b. Earlier trials on the interviewer side here: Qwen3.5-9B named the intended data structure unasked, and Qwen3-14B followed the hint rules best but did not fit fully in 16 GB beside the speech models.
Hardware: 8 GB and 12 GB NVIDIA cards, AMD through ROCm or Vulkan, and Apple Silicon. On an M1 Pro, my estimate is that a report call takes 90 s or more, which would hit the 45 s limit; a measurement would settle it. If it does, the deadlines should probably be configurable for local setups, and the results here are what would size them.
Related: #105 (running CodeTrial on a local model), #110.
#110 lets the report, the pause review, the phase judge and the interviewer behaviour check run on a model you host yourself, through
scripts/gemini-shim.pyin front of llama.cpp'sllama-server. So far it has been measured on one GPU with one model. This issue collects results from other models and other hardware, so that we can answer two questions:log_hint, or serves a hint past its rung.The live interviewer is not covered. In #110 it still talks to Gemini Live, so the steps below need neither a Gemini key nor LiveKit.
What you need
make builddownloads the Clang that WebRTC needs; on other hosts, see the README's "Dependencies by lifecycle".target/.Steps
1. Check out #110 and build it once
2. Build llama.cpp and measure the raw speed
Keep the
llama-benchtable for the results.3. Start the model and the shim, each in its own terminal:
build/bin/llama-server -m MODEL.gguf --host 127.0.0.1 --port 8080 \ -ngl 99 -c 32768 -fa on -ctk q8_0 -ctv q8_0 --jinja--jinjais required, or the tool calls do not work. Use--thinking offfor models that think by default (Gemma 4, Qwen3): with thinking on, gemma-4-12b sometimes repeated itself until the output limit and took about ten times as long. If the model does not fit, lower-cor-ngland note what you used.4. Run the bench from the CodeTrial checkout. Save this as
local-bench.sh:Please run it twice: results vary between runs even though the calls are seeded (see the reference below). On an RTX 5070 Ti one run takes about 5 minutes. The report test uses the Two Sum golden prompt by default, and
REPORT_PROBLEMchanges the problem. The shim's own log shows each call's tokens and time; please include the report calls (... in, ... out, STOP, ...s).Results template
Post a comment with this filled in, and attach the full log if anything failed. These commands print most of the hardware fields:
Reference result
RTX 5070 Ti 16 GB, i7-14700, 64 GB RAM, Ubuntu 24.04, driver 580.173, llama.cpp bd43117, gemma-4-12b-it Q4_K_M with the flags above and
--thinking off:live_interviewer_poses_the_variant_and_serves_hints_in_order)log_hint, on 3sum and on two-sum with the whiteboardlog_hint, and limits volunteered (3sum 100,000, coin-change 10,000, two-sum 10^9)The failures are model behaviour, not timeouts. Every report call stayed inside the 45 s limit, the longest at 24 s.
In a five-minute interview by hand on the same machine, with the interviewer also local, a report took one call and 17 s.
Worth trying
Related: #105 (running CodeTrial on a local model), #110.