Skip to content

Benchmark local models and GPUs for the reports and the behaviour check #257

Description

@alanhc

#110 lets the report, the pause review, the phase judge and the interviewer behaviour check run on a model you host yourself, through scripts/gemini-shim.py in front of llama.cpp's llama-server. So far it has been measured on one GPU with one model. This issue collects results from other models and other hardware, so that we can answer two questions:

  1. Which models follow the interview rules? The behaviour check fails a model that names the source problem, volunteers a limit, answers a hint request without calling log_hint, or serves a hint past its rung.
  2. Which hardware is fast enough? The local deadlines in Run reports and the behaviour check on a local model #110 are fixed: 45 s for each report call, 250 s for the whole report, and 12 s for each pause review or phase judgment. A slower GPU, or a model that writes longer answers, will hit them.

The live interviewer is not covered. In #110 it still talks to Gemini Live, so the steps below need neither a Gemini key nor LiveKit.

What you need

  • A GPU that llama.cpp supports (CUDA, ROCm, Vulkan or Metal), or a CPU if you are patient. gemma-4-12b Q4_K_M with a 32k context takes about 8.2 GB of VRAM.
  • Rust 1.98 or newer and Python 3.9 or newer. On Linux x86_64, make build downloads the Clang that WebRTC needs; on other hosts, see the README's "Dependencies by lifecycle".
  • About 20 GB of disk: the model, a llama.cpp build and CodeTrial's target/.

Steps

1. Check out #110 and build it once

git clone https://github.com/sysprog21/codetrial && cd codetrial
gh pr checkout 110        # or: git fetch origin pull/110/head:pr-110 && git checkout pr-110
make build

2. Build llama.cpp and measure the raw speed

git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON     # -DGGML_HIP=ON, -DGGML_VULKAN=ON, or nothing on macOS for Metal
cmake --build build -j --config Release
build/bin/llama-bench -m MODEL.gguf -p 512,4096 -n 256 -fa 1

Keep the llama-bench table for the results.

3. Start the model and the shim, each in its own terminal:

build/bin/llama-server -m MODEL.gguf --host 127.0.0.1 --port 8080 \
    -ngl 99 -c 32768 -fa on -ctk q8_0 -ctv q8_0 --jinja
scripts/gemini-shim.py --listen 127.0.0.1:8090 --llama http://127.0.0.1:8080 --thinking off

--jinja is required, or the tool calls do not work. Use --thinking off for models that think by default (Gemma 4, Qwen3): with thinking on, gemma-4-12b sometimes repeated itself until the output limit and took about ten times as long. If the model does not fit, lower -c or -ngl and note what you used.

4. Run the bench from the CodeTrial checkout. Save this as local-bench.sh:

#!/bin/sh
# Report x3, then the behaviour check (which includes the phase judge).
set -u
export CODETRIAL_GEMINI_REST_BASE=${SHIM:-http://127.0.0.1:8090}
export BEHAVIOR_CANDIDATE_BASE=${LLAMA:-http://127.0.0.1:8080}
export GOOGLE_API_KEY=local
export CXX=${CXX:-$PWD/target/clang/bin/clang++}
log=local-bench-$(date +%Y%m%d-%H%M).log
for i in 1 2 3; do
    echo "== report $i"
    cargo test -q --lib a_local_model_writes_a_report -- --ignored --nocapture
done >"$log" 2>&1
echo "== behaviour" >>"$log"
cargo test -q --test interview_behavior -- --ignored --nocapture >>"$log" 2>&1
grep -E '^== |^elapsed|--- FAILED|^test result' "$log"
grep -E '^[a-z0-9-]+(/[a-z-]+| \([A-Za-z]+\)): ' "$log"
grep -A1 'panicked at' "$log" | grep -vE 'panicked at|^--$|behaviour failures'
echo "full log: $log"
sh local-bench.sh
sh local-bench.sh

Please run it twice: results vary between runs even though the calls are seeded (see the reference below). On an RTX 5070 Ti one run takes about 5 minutes. The report test uses the Two Sum golden prompt by default, and REPORT_PROBLEM changes the problem. The shim's own log shows each call's tokens and time; please include the report calls (... in, ... out, STOP, ...s).

Results template

Post a comment with this filled in, and attach the full log if anything failed. These commands print most of the hardware fields:

# Linux
nvidia-smi --query-gpu=name,memory.total,driver_version --format=csv   # or rocm-smi / vulkaninfo --summary
lscpu | grep "Model name"; free -g | grep Mem; cat /etc/os-release | grep PRETTY
# macOS
system_profiler SPHardwareDataType SPDisplaysDataType | grep -E "Chip|Memory|Total Number of Cores"; sw_vers
GPU (model, VRAM):
CPU:
RAM:
OS:
GPU backend and driver:          (e.g. CUDA 12.9 / driver 580, ROCm 6.4, Vulkan, Metal)
llama.cpp commit:
Model file:                     (e.g. gemma-4-12b-it-Q4_K_M.gguf)
llama-server flags:
llama-bench pp512 / pp4096 / tg256:
VRAM in use while serving:

For each of the two runs:
  Report x3: valid? / seconds each / calls each (from the shim log)
  Behaviour check: tests passed of 5 (the phase judge is one of them),
                   and the failure lines the script prints
Anything else you saw:

Reference result

RTX 5070 Ti 16 GB, i7-14700, 64 GB RAM, Ubuntu 24.04, driver 580.173, llama.cpp bd43117, gemma-4-12b-it Q4_K_M with the flags above and --thinking off:

  • Throughput from the server log: prompt about 3,400 tokens/s, generation about 77 tokens/s.
Run 1 Run 2
Report x3 all valid, 44.5 s each, two calls each: the first answer (3,267 tokens in, 1,846 out, 24 s) was refused by validation and the repair (5,431 in, 1,530 out, 20.5 s) was accepted all valid, 15 to 16 s each, one call each
Behaviour check 2 of 5 3 of 5
Phase judge test pass pass
Hint order (live_interviewer_poses_the_variant_and_serves_hints_in_order) fail: a second hint request answered without log_hint, on 3sum and on two-sum with the whiteboard pass
Uncertain speech fail: wrong indices not challenged against the input fail, the same way
Played candidates fail, 5 lines fail, 8 lines: hint requests answered without log_hint, and limits volunteered (3sum 100,000, coin-change 10,000, two-sum 10^9)

The failures are model behaviour, not timeouts. Every report call stayed inside the 45 s limit, the longest at 24 s.

In a five-minute interview by hand on the same machine, with the interviewer also local, a report took one call and 17 s.

Worth trying

  • Models: Qwen3-14B, Qwen3.5-9B, gpt-oss-20b, Mistral Small 3.x, Llama 3.x 8B, and larger quantizations of gemma-4-12b. Earlier trials on the interviewer side here: Qwen3.5-9B named the intended data structure unasked, and Qwen3-14B followed the hint rules best but did not fit fully in 16 GB beside the speech models.
  • Hardware: 8 GB and 12 GB NVIDIA cards, AMD through ROCm or Vulkan, and Apple Silicon. On an M1 Pro, my estimate is that a report call takes 90 s or more, which would hit the 45 s limit; a measurement would settle it. If it does, the deadlines should probably be configurable for local setups, and the results here are what would size them.

Related: #105 (running CodeTrial on a local model), #110.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions