Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 19 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -259,7 +259,8 @@ actually goes are in [docs/development.md](docs/development.md).
## Configuration

`config/codetrial.env.example` documents the variables that belong in a config
file; `NODE_ENV` and `INTERVIEW_ROOM_NAME` are set in the environment instead.
file; `NODE_ENV`, `INTERVIEW_ROOM_NAME` and `CODETRIAL_GEMINI_REST_BASE` are set
in the environment instead.
The common ones:

| Variable | Default | Purpose |
Expand Down Expand Up @@ -294,6 +295,23 @@ save.
Serving more than one LiveKit project from one deployment is in
[docs/providers.md](docs/providers.md).

The report and the quiet-pause reviews can be written by a model on your own
hardware instead. Point `CODETRIAL_GEMINI_REST_BASE` at a server that answers
Gemini's `generateContent`, such as `scripts/gemini-shim.py` in front of
llama.cpp's `llama-server`, which needs `--jinja` for the tool calls; the
shim's docstring has the commands. The live interviewer still talks to
Gemini. Any base other than Google's gets longer report deadlines, 45 seconds
a call and 250 in all instead of 20 and 125, since a 12B model on one 16 GB GPU
takes 14 to 32 seconds per report.

The shim carries function calls too, so the interviewer behaviour check in
[docs/development.md](docs/development.md#checks-outside-the-gate) runs against
the same base. With Gemma 4, start it with `--thinking off`. That check names no
thinking budget, and with thinking on gemma-4-12b sometimes repeated itself to
the output limit or put its reasoning in the reply. With it off the check
passed 18 of 21 problem runs across seven runs, about 24 seconds a run against
240; each miss was a second hint request answered without calling `log_hint`.

## Recording

Recording is off by default. Enabling it requires a separate LiveKit project,
Expand Down
23 changes: 18 additions & 5 deletions docs/development.md
Original file line number Diff line number Diff line change
Expand Up @@ -238,11 +238,11 @@ reads them.
## Checks outside the gate

Whether the interviewer actually follows the live prompt, rather than whether
the prompt says the right things, needs a Gemini key. The check scripts a
candidate through three problems against a text model given the same
instructions, greeting and tools, and fails on a named source, a volunteered
limit, an unanswered size question, or a hint that goes past the rung it was
served:
the prompt says the right things, needs a Gemini key or a local model. The
check scripts a candidate through three problems against a text model given
the same instructions, greeting and tools, and fails on a named source, a
volunteered limit, an unanswered size question, or a hint that goes past the
rung it was served:

```bash
scripts/interview-behavior-check.sh
Expand All @@ -252,6 +252,19 @@ BEHAVIOR_PROBLEMS=3sum,lru-cache scripts/interview-behavior-check.sh
A free key allows fifteen requests a minute, so the check waits out rate
limits; three problems take about two minutes.

Against a local model, point it at `scripts/gemini-shim.py` the way the report
is pointed, so the check sends what production sends through the shim the
report uses. The key is still read but goes no further than the shim, which
ignores it, so any non-empty value does when the config file has none:

```bash
CODETRIAL_GEMINI_REST_BASE=http://127.0.0.1:8090 GOOGLE_API_KEY=local \
scripts/interview-behavior-check.sh
```

`BEHAVIOR_LOCAL_BASE` instead talks to an OpenAI-compatible server directly,
without the shim; it is what the played candidates use for their own side.

The end-to-end browser check additionally needs Playwright and Chromium:

```bash
Expand Down
Loading