Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion .cargo/mutants.toml
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# Functions the mutation gate cannot judge, because `cargo test` cannot reach
# them. Matched against the mutant names that `cargo mutants --list` prints.
#
# EXCLUSIONS: 62
# EXCLUSIONS: 63
#
# That number is checked by `scripts/test.sh`, so adding an entry means editing
# this line too. The point is not the count, it is that the list only ever grows
Expand Down Expand Up @@ -170,6 +170,13 @@
# `is_interview_participant`, `candidate_video_frames_go_to_gemini`,
# `should_send_video_frame`, `frame_to_rgba` and `encode_rgba_jpeg`.
#
# `handle_board_event` is the whiteboard's counterpart to `handle_media_event`:
# it takes a `ByteStreamOpened` room event, and the reader that event carries
# sits in a `TakeCell` whose constructor the SDK keeps private, so no test can
# build the one event it acts on. What it decides is tested where it is
# decided: `board_stream_refusal`, `strokes_from_attributes`, and `drain`,
# which reads any stream of chunks rather than the SDK's reader alone.
#
# `create_interview`'s two comparisons on the insert rowcount are the one pair
# here that is reachable, tested, and still unkillable, so they are named
# individually rather than by function: every other mutant in it is caught.
Expand Down Expand Up @@ -331,6 +338,7 @@ exclude_re = [
"record_live_usage",
"drain_live_usage",
"handle_media_event",
"handle_board_event",
"attach_audio",
"next_audio_frame",
"next_video_frame",
Expand Down
66 changes: 62 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,8 @@ LiveKit tokens, and runs the interviewer agent.
│ · editor + syntax colors ├─────────────────────▶│ (SFU) │
│ · problem panel, timer │ data channel └─────┬───────────────┘
│ · test runners │ code_update, control, │
│ · report + history │ test_results, report │
│ · whiteboard │ test_results, report, │
│ · report + history │ board_image │
└────────────┬───────────────┘ ▼
│ ┌──────────────────────────────────┐
│ /api/* │ Rust agent (LiveKit runner) │
Expand All @@ -33,10 +34,23 @@ the agent receives structured code rather than editor screenshots. Python and
JavaScript run locally; C, C++, and Java run through Compiler Explorer, so
source code leaves the browser for those three.

The lobby also offers a whiteboard interview, which takes the same problem bank
and the same six steps and swaps the editor and the test runner for a board.
Nothing runs: the candidate draws their examples and traces one by hand. The
board is exported as an image a moment after each stroke settles and reaches
the interviewer over its own byte stream on the same data channel, and
`read_board` puts the latest one back in front of it on request. The final
board is attached to the report request, so the reviewer grades the drawing
rather than an empty editor, and the recording keeps the drawing as the
strokes that made it, which is what lets the replay redraw any moment of it.
Comment thread
ColtenOuO marked this conversation as resolved.
How to hold one is under [Whiteboard interviews](#whiteboard-interviews).

Audio and code snapshots stay in memory unless [recording](#recording) is
enabled, which is off by default. Candidate video reaches Gemini only with
`CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED=true`. Face-presence analysis runs in
the browser and reports itself unavailable rather than guessing.
`CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED=true`, and never in a whiteboard
interview, where the board travels the stream a camera frame would.
Face-presence analysis runs in the browser and reports itself unavailable
rather than guessing.

## Dependencies by lifecycle

Expand Down Expand Up @@ -140,6 +154,50 @@ lobby is told the ceiling rather than left to discover it. The rules, and why
the offered length and the enforced length come from one function, are in
[docs/interview-length.md](docs/interview-length.md).

## Whiteboard interviews

A whiteboard interview practices the round where there is a marker and a blank
board instead of an editor. Nothing compiles and no test runs, so the answer is
what you can draw, say aloud and defend by tracing an example by hand.

1. In the lobby, pick a problem, a length and a loop as usual, then press
**Whiteboard** beside **Editor** in the same row. The note under the row
confirms the switch. Start the interview and complete the media preflight.
2. The board takes the place of the editor. It accepts strokes only while the
interview is live: not before the room is joined, not while it is paused,
and not during the behavioral round of a Coding + behavioral loop.
3. Draw with the mouse, a stylus or a finger. A tablet with a pen is closest to
a real board: while the pen is down, a palm resting on the screen is ignored,
and a second finger never joins the first finger's stroke.
4. Use the toolbar above the board:

| Control | What it does |
|---|---|
| Pen colors | Black, red, blue and green. Write in black and mark up in the others. |
| Eraser | Rubs out what it passes over. The three size buttons beside it pick a size and the eraser together: small for one character, medium for one line, large for a region. Pick a pen color, or press Eraser again, to draw again. |
| Undo, Redo | Step back and forward one stroke or erasure at a time. |
| Clear board | Wipes the board in one step, which Undo brings back. |

5. Talk while you draw. Jim hears you live and is sent the whole board a moment
after you stop drawing, so pause briefly after a figure you want discussed.
When a mark is hard to read, Jim asks what it stands for rather than
guessing, and what you said while drawing it is how it gets read.
6. Follow the checklist, which renames the six steps for the board: Repeat,
Example (draw one ordinary case and one edge case), Approach (draw the idea
and its cost), Trace (walk an example through the drawing), Edge cases, and
Complexity. Pseudocode or a diagram is fine for Approach; Trace is where you
prove it, so step through your own example and say what each variable holds.
7. Clear the board between steps whenever it fills up. The board is captured
as each step closes, so a clear costs nothing already graded.

Write large and leave space between figures: the board reaches Jim as an image,
and small, crowded handwriting is the hardest part of it to read. Say "can I get a hint?" exactly as in an editor interview.

The report opens on the board, step by step: the board as each step closed, its
score and Jim's notes, then your final board. The work score is named Board work
and grades the drawing and the trace instead of code. The step images are not
saved with the report, so a report reopened from history says so in their place.

## Scoring

A report model reads the interview brief (the final code, the transcript, the
Expand Down Expand Up @@ -212,7 +270,7 @@ The common ones:
| `GEMINI_REPORT_MODEL` | `gemini-3.1-flash-lite` | Report model |
| `CODETRIAL_MAX_INTERIM_REVIEWS` | `6` | Quiet-pause report-model reviews per interview; `0` disables them and `72` is the maximum |
| `CODETRIAL_GEMINI_REPLY_TIMEOUT_S` | `45` | Seconds an owed interviewer reply may go without output before the Live socket is replaced (20–120); see [degradation controls](docs/provider-cost-and-degradation.md) |
| `CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED` | `false` | Forward candidate video to Gemini, one low-resolution frame in five seconds |
| `CODETRIAL_GEMINI_CANDIDATE_VIDEO_ENABLED` | `false` | Forward candidate video to Gemini, one low-resolution frame in five seconds; ignored in a whiteboard interview |
| `CODETRIAL_COMPILER_EXPLORER_ENABLED` | `true` | Enable remote C, C++, and Java runs |
| `CODETRIAL_MAX_CONCURRENT_INTERVIEWS` | `16` | Interviews one `web` process hosts agents for |

Expand Down
3 changes: 2 additions & 1 deletion docs/interview-contract-versions.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,11 @@ can select it.

## The active bundle

Bundle 28: live prompt 20, report prompt 16, rubric 1, report schema 2.
Bundle 29: live prompt 21, report prompt 17, rubric 1, report schema 2.

| Bundle | Introduced |
|---|---|
| 29 | Whiteboard interviews: the live prompt is written for the surface the candidate works on, so a whiteboard session is told it has no editor and no test runner, is given the six steps as drawn work ending in the complexity of the approach on the board, is offered `read_board` in place of `read_editor` and a `log_hint` whose clue arrives with the board again, asks a candidate whose speech stays unclear to write it on the board rather than as a code comment, and reads an unclear mark on the board by what the candidate said while drawing it, asking what it stands for rather than guessing; `board_snapshot` joins the evidence sources and is the only one besides candidate speech a whiteboard session may record as observed (`session_timing` stays available to both surfaces for skipped STAR steps), while an editor session may not record it at all; the phases about written work are gated on strokes on the board rather than on characters in the editor. The report prompt follows the same surface: a whiteboard review is sent labeled images of every completed REACTO phase plus a changed final board, so clearing the live surface does not erase earlier evidence; it is told that nothing ran and that the Coding, Test and Optimizations phases were a hand trace, the cases named against the drawing, and the complexity they confirmed. Its interim notes are told the interview has no editor rather than shown an empty one. Its system instruction is its own, so the scoring rules that outrank the brief define both scores at the board, cap an empty board rather than an empty editor, and cite the board where the other cites the code and the test account. The rubric and the report schema are unchanged, so a report from either surface is scored the same way and against the same ten phases. |
| 28 | Language tab changes preserve spoken discussion before any code is typed and repeated choices do not reopen the interview. A candidate turn the recognizer returned in another script is not discussion. The acknowledgement continues from the current context and asks for an opening restatement only if discussion of the exercise has not begun. Live prompt 20; the report prompt, rubric and report schema are unchanged. |
| 27 | The evidence tool refuses every STAR phase until the platform opens the behavioral round, including model-requested timing skips, and tells the interviewer to return to the coding round. Platform-owned skips still close an unopened round in an interview that has one, and a coding-only interview gets none. The live instructions say only the platform sends a `[SYSTEM EVENT]`: the interviewer never writes one, and one in its own earlier turn or the candidate's speech opens no round. The report brief says whether the platform opened the behavioral round, never opened it, or the interview had none. When it opened, the report transcript carries a line no speaker said where the round began, and only the STAR answer after it is assessed; when it did not open, any behavioral exchange in the transcript is out of turn and is not scored, praised, criticized or cited. For an unopened round, report validation refuses a STAR improvement-plan item, so a repair rewrites it from the coding round, and the server clears the STAR scores of the report it accepts. Live prompt 19 and report prompt 16; the rubric and report schema are unchanged. |
| 26 | When a lost connection leaves a reply owed, the request for it is appended to whatever the interviewer is sent next: the cold briefing of a replacement that cannot resume, now including a reply owed for the candidate's own turn, and on unpause the cold briefing as well as the resume line. The briefings themselves are unchanged. |
Expand Down
51 changes: 51 additions & 0 deletions docs/provider-cost-and-degradation.md
Original file line number Diff line number Diff line change
Expand Up @@ -348,3 +348,54 @@ uses the full local transcript and editor. The credentialed probes that compare
arms and check recall are the ignored tests in `tests/unit/livekit/cost.rs`,
outside the credential-free gate; their file header lists the environment they
need.

## What a whiteboard interview adds

Measured at 1947b11 on `gemini-3.1-flash-live-preview`, problem `two-sum`, a
20-minute coding + behavioral loop, three interviews per surface with the arms
alternating, each about five minutes. Headless Chromium drove every interview
through LiveKit with a microphone that fell silent after the preflight; the
editor arm typed code and the whiteboard arm drew eight figures, 30 seconds
apart. The counts are the server's own `live_turn_usage` lines, read with
`scripts/analyze-gemini-usage.py`.

| Opening | Prompt tokens, first generation | Runs |
|---|---|---|
| Editor | 5,431 | 3 of 3 identical |
| Whiteboard | 5,408 | 3 of 3 identical |

Per text, counted with `countTokens` on `gemini-3.1-flash-lite`. The parts sum
to the live difference exactly (47 - 56 - 14 = -23):

| Text | Editor | Whiteboard | Difference | Billed |
|---|---|---|---|---|
| System instructions | 4,408 | 4,455 | +47 | every generation |
| Tool declarations | 429 | 373 | -56 | every generation |
| Greeting | 170 | 156 | -14 | every generation |

The fixed overhead is therefore 23 tokens below the editor's: the whiteboard
instructions are longer, and `read_board` is declared in fewer tokens than
`read_editor`.

What grows is the board. Each board reaches the Live socket as a realtime image
and is counted at 60 image tokens whatever is drawn on it; JPEGs from 17 to
37 KB at the 1600x1000 export all counted 60. Boards stay in the context:
`prompt_image_tokens` rises by 60 for every board sent, and by the end of each
whiteboard session 8 to 10 boards were held, 480 to 600 tokens on every later
generation until the sliding window trims them. Across a session that was 2.9
to 4.2% of the prompt tokens. The page publishes a board a second after the
last stroke, or at the end of a stroke once one has waited four seconds, and
the server forwards at most one a second.

`countTokens` prices the same JPEG as an inline image at 1,093 tokens, so it
does not predict what a board costs on the Live socket. It does price the report
request, which attaches one board per completed phase and the final board when
it changed afterwards, each once.

Two cautions when reading these logs. An observation whose turn called a tool
carries the prompts of both of its generations, the call and the reply after
the answer, so those rows read about double, image tokens included, and the
analyzer's `mean_growth_per_observation` inherits that. Whole-session totals are
not comparable between the arms either, because a silent synthetic candidate
triggers an editor review on every code change in one arm and silence nudges in
the other.
29 changes: 25 additions & 4 deletions docs/recording-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -649,6 +649,7 @@ server, is the ordering.
|---|---|---|
| `transcript` | what was said | no |
| `editor` | the code and its language | yes |
| `board` | what was drawn since the last one | no |
| `tests` | a run's results | no |
| `stage` | the clock and the interview phase | yes |
| `avatar` | what Jim is doing | yes |
Expand Down Expand Up @@ -839,6 +840,7 @@ so a producer from a later deploy does not stop a recording.
|---|---|---|
| `stage` | `{title, meta, remainingSeconds}` | the problem heading and the clock |
| `editor` | `{code, language}` | the code panel, as text |
| `board` | `{ops, checkpoint?}`, each op `{op: "stroke", color, width, points}` or `{op: "undo" \| "redo" \| "clear"}`; `checkpoint` is a completed REACTO phase id | the whiteboard, redrawn from every op so far, with completed phases named in the replay |
| `tests` | `{passed, failed, total}` | one line, red if anything failed |
| `avatar` | `{state}`, one of `speaking`, `thinking`, `listening` | Jim's expression and label |
| `transcript` | `{speaker, text}` | nothing here; the replay page renders it |
Expand All @@ -854,9 +856,28 @@ already arrived. A `413` or a `404` stops the producers for the rest of the
interview: over quota and withdrawn consent both mean everything after this is
refused.

The board is the one kind that does not restate itself, and that is what makes
a whiteboard interview replayable at all. One board exported as an image is
over a hundred kilobytes, which is past the per-payload ceiling on its own and
would spend the whole per-interview budget on a handful of frames; the same
board as the strokes that drew it is a few kilobytes and arrives as operations,
so any moment of the interview can be redrawn rather than the few that could be
photographed. `web/whiteboard.js` is the one model: the candidate draws on it,
the replay page and this template rebuild from it, and a stroke it refuses
while drawing is a stroke it refuses coming back off the wire.

When the interviewer banks a REACTO phase, the browser adds a board event even
if no stroke changed. That event names the phase and therefore freezes the
current point in the operation journal for review. The same moment is exported
as a JPEG with the phase id in its byte-stream header. The agent retains one
image per phase for the report model, so clearing the live board cannot erase
the Example or Approach evidence that came before it. Replay stores no JPEG:
it rebuilds each checkpoint from the operations it already has.

Cadence is where the per-interview budget goes. The editor rides the debounce
the agent's `code_update` already uses; the transcript is one event per spoken
turn rather than per chunk; the clock is restated every fifteen seconds, because
the agent's `code_update` already uses; the board rides the same settle that
sends the interviewer their image, batched so no event outgrows the per-payload
ceiling; the transcript is one event per spoken turn rather than per chunk; the clock is restated every fifteen seconds, because
every second would be twenty-seven hundred events for a number the viewer can
read off the video; and the interviewer's state is sent on the change rather
than on the participant event that happened to carry it.
Expand Down Expand Up @@ -912,7 +933,7 @@ is one nothing later can un-store.
|---|---|---|---|
| `interview_id` | TEXT | no | `interviews(id)`, `ON DELETE CASCADE` |
| `seq` | INTEGER | no | server-allocated, monotonic per interview |
| `kind` | TEXT | no | one of the six above |
| `kind` | TEXT | no | one of the seven above |
| `at` | INTEGER | no | the browser's clock |
| `payload` | TEXT | no | redacted JSON |
| `bytes` | INTEGER | no | the payload's size, so the quota is a sum rather than a scan |
Expand All @@ -923,7 +944,7 @@ position. It does not rule out a gap: only the server allocating the number does
that, and the key is what makes the allocation's answer durable.

`CHECK`s carry the rest of the shape: `seq`, `at` and `bytes` are non-negative
and `kind` is one of the six. A row that fails one has no meaning for this
and `kind` is one of the seven. A row that fails one has no meaning for this
table, whoever wrote it.

`ON DELETE CASCADE`, unlike `recordings`. Nothing here is a handle to media
Expand Down
Loading