From 48866a3e823e8782f0095d2bf852959e47853193 Mon Sep 17 00:00:00 2001 From: shekharprateek Date: Mon, 14 Sep 2026 23:08:54 -0600 Subject: [PATCH] Add docs/: setup, concepts, and methodology pages Every link in the README's documentation map was a 404 -- docs/ did not exist. This adds the 12 pages that are setup and explanation rather than published results, so a new reader can get from the landing page to a first run. Imported: getting-started, hosting-paths, how-a-run-works, benchmark-your-own-repo, repository-structure, why-this-exists, vision, diagram-ascii, cost-per-task-methodology, serving-optimization-notes, omp-setup, kiro-cli-setup. Not imported, and their README rows removed: - gpu-selection-h200-vs-l40s.md -- 10 of its links point into self-hosted/vllm/benchmark-output/, the committed throughput dashboards this repo does not ship. The doc is a report on that data, so it cannot stand without it. - release-notes/ -- 0.1.0.md states that every score in it comes from docs/metrics/pareto-frontier-omp-swe3.json. It is a results release note. Edits to the imported pages, all to remove references to content this repo does not carry: - benchmark-your-own-repo.md: the "use the frontier we already published" path pointed at agentic-coding-swe-comparison-swe3.md. Repointed at the reference frontier that actually ships, vend/swe-router/models.json, with its own single-repository caveat noted. - why-this-exists.md, cost-per-task-methodology.md, omp-setup.md, diagram-ascii.md: dropped links to excluded comparison/results docs, the slides directory, and two raw run-data files. - cost-per-task-methodology.md: two sentences referred to a README leaderboard that no longer exists. - vision.md, kiro-cli-setup.md: dropped three issue links into the aarora79 personal repo; those issue numbers do not exist here. - repository-structure.md: the tree still said claude-code-multi-model/ and omitted docs/ and vend/. It also claimed a committed Hello-World example under swe-benchmark-data/, which is not present -- that path is gitignored run output. Verified: every relative link in README.md resolves, and every internal link across docs/ resolves. --- README.md | 2 - docs/benchmark-your-own-repo.md | 65 +++++++ docs/cost-per-task-methodology.md | 262 +++++++++++++++++++++++++++++ docs/diagram-ascii.md | 54 ++++++ docs/getting-started.md | 62 +++++++ docs/hosting-paths.md | 50 ++++++ docs/how-a-run-works.md | 38 +++++ docs/kiro-cli-setup.md | 177 +++++++++++++++++++ docs/omp-setup.md | 87 ++++++++++ docs/repository-structure.md | 50 ++++++ docs/serving-optimization-notes.md | 78 +++++++++ docs/vision.md | 82 +++++++++ docs/why-this-exists.md | 27 +++ 13 files changed, 1032 insertions(+), 2 deletions(-) create mode 100644 docs/benchmark-your-own-repo.md create mode 100644 docs/cost-per-task-methodology.md create mode 100644 docs/diagram-ascii.md create mode 100644 docs/getting-started.md create mode 100644 docs/hosting-paths.md create mode 100644 docs/how-a-run-works.md create mode 100644 docs/kiro-cli-setup.md create mode 100644 docs/omp-setup.md create mode 100644 docs/repository-structure.md create mode 100644 docs/serving-optimization-notes.md create mode 100644 docs/vision.md create mode 100644 docs/why-this-exists.md diff --git a/README.md b/README.md index 7de9f17f..891d8d5f 100644 --- a/README.md +++ b/README.md @@ -124,10 +124,8 @@ Where to read more, by topic: | [benchmarks/docs/end-to-end-self-hosted-run.md](benchmarks/docs/end-to-end-self-hosted-run.md) | The full manual run-book for an end-to-end self-hosted benchmark. | | [self-hosted/vllm/README.md](self-hosted/vllm/README.md) | Standing up a vLLM server: install, tensor parallelism, tool-call parsers, and the serving-config reference. | | [self-hosted/vllm/models/](self-hosted/vllm/models/) | Per-model serving guides (HF repo, context window, TP size, tool parser, hardware fit) for every benchmarked model. | -| [docs/gpu-selection-h200-vs-l40s.md](docs/gpu-selection-h200-vs-l40s.md) | Which GPU to serve on: one H200 slice of a p5en vs a whole g6e.4xlarge (1x L40S), same model and config. Why the H200 slice is 41% cheaper per unit of work despite costing 2.6x per hour, and what it costs to serve N developers. Public on-demand prices, no discounts. | | [docs/cost-per-task-methodology.md](docs/cost-per-task-methodology.md) | How the cost numbers are derived: the two cost lenses, prompt-caching accounting (API vs self-hosted), and why agentic coding is prefill-bound. | | [docs/serving-optimization-notes.md](docs/serving-optimization-notes.md) | Portable vLLM serving defaults and why we do not tune the prefill knobs per model. | -| [docs/release-notes/](docs/release-notes/) | Release notes per version, newest first, and the versioning scheme: new dataset or model is a minor, a methodology change or new functionality is a major. | | [CONTRIBUTING.md](CONTRIBUTING.md) / [SECURITY.md](SECURITY.md) / [SUPPORT.md](SUPPORT.md) | How to contribute, report a vulnerability, and get help. | ## See also diff --git a/docs/benchmark-your-own-repo.md b/docs/benchmark-your-own-repo.md new file mode 100644 index 00000000..18708fb8 --- /dev/null +++ b/docs/benchmark-your-own-repo.md @@ -0,0 +1,65 @@ +# Benchmark your own repositories + +The harness works against any GitHub repository. Name your own repos in a dataset YAML and the models, judge and cost math are the ones behind every published result here. + +## Datasets + +A dataset is a single YAML file: a metadata header plus a list of tasks, each pointing at a GitHub repo and a problem. Two datasets ship in [benchmarks/dataset/](../benchmarks/dataset/): + +- [hello-world.yaml](../benchmarks/dataset/hello-world.yaml) -- a trivial sanity dataset (the [octocat/Hello-World](https://github.com/octocat/Hello-World) repo) for kicking the tires of a new model or endpoint. +- [mcp-gateway-registry.yaml](../benchmarks/dataset/mcp-gateway-registry.yaml) -- the reference dataset, whose tasks are drawn from real upstream issues in [agentic-community/mcp-gateway-registry](https://github.com/agentic-community/mcp-gateway-registry). + +**Nothing in the harness is specific to a particular repository.** Adding your own benchmark dataset is just writing another YAML file in the same format -- point tasks at any public repo and pinned ref. The dataset format is documented in the [harness reference](../benchmarks/docs/harness-reference.md#the-dataset). + +### What do I do with this? + +The point of this repo is to help you **pick the right coding agent and model for your tasks** -- the pairing that lands the quality you need at the cost and latency you can live with, instead of defaulting to the most expensive option. There are two ways to get there: + +1. **Start from the reference frontier that ships with `swe-router`.** [`vend/swe-router/models.json`](../vend/swe-router/models.json) carries a measured per-tier score and cost per task for each model, so the skill can make a recommendation out of the box. Treat it as a starting point: it was measured on one Python/FastAPI service, and its own `single_repository_warning` says applying it to a very different codebase is extrapolation. +2. **Build your own frontier on your own code.** When you want numbers on **work that looks like yours** rather than a reference repo, use the benchmarking harness here: write a dataset YAML pointing at your repositories and run it. The models, harnesses, judge, and cost math are the same ones that produced the reference numbers. This is the rest of this section. + +**Then put it in front of developers.** [`swe-router`](../vend/swe-router/) reads whichever frontier you point it at -- ours or the one you just built -- and names the cheapest model clearing the bar for each task. Five files, no dependencies, works in any assistant that reads a skill. Point it at your own `models.json` and the recommendations are grounded in your code rather than our example repo. + +### Benchmark your own code repositories + +This is option 2 above -- building your own frontier on your own code. It is a few steps: + +1. **Create a dataset file** under [benchmarks/dataset/](../benchmarks/dataset/), for example `my-team.yaml`. Copy [mcp-gateway-registry.yaml](../benchmarks/dataset/mcp-gateway-registry.yaml) as a template. Minimal shape: + + ```yaml + schema_version: "1.0" + name: my-team + title: My team's benchmark + description: Real tasks from our own repositories. + default_ref: main # pin a tag/commit per task for reproducibility + metrics: [input_tokens, output_tokens, num_turns] + complexity_levels: [low, medium, high] + tasks: + - id: add-rate-limiting-to-gateway + repo: https://github.com/your-org/your-repo + ref: v2.3.0 # pin so every run clones the same code + complexity: medium + tags: [python, api, feature] + problem_statement: | + Describe the task in enough detail for an agent to act on it without + you present -- what to change, constraints, and what "done" means. + ``` + +Each task points at a repo + pinned ref + a problem statement (from a real ticket or issue). Full field reference: [harness reference -> The dataset](../benchmarks/docs/harness-reference.md#the-dataset). Any repo the runner can `git clone` works (public, or private with credentials available to your shell). + +2. **Run it** against whichever model/harness/path you want -- same commands as the example, just swap the dataset: + + ``` + /benchmark provider=bedrock model=claude-opus-5 dataset=dataset/my-team.yaml + ``` + +or headless: `benchmarks/scripts/run-e2e-benchmark.sh --provider bedrock --model ... --dataset dataset/my-team.yaml`. Pick the harness with `--agent claude|pi|omp|kiro` and the skill with `--skill swe2|swe3` (`--agent kiro` drives Kiro's managed models and sets `--provider kiro` for you). + +3. **Read your results.** Artifacts and scores land under `benchmarks/swe-benchmark-data//////`, and the same generators build your own cost/quality frontier (`gen_swe_comparison.py`, `plot_cost_quality.py`). Your runs are gitignored, so a customer's private code never lands in version control. + +> **Tips for good tasks:** pin a `ref` so reruns are comparable; write the `problem_statement` like a well-scoped ticket; use `tags` to slice results by language/domain/change-type; and add optional `ground_truth` (reviewer-only, never shown to the agent) if you want the judge to check against a known-good approach. + + +--- + +[< Back to the README](../README.md) diff --git a/docs/cost-per-task-methodology.md b/docs/cost-per-task-methodology.md new file mode 100644 index 00000000..5a53379b --- /dev/null +++ b/docs/cost-per-task-methodology.md @@ -0,0 +1,262 @@ +# Cost per task: methodology, the two lenses, and what agentic coding does to it + +How the throughput skill turns a fixed instance price into a **cost per token** and **cost per task** for a self-hosted (vLLM) model, why there are two ways to express it, and what the agentic-coding workload shape means for user experience and for scaling. The core model is **`run_cost = GPU-seconds x $/second`** — price tokens by dividing them by measured throughput at a stated concurrency, never by wall-clock; and the only lever that lowers that cost on a KV-bound model is KV-cache headroom, which trades against context window. A worked GLM-5.2 example runs through both. Companion to [serving-optimization-notes.md](serving-optimization-notes.md); produced by [clients/build_performance_summary.py](../self-hosted/vllm/clients/build_performance_summary.py) and surfaced by [clients/build_performance_dashboard.py](../self-hosted/vllm/clients/build_performance_dashboard.py) and [clients/cost_for_task.py](../self-hosted/vllm/clients/cost_for_task.py). + +> [!IMPORTANT] +> **Self-hosted GPU pricing basis (read before quoting any self-hosted dollar figure).** Every self-hosted cost on this page and in the charts is derived from an hourly rate in [`self-hosted/vllm/pricing.json`](../self-hosted/vllm/pricing.json). Both instance families are based at their **3-year commitment rate**, so the whole fleet is one commitment term and the cost columns are comparable across it: +> - **g6e.12xlarge**: **$4.533/hr** (on-demand $10.493, 1-year $6.61). +> - **g6e.4xlarge**: **$1.298/hr** (on-demand $3.004, 1-year $1.893). +> - **p5en.48xlarge**: **$27.72/hr** (on-demand $63.296, 1-year $40.43). +> +> `dollars_per_hour` **is** the rate charged -- nothing is applied on top -- then prorated by `tp / gpus_per_instance` for a partial-box run, so a TP=4 model on the 8-GPU p5en is charged $13.86/hr and a TP=1 model $3.465/hr. The alternative terms sit in each entry's `rates` map for reference only. +> +> **To price at a different term**, move that value into `dollars_per_hour` and regenerate: every cost number, chart, and frontier scales linearly, so [`clients/reprice_performance_summary.py`](../self-hosted/vllm/clients/reprice_performance_summary.py) rescales a committed summary exactly, without needing the (local, gitignored) sweep DuckDB. +> +> There is deliberately **no discount multiplier**. `pricing.json` used to carry one, and p5en was based at on-demand times a `0.35` **placeholder** -- an assumption sitting in the same field, and flowing into the same charts, as measured prices. `pricing.py` now rejects a leftover `discount` key rather than honouring it, so the concept cannot creep back. + +### Which AWS discount instrument these rates name + +The g6e rates are a "**3-year commitment rate**" without naming an instrument, because the instrument does not change the number: AWS's [own comparison](https://docs.aws.amazon.com/savingsplans/latest/userguide/sp-ris.html) puts an EC2 Instance Savings Plan and a Standard RI in the same **up to 72% off** tier, and a Compute Savings Plan and a Convertible RI in the same **up to 66%** tier. Pick either; the rate lands in the same place. The p5en rates are quoted as **EC2 Instance Savings Plan** rates because that is where they were read from, and they sit at the same ratios to on-demand as g6e (1-year 63.9%, 3-year 43.8%) -- which is the check that the two families are on the same basis. + +Reserved Instances are not retired. AWS [recommends Savings Plans over them](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-reserved-instances.html) and still documents buying, modifying, exchanging and reselling RIs. Two capabilities remain RI-only: reselling an unwanted Standard RI on the RI Marketplace, and a zonal RI's capacity reservation. Savings Plans provide no capacity, so pair one with an On-Demand Capacity Reservation if supply is tight. + +For a **steady inference endpoint** like the coding-agent workload these benchmarks model, an **EC2 Instance Savings Plan on the GPU family** is the closest fit: the model has to fit the GPU, so cross-family flexibility buys little and costs about six points of discount. Commit only the always-on baseline. Benchmark sweeps like the ones behind these tables are bursty and interruption-tolerant, so they belong on on-demand or Spot rather than under a commitment. Pull real rates with `aws savingsplans describe-savings-plans-offering-rates --service-codes AmazonEC2 --filters name=instanceType,values=` rather than estimating them. + +## The starting point: a fixed-cost machine, not a per-token bill + +A self-hosted model has **no per-token price**. You rent a GPU instance by the hour (e.g. `g6e.12xlarge` at $4.533/hr on a 3-year commitment) and it processes whatever tokens it can. So the only honest cost is derived, not quoted: + +``` +dollars_per_second = dollars_per_hour / 3600 +``` + +Everything below attributes that fixed $/second to tokens. The number is real and defensible — it is what the hardware actually costs to run — unlike the fictional token-priced `total_cost_usd` the quality harness records for self-hosted models. + +## The two lenses (why there are two) + +On a fixed-cost machine the $/hr is **not itemized per token**, so "cost per token" depends on how you attribute the GPU-second. We report both; they answer different questions. + +### Lens A — blended / measured (no assumption) + +Every processed token — prompt or generation — costs the same slice of GPU time: + +``` +blended_cost_per_token = dollars_per_second / (prompt_tokens_per_second + generation_tokens_per_second) +``` + +Both throughputs are measured server-side from vLLM counters over the concurrency window. Input and output cost the **same** per token. This makes no pricing assumption — it just divides the bill by the work done. **This is the primary lens** (see "Why blended is the honest headline" below). + +### Lens B — lab-style split (one convention `w`) + +The commercial-API shape, where input tokens are billed cheaper than output. We introduce one convention: an input token counts as `w` of an output token when splitting GPU time (default `w = 0.25`). + +``` +cost_per_output_token = dollars_per_second / (generation_tps + w * prompt_tps) +cost_per_input_token = w * cost_per_output_token +``` + +`w` is a **chosen convention, not a measurement** — it exists only to produce the familiar input-cheaper-than-output shape for comparison with hosted APIs. It is a secondary view. + +### Cost per task (either lens) + +A "task" is defined by its token counts `N` input : `M` output, taken from real agentic runs (`metrics.json` of the model's `/swe3` sessions): + +``` +task_cost = cost_per_input_token * N + cost_per_output_token * M # Lens B +task_cost = blended_cost_per_token * (N + M) # Lens A (per-token equal) +``` + +Here `N`/`M` are **all processed tokens** — for the blended lens, `N + M` means every token the server handled: fresh input + output **plus cache-read + cache-write**. This matters because the blended rate was measured over every token vLLM processed (cached prefixes included), so the count it multiplies must include them too; and because it is the only token basis that agrees across harnesses (pi reports cache tokens separately, Claude Code folds them into `input` — see the [appendix](#appendix-prompt-caching-and-why-self-hosted-vs-api-costs-are-not-measured-the-same-way)). Pricing only fresh `input + output` would undercount a pi run by ~half while leaving a Claude Code run unchanged. + +The dashboard exposes `N`, `M`, and `w` as adjustable inputs so you can price any task shape. [clients/cost_for_task.py](../self-hosted/vllm/clients/cost_for_task.py) does the same from the CLI for any task run **outside** the harness — because per-token cost is a property of the model + hardware + load, not of the tasks the sweep happened to run. + +### Cost is GPU-seconds times dollars-per-second — NOT wall-clock times dollars-per-hour + +Substitute the blended rate into the task-cost formula and it collapses to something more intuitive: + +``` +run_cost = (dollars_per_second / combined_tokens_per_second) * tokens_processed + = dollars_per_second * (tokens_processed / combined_tokens_per_second) + = dollars_per_second * GPU-seconds +``` + +The blended lens is just **"how many GPU-seconds did this work occupy, times the price of a GPU-second."** That is the whole model. Two things flow from it: + +- **Throughput converts token *work* into GPU-seconds.** `tokens_processed / combined_tok_per_sec` is the time the GPU actually spent crunching tokens — with the idle gaps (agent thinking, tool calls, network) removed. This is why we divide by *measured* throughput and never by wall-clock. +- **Concurrency selects *which* throughput you divide by,** and that is where amortization lives. At concurrency 1 the box sustains a low combined rate (one lonely stream); at its saturation concurrency it sustains a much higher one, because idle gaps in one session are backfilled by other sessions' tokens. Pricing at the saturated rate assumes you keep the box busy — i.e. the hourly cost is shared across concurrent users. **The self-hosted cost is therefore an operating-point assumption, not a free number.** + +**Why not just pro-rate wall-clock?** Because a single agentic session leaves the GPU idle most of the time. Wall-clock pricing charges you full box price for that idle time and, worse, assumes one user owns the whole box. See the GLM-5.2 worked example below. + +#### Worked example: GLM-5.2, 5 SWE tasks on p5en.48xlarge (8xH200, $27.72/hr = the 3-year EC2 Instance Savings Plan rate for the full box) + +From the throughput sweep (the throughput sweep's `performance-summary.json`) and the SWE run (the run's `run-summary.json`). The sweep's cheapest sustainable point is **concurrency 5**, where the server sustains **15,679 prompt + 189 decode = 15,867 combined tok/s** and the KV cache is already 100% full. The 5 tasks processed **41,519,145 tokens** (input+output). On this self-hosted vLLM run the `cache_read`/`cache_write` tokens are a *partition* of `input_tokens`, not additions to it, so they are already counted inside input and are NOT added again (adding them back would ~2x the count; see issue #136 and the appendix). This is measured over **6,288 wall-clock seconds**. + +| Costing method | Calculation | Result | What it assumes | +|---|---|--:|---| +| **Blended / GPU-seconds (what we use)** | `41,519,145 / 15,867 = 2,617 s = 0.727 hr; x $27.72` | **$20.15** | box kept busy at c=5 (shared across ~11 concurrent requests) | +| Wall-clock pro-rate (rejected) | `6,288 s = 1.747 hr; x $27.72` | $48.42 | one user owns the box; idle time billed | +| GPU-seconds at concurrency 1 (rejected) | `41,519,145 / 7,019 = 1.64 hr; x $27.72` | $45.55 | dedicated box, serial single user | + +The $20.15 figure is the **c=5** number. Note it is *lower* than the naive wall-clock estimate ($48.42) — because dividing by measured throughput removes the idle wall-clock (6,288 s elapsed vs 2,617 s of actual token-crunching, ~1.0 hr idle) the agent spent thinking and tool-calling. And it is below the c=1 figure ($45.55) — the gap between those two is exactly the amortization concurrency buys. If you cannot actually run the box at c=5 (e.g. you dedicate it to one serial developer), your true cost is nearer $46, not $20. Always state the operating point. + +Every figure in that table is linear in the hourly rate, so the whole example rescales by one factor: at 1-year ($40.43/hr) the blended figure is $29.39, and at on-demand ($63.296/hr) it is $46.01. That linearity is what [`clients/reprice_performance_summary.py`](../self-hosted/vllm/clients/reprice_performance_summary.py) exploits. + +### The trade-off between the lenses + +- **Blended** is assumption-free and workload-honest, but it prices input and output identically — which looks unfamiliar next to a commercial API invoice. +- **Split** looks like an API bill, but its output-heavy pricing (`cost_out = 4x cost_in` at `w=0.25`) **understates the true cost of an input-heavy workload**. Agentic coding is exactly that (see below), so the split lens's headline "cost per output token" can be misleading — a task that is 98% input tokens is cheap under split's logic but is really consuming the machine's prefill capacity. Report both; **trust blended for this workload.** + +## What agentic coding does to all of this + +Agentic coding has an extreme request shape: **very large, read-heavy prompts and small outputs.** Measured on the **pi** agent driving the single-agent `/swe3` skill (5 real runs each on `mcp-gateway-registry`; token counts from each run's `metrics.json`), the input:output ratio is extremely lopsided, because a single agent replays the whole growing transcript on every turn. (These are pi numbers, not Claude Code: only pi splits prompt tokens into `input + cache_read`, so its prompt-side count is directly comparable across models — see the [appendix](#appendix-prompt-caching-and-why-self-hosted-vs-api-costs-are-not-measured-the-same-way).) + +| Model (pi `/swe3`) | Turns/task | Prompt tokens/task (input+cache) | Output/task | Ratio | +|---|--:|--:|--:|--:| +| glm-5.2 | 75 | ~16.44M | ~107.1K | **153:1** | +| qwen3.6-35b | 41 | ~7.54M | ~41.9K | **180:1** | +| deepseek-v3.2 | 84 | ~8.87M | ~33.1K | **268:1** | +| kimi-k2.7-code | 117 | ~20.36M | ~50.4K | **404:1** | +| nemotron-ultra-550b | 112 | ~27.43M | ~41.4K | **663:1** | + +The model reads the repo, tool-calls, reasons over big files, and emits a comparatively tiny design/patch — so the ratio sits in the **~150:1 to ~660:1** band. This is the extreme prefill-heavy end of the workload spectrum, and it is why the server is prefill/KV-bound (see [References](#references) for external corroboration at 180-220:1). + +**What drives the spread is turn count, not reasoning-token output.** The ratio's numerator grows with **turns**: every turn re-feeds the entire growing transcript as fresh prompt, so a model that takes 112 turns (nemotron) accumulates ~27M prompt tokens while a model that finishes in 41 (qwen3.6-35b) accumulates ~7.5M. It is *not* that reasoning models emit more output — the opposite: the highest-ratio models (nemotron, kimi) emit the *fewest* output tokens per turn (~370-430), while glm-5.2 and qwen3.6-35b emit the *most* (~1,000-1,400/turn) and have the *lowest* ratios. There is no separate reasoning/thinking token field in the metrics — whatever a model streams as thinking is already inside `output_tokens` — so a verbose reasoner would *lower* the ratio (bigger denominator), not raise it. The lever to cut these models' cost is therefore **fewer turns** (better tool use, less thrashing), not suppressing reasoning. + +That shape has three consequences: + +1. **The server is prefill-heavy: it spends far more compute on reading prompts than on generating.** Prompt throughput (~9-13K tok/s) dwarfs generation (~90-145 tok/s). But "prefill-heavy" is not the same as "saturated" — on a healthy server the prefill still completes in a few seconds (measured **prefill ~2-4s mean** across the sweep), because the prompt is processed in large batched chunks, not token-by-token. + +2. **User experience = time-to-first-token, and on a healthy server it is fine and degrades gracefully.** Because each request must prefill ~100K+ tokens before the first output token, **TTFT is the felt latency**. Measured (median) it is **~1-2s uncontended and rises to ~5-8s at concurrency 20** — good, and it degrades gracefully, not off a cliff. TTFT = **queue wait + prefill**; the decomposition shows prefill is a flat ~2-4s and queue wait stays near zero until higher concurrency (p50 0s up to c=7, ~2s by c=20). Report **TTFT as p50/p90**, never the mean: the mean is distorted by a few cold-cache outliers and by how many requests complete in the window, which can make it move in the wrong direction. + + Full curve (qwen3.6-35b, multi-repo, 25 repos, 10-min windows per level, healthy server): + + | c | gen t/s | prompt t/s | TTFT p50 | TTFT p90 | queue p50 | prefill mean | blended $/1M | task $ | + |--:|--:|--:|--:|--:|--:|--:|--:|--:| + | 1 | 145 | 12202 | 2s | 20s | 0s | 4s | 0.10 | 0.16 | + | 2 | 114 | 13252 | 1s | 20s | 0s | 3s | 0.10 | 0.14 | + | 5 | 87 | 9441 | 2s | 20s | 0s | 3s | 0.14 | 0.21 | + | 7 | 111 | 10168 | 5s | 20s | 0s | 3s | 0.12 | 0.19 | + | 10 | 103 | 10326 | 5s | 40s | 1s | 3s | 0.12 | 0.19 | + | 15 | 109 | 10667 | 8s | 80s | 2s | 3s | 0.12 | 0.18 | + | 20 | 134 | 11705 | 5s | 40s | 2s | 2s | 0.11 | 0.16 | + + > **Watch the server state when you measure.** An earlier run of this same sweep reported TTFT p50 pegged at the histogram ceiling (>640s) with queue-wait means of ~135-240s. That was **not** the model's true behavior — the server was in a backed-up state (a stale scheduler backlog from prior experimentation): it completed ~7x fewer requests per window (23 vs 171 at c=10) because they sat queued. A clean re-run gave the healthy 1-8s numbers above. The lesson: the c=1 baseline and the queue-vs-prefill split exist precisely to catch this — if the c=1 TTFT is not a small number, or queue-wait dominates prefill at low concurrency, the server is not in a clean state and the run should be discarded. + +3. **Report cost on the blended (per processed token) lens.** Generation tok/s alone is a poor headline for an input-heavy workload; blended cost counts prompt + generation, i.e. the work the machine actually does. Blended cost stayed in a tight **$0.10-0.14/1M** band across the whole sweep, cheapest at low concurrency, which is the stable, comparable figure. + +## The only lever that lowers self-hosted cost: KV-cache headroom (which trades against context window) + +Since `run_cost = GPU-seconds x $/sec` and `$/sec` is fixed by the instance, the **only** way to lower cost is to shrink GPU-seconds — i.e. raise the sustained combined throughput. And throughput is capped by whichever runs out first: compute (prefill-bound) or KV-cache space (memory-bound). For a large model on a memory-tight box, it is the **KV cache**, and this is where the interesting trade-off lives. + +**The caching benefit is already in the throughput — do not try to credit it again.** vLLM prefix caching (98-99% hit rate on the GLM SWE run) lets the server skip prefill recompute for the repeated conversation prefix; that is *why* prompt throughput is ~15.7K tok/s rather than a fraction of it. Higher throughput → fewer GPU-seconds → lower cost. The cache payoff shows up as the low blended rate, not as a per-token discount — see the [appendix](#appendix-prompt-caching-and-why-self-hosted-vs-api-costs-are-not-measured-the-same-way) for why applying both would double-count. + +**But prefix cache and live-request KV compete for the same HBM, and that caps concurrency.** GLM-5.2-FP8 is ~744 GB of weights on a 1,128 GB p5en, leaving only ~380 GB for KV. At a **300K-token context window** each concurrent agentic session reserves a large KV slab, so the box saturates KV at only ~11 running requests (**concurrency 5** — where the sweep shows `kv_cache_usage.peak = 1.00` and `requests_waiting` starts climbing). Beyond c=5 throughput does *not* rise; extra sessions just queue. So GLM on p5en has **no vertical headroom** — it is already at its cheapest sustainable point at c=5. + +**The one real lever is a smaller context window.** Reduce the per-session KV footprint (e.g. 300K → 128K) and more sessions fit before KV saturates → higher combined throughput → fewer GPU-seconds per task → lower `$/task`. On a KV-bound model this is essentially the *only* knob short of cheaper hardware (reserved/spot) or a smaller model. + +**There is no free lunch — the window is also an accuracy lever.** Agentic coding *uses* long context: input-per-call climbs from ~50K early in a session to ~200K deep in it as the transcript, tool outputs, and read files accumulate (this is the ~150:1 input:output ratio, corroborated by external agentic-coding data at 180-220:1 — the GLM run's 153:1 is normal, not anomalous; see [References](#references)). Truncate the window and the model loses earlier reasoning and file context, which can lower task quality. So **cheaper serving (more KV headroom → more concurrency → higher throughput) is bought with a shorter context window that risks lower accuracy.** The right operating point is workload-specific: measure score vs window, don't assume. + +## Scaling: it depends whether the model is KV-bound or has headroom + +**Two regimes, and you must know which one you are in — do not assume vertical headroom.** + +**Regime A — KV-bound (e.g. GLM-5.2 on p5en, above):** the model nearly fills the box, KV saturates at low concurrency, and there is *no* vertical headroom. Concurrency is capped by memory, not latency budget. Your levers are context-window reduction, cheaper hardware, a smaller/right-sized model (a small-active MoE that fits at TP=4 leaves far more KV room and serves many more concurrent sessions), or horizontal replicas — **not** "push concurrency higher on this box." + +**Regime B — headroom (e.g. qwen3.6-35b on g6e.12xlarge):** on a **healthy** server this workload is not at a hard ceiling — TTFT p50 is only 5-8s at c=20, KV cache never saturates, and there are no preemptions. So there is genuine vertical headroom: this instance can take more concurrency before latency becomes a problem, and the serving defaults (chunked prefill on, batch budget auto) already use it well. Push concurrency up until p90 TTFT or KV pressure crosses your latency budget — that is the per-instance operating point (the dashboard's recommended-concurrency banner picks the cheapest-blended point; temper it with the p90 TTFT you can tolerate). + +**How to tell which regime you are in:** look at `kv_cache_usage.peak` across the sweep. If it hits 1.00 at low concurrency and throughput stops rising there, you are KV-bound (Regime A). If it stays well under 1.0 as concurrency climbs, you have headroom (Regime B). + +Beyond that per-instance limit, scale **horizontally**: the workload is embarrassingly parallel (N developers on N repos are N independent sessions, nothing to synchronize), so add replicas behind a load balancer. The **blended cost per token is roughly flat across replicas** — each instance has the same $/hr and the same throughput profile — so cost scales linearly with load and the per-task cost measured on one instance is the per-task cost at fleet scale: **measure once, multiply by replicas for capacity.** + +## A third cost basis: managed-model credits (kiro-cli) + +The two lenses above turn a GPU's hourly price into a cost per task for **self-hosted** models. A hosted API (Anthropic on Bedrock) uses its **metered per-token bill**. The **kiro-cli** harness introduces a third basis again: kiro-cli drives Kiro's managed, Bedrock-backed models and bills in **credits**, not tokens or GPU-seconds. See [kiro-cli-setup.md](kiro-cli-setup.md) for install and the harness constraints. + +**What kiro-cli reports.** A non-interactive kiro-cli run emits no token counts. It prints a one-line summary to stderr on completion -- `▸ Credits: • Time: s` -- so the per-run cost signal is **credits consumed** (and wall-clock time). The credits figure already includes the model's `rate_multiplier` (from `kiro-cli chat --list-models`: claude-opus-5 at 2.2 burns credits faster than qwen3-coder-next at 0.05), so it is not multiplied by the rate again. + +**Credits to dollars.** + +``` +cost_per_task_usd = credits_consumed_for_the_run x DOLLARS_PER_CREDIT +``` + +Kiro's published pricing gives two defensible per-credit rates: + +| Basis | $/credit | Derivation | +|---|--:|---| +| Blended (included monthly allotment) | **$0.02** | Every paid tier is the same rate: Pro $20/1,000, Pro+ $40/2,000, Pro Max $100/5,000, Power $200/10,000 | +| Marginal (add-on / overage) | **$0.04** | "Add-on credits $0.04/credit" once the monthly allotment is spent | + +`DOLLARS_PER_CREDIT` is a **configurable rate**, the same stance this page takes on the self-hosted GPU rate (a documented commitment term you swap for your own). The default is the **$0.04 marginal** rate -- the honest "what does one more task cost" figure -- with $0.02 available for an all-you-can-use blended view. Example: a run reporting `Credits: 0.21` costs `0.21 x $0.04 = $0.0084` (or `$0.0042` blended); a real swe task at 50-300 turns consumes far more. + +**How the harness records it.** `run-swe-headless.py` captures kiro-cli's output, strips the ANSI color codes, and regex-parses the `Credits:` value from the summary line. It multiplies that by `kiro_dollars_per_credit` -- a config knob (default 0.04) settable in `runner.yaml` or with `--kiro-dollars-per-credit` -- to get the run's `total_cost_usd`. Each task's `metrics.json` stores **both** the raw `kiro_credits` (provenance) and the derived `total_cost_usd`, and `summarize_run.py` averages `total_cost_usd` across the run's tasks into the reported `$/task` (the `mean_cost_usd_excl_failed` field). So the dollar figure is always traceable back to the exact credits kiro-cli charged. + +**Important: Kiro is a per-developer monthly subscription, not pure usage-based pricing -- and this figure ignores that.** Kiro sells seats (see [kiro.dev/pricing](https://kiro.dev/pricing/)): Free ($0/mo, 50 credits), Pro ($20/mo, 1,000), Pro+ ($40/mo, 2,000), Pro Max ($100/mo, 5,000), Power ($200/mo, 10,000). Those credits are **included in the seat**, and the $0.04/credit rate applies **only to add-on/overage credits once the monthly allotment is spent**. The `credits x $0.04` cost this repo reports therefore treats **every credit as marginal overage** -- as if the monthly allotment were already exhausted -- which is the conservative worst case. For a developer working **within** their allotment, the marginal dollar cost of one more task is effectively already paid by the seat (up to the cap); the amortized rate is closer to the blended **$0.02/credit** (seat price / included credits). Set `kiro_dollars_per_credit` to reflect your plan and expected volume. + +**This makes the cross-harness dollar comparison structurally uneven, not just a different unit.** pi and Claude Code on Bedrock are **pure usage-based, per-token** billing -- no seat, no monthly commitment; you pay only for the tokens you consume. Kiro bundles a **fixed monthly seat plus an included credit allotment**. So comparing kiro's per-task credit cost against Bedrock metered dollars compares a *subscription-plus-credits* model against a *usage-based* one. For a real total-cost comparison, model kiro's **monthly seat cost + expected task volume** against the others' metered (Bedrock) or hardware-derived (self-hosted) spend -- do not read the single per-task dollar figure as directly equivalent. + +**Do not compare raw dollars across the three bases.** Metered Bedrock dollars, hardware-derived self-hosted GPU-seconds, and Kiro credits are measured on different footings (and, per the note above, the credit-to-dollar conversion depends on your Kiro plan and whether you are within your monthly allotment). As with the metered-vs-self-hosted comparison, treat any cross-basis dollar tie as an order-of-magnitude result and state the provenance; compare within a basis. + +## Summary + +- Cost is derived from `instance $/hr / measured tokens/sec` — real, not a quoted price. The formula collapses to **`GPU-seconds x $/second`**: price the tokens by dividing them by *measured throughput*, never by wall-clock (which over-charges idle agent-thinking time and assumes one user owns the box). Worked example: GLM-5.2's 5 SWE tasks cost **$29.90 at c=5**, vs $72 naive wall-clock and $68 at c=1 — the operating point *is* the number, so always state it. +- **Blended** (per processed token, input == output) is the honest primary lens; **split** (`w`-weighted, API-shaped) is a familiar-but-misleading secondary lens for input-heavy work. Caching is already baked into the blended rate (it raises throughput) — do not also discount cache-read tokens per-token, that double-counts. +- Agentic coding is input-heavy (~50:1 early, ~150-220:1 deep in a session), so the server is prefill-heavy — but on a healthy instance TTFT is a few seconds and degrades gracefully; **report TTFT as p50/p90 (not mean) and decompose queue vs prefill**, and always include a c=1 baseline to catch a backed-up server. +- **The only lever that lowers self-hosted cost is raising sustained throughput, which on a KV-bound model means more KV headroom — bought by shrinking the context window, at a possible accuracy cost. No free lunch.** Check `kv_cache_usage.peak`: if it pegs at 1.00 at low concurrency (Regime A, e.g. GLM-5.2 on p5en) there is no vertical headroom — right-size the model/window or scale horizontally; if it stays low (Regime B, e.g. qwen3.6-35b on g6e) push concurrency up to your latency budget first. +- Blended per-task cost is flat across replicas: plan capacity from the measured per-task cost — **measure once, multiply by replicas.** + +## Companion: picking the best harness per model + +The combined cost/quality chart plots one point per model rather than one per model-and-harness, which means something has to choose between a model's Claude Code run and its pi run. That choice uses **both** axes -- Pareto dominance first, then the lower cost per point as the tie-break --. Note the plotted point is therefore not always the model's highest score. + +## References + +Public data on the input-heavy / prefill-heavy shape of agentic-coding workloads — the external corroboration for the ~150:1 input:output ratio and the "prefill-bound, not decode-bound" framing used above: + +- [Together.ai — Benchmarking inference at scale: coding agents](https://www.together.ai/blog/coding-agent-benchmarks) — per-request accumulated context ~80-100K input vs ~450 avg output; implied **~180:1 to 220:1**. +- [Requesty — The Coding Agent Economy](https://www.requesty.ai/coding-agent-economy) — avg input per call ~84K (Claude Code), ~95K (OpenCode); output "negligible." +- [Applied Compute — Benchmarking inference on agentic workloads](https://www.appliedcompute.com/research/inference-benchmark) — per single assistant turn ~10K input : ~200-300 output (**~33:1 to 50:1** early-session). +- [dstack — Benchmarking prefill-decode ratios](https://dstack.ai/blog/benchmarking-pd-ratios/) — prefill/decode contrast; reasoning workloads sit at the opposite (~1:3) regime. + +The ratio climbs through a session because each tool call replays the growing transcript as fresh input while generating only a small edit — ~30-50:1 early, ~180-220:1+ deep in a session — which is why the server is prefill/KV-bound and why KV-cache headroom (not decode speed) is the cost lever. + +## Appendix: prompt caching, and why self-hosted vs API costs are not measured the same way + +*Measurement-plumbing detail, not part of the core cost model. It explains how `total tokens processed` — the count the blended cost multiplies — is derived, and why it depends on whether the backend's cache fields are a partition of `input_tokens` or additions to it. Numbers below are from the `/swe3` runs on `mcp-gateway-registry`.* + +Comparing a self-hosted model's `$/task` against a hosted API model's (e.g. Claude on Bedrock) is the most error-prone part of this analysis, because **the two paths account for cached tokens completely differently.** This is not a modeling choice — it is what each backend reports back to the client. + +- **Anthropic API / Bedrock** implements explicit prompt caching and returns `cache_read_input_tokens` / `cache_creation_input_tokens` in every response's `usage`. Claude Code records those, so on a Bedrock run the reused context lands in `cache_read_tokens` (billed at ~10% of the input rate) and the fresh, full-price `input_tokens` is tiny. Example from an Opus-4.8 `/swe3` task: **`input_tokens: 461`, `cache_read_tokens: 25,260,499`** — ~99.99% of the prompt served from cache. Opus's per-task cost is therefore dominated by *output* tokens (verbose, priced high), not input. +- **On the self-hosted path the two harnesses report cache tokens differently, because they read different fields off vLLM's OpenAI-compatible response:** + - **Claude Code** does not populate the Anthropic-specific cache fields for a vLLM endpoint, so it records **`cache_read_tokens: 0`** and books the entire (re-fed, growing) conversation as fresh `input_tokens`. Example from a GLM-5.2 Claude Code `/swe3` run: **`input_tokens: 86,412,298`, `cache_read_tokens: 0`.** + - **pi** *does* surface vLLM's prefix-cache accounting, so it reports a `prefix_cache_hit_rate` and splits the prompt into `cache_read_tokens` + `cache_write_tokens`, both of which are a **partition of** `input_tokens` — not additions to it. The same GLM-5.2 model under pi `/swe3`: **`input_tokens: 40,983,637`, `cache_read_tokens: 40,663,808`** — i.e. `cache_read` is **99.2% of `input`** (the prefix-cache hit rate), because `input_tokens` already counts the full prompt and `cache_read`/`cache_write` merely describe how much of it was served from cache. `cache_read + cache_write ≈ input_tokens` is the signature of this partition. + +**The two harnesses do NOT process the same number of tokens, and the total-processed basis depends on how each backend accounts for cache.** Claude Code's GLM-5.2 prompt total is `input` ≈ 86.4M (it folds cache into input, `cache_read` = 0); pi's is `input` ≈ 41.0M (of which ~99% was served from cache). These are genuinely different amounts of work — the two runs took different numbers of turns / grew conversations of different lengths — not the same prompt "categorized differently," which is what an earlier version of this note wrongly claimed by adding pi's `cache_read` on top of its `input` (`input + cache_read` ≈ 81.6M) to force an apparent match. That double-counted the cached prompt and inflated every self-hosted total (and therefore cost) by ~2x (issue #136). The correct **total tokens processed** is partition-aware: + +- **Self-hosted vLLM (cache is a partition of `input`):** `total = input + output`. The cache is already inside `input`, so it must not be added again. (GLM-5.2 pi `/swe3`: 41.0M input + 0.5M output ≈ 41.5M, not 82.7M.) +- **Anthropic / Bedrock (cache is additive to `input`):** `total = input + output + cache_read + cache_write`. Here `input_tokens` is only the fresh, uncached tokens (often ~2), and the reused prompt lives separately in `cache_read`, so it must be added. + +The blended `$/token` rate is measured server-side over each token counted **once** (`vllm:prompt_tokens_total` + `vllm:generation_tokens_total`), so the count it multiplies must also count each token once — which is exactly what the partition-aware total does. + +**This does NOT mean either model failed to cache.** vLLM's `--enable-prefix-caching` (on by default here) caches the KV of repeated prefixes server-side and reuses them across the growing agentic conversation. pi surfaces that reuse in `cache_read_tokens`; Claude Code does not, so its client-side `input_tokens` overstates the fresh work — but the GPU did the same thing under both. + +**The caching is measured server-side, and `/swe3` pi captures it directly.** Each pi `/swe3` task records a true server-side `prefix_cache_hit_rate` (from vLLM's `vllm:prompt_tokens_cached_total` / `prompt_tokens_total`). Over the `/swe3` runs on `mcp-gateway-registry` the mean prompt-cache hit rate — the self-hosted analogue of Anthropic's `cache_read_tokens` fraction — was uniformly high across **every** self-hosted model, on both node types: + +| Model (self-hosted, pi `/swe3`) | Node | Mean `prefix_cache_hit_rate` | +|---|---|--:| +| qwen3-coder-480b | p5en.48xlarge | **98.8%** | +| glm-5.2 | p5en.48xlarge | **98.7%** | +| minimax-m2.5 | p5en.48xlarge | **98.3%** | +| deepseek-v3.2 | p5en.48xlarge | **98.2%** | +| devstral-2-123b | p5en.48xlarge | **98.2%** | +| kimi-k2.7-code | p5en.48xlarge | **97.5%** | +| nemotron-ultra-550b | p5en.48xlarge | **96.7%** | +| gemma-4-31b | g6e.12xlarge | **96.6%** | +| qwen3-coder-30b | g6e.12xlarge | **95.9%** | +| qwen3.6-35b | g6e.12xlarge | **95.2%** | + +So on `/swe3`, **~95-99% of the prompt tokens Claude Code would have counted as fresh input were actually served from vLLM's prefix cache** — exactly the tokens Anthropic would have reported (and discounted) as `cache_read`. The single-agent `/swe3` shape replays a stable prefix on every turn, so it caches uniformly well across models and node types. To estimate an API-style billable-input for a self-hosted run: `billable_input ~= input_tokens x (1 - hit_rate)`. + +**Why the hardware-derived cost is still fair despite the 0 in the Claude Code client.** The blended `$/token` already bakes the caching in: it comes from *measured throughput* (tokens/sec the server actually sustained), and that throughput was achieved *with* prefix caching active. So a self-hosted `$/task` is not penalized for the un-credited `input_tokens` — the cheap per-token rate reflects a GPU that was mostly reusing cached prefills. The client-side token *count* is inflated; the *cost* is not. + +**Bottom line for cross-path comparison.** When a hosted-API model (small billable input, cached) lands near a self-hosted model (huge counted input, cheap per token) on `$/task`, treat it as an **order-of-magnitude** result, not an exact tie — the two token counts are measured on different bases. State the provenance (metered API bill vs hardware-derived) alongside the number. diff --git a/docs/diagram-ascii.md b/docs/diagram-ascii.md new file mode 100644 index 00000000..71a925ae --- /dev/null +++ b/docs/diagram-ascii.md @@ -0,0 +1,54 @@ +# The workflow diagram, in ASCII + +An ASCII rendering of how the benchmark and `swe-router` fit together, for terminals and for anywhere an image will not render. + +Change one and change the others: the HTML is the source, the PNG is a screenshot of it, and this is a hand-composed replica. + +``` +MEASURE ONCE, SPEND LESS ON EVERY TASK +Everyone knows the top model is overkill for most tasks. +This makes the cheaper choice the automatic one. + ++- PLATFORM TEAM: a cron job, weekly, unattended ------------+ +- EVERY DEVELOPER: every task ----+ +| 1 . MEASURE SHIPS AS THE | | 2 . ROUTE | +| self-hosted . managed service swe-router SKILL | | Claude Code . Codex . pi . ... | +| . vendor API | | | +| +------------------------+ | | +------------------------------+ | +| +----------------------------+ | models.json | | | | A task begins | | +| | Available models | | score + cost, per | | | | swe-router is installed in | | +| | whatever security | | tier | | | | the harness - it engages | | +| | approved | | | | | | on its own, nobody invokes | | +| +----------------------------+ | allowed-models.txt | | | | it | | +| | the approved list, | | | +------------------------------+ | +| +----------------------------+ | now the filter | | | | | +| | Your dataset | | | | | +------------+------------+ | +| | tasks from your repo | | route.py | | | v v | +| +----------------------------+ | the decision, as | | | +------------+ +------------+ | +| | | code | | | | How bad if | | How hard | | +| v | | | | | wrong? | | is it? | | +| +----------------------------+ | One curl to install. | | | | -> a | | -> a | | +| | Run the benchmark | | Every developer | | | | floor | | table | | +| | every model x every | | reads the same | | | +------------+ +------------+ | +| | task, judged | | numbers on the same | | | +------------+------------+ | +| +----------------------------+ | day. | | | v | +| | +------------------------+ | | +==============================+ | +| v | | | CHEAPEST MODEL OVER THE FLOOR| | +| +============================+ | | | ranked over what this | | +| | YOUR FRONTIER | | | | developer can actually | | +| | not a vendor claim, not | | | | select | | +| | a public set that leaked | | | +==============================+ | +| | into training data | | | | +| +============================+ | | | ++------------------------------------------------------------+ +----------------------------------+ + +==================================================================================================== +PLATFORM TEAM GETS Up to 88% less per task, against running the top model on every task + - measured on your own repo, not claimed. + +DEVELOPERS GET No decision. The right model arrives with the task; nobody weighs + quality against the bill, twenty times a day. +``` + +--- + +[< Back to the README](../README.md) diff --git a/docs/getting-started.md b/docs/getting-started.md new file mode 100644 index 00000000..6b18bc40 --- /dev/null +++ b/docs/getting-started.md @@ -0,0 +1,62 @@ +# Getting started + +What to install, and the order to do it in, from a fresh box to a first benchmark run. + +## Prerequisites + +> **On a fresh machine, start with the [`/setup-machine` skill](../.claude/skills/setup-machine/SKILL.md).** It inspects the box, prints exactly what it will install and why, installs it, and summarizes -- so you do not have to work through the list below by hand. It also installs the GPU stack (vLLM, nvtop, nvitop) only when a GPU is present, and puts the vLLM venv and its caches on the large ephemeral NVMe when the root disk is too small. + +- An **AWS account** with [Amazon Bedrock model access](https://console.aws.amazon.com/bedrock/home#/modelaccess) enabled for the models you want (Paths 1 and 2). +- **AWS credentials** configured locally (`aws configure`, an IAM role, or AWS SSO). +- **[Claude Code CLI](https://docs.anthropic.com/en/docs/claude-code)** installed. +- **[uv](https://docs.astral.sh/uv/)** and **Python 3.10+** for the harness. +- For Path 3: permission to launch an **EC2 GPU instance** (for example `g6e.12xlarge`). + +> The `bedrock-mantle` endpoint used for Path 2 (third-party models) is available in **`us-east-1`**. + +## Get started + +1. **Set up the machine -- do this first on any new box.** Run the **`/setup-machine` skill** from Claude Code (or its script directly). It reports what the instance is, lists every missing dependency with the reason each one is needed, installs them, and prints a summary table: + + ```bash + # Dry run: report only, install nothing + .claude/skills/setup-machine/setup-machine.sh --check + + # Install everything missing (git identity is required, never guessed) + .claude/skills/setup-machine/setup-machine.sh --install \ + --git-name "Your Name" --git-email "you@example.com" + ``` + +Add `--with-omp` / `--with-kiro` to include those two harnesses (opt-in: both ship third-party install scripts, and kiro-cli needs an interactive sign-in). See [.claude/skills/setup-machine/SKILL.md](../.claude/skills/setup-machine/SKILL.md) for the full component list and flags. + +2. **Set up the harness** (its own isolated virtual environment): + + ```bash + cd benchmarks + uv sync + cp config/runner.example.yaml config/runner.yaml + ``` + +3. **Wire the agent CLIs to Amazon Bedrock.** Installing `claude` and `codex` does not configure them -- an unconfigured `codex` silently calls `api.openai.com` and 401s mid-run. Follow [benchmarks/docs/agent-cli-bedrock-setup.md](../benchmarks/docs/agent-cli-bedrock-setup.md). + +4. **Run a benchmark.** The fastest way is the **`/benchmark` skill** from Claude Code, which drives the whole flow interactively -- pre-flight checks, the harness run over a dataset, and the judge -- for any of the three paths. It even manages the vLLM server and metrics collector for the self-hosted path: + + ``` + /benchmark provider=vllm model=qwen3.6-35b dataset=dataset/mcp-gateway-registry.yaml + ``` + +Prefer a script? The same flow runs headless via [benchmarks/scripts/run-e2e-benchmark.sh](../benchmarks/scripts/run-e2e-benchmark.sh) (`--provider bedrock|litellm|vllm --model ... --dataset ...`). + +5. **Pick a path and follow its guide** for the setup details each one needs -- every guide ends with a copy-pasteable run command: +- [Path 1 - Anthropic models directly on Amazon Bedrock](../benchmarks/docs/path-anthropic-on-bedrock.md) +- [Path 2 - open-weight models on Amazon Bedrock via a LiteLLM proxy](../benchmarks/docs/path-open-weight-on-bedrock-litellm.md) +- [Path 3 - self-hosted open-weight models on EC2 with vLLM](../benchmarks/docs/path-self-hosted-vllm.md) + +6. **Read the shared mechanics** once (they apply to every path): the [harness reference](../benchmarks/docs/harness-reference.md) covers the dataset format, the runner config, running the benchmark, the metrics file, and the judge. + +For Path 3 you must first stand up the vLLM server itself -- see [self-hosted/vllm/README.md](../self-hosted/vllm/README.md) (or let the `/benchmark` skill start it for you). + + +--- + +[< Back to the README](../README.md) diff --git a/docs/hosting-paths.md b/docs/hosting-paths.md new file mode 100644 index 00000000..6d954537 --- /dev/null +++ b/docs/hosting-paths.md @@ -0,0 +1,50 @@ +# The three hosting paths + +Where the model runs and how the request reaches it. All three paths run the same agent, tasks, skill and scoring. + +## The three hosting paths + +Whichever path you choose, the agent (Claude Code), the tasks, the `/swe` skill, and the scoring are identical -- only *where the model runs and how the request reaches it* changes. + +```mermaid +flowchart TD + subgraph Harness["Benchmark harness (benchmarks/)"] + CC["Claude Code CLI
(the coding agent)
speaks Anthropic Messages API"] + end + + BedrockA["Path 1
Amazon Bedrock
Anthropic route
───────────────
Claude Opus · Sonnet · Haiku"] + Proxy["LiteLLM proxy (we run it)
Anthropic ↔ OpenAI translation"] + BedrockM["Path 2
Amazon Bedrock (mantle endpoint)
───────────────
Kimi · Qwen · DeepSeek · Mistral …
(any open-weight model on Bedrock)"] + VLLM["Path 3
EC2 GPU node · vLLM
───────────────
your self-hosted open-weight model"] + + CC -- "Anthropic Messages
(provider: bedrock)" --> BedrockA + CC -- "Anthropic Messages
(provider: endpoint)" --> Proxy + Proxy -- "/v1/chat/completions" --> BedrockM + CC -- "Anthropic Messages
(provider: endpoint, SSH tunnel)" --> VLLM + + classDef agent fill:#E5E7EB,stroke:#6B7280,color:#111827 + classDef proxy fill:#EDE9FE,stroke:#7C3AED,color:#3B0764 + classDef bedrock fill:#FFF3E0,stroke:#FF9900,color:#1F2937 + classDef ec2 fill:#E0F2FE,stroke:#0284C7,color:#0C4A6E + class CC agent + class Proxy proxy + class BedrockA,BedrockM bedrock + class VLLM ec2 +``` + +| | Path 1 - Anthropic on Bedrock | Path 2 - open-weight on Bedrock (LiteLLM) | Path 3 - self-hosted on EC2 (vLLM) | +| --- | --- | --- | --- | +| **Which models** | Anthropic family (Claude Opus, Sonnet, Haiku) | Any open-weight model on Bedrock (Kimi, Qwen, DeepSeek, Mistral, GLM, …) | Any open-weight model you can serve (Qwen3-Coder, GLM, Kimi, …) | +| **Where the model runs** | Amazon Bedrock | Amazon Bedrock | Your EC2 GPU instance | +| **How Claude Code reaches it** | Directly, native Anthropic route | Through a [LiteLLM](https://github.com/BerriAI/litellm) proxy we run that translates Anthropic ↔ OpenAI | Directly to your vLLM server (over an SSH tunnel) | +| **Cost model** | Pay-per-token | Pay-per-token | Fixed hourly GPU cost | +| **Extra infrastructure** | None | The LiteLLM proxy ([one script](../benchmarks/scripts/bedrock-mantle-proxy.sh)) | An EC2 GPU node running vLLM | +| **Best for** | Benchmarking the Anthropic family | Model variety with zero infrastructure to manage | Data sovereignty, air-gapped, and high-volume workloads where fixed GPU cost beats per-token pricing | +| **Operational guide** | [Path 1](../benchmarks/docs/path-anthropic-on-bedrock.md) | [Path 2](../benchmarks/docs/path-open-weight-on-bedrock-litellm.md) | [Path 3](../benchmarks/docs/path-self-hosted-vllm.md) | + +The key enabler for Path 2 is the LiteLLM proxy. Claude Code speaks the [Anthropic Messages API](https://docs.aws.amazon.com/bedrock/latest/userguide/inference-messages-api.html), which on Bedrock reaches **only** Claude/Anthropic models; the open-weight models are reachable only through Bedrock's OpenAI-compatible [`bedrock-mantle` endpoint](https://docs.aws.amazon.com/bedrock/latest/userguide/inference.html) (Chat Completions). The proxy sits between the two and translates in both directions, so **any open-weight model on Bedrock can be wired into Claude Code** without changing the agent. All 38 third-party models on `bedrock-mantle` support tool calling and streaming. + + +--- + +[< Back to the README](../README.md) diff --git a/docs/how-a-run-works.md b/docs/how-a-run-works.md new file mode 100644 index 00000000..28946de2 --- /dev/null +++ b/docs/how-a-run-works.md @@ -0,0 +1,38 @@ +# What a single benchmark run does + +One task from start to finish: clone the repo, drive the agent, record the metrics, score the artifacts. + +## What a single benchmark run does + +The flow below is identical across all three paths; only the box the request lands in (Bedrock's Anthropic route, the LiteLLM proxy, or your vLLM server) changes. + +```mermaid +sequenceDiagram + participant H as Harness
(run-swe-headless.py) + participant G as GitHub repo + participant CC as Claude Code
(/swe skill) + participant M as Model
(path 1/2/3) + participant J as Judge
(codex_judge.py) + + H->>G: clone repo at pinned ref (temp dir) + H->>CC: claude -p "/swe repo … problem … model …" + loop agent loop (bounded by max_turns) + CC->>M: Anthropic Messages request + M-->>CC: reply (text and/or tool_use) + CC->>CC: run tools (read repo, write artifacts) + end + CC-->>H: 4 artifacts + JSON result (tokens, latency, turns) + H->>H: write metrics.json beside artifacts + H->>G: remove temp clone + J->>J: score the 4 artifacts against the rubric + J-->>H: eval.json (quality scores) merged into metrics.json +``` + +The `/swe2` and `/swe3` skills land **six artifacts** -- four design docs plus the implemented change (`patch.diff`, `implementation.md`) -- but they do **not** run tests or open PRs; whether the design and code are any good is the downstream evaluation the judge (or a human) performs on the artifacts. Full mechanics are in the [harness reference](../benchmarks/docs/harness-reference.md). + +> **"SWE" here means software engineering in general -- not [SWE-bench](https://www.swebench.com/), the specific benchmark dataset.** The `/swe` skill lets you run any model against any task in any repo of your choosing. It is a *harness*, not a fixed benchmark set: compare results across models on the same task, or a single model across tasks of varying difficulty. + + +--- + +[< Back to the README](../README.md) diff --git a/docs/kiro-cli-setup.md b/docs/kiro-cli-setup.md new file mode 100644 index 00000000..427b780d --- /dev/null +++ b/docs/kiro-cli-setup.md @@ -0,0 +1,177 @@ +# kiro-cli setup and benchmark-integration notes + +Install and configuration reference for **kiro-cli** as a coding agent (harness) for this benchmark. This file covers how to install kiro-cli, how it authenticates, how to drive it headlessly, and -- importantly -- the two ways it differs from the `pi` and `claude-code` harnesses that constrain how it can be wired into the benchmark. Read the "Benchmark integration status" section before starting the harness wiring in issue #73. + +## What kiro-cli is + +kiro-cli is the command-line agent from [kiro.dev](https://kiro.dev), and is the successor to the Amazon Q Developer CLI (the installed binary still ships a `q` shim and reports internal `fig_*` / AWS SDK components). It runs an interactive or non-interactive coding agent in the terminal, backed by **Kiro's managed models** (powered by Amazon Bedrock), selected through the CLI rather than a user-supplied endpoint. + +> **As of this version (kiro-cli 2.18.1), kiro-cli supports only Amazon Bedrock-backed managed models.** There is no support for any other backend -- no self-hosted vLLM, no OpenAI-compatible base URL, no third-party provider. Model selection is limited to the managed list returned by `kiro-cli chat --list-models`, all served through Amazon Bedrock. This is the key constraint that shapes the benchmark integration below. + +Verified version at time of writing: **kiro-cli 2.18.1**. + +## Prerequisites + +- Linux, macOS, or Windows. This repo's node is Linux (Amazon Linux / Ubuntu on EC2). +- An AWS sign-in for kiro-cli: a free **Builder ID** (or Google/GitHub social login), or a **pro** IAM Identity Center license. kiro-cli will not run a chat turn until you are logged in. +- Network egress to `cli.kiro.dev`, `prod.download.cli.kiro.dev`, and the AWS sign-in endpoints. + +## Install + +The official installer covers macOS, Linux, and Windows: + +```bash +curl -fsSL https://cli.kiro.dev/install | bash +``` + +It installs three binaries into `~/.local/bin` (ensure that is on your `PATH`): + +- `kiro-cli` -- the launcher and subcommand entry point. +- `kiro-cli-chat` -- the chat/agent backend (`kiro-cli chat` dispatches to it). +- `kiro-cli-term` -- the terminal integration. + +To inspect the installer before running it (recommended on a shared node): + +```bash +curl -fsSL https://cli.kiro.dev/install -o /tmp/kiro-install.sh +less /tmp/kiro-install.sh # review, then: +bash /tmp/kiro-install.sh +``` + +### Verify + +```bash +export PATH="$HOME/.local/bin:$PATH" +kiro-cli --version # -> kiro-cli 2.18.1 +kiro-cli --help-all # full subcommand list +``` + +## Authenticate + +kiro-cli requires a sign-in before any chat turn. Any command that needs a model (for example `kiro-cli chat ... --list-models`) will otherwise drop into an interactive device-code login and block. + +The **free** license accepts a Builder ID or a social login (Google or GitHub). Signing in with a **Google account** via the free-license device flow is confirmed working on this node -- no AWS account of your own is required. + +```bash +# Free license: Builder ID, or Google / GitHub social login (Google confirmed working): +kiro-cli login --license free --use-device-flow + +# Pro (IAM Identity Center): +kiro-cli login --license pro \ + --identity-provider https://.awsapps.com/start \ + --region us-east-1 --use-device-flow + +kiro-cli whoami # confirm the signed-in identity +``` + +`--use-device-flow` prints a URL and a code to confirm in a browser -- use it on headless servers where a browser redirect cannot be handled. + +`KIRO_HOME` redirects the global `~/.kiro` directory (config, settings, session store) to another location -- useful for keeping a per-run or per-profile config isolated from a developer's global setup: + +```bash +KIRO_HOME=/path/to/run-config kiro-cli whoami +``` + +## Headless (non-interactive) use + +The non-interactive form takes a prompt argument, selects a model with `--model`, and pre-approves tools (no operator is present to confirm tool calls): + +```bash +kiro-cli chat --no-interactive --trust-all-tools --model claude-sonnet-5 "Find all TODO comments in src/" +``` + +`--model` is optional (kiro-cli falls back to the `auto` default), but the benchmark always passes it so a run is pinned to a known model. List the model names you can pass to `--model` (requires login): + +```bash +kiro-cli chat --list-models # human-readable +kiro-cli chat --list-models --format json # machine-readable (name, context, rate_multiplier) +``` + +Relevant `kiro-cli chat` flags: + +| Flag | Purpose | +|---|---| +| `--no-interactive` | Run without an interactive session; requires a prompt argument. | +| `--model ` | Select a model from Kiro's managed list (see `--list-models`). | +| `--effort ` | Reasoning effort: `low`, `medium`, `high`, `xhigh`, `max`. | +| `--agent ` | Use a named agent config (see `kiro-cli agent`). | +| `--trust-all-tools` | Auto-approve every tool call (required when non-interactive). | +| `--trust-tools=` | Auto-approve only specific tools, e.g. `fs_read,fs_write`. | +| `--list-models -f json` | List available managed models as JSON (requires login). | + +Prompt context can also be piped in on stdin: + +```bash +cat build-error.log | kiro-cli chat --no-interactive --trust-all-tools "Explain this failure and suggest a fix" +``` + +### Headless output and metrics + +A non-interactive chat turn streams **ANSI-colored narration to stdout** (the agent's edits and tool calls) -- not JSON. It does **not** report input/output token counts. It **does** print a one-line run summary to **stderr** on completion: + +``` + ▸ Credits: 0.21 • Time: 17s +``` + +So the two capturable per-run metrics are **Credits** (Kiro's billing unit) and **wall-clock time**; success is gated on the exit code and the presence of expected artifacts. Credits, not tokens, are the cost signal for a kiro-cli harness. + +### Available managed models + +`kiro-cli chat --list-models --format json` (requires login) returns the managed models and a per-model `rate_multiplier` in Credits -- a relative cost weight. As of this writing the list includes the same families this benchmark uses, for example: + +| Model | Context | Rate (Credit) | +|---|---|---| +| qwen3-coder-next | 256k | 0.05 | +| gpt-5.6-luna | 272k | 0.10 | +| deepseek-3.2 | 164k | 0.25 | +| minimax-m2.5 | 196k | 0.25 | +| claude-haiku-4.5 | 200k | 0.40 | +| glm-5 | 200k | 0.50 | +| auto (default) | 1M | 1.00 | +| claude-sonnet-5 | 1M | 1.30 | +| claude-opus-5 | 1M | 2.20 | +| gpt-5.6-sol | 272k | 2.40 | + +(Full list also includes claude-opus-4.5/4.6/4.7/4.8, claude-sonnet-4/4.5/4.6, gpt-5.6-terra, minimax-m2.1, and internal-only entries. `default_model` is `auto`.) + +### Translating credits to dollars + +kiro-cli bills in **credits**, so a dollar cost per task is `credits_consumed x $/credit`. The per-run credits come from the stderr summary line and **already include the model's `rate_multiplier`** (a run on claude-opus-5 at 2.2 burns credits faster than one on qwen3-coder-next at 0.05), so do not multiply by the rate again. + +Kiro's published pricing gives two defensible per-credit rates: + +| Basis | $/credit | Derivation | +|---|---|---| +| Blended (included monthly allotment) | **$0.02** | Every paid tier is the same rate: Pro $20/1,000, Pro+ $40/2,000, Pro Max $100/5,000, Power $200/10,000 | +| Marginal (add-on / overage) | **$0.04** | "Add-on credits $0.04/credit" once the monthly allotment is spent | + +``` +cost_per_task_usd = credits_consumed_for_the_run x DOLLARS_PER_CREDIT +``` + +Treat `DOLLARS_PER_CREDIT` as a **configurable rate**, the same way the self-hosted GPU rate is a documented commitment term in this repo, swappable for your own (see [cost-per-task-methodology.md](cost-per-task-methodology.md)). The default is the **$0.04 marginal** rate -- the honest "what does one more task cost" figure -- with $0.02 available for an all-you-can-use blended view. Example: a run reporting `Credits: 0.21` costs `0.21 x $0.04 = $0.0084` (or `$0.0042` at the blended rate). Trivial tasks cost cents; a real swe task at 50-300 turns consumes far more credits. + +Kiro credits are a **third cost basis**, alongside metered Bedrock dollars and hardware-derived self-hosted GPU-seconds. As with those, compare within a hosting basis rather than reading raw dollars across bases; the credit-to-dollar conversion depends on your Kiro plan. + +### Caveat: Kiro is a per-developer subscription, and this figure ignores that + +Kiro's real pricing is a **per-developer monthly subscription** ([kiro.dev/pricing](https://kiro.dev/pricing/)): Free ($0/mo, 50 credits), Pro ($20/mo, 1,000), Pro+ ($40/mo, 2,000), Pro Max ($100/mo, 5,000), Power ($200/mo, 10,000). Those credits are **included in the seat**; the **$0.04/credit** default applies **only to overage** beyond the monthly allotment. So the `credits x $0.04` cost the harness reports treats **every credit as if it were add-on overage** -- the conservative worst case. A developer working within their monthly allotment has effectively already paid for those credits via the seat; the amortized rate is nearer the blended **$0.02/credit**. + +This matters when comparing kiro-cli to the **pi** and **Claude Code** harnesses: driving those through **Amazon Bedrock is pure usage-based, per-token** billing -- no seat, no monthly commitment. kiro-cli instead bundles a **fixed monthly seat with an included credit allotment**. A fair total-cost comparison models kiro's **seat cost + expected monthly volume** against the others' metered/hardware spend, rather than treating the single per-task credit-dollar figure as equivalent. See [cost-per-task-methodology.md](cost-per-task-methodology.md). + +## Benchmark integration status + +kiro-cli is **not yet wired** into the benchmark harness. Two properties of the tool differ from the `pi` and `claude-code` harnesses and shape how it can be integrated (tracked in issue #73): + +1. **No self-hosted / OpenAI-compatible endpoint.** Unlike `pi` (which points at a vLLM endpoint via `--provider vllm`), kiro-cli talks only to Kiro's managed, Bedrock-backed models and authenticates through AWS. There is no base-URL or custom-endpoint setting. **A "kiro-cli against self-hosted vLLM" run -- the original framing of issue #73 -- is therefore not possible.** kiro-cli can only be benchmarked driving its own managed models. + +2. **No token counts, but a credits + time signal.** `--format json` applies to `--list-models`/`--list-sessions` only; a non-interactive chat turn streams ANSI-colored text to stdout (not JSON) and does **not** report input/output token counts. It **does** print `▸ Credits: • Time: s` to stderr on completion (see "Headless output and metrics"). An integration would normalize **credits** (parsed from stderr), not tokens, as the cost metric, gating success on the exit code and artifact presence. This differs from `pi`/`claude-code`, which emit token- and dollar-level accounting. + +The practical consequence: if kiro-cli is added as a third harness, it would be a **managed-model** harness (its own Kiro models, priced in **Kiro credits** -- a third cost basis alongside metered Bedrock dollars and hardware-derived self-hosted GPU-seconds, so compare within a hosting basis as the repo already does). Per-model credit weights come from `--list-models` (`rate_multiplier`); per-run credits come from the stderr summary line. It cannot be a self-hosted-vLLM harness. Confirm the scope before mirroring the `pi` wiring described in issue #73. + +## References + +- Install and CLI docs: https://kiro.dev/docs/cli/ +- Headless mode: https://kiro.dev/docs/cli/headless/ +- Models: https://kiro.dev/docs/models/ +- Benchmark harness reference: [benchmarks/docs/harness-reference.md](../benchmarks/docs/harness-reference.md) diff --git a/docs/omp-setup.md b/docs/omp-setup.md new file mode 100644 index 00000000..d87b8aad --- /dev/null +++ b/docs/omp-setup.md @@ -0,0 +1,87 @@ +# omp (oh-my-pi) setup + +Install and configuration reference for **[oh-my-pi](https://github.com/can1357/oh-my-pi)** (`omp`, [omp.sh](https://omp.sh)) as a coding-agent harness for this benchmark. omp is a fork of the [pi coding agent](https://github.com/earendil-works/pi-coding-agent) and speaks the same JSON-lines event stream, so the harness reuses pi's result parser. It differs in three ways the harness handles for you, listed under [How the harness drives it](#how-the-harness-drives-it). + +Pick it per run with `--agent omp`. Results land under `swe-benchmark-data//omp////`, separate from every other harness. + +## Install + +```bash +curl -fsSL https://omp.sh/install | sh +omp --version +``` + +The installer puts a single binary in `~/.local/bin/omp` (override with `PI_INSTALL_DIR`). Add that directory to `PATH` if it is not there already -- the pre-flight check fails with `omp CLI not found on PATH` otherwise. + +Two install modes, chosen for you: + +- **Prebuilt binary** (the default when `bun` is absent). Downloads `omp-linux-x64` from the GitHub release, ~186 MB, and runs `omp --version` to prove the binary starts before reporting success. +- **From source via bun** (`--source`, or automatic when a matching-architecture `bun` >= 1.3.14 is already installed). Add `--binary` to force the prebuilt path. + +To pin a version, pass `--ref `. The results in this repo were produced with **v18.0.10**. + +If you would rather not pipe a remote script to a shell, read it first: + +```bash +curl -fsSL https://omp.sh/install -o /tmp/omp-install.sh +less /tmp/omp-install.sh +sh /tmp/omp-install.sh +``` + +### Under a systemd unit + +A `systemd --user` unit starts with a minimal `PATH` that excludes `~/.local/bin`, so `omp` and `uv` both disappear and the run dies at pre-flight. Export a `PATH` inside any script you launch that way: + +```bash +export PATH="$HOME/.local/bin:/usr/local/bin:/usr/bin:/bin" +``` + +## Authentication + +omp reaches models two ways, matching the harness's `provider` setting: + +- **`--provider vllm` / `endpoint`** -- an OpenAI-compatible base URL. The harness writes the provider block for you (see below); no credentials beyond the endpoint's own API key, which is usually the throwaway `local`. +- **`--provider bedrock`** -- Amazon Bedrock, as `amazon-bedrock/`. omp resolves AWS credentials itself, including an EC2 instance role. + +On the Bedrock path the harness logs `could not resolve AWS credentials via 'aws configure export-credentials'`. That warning is harmless when an instance role is present: omp finds the role on its own. Confirm by watching the artifacts appear, or `tail` the event stream. + +## How the harness drives it + +```bash +omp -p --mode json --no-session --auto-approve --max-time=1800 \ + --model / -- "" +``` + +Three differences from pi that the harness papers over: + +1. **Config is YAML, not JSON.** pi reads `models.json`; omp reads `models.yml` (custom providers) and `config.yml` (settings). The harness writes both per run into a private agent directory pinned by `PI_CODING_AGENT_DIR`, which omp inherits from pi, so a run never touches a developer's global `~/.omp`. +2. **No `--skill` flag.** omp's `--skills` is a glob filter over discovered skills, not a path, so the harness inlines the whole `SKILL.md` ahead of the task prompt the way it does for kiro-cli. The trailing `--` ends option parsing, which matters because the inlined SKILL.md opens with `---` (YAML frontmatter) that omp would otherwise reject as a flag. +3. **stdin must be closed.** omp treats an inherited stdin as a piped prompt and blocks waiting for EOF, ignoring the positional prompt. The harness passes `stdin=DEVNULL`. Without it a task hangs until the timeout with no output. + +### Compaction + +omp expresses its compaction trigger as an absolute `compaction.thresholdTokens`, where pi uses `reserveTokens`. The harness converts, reserving a full response plus ~8K of headroom. Without it omp fills the context window to within its default reserve and one capped response overflows, killing the run before the last artifacts are written. + +### The wall-clock cap + +omp has no turn cap, so a model that finishes the work and keeps emitting tokens would run until the harness timeout and then burn a retry. `agent_max_time_seconds` (default **1800**, in `config/runner.yaml`) becomes `--max-time`, letting omp stop itself first. Set it to `0` to disable and rely on the harness timeout alone. + +## Watching a run + +omp buffers nothing to the terminal for tens of minutes at a time. The harness mirrors its events to `/omp-stream.jsonl`, which is the only way to watch a task in flight: + +```bash +uv run benchmarks/scripts/omp_stream_view.py --latest +tail -f path/to/omp-stream.jsonl | uv run benchmarks/scripts/omp_stream_view.py - +``` + +The stream is one line per token, so read it through the viewer rather than raw. omp's own `~/.omp/logs` holds lifecycle debug lines, not the event stream. + +## Known behaviour + +- **Partial artifacts on the first pass.** omp sometimes exits after four of the six artifacts. The harness's top-up pass re-prompts for the missing ones and recovers; a task that needs it costs two agent invocations. Across 63 Bedrock tasks this did not recur systematically. +- **Stream files are large.** A single task's `omp-stream.jsonl` runs from 1 MB to 47 MB. They stay gitignored, including inside the committed `Hello-World` worked examples. + +## See also + +- [agent-cli-bedrock-setup.md](../benchmarks/docs/agent-cli-bedrock-setup.md) -- wiring `omp` and the `codex` judge to Amazon Bedrock. Working AWS credentials are not enough for `codex`; prove it with a real call before a long run. diff --git a/docs/repository-structure.md b/docs/repository-structure.md new file mode 100644 index 00000000..80a59a12 --- /dev/null +++ b/docs/repository-structure.md @@ -0,0 +1,50 @@ +# Repository structure + +What lives where in this repository. + +## Repository structure + +```text +sample-agentic-coding-harness-benchmarks/ +├── README.md ← Start here (concepts, the two steps, the three hosting paths) +├── LICENSE MIT-0 +├── CODE_OF_CONDUCT.md +├── CONTRIBUTING.md +├── SECURITY.md +├── SUPPORT.md +├── THIRD_PARTY Third-party dependency attributions +├── .github/ Issue and pull-request templates +├── .claude/ ← Claude Code skills shipped with the repo +│ └── skills/ +│ ├── setup-machine/ /setup-machine — inspect a fresh box, install every dependency (start here) +│ ├── benchmark/ /benchmark — run one end-to-end benchmark (service + harness + judge) +│ ├── swe/, swe2/, swe3/ /swe* — drive a model through a SWE task on any repo (swe3 is the default) +│ ├── swe-router/ /swe-router — recommend the right model for a task, from these measurements +│ ├── throughput/ /throughput — sweep a served model's throughput +│ ├── security-check/ /security-check — Cipher security review + fix before any commit +│ └── vllm-setup/ /vllm-setup — stand up the EC2 vLLM server (Path 3) +├── docs/ ← Concepts, setup guides, and methodology +├── vend/ ← Vendored, installable artifacts +│ └── swe-router/ /swe-router skill: route.py, models.json, install.sh +├── benchmarks/ ← The benchmark harness +│ ├── README.md Harness landing page +│ ├── docs/ Shared harness reference + one guide per hosting path +│ ├── config/ runner.example.yaml, litellm-mantle.yaml (Path 2 proxy) +│ ├── dataset/ Benchmark dataset YAML files +│ ├── scripts/ Run harness, dataset/config loaders, judges, proxy launcher +│ ├── tests/ Unit tests +│ └── swe-benchmark-data/ Where your runs land. Gitignored — no run output is committed. +└── self-hosted/ ← Path 3: EC2 self-hosted serving (vLLM) + └── vllm/ + ├── README.md Full EC2 + vLLM setup guide + ├── models/ Per-model serving guidelines (one .md per model) + ├── scripts/ vllm-install.sh, vllm-serve.sh, tunnel.sh, … + ├── clients/ Inference + metrics-collection Python clients + ├── tests/ unittest suite for the clients + └── config/ claude-code.json, opencode.json +``` + + +--- + +[< Back to the README](../README.md) diff --git a/docs/serving-optimization-notes.md b/docs/serving-optimization-notes.md new file mode 100644 index 00000000..ca7203e0 --- /dev/null +++ b/docs/serving-optimization-notes.md @@ -0,0 +1,78 @@ +# vLLM serving optimization notes + +A running log of serving-configuration findings for the self-hosted (vLLM) path. Append new entries at the top with a date and the model/hardware they came from. The goal is a small set of **portable, common-sense defaults** we do not have to re-tune per model, plus a record of *why* — so the next model does not repeat the investigation. + +> Companion doc: [cost-per-task-methodology.md](cost-per-task-methodology.md) — how the fixed instance price becomes a cost per token / per task, the two cost lenses, and why the prefill-bound shape found below points to horizontal scaling. + +## TL;DR — the portable defaults (do not tune per model) + +| Knob | Setting | Why it generalizes | +|---|---|---| +| `--enable-chunked-prefill` | **on** (vLLM V1 default) | Interleaves large prefills with decode so one giant prompt cannot freeze generation. Universal win; never turn off. | +| `--max-num-batched-tokens` | **unset** (V1 auto) | V1 derives it from KV-cache size and the model. It self-scales per hardware/model — hand-setting it is what *creates* a per-model chore. | +| `--max-num-seqs` | **unset** | Only set it for the specific boot-time "can't fit N sequences in KV cache" error (`vllm-serve.sh` documents this). That is a correctness fix, not a throughput tune. | +| `--enable-prefix-caching` | **on** | Free KV reuse when prompts share a prefix. Helps a lot in single-repo/single-user workloads; helps little across diverse repos (see below) — but never hurts. | + +**We leave `vllm-serve.sh` on these defaults.** Full config resolved at boot confirms the two prefill knobs are already `enable_chunked_prefill=True` and `max_num_batched_tokens` auto. + +--- + +## 2026-07-26 — CORRECTION: the "prefill-saturated" run below was a backed-up server, not the model + +A clean re-run of the same sweep (same model, hardware, dataset, windows) — this time with a **c=1 baseline** and **TTFT reported as p50/p90 with a queue-vs-prefill split** — showed the server is **healthy**, not saturated: + +- TTFT p50 **~1-2s uncontended, ~5-8s at c=20** (not minutes); prefill mean a flat ~2-4s; queue-wait p50 ~0s until c=7, ~2s by c=20. +- Generation throughput ~90-145 tok/s, prompt ~9-13K tok/s, KV never saturated, zero preemptions. +- At c=10 the clean run completed **171 requests** in the window vs **23** in the run below — the earlier run's requests were stuck in a **stale scheduler backlog** (queue mean 135s vs 5.5s). + +So the entry below correctly ruled out KV pressure, but its "prefill saturation / 2-4 minute TTFT" conclusion was measuring a **transient backed-up server state**, not the workload's true behavior. **Lesson baked into the tooling:** always run a **c=1 baseline** and watch the **queue-vs-prefill decomposition** — if c=1 TTFT is not small, or queue-wait dominates prefill at low concurrency, the server is not clean and the run must be discarded. See [cost-per-task-methodology.md](cost-per-task-methodology.md) for the corrected curve. The portable serving defaults (top of this file) are unchanged and remain correct. + +--- + +## 2026-07-25 — Agentic coding is prefill-bound, not KV-bound (qwen3.6-35b on g6e.12xlarge, 4xL40S) + +> **Superseded — see the 2026-07-26 correction above.** This investigation's KV-vs-prefill reasoning is sound, but its headline numbers came from a backed-up server and overstate the problem. Kept for the reasoning and the metric-artifact lesson. + +**Context.** Running the throughput skill against qwen3.6-35b. Switched the load dataset from `mcp-gateway-registry` (5 tasks, ONE repo) to `multi-repo-throughput` (25 tasks, 25 DIFFERENT repos) and saw generation throughput drop sharply and *fall* with concurrency. Investigated whether it was KV-cache pressure. **It is not** — it is prefill (prompt-processing) saturation, and it is the true cost of this workload, not a mistuning. + +### What the data showed + +Single-repo run (RUN1) vs multi-repo run (RUN2), server-side counters from the DuckDB collector: + +| | RUN1 single-repo | RUN2 multi-repo | +|---|---|---| +| gen tok/s @ c2 | 53.5 | 44.5 | +| gen tok/s @ c5 | 79.8 (rising) | 19.6 (collapsing) | +| prefix-cache hit % | 32-60% | ~14% | +| KV% peak @ c2 (same 2 sessions) | 28.7% | 62.6% | + +Multi-repo, per concurrency level (RUN2): + +| c | gen t/s | prompt t/s | prefill:decode | KV% peak | preempt | wait peak | +|--:|--:|--:|--:|--:|--:|--:| +| 2 | 44.5 | 5936 | 133:1 | 62.6 | 0 | 28 | +| 5 | 19.6 | 5568 | 283:1 | 48.0 | 0 | 35 | +| 7 | 16.4 | 4544 | 277:1 | 28.8 | 0 | 40 | + +### Why it is NOT KV-cache pressure + +- KV usage **falls** as concurrency rises (62 -> 48 -> 29%). Under KV pressure it would pin near 100%. +- **Zero preemptions, zero swaps** at every level. KV pressure shows up as preemptions. +- Requests wait for `reason="capacity"` (the scheduler's per-step prefill token budget), not for KV blocks. + +### Why it IS prefill saturation + +- prefill:decode token ratio is **133:1 at c2, rising to ~280:1** — the GPU spends ~99.5% of its time processing prompts (~5-6K prompt tok/s) and almost none generating (16-44 gen tok/s). +- Agentic coding has an extreme request shape: ~100K+ token read-heavy prompts, tiny outputs. Measured per-task ratio on qwen3.6-35b was **~50:1 input:output** (1.51M in : 30K out across 5 real /swe2 runs). +- The prefix cache had been *masking* this: in the single-repo run, 32-60% of each prompt's KV was already computed and reused, so the fixed prefill budget covered far more effective prompt and left GPU time for decode — gen throughput even *rose* with concurrency. Across 25 distinct repos the hit rate drops to ~14%, nearly every prompt token must be prefilled from scratch, and more concurrency just means more giant prompts contending for the same prefill budget -> gen tok/s *falls*. + +### The takeaways + +1. **The single-repo number was an optimistic best case.** N developers all in one repo with a warm shared cache is the friendliest possible workload, not a representative one. The **multi-repo number is the honest cost** of N developers on N different projects — which is what the cost model should price. +2. **No serving knob fixes this**, because it is not broken. No setting makes 25 distinct 100K-token prefills free. `max_num_batched_tokens` only changes how the fixed prefill work is packed per step; it cannot create FLOPs. Raising it trades latency for throughput but does not change the prefill-bound regime. This is a property of the **workload shape**, not the config. +3. **Do not headline the cost on generation tok/s.** At a ~50:1 input:output ratio it is a misleading denominator that makes a healthy, 99%-busy server look broken. **Headline the blended (per-processed-token) cost** — cost per (prompt + generation) token. That number was stable at ~5-8K processed tok/s across the whole concurrency sweep in *both* runs, because it measures the work the machine actually does. It is model- and hardware-portable by construction and needs zero per-model tuning. The lab-style split lens (input priced at `w` x output) understates cost for this workload and should stay a secondary view. + +### Constraints honored + +- **200K context window is fixed** (`MAX_MODEL_LEN=200000`) — not reduced. The bottleneck is prefill compute, not the window or KV size anyway, so lowering it would not have helped. +- No changes made to `vllm-serve.sh`; defaults are the recommended values. diff --git a/docs/vision.md b/docs/vision.md new file mode 100644 index 00000000..53032db1 --- /dev/null +++ b/docs/vision.md @@ -0,0 +1,82 @@ +# The vision: a cost-aware coding harness that routes across the frontier + +This repository measures a [cost/quality Pareto frontier](../README.md#results-a-worked-example) -- the set of models where nothing else is both better and cheaper. Building that frontier is not the end goal. It is the **lookup table** for the thing we actually want to build: a coding harness (or agent) that, given a task, **intelligently picks the right model from the frontier for that task** -- and switches models mid-task when the situation changes -- so a developer gets frontier-level results at workhorse-level cost without ever having to think about model selection. + +## The problem this solves + +Today a developer picks one model and uses it for everything. That is wasteful in both directions: + +- Using a **frontier model** (e.g. Claude Opus) for every step is expensive -- most of a coding task is routine work a cheaper model does just as well. +- Using a **budget model** for everything risks quality -- hard planning, subtle correctness, and thorny debugging are exactly where the cheap model falls short, and you often do not find out until the work is already going wrong. + +The frontier this repo measures says these are not either/or choices. Different models win in different regions of the cost/quality plane, and different *phases of a single task* live in different regions. + +## What the harness would do + +Given a task, the harness classifies it and routes each phase to the cheapest model on the frontier that clears the quality bar for that phase: + +- **Plan with a frontier model.** Decompose the problem, write the design, decide the approach -- the high-leverage step where quality matters most and token volume is smallest, so paying frontier prices here is cheap in absolute terms. +- **Execute with an appropriate open-weight model.** Generate the artifacts and land the code with a **workhorse** (mid-frontier, e.g. GLM-5.2 / Kimi) or a **budget** model (e.g. a 3B-active MoE at ~$1/task) -- this is the bulk of the tokens, so this is where routing saves real money. +- **Escalate when it matters.** The developer can always ask the harness to switch back to a frontier model; and the harness can decide *on its own* that a run is going badly -- "my initial assessment was wrong, this is harder than it looked" -- and escalate mid-task before it wastes a budget model's turns on something it cannot finish. + +The result: the developer states a task and a budget posture ("cheap", "balanced", "best"), and the harness handles model selection and switching underneath. Three tiers -- **frontier**, **workhorse**, **budget** -- picked and swapped per task and per phase, automatically. + +## The first concrete step: `/swe-auto` + +The first slice of this vision is **`/swe-auto`** -- a router skill that runs on **either Claude Code or pi**. The developer does not choose the model. Given a repo + ref + problem, a configurable **router model** triages the task read-only, classifies it as **frontier / workhorse / budget**, consults the measured [cost/quality Pareto frontier](../README.md#results-a-worked-example) to pick the cheapest non-dominated model that clears that tier's quality band, then shells out to the existing headless runner to run the **`/swe3`** skill with the selected model and harness -- producing the six artifacts (and, optionally, an `eval.json`). If the first pick fails to complete or scores below its band, it escalates one tier and re-runs, bounded by `max_escalations`. + +The key design decision: the skill does the triage and frontier lookup **inline** (cheap), then **executes via `run-swe-headless.py`** rather than spawning a subagent -- so the same executor path works whether the router runs under Claude Code or pi (pi cannot fan out). Note the two independent harness choices: which agent runs the *router skill*, and which agent the *executor* drives `/swe3` under (`--agent`); they can match or differ. + +The sequence below shows one `/swe-auto` invocation end to end -- launch, triage, frontier lookup, execution, optional scoring, and the escalation loop: + +```mermaid +sequenceDiagram + participant Dev as Developer
(Claude Code or pi) + participant SA as /swe-auto skill
(router) + participant RM as Router model
(e.g. claude-opus-5) + participant G as GitHub repo + participant F as Pareto frontier JSON
(GitHub main, raw) + participant R as Headless runner
(run-swe-headless.py) + participant M as Selected model + harness
(runs /swe3) + participant J as Judge
(optional) + + Dev->>SA: /swe-auto repo, ref, problem + SA->>G: clone repo at pinned ref (read-only triage) + SA->>RM: classify this task (problem + relevant code) + RM-->>SA: tier (frontier, workhorse, or budget) + rationale + SA->>F: fetch pareto-frontier for harness + swe3 + F-->>SA: non-dominated models (by frontier_scope) + SA->>SA: map tier to quality band, pick cheapest model that clears it + + loop until artifacts complete and score in band (max_escalations) + SA->>R: run-swe-headless.py --agent HARNESS --model SELECTED --skill swe3 + R->>M: drive /swe3 over the task (bounded agent loop) + M-->>R: six artifacts (issue, lld, review, testing, patch.diff, implementation) + opt judge enabled + R->>J: score the artifacts + J-->>R: eval.json (quality score) + end + R-->>SA: artifacts + metrics (+ eval.json) + alt incomplete or scored below band + SA->>SA: escalate one tier up, re-select model + else complete and in band + SA->>SA: done + end + end + + SA-->>Dev: artifacts + routing.json
(tier, candidates, selected model, rationale, escalations, cost/score) +``` + +## Why this repo is the foundation + +You cannot route intelligently without knowing, per model: + +1. **How good it is** -- the quality benchmark (judge scores, and the [per-dimension breakdown](../README.md#quality-by-dimension-where-models-are-strong-or-weak) that tells you *which kinds* of work each model is reliable at). +2. **What it costs** -- the throughput benchmark, turned into a hardware-derived [cost per task](cost-per-task-methodology.md). +3. **Where it sits relative to every other model** -- the Pareto frontier that combines the two. + +That is exactly what this harness produces today. The routing agent is the next layer built on top of it. Every model we add and every sweep we run sharpens the table the router will read. + +## Status + +The measurement half -- quality benchmarking, throughput benchmarking, and the combined frontier across three hosting paths -- is what exists in this repo now. The routing agent described above is the direction this work is heading, not a shipped feature; its first concrete slice is tracked as `/swe-auto` -- per-task routing that picks one model for the whole task and escalates a tier between runs. Per-*phase* routing (plan with a frontier model, execute with a workhorse) and in-flight mid-run model switching are later steps that build on it. This document is the north star that the benchmark harness is built to serve. diff --git a/docs/why-this-exists.md b/docs/why-this-exists.md new file mode 100644 index 00000000..61fc4743 --- /dev/null +++ b/docs/why-this-exists.md @@ -0,0 +1,27 @@ +# Why this exists + +Why this repo measures harness x model on real repositories, and what the quality and throughput benchmarks each contribute to a cost per task. + +## Why this exists + +Enterprises are adopting coding agents and models at scale, and the bill grows with every developer and every task. The two big levers on that bill -- **which harness** drives the work and **which model** it drives -- get chosen on gut feel or on public leaderboards that may already be **saturated**: models can be tuned toward well-known public test sets, so a high headline number is a poor predictor of performance on a team's actual, messy, long-horizon coding work. + +This repo measures what decides the bill instead: **harness x model, on real agentic software-engineering tasks against real repositories**, reporting all three axes a buyer trades off -- **cost, latency, and accuracy**. Crossing harnesses with models gives real **optionality**: the same model can be a few points more accurate under one agent yet several times cheaper and faster under another. With those numbers in hand, an organization can make an **informed, defensible decision** about the cost/latency/accuracy trade-off for its own workload -- and, very often, **cut its coding bill** by picking a cheaper harness-and-model pairing that is more than good enough, rather than defaulting to the most expensive option. That is the deliverable: an evidence base for smart, budget-aware choices on work that looks like yours, not like a leaderboard. + +## Overview + +This repository is a **benchmark and harness for measuring how well different LLMs perform real-world software-engineering tasks** when driven by a coding agent. It supports **four coding agents (harnesses)** today -- [Claude Code](https://docs.anthropic.com/en/docs/claude-code), Anthropic's command-line coding agent; [pi](https://github.com/earendil-works/pi-coding-agent), a lightweight open-source agent; [oh-my-pi](https://github.com/can1357/oh-my-pi) (`omp`, [omp.sh](https://omp.sh)), a fork of pi -- see [omp setup](omp-setup.md); and [kiro-cli](https://kiro.dev) (the successor to the Amazon Q Developer CLI), which drives Kiro's own managed, Amazon Bedrock-backed models -- with [opencode](https://opencode.ai) being added soon. Claude Code and pi are each wired to run with a model hosted in any of **three different places**, so you can put many models through the *same* tasks with the *same* harness and compare them on both quality and cost; kiro-cli instead runs Kiro's managed models directly (it cannot target a self-hosted endpoint -- see [kiro-cli setup](kiro-cli-setup.md)). Pick the harness per run with `--agent claude` (default), `--agent pi`, `--agent omp`, or `--agent kiro`, and the skill with `--skill swe2`/`--skill swe3`; results are kept separate on disk (`////`) so neither the agents nor the two skills ever overwrite each other. + +It runs **two complementary benchmarks**, and combining them is the whole point: + +1. **Quality** -- how well a model does a real coding task (scored 0-100 by an independent LLM judge). +2. **Throughput** -- how many tokens per second a self-hosted model sustains on a given GPU instance, which turns the instance's hourly price into a **hardware-derived cost per task**. + +Quality alone tells you which model is best; cost alone tells you which is cheapest. Plotting one against the other yields the **cost/quality Pareto frontier** (the chart below) -- the set of models where nothing else is both better *and* cheaper. That frontier is the deliverable: it is what lets you choose a model for a real budget, and it exists only because this repo measures both halves. + +**The two benchmarks must be combined over the *real agentic coding tasks*, not over a synthetic input:output token ratio.** Agentic coding is a **prefill-heavy, long-horizon** workload: each task replays a large, growing transcript as fresh input on every turn and emits a much smaller edit, so the real input:output ratio runs ~150:1 up to ~660:1 -- far more lopsided than the ~3:1 or ~4:1 assumed by generic pricing. A model's cost per task therefore depends on *how* it drives the task (how many turns, how much context it re-reads, whether prefix caching hits), which a lab-style token-count estimate cannot capture. This repo measures throughput on that same prefill-heavy shape and multiplies it by the tokens each run processed, so the cost on the frontier is the cost of the *work as it happens* -- see [cost-per-task-methodology.md](cost-per-task-methodology.md). + + +--- + +[< Back to the README](../README.md)