A test runner for people working with coding agents.
Test whether real GitHub Copilot sessions can use your MCP servers, CLI tools, skills, system prompts, and custom agents. Write a user task, execute it through pytest, and verify the actual output with ordinary assertions.
The framework runs the task and records what happened. Your existing coding agent investigates the evidence, reads the source, and helps you improve the interface. Version 1.0 removes AI judging, report dashboards, rankings, and skill refinement. See the migration guide for this breaking change.
Your tool can pass its own tests while Copilot chooses the wrong operation, supplies the wrong arguments, or misunderstands its descriptions. That is a different question from whether the tool implementation works.
Use real task execution to test the AI-facing interface, and independent output checks to establish correctness. A completed session or a tool call by itself does not prove that the requested task succeeded.
You need Python 3.11+, uv, and a GitHub account with Copilot access.
uv add pytest-skill-engineering
gh auth login --hostname github.com
uv run pytest-skill-engineering init
uv run pytest-skill-engineering doctor
uv run python -m pytest "tests\test_copilot_eval.py" -vinit creates a Todo MCP starter test, explicit JSON evidence settings, and
explicit starter-model rates in pricing.toml. It refuses conflicting settings
or an existing starter. Live execution can consume Copilot premium requests.
pytest shows ordinary test results and assertion failures. Give your coding
agent aitest-reports\results.json and the relevant source to investigate what
happened. No dashboard or second AI judgment is generated.
We evaluated a companion skill for coding agents using this framework and chose not to ship it. The comparison did not demonstrate a benefit; adding general test-writing advice was not evidence that users needed another product layer. Use the documentation, examples, and native evidence directly.
The case study records the actual results, failures, limitations, and removal decision. Its historical experiment remains reproducible with a frozen test fixture, not an installable companion skill. Default sample checks are offline; live execution is explicit. The quickstart is the smaller first example.
Testing your own domain skills remains a supported framework capability.
from __future__ import annotations
import subprocess
import sys
from pytest_skill_engineering.copilot import CopilotEval
async def test_addition(copilot_eval, tmp_path):
agent = CopilotEval(
name="addition",
model="gpt-5.6-luna",
instructions="Write Python code and save the requested file.",
working_directory=str(tmp_path),
)
result = await copilot_eval(agent, "Create calc.py with add(a, b) returning a + b.")
assert result.success, result.error
assert (tmp_path / "calc.py").is_file()
checked = subprocess.run(
[
sys.executable,
"-c",
"from calc import add; assert add(2, 3) == 5; assert add(-2, 2) == 0",
],
cwd=tmp_path,
capture_output=True,
text=True,
timeout=10,
)
assert checked.returncode == 0, checked.stderrUse your approved isolation when executing generated code. A temporary directory is not a security sandbox.
| Capability | What it provides |
|---|---|
CopilotEval and copilot_eval |
Real SDK sessions with explicit configuration |
| MCP, CLI, skills, plugins, custom agents | The interfaces and definitions under test |
| Execution controls | Time, usage, request, tool, and permission controls; one attempt per execution |
| Captured evidence | Configuration, calls, arguments, outputs, completion flags, errors, and usage |
| Ordinary pytest checks | Consumer-owned verification and recorded properties |
ab_run and repetitions |
Isolated working directories and repeated observations |
| Structured JSON | Captured execution and ordinary pytest outcomes, without rankings or advice |
| Pricing | Explicit USD estimates and recorded Copilot premium requests |
A/B entries share one pytest outcome. Record side-specific checks; do not read that shared outcome as independent per-side success or causal improvement. Missing prices and incomplete evidence are recorded explicitly.
Ask your existing coding agent to inspect a failed test, its saved JSON evidence, and the source. It can explain the supported cause and help fix it, without changing the test's success criteria after seeing the failure.
Reading current schema-4.0 evidence needs no authentication or paid call. The framework records results; the coding agent interprets them. Subjective review is advice, not an automatic pytest verdict.
Read the full documentation.
MIT. Inspired by agent-benchmark.