Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
31 commits
Select commit Hold shift + click to select a range
2c321a9
feat: v0.3.0 managed advisory Jev runtime, multi-harness integration,…
Coding-Dev-Tools Sep 28, 2026
5943d75
feat: complete portable Jev integration and qualified evidence selection
Coding-Dev-Tools Sep 28, 2026
a100ae7
test: preserve evidence bytes on Python 3.9
Coding-Dev-Tools Sep 28, 2026
d6c4530
test: make deadline and contention checks portable across Windows run…
Coding-Dev-Tools Sep 28, 2026
9e4275b
fix: bind evidence qualification and add explicit Command Code workflow
Coding-Dev-Tools Sep 28, 2026
6ece687
test: retain case-insensitive Windows environment for capture subproc…
Coding-Dev-Tools Sep 28, 2026
2b957e9
fix: preserve existing keyring credentials during guided reconfiguration
Coding-Dev-Tools Sep 28, 2026
95048e5
test: isolate response contract fixtures from disk accounting timing
Coding-Dev-Tools Sep 28, 2026
27da47d
fix: redact wire question IDs and preserve caller mappings
Coding-Dev-Tools Sep 28, 2026
adc337c
fix: close release review gaps and ship portable evidence capture
Coding-Dev-Tools Sep 29, 2026
09cbde8
fix: preserve caller score legends and default evidence calls off
Coding-Dev-Tools Sep 29, 2026
d60f21f
test: verify bounded ledger waits without a spurious wall deadline
Coding-Dev-Tools Sep 29, 2026
05dfb54
fix: reject unusable credentials before protected storage
Coding-Dev-Tools Sep 29, 2026
f9ed66c
fix: reject anomalous usage with conservative accounting
Coding-Dev-Tools Sep 29, 2026
6249fe0
fix(harness): keep the environment interpreter in generated entries
Coding-Dev-Tools Sep 30, 2026
85b4f50
fix(credentials): treat unexpanded harness references as absent keys
Coding-Dev-Tools Sep 30, 2026
d7d0a70
fix(evidence): screen private names only below the approved root
Coding-Dev-Tools Sep 30, 2026
8d4af48
fix(client): honor an explicit library key before setup
Coding-Dev-Tools Sep 30, 2026
fab5643
build: single-source the package version
Coding-Dev-Tools Sep 30, 2026
685524c
feat(hooks): escalate-only pre-execution shell guard for harness hooks
Coding-Dev-Tools Sep 30, 2026
873b32b
perf(mcp): advertise a compact jev_decide input schema
Coding-Dev-Tools Sep 30, 2026
660a79f
ci: test Python 3.14 and Node 24; neutral credential dialog wording
Coding-Dev-Tools Sep 30, 2026
a6af974
fix(cli): explain why a configured runtime is disabled
Coding-Dev-Tools Sep 30, 2026
2344832
feat(client): accept plain question objects in Python
Coding-Dev-Tools Sep 30, 2026
b640bef
docs: task-first quickstart, hook guide and review migration notes
Coding-Dev-Tools Sep 30, 2026
eb87ad5
fix(hooks): guard Claude Code's PowerShell tool on Windows
Coding-Dev-Tools Sep 30, 2026
1373297
fix(hooks): check reads outside the working directory
Coding-Dev-Tools Sep 30, 2026
22e46de
docs(validation): record the 2026-09-30 merge-readiness review
Coding-Dev-Tools Oct 1, 2026
c0e9165
feat: add structured memory advice and harden Command Code integrations
Coding-Dev-Tools Oct 1, 2026
837d909
Merge local harness updates and preserve explicit Jev consent and evi…
Coding-Dev-Tools Oct 1, 2026
8b79793
fix(tests,ts): add comprehensive client tests, runtime validation, an…
Coding-Dev-Tools Oct 3, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .gitattributes
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
* text=auto eol=lf
*.py text eol=lf
*.md text eol=lf
*.json text eol=lf
*.toml text eol=lf
*.yml text eol=lf
*.sh text eol=lf
*.ps1 text eol=lf
*.ts text eol=lf
*.cjs text eol=lf
55 changes: 52 additions & 3 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ jobs:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python-version: ["3.9", "3.10", "3.11", "3.12", "3.13"]
python-version: ["3.9", "3.10", "3.11", "3.12", "3.13", "3.14"]

steps:
- uses: actions/checkout@v4
Expand All @@ -26,13 +26,62 @@ jobs:
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -e ".[test]"
pip install -e ".[test,mcp,setup]"

- name: Run Pytest
run: |
python -m pytest tests/ -v

- name: Check configured Python lint rules
if: matrix.os == 'ubuntu-latest' && matrix.python-version == '3.12'
run: |
python -m pip install ruff==0.16.9
python -m ruff check .

- name: Test CLI
run: |
jev --help
jev guard "git status"
jev doctor --json

- name: Clean wheel and source installation
if: matrix.python-version == '3.12'
run: |
python -m pip install build
python scripts/check_packages.py --output "${{ runner.temp }}/jev-artifacts"

- name: Retain Python release candidates
if: matrix.python-version == '3.12'
uses: actions/upload-artifact@v4
with:
name: python-release-${{ matrix.os }}
path: |
${{ runner.temp }}/jev-artifacts/dist/*
${{ runner.temp }}/jev-artifacts/verification.json

typescript:
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
node-version: ["22", "24"]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: ${{ matrix.node-version }}
- run: npm ci
working-directory: ts
- run: npm test
working-directory: ts
- run: npm run check:package
working-directory: ts
- uses: actions/upload-artifact@v4
with:
name: npm-release-${{ matrix.os }}-node${{ matrix.node-version }}
path: ts/*.tgz
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: python -m pip install -e ".[test]"
- run: python -m pytest tests/test_parity.py tests/test_question_ids.py -q
8 changes: 8 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -8,3 +8,11 @@ __pycache__/
.venv/
venv/
.env
build/
dist/
ts/node_modules/
ts/dist/
ts/*.tgz
.coverage
# Local agent scratch output; never part of a release.
Claude outputs/
21 changes: 21 additions & 0 deletions LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) 2026 Coding-Dev-Tools and contributors

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
9 changes: 9 additions & 0 deletions MANIFEST.in
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
include LICENSE README.md
recursive-include docs *.md *.json
recursive-include examples *.py *.ps1 *.sh *.json *.log
recursive-include scripts *.py *.ps1
recursive-include tests *.py
include ts/package.json ts/package-lock.json ts/tsconfig.json ts/README.md ts/LICENSE
recursive-include ts/src *.ts
recursive-include ts/test *.cjs *.json
recursive-include ts/scripts *.cjs
211 changes: 113 additions & 98 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,135 +1,150 @@
# jev-decision
# Jev Decision

[![CI](https://github.com/Coding-Dev-Tools/jev-decision/actions/workflows/ci.yml/badge.svg)](https://github.com/Coding-Dev-Tools/jev-decision/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](https://opensource.org/licenses/MIT)
[![Python: >=3.9](https://img.shields.io/badge/python-3.9+-blue.svg)](https://www.python.org/)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](LICENSE)

Zero-dependency System 1 decision engine, calibrated guardrails, MCP server, and token optimization client for **Jev (TypeSafe AI)**.
Fast, typed, budgeted [TypeSafe Jev](https://docs.typesafe.ai/api) decisions for coding-agent harnesses: Command Code, Claude Code, Codex, Cursor, Gemini CLI, Antigravity, OpenCode, and any MCP or shell-capable client.

Based on the architecture by Diogo Almeida (@CompleteSkeptic) and validated against empirical benchmarks in *arXiv:2609.29429*.
Jev is a "System 1" model: it returns typed probabilities for yes/no (Noul), multiple-choice (Choice), and rubric (Score) questions. This repository provides a portable client, a shared Python budget ledger, and optional harness and memory-system recipes. Quality, latency and net savings depend on the workload and need separate measurement.

---
**Advice, never authority.** Jev results never grant a permission, approve a command, or certify that a task is complete. Your harness's permission rules and your executed tests stay in charge. Any failure (no key, budget reached, timeout, provider error) returns an explicit `unavailable` result, and your workflow carries on as if Jev were absent.

## Key Features
## Quickstart

- **Zero External Dependencies**: Pure standard library (`urllib.request`, `json`, `math`). Runs anywhere with zero pip bloat.
- **Universal Multi-Harness Support**:
- **Python**: Drop-in client + guardrail helpers.
- **Model Context Protocol (MCP)**: Native stdio MCP server for Cursor, Claude Desktop, Antigravity, Windsurf, Cline.
- **TypeScript / Node.js**: Zero-dependency TS package under `ts/`.
- **CLI**: Fast terminal inspection command (`jev guard`, `jev prune`, `jev verify`).
- **Disruptive Token & Latency Economics**:
- 70–300ms single forward-pass latency.
- $0.042 per million input tokens, **$0 output tokens**.
- Upstream context pruning saving **80%–92%** of ongoing LLM prompt tokens.
- **Calibrated Probabilities (RLCD)**:
- Strict high-stakes threshold ($p \ge 0.95$ for destructive commands).
- Balanced medium-stakes threshold ($p \ge 0.85$ for task verification).
- Permissive low-stakes threshold ($p \ge 0.40$ for context pruning).
- **Built-in Deterministic Offline Fallback**: Operates 100% offline with zero network calls when no API key is provided.
### 1. Install

---
MCP support needs Python 3.10+; the core library and CLI run on 3.9+.

## Installation
The v0.3 changes are currently reviewed in [PR #1](https://github.com/Coding-Dev-Tools/jev-decision/pull/1). These commands install that development branch; they do not install a published v0.3 release.

### Python
```bash
pip install git+https://github.com/Coding-Dev-Tools/jev-decision.git
```
*(Or clone locally and run `pip install -e .`)*
```sh
# Recommended: an isolated tool install that puts `jev` and `jev-mcp` on PATH
uv tool install "jev-decision[mcp,setup] @ git+https://github.com/Coding-Dev-Tools/jev-decision@codex/jev-harness-integration"

### TypeScript / Node.js
```bash
cd ts && npm install && npm run build
# Or into a virtual environment you manage
python -m venv .venv
# Activate on macOS/Linux: . .venv/bin/activate
# Activate on PowerShell: .\.venv\Scripts\Activate.ps1
python -m pip install "jev-decision[mcp,setup] @ git+https://github.com/Coding-Dev-Tools/jev-decision@codex/jev-harness-integration"
```

---
`setup` adds OS credential storage and timezone data. Generated harness entries point at this exact interpreter, so keep the environment in place after installing.

## MCP Server Setup (Cursor / Claude Desktop / Antigravity)
### 2. Configure a key and a daily budget

Add to your `claude_desktop_config.json`, `antigravity.json`, or `.cursor/mcp.json`:
A fresh install makes **no** provider calls until you run setup.

```json
{
"mcpServers": {
"jev-decision": {
"command": "jev-mcp",
"env": {
"TYPESAFE_API_KEY": "your-api-key"
}
}
}
}
```sh
jev setup # interactive: key source, daily budget, workspace, harness
jev doctor --live # one tiny budgeted request that checks authentication
```

Exposes four high-speed tools to your agent:
1. `jev_guard_command(command, cwd)`: Evaluates shell command safety in ~100ms.
2. `jev_prune_output(raw_output, current_goal)`: Prunes verbose boilerplate from command outputs.
3. `jev_verify_completion(goal, recent_actions, last_output)`: Verifies test proof before task completion.
4. `jev_decide(state, questions)`: Arbitrary parallel evaluations.
`jev setup` stores the key in Windows DPAPI or the macOS/Linux keychain. It never writes the key to a config file. On headless machines, point setup at an environment variable *name* instead:

```sh
export TYPESAFE_API_KEY=... # set this in the harness's launch environment
jev setup --non-interactive --credential-source env --daily-budget 1 --timezone UTC --workspace "$PWD"
```

A daily budget of `0` keeps requests disabled. Before each request the shared ledger reserves the worst-case cost of that request, so the cap holds across every harness that uses the same runtime.

### 3. Connect your harness

---
Preview the change, apply it, then reload the client:

## CLI Usage
```sh
jev harness install --target command-code --dry-run
jev harness install --target command-code --apply
```

| Harness | `--target` | What gets installed |
| --- | --- | --- |
| Command Code | `command-code` | `jev` MCP server plus the `/jev-advice` skill ([guide](docs/COMMAND_CODE.md)) |
| Claude Code | `claude-code` | MCP server and skill |
| Codex | `codex` | `[mcp_servers.jev]` and skill |
| Cursor | `cursor` | MCP server and skill |
| Gemini CLI | `gemini-cli` | MCP server and skill |
| Antigravity CLI / IDE | `antigravity`, `antigravity-ide` | MCP server and skill |
| OpenCode | `opencode` | MCP server and skill |
| Claude Desktop, Crush | `claude-desktop`, `crush` | MCP server |
| Pi, Hermes, OMP, OpenClaude, Copilot | `pi`, `hermes`, `omp`, `openclaude`, `copilot` | CLI-based skill |

Add `--scope project --project-root /abs/project` for a project-level configuration where the harness supports one. `jev harness restore --target NAME --apply` removes only what Jev added and reports any entries you changed yourself. File locations and verification status are in the [integration matrix](docs/INTEGRATIONS.md).

```bash
# Evaluate bash command safety
jev guard "git status"
# [ALLOWED (Auto-Execute)] Category: read_only | Safety: 0.98
### 4. Optional: guard shell commands before they run

jev guard "rm -rf / --no-preserve-root"
# [BLOCKED (Requires Approval)] Category: destructive_or_leak | Safety: 0.01
The optional pre-tool hook can assess a shell command before it runs:

# Token-prune large build or test logs
pytest | jev prune --goal "fix authentication bug" --stats
```sh
jev hook config claude-code # prints the exact JSON to merge into your settings
```

---
The hook is **escalate-only**. It can force the harness's normal approval prompt (Claude Code, Cursor), or block a flagged command with a reason (Command Code, Codex, Gemini CLI). By default it blocks only in sessions that run without approval prompts. It never approves anything. Simple read-only commands such as `ls` or `git status` skip the request. Any local failure leaves the harness unchanged, and `JEV_HOOK=off` turns the hook off. See [docs/HOOKS.md](docs/HOOKS.md).

## Python API Usage
## Use it from code

```python
from jev_decision import JevClient, guard_bash_command, prune_tool_output, verify_turn_completion

client = JevClient()

# 1. Shell Safety Guard
safety = guard_bash_command("git status", cwd="/repo", client=client)
if safety["allow_auto"]:
# Execute immediately without prompting user
pass

# 2. Context Pruning
pruned, stats = prune_tool_output(huge_log, current_goal="fix auth endpoint", client=client)
print(f"Omitted {stats['saved_lines']} lines of boilerplate tokens!")

# 3. Task Completion Verification
check = verify_turn_completion(
goal="Fix issue #42",
recent_actions="edited file.py and ran pytest",
last_output="100% green, 45 passed in 0.2s",
client=client
from jev_decision import JevClient

client = JevClient() # uses the saved `jev setup` runtime; fresh defaults stay offline
batch = client.evaluate(
{"request": "Export needs a preview before downloading."},
[{"id": "intent", "type": "choice", "instructions": "Classify the requested change.",
"criteria": {"feature": "Adds new behavior", "bug": "Fixes broken existing behavior", "unclear": None}}],
)
if check["is_complete"]:
# Safe to conclude session
pass
if batch.status == "ok":
print(batch.get_choice("intent").selected, batch.get_choice("intent").probabilities)
else:
print(batch.status, batch.error_code) # carry on without Jev
```

---
The native ID-keyed map (`{"intent": {"type": "choice", ...}}`) and the typed `NoulQuestion`/`ChoiceQuestion`/`ScoreQuestion` classes also work. Ready-made helpers include `guard_bash_command`, `verify_turn_completion`, `assess_memory_relation`, `assess_memory_relevance`, and `prune_tool_output`. Constructing `JevClient(api_key=...)` before any saved configuration explicitly opts in with the default $1/day cap on the shared ledger. Ambient environment keys do not enable fresh default clients or hooks; saved configuration always wins.

## Agent Skill
The [memory-system guide](docs/MEMORY_SYSTEMS.md) and [native JSON recipe](examples/memory-advice.json) cover scoped memory assessments and the optional Engraphis question bridge. Memory writes and evidence omission remain governed by the caller.

To register the Jev skill with Claude Code or Antigravity:
```bash
# Claude Code:
npx skills add Coding-Dev-Tools/jev-decision
From a shell or any harness without MCP, run `jev decide --file request.json` (see [examples](examples/)). TypeScript users have an explicit-key client in [`ts/`](ts/README.md). It uses the same wire contract, but the Python budget ledger does not cover its calls.

# Antigravity / Gemini:
cp SKILL.md ~/.gemini/antigravity/skills/jev-decision/SKILL.md
```
### MCP tools

---
| Tool | Purpose |
| --- | --- |
| `jev_decide` | Batch of typed Noul/Choice/Score questions over one state |
| `jev_guard_command` | Risk category and probability for a shell command (advisory) |
| `jev_verify_completion` | Gaps between a goal and the supplied verification evidence |
| `jev_read_evidence` | Read a saved log page by page, with source hashes and line references |
| `jev_prune_output` | Relevance measurement for text already in context |
| `jev_status` | Local configuration and budget, with no network call |

## License
### Getting good answers

Follow the provider's [Jev 1.13 guidance](https://docs.typesafe.ai/model-jaggedness/jev-1.13):

- **Batch** related questions into one call. Extra questions add little latency.
- **Describe every option.** Give Choice labels and Score levels explicit meanings and boundary conditions.
- **Send only the relevant state.** Unrelated text acts as a distractor.
- **Keep deterministic work in code:** arithmetic, counting, date comparisons, parsing, and exit codes.
- **Calibrate thresholds for each question** on your own data. Don't reuse one question's threshold for another.

## Evidence capture and selection

A large log only saves context tokens if it never enters the model's context. `jev capture --directory /abs/new-dir -- pytest` runs the command and saves stdout, stderr, hashes, and the exit status. It prints only a small reference. `jev evidence` / `jev_read_evidence` then pages through the saved file inside approved workspace roots, with secrets redacted and original line numbers preserved.

Evidence **selection** (dropping low-relevance log spans) is off by default. **No workload ships qualified for automatic omission.** `shadow` mode measures, and `select` requires a locally qualified profile built with the [evaluation workflow](docs/EVALUATION.md). See [docs/EVIDENCE.md](docs/EVIDENCE.md). Savings depend on your logs, model, and harness, so this project makes no universal savings claim.

## Safety and accounting

- The model (`jev-1.13.0`) and the official HTTPS endpoint are pinned. There is no proxy discovery and no redirect following.
- Each attempt reserves its worst-case cost in a SQLite ledger shared by every process with the same `JEV_HOME`. There is at most one transient retry within a single deadline.
- Keys live in DPAPI, the OS keychain, or an environment variable you name. Harness configs contain only variable references. Unexpanded `${NAME}` placeholders are treated as missing keys.
- Question IDs, labels, and state are scanned for recognizable secrets and redacted before they leave the machine.

## Development

```sh
python -m pip install -e ".[test,mcp,setup]" build
python -m pytest -q
python scripts/check_packages.py --output /abs/new-dir # wheel/sdist installed outside the checkout
cd ts && npm ci && npm test && npm run check:package
```

MIT © Coding-Dev-Tools
CI runs Python 3.9–3.14 on Windows, macOS and Linux, plus Node 22 and 24. [Migration from 0.2](docs/MIGRATION_0_3.md) · [Runtime contract](docs/SPECIFICATION.md) · [Validation records](docs/validation/README.md) · [MIT license](LICENSE)
Loading
Loading