Skip to content

Add experimental Jev semantic observer boundary - #35

Merged
halqme merged 42 commits into
mainfrom
feat/semantic-observer-jev
Sep 19, 2026
Merged

halqme merged 42 commits into
mainfrom
feat/semantic-observer-jev

Conversation

@halqme

@halqme halqme commented Sep 18, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • add packages/semantic-predicate as a Pi-independent workspace package
  • call Jev through OpenRouter's native Decisions API (POST /api/alpha/decisions)
  • model the native Jev primitives and response shapes: Noul, Choice, and Score
  • add extensions/semantic-observer as a thin, explicitly invoked Pi adapter
  • derive Jev state from Pi Kit runtime evidence instead of caller-authored summaries
  • keep all semantic observations advisory; they cannot mutate task, satisfy verify, or unlock completion

Boundary

task / context / code / edit / verify runtime evidence
                    |
                    v
extensions/semantic-observer
  - selects a predicate-specific evidence projection
  - asks focused semantic questions
                    |
                    v
packages/semantic-predicate
  - no Pi dependency
  - OpenRouter Decisions API
  - typed Jev questions / answers

extensions/* and packages/* are both root workspaces. The evaluator package knows nothing about Pi.

Pi runtime evidence

The caller now supplies only the observation IDs to run:

semantic_observe({
  observations: ["scopeDrift", "verificationGap", "consistencyRisk"]
})

It does not supply request, changes, verification prose, or repository context.

A new side-effect-free task/evidence.ts packet exposes the same runtime facts used by task.review_context without creating a review request:

  • task goal and acceptance criteria
  • latest checkpoint / current stage
  • tracked read/context/code/edit/write resource provenance
  • workspace delta since the task baseline
  • verification evidence and provenance

The semantic observer augments that packet with:

  • a bounded Git diff from the task's captured baseline
  • bounded excerpts from the exact successful read / context tool results already seen in the session, joined back to resource events via toolCallId

The working plan is included only as current-stage context and explicitly marked as a hypothesis, not as authority to expand task scope.

Context projections

scopeDrift
  task goal + acceptance
  + current checkpoint
  + tracked mutations + task-baseline diff

verificationGap
  task goal + acceptance
  + tracked mutations + task-baseline diff
  + executed verification only

consistencyRisk
  tracked mutations + task-baseline diff
  + observed repository paths
  + bounded excerpts from the exact read/context results already observed

The context policy remains:

  1. Filter before inference.
  2. Prefer runtime facts and repository evidence over model-authored summaries.
  3. Give each judgment only the evidence it needs.
  4. Keep deterministic facts in code.
  5. Preserve Jev probabilities; thresholds and actions remain outside the observer.
  6. Do not automatically accumulate broad session history, repository dumps, or prior Jev outputs.

Jev API

  • endpoint: https://openrouter.ai/api/alpha/decisions
  • request: model + state + questions
  • default model pinned to typesafe/jev-1.13
  • Noul returns noul as the yes-probability
  • Choice returns choice + probabilities + confidence
  • Score returns score + legend + probabilities + confidence

Safety / scope

  • explicit call only; no automatic hooks
  • no changes to task, verify, repository, or delegate authority
  • no action is taken from a probability
  • missing OPENROUTER_API_KEY only fails the observer call
  • no caller-authored semantic evidence payload
  • no broad session or repository context is sent by default

Validation

  • semantic-predicate tests cover native response parsing and Decisions API request shape
  • semantic-observer tests cover the three runtime evidence projections
  • task review-context tests continue to exercise the shared evidence packet path
  • root workspace CI runs lint, typecheck, and tests on Linux/macOS

halqme and others added 30 commits September 18, 2026 22:18
@halqme
halqme marked this pull request as ready for review September 19, 2026 11:04
@halqme
halqme merged commit 62b8eee into main Sep 19, 2026
3 checks passed
@halqme
halqme deleted the feat/semantic-observer-jev branch September 19, 2026 11:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant