Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
42 commits
Select commit Hold shift + click to select a range
ec28cf8
Add packages/semantic-predicate/package.json
halqme Sep 18, 2026
c7964f2
Add packages/semantic-predicate/src/index.ts
halqme Sep 18, 2026
47617fb
Add packages/semantic-predicate/src/index.test.ts
halqme Sep 18, 2026
86ebf83
Add extensions/semantic-observer/package.json
halqme Sep 18, 2026
2b6ab71
Add extensions/semantic-observer/index.ts
halqme Sep 18, 2026
5f79ba0
Keep semantic observer outside workspace registration
halqme Sep 18, 2026
9ac671c
Decouple semantic observer from workspace packaging
halqme Sep 18, 2026
906143d
Keep semantic predicate package standalone
halqme Sep 18, 2026
5eeee6b
Register semantic observer extension
halqme Sep 18, 2026
9b356d5
Document experimental semantic package boundary
halqme Sep 18, 2026
ed6e030
feat: Add semantic observer package
halqme Sep 18, 2026
140336c
style(background-process): wrap option description line
halqme Sep 18, 2026
433576a
Use OpenRouter Decisions API for semantic predicates
halqme Sep 18, 2026
26a6bda
Test native Jev decision response shapes
halqme Sep 18, 2026
9ee52eb
Expose semantic predicate workspace package
halqme Sep 18, 2026
dd5f939
Depend on semantic predicate workspace package
halqme Sep 18, 2026
fd0854e
Make semantic observer context explicit and minimal
halqme Sep 18, 2026
5a7d909
Keep semantic observer package dependency-free from workspace contract
halqme Sep 18, 2026
3364394
Keep semantic predicate boundary as relative adapter import
halqme Sep 18, 2026
0845115
Use Node test runner for semantic predicate package
halqme Sep 18, 2026
fe077ad
Use Node test assertions for semantic predicate
halqme Sep 18, 2026
3d00e97
Fix semantic predicate tests for Node runner
halqme Sep 18, 2026
a9f35d9
Document semantic observation context policy
halqme Sep 18, 2026
ee9cb7a
Define semantic observer evidence boundary
halqme Sep 18, 2026
3c4317f
Fix semantic predicate package manifest newline
halqme Sep 18, 2026
557b5b5
Keep semantic observation tuples typed
halqme Sep 18, 2026
70510de
Normalize semantic predicate package manifest
halqme Sep 18, 2026
a21a840
Skip empty semantic observer test workspace
halqme Sep 18, 2026
1ede9f5
Add semantic observer TypeScript config
halqme Sep 18, 2026
8db62e9
Add semantic predicate TypeScript config
halqme Sep 18, 2026
0aa9df5
Build optional semantic response fields explicitly
halqme Sep 18, 2026
be08fd7
Expose resource tool calls in task evidence
halqme Sep 18, 2026
8cd1414
Expose side-effect-free task evidence packet
halqme Sep 18, 2026
4b6a41e
Reuse task evidence packet for review context
halqme Sep 18, 2026
31a9237
Build semantic state from Pi runtime evidence
halqme Sep 18, 2026
891857a
Derive semantic observations from Pi runtime state
halqme Sep 18, 2026
d9220fa
Test semantic runtime evidence projections
halqme Sep 18, 2026
6592936
Run semantic observer evidence tests
halqme Sep 18, 2026
acc24cc
Typecheck semantic observer evidence module
halqme Sep 18, 2026
5a51235
Document runtime-derived semantic evidence
halqme Sep 18, 2026
703aa18
Describe Pi runtime evidence projections for Jev
halqme Sep 18, 2026
a93e7d9
Respect exact optional semantic diff typing
halqme Sep 18, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 21 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,18 +35,37 @@ extensions/
code/
syntax/
session-metrics/
semantic-observer/
task/
terminal/
packages/
semantic-predicate/
skills/
prompts/
docs/
tsconfig.json
```

Every runtime workspace now lives under `extensions/`; there is no separate `packages/` layer. `session-metrics` owns both the Pi extension and its offline CLI/analysis kernel. Multi-word extension directories use kebab-case, and the shared TypeScript configuration lives at the repository root.
`extensions/` remains the home of Pi runtime integration. `packages/` is reserved for code that is meaningful without Pi; both are root workspaces. The experimental `semantic-predicate` package lives under `packages/` so Jev/OpenRouter evaluation can be removed or reused without changing Pi runtime contracts. `session-metrics` continues to own both its Pi extension and offline CLI/analysis kernel. Multi-word extension directories use kebab-case, and the shared TypeScript configuration lives at the repository root.

The repository extension exposes `context` and `code`, and transparently strengthens the built-in `edit` path for supported source files. The old standalone Astrolabe and BM25 tool surfaces are gone; their useful structural and lexical mechanisms are internal implementation details under `src/syntax` and `src/context`.

Additional independent utilities remain available through the extensions listed above. `ask` provides synchronous structured user decisions in the interactive TUI; offline session analysis is provided by the `session-metrics` CLI in `extensions/session-metrics`.
Additional independent utilities remain available through the extensions listed above. `ask` provides synchronous structured user decisions in the interactive TUI; offline session analysis is provided by the `session-metrics` CLI in `extensions/session-metrics`. `semantic-observer` is an experimental, explicitly invoked observer: the caller selects semantic judgments, while Pi Kit builds compact evidence from task state, tracked reads/context, actual workspace changes, and executed verification. It returns advisory probabilities without changing task, verification, or completion state.


## Experimental semantic observation

`semantic-observer` treats Jev as a sensor, not an authority. It uses OpenRouter's Decisions API through the Pi-independent `packages/semantic-predicate` package and keeps thresholds or actions outside the model boundary.

Context is assembled as evidence, not as a transcript:

- Prefer runtime-captured primary evidence: the task goal and acceptance criteria, tracked mutations and the task-baseline Git diff, executed verification, and the successful `read`/`context` results the agent actually observed.
- Keep fields named and structured. Questions refer to the state fields they judge rather than relying on one opaque prompt.
- Give each judgment only the fields it needs. Questions that need different evidence are evaluated against separate minimal states; questions with the same state may be batched.
- Keep deterministic facts in code. Jev is for semantic judgments such as scope drift or whether verification meaningfully covers a change, not whether a check exists or how many files changed.
- Preserve probabilities. The observer does not turn Jev output into a pass/fail result; later policy may choose thresholds after the behavior has been measured.
- Do not feed broad session history, repository dumps, or previous Jev outputs back into later state by default. Add context only when it is evidence for the next judgment.

The current observer accepts only an `observations` list (`scopeDrift`, `verificationGap`, `consistencyRisk`). Evidence payloads are not authored by the calling model. It is deliberately explicit-call and advisory while the experiment is being evaluated.

See [`docs/architecture.md`](docs/architecture.md) for the design rationale and runtime contracts.
23 changes: 23 additions & 0 deletions bun.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

39 changes: 39 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,45 @@ This follows the centralized asynchronous isolated delegation pattern evaluated

Stable behavior belongs in tools and runtime state. `AGENTS.md` therefore contains only repository invariants and development mechanics; tool-routing and workflow state are not encoded as an always-on prompt layer. This is consistent with the repository-context results in arXiv:2602.11988.


## Experimental semantic observation

`semantic-observer` is outside the mechanical authority path. Its outputs are observations only: they cannot mutate repository state, satisfy `verify`, or unlock `task.finish`. The Pi-facing extension adapts runtime evidence; `packages/semantic-predicate` owns the Pi-independent OpenRouter Decisions API client and typed Jev primitives.

The context boundary is intentionally narrower than the model context window. Jev 1.13 degrades when state contains irrelevant detail, so the observer does not treat the current conversation or repository as a default context blob. Each semantic judgment declares the evidence it needs and receives a small structured state with named fields. Primary runtime or repository evidence is preferred over a model-authored narrative summary.

The observer does not ask the calling model to summarize its own work. It projects existing Pi Kit runtime state instead. `task/evidence.ts` exposes the same side-effect-free packet used by `task.review_context`: task contract and latest checkpoint, resource provenance, workspace delta, and verification evidence. The semantic observer augments that packet with a bounded Git diff from the task's captured baseline and bounded excerpts from successful `read`/`context` tool results identified by their tracked tool-call IDs.

The current context views are:

```text
scopeDrift
task goal + acceptance
+ current checkpoint (plan marked as hypothesis)
+ tracked mutations + task-baseline diff

verificationGap
task goal + acceptance
+ tracked mutations + task-baseline diff
+ executed verification only

consistencyRisk
tracked mutations + task-baseline diff
+ paths observed during the task
+ bounded excerpts from the exact read/context results already seen
```

The caller supplies only which observation IDs to run. These are separate requests because their evidence sets differ. If future questions genuinely share the same state, they should be batched into one Decisions API request; Jev evaluates questions independently and batching avoids sending the same state repeatedly.

This boundary follows four rules:

1. **Filter before inference.** Retrieval and runtime state select evidence before Jev sees it.
2. **Semantic only.** Exact checks, counts, dates, presence tests, and arithmetic stay in code.
3. **Probabilities before policy.** Raw Noul probabilities or Choice/Score distributions are recorded first; thresholds and actions belong to deterministic policy outside the package.
4. **No ambient accumulation.** Session history, broad diffs, repository dumps, and prior semantic answers are not automatically carried forward. A second-stage request receives earlier output only when code needs that result to construct genuinely new state.

The experiment is intentionally explicit-call. Automatic hooks, escalation, or review routing should be added only after session evidence shows which judgments are useful and how their probabilities calibrate on Pi Kit work.

## Evaluation

`session-metrics` reconstructs runtime behavior from Pi session JSONL without active instrumentation. In addition to generic tool/action metrics, it records the `context`, `code`, `task`, `delegate`, and `verify` surfaces and verification provenance so harness changes can be compared against historical trajectories.
Expand Down
3 changes: 2 additions & 1 deletion extensions/background-process/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -130,7 +130,8 @@ export default function backgroundProcessExtension(pi: ExtensionAPI): void {
inspectRunning: Type.Optional(
Type.Boolean({
default: false,
description: "For check only: include stdout/stderr while a process is pending or running.",
description:
"For check only: include stdout/stderr while a process is pending or running.",
}),
),
}),
Expand Down
122 changes: 122 additions & 0 deletions extensions/semantic-observer/evidence.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,122 @@
import assert from "node:assert/strict";
import test from "node:test";

import type { TaskEvidencePacket } from "../task/evidence.ts";
import { projectObservationState } from "./evidence.ts";

const packet: TaskEvidencePacket = {
task: {
id: "task-1",
goal: "Update the parser only",
acceptance: ["Parser accepts the new syntax"],
status: "active",
latestCheckpoint: {
at: "2026-09-18T00:00:00.000Z",
summary: "Parser implementation is in progress",
plan: ["Rewrite unrelated renderer", "Update parser"],
completed: ["Located parser"],
},
},
resources: {
observed: ["src/parser.ts", "src/parser.test.ts"],
mutated: ["src/parser.ts"],
changedDuringTask: ["src/parser.ts"],
preexistingDirty: [],
timeline: [
{
operation: "observe",
path: "src/parser.ts",
tool: "context",
action: "inspect",
toolCallId: "context-1",
},
{
operation: "mutate",
path: "src/parser.ts",
tool: "code",
action: "edit",
toolCallId: "code-1",
},
],
coverage: {
observations: "explicit-tools",
mutations: "explicit-tools",
workspaceDelta: "git",
opaqueToolEffects: "not-attributed",
},
},
verification: [
{
id: "verify-1",
taskId: "task-1",
provenance: "typecheck",
origin: "executed",
passed: true,
summary: "tsc --noEmit",
at: "2026-09-18T00:01:00.000Z",
},
{
id: "verify-2",
taskId: "task-1",
provenance: "self_review",
origin: "reported",
passed: true,
summary: "looks good",
at: "2026-09-18T00:02:00.000Z",
},
],
workspace: {
baselineHead: "abc",
currentHead: "abc",
currentDirty: ["src/parser.ts"],
changedDuringTask: ["src/parser.ts"],
uncommittedTaskChanges: ["src/parser.ts"],
taskCommitRequired: true,
taskCommitPresent: false,
},
};

test("scope drift state uses task authority and actual changes", () => {
const state = projectObservationState("scopeDrift", packet, {
diff: "@@ parser diff @@",
context: [],
}) as any;

assert.equal(state.task.goal, "Update the parser only");
assert.deepEqual(state.task.acceptance, ["Parser accepts the new syntax"]);
assert.equal(state.task.current_stage.summary, "Parser implementation is in progress");
assert.equal(state.task.current_stage.plan_authority, "hypothesis");
assert.deepEqual(state.changes.changed_paths, ["src/parser.ts"]);
assert.equal(state.changes.diff, "@@ parser diff @@");
});

test("verification gap state includes only executed verification", () => {
const state = projectObservationState("verificationGap", packet, {
context: [],
}) as any;

assert.equal(state.verification.length, 1);
assert.equal(state.verification[0].provenance, "typecheck");
assert.equal(state.verification[0].summary, "tsc --noEmit");
});

test("consistency state reuses observed repository evidence", () => {
const state = projectObservationState("consistencyRisk", packet, {
context: [
{
tool: "context",
paths: ["src/parser.ts"],
text: "export function parse() {}",
},
],
}) as any;

assert.deepEqual(state.repository_evidence.observed_paths, [
"src/parser.ts",
"src/parser.test.ts",
]);
assert.equal(
state.repository_evidence.excerpts[0].text,
"export function parse() {}",
);
});
Loading
Loading