Skip to content

Observability, a human in the loop and the Skillhook Cloud link (0.5.0) - #12

Merged
JOsacky merged 19 commits into
mainfrom
claude/skillhook-saas-plan-3efb17
Sep 29, 2026
Merged

JOsacky merged 19 commits into
mainfrom
claude/skillhook-saas-plan-3efb17

Conversation

@JOsacky

@JOsacky JOsacky commented Sep 29, 2026

Copy link
Copy Markdown
Member

Everything skillhook needs to be observed and driven from outside the machine, released as 0.5.0 (0.4.0 is folded in and gets no separate npm release). It is the machine side of Skillhook Cloud (MeterApp/skillhook-cloud).

0.4.0: local observability and a human in the loop

  • Event bus with SSE streams (GET /events, GET /jobs/<id>/events) and artifact reads.
  • Delivery log: every webhook request is recorded, including the ones that were rejected, skipped, deduplicated or folded into a running job (deliveries list|show, GET /deliveries).
  • Replay of a recorded delivery or an earlier job, and ad-hoc runs of a SKILL.md that is not installed (run --file, POST /skills/test).
  • Task outcomes: response.json / structured output gives each job completed | partial | needs_human | nothing_to_do | failed | unknown.
  • Agent job API and a human in the loop: a per-run MCP server (job_progress, job_ask_human, job_set_outcome, job_note, job_context) and the same through skillhook job …. An answer reaches the waiting run, or resumes the Claude/Codex session (claude -p --resume, codex exec resume) when the run already ended (jobs answer, POST /jobs/<id>/answer).
  • Deep health: MCP servers, plugins, CLI logins and codex doctor (health, GET /health/checks).
  • Runner readiness and failure kinds: runners are checked before a job; failures are classified (auth, usage_limit, rate_limit, …); opt-in fallback and retry.
  • Stats per skill, and live configuration and control (PATCH /config, restart, update, logs).

0.5.0: the Skillhook Cloud link (opt-in)

  • skillhook cloud connect --code XXXX-XXXX [--control] pairs the machine; cloud status, cloud disconnect. Observe mode by default. cloud.allow_commands / cloud.deny_commands narrow what the cloud may do, and the cloud can never change host, port, trust_proxy, runners, env_passthrough, projects or cloud.*.
  • An outbound long-poll only (the cloud never connects to the machine), with an outbox that survives restarts. Headers are redacted and every .env value is scrubbed before anything leaves. SKILLHOOK_NO_CLOUD=1 is a kill switch.
  • Control commands: runs, tests, replays, answers, config patches within bounds, restart, update. Secrets are generated sealed (X25519 + AES-GCM) to whoever asked. Live output streaming, chunked artifact uploads, and hosted-URL deliveries that go through the normal webhook pipeline, so signatures are still checked with the machine's own secret.
  • The runners are checked as soon as the link connects, so the cloud knows at once whether claude and codex are signed in.
  • MCP cloud_status and cloud_disconnect. There is deliberately no cloud_connect: pairing hands the machine to an account, so a person runs it.
  • AGENTS.md's outbound-requests rule now lists the link as the second opt-in outbound connection; docs/cloud.md and docs/cloud-protocol.md say what is sent and what never leaves.

Landing this

Merging bumps package.json to 0.5.0, so the Release workflow tags v0.5.0 and Publish ships @meterapp/skillhook@0.5.0 to npm.

Tests

  • npm run check: 41 files, 267 tests, build, schemas and release metadata.
  • End to end against the new cloud on a local stack: pairing, a run that asked a person and continued after the answer, a needs_human run resumed by an answer, direct, rejected and hosted-URL deliveries, live output, a config patch applied live, and a secret generated sealed to the browser.

🤖 Generated with Claude Code

JOsacky and others added 18 commits September 28, 2026 14:41
- src/events.ts: typed in-process `Events` bus (job.*, schedule.*,
  skill.changed, server.*) with seq, timestamps and full records
- queue, scheduler, registry and `serve` publish on it; the queue calls
  setMaxListeners(0) so many `?wait=` callers no longer trigger Node's
  MaxListenersExceededWarning
- GET /events and GET /jobs/<id>/events (server-sent events),
  GET /jobs/<id>/artifacts/<name> (raw artifact, ?tail=)
- SseParser + openAdminEventStream client; `jobs logs -f` follows a
  running job through the server when one is running
- docs: api.md, AGENTS.md, CHANGELOG, llms.txt, README

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- src/delivery-log.ts: jobs/.delivery-log/deliveries.jsonl (ring of
  deliveries.max records) + bodies/ for refused deliveries (capped by
  deliveries.body_max_bytes, off with deliveries.store_bodies: false)
- server: every POST|PUT /hooks/<skill> is recorded with its outcome
  (accepted|duplicate|in_flight|skipped|rejected|challenge|error), the
  status and code the sender got, the reason, redacted headers and the
  job; rate-limited and unknown-skill deliveries included; the event
  bus publishes delivery.received
- GET /deliveries (filters, cursor), GET /deliveries/<id>?include=body,
  deliveries in GET /health; GET /jobs pages with after/next_after,
  filters by trigger/since, caps limit at 500, 400 on bad filters;
  a malformed hook name is 404 instead of 500
- CLI `deliveries list|show`, `jobs list --trigger --since --after`;
  MCP list_deliveries, get_delivery, recent_deliveries in status,
  paging in list_jobs
- config deliveries.* (schema regenerated); docs: api, operations,
  security, mcp, README, llms.txt, CHANGELOG

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- src/response.ts: JobOutcome (completed|partial|needs_human|
  nothing_to_do|failed|unknown), JobResponse, the default response
  schema, parsing of what an agent reports, outcome derivation
- skillhook block field `response: { mode: text|file|structured,
  schema? }`; claude adds --json-schema and reads structured_output,
  codex writes response.schema.json and passes --output-schema; the
  answer is stored as response.json; SKILLHOOK_RESPONSE_PATH and
  {{response_path}}; guardrails tell the agent how to report
- queue derives job.outcome at finish (shell exit 0 = completed,
  nothing reported = unknown, any non-succeeded status = failed);
  interrupted/cancelled jobs are failed
- surfaces: outcome/response in the ?wait= response, GET /jobs?outcome=,
  ?include=response and the response artifact, jobs list --outcome
  (new column), jobs show --response, MCP list_jobs outcome filter,
  skillhook_status recent_jobs.outcome
- fixtures emulate --json-schema / --output-schema; docs: skills
  (Reporting the outcome), runners, api, operations, mcp, authoring
  skill, README, llms.txt, CHANGELOG; yaml schema regenerated

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- src/replay.ts: planReplay resolves a delivery-log record (its job's
  event.json/body.bin when it was accepted, the stored body otherwise)
  or a job into a ManualRunInput with trigger `replay`,
  source.method REPLAY, the original sender IP, redacted headers plus
  x-skillhook-replay-of, and replay_of {delivery, job}; no signature
  check (force for a rejected/error delivery), `when` filters unless
  skipped, never de-duplicated (no delivery id, no fingerprint)
- src/manual.ts: buildManualEvent/createManualJob moved out of ops.ts
  (leaf module, re-exported) so the server can create replay jobs
  without an import cycle; ManualRunInput gains sourceIp, replayOf,
  body; jobs carry replay_of
- POST /deliveries/<id>/replay and POST /jobs/<id>/replay; CLI
  `deliveries replay` / `jobs replay` (through the server when one
  runs, in-process otherwise); MCP replay_delivery / replay_job;
  ops.postToServer; guardrails say the run is a replay
- docs: api, operations, skills, runners, security, mcp, README,
  llms.txt, CHANGELOG

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- manual.createAdhocJob validates the document, creates the job with
  trigger `test`, adhoc: true, skill_file and source.method TEST, and
  writes it to jobs/<id>/skill/<name>/SKILL.md; the queue loads
  ad-hoc skills from there (loadAdhocSkill), also after a restart;
  SkillSource gains { type: "adhoc", job }
- POST /skills/test (400 invalid_skill_document), `skillhook run
  --file SKILL.md | --stdin` (with --dry-run), MCP test_skill;
  ops.runJobLocally / runAdhocLocally
- every job records skill_file; manual jobs persist a cwd override so
  `run --cwd` applies to real runs, not only --dry-run
- docs: api, skills (Testing a skill), operations, runners, mcp,
  README, llms.txt, authoring skill, CHANGELOG; new src/queue.test.ts

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every Claude and Codex run gets a per-run MCP server (`skillhook mcp --job`,
injected with `claude --mcp-config` / `codex -c mcp_servers.skillhook_job.*`)
with job_progress, job_ask_human, job_set_outcome, job_note and job_context;
`skillhook job progress|ask|outcome|note|context` ($SKILLHOOK_BIN) is the same
API for shell skills. Both write files in the job directory (progress.jsonl,
progress.json, question.json, answer.json) which the queue watches: they
become the job's progress/question/answer fields, the events job.progress,
job.waiting_human and job.answered, GET /jobs/<id>/progress and the timeline
in `jobs show`. Asking pauses the job's timeout clock (human_wait_seconds).

A person answers with `skillhook jobs answer`, POST /jobs/<id>/answer or the
MCP tool answer_job: live when the agent still waits, otherwise as a new job
with trigger `resume` that continues the session (`claude -p --resume`,
`codex exec resume`) with the answer in a <human_answer> block, linked by
resume_of / resolved_by. `jobs list --waiting`, `GET /jobs?waiting=1` and
`list_jobs {waiting}` show what waits for a person; a run that ends with its
question unanswered counts as needs_human. New block fields agent_api and
human_wait_seconds.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`skillhook health` (GET /health/checks, MCP get_health) is the doctor plus
what the agents depend on, grouped (system, skillhook, runners, tools, skills,
exposure): claude/codex versions and logins, one check per MCP server either
CLI knows (connected, needs authentication, failed), Claude's MCP config
diagnostics and plugins, `codex doctor`, disk space, and per skill the last
run and unset `env:` names. src/tools.ts probes the CLIs with the job
environment (baseRunEnv) and holds the parsers of their output, pinned by
captured real samples; the fakes answer the same subcommands, with state
files under CLAUDE_CONFIG_DIR / CODEX_HOME.

src/health.ts owns runHealth and HealthCache (one report per flavour for
health.cache_seconds, shared between concurrent callers, health.changed on
status changes); doctor.ts is the quick flavour printed flat and gained a
disk check and CLI versions. GET /doctor and GET /health/checks (?deep, ?network,
?refresh) serve the cache; the server makes no outbound request unless asked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Before a job spawns, the queue checks that its runner is installed and
logged in (or has an API key), with the job environment, cached for
health.readiness_cache_seconds: `skillhook runners`, GET /runners, MCP
get_runners, event runners.changed (src/readiness.ts). A runner that is not
ready fails the job at once (failure.kind auth, no process) unless the
skill's new `fallback: { runners: [codex] }` (or defaults.fallback) names a
ready runner, which takes over (runner_requested, runner_reason).

Every failed or timed_out job carries failure {kind, code, retryable,
message} classified from the CLI's output by src/runners/failure.ts (auth,
usage_limit, rate_limit, budget, max_turns, not_found, timeout, crash,
unknown), pinned by the captured real lines; `jobs list --failure`,
GET /jobs?failure= and list_jobs filter by it. fallback.on may add auth,
usage_limit, rate_limit and crash, and the new `retry: {attempts, on,
backoff_seconds}` repeats a run on the same runner; both act only on a run
that failed before the agent produced anything and keep the earlier runs in
attempts[]. The queue's execute is now pre-flight plus an attempt loop.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…kill

`skillhook stats [--since 24h|7d|ISO] [--until ISO] [--skill S]`, GET /stats
and the MCP tool get_stats sum up the job directories and the delivery log:
jobs by status, outcome, trigger, runner and failure kind, success and
completion rates, duration and queue-wait percentiles, cost and tokens
(Claude and Codex usage added up), deliveries by outcome and HTTP status,
and the same per skill. src/stats.ts is pure aggregation (computeStats) plus
collectStats over the store and the log; parseSince reads 24h-style windows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The running server holds one live skillhook.json (ConfigRef): `skillhook
config set`/`unset` tell it to re-read the file (`config reload`,
POST /config/reload, PATCH /config {set, unset}, MCP update_config; a hand
edit is noticed within five seconds), every key but host and port applies in
place at once, those two are reported as pending_restart (GET /config, MCP
get_config), and an invalid change is refused without writing. Event
config.changed. The logger, the job store and the rate limiter follow the
live values; the queue can drain.

POST /control/restart (MCP restart_server) stops a service-run server
gracefully and lets launchd/systemd start it again (409 not_a_service
otherwise); GET /service, GET /logs and POST /update (MCP check_update)
expose the service status, its log and the update check. applyUpdate in
src/update.ts is now shared by `skillhook update`, the route and the tool.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Delivery log and replay, task outcomes, the agent job API with a human in
the loop, deep health, runner readiness with failure kinds and fallback,
stats, live configuration and remote control, and the event bus behind all
of it. CI smoke tests cover the new commands.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`cloud.*` settings (enabled false, url, machine_id, mode observe|control,
allow/deny command lists, upload switches, ingress, intervals, outbox size)
and the wire protocol between a machine and Skillhook Cloud as pure zod
schemas in src/cloud/protocol.ts, exported as @meterapp/skillhook/protocol
for the cloud repo: pairing, sync request/response, event envelopes, commands
with per-type argument schemas and classes, hosted-ingress items and acks,
hints, errors. A test pins the repeated vocabulary to skillhook's own.
src/cloud/config.ts has the URL rules, the kill switch and commandAllowed.
SKILLHOOK_CLOUD_* never reaches a run's environment, even when listed in
env:. No link yet: nothing leaves the machine.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`skillhook cloud connect --code XXXX-XXXX [--control]` pairs the machine
(token to .env as SKILLHOOK_CLOUD_TOKEN, cloud.* to skillhook.json, observe
mode unless --control); `cloud status` and `cloud disconnect` (revokes the
token) complete it. `serve` runs a CloudLink that idles until cloud.enabled
and then keeps one outbound HTTPS long-poll to cloud.url:

- events from the bus, redacted (headers, command lines, environments) and
  scrubbed of every .env value, spooled in jobs/.cloud/outbox.jsonl until
  acknowledged; snapshots and periodic deep health reports;
- read commands through a dispatcher that validates arguments against the
  protocol and applies cloud.mode and the allow/deny lists (control commands
  answer unsupported_command for now), each command run once;
- hosted-ingress deliveries replayed to the local server so signatures are
  checked with the local secret; the delivery record says via: ingress.

Backoff with jitter, degraded after three failures, stop on 401/403, halve
on 413 and give up on an event that never fits, honour 429, wait on 426,
token rotation. /health (admin) reports the link; doctor/health gain a
`cloud link` check. AGENTS.md's outbound-request rule and security.md's
outbound section are rewritten for the opt-in link. Tests run against
src/test-support/fake-cloud.ts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In cloud.mode: control (or when allow-listed) the link now runs the commands
that act on the machine, through the same functions the CLI and admin API
use (src/cloud/control.ts): skill.run and skill.test (source.method CLOUD,
x-skillhook-cloud-user), delivery/job replay, job.cancel, job.answer (live or
resumed), config.patch (never host, port, trust_proxy, runners,
env_passthrough, projects or cloud), schedule.run, update.install,
service.restart once the cloud has acknowledged the answer, skill.put
(validated, never over a repository's hook) and skill.delete (moved to
jobs/.removed-skills/). secret.generate requires recipient_key and returns
the value only sealed (src/cloud/seal.ts: X25519, HKDF-SHA256, AES-256-GCM,
opened by WebCrypto in a browser as tested); secret.set opens a value sealed
to the machine key that pairing now creates (SKILLHOOK_CLOUD_PRIVATE_KEY),
allow-list only.

job.watch streams complete lines of a job's output as transient job.output
events; job.artifact uploads artifacts over 256 KiB in 1 MiB chunks with the
whole file's sha256. docs/cloud.md states plainly that control mode amounts
to shell access and how deny_commands narrows it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cloud_status and cloud_disconnect over MCP, sharing their logic with the CLI
(src/cloud/pair.ts). There is deliberately no cloud_connect: pairing hands
the machine to whichever account issued the code, and control mode amounts
to shell access, so an agent must never be able to do it from a code it read
somewhere; the MCP instructions and the setup skill's new Skillhook Cloud
step say the person runs `skillhook cloud connect` themselves.

The operator MCP server gets its first tests (src/mcp.test.ts) through a
shared in-memory client (src/test-support/mcp-client.ts), covering the tools
added in 0.4.0 and the cloud tools; CI smoke runs `cloud status`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The opt-in Skillhook Cloud link: pairing, the sync loop with its spool,
read and control commands under the machine's own policy, sealed secrets,
live output and artifact uploads, hosted-ingress delivery through the local
webhook pipeline, the wire protocol as @meterapp/skillhook/protocol, and the
cloud_status / cloud_disconnect MCP tools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Nothing checked the runners until a job started, so a freshly paired machine's
dashboard could not say whether claude and codex were installed and signed in.
The link now asks the readiness cache once on connect; first answers become
runners.changed events.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0.5.0 is not published yet, so the connect-time readiness check belongs in its notes rather
than under Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JOsacky
JOsacky enabled auto-merge (squash) September 29, 2026 01:43
Comment thread src/cloud/config.ts Fixed
CodeQL (js/polynomial-redos, high) flagged /\/+$/ on the configured cloud URL: a long run of
slashes makes it backtrack quadratically. A single backwards scan does the same in linear time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@JOsacky
JOsacky merged commit 0f6a910 into main Sep 29, 2026
6 checks passed
@JOsacky
JOsacky deleted the claude/skillhook-saas-plan-3efb17 branch September 29, 2026 01:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants