diff --git a/.claude-plugin/plugin.json b/.claude-plugin/plugin.json index 8c9271f..689d499 100644 --- a/.claude-plugin/plugin.json +++ b/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "hermes-helmet", "displayName": "Hermes Helmet", - "version": "0.5.0", + "version": "0.6.0", "description": "First-officer skills for Hermes Helmet: setup, single-issue delivery, and dependent-issue coordination.", "author": { "name": "Machine Wisdom", diff --git a/.codex-plugin/plugin.json b/.codex-plugin/plugin.json index 81ca114..74ac5e7 100644 --- a/.codex-plugin/plugin.json +++ b/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "hermes-helmet", - "version": "0.5.0", + "version": "0.6.0", "description": "First-officer skills for Hermes Helmet: setup, single-issue delivery, and dependent-issue coordination.", "author": { "name": "Machine Wisdom", diff --git a/CHANGELOG.md b/CHANGELOG.md index 3b2e951..53c6426 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,20 @@ All notable changes to Hermes Helmet are recorded here. The project follows ## Unreleased +### 0.6.0 + +- Captain’s Bridge can prepare an updated walkthrough in the background. The + panel captures the originating chat and a verifiable snapshot (delivery is + rejected if captured records changed or digests are inconsistent), asks the first + officer to dispatch one bounded read-only agent without waiting, and keeps the + current walkthrough visible. A timeout asks the first officer to stop the agent + and reports the stop as requested, not confirmed. Cancel, supersession, timeout, failure, duplicate + submission and late completion preserve the last useful view. Verified with + synthetic protocol and panel tests only; acceptance in an installed Codex host + is recorded separately in the pull request. Measured Codex limit: after the + parent turn finishes, a delivery reaches the server but the originating panel does + not render it, so the panel now says so instead of implying background success. + ### 0.5.0 - Captain’s Bridge Refresh records now rereads the exact bound chat directly diff --git a/docs/first-officer-plugins.md b/docs/first-officer-plugins.md index ff2f441..05fb721 100644 --- a/docs/first-officer-plugins.md +++ b/docs/first-officer-plugins.md @@ -29,8 +29,9 @@ path. Do not delete unrelated host skills. Canonical behavior stays in The Codex plugin also carries Captain’s Bridge: a read-only panel that explains one recorded piece of work in the invoking chat and opens its supporting records. It is Codex-only and starts no second server beyond the -plugin's own MCP server. Background preparation and cancellation are not part of -this release. See the +plugin's own MCP server. From an open Bridge, “Update walkthrough” can ask the +first officer to start one read-only background preparation, cancellable from the +panel; see the lifecycle in the Bridge guide. See the [source and test guide](../mcp/captains-bridge/README.md). The Claude Code port is tracked in [issue #55](https://github.com/MachineWisdomAI/hermes-helmet/issues/55); the diff --git a/mcp/captains-bridge/README.md b/mcp/captains-bridge/README.md index e75299c..a069a8e 100644 --- a/mcp/captains-bridge/README.md +++ b/mcp/captains-bridge/README.md @@ -31,12 +31,71 @@ not load it. background work or project task), keeps the explanation with its original read time and flags it as older when the fingerprint changed. Failures, foreign or out-of-order responses and a partially written final record keep - the last view and report the limit. “Update walkthrough” - sends a request to the first officer. Delegated background preparation and - cancellation are not implemented. + the last view and report the limit. “Update walkthrough” rereads only; see + Background preparation below. - Show Me and Retro request separately installed skills and report if they are unavailable. +## Background preparation lifecycle + +“Update walkthrough” (and its stale notice) runs this lifecycle. Opening the +Bridge never starts it. + +1. Capture: the panel calls `request_walkthrough_update` with its portable view. + The server rereads that exact chat and returns a request: a random + `requestId`, the `threadId`, and a snapshot (source fingerprint, read time, + record count). No server-side state is kept, so nothing is process-local. +2. Dispatch: the panel sends one `ui/message` asking the first officer to start + exactly one read-only subagent, not wait or poll, and not pause project work. + The subagent reads only `threadId` with `read_chat_work`, never its own chat. + It must not run project tasks or repair the viewer. +3. Deliver: the subagent calls `deliver_walkthrough_update` once with the request + unchanged, or with `failure`. Citations are validated against the snapshot’s + records only; the result keeps the snapshot’s fingerprint and read time, so a + later chat change marks it older instead of fresh. The snapshot also carries a + digest of the exact captured records and a seal over every snapshot field. On + delivery the server rereads the chat and rejects the request if a captured + record has changed, if the digests are inconsistent, or if the chat names no + valid chat; records appended later are allowed and only mark the result older. + The seal is a consistency check, not authentication: the server keeps no + secret or state, so it does not stop a caller who can read the chat from + building a consistent snapshot. It also does not preserve the captured text; + it detects that the text changed. +4. Settle: the panel owns the single active request. It accepts a delivery only + when `requestId` matches and the request is still active, and ignores every + other result. Delivery, failure, cancel, supersession or the 10-minute timeout + end the request; each preserves the last useful view, and a second delivery + for a settled request is dropped. The panel never retries. +5. Cancel: “Cancel update” ends the request at once and asks the first officer to + interrupt the subagent. If the host cannot be told, the panel says so and still + discards any later result. The 10-minute timeout does the same: it invalidates + the request at once, then asks the first officer to interrupt the subagent. The + panel says the stop was requested, not confirmed, and says so plainly if the + host could not be told. Each async effect (acknowledgement failure, timeout + handoff) applies only to the request that owns it, so a late failure of a + cancelled request never clears a newer one. A substantive new instruction from the Captain + supersedes the request the same way unless the Captain says to keep it. + +Measured host limit (installed Codex, parent turn already finished): the preparation +agent read the originating chat and `deliver_walkthrough_update` returned a validated +walkthrough with the captured fingerprint, read time and request identity, so delivery +to the server worked. The originating panel did not render it: it kept the old +explanation and its progress notice until its own 10-minute timeout, and interrupting +the agent reported it had already completed. This server keeps no state and cannot push +to a panel, so rendering in the original panel after the parent finishes is unsupported +by the host seam as measured; no background rendering is claimed. The panel therefore +says a result may not be shown if the chat turn has finished, offers Cancel to stop +waiting, and on timeout says a finished agent's result reached the server but cannot be +shown. A failure result is subject to the same boundary. Foreground delivery while the +parent turn is still active is unchanged. No second queue or observer was added. + +Unsupported boundaries, reported rather than hidden: the server cannot itself stop +a subagent or know the Captain issued a new instruction; the first officer does +both from the messages above. Whether a delivery reaches the panel when the parent +turn has already finished, and whether work continues during preparation, depend +on Codex host behavior that synthetic tests cannot establish; they require the +installed-host acceptance recorded in the pull request. + ## Checks Python 3.11 or newer and Node.js. The tests use synthetic records only and do @@ -47,6 +106,7 @@ python3 -B -m unittest discover -s . -p 'test_*.py' -v node test_view.cjs node test_delivery.cjs node test_actions.cjs +node test_preparation.cjs ``` `scripts/verify.sh` runs these. Protocol and fixture checks do not establish diff --git a/mcp/captains-bridge/preparation.py b/mcp/captains-bridge/preparation.py new file mode 100644 index 0000000..4e3f958 --- /dev/null +++ b/mcp/captains-bridge/preparation.py @@ -0,0 +1,98 @@ +"""Background walkthrough preparation: portable request identity and settlement. + +The server keeps no request state. The panel owns the single active request and +discards any settlement that is not for it, so cancellation, supersession, timeout +and late completion need no cross-process coordination. Nothing is written. +""" +import copy +import hashlib +import json +import re +import secrets +from walkthrough import fingerprint, validate + +OUTCOMES = ('delivered', 'failed', 'cancelled', 'superseded') +CHAT = re.compile(r'[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}') + + +def records_digest(records): + """Digest of the exact captured records, in order.""" + return hashlib.sha256(json.dumps(records, sort_keys=True).encode()).hexdigest() + + +def seal(request_id, thread_id, snapshot): + """Bind every snapshot field so an edited or inconsistent field is detectable. + + This is a consistency check, not authentication: the server keeps no secret + and no state, so it detects corruption and inconsistency, not a deliberate + forger who can read the chat. The captured records are verified separately. + """ + fields = [request_id, thread_id, snapshot['fingerprint'], snapshot['readAt'], + snapshot['recordCount'], snapshot['evidence']] + return hashlib.sha256(json.dumps(fields).encode()).hexdigest() + + +def new_request(thread_id, data): + """Capture the originating chat and a verifiable snapshot of what was read.""" + request_id = secrets.token_urlsafe(18) + snapshot = {'fingerprint': fingerprint(data), 'readAt': data['readAt'], + 'recordCount': len(data['records']), 'evidence': records_digest(data['records'])} + snapshot['seal'] = seal(request_id, thread_id, snapshot) + return {'requestId': request_id, 'threadId': thread_id, 'snapshot': snapshot} + + +def check_request(request): + """Return a normalized copy of an untrusted request, or raise ValueError.""" + if not isinstance(request, dict): + raise ValueError('The walkthrough request is missing.') + snapshot = request.get('snapshot') + request_id, thread_id = request.get('requestId'), request.get('threadId') + if not isinstance(request_id, str) or not re.fullmatch(r'[A-Za-z0-9_-]{16,64}', request_id): + raise ValueError('The walkthrough request has no valid identity.') + if not isinstance(thread_id, str) or not CHAT.fullmatch(thread_id): + raise ValueError('The walkthrough request has no valid originating chat.') + if not isinstance(snapshot, dict): + raise ValueError('The walkthrough request has no preparation snapshot.') + digest, read_at, count = snapshot.get('fingerprint'), snapshot.get('readAt'), snapshot.get('recordCount') + evidence, sealed = snapshot.get('evidence'), snapshot.get('seal') + if (not all(isinstance(h, str) and re.fullmatch(r'[0-9a-f]{64}', h) for h in (digest, evidence, sealed)) + or not isinstance(read_at, str) or not 0 < len(read_at) <= 80 + or not isinstance(count, int) or isinstance(count, bool) or count < 1): + raise ValueError('The walkthrough request snapshot is invalid.') + clean = {'fingerprint': digest, 'readAt': read_at, 'recordCount': count, 'evidence': evidence} + if not secrets.compare_digest(sealed, seal(request_id, thread_id, clean)): + raise ValueError('The walkthrough request snapshot is inconsistent. Request a new update.') + clean['seal'] = sealed + return {'requestId': request_id, 'threadId': thread_id, 'snapshot': clean} + + +def prepared_walkthrough(request, walkthrough, data): + """Validate a delivered explanation against its snapshot, not the current read. + + Citations must belong to the records that existed at snapshot time. The result + keeps the snapshot's fingerprint and read time, so records read later flag it + as older instead of presenting it as freshly interpreted. + """ + snapshot = request['snapshot'] + if len(data['records']) < snapshot['recordCount']: + raise ValueError('The chat has fewer records than the preparation snapshot. Request a new update.') + captured = data['records'][:snapshot['recordCount']] + if not secrets.compare_digest(records_digest(captured), snapshot['evidence']): + raise ValueError('The captured records have changed since the preparation snapshot. Request a new update.') + if len(data['records']) == snapshot['recordCount'] and fingerprint(data) != snapshot['fingerprint']: + raise ValueError('The snapshot fingerprint does not match the captured chat. Request a new update.') + bounded = dict(data, records=captured) + checked = validate(copy.deepcopy(walkthrough), bounded) + checked.update(fingerprint=snapshot['fingerprint'], explainedAt=snapshot['readAt']) + return checked + + +def settlement(request, outcome, message=None): + if outcome not in OUTCOMES: + raise ValueError('Unknown settlement outcome.') + value = {'requestId': request['requestId'], 'outcome': outcome} + if message is not None: + if not isinstance(message, str) or len(message) > 500: + raise ValueError('A settlement message must be text, at most 500 characters.') + value['message'] = message + return value diff --git a/mcp/captains-bridge/server.py b/mcp/captains-bridge/server.py index 031c890..d943321 100644 --- a/mcp/captains-bridge/server.py +++ b/mcp/captains-bridge/server.py @@ -6,6 +6,7 @@ from pathlib import Path from chat_reader import read_chat from walkthrough import validate, fingerprint +from preparation import new_request, check_request, prepared_walkthrough, settlement ROOT = Path(__file__).resolve().parent URI = 'ui://hermes-helmet/captains-bridge-v1' @@ -42,6 +43,16 @@ def tool(name, title, description, properties, required=(), app=False, entry=Fal VIEW['_meta'] = {'ui': {'visibility': ['app']}} +REQUEST = tool('request_walkthrough_update', 'Capture walkthrough request', + 'Read the exact chat in the portable view and capture an immutable preparation snapshot and request identity. Starts no work.', + {'view': {'type': 'object'}}, ('view',)) +REQUEST['_meta'] = {'ui': {'visibility': ['app']}} +DELIVER = tool('deliver_walkthrough_update', 'Deliver prepared walkthrough', + 'For a read-only preparation agent: deliver the explanation for one captured request, or report failure. Pass the request unchanged. Cites record IDs from the request chat only. Delivery is discarded by the panel unless the request is still active. Never retry.', + {'request': {'type': 'object'}, 'walkthrough': {'type': 'object'}, + 'failure': {'type': 'string', 'maxLength': 500}}, ('request',), app=True) + + def bind(thread_id): data = read_chat(thread_id) key = secrets.token_urlsafe(24) @@ -93,13 +104,38 @@ def read_view(view): return {'thread_id': view['threadId'], 'data': data, 'walkthrough': account} +def capture_request(view): + state = read_view(view) + request = new_request(state['thread_id'], state['data']) + text = 'Captured walkthrough request ' + request['requestId'] + '. No work was started.' + return {'content': [{'type': 'text', 'text': text}], 'structuredContent': {'request': request}} + + +def deliver(args): + request = check_request(args.get('request')) + failure = args.get('failure') + if failure is not None or args.get('walkthrough') is None: + note = failure if isinstance(failure, str) and failure.strip() else 'The preparation agent returned no explanation.' + outcome = settlement(request, 'failed', note[:500]) + return {'content': [{'type': 'text', 'text': 'Reported failure for request ' + request['requestId'] + '.'}], + 'structuredContent': {'preparation': outcome}} + # The reader opens only the request's own chat; the child never picks one. + data = read_chat(request['threadId']) + account = prepared_walkthrough(request, args['walkthrough'], data) + state = {'thread_id': request['threadId'], 'data': data, 'walkthrough': account} + result = present(None, state) + result['structuredContent']['preparation'] = settlement(request, 'delivered') + result['_meta']['preparation'] = result['structuredContent']['preparation'] + return result + + def handle(method, params): if method == 'initialize': return {'protocolVersion': params.get('protocolVersion', '2025-06-18'), 'capabilities': {'tools': {}, 'resources': {}}, 'serverInfo': {'name': 'hermes-helmet-captains-bridge', 'version': VERSION}} if method == 'ping': return {} if method == 'tools/list': - return {'tools': [OPEN, READ, PRESENT, REFRESH, VIEW]} + return {'tools': [OPEN, READ, PRESENT, REFRESH, VIEW, REQUEST, DELIVER]} if method == 'resources/list': return {'resources': [{'uri': URI, 'name': 'Captain’s Bridge', 'mimeType': 'text/html;profile=mcp-app'}]} if method == 'resources/templates/list': @@ -109,6 +145,10 @@ def handle(method, params): if method != 'tools/call': raise ValueError('Unknown method or resource') args, name = params.get('arguments') or {}, params.get('name') + if name == REQUEST['name']: + return capture_request(args.get('view')) + if name == DELIVER['name']: + return deliver(args) if name in (VIEW['name'], REFRESH['name']): return present(None, read_view(args.get('view')), for_app=True) if name in (OPEN['name'], READ['name']) and args.get('thread_id'): diff --git a/mcp/captains-bridge/test_preparation.cjs b/mcp/captains-bridge/test_preparation.cjs new file mode 100644 index 0000000..bb0f3f1 --- /dev/null +++ b/mcp/captains-bridge/test_preparation.cjs @@ -0,0 +1,154 @@ +const fs=require('fs'),vm=require('vm'),assert=require('node:assert/strict'); +class Element{ + constructor(tag){this.tagName=tag;this.children=[];this.textContent='';this.hidden=false;this.disabled=false;this.style={setProperty(){}}} + append(c){this.children.push(c)} prepend(c){this.children.unshift(c)} replaceChildren(...c){this.children=c} focus(){} +} +const roots=Object.fromEntries(['app','error','connect','connection','request-status'].map(k=>[k,new Element(k)])); +const document={getElementById:id=>roots[id],createElement:t=>new Element(t),documentElement:new Element('html')}; +let listener;const sent=[];const timers=[]; +const parent={postMessage(m){if(m.method)sent.push(m)}}; +const context=vm.createContext({document,parent,window:{scrollY:0,scrollTo(){}}, + setTimeout:(fn,ms)=>{timers.push({fn,ms});return timers.length},clearTimeout:id=>{if(timers[id-1])timers[id-1].fn=null}, + requestAnimationFrame:f=>f(),addEventListener:(t,fn)=>{if(t==='message')listener=fn},console}); +vm.runInContext(fs.readFileSync('view.html','utf8').match(/ diff --git a/pyproject.toml b/pyproject.toml index dabe017..a5466bd 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "hermes-helmet" -version = "0.5.0" +version = "0.6.0" description = "Open-source software factory for coding agents: delegation, review, and repair" readme = "README.md" requires-python = ">=3.11" diff --git a/scripts/verify.sh b/scripts/verify.sh index 0f2d286..f6d99e2 100755 --- a/scripts/verify.sh +++ b/scripts/verify.sh @@ -163,7 +163,7 @@ PYTHONPATH=src "$PYTHON_BIN" -m unittest discover -s tests -v cd mcp/captains-bridge "$PYTHON_BIN" -B -m unittest discover -s . -p 'test_*.py' -v if command -v node >/dev/null 2>&1; then - for check in test_view.cjs test_delivery.cjs test_actions.cjs test_refresh.cjs; do + for check in test_view.cjs test_delivery.cjs test_actions.cjs test_refresh.cjs test_preparation.cjs; do node "$check" done elif [ -n "${CI:-}" ]; then diff --git a/skills/observe-chat/SKILL.md b/skills/observe-chat/SKILL.md index 66185ca..e7c3bc9 100644 --- a/skills/observe-chat/SKILL.md +++ b/skills/observe-chat/SKILL.md @@ -73,6 +73,18 @@ references, retaining consequential failures. Keep async handoffs and later wakeups related to the same work without inventing a parent-child execution tree. Do not assign turn timestamps to individual operations. Gaps are not proof of a stall. +## Background update (preparation agent) + +When the Captain’s panel requests an update, the first officer dispatches one +read-only subagent and continues authorized work without waiting or polling. The +subagent reads only the `threadId` in the request with `read_chat_work`, never its +own chat, runs no project task, does not repair the viewer, and never retries. It +cites only IDs from that read and calls `deliver_walkthrough_update` once with the +request unchanged, or with `failure`. On a cancel message or a substantive new +instruction from the Captain, the first officer interrupts the subagent and +discards its result unless the Captain keeps the request. Present a delivery +briefly at an available boundary; do not hold project delivery open for it. + ## Refresh and limitations The reader opens Codex’s existing local chat index in read-only mode and reads @@ -85,7 +97,7 @@ The prepared result carries an explicit chat reference and explanation so app calls do not depend on the model tool process’s memory. Each app read validates the citations again; it writes no transcript or standalone viewer file. If records changed it labels the explanation as older, showing the record-read time apart from the explanation’s own source time, and never restamps or rewrites it. Refresh calls the read-only app tool directly: it sends no message to the first officer. A failed, timed-out, foreign or out-of-order refresh keeps the last view and reports why. “Update walkthrough” -requests this skill again in this same chat. No automatic monitoring is added. +starts the background update above in this same chat. No automatic monitoring is added. The extension does not directly collect Hermes Docker events: use recorded worker results only when explicitly connected to this chat, and disclose missing coverage at the affected item. A source-bound explanation is not independent diff --git a/src/hermes_helmet/__init__.py b/src/hermes_helmet/__init__.py index 73063cf..2adbd02 100644 --- a/src/hermes_helmet/__init__.py +++ b/src/hermes_helmet/__init__.py @@ -7,7 +7,7 @@ from hermes_helmet.authority import Policy, load_authority, load_policy, render_crew_contract -__version__ = "0.5.0" +__version__ = "0.6.0" __all__ = [ "Policy",