Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"name": "hermes-helmet",
"displayName": "Hermes Helmet",
"version": "0.5.0",
"version": "0.6.0",
"description": "First-officer skills for Hermes Helmet: setup, single-issue delivery, and dependent-issue coordination.",
"author": {
"name": "Machine Wisdom",
Expand Down
2 changes: 1 addition & 1 deletion .codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "hermes-helmet",
"version": "0.5.0",
"version": "0.6.0",
"description": "First-officer skills for Hermes Helmet: setup, single-issue delivery, and dependent-issue coordination.",
"author": {
"name": "Machine Wisdom",
Expand Down
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,18 @@ All notable changes to Hermes Helmet are recorded here. The project follows

## Unreleased

### 0.6.0

- Captain’s Bridge can prepare an updated walkthrough in the background. The
panel captures the originating chat and a verifiable snapshot (delivery is
rejected if captured records changed or digests are inconsistent), asks the first
officer to dispatch one bounded read-only agent without waiting, and keeps the
current walkthrough visible. A timeout asks the first officer to stop the agent
and reports the stop as requested, not confirmed. Cancel, supersession, timeout, failure, duplicate
submission and late completion preserve the last useful view. Verified with
synthetic protocol and panel tests only; acceptance in an installed Codex host
is recorded separately in the pull request.

### 0.5.0

- Captain’s Bridge Refresh records now rereads the exact bound chat directly
Expand Down
5 changes: 3 additions & 2 deletions docs/first-officer-plugins.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,8 +29,9 @@ path. Do not delete unrelated host skills. Canonical behavior stays in
The Codex plugin also carries Captain’s Bridge: a read-only panel that
explains one recorded piece of work in the invoking chat and opens its
supporting records. It is Codex-only and starts no second server beyond the
plugin's own MCP server. Background preparation and cancellation are not part of
this release. See the
plugin's own MCP server. From an open Bridge, “Update walkthrough” can ask the
first officer to start one read-only background preparation, cancellable from the
panel; see the lifecycle in the Bridge guide. See the
[source and test guide](../mcp/captains-bridge/README.md). The Claude Code port
is tracked in
[issue #55](https://github.com/MachineWisdomAI/hermes-helmet/issues/55); the
Expand Down
53 changes: 50 additions & 3 deletions mcp/captains-bridge/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,12 +31,58 @@ not load it.
background work or project task), keeps the explanation with its original
read time and flags it as older when the fingerprint changed. Failures,
foreign or out-of-order responses and a partially written final record keep
the last view and report the limit. “Update walkthrough”
sends a request to the first officer. Delegated background preparation and
cancellation are not implemented.
the last view and report the limit. “Update walkthrough” rereads only; see
Background preparation below.
- Show Me and Retro request separately installed skills and report if they are
unavailable.

## Background preparation lifecycle

“Update walkthrough” (and its stale notice) runs this lifecycle. Opening the
Bridge never starts it.

1. Capture: the panel calls `request_walkthrough_update` with its portable view.
The server rereads that exact chat and returns a request: a random
`requestId`, the `threadId`, and a snapshot (source fingerprint, read time,
record count). No server-side state is kept, so nothing is process-local.
2. Dispatch: the panel sends one `ui/message` asking the first officer to start
exactly one read-only subagent, not wait or poll, and not pause project work.
The subagent reads only `threadId` with `read_chat_work`, never its own chat.
It must not run project tasks or repair the viewer.
3. Deliver: the subagent calls `deliver_walkthrough_update` once with the request
unchanged, or with `failure`. Citations are validated against the snapshot’s
records only; the result keeps the snapshot’s fingerprint and read time, so a
later chat change marks it older instead of fresh. The snapshot also carries a
digest of the exact captured records and a seal over every snapshot field. On
delivery the server rereads the chat and rejects the request if a captured
record has changed, if the digests are inconsistent, or if the chat names no
valid chat; records appended later are allowed and only mark the result older.
The seal is a consistency check, not authentication: the server keeps no
secret or state, so it does not stop a caller who can read the chat from
building a consistent snapshot. It also does not preserve the captured text;
it detects that the text changed.
4. Settle: the panel owns the single active request. It accepts a delivery only
when `requestId` matches and the request is still active, and ignores every
other result. Delivery, failure, cancel, supersession or the 10-minute timeout
end the request; each preserves the last useful view, and a second delivery
for a settled request is dropped. The panel never retries.
5. Cancel: “Cancel update” ends the request at once and asks the first officer to
interrupt the subagent. If the host cannot be told, the panel says so and still
discards any later result. The 10-minute timeout does the same: it invalidates
the request at once, then asks the first officer to interrupt the subagent. The
panel says the stop was requested, not confirmed, and says so plainly if the
host could not be told. Each async effect (acknowledgement failure, timeout
handoff) applies only to the request that owns it, so a late failure of a
cancelled request never clears a newer one. A substantive new instruction from the Captain
supersedes the request the same way unless the Captain says to keep it.

Unsupported boundaries, reported rather than hidden: the server cannot itself stop
a subagent or know the Captain issued a new instruction; the first officer does
both from the messages above. Whether a delivery reaches the panel when the parent
turn has already finished, and whether work continues during preparation, depend
on Codex host behavior that synthetic tests cannot establish; they require the
installed-host acceptance recorded in the pull request.

## Checks

Python 3.11 or newer and Node.js. The tests use synthetic records only and do
Expand All @@ -47,6 +93,7 @@ python3 -B -m unittest discover -s . -p 'test_*.py' -v
node test_view.cjs
node test_delivery.cjs
node test_actions.cjs
node test_preparation.cjs
```

`scripts/verify.sh` runs these. Protocol and fixture checks do not establish
Expand Down
98 changes: 98 additions & 0 deletions mcp/captains-bridge/preparation.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
"""Background walkthrough preparation: portable request identity and settlement.

The server keeps no request state. The panel owns the single active request and
discards any settlement that is not for it, so cancellation, supersession, timeout
and late completion need no cross-process coordination. Nothing is written.
"""
import copy
import hashlib
import json
import re
import secrets
from walkthrough import fingerprint, validate

OUTCOMES = ('delivered', 'failed', 'cancelled', 'superseded')
CHAT = re.compile(r'[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}')


def records_digest(records):
"""Digest of the exact captured records, in order."""
return hashlib.sha256(json.dumps(records, sort_keys=True).encode()).hexdigest()


def seal(request_id, thread_id, snapshot):
"""Bind every snapshot field so an edited or inconsistent field is detectable.

This is a consistency check, not authentication: the server keeps no secret
and no state, so it detects corruption and inconsistency, not a deliberate
forger who can read the chat. The captured records are verified separately.
"""
fields = [request_id, thread_id, snapshot['fingerprint'], snapshot['readAt'],
snapshot['recordCount'], snapshot['evidence']]
return hashlib.sha256(json.dumps(fields).encode()).hexdigest()


def new_request(thread_id, data):
"""Capture the originating chat and a verifiable snapshot of what was read."""
request_id = secrets.token_urlsafe(18)
snapshot = {'fingerprint': fingerprint(data), 'readAt': data['readAt'],
'recordCount': len(data['records']), 'evidence': records_digest(data['records'])}
snapshot['seal'] = seal(request_id, thread_id, snapshot)
return {'requestId': request_id, 'threadId': thread_id, 'snapshot': snapshot}


def check_request(request):
"""Return a normalized copy of an untrusted request, or raise ValueError."""
if not isinstance(request, dict):
raise ValueError('The walkthrough request is missing.')
snapshot = request.get('snapshot')
request_id, thread_id = request.get('requestId'), request.get('threadId')
if not isinstance(request_id, str) or not re.fullmatch(r'[A-Za-z0-9_-]{16,64}', request_id):
raise ValueError('The walkthrough request has no valid identity.')
if not isinstance(thread_id, str) or not CHAT.fullmatch(thread_id):
raise ValueError('The walkthrough request has no valid originating chat.')
if not isinstance(snapshot, dict):
raise ValueError('The walkthrough request has no preparation snapshot.')
digest, read_at, count = snapshot.get('fingerprint'), snapshot.get('readAt'), snapshot.get('recordCount')
evidence, sealed = snapshot.get('evidence'), snapshot.get('seal')
if (not all(isinstance(h, str) and re.fullmatch(r'[0-9a-f]{64}', h) for h in (digest, evidence, sealed))
or not isinstance(read_at, str) or not 0 < len(read_at) <= 80
or not isinstance(count, int) or isinstance(count, bool) or count < 1):
raise ValueError('The walkthrough request snapshot is invalid.')
clean = {'fingerprint': digest, 'readAt': read_at, 'recordCount': count, 'evidence': evidence}
if not secrets.compare_digest(sealed, seal(request_id, thread_id, clean)):
raise ValueError('The walkthrough request snapshot is inconsistent. Request a new update.')
clean['seal'] = sealed
return {'requestId': request_id, 'threadId': thread_id, 'snapshot': clean}


def prepared_walkthrough(request, walkthrough, data):
"""Validate a delivered explanation against its snapshot, not the current read.

Citations must belong to the records that existed at snapshot time. The result
keeps the snapshot's fingerprint and read time, so records read later flag it
as older instead of presenting it as freshly interpreted.
"""
snapshot = request['snapshot']
if len(data['records']) < snapshot['recordCount']:
raise ValueError('The chat has fewer records than the preparation snapshot. Request a new update.')
captured = data['records'][:snapshot['recordCount']]
if not secrets.compare_digest(records_digest(captured), snapshot['evidence']):
raise ValueError('The captured records have changed since the preparation snapshot. Request a new update.')
if len(data['records']) == snapshot['recordCount'] and fingerprint(data) != snapshot['fingerprint']:
raise ValueError('The snapshot fingerprint does not match the captured chat. Request a new update.')
bounded = dict(data, records=captured)
checked = validate(copy.deepcopy(walkthrough), bounded)
checked.update(fingerprint=snapshot['fingerprint'], explainedAt=snapshot['readAt'])
return checked


def settlement(request, outcome, message=None):
if outcome not in OUTCOMES:
raise ValueError('Unknown settlement outcome.')
value = {'requestId': request['requestId'], 'outcome': outcome}
if message is not None:
if not isinstance(message, str) or len(message) > 500:
raise ValueError('A settlement message must be text, at most 500 characters.')
value['message'] = message
return value
42 changes: 41 additions & 1 deletion mcp/captains-bridge/server.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
from pathlib import Path
from chat_reader import read_chat
from walkthrough import validate, fingerprint
from preparation import new_request, check_request, prepared_walkthrough, settlement

ROOT = Path(__file__).resolve().parent
URI = 'ui://hermes-helmet/captains-bridge-v1'
Expand Down Expand Up @@ -42,6 +43,16 @@ def tool(name, title, description, properties, required=(), app=False, entry=Fal
VIEW['_meta'] = {'ui': {'visibility': ['app']}}


REQUEST = tool('request_walkthrough_update', 'Capture walkthrough request',
'Read the exact chat in the portable view and capture an immutable preparation snapshot and request identity. Starts no work.',
{'view': {'type': 'object'}}, ('view',))
REQUEST['_meta'] = {'ui': {'visibility': ['app']}}
DELIVER = tool('deliver_walkthrough_update', 'Deliver prepared walkthrough',
'For a read-only preparation agent: deliver the explanation for one captured request, or report failure. Pass the request unchanged. Cites record IDs from the request chat only. Delivery is discarded by the panel unless the request is still active. Never retry.',
{'request': {'type': 'object'}, 'walkthrough': {'type': 'object'},
'failure': {'type': 'string', 'maxLength': 500}}, ('request',), app=True)


def bind(thread_id):
data = read_chat(thread_id)
key = secrets.token_urlsafe(24)
Expand Down Expand Up @@ -93,13 +104,38 @@ def read_view(view):
return {'thread_id': view['threadId'], 'data': data, 'walkthrough': account}


def capture_request(view):
state = read_view(view)
request = new_request(state['thread_id'], state['data'])
text = 'Captured walkthrough request ' + request['requestId'] + '. No work was started.'
return {'content': [{'type': 'text', 'text': text}], 'structuredContent': {'request': request}}


def deliver(args):
request = check_request(args.get('request'))
failure = args.get('failure')
if failure is not None or args.get('walkthrough') is None:
note = failure if isinstance(failure, str) and failure.strip() else 'The preparation agent returned no explanation.'
outcome = settlement(request, 'failed', note[:500])
return {'content': [{'type': 'text', 'text': 'Reported failure for request ' + request['requestId'] + '.'}],
'structuredContent': {'preparation': outcome}}
# The reader opens only the request's own chat; the child never picks one.
data = read_chat(request['threadId'])
account = prepared_walkthrough(request, args['walkthrough'], data)
state = {'thread_id': request['threadId'], 'data': data, 'walkthrough': account}
result = present(None, state)
result['structuredContent']['preparation'] = settlement(request, 'delivered')
result['_meta']['preparation'] = result['structuredContent']['preparation']
return result


def handle(method, params):
if method == 'initialize':
return {'protocolVersion': params.get('protocolVersion', '2025-06-18'), 'capabilities': {'tools': {}, 'resources': {}}, 'serverInfo': {'name': 'hermes-helmet-captains-bridge', 'version': VERSION}}
if method == 'ping':
return {}
if method == 'tools/list':
return {'tools': [OPEN, READ, PRESENT, REFRESH, VIEW]}
return {'tools': [OPEN, READ, PRESENT, REFRESH, VIEW, REQUEST, DELIVER]}
if method == 'resources/list':
return {'resources': [{'uri': URI, 'name': 'Captain’s Bridge', 'mimeType': 'text/html;profile=mcp-app'}]}
if method == 'resources/templates/list':
Expand All @@ -109,6 +145,10 @@ def handle(method, params):
if method != 'tools/call':
raise ValueError('Unknown method or resource')
args, name = params.get('arguments') or {}, params.get('name')
if name == REQUEST['name']:
return capture_request(args.get('view'))
if name == DELIVER['name']:
return deliver(args)
if name in (VIEW['name'], REFRESH['name']):
return present(None, read_view(args.get('view')), for_app=True)
if name in (OPEN['name'], READ['name']) and args.get('thread_id'):
Expand Down
Loading
Loading