Skip to content

DO NOT MERGE — merge train: 3292,3299,3297,3303,3294,3268,3235,3269,3327,3310 - #3329

Closed
vybe wants to merge 106 commits into
devfrom
train/20261007-0858
Closed

vybe wants to merge 106 commits into
devfrom
train/20261007-0858

Conversation

@vybe

@vybe vybe commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Integration surface for #3292, #3299, #3297, #3303, #3294, #3268, #3235, #3269, #3327, #3310. Never merged; members merge individually once green.

dolho and others added 30 commits October 5, 2026 16:25
…tHub token (#3164)

Fork-to-own and the post-creation repo binding required a GitHub token in
the form even when the user had a personal token saved in Settings; the
create path already resolved it (resolve_github_pat, tier per_user) but
only used it to read the template.

- ForkToOwnRequest.github_pat / BindAgentRepoRequest.github_pat are
  optional; a supplied token still wins and is still charset-validated.
- Omitted: fork-to-own uses the resolver's per_user token; binding reads
  the caller's own saved token by user id (never the agent's current
  per-agent token). Persisted as the agent's per-agent PAT, tier `fork`,
  as before.
- The platform token is never a fork or bind identity: with no personal
  token the request is a named 400 FORK_PAT_REQUIRED pointing at Settings
  (ent#162 Decision 2 — the global PAT must not land in a per-agent row).
- A token refusal raised while using the saved token says so
  (token_source "saved", message naming Settings).
- UI: SavedGithubTokenField shows "Using your saved GitHub token" with an
  explicit override, presence read from GET /api/users/me/github-pat;
  the field is required only when nothing is saved.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… not a literal "admin" (#3262)

The system agent, the Cornelius seed and the default system seed each
hard-coded the owner username "admin". The admin account is created as
utils.admin_identity.admin_username(), which honours ADMIN_USERNAME, so
on an ADMIN_USERNAME=root install the system agent was never created
("Admin user 'admin' not found") and both seeds deferred forever as if
setup had not run. #2381 fixed the same mismatch in routers/setup.py.

All three now call admin_username() at use time. Existing installs are
unaffected: register_agent_owner never re-owns an existing row.

Fixes #3262

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…after a reload (#3265)

The composer uploaded a file on attach, then cleared its chip when the
turn went out, and nothing about the attachment reached the user's row.
The bubble showed only the typed text, so the person could not tell
whether the file went with the message, and a reload could not show it.

The 1:1 send now waits for in-flight uploads, puts the settled files on
the message and sends their names. The server keeps a sent file only
when that name is in the caller's own uploads to the agent (size and
type from there), records a failed upload with its reason, and stores
the list on the user row (new nullable `attachments` column, SQLite
migration + Alembic 0090). History returns it, so a reload shows it.
The bubble renders a thumbnail per image, a chip per other file and a
failed chip; clicking downloads through the existing upload route.
Thumbnails are fetched once per tab and only when in view, because
that route is rate-limited per person.

Fixes #3265

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
) — mechanical, per the merge-train note on the PR

tests/registry.json was the only conflict; rebuilt from the git stages
(dev's list deduped + this branch's one new entry), never spliced.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…) — mechanical, per the merge-train note on the PR

`build` was red on roomEscalationAttachments.spec.js "does NOT clear
them". The slice ran from `async function send()` to `submitUserText`,
so it now read #3265's ordinary-send `clearAttachments()`, which runs
only after the escalation branch has returned. Behaviour was right; the
pin matched by accident. The slice now ends at `const reply =
replyTo.value`, the first line of the ordinary-send tail. Moving a
`clearAttachments()` into the escalation branch still turns it red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#3265) — mechanical, per the merge-train note on the PR

#3255, #3145 and this PR each added an 0090 revision off
0089_supersede_queue_flood_backlog (the #2068 two-heads fork). The
tables are disjoint (operator_queue, chat_sessions,
enterprise_portal_messages), so this is a re-parent:
0090_portal_messages_attachments becomes 0092_portal_messages_attachments
with down_revision 0091_chat_session_claude_id. The SQLite entry is
keyed by name and needs nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
) — mechanical

#3255 (0090) and #3145 (0091) landed, so this branch's 0092 extends a
single Alembic head. The SQLite MIGRATIONS tail keeps dev's entries first;
tests/registry.json is a per-entry three-way merge (no dev entry lost or
changed).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…still succeed (#3245)

Requirements §10.10 and execution.md §Idempotency describe the replay-time
check: ended/gone runs go fresh, a lease_expired FAILED is held for the
agent timeout + 5 min after the failure (a heuristic, not a bound), the
compare-and-delete that keeps two retries to one new run, and the residuals.
…ontract (ent#730)

Requirements §49.6/§49.7, the backend and frontend area files, the dashboard
and custom-metrics feature flows, and the user doc (a copyable per-channel
recipe, the selector rules and what each refusal means) for a dashboard.yaml
widget naming one series of a dimensioned metric with dims:.

Refs Abilityai/trinity-enterprise#730
…; narrow failure-stamp read (#3245)

db.idempotency_discard_completed_if_execution deletes a completed row only
while it still names the given execution, so a slow retry cannot erase a
faster retry's fresh record. db.get_execution_failure_stamp reads (status,
error, completed_at) by primary key for the replay-liveness check.
…_dims

The read path (a dashboard.yaml widget's dims: selector) reuses it, so a
selector is valid exactly when a recorded point's dims would be. No alias,
no behaviour change; its one caller is updated.

Refs Abilityai/trinity-enterprise#730
…lt, opt-out, kill switch (#3232)

§15.1h gains the MCP caller contract: async dispatches carry the platform
turn as parent by default, a typed "manual" opts out, sync stays opt-in,
sequential /chat carries nothing, MCP_REPORT_BACK_ENABLED stops it all.
Known limits drop the stale pull-sink entry (#3114) and add the shutdown and
cleanup-sweep terminals, the cross-turn replay and plain sequential /chat.
§45 FR-5b states execution_id parity for dedicated tools; the mcp-server
architecture row names resolveReportBack and the kill-switch plumbing.

Refs #3232
… lease_expired hold window (#3245)

execution_liveness(execution_id, agent_name) answers whether the run a stored
receipt names can still succeed: live / succeeded / ended / gone /
maybe_alive / indeterminate. A FAILED whose error starts lease_expired: (the
slot reaper's mark) is maybe_alive until completed_at + the agent's timeout +
SLOT_TTL_BUFFER; the tag is checked before the stamp or the timeout. Every
doubt resolves to indeterminate (do not re-run). discard_replay_of wraps the
compare-and-delete and swallows a delete error to None.

Guards pin the tag parity with cleanup_service and that every
fail_stale_slot_execution writer carries the tag.
…oute (#3232) — red

Real createServer + real MCP client over streamable HTTP + the real
reconciler, with a fake backend recording each /task and /chat body, the
forwarded X-Trinity-Execution-Id and the Idempotency-Key. Covers the async
default, the typed opt-in and the manual opt-out, header-first, the
dedicated tool, the report_back result fields, the kill switch and its env
wiring, and a table test of resolveReportBack. Every absent assertion is
paired with a forwarding sibling so the file is red on dev.

Red on its own by design; the next commits turn it green.

Refs #3232
…, and the list says its total (Abilityai/trinity-enterprise#815)

The platform's heads-ups about a person were dropped for every machine key
AFTER the SQL limit, so a page of them could come back empty while the
caller's own rows sat just below the cut. The exclusion is now a condition
in `_list_conditions` (case-insensitive, leading-whitespace-tolerant, NULL
request_id kept) on both list routes, and the Python filter stays as a belt.

`GET /api/operator-queue` now returns `total` (a COUNT over the same WHERE),
`has_more` (read off a `limit + 1` page, never the count) and `next_offset`;
a belt drop nulls `total` and adds a `warnings` entry. The router delegates
to `operator_queue_service.list_for_principal`.

Refs Abilityai/trinity-enterprise#815
A dashboard.yaml widget bound to a dimensioned metric could only show the
cross-series fold, so per-channel tiles all showed the same number with no
label. bind_dashboard_widgets now picks a source (the one series a dims:
selector names, or today's fold) and fills the widget once, so value,
freshness, colour and sparkline come from the same place on both paths.

- parse_dims_selector (public, pure, total) validates the selector with the
  write path's validate_dims and maps its codes; {} and null are no selector.
- Matching is canonical_dims, exact match only.
- Three named refusals that never fall back to the fold:
  metric_dimension_invalid, metric_dimension_undeclared,
  metric_series_not_found (facts in binding_detail, never 'does not exist').
- Every bound widget carries bound_series facts saying what its number is.
- _latest_entry, fold, _chart and the route are untouched.

Refs Abilityai/trinity-enterprise#730
…ry /task route (#3232)

resolveReportBack decides once per call whether parent_execution_id goes on
the /task body and which report_back fields the result gets: async
dispatches (parallel+async, #946 pull-routed) default to the platform header
turn; a typed "manual" opts out; a self-task with inject_result is
excluded; sync parallel is opt-in by a typed id; the header turn wins over a
typed id, which must pass the header's own format check
(isWellFormedExecutionId); sequential /chat never carries one. One
[Report-Back #3232] log line per call. callerTurn and the idempotency key
are unchanged. createChatTools / runAgentChat take reportBackEnabled
(default on), plumbed from createServer.

Refs #3232
…ution_id (#3232)

zod drops an undeclared key before execute, so the dedicated tools lost the
typed id (and the manual opt-out) silently. They now declare it and pass it
to runAgentChat, with the same report-back rule as chat_with_agent.
reportBackEnabled threads through makeDedicatedChatTool (trailing optional)
and ReconcilerOptions (required, so tsc fails if a start site forgets it);
index.ts hands createServer's value to the reconciler.

Refs #3232
…declared direction

_threshold_verdict is now the one threshold rule; _threshold_color is a
two-line mapping of it, so a tile's colour and its verdict cannot disagree.
A successful bind writes threshold_verdict ({level: critical|warning|ok,
threshold}) whenever the metric is judgeable (not status, up_good/down_good,
at least one threshold) and the value is a number, so the field's presence
never changes between polls. Every successful bind writes the registry's
direction (neutral when none is declared). Refusals drop both.

Refs Abilityai/trinity-enterprise#730
… value on a bound tile

The user doc recommends a value: 0 placeholder for agents on an older base
image. The metric_store_unavailable arm left it in place, so an outage showed
a believable 0 under the refusal. The arm now drops value; the cached
dashboard path re-binds through the same function and gets it too.

Refs Abilityai/trinity-enterprise#730
DashboardPanel prefers an author-typed trend/trend_value over the computed
history.trend, so on an unselected bound tile an author arrow could
contradict the sparkline beside it (with direction-aware colours, a green
arrow over a red line). Every successful bind now drops them; the selected
path already did.

Refs Abilityai/trinity-enterprise#730
…ery chat_with_<agent> (#3232)

EXECUTION_ID_PARAM_DESCRIPTION states the async default, the "manual"
opt-out, how a sync call opts in, that requested is not a guarantee, that a
plain sequential call does not report back, and that a replay reports where
the first run was asked to. Parameter text is not cut at the 2,048-char
tool-description cap; the tool descriptions and DELEGATION_CONTRACT are
unchanged.

Refs #3232
…g the dashboard read

A YAML metric: [ad_spend] or metric: {a: b} is truthy, so the widget is
bound, and by_name.get(<list>) raised TypeError out of the per-widget loop,
which runs outside the store try: one bad line took the whole dashboard read
down. That widget now gets a named refusal (metric_name_invalid, naming the
YAML kind, never the value) and the rest of the dashboard renders.

Refs Abilityai/trinity-enterprise#730
…rt-back (#3232)

createServer reads MCP_REPORT_BACK_ENABLED (default on; only "false"
turns it off) and logs the mode at startup beside the #946 line. Wired into
the mcp-server service of all three compose files and documented in
.env.example. Turning it off stops every parent_execution_id, default and
typed, without an image revert.

Refs #3232
_reclaim_ended_receipt runs between begin() and the replay branch in both
admission seams (admit_chat_request for /chat, begin_task_idempotency for
/task). When the replay would hand back a dispatch receipt whose run ended
failed/cancelled/skipped, or whose row is gone, the row is compare-and-deleted
and the key claimed again once, so an identical retry dispatches a new run.
Live, succeeded, maybe-alive (lease_expired inside its window) and
indeterminate runs replay as before; sched: keys and non-receipt snapshots are
never checked. Each reclaim and each held replay is logged without the key.

Tests cover AC1 on both routes (incl. #3145's pulled /chat receipt written by
the real run_pulled_chat_turn), AC2 over the real /task endpoint, the race in
both interleavings (and that an unconditional delete would start two runs),
retry generations, the gate-first order, and structural pins for the seam
wiring and the no-late-writer precondition.
#3232)

channel-completion-report gains an Entry Points row, the per-route table,
a copy-paste example, the opt-out, the chain behaviour and a 'no note
arrived' runbook; its pull-sink row is corrected (#3114) and the shutdown /
cleanup-sweep terminals are listed. agent-to-agent-collaboration: the pull
branch forwards the parent, and async delegation reports back by default.
User docs replace the unconditional 'yes, you'll hear back' with the async
vs sequential split. feature-flows.md index row.

Refs #3232
…ter and hints

BoundMetricMark captions every bound tile whose number needs qualifying, from
the backend's bound_series facts: channel=meta for a selected series, 'sum of
3 channel values' (with '· N stale') for a fold, 'newest of 3: channel=meta'
for a last metric over several series. A metric_series_not_found refusal is a
neutral footer row (chip, selector, 'no recent data') with a deterministic
hint built from binding_detail; every dims: refusal links the docs. The copy
lives in metricFormat.js (formatDims, boundSeriesNote, refusalHint), and
chartBasisNote now shares formatDims (multi-key labels join with ', ').

Refs Abilityai/trinity-enterprise#730
trinity-ability and others added 15 commits October 7, 2026 09:31
…ases 2026-10-06)

Adds parametrized edge-case and Hypothesis property tests for capacity/slots,
idempotency, dispatch breaker + redelivery governor, loop service, credential
encryption + secret settings, URL validation, safe YAML + credential paths +
sanitizer, backlog + dispatch admission, operator queue, and pull coordination.

Tests only — no product code changes. 22 real bugs are pinned as strict xfails,
each naming its issue: #3311 #3312 #3313 #3314 #3315 #3316 #3317 #3318 #3319
#3320 #3321 #3322 #3323 #3324 #3325. Lua-tier breaker tests skip until lupa is
added (#3326).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…o 20 files (#3265 review)

- send(): held behind `settlingUploads` from the moment the composer clears
  until the turn is under way, so a second Enter while a big file uploads no
  longer starts a parallel turn (mounted spec counts startPortalChat calls).
- Composer caps the TOTAL across drops and pastes at MAX_BATCH_FILES and
  says how many were added; the server keeps the first 20 attachments and
  trims a long upload error to 300 instead of rejecting the turn with a 422.
- resolve_turn_attachments tells "uploads could not be read" (agent not
  running, Docker unreadable) from "not in your uploads": the name is kept
  without size/type rather than stored as a permanent failure.
- Tests for the two untested hops (start_portal_turn -> portal_chat,
  portal_chat -> _persist_user_turn) and the sync /chat route; each goes red
  with its kwarg removed.
- A refused thumbnail keeps its fixed 160x112 box, named, instead of
  collapsing to a chip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ork, no empty override (#3164 review)

- Bind's unexpected-error handler scrubs with the token actually used;
  `body.github_pat` is None on the saved path and raised before the scrub,
  skipping the audit row and the idempotency release.
- Bind falls back to the saved token only when the caller owns the agent;
  an admin binding someone else's agent must type a token.
- Fork-to-own refuses the saved token for an agent key (an agent key
  resolves to its owner, so it could create a repo in the owner's account).
- SAVED_TOKEN_TIERS is per_user only; per_agent cannot occur on the create
  path and would hand an agent's identity to a user-owned repo.
- SavedGithubTokenField emits `update:overriding`; BindRepoPanel and
  CreateAgentModal refuse an override submitted empty instead of sending
  the saved token. The input is BaseInput.
- /me/github-pat adds `usable` (present AND decryptable); the forms read it.
- Docs: agent-lifecycle, agent-repo-binding and template-processing
  describe the optional token and where the saved one is validated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- sanitize_text idempotence: drop the PASSWORD=' glue from the property
  alphabet; pin the real counterexample as a strict xfail on #3328.
- operator-queue marker: a continuing hold keeps the since the file already
  carried (even a blank one), not always T1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	docs/memory/feature-flows.md
#	tests/registry.json
# Conflicts:
#	docs/memory/feature-flows.md
#	docs/memory/feature-flows/agent-to-agent-collaboration.md
#	tests/registry.json
# Conflicts:
#	docs/memory/feature-flows.md
# Conflicts:
#	docs/memory/feature-flows.md
#	tests/registry.json
# Conflicts:
#	tests/registry.json
# Conflicts:
#	tests/registry.json
# Conflicts:
#	tests/registry.json
# Conflicts:
#	docs/user-docs/advanced/dynamic-dashboards.md
#	docs/user-docs/api-reference/chat-api.md
#	docs/user-docs/faq/mcp-and-api.md
#	docs/user-docs/integrations/mcp-server.md
@vybe vybe added the ui PR touches the frontend UI — triggers Playwright e2e tests label Oct 7, 2026
vybe and others added 2 commits October 7, 2026 10:20
… — mechanical

The exact-dict assertion predates the #3164 `usable` field; the no-echo
checks are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vybe

vybe commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

merge-train done: all 10 members merged individually.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants