Skip to content

[G16] Build repeatable two-organization runtime qualification #16

Description

@jjangg96

Goal

Provide a repeatable, opt-in hardware/live test suite proving real GitHub workflow routing, Linux DB services and macOS jobs across two organizational installations.

Create one active Codex goal from the statement above when this issue is dispatched. The Project Goal field is a work specification; it does not start an agent. Do not invent a token budget.

Execution contract

Field Value
Goal key G16
Stage M3 - Reliability qualification
Initial status Backlog
Primary agent gpt-5.6-luna / max
Priority / risk P1 / Medium
Test profiles offline, trusted-runtime, trusted-live-github

Use one issue branch/worktree and one focused PR. Independent Luna max review is required for authentication, protocol, concurrency, resource ownership, cleanup or service identity boundaries; other changes need independent contract review. Model fields are routing instructions, not GitHub user assignments.

Dependencies

Dependencies must be Done before implementation begins. A new issue is not blocked simply because its future evidence has not been collected.

Scope

Fixtures, scripts and sanitized evidence format from reviewed contracts; no unauthenticated public PR access to local hardware.

TDD and failure evidence

  1. Red: tests detect wrong org/label routing, leftover DB/container/workdir and busy-job termination.
  2. Execute Linux shell/job-container/services workflows and native macOS workflows, queue bursts and cap changes.
  3. Use immutable reviewed commit and disposable controlled resources; record every skipped profile.

Capture a meaningful failing case before the implementation, then green evidence and relevant refactor checks. Tooling/prose-only work uses appropriate negative checks without artificial application tests. Live/runtime profiles require reviewed commits, a dedicated trusted test environment and explicit authorization for the concrete experiment. Public PR CI uses hosted environments without credentials. Planned or skipped tests never count as passed.

Acceptance criteria

  • Both orgs and both platforms appear in evidence, with actual job URLs sanitized if private.
  • Parallel service tests catch port/network/workspace conflicts.
  • Cleanup uses owned inventory and preserves existing manual infrastructure.
  • Luna max signs off contract coverage; credentials stay outside fixtures/artifacts.
  • Record exact validation commands, actual results, skipped/live-test gaps and applicable rollback notes in the PR.
  • Independent review is resolved and the focused PR is merged under the repository execution policy.
  • Update issue/Project accurately; mark the active goal complete only after all required evidence and work are complete.

Safety invariants

Preserve existing manual runners; no global Docker prune/context switching, broad process kill, implicit App enrollment, busy-job cancellation during ordinary scale-down or transparent workflow replay. Use only verifiably owned resources. Keep management credentials and raw secret-bearing SDK errors out of worker environments, logs, fixtures and commits; per-worker JIT transport follows G01. Native pools remain trusted-only.

Design references

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent:luna-maxPrimary implementer: gpt-5.6-luna, max reasoningpriority:P1Required delivery workrelease:distributionApproved delivery sequencing; does not change acceptance or dependency gatesrisk:mediumBounded contract and validation reviewtype:implementationBounded implementation goal with TDD evidence

    Type

    No type

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions