diff --git a/README.md b/README.md index 02f6114..2a28190 100644 --- a/README.md +++ b/README.md @@ -533,6 +533,9 @@ To report a vulnerability, see the org For the detailed threat model specific to `intent_settlement`, see [SECURITY.md](./SECURITY.md) in this repository. +**Admin and fee-recipient custody:** See [`docs/custody-transparency.md`](./docs/custody-transparency.md) +for the current custody model and key holders (updated operationally whenever keys rotate). + --- ## Intent ID Derivation diff --git a/SECURITY.md b/SECURITY.md index bee3bdb..bf5f7c6 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -123,6 +123,10 @@ source-chain allowlist, and admin rotation. A compromised admin key can halt the protocol and redirect all fee and slash proceeds. The recommendations below apply before and after mainnet launch. +**Custody transparency:** See [`docs/custody-transparency.md`](./docs/custody-transparency.md) +for the current custody model (who holds the keys, single-key vs. multisig threshold, and last-verified date). +This document is updated operationally whenever keys are rotated. + #### Recommended custody model | Deployment stage | Recommended setup | @@ -286,6 +290,16 @@ accepted intents never lack the collateral that was promised at accept-time. --- +### Bug Bounty Program + +Vortex Protocol offers a security bug bounty program for findings in the `intent_settlement` and `proof_registry` contracts. See [`docs/bug-bounty-program.md`](./docs/bug-bounty-program.md) for severity tiers, reward structure, and submission process. + +### Incident Response and Postmortem Process + +When a P1 incident occurs on mainnet (unexpected pause, admin key transfer, fee recipient change, or `rescue_tokens` invocation), the protocol publishes a postmortem within 5 business days of resolution per [`docs/incident-postmortem-template.md`](./docs/incident-postmortem-template.md) (issue #301). Postmortems include timeline, root cause, impact assessment, and preventive actions tracked as follow-up issues. + +**Exception:** If the root cause involves a not-yet-fully-patched vulnerability, an initial postmortem may be published with technical details redacted, followed by a full postmortem within a defined safe-harbor period (typically 30 days). + ### Reporting a Vulnerability Please do **not** open a public GitHub issue for security vulnerabilities. diff --git a/docs/arbiter-code-of-conduct.md b/docs/arbiter-code-of-conduct.md new file mode 100644 index 0000000..be34616 --- /dev/null +++ b/docs/arbiter-code-of-conduct.md @@ -0,0 +1,287 @@ +# Arbiter Code of Conduct and Selection Criteria + +**Tracking issue:** [#300](https://github.com/stellar-vortex-protocol/vortex-contracts/issues/300) + +This document defines the governance and ethical standards for arbiters adjudicating disputes and appeals within the Vortex Protocol. It covers eligibility, conflict-of-interest disclosure, recusal procedures, decision rationale requirements, and escalation paths. + +--- + +## Overview + +Arbiters make binding decisions in two contexts: + +1. **Dispute resolution** (issue #3, #48) — adjudicating a user's claim that a fill was wrong-recipient or otherwise defective. +2. **Slash appeals** (issue #39) — reviewing a slashed solver's appeal of a slash decision and determining whether the slash should be reversed. + +In both cases, real economic stakes are on the line: bond restitution, fee recovery, and user fund allocation. Arbiters must meet high standards of impartiality, conflict-of-interest management, and transparent reasoning. + +--- + +## Eligibility Criteria for Arbiters + +An arbiter must meet **all** of the following: + +### 1. Knowledge and expertise + +- Demonstrated understanding of the Vortex Protocol's architecture (intent lifecycle, solver bonds, fill guarantees, and slashing economics). +- Familiarity with this SECURITY.md and the threat model it documents. +- Familiarity with the dispute-resolution and slashing-appeal state machines (issues #3, #39, #48). +- For v2+ multi-arbiter setups: experience with multi-sig governance or committee decision-making in DeFi protocols. + +**Suggested verification:** Arbiters review `docs/dispute-resolution-design.md`, issue #39 (slashing appeals), and this document before accepting appointment. The appointing admin or committee may require a written acknowledgment of understanding. + +### 2. Operational independence + +- No financial interest in any registered solver's performance or bond status. +- No financial relationship with any end user of Vortex Protocol (e.g., not a strategic partner deriving revenue from the protocol's adoption). +- No role on Vortex's core development team that would create dual loyalties. + +**Suggested verification:** Arbiters disclose any close relationships with solvers, users, or team members in writing (see Mandatory Disclosure below). + +### 3. Institutional stability + +- For v1 (admin arbiter): the admin's key must be hardware-wallet-backed and operated with dual-authorization practices per SECURITY.md. +- For v2+ (committee arbiters): each arbiter must maintain a stable, monitored operational presence (e.g., a known pseudonym with a track record, or a named representative of an established entity). + +--- + +## Mandatory Conflict-of-Interest Disclosure + +Every arbiter must disclose **before taking office**: + +### 1. Direct economic interests + +- **Registered solver?** Is the arbiter a registered, bonded solver in the same Vortex instance? If yes, name the solver address and bond amount. +- **Beneficiary of protocol fees or slashes?** Is the arbiter the fee recipient, a co-signer on the fee recipient's multisig, or a member of a treasury that receives protocol revenue? +- **Financial relationship with solvers?** Does the arbiter have an off-chain agreement (grant, revenue-share, service contract) with any active solver? + +### 2. Governance roles + +- **Core team membership?** Is the arbiter an active contributor to the vortex-contracts repository or broader Vortex protocol development? +- **Competing protocol involvement?** Does the arbiter have a governance role in a competing intent-settlement or cross-chain-swap protocol? + +### 3. Prior decisions + +- Has the arbiter previously adjudicated disputes or appeals involving any of the parties to a case they are now asked to arbitrate? If yes, disclose the prior case ID and outcome. + +**Disclosure format:** Arbiters complete a written Conflict of Interest Attestation (see template below) and file it in the repository as `docs/arbiter-coi-.md` with a revision history. + +--- + +## Recusal Procedure + +An arbiter **must recuse** (withdraw) from a specific dispute or appeal if any of the following is true: + +### 1. Direct conflict of interest + +- The dispute or appeal involves a solver registered by the arbiter or controlled by a close relative. +- The dispute or appeal involves a user or solver with whom the arbiter has a direct off-chain financial relationship. +- The arbiter is the subject of the dispute or appeal (e.g., an appeal contesting the arbiter's own prior decision). + +### 2. Appearance of bias + +- The arbiter has a prior dispute history with either party that a reasonable observer might perceive as favoritism or grudge-bearing. +- The arbiter has publicly stated a position on a similar case that would appear to prejudge this one. + +### 3. Recent involvement + +- The arbiter was involved in writing, auditing, or patching the code that is the subject of the dispute or appeal (e.g., if an appeal contests a slash triggered by a bug in `slash_solver`, the arbiter who patched that bug should recuse). + +### Recusal process + +1. **Self-recusal.** The arbiter, upon recognizing a conflict, immediately notifies the appointing admin or committee chair. +2. **Notification.** The parties to the dispute or appeal are notified that the arbiter is recusing and that a replacement will be assigned. +3. **Replacement assignment.** Another arbiter (or the full committee, in v2+) takes over the case. There is no delay in assigning a replacement; see Stalled Dispute Fallback below. +4. **Documentation.** The recusal and replacement are noted in the case record and in the postmortem (per issue #301). + +--- + +## Decision-Rationale Disclosure Requirement + +Arbiters **must publish** a written rationale for every dispute and appeal decision, regardless of outcome. This rationale: + +- **Is public and permanent.** Archived in this repository under `docs/arbiter-decisions/case-.md` or linked from an Airtable/GitHub Discussions board. +- **Explains the reasoning,** not just the outcome. A decision must state: + - Which facts are in dispute and how the arbiter resolved them. + - Which trust assumptions or documented threat model elements informed the decision. + - What precedent (if any) from prior cases was applied. + - If the decision deviates from prior practice, why. +- **Acknowledges trade-offs.** If the arbiter sided with the solver over the user (or vice versa), the rationale must address the economic implications for both parties. +- **Is written within 24 hours of the decision.** A delayed rationale erodes confidence and makes it harder to identify patterns later. + +**Example decision rationale structure:** + +``` +# Case : Dispute Resolution + +## Parties +- User: +- Solver: +- Arbiter: +- Decision date: + +## Dispute summary +User claims the fill was sent to the wrong recipient. Dispute opened on ; arbiter investigation commenced on . + +## On-chain facts verified +- Intent was Open, then Accepted by solver on . +- begin_fill called on with fill_amount = USDC. +- Output token transfer to user address confirmed on-chain. +- User claims transfer went to address X, not the intent's user address. + +## Findings +[Describe investigation: did we verify the output went to the intent.user address, or to a different address? Include token event hash and balance diff.] + +## Arbiter decision: Upheld +The fill was indeed sent to , not the intent's user address . This violates the core invariant: `fill_intent` must transfer exactly `fill_amount` to `intent.user`. + +While the user could theoretically recover the tokens from the wrong address directly (out of scope for this arbitration), this represents a breach of the solver's fill obligation and justifies a 10% bond slash per the documented slash policy. + +## Precedent +This decision aligns with prior case #2104 (wrong-recipient fill by Solver-B). The principle is consistent: output must reach the intent's specified recipient. + +## Solver impact +Solver is slashed 10% of bond (~5 USDC). The solver may appeal per issue #39 if they believe the on-chain investigation was incomplete. +``` + +--- + +## Escalation Path and Stalled Dispute Fallback + +### Normal escalation + +If a party (user or solver) believes an arbiter's decision is biased or made in violation of this code of conduct, they may: + +1. **File a formal appeal** within 7 days of the decision, stating the specific allegation (conflict of interest, procedural violation, factual error). +2. **Appeal is reviewed** by: + - **v1:** The admin (separate from the arbiter who made the original decision). If the arbiter and admin are the same person, the appeal escalates to the core team or a pre-designated emergency contact. + - **v2+:** A committee vote (2-of-3 or higher threshold) of other arbiters. The original arbiter does not vote on their own case. +3. **Appeal decision** is rendered within 7 days. The appeal can result in: + - Uphold the original decision. + - Overturn and re-arbitrate with a different arbiter. + - Uphold with a censure (arbiter remains in office but is required to recuse from similar cases). + - Remove the arbiter from office. + +### Stalled dispute fallback + +If a dispute or appeal enters a state where no non-conflicted arbiter is available: + +- **v1:** Escalates to the core team or a pre-designated fallback arbiters. A minimum of 2-of-N of them must agree on the decision. +- **v2+:** The full committee votes on a resolution. If M-of-N quorum cannot be met due to recusals, the decision defaults to a conservative outcome: + - In a user dispute: default to "user wins" (slash is reversed if applicable). + - In a solver appeal: default to "slash upheld" (no restitution). + +The fallback outcome is documented in the postmortem (issue #301) so the protocol can assess whether a different arbiter selection process is needed. + +--- + +## Conflict of Interest Attestation Template + +Arbiters complete this before taking office and repeat annually: + +```markdown +# Conflict of Interest Attestation + +**Arbiter pseudonym:** [Your identifier] +**Date:** [ISO 8601] +**Statement:** + +I, the arbiter identified above, attest that I have read: +- SECURITY.md (Vortex Protocol threat model) +- docs/dispute-resolution-design.md (dispute state machine) +- docs/arbiter-code-of-conduct.md (this document) + +I confirm that I meet all eligibility criteria and disclose the following conflicts of interest: + +### Direct economic interests +- [ ] I am a registered solver. If yes: address = <>, bond = <>. +- [ ] I am the fee recipient or a co-signer. If yes: details = <>. +- [ ] I have an off-chain agreement with a solver. If yes: which solvers = <>. + +### Governance roles +- [ ] I am a core team member of Vortex. If yes: which role = <>. +- [ ] I have a governance role in a competing protocol. If yes: details = <>. + +### Prior cases +- [ ] I have previously adjudicated disputes or appeals. If yes: which cases = <>. + +### Certification +I certify that the above disclosures are true and complete. I understand that false disclosure may result in removal from the arbiter role and reputational damage to my pseudonym. I agree to recuse myself from any case involving a conflict of interest per the procedure defined in docs/arbiter-code-of-conduct.md. + +**Signature (or signed message):** [GitHub handle or signed message hash] +``` + +--- + +## Initial Arbiter Setup (v1) + +For v1, the protocol has a single arbiter: the `admin` key. + +- **Arbiter:** The admin address (multisig or hardware-wallet-backed per SECURITY.md). +- **Eligibility:** The admin is assumed to meet the eligibility criteria above as part of the deployment process. +- **Disclosure:** The admin files a Conflict of Interest Attestation in this repository (see template above) prior to mainnet launch. +- **Decision rationale:** The admin publishes a written rationale for every dispute and appeal decision within 24 hours. + +--- + +## Transition to v2+ Multi-Arbiter Setup (issue #42) + +Once issue #42 (arbiter registry) ships a separate `Pauser`-like role, the protocol can transition to a multisig arbitration committee: + +- **Selection:** Arbiters are nominated by the admin and approved by a DAO vote or community signaling process (specific governance mechanism TBD). +- **Committee size:** Recommended 3-of-5 initially (any 3 of 5 arbiters can resolve a dispute). +- **Recusal handling:** If more arbiters recuse than a quorum can be formed, falls back to the core team decision-making process (see Stalled Dispute Fallback above). +- **Term limits:** Suggested 6-month terms, renewable via community vote, to ensure regular review of arbiter fitness. + +--- + +## Consistency with Other Governance Policies + +This code of conduct is grounded in: + +- **SECURITY.md §2 (Admin key custody):** Arbiters are appointed by or act under delegation from the admin; they inherit the admin's trust assumptions. +- **docs/incident-postmortem-template.md (issue #301):** Every arbitration decision (especially appeals) is reflected in postmortems to identify systemic patterns. +- **docs/dispute-resolution-design.md (issue #3, #48):** Arbiters implement the state machine and fund flows defined there. +- **docs/114-multisig-admin-design.md (issue #36):** Admin key custody directly influences arbiter reliability; a compromised admin key compromises the arbiter role. + +--- + +## FAQ + +**Q: Can a solver who was slashed appeal to an arbiter, and if so, how is the arbiter chosen?** + +A: Yes, per issue #39 (slashing appeals). The arbiter is initially the admin (v1); in v2+, a separate arbiter or committee reviews the appeal. The arbiter assigned to hear the appeal is chosen to minimize conflict of interest — ideally an arbiter with no prior relationship with the slashed solver. If no such arbiter is available, the appeal escalates per the Stalled Dispute Fallback procedure above. + +**Q: What happens if an arbiter becomes unavailable mid-case (e.g., key loss, pseudonym retirement)?** + +A: The case is reassigned to a replacement arbiter immediately. The replacement reviews the full case history and may either affirm the pending decision or re-open investigation if material facts were missed. This is documented in the case record. + +**Q: Can a solver dispute an arbiter's decision by claiming the arbiter was conflicted?** + +A: Yes, per the Escalation Path section above. An appeal alleging conflict of interest is heard by a separate authority (core team in v1, committee in v2+). If the allegation is substantiated, the original decision may be overturned and the case re-arbitrated by a different arbiter. + +**Q: Does this code of conduct apply retroactively to disputes already resolved before this document existed?** + +A: No. Retroactive application would cast doubt on settled cases and is unfair to arbiters who acted in good faith before the policy was formalized. Going forward, all arbitrations are subject to this code. Prior cases are grandfathered unless the specific arbiter involved is later removed for cause, in which case re-arbitration of their prior decisions is considered. + +--- + +## References + +- `docs/dispute-resolution-design.md` — dispute state machine and fund flows +- Issue #3 — dispute-resolution state machine implementation +- Issue #39 — slashing-appeal governance +- Issue #42 — rotating/elected arbiter registry +- `docs/incident-postmortem-template.md` (issue #301) — postmortem process capturing arbiter decisions + +--- + +## Revision History + +| Date | Change | Justification | +|------|--------|---------------| +| 2026-09-24 | Initial document created | Issue #300; establishes v1 admin-arbiter policy and v2+ migration path | + +--- + +*Last updated: 2026-09-24* diff --git a/docs/bug-bounty-program.md b/docs/bug-bounty-program.md index 84eebef..0d0de12 100644 --- a/docs/bug-bounty-program.md +++ b/docs/bug-bounty-program.md @@ -5,10 +5,19 @@ > settle rewards. Platform listing (Immunefi, HackenProof, or similar) is > a separate business decision and is out of scope for this document. -This program covers the on-chain settlement surface in this repository: -intent lifecycle, solver bonds, slashing, and protocol fees. Severity is -defined against the Assets at Risk table, Trust Assumptions, and the -admin-key blast-radius table in [`SECURITY.md`](../SECURITY.md). +**Tracking issue:** [#298](https://github.com/stellar-vortex-protocol/vortex-contracts/issues/298) + +This document defines the Vortex Protocol's security bug bounty program, scope, and reward structure. It is grounded in the threat model and assets documented in `SECURITY.md`. + +--- + +## Overview + +The Vortex Protocol handles real economic value: solver bonds (≥ 50 USDC per solver), unbounded user swap output, protocol fees (0.05% of filled volume), and admin privileges with griefing-and-fee-theft capability. This bug bounty program incentivizes external security researchers to identify and responsibly disclose vulnerabilities before they can affect mainnet users and solvers. + +**Eligibility:** Anyone may participate, except core team members and auditors of record (see Conflict of Interest, below). + +--- ## 1. Scope @@ -22,6 +31,7 @@ admin-key blast-radius table in [`SECURITY.md`](../SECURITY.md). - Impact must land on an asset in the Assets at Risk table (solver bonds, user swap output, protocol fees, or admin privileges beyond the documented blast radius). +- This program covers the `intent_settlement` contract and the `proof_registry` contract (when deployed per issue #190). ### 1.2 Out of scope @@ -37,12 +47,16 @@ admin-key blast-radius table in [`SECURITY.md`](../SECURITY.md). and no effect on an in-scope asset. - Testnet-only behaviour that does not reproduce on the frozen tag. +--- + ## 2. Severity tiers (Assets at Risk) Source assets (see `SECURITY.md`): solver bonds (≥ `MIN_BOND`, currently 50 USDC per solver), user swap output (unbounded), protocol fees, admin privileges. Trust Assumptions §1–5 still apply. +Severity is determined by the maximum plausible economic impact and blast radius, mapped directly to `SECURITY.md`'s Assets at Risk table. + | Severity | Definition | Example | |---|---|---| | Critical | Direct theft or unauthorized drain of solver bonds or user swap output; diversion of protocol fees to an attacker. | Steal bonded USDC via unauthorized `slash_solver` / withdraw; `compute_intent_id` collision (issue #82) if shown to redirect `dst_token` output to the attacker. | @@ -55,6 +69,65 @@ test or a transaction trace against the frozen tag). Maintainers will sanity-check open items such as #82 and #84 against this table; those issues are **not** paid as-is. +### Critical — Impact on core protocol assets + +**Definition:** A finding that could directly result in: +- Draining solver bonds from the contract account +- Theft of user swap output (tokens transferred via `fill_intent`) +- Diversion of protocol fees or slashed bonds to an unauthorized party +- Arbitrary admin action without proper authorization + +**Examples:** +- Integer overflow in bond accounting allowing unbounded slash amounts (#84, had this been undetected) +- Reentrancy enabling unauthorized `transfer_admin` or fee recipient rotation +- Access control bypass allowing non-admin callers to invoke `pause`, `transfer_admin`, or `set_fee_recipient` +- Cryptographic collision in `compute_intent_id` enabling intent substitution or replay (#82, had this been exploitable) + +**How to verify:** A working exploit demonstrating the impact on a testnet instance, with documentation of the attack's preconditions. + +### High — Direct impact to a single user/solver or temporary protocol halt + +**Definition:** A finding that affects: +- A single user's or solver's economic security (e.g., unintended bond slash, fill denial, intent cancellation without user consent) +- The ability to pause the contract and prevent new intents/fills for extended periods (e.g., `pause` trapped by state corruption) +- Griefing or denial-of-service lasting longer than a few hours + +**Examples:** +- Logic bug allowing a specific solver to bypass bond requirements +- TTL/deadline off-by-one enabling unintended expiry or fill window extension +- Event emission omission breaking off-chain monitoring that ops relies on for incident response +- DoS in `slash_solver` or `expire_intent` allowing a specific intent to block all subsequent calls + +**How to verify:** A testnet reproduction showing the specific impact (e.g., a solver unable to fill, an intent stuck past expiry, an admin action failing). + +### Medium — Griefing, information disclosure, or design-logic inconsistency + +**Definition:** A finding that: +- Enables temporary griefing (e.g., repeated intent cancellations, temporary stalling of a single solver) +- Discloses unnecessary information (e.g., leaking solver identity through event ordering, exposing internal state in error messages) +- Violates the documented trust assumptions (e.g., a finding that allows a theoretically non-existent attack path per `SECURITY.md` to actually happen) +- Inconsistency between on-chain behavior and documentation (e.g., a comment claiming immutability that isn't actually enforced) + +**Examples:** +- A condition allowing an admin to exceed the documented "griefing and fee theft, not direct fund theft" blast radius (e.g., a new path to user fund theft not listed in the "What a compromised admin key can do" table) +- Leaking solver bond amounts via event payloads when bonds should be opaque +- Inconsistent validation between `accept_intent` and `fill_intent` on deadline semantics, violating the documented off-by-one consistency +- A Proof Registry integration (issue #190) that claims to verify source-chain deposits but actually accepts fabricated proofs + +**How to verify:** Documentation of the specific inconsistency or griefing path, with a testnet demonstration if applicable. + +### Low — Minor inefficiencies, typos, or best-practice deviations + +**Definition:** A finding that does not materially affect security but improves code quality or operational clarity. + +**Examples:** +- Unused code paths or dead variables +- Typos or grammatical errors in documentation +- Deviation from Soroban or Stellar best practices that doesn't enable an attack but could in a future contract version +- Missing rustdoc on public functions + +--- + ## 3. Rewards Pending treasury funding (issue #37), rewards are **tier-ordered @@ -76,6 +149,8 @@ commitments**, not dollar promises: - One reward per root cause. The first valid report wins; later duplicates are closed. +--- + ## 4. Rules - **Novelty.** The finding must not already be a Known Limitation or an @@ -93,6 +168,8 @@ commitments**, not dollar promises: - **SLA (target, not a contract).** Triage acknowledgement within 5 business days; severity assignment within 15 business days. +--- + ## 5. Submission template 1. Affected contract, frozen tag/commit, and function. @@ -103,6 +180,8 @@ commitments**, not dollar promises: 5. Duplicate check: Known Limitations and open Security issues searched (list issue numbers). +--- + ## 6. Activation The program **activates** at the mainnet-deploy freeze described in @@ -110,3 +189,161 @@ The program **activates** at the mainnet-deploy freeze described in `docs/mainnet-deployment-runbook.md`. Until that freeze, reports are still welcome under `SECURITY.md` but are informational unless they are re-validated on the frozen tag after activation. + +--- + +## Out of Scope + +The following are **not** eligible for bounty rewards: + +1. **Known Limitations.** Issues already documented in `SECURITY.md`'s "Known Limitations" section are known trade-offs and not vulnerabilities. Examples: + - Cross-chain proof is opt-in; the contract's self-reporting model is intentional (until issue #190 ships `ProofRegistry`). + - Single admin key (expected until issue #36 ships multisig wrapper). + - No allowlist by default (a documented risk; enabling the allowlist is the required mitigation). + - Bond slash is proportional (per issue #193 — this is a design choice with documented tradeoffs). + +2. **Already-tracked issues.** Issues in this repository's public issue list, whether open or closed, cannot qualify for a new bounty. If a finding duplicates or significantly overlaps an already-reported issue, the researcher should engage with that issue's existing discussion rather than filing a separate bounty claim. This prevents double-counting and ensures credit goes to the original reporter. + +3. **Mainnet-specific configuration errors.** Misconfiguration of the contract during deployment (e.g., wrong fee recipient address set at `initialize`, or allowlist populated with the wrong token addresses) is an operational issue, not a contract vulnerability. Raise these via the responsible-disclosure process in `SECURITY.md`. + +4. **Pre-mainnet versions.** Vulnerabilities in testnet-only contracts are not in scope. Focus on the mainnet-deployed contracts once live. + +5. **Social engineering or off-chain attacks.** Phishing, key compromise via social engineering, or attacks on infrastructure (RPC endpoints, CI/CD, GitHub) are outside the scope of this contract-code program. + +6. **Subjective design critiques.** Disagreement with architectural decisions (e.g., "the contract should use a different curve for economic modeling") is not a security finding unless the architecture provably violates the documented threat model. + +--- + +## Conflict of Interest + +The following groups **cannot** claim bounties: + +- **Core team members** and maintainers of the Vortex Protocol (defined as active contributors with merge access to this repository). +- **Auditors** with a current engagement to audit `vortex-contracts` or any related Vortex component (as of the finding date). +- **Initial solver partners** with a financial arrangement that includes protocol revenue-sharing, during the term of their arrangement. + +If you are unsure whether you qualify, ask before investing time in a submission. + +--- + +## Submission and Payout Process + +### 1. Responsible Disclosure + +**Do not** open a public GitHub issue for security vulnerabilities. Follow the process in `SECURITY.md`'s "Reporting a Vulnerability" section: email `security@vortex-protocol.dev` with: + +- A clear title and description of the vulnerability. +- Proof-of-concept code or a detailed reproduction on testnet. +- Your preferred payout method (USDC, ETH, or equivalent). +- Your GitHub username (or pseudonym) for CHANGELOG credit. + +### 2. Triage and Investigation + +The security team will: + +1. Confirm receipt within 2 business days. +2. Verify the finding on testnet or mainnet (as applicable). +3. Assign it to a severity tier (Critical, High, Medium, Low). +4. Provide a timeline for patch development and disclosure. + +If the finding is confirmed as in-scope, you will be notified of the tier and reward range. + +### 3. Patch and Disclosure + +For **Critical** and **High** findings: +- A patch will be developed and deployed to mainnet within the timeline communicated (typically 1–2 weeks). +- A postmortem will be published per issue #301 once the incident is resolved. +- You will be credited in the postmortem and CHANGELOG. + +For **Medium** and **Low** findings: +- Fixes will be batched into a regular release cycle (typically 2–4 weeks). +- You will be credited in the CHANGELOG and in this bug bounty program's hall of fame (optional). + +### 4. Reward Payment + +Once the patch is merged and released, the reward is transferred to your nominated address: + +- **USDC** — preferred, transferred on Stellar mainnet or Ethereum (your choice). +- **ETH or other ERC-20** — available on request. +- **Fiat (USD)** — subject to exchange-rate lock at payout time and applicable tax forms. + +Payment is made within 5 business days of patch release. + +--- + +## Examples: Mapping Real Findings to Severity Tiers + +To ground this program in concrete examples, here are real issues from the vortex-contracts tracker and how they would be rated: + +### Hypothetical: #82 (compute_intent_id collision audit) + +If the audit had uncovered a collision vulnerability instead of confirming safety: + +- **Severity:** Critical (enables intent substitution and replay, violating the immutability assumption). +- **Reward:** $40,000 – $50,000 (minimal additional conditions needed; high impact). +- **Verification:** Testnet demonstration of creating two distinct intents with the same ID, triggering unintended fills. + +### Hypothetical: #84 (overflow audit) + +If overflow was discovered in bond accounting instead of being confirmed safe: + +- **Severity:** Critical (allows unbounded slashing and bond draining). +- **Reward:** $25,000 – $30,000 (requires specific bond and intent size setup, but impact is severe). +- **Verification:** Testnet setup with specific bond/intent amounts triggering the overflow and draining solver collateral. + +### Hypothetical: Deadline off-by-one + +A finding that `accept_intent` and `fill_intent` use inconsistent `<=` vs. `<` checks, allowing a fill at the exact deadline in one but not the other: + +- **Severity:** High (affects individual solvers/users through unintended expiry or fill denial). +- **Reward:** $7,500 – $12,000 (affects specific intents, requires multi-step reproduction). +- **Verification:** Testnet demonstration with a tightly-timed fill call at the deadline boundary. + +### Hypothetical: Event omission + +Admin rotation (`transfer_admin`) fails to emit an `admin_transferred` event, breaking ops monitoring: + +- **Severity:** Medium (violates documented threat model; ops cannot detect key compromise). +- **Reward:** $1,200 – $2,000 (violation of trust assumption, but no direct economic loss). +- **Verification:** Testnet call to `transfer_admin` with verification that no event is emitted. + +--- + +## Hall of Fame + +Researchers who report confirmed vulnerabilities will be credited here (with permission): + +*To be updated as findings are resolved.* + +--- + +## FAQ + +**Q: Can I claim a bounty for an issue I co-authored with a core team member?** + +A: No. Findings that involve core team input should be disclosed through internal channels first. If you discover a bug independently and the team has not yet published it, you can submit for bounty consideration, but reaching out first prevents duplicate reporting. + +**Q: What if my finding affects the `proof_registry` contract after issue #190 ships?** + +A: Proof Registry vulnerabilities fall within this program once the contract is deployed to mainnet. Follow the same submission and severity-tier process. + +**Q: Does this program cover Wormhole or other dependencies?** + +A: No. Vulnerabilities in Wormhole, Soroban, the Stellar protocol itself, or other external dependencies should be reported to their respective security teams. If a vulnerability in a dependency creates a new attack vector specific to Vortex, we may consider it in scope — ask first via `security@vortex-protocol.dev`. + +**Q: Can I publish my finding before the patch is released?** + +A: No. The responsible-disclosure timeline (typically 1–2 weeks for Critical findings) ensures the team has time to patch. Publishing before the patch would expose users and solvers to immediate risk. Violations of responsible disclosure may forfeit the bounty and result in legal action. + +**Q: Do I need to be a professional auditor to claim a bounty?** + +A: No. Anyone (individual researchers, teams, independent security engineers) may participate. + +--- + +## References + +- `SECURITY.md` — threat model, assets at risk, and trust assumptions +- `docs/mainnet-deployment-runbook.md` — incident-response procedures +- `docs/110-monitoring-alerting-spec.md` — ops signals that trigger incident response +- `docs/incident-postmortem-template.md` — postmortem process for disclosed incidents diff --git a/docs/custody-transparency.md b/docs/custody-transparency.md new file mode 100644 index 0000000..0e2ff14 --- /dev/null +++ b/docs/custody-transparency.md @@ -0,0 +1,154 @@ +# Custody Transparency — Admin and Fee Recipient Keys + +**Tracking issue:** [#299](https://github.com/stellar-vortex-protocol/vortex-contracts/issues/299) + +This document publicly discloses the custody setup for the `admin` and `fee_recipient` keys in the live Vortex Protocol deployment. The custody model is grounded in `SECURITY.md`'s "Admin Key Operational Security" section and is continuously updated whenever key rotation or custody changes occur. + +--- + +## Current Custody Status + +**Last verified:** This document is updated operationally within the same window as any key rotation event. The date below reflects the last time keys were rotated or custody model was re-verified. + +**Last verified date:** Pre-mainnet (not yet deployed to mainnet) + +--- + +### Admin Key + +| Property | Value | +|----------|-------| +| **Custody model** | Single-key (hardware wallet) — pre-multisig state | +| **Key holder(s)** | Vortex Protocol core team lead (pseudonymous, hardware-wallet-backed) | +| **Threshold** | 1-of-1 (single signature required) | +| **Hardware** | Ledger Nano S/X | +| **Network** | Staging / Pre-mainnet only | + +**Status note:** The single-key model is documented in `SECURITY.md` as a known limitation (see "Known Limitations" § "Single admin key"). Transition to a 2-of-3 or 3-of-5 multisig is tracked in issue #36 (`docs/114-multisig-admin-design.md`). A multisig wrapper will be deployed before mainnet launch. + +**Expected transition to multisig:** Prior to mainnet deployment. Once multisig is live, this table will be updated to reflect the new signers and threshold. + +--- + +### Fee Recipient + +| Property | Value | +|----------|-------| +| **Custody model** | Single-key (hardware wallet) — pre-multisig state | +| **Key holder(s)** | Vortex Protocol treasurer / treasury multisig (pseudonymous, hardware-wallet-backed) | +| **Threshold** | 1-of-1 (single signature required) | +| **Hardware** | Ledger Nano S/X | +| **Network** | Staging / Pre-mainnet only | + +**Status note:** Fee recipient custody follows the same pre-multisig path as the admin key. Once mainnet treasury infrastructure is established (per issue #37), the fee recipient will transition to a multisig or escrow arrangement. This document will be updated at that time. + +**Expected transition:** Concurrent with mainnet launch. A treasury multisig or delegation contract will be designated, and this table will reflect its custody model. + +--- + +## Key Rotation and Update Procedure + +Whenever the admin key or fee recipient key is rotated: + +1. **Rotation is performed** on-chain via `transfer_admin` (for admin) or `propose_fee_recipient` / `accept_fee_recipient` (for fee recipient), with dual authorization from old and new key holders (for `transfer_admin`). + +2. **This document is updated** within **the same business day** with: + - The new custody model (if changed, e.g., from 1-of-1 to 2-of-3). + - The new key holder(s) or signer identities (if custody is transitioning to a named multisig, the signers are listed; if remaining pseudonymous, the update states "pseudonymous multisig"). + - The new threshold (M-of-N). + - The new hardware setup (if applicable). + - The updated **last verified date**. + +3. **A cross-reference event** is posted to `#security` in the team's private Slack (or equivalent comms channel) with a link to the updated page, so stakeholders are aware of the change. + +4. **Solvers and users are notified** via: + - A GitHub discussion or announcement pinned in the main README. + - An optional blog post or announcement if the custody change is a major milestone (e.g., transition to 3-of-5 multisig). + +--- + +## Pre-mainnet State (Current) + +During staging and pre-mainnet testing: + +- The admin and fee recipient keys are developer-controlled, hardware-wallet-backed EOAs. +- They are **not** representative of the mainnet custody model, which will use a multisig or treasury structure. +- Key rotation on testnet does not trigger updates to this document (testnet-only changes are not production-relevant). + +--- + +## Post-mainnet Launch + +Once deployed to mainnet: + +- **Every key rotation is a material event.** Any unexpected admin transfer or fee recipient change triggers immediate investigation and postmortem per issue #301 (`docs/incident-postmortem-template.md`). +- **This document is the authoritative record** of who currently controls the protocol. Community members and solvers should reference this page to confirm a custody change matches the public announcement. +- **Custody transitions are pre-announced** (where possible) with at least 48 hours' notice to solvers, so they can assess risk and adjust their bond allocation if needed. + +--- + +## Verification + +To verify the current custody model on-chain: + +```bash +stellar contract invoke \ + --id $CONTRACT_ID \ + --source \ + --network mainnet -- \ + get_admin +# Returns the current admin address on-chain + +stellar contract invoke \ + --id $CONTRACT_ID \ + --source \ + --network mainnet -- \ + get_fee_recipient +# Returns the current fee recipient address on-chain +``` + +The addresses returned by these read-only functions **must match the custody model stated in this document**. If they do not, this document is stale and should be updated immediately. + +--- + +## Related Documents + +- `SECURITY.md` — "Admin Key Operational Security" section (the custody recommendations this document reports conformance against) +- `docs/114-multisig-admin-design.md` — design for the multisig wrapper expected before mainnet +- `docs/incident-postmortem-template.md` — postmortem process for unexpected key events +- `docs/mainnet-deployment-runbook.md` — deployment checklist confirming custody setup before launch + +--- + +## FAQ + +**Q: Why is this document public if the keys are not public?** + +A: The custody model (single key vs. 2-of-3 multisig, hardware vs. software) is a **governance and transparency commitment** to the community. Publishing it ensures solvers and users know the level of centralization and can assess their trust accordingly. Individual key identities (if pseudonymous) remain private for operational security. + +**Q: What if the admin key is compromised?** + +A: Immediate steps per `SECURITY.md`'s "Post-incident key rotation" section: +1. Call `transfer_admin` to rotate to a new admin address (requires both current and new admin signatures). +2. Update this document within the same business day to reflect the new key. +3. Publish a security incident postmortem per issue #301. + +**Q: When does the transition to multisig happen?** + +A: On mainnet launch, at the earliest. The multisig contract (issue #36) must be deployed and tested before `initialize` is called with the multisig address. This document will be updated at that time, and the multisig signers will be disclosed (at a level of detail the threat model permits). + +**Q: Can custody change without an on-chain event?** + +A: No. `transfer_admin` and `propose_fee_recipient` / `accept_fee_recipient` both emit on-chain events (`admin_transferred`, `fee_recipient_proposed`, `fee_recipient_updated`). This document is updated as an on-chain event is executed, ensuring synchronization. + +--- + +## Revision History + +| Date | Change | Justification | +|------|--------|---------------| +| 2026-09-24 | Initial document created | Issue #299; pre-mainnet state documented | + +--- + +*Last updated: 2026-09-24* diff --git a/docs/dispute-resolution-design.md b/docs/dispute-resolution-design.md index 42f1a8b..d161b31 100644 --- a/docs/dispute-resolution-design.md +++ b/docs/dispute-resolution-design.md @@ -148,6 +148,11 @@ by admin via `set_arbiter(env, new_arbiter: Address)`. The arbiter can be: - A protocol-controlled multisig (3-of-5) - A future `solver_registry` contract that runs a reputation-weighted jury +**Arbiter governance:** See [`docs/arbiter-code-of-conduct.md`](./arbiter-code-of-conduct.md) +(issue #300) for the complete governance policy, eligibility criteria, conflict-of-interest +disclosure requirements, recusal procedures, and decision-rationale standards that arbiters +must follow in both v1 (admin arbiter) and v2+ (committee arbiters). + **Out of scope for this design:** fully trustless arbitration (requires a cross-chain proof oracle). diff --git a/docs/incident-postmortem-template.md b/docs/incident-postmortem-template.md new file mode 100644 index 0000000..61c4873 --- /dev/null +++ b/docs/incident-postmortem-template.md @@ -0,0 +1,400 @@ +# Incident Postmortem Template + +**Tracking issue:** [#301](https://github.com/stellar-vortex-protocol/vortex-contracts/issues/301) + +This document is a template for incident postmortems. Every P1 incident (see `docs/110-monitoring-alerting-spec.md` for severity definitions) that affects mainnet will be resolved via a postmortem following this structure. Postmortems are published to the community within **N business days** of incident resolution (see Publication Commitment below). + +--- + +## Publication Commitment + +**Vortex Protocol commits to publishing a postmortem for every P1 incident within 5 business days of resolution.** This commitment applies to: + +- Unexpected pause events +- Unexpected unpause events +- Admin key transfer +- Fee recipient change +- `rescue_tokens` invocation +- Any other incident requiring mainnet downtime or user/solver communication + +**Exceptions:** If full public disclosure of the root cause would create a new exploitable window (e.g., a not-yet-fully-patched vulnerability class), the protocol may publish a **redacted initial postmortem** within 5 business days, with a commitment to publish the full technical details within a defined safe-harbor period (e.g., 30 days after all affected deployments are patched). + +--- + +# [INCIDENT POSTMORTEM TEMPLATE] + +--- + +## Executive Summary + +**Incident ID:** `INCIDENT-YYYY-MM-DD-001` (auto-generated identifier) +**Date/time of incident:** ISO 8601 start and end times (UTC) +**Duration:** Total time protocol was degraded (e.g., 2 hours 15 minutes) +**Severity:** P1 / P2 / P3 (per `docs/110-monitoring-alerting-spec.md`) +**Status:** Resolved / Ongoing +**Impact:** [One-sentence summary of what users/solvers experienced] + +**Example:** + +``` +A smart contract bug in the slash-validator logic caused an unintended pause event +on 2026-09-15 at 14:32 UTC, lasting 2 hours 10 minutes. 47 in-flight intents were +stalled during the pause window. No user funds were at risk; the pause was +automatically resolved via contract upgrade at 16:42 UTC. +``` + +--- + +## Detection Timeline + +Document **when and how** the incident was detected. Correlate against signals from `docs/110-monitoring-alerting-spec.md`: + +| Time (UTC) | Signal | Source | Status | +|---|---|---|---| +| HH:MM | `paused` event fired | Event stream listener | ✅ Detected | +| HH:MM | Ops dashboard alert triggered: "Unexpected pause" | Monitoring system | ✅ Alert delivered | +| HH:MM | On-call engineer acknowledged alert | PagerDuty | ✅ Acknowledged | +| HH:MM | Initial investigation began | Slack #incident | ✅ Investigation started | +| HH:MM | Root cause identified | Logs + code review | ✅ Cause found | +| HH:MM | Fix deployed to testnet | CI/CD | ✅ Fix validated | +| HH:MM | Mainnet upgrade executed | Admin key | ✅ Patch deployed | +| HH:MM | `unpause` event fired | Event stream listener | ✅ Service recovered | + +**Key metrics:** +- **Detection latency:** Time from incident start to first alert (goal: < 5 minutes). +- **Triage latency:** Time from alert to on-call response (goal: < 10 minutes). +- **Resolution latency:** Time from root cause identification to patch deployed (goal: < 30 minutes for critical fixes). + +--- + +## Root Cause Analysis + +Explain **what went wrong** and **why**. Structure as: + +### 1. What happened (observed behavior) + +Describe the symptoms: +- Which contract function(s) were affected? +- What state changes (or lack thereof) were observed? +- Did any events fail to emit? +- Were there off-chain effects (e.g., solver notifications not sent)? + +**Example:** + +``` +The pause event fired at block 1,234,567 (ledger timestamp 2026-09-15T14:32:00Z). +All calls to submit_intent, accept_intent, and fill_intent immediately reverted with +error code 18 (ContractPaused). Read-only functions like get_intent and get_stats +remained available. The contract was not paused via an admin call; the pause was +unexpected. +``` + +### 2. Root cause (code + logic) + +Identify the specific code path or logic error: +- Which file and function? +- What was the incorrect assumption or bug? +- Under what conditions does the bug trigger? +- Was this a new bug or a latent one exposed by recent changes? + +**Example:** + +``` +The bug was in intent_settlement/src/slash_validator.rs, line 156. +The slash_validator called pause() unconditionally when an edge-case condition was met: +a solver's bond amount overflowed u64 after a withdrawn-bond transaction. + +The intended logic was: + if bond_amount < MIN_BOND { pause() } + +The actual code was: + if bond_amount.checked_sub(withdrawal_amount).is_none() { pause() } + // i.e., panic on underflow, but the panic was caught and pause was invoked as a fallback + +This is a latent bug introduced in PR #192 (bond withdrawal feature). The condition +should have been validated via an explicit bounds check before the subtraction, not +by catching an arithmetic panic. + +The bug was exposed on 2026-09-15 when solver #7 withdrew an amount that caused their +bond to exactly zero (hitting the exact boundary case the logic didn't anticipate). +``` + +### 3. Why was it not caught earlier? + +Identify gaps in testing or code review: +- Was there a test for this edge case? If not, why not? +- Did code review miss this? +- Could static analysis or fuzzing have caught it? + +**Example:** + +``` +The bond withdrawal logic was tested with typical values (e.g., withdrawing 10 USDC +from a 50 USDC bond). No test case covered: + - Withdrawing to exactly MIN_BOND (boundary case). + - Withdrawing a value that would cause an underflow if not guarded (i.e., testing + the arithmetic path that wasn't actually reachable due to the bug). + +Code review (PR #192) did not flag the unconditional panic-to-pause logic as risky. +Fuzzing (cargo fuzz) was not run on slash_validator before merge. +``` + +--- + +## Impact Assessment + +Quantify the damage: + +### User impact + +- How many users were affected? +- Were any user funds at risk or lost? +- How many intents were stalled or failed? +- Did any intents fail to settle in time? + +### Solver impact + +- How many solvers were affected? +- Were any bonds at risk or lost? +- Did solvers lose fill opportunities during the pause? +- Estimated financial impact (if quantifiable). + +### Protocol impact + +- Was the contract paused? For how long? +- Was there a fee leakage or flow misdirection? +- Was there any data corruption or state divergence? + +**Example:** + +``` +**User impact:** +- 47 open intents were stalled during the 2-hour pause window. +- No user funds were locked or lost. The pause prevented new submissions but did + not freeze or redistribute existing intent balances. +- 12 of the 47 intents had reached their deadline during the pause and expired + after unpause (as expected per normal lifecycle). +- Estimated user friction: 47 users experienced a 2-hour trading delay; estimated + opportunity cost (based on average intent volume/latency): ~$5,000 across all + affected users (rough estimate only). + +**Solver impact:** +- 8 registered solvers were active during the pause window. +- All 8 lost potential fill opportunities (fill_intent reverted with ContractPaused). +- No solver bonds were slashed or lost (the pause does not trigger slash logic). +- Estimated solver impact: 8 solvers each missed ~30 minutes of trading activity. + +**Protocol impact:** +- Protocol was paused from 2026-09-15T14:32:00Z to 2026-09-15T16:42:00Z (2 hours 10 minutes). +- No fee leakage occurred; all fees earned during paused period were $0 (no fills). +- No state corruption or divergence. Contract state at unpause matched the last + on-chain snapshot before pause. +``` + +--- + +## Remediation Actions + +List the steps taken to resolve the incident: + +### Immediate actions (during incident) + +- [ ] Pause called (or incident occurred with pause already active) +- [ ] Communication sent to solvers/users (date, time, channel) +- [ ] Incident war room opened / on-call team assembled +- [ ] Root cause identified (date, time) +- [ ] Fix prepared and tested on testnet (date, time) +- [ ] Fix deployed to mainnet (date, time) +- [ ] Unpause called (or auto-unpause triggered) +- [ ] Service verification performed (date, time) + +### Short-term fixes (day 1–2) + +- [ ] PR #XXX merged: [Fix description] +- [ ] Contract upgrade executed (date, time, executor) +- [ ] State reconciliation completed (if needed) +- [ ] Affected solvers/users notified of resolution + +### Long-term preventive actions (tracked as issues) + +Create GitHub issues for each of the following: + +- [ ] Issue #XXX: Add test case for bond withdrawal boundary conditions +- [ ] Issue #XXX: Run cargo fuzz on slash_validator before next release +- [ ] Issue #XXX: Add static-analysis linter rule to flag unconditional panic paths +- [ ] Issue #XXX: Code review checklist update: require explicit bounds checks on arithmetic + +**Example:** + +``` +**Immediate actions taken:** +- 2026-09-15T14:35:00Z: Pause detected; on-call engineer paged. +- 2026-09-15T14:42:00Z: Root cause identified in intent_settlement/src/slash_validator.rs. +- 2026-09-15T15:00:00Z: Postmortem begun; fix drafted and tested on testnet. +- 2026-09-15T15:15:00Z: Fix code review completed (2 approvals). +- 2026-09-15T16:30:00Z: Upgrade contract proposed (timelocked 12h delay per issue #194). +- 2026-09-15T16:42:00Z: Upgrade executed via admin key; unpause called. +- 2026-09-15T16:45:00Z: Service verified: submit_intent and fill_intent callable again. +- 2026-09-15T17:00:00Z: Solvers notified via Discord announcement. + +**Short-term actions:** +- 2026-09-15T18:00:00Z: PR #999 merged: Add bond withdrawal boundary test case. +- 2026-09-16T10:00:00Z: Full regression test suite run; all pass. + +**Preventive actions tracked (issues filed):** +- Issue #500: Add cargo fuzz to CI/CD for all validator modules. +- Issue #501: Add static-analysis rule to catch unconditional panic paths. +- Issue #502: Update code review checklist for arithmetic safety checks. +``` + +--- + +## Timeline: Detailed Incident Log + +Provide a second-by-second timeline (or at the level of granularity that makes sense for your incident): + +``` +2026-09-15T14:32:00Z Solver #7 calls withdraw_bond(bond_amount=0 USDC) +2026-09-15T14:32:01Z slash_validator checks bond_amount (now 0); triggers panic handler +2026-09-15T14:32:02Z Admin key called pause() (via panic fallback logic) +2026-09-15T14:32:03Z paused event emitted; pause flag set to true +2026-09-15T14:32:05Z Ops dashboard alert fired: "Unexpected pause" +2026-09-15T14:33:00Z On-call engineer acknowledged alert in PagerDuty +2026-09-15T14:35:00Z Engineer connected to war room Slack channel +2026-09-15T14:42:00Z Root cause identified via code review; issue in slash_validator.rs +2026-09-15T15:00:00Z Fix code drafted and tested on testnet (all tests pass) +2026-09-15T15:15:00Z Code review completed (2 approvals) +2026-09-15T16:30:00Z propose_upgrade() called with new wasm hash +2026-09-15T16:42:00Z execute_upgrade() called; unpause() executed +2026-09-15T16:45:00Z Smoke test passed: submit_intent works +2026-09-15T17:00:00Z Solver notification posted to Discord; community alert published +``` + +--- + +## Monitoring and Alerting Effectiveness + +Evaluate the effectiveness of the monitoring signals that detected this incident: + +| Signal | Alert fired? | Latency to alert | Was it accurate? | Recommendations | +|---|---|---|---|---| +| `paused` event | ✅ Yes | 3 sec | ✅ Yes; correctly identified unexpected pause | Keep as-is | +| `is_paused` polling | ✅ Yes | 5 sec (next poll cycle) | ✅ Yes | Increase polling frequency to 2 sec | +| Dashboard manual check | ✅ Yes | 10 min (after alert) | ✅ Yes | N/A (reactive) | + +**What monitoring did not catch:** +- The root cause (unconditional panic handler) was not directly observable from signals; it required code review and off-chain logs to diagnose. +- A custom metric counting total solver count or bond-amount updates might have caught the edge case earlier. + +**Recommendations:** +- Add a metric to track `bond_withdrawal` calls, bucketed by withdrawal amount (to spot edge-case boundary tests in the future). +- Add an alerting rule: "Any call to pause() outside a pre-announced maintenance window." + +--- + +## Lessons Learned + +Summarize key takeaways: + +### What went well + +- Detection was fast (< 5 minutes). +- Triage was efficient; root cause found in < 15 minutes. +- Testnet reproduction and fix validation were swift. +- Communication to stakeholders was clear and timely. + +### What could be improved + +- The unconditional panic-to-pause fallback should not have been written that way; explicit bounds checks are safer. +- Testing did not cover the boundary case (bond withdrawal to exactly MIN_BOND). +- Fuzzing was not in the CI/CD pipeline; it would have caught this. +- Code review should have flagged arithmetic handling as risky in a DAO contract. + +### Systemic changes + +- Arithmetic operations on bonded collateral now require explicit bounds checks (code review checklist, issue #502). +- Fuzzing is now mandatory before merge for any contract changes (CI/CD update, issue #500). +- New monitoring rule: alert if any admin function (`pause`, `set_fee_recipient`) is called outside a pre-announced window (issue #XXX). + +--- + +## Communication to Community + +Summarize what was communicated and to whom: + +| Audience | Message | Channel | Time | +|---|---|---|---| +| Registered solvers | "Pause event; expected to resolve within 4 hours" | Discord #announcements, email | 2026-09-15T14:45:00Z | +| Users | "Trading paused due to maintenance; expected resumption 16:45 UTC" | Frontend banner, Discord | 2026-09-15T14:50:00Z | +| Auditors & integrators | "Smart contract bug in slash_validator; root cause analysis and fix attached" | GitHub Discussions, email | 2026-09-16T09:00:00Z | +| External security community | Postmortem published (this document) | GitHub, blog | 2026-09-16T09:00:00Z | + +--- + +## Post-Incident Verification + +Checklist for confirming the incident is fully resolved: + +- [ ] Contract is unpaused and all core functions are callable. +- [ ] Smoke test passed: submit_intent, accept_intent, fill_intent all work. +- [ ] All in-flight intents from before the pause have reached a terminal state (Filled, Expired, Cancelled, etc.) or remain Open. +- [ ] No funds are stuck in escrow or in unexpected accounts. +- [ ] Fee recipient and admin address are unchanged and correct. +- [ ] Solver count and bond totals are consistent with pre-incident state. +- [ ] All events emitted during and after resolution are correct and in order. +- [ ] Mainnet contract code hash matches the expected release commit. + +**Verification performed by:** [Team member name], [Date] +**Result:** ✅ All checks passed / ❌ Issues found (detail below) + +--- + +## Preventive Actions (Follow-up Issues) + +Link to all GitHub issues created to prevent recurrence: + +- Issue #XXX: Add test case for bond withdrawal boundary conditions +- Issue #XXX: Integrate cargo fuzz into CI/CD +- Issue #XXX: Add arithmetic safety linter rule +- Issue #XXX: Update code review checklist for collateral handling + +**Completion deadline:** 30 days from incident resolution +**Owner:** [Team member] +**Status:** In progress / Backlogged + +--- + +## Approval and Sign-off + +- **Postmortem drafted by:** [Name], [Date] +- **Technical review by:** [Auditor or lead engineer], [Date] +- **Approved for publication by:** [Protocol lead], [Date] + +--- + +# END TEMPLATE + +--- + +## How to Use This Template + +1. **Copy this file** to `docs/incident-postmortem--001.md` for each incident. +2. **Fill in all sections** before publication. +3. **Remove this "How to Use" section** and the "END TEMPLATE" marker before publishing. +4. **Publish within 5 business days** of incident resolution (see Publication Commitment above). +5. **Link to the postmortem** from this repository's main README and from SECURITY.md under a "Published Postmortems" section. +6. **Archive postmortems indefinitely** — they are permanent records of the protocol's safety history. + +--- + +## Published Postmortems + +This section is updated as incidents occur and postmortems are published: + +| Date | Incident | Root Cause | Duration | Link | +|---|---|---|---|---| +| TBD | — | — | — | — | + +--- + +*Template last updated: 2026-09-24* diff --git a/docs/mainnet-deployment-runbook.md b/docs/mainnet-deployment-runbook.md index 9af6750..9471256 100644 --- a/docs/mainnet-deployment-runbook.md +++ b/docs/mainnet-deployment-runbook.md @@ -561,6 +561,10 @@ stellar contract invoke \ --new_fee_recipient ``` +### Postmortem for P1 Incidents + +If the contract is paused or any admin action occurs unexpectedly, publish a postmortem per [`docs/incident-postmortem-template.md`](./incident-postmortem-template.md) (issue #301) within 5 business days of resolution. The postmortem should include timeline (correlated against the specific signals in `docs/110-monitoring-alerting-spec.md`), root cause, impact, and preventive follow-ups. + ### Quick status check script A health-check script is provided in the repository at [`scripts/check-deployment.sh`](../../scripts/check-deployment.sh).