From a25a51589eeff403a947885350d25720ff94f198 Mon Sep 17 00:00:00 2001 From: yaaertu Date: Wed, 2 Sep 2026 08:09:26 +0300 Subject: [PATCH 1/6] docs: re-audit competitive landscape --- docs/product/LANDSCAPE.md | 136 +++++++++++++++++++++++++++++++++----- 1 file changed, 119 insertions(+), 17 deletions(-) diff --git a/docs/product/LANDSCAPE.md b/docs/product/LANDSCAPE.md index 1cee37a..6b2f5c4 100644 --- a/docs/product/LANDSCAPE.md +++ b/docs/product/LANDSCAPE.md @@ -1,33 +1,135 @@ # Competitive landscape -FixBundle should not pretend the surrounding problems are unsolved. Several strong tools already cover adjacent pieces. The product boundary is deliberately narrower: **portable failure evidence**, not generic repository packing or generic agent memory. +> Last re-audited: 2026-09-02 -## Repomix -Repository: https://github.com/yamadashy/repomix +FixBundle must not sell novelty that the market does not support. The surrounding problem is real, but several products now cover major pieces of the same workflow. The product only deserves to continue if users repeatedly choose the **portable, cross-source evidence-diff** job strongly enough to justify a standalone tool. -Repomix is the mature reference for **codebase → LLM-friendly context**. It packs local/remote repositories, counts tokens, supports compression and MCP integrations, and has a substantial user base. +## Market reality -FixBundle should not compete by becoming a worse repo packer. +The broad message **“turn failures into agent-ready debug bundles” is not unique**. -**Different job:** FixBundle packages a *failure event*: command output/exit state, exact Git identity, incident-vs-current revision, environment/stack evidence, selected source/config, redaction and checksums. +The broad message **“let an AI inspect a failed GitHub Actions run” is not unique**. -## temporal-debug-skill -Repository: https://github.com/MeherBhaskar/temporal-debug-skill +The broad message **“compare two runs and show what regressed” is not unique**. -This agent skill independently validates the temporal-debugging pain: production incidents may belong to an older commit, and temporary Git worktrees are safer than checking out over current work. +Therefore FixBundle must not position itself as a generic AI debugging platform, observability product, CI doctor, test platform, or repository packer. -FixBundle v0.3 overlaps on isolated worktree lifecycle. We should acknowledge that overlap instead of claiming novelty. +The only currently defensible wedge is narrower: -**Different job:** the skill teaches an agent *how to inspect history*. FixBundle creates a standalone, inspectable archive that can cross agent/vendor/support boundaries and carries logs, environment, source/config snapshots, redaction and integrity metadata. +> **Capture a known-good and a failing state as portable evidence artifacts, verify their integrity, and deterministically show which evidence changed across local, historical Git, GitHub Actions, and OpenTelemetry sources before anyone makes a causal claim.** -## GitHub Actions + Copilot -GitHub can explain failed workflow checks with Copilot and exposes downloadable workflow logs. +That wedge is still a hypothesis, not a validated market. -**Different job:** FixBundle v0.4 should not merely re-explain the same log. It should normalize a failed CI run into the same portable evidence schema used locally, tie it to commit/diff/config context, sanitize it, and make the archive usable outside GitHub Copilot. +## Direct and adjacent competitors + +### DebugBundle + +- Product: https://debugbundle.com/ +- Repository: https://github.com/debugbundle/debugbundle +- Public repo observed 2026-09-02: created 2026-05-06, AGPL-3.0, TypeScript, 6 stars, 0 forks. +- Positioning: production errors → deterministic agent-ready bundles through SDKs, CLI, API, MCP, dashboard and hosted/self-hosted workflows. +- Capabilities include production event capture, redaction, incident grouping, reproduction artifacts, deploy metadata, probes, analytics and deploy-comparison analysis. + +**Overlap:** deterministic bundles, redaction, agent-readable evidence, local-first workflow, production incident context. + +**Current difference:** FixBundle does not require application SDK instrumentation or a hosted incident backend. It can package arbitrary local command failures, isolated historical Git commits, GitHub Actions runs and OTLP file exports into portable checksummed ZIPs, then compare two artifacts offline. That difference matters only if users value it enough to retain and reuse artifacts. + +**Threat level: HIGH.** If DebugBundle or another product exposes equally low-friction arbitrary cross-source offline artifact comparison, FixBundle loses a major differentiation claim. + +### GitHub Agentic Workflows / CI Failure Doctor + +- Project: https://github.com/github/gh-aw +- CI failure investigation example: https://github.github.com/gh-aw/gallery/ci-failure-investigation/ +- Audit/diff reference: https://github.github.com/gh-aw/reference/audit/ + +GitHub Agentic Workflows can start an agent after a failed workflow, preload failed jobs/logs/artifacts, correlate failures with repository changes and produce an actionable diagnosis. Its audit tooling can also compare workflow runs for behavioral drift. + +**Overlap:** failed-run evidence collection, run comparison, agent consumption. + +**Current difference:** GitHub's workflow is GitHub-native and agent-execution-oriented. FixBundle's claim is a portable, vendor-neutral evidence artifact that can leave GitHub and be compared with local or OTLP evidence without an AI being required. + +**Threat level: HIGH for GitHub-only use cases.** FixBundle should never become a worse CI Doctor. + +### TestSprite CLI + +- Repository: https://github.com/TestSprite/testsprite-cli +- Product: https://www.testsprite.com/ + +TestSprite is a cloud testing platform. Its CLI exposes self-consistent failure bundles and `test diff `, including verdict changes, failure-kind changes, failed-step shifts, per-step status flips and code-version drift. + +**Overlap:** failure bundles, durable run identity, before/after comparison, agent-friendly output. + +**Current difference:** TestSprite owns the test execution platform and compares TestSprite test runs. FixBundle does not own execution; it imports evidence from arbitrary local, historical, CI and telemetry sources. + +**Threat level: MEDIUM/HIGH.** It independently validates the “last green vs current red” job, but also proves that run-diff can be a feature inside a larger product rather than a standalone business. + +### temporal-debug-skill + +- Repository: https://github.com/MeherBhaskar/temporal-debug-skill + +The skill addresses historical debugging by resolving an incident commit and inspecting it in an isolated Git worktree rather than changing the user's active workspace. + +**Overlap:** historical revision safety and temporal debugging. + +**Current difference:** the skill teaches an agent how to inspect history. FixBundle emits an inspectable portable artifact with command/log/environment/source evidence, redaction and integrity metadata. + +**Threat level: MEDIUM.** Historical Git isolation by itself is not a product moat. + +### Repomix + +- Repository: https://github.com/yamadashy/repomix + +Repomix is the mature reference for repository → LLM-friendly context packaging. + +**Overlap:** portable context for AI tools. + +**Current difference:** Repomix packages code. FixBundle packages a failure state and its evidence identity. + +**Threat level: LOW if FixBundle stays failure-first. HIGH if FixBundle drifts into generic repository packing.** + +### Support-bundle ecosystems + +Examples include Replicated Troubleshoot and product-specific support bundles. These systems collect diagnostics, redact sensitive data and produce archives for support/debugging. + +**Overlap:** bounded diagnostic collection, privacy, archive handoff. + +**Current difference:** most support-bundle systems are product/platform specific and do not treat two arbitrary evidence archives as a cross-source temporal comparison contract. + +**Threat level: MEDIUM.** “Make a redacted ZIP” is commodity functionality. + +## What the market evidence says + +Current public discussions repeatedly show engineers comparing last-green/current-red states manually, checking exact failed steps, recent changes, environment drift and historical failure signatures. This validates the problem shape. + +It does **not** validate FixBundle as a standalone product. A real market test must show that users: + +1. accept the capture friction, +2. intentionally retain an artifact, +3. later produce or obtain a second artifact, +4. run a comparison, +5. say the comparison reduced evidence reconstruction, uncertainty or tool switching, +6. choose this workflow over built-in GitHub/TestSprite/observability alternatives for a concrete reason. ## Product boundary -FixBundle wins only if this stays true: -> **Repomix packages code. Temporal Debug navigates time. Copilot explains a GitHub failure. FixBundle packages the evidence of a failure so any debugger can work from the same facts.** +FixBundle is not: + +- an AI root-cause oracle, +- an observability SaaS, +- an AI SRE, +- a cloud testing platform, +- a generic repo packer, +- a replacement for GitHub Actions logs, +- a dashboard that happens to export JSON. + +The candidate job is: + +> **Prove what changed in failure evidence before anyone guesses why.** + +## Kill rule + +If external validation cannot prove that artifact portability + cross-source deterministic compare creates repeated value, do not solve the problem by adding more adapters, dashboards, MCP surfaces or AI summaries. + +At that point FixBundle should be narrowed into a library/protocol/CI utility, pivoted to a stronger adjacent job, or archived. -If a new feature does not strengthen that result, it probably does not belong in the core. +See `docs/product/VALIDATION.md` for the falsification protocol. From ed165fd7c12e8ea629d84a772df2945096b2dd8f Mon Sep 17 00:00:00 2001 From: yaaertu Date: Wed, 2 Sep 2026 08:10:12 +0300 Subject: [PATCH 2/6] docs: add v0.6 market falsification protocol --- docs/product/VALIDATION.md | 330 +++++++++++++++++++++++++++++++++++++ 1 file changed, 330 insertions(+) create mode 100644 docs/product/VALIDATION.md diff --git a/docs/product/VALIDATION.md b/docs/product/VALIDATION.md new file mode 100644 index 0000000..cdb8ba5 --- /dev/null +++ b/docs/product/VALIDATION.md @@ -0,0 +1,330 @@ +# v0.6 Market Reality & Falsification Protocol + +> Status: ACTIVE +> Started: 2026-09-02 +> Code freeze: v0.7 feature work is blocked until this protocol produces evidence. + +## 1. Decision we are trying to make + +The question is not whether FixBundle can be engineered further. It can. + +The question is whether **portable cross-source failure evidence + deterministic before/after comparison** is valuable enough to deserve a standalone product and repeated user behavior. + +This protocol is designed to falsify that thesis quickly. + +## 2. Current thesis + +Candidate job-to-be-done: + +> **I have a known-good state and a failing state. Give me a trustworthy, portable answer to “what evidence changed?” before an AI or human starts guessing causality.** + +Candidate differentiators: + +1. **Cross-source:** local command, historical Git, GitHub Actions and OTLP file evidence share one comparison model. +2. **Portable:** the result is an inspectable ZIP/JSON artifact, not a dashboard-only view. +3. **Offline compare:** two retained artifacts can be compared without an account, cloud backend or LLM. +4. **Integrity first:** checksum validation and fail-closed ZIP handling happen before evidence is interpreted. +5. **Non-causal contract:** compare reports observed deltas and unavailable evidence instead of manufacturing root-cause confidence. + +Every one of these is a hypothesis until an unrelated user values it. + +## 3. Known market facts that weaken the thesis + +### Direct overlap exists + +- DebugBundle already markets deterministic, agent-ready production debug bundles through SDK, CLI, API, MCP and dashboard surfaces. +- TestSprite already ships self-consistent failure bundles and a two-run `test diff` command. +- GitHub Agentic Workflows already provides CI failure investigation and workflow-run audit/diff capabilities. +- Product-specific support-bundle ecosystems already normalize/redact diagnostics into archives. + +Therefore the following claims are prohibited: + +- “Nobody has done this.” +- “First AI debugging bundle.” +- “Unique failure bundle format.” +- “Only tool that compares failed runs.” +- “The market is empty.” + +### Category pull is not proven + +A direct competitor existing is not proof that customers care. At the 2026-09-02 audit, the public `debugbundle/debugbundle` repository had 6 stars and 0 forks. That is evidence of active competition, not evidence of broad pull. + +Likewise, FixBundle currently has no unrelated-user retention signal. + +## 4. What must be true for FixBundle to continue + +### H1 — Evidence delta adds material information + +For qualified public failures, FixBundle-style evidence comparison must surface at least one fact that materially narrows the investigation beyond “the workflow is red.” + +Examples of material information: + +- an overall-red run actually failed for a different job/step than the incident under investigation, +- the alleged suspect commit is already present in a later green run, +- the first actual red occurs later than the issue author believes, +- environment/runtime/dependency identity changed while source did not, +- a production trace belongs to a different release/deployment than the current source tree, +- a required evidence class is absent, making a causal claim unjustified. + +Non-material output: + +- reformatting the same stack trace, +- restating the issue title, +- generic AI root-cause suggestions, +- a giant diff without a narrowed evidence question. + +### H2 — Portability matters + +At least one unrelated user must prefer retaining a portable artifact because the investigation crosses tool boundaries: GitHub → local IDE/agent, production telemetry → repository, vendor support → engineering, or one AI tool → another. + +If everyone is satisfied with the native platform view, portable ZIPs are a feature looking for a market. + +### H3 — Retention creates second-use value + +At least one unrelated user must intentionally keep a first artifact and later compare it with a second real incident/run. + +This is the north-star proof. + +A star, clone, install or one-off compliment does not satisfy H3. + +### H4 — FixBundle beats “just use the platform” somewhere + +For at least one repeated workflow, the user must be able to state why GitHub-native CI diagnosis, TestSprite, DebugBundle, Sentry/observability AI, or manual log comparison is insufficient or more cumbersome. + +If there is no crisp answer, FixBundle has no defensible product boundary. + +## 5. Qualification filter for public field cases + +A field case counts only when all mandatory criteria are true: + +- opened by a human maintainer/contributor, not a reporting bot, +- currently unresolved or still useful as a live investigation, +- contains a real failure identity: run, job, command, trace, exception or reproducible event, +- has at least two states worth comparing or a credible missing-baseline problem, +- FixBundle-style analysis can be performed read-only before outreach, +- the resulting note adds evidence rather than promotion. + +Preferred cases: + +- last-green / first-red uncertainty, +- “works locally, fails in CI,” +- repeated overall-red runs with uncertain failure identity, +- dependency/environment drift, +- production incident tied to historical code, +- CI vs production mismatch, +- recurring incident where a retained baseline would matter. + +Reject: + +- already-bisected issues with exact root cause and fix, +- bot-generated diagnosis reports, +- giant projects where a drive-by note adds noise, +- cases requiring private data we cannot inspect, +- issues where the only contribution would be a FixBundle link. + +## 6. Field experiment ladder + +### Experiment A — 10 evidence-first public cases + +Goal: test H1 before asking anyone to install anything. + +For 10 qualified cases, record: + +- repository + issue, +- human/bot qualification, +- baseline state, +- incident state, +- evidence classes available, +- evidence classes missing, +- exact new fact found, +- whether the new fact changes the suspect window or next debugging step, +- whether the same conclusion was already explicit in the issue. + +Working pass threshold: + +- **PASS:** at least 5/10 produce a material new fact or a precise missing-evidence finding. +- **WEAK:** 3–4/10. +- **FAIL:** fewer than 3/10. + +This is a product-learning threshold, not a statistical claim. + +### Experiment B — Evidence-first outreach + +Only after a case has a useful read-only result. + +Outreach format: + +1. lead with the exact evidence correction/narrowing, +2. include direct run/job/commit references, +3. separate FACT from HYPOTHESIS, +4. identify the missing evidence required for stronger causality, +5. do not lead with a product link, +6. mention FixBundle only if it is relevant to retaining/capturing that evidence. + +Initial cap: **5 contextual outreaches**. No bulk promotion. + +Record: + +- response / no response, +- correction accepted / disputed, +- did maintainer change investigation direction, +- did maintainer ask how the evidence was produced, +- did maintainer install/capture, +- friction encountered. + +Working failure signal: 0 useful responses or capture attempts after 5 genuinely qualified, evidence-first contacts. + +### Experiment C — External activation + +A user counts as activated only when they generate a valid FixBundle from their own real failure. + +Do not count: + +- our own repo, +- our own machine-only demos, +- fixture bundles, +- a user merely starring/cloning the repository. + +For each activation measure: + +- install path used, +- time from install start to first valid bundle, +- capture mode, +- bundle size, +- redaction/preview concern, +- command/token friction, +- whether user could explain what the bundle was for. + +Target learning state: 3 unrelated real captures. + +### Experiment D — Baseline retention + +After a useful first capture, ask one simple behavioral question: + +> “Would you keep this ZIP as the baseline for the next time this fails?” + +Record actual behavior, not intention. + +Retention counts only when the user keeps or references the artifact later. + +### Experiment E — Repeat-use compare + +This is the north-star test. + +Success event: + +```text +same unrelated user + ↓ +real incident/run #1 captured + ↓ +artifact intentionally retained + ↓ +real incident/run #2 occurs + ↓ +fixbundle compare baseline incident + ↓ +user reports a concrete reduction in evidence reconstruction, +uncertainty, or tool switching +``` + +Capture the exact sentence or behavior. Do not translate it into fabricated “minutes saved” unless the user actually measures time. + +## 7. User #1 candidate: tanbamboo/rusql#159 + +Current read-only finding: + +```text +M59/M61 2d446a0 mysql-diff PASS +last green 2574725 mysql-diff PASS +M64 ad79f15 mysql-diff PASS +M63 a3d0372 mysql-diff FAIL +later 8f9ba52 mysql-diff FAIL +``` + +Two overall-red runs listed as the “same failure” actually had `mysql-diff` success and were red because the Rust formatting job failed. The first verified `mysql-diff` failure is the M63 merge run. + +The harness also launches `rusql-server` with child stdio ignored, so the current evidence cannot distinguish server panic/exit from a connection/lifecycle failure. That is a concrete evidence gap, not a root-cause guess. + +Status: candidate prepared; external contact requires explicit approval. + +## 8. Competitive displacement questions + +When a user activates, ask only questions tied to actual workflow: + +- Why not use the GitHub Actions UI / Copilot / CI Doctor for this case? +- Why not keep the evidence in Sentry/Datadog/other observability backend? +- Would JSON/Markdown alone be enough, or does an inspectable ZIP matter? +- Do you need to compare across different sources, or only two runs from the same platform? +- Is integrity/checksum validation valuable or invisible ceremony? +- Would you install an SDK to get richer evidence, or is “no instrumentation” the point? +- Which evidence field made the comparison useful? + +Do not turn every answer into a feature request. Look for repeated independent demand. + +## 9. Kill / pivot criteria + +FixBundle should be narrowed, pivoted or archived if any of these become true: + +### Kill signal K1 — No evidence advantage + +After 10 qualified public cases, fewer than 3 produce a material new fact or a precise missing-evidence finding. + +### Kill signal K2 — No activation pull + +After 5 high-quality evidence-first outreaches, nobody attempts a real capture or asks for the workflow. + +### Kill signal K3 — No retention + +After 3 unrelated real captures, nobody intentionally retains an artifact or sees a reason to compare later. + +### Kill signal K4 — Native platforms erase the wedge + +A major existing tool provides the same arbitrary local + historical + CI + telemetry artifact model, offline cross-source compare, comparable integrity/privacy guarantees and lower adoption friction, and target users prefer it. + +### Kill signal K5 — Value collapses into a feature + +Users consistently like the comparison but only inside another tool they already use. In that case, stop pretending the CLI is the business. Consider a library, GitHub Action, schema/protocol or integration surface instead. + +## 10. Pivot hypotheses allowed after failure + +These are research directions, not roadmap commitments: + +1. **Evidence Diff CLI/library** — strip the product to deterministic run/incident delta generation. +2. **CI artifact action** — failed workflow → redacted evidence artifact, no standalone capture UX. +3. **Support-bundle verifier** — validate/redact/compare third-party diagnostic archives. +4. **Evidence protocol** — schema + integrity tooling that other products embed. +5. **Temporal incident linker** — map old production evidence to exact historical source without mutating the active workspace. + +Do not build any of these before the current thesis is tested. + +## 11. Change-control during validation freeze + +Allowed without unlocking v0.7: + +- documentation corrections, +- packaging/install friction fixes that block a real user, +- security/privacy fixes, +- broken release/CI fixes, +- evidence needed to run the validation experiments. + +Not allowed without evidence: + +- new source adapters, +- dashboard/UI expansion, +- MCP/agent integrations, +- AI root-cause summaries, +- hosted cloud features, +- speculative enterprise controls, +- version-number-driven feature work. + +## 12. Decision log + +Every meaningful field result should end in one of four labels: + +- `THESIS_STRENGTHENED` +- `THESIS_WEAKENED` +- `FRICTION_ONLY` +- `NO_SIGNAL` + +The goal is not to protect FixBundle. The goal is to reach a correct product decision faster. From 4b254a33b74914b5d8249e865b7a43c5c1cf6554 Mon Sep 17 00:00:00 2001 From: yaaertu Date: Wed, 2 Sep 2026 08:10:32 +0300 Subject: [PATCH 3/6] docs: lock next move to validation freeze --- docs/product/NEXT.md | 114 ++++++++++++++++++++++++++++++++++++------- 1 file changed, 97 insertions(+), 17 deletions(-) diff --git a/docs/product/NEXT.md b/docs/product/NEXT.md index a1b55fa..c85306d 100644 --- a/docs/product/NEXT.md +++ b/docs/product/NEXT.md @@ -1,8 +1,8 @@ # Next move -## Distribution gate — prove repeat use before v0.7 +## v0.6 validation freeze — prove a standalone job before v0.7 -v0.3–v0.6 now cover the evidence lifecycle: +v0.3–v0.6 already cover the evidence lifecycle: ```text local / historical / GitHub Actions / OTLP production @@ -14,22 +14,102 @@ local / historical / GitHub Actions / OTLP production deterministic what-changed ``` -The next highest-value move is **not another adapter**. It is proving that an unrelated maintainer can install FixBundle, capture a real failure, keep the ZIP, and use FixBundle again when the incident changes or recurs. +The next move is **not another adapter and not Agent Handoff**. -### Definition of done for the distribution gate -- publish a GitHub v0.6.0 release only after main CI is green, -- README first screen explains capture + compare in one glance, -- repository About/topics reflect GitHub Actions, OpenTelemetry and regression comparison, -- provide one copy/paste install path and one copy/paste compare path, -- show reproducible local, historical, live GitHub, OTLP and compare proof without fabricated metrics, -- ask for real issue/discussion feedback around failed CI and production incident handoff, -- record only observed stars/forks/issues/downloads; no vanity projections, -- do not begin v0.7 solely because v0.6 is merged. +The market re-audit on 2026-09-02 found direct overlap: -### Adoption question -**Would someone keep a FixBundle artifact because comparing it with the next incident saves time?** +- DebugBundle already sells deterministic agent-ready debug bundles for production incidents, +- GitHub Agentic Workflows already investigates failed CI runs and exposes run-diff/audit flows, +- TestSprite already ships failure bundles plus two-run `test diff`, +- support-bundle and temporal-debugging tools already cover adjacent pieces. -If the answer is not demonstrated, improve packaging, docs, discovery and workflow friction before adding more sources. +That means the old question “can we build this?” is closed. The new question is harder: -### Candidate after distribution proof -v0.7 Agent Handoff can add Codex / Claude Code / Cursor export profiles, but only if users need tool-specific handoff beyond the common portable evidence contract. +> **Will unrelated users repeatedly choose portable, cross-source evidence comparison instead of staying inside the platform that already owns the failure?** + +See [`VALIDATION.md`](VALIDATION.md) for the falsification protocol. + +## Current candidate positioning + +Do not lead with “AI-ready bundle.” That phrase is already occupied. + +Candidate result: + +> **Capture a known-good state and a failing state. Prove what changed in the evidence before anyone guesses why.** + +This is still a hypothesis. Do not harden it into branding until field cases support it. + +## Gate sequence + +```text +v0.6.0 public release + ↓ +competitive reality audit + ↓ +10 qualified evidence-first field cases + ↓ +≤5 contextual maintainer outreaches + ↓ +first unrelated real capture + ↓ +artifact intentionally retained + ↓ +same user returns with incident #2 + ↓ +compare used on retained artifacts + ↓ +user states concrete value + ↓ +ONLY THEN decide whether v0.7 exists +``` + +## What counts as progress now + +- finding a public failure where evidence comparison changes the suspect window, +- proving an alleged “same failure” is actually a different failed job/step, +- proving a suspected commit already existed in a green run, +- finding environment/runtime/dependency drift missed by source-only diagnosis, +- identifying a specific missing evidence class that blocks causality, +- learning why a user would or would not keep the artifact, +- learning why GitHub/TestSprite/DebugBundle/native observability is or is not enough. + +## What does not count + +- another version number, +- another source adapter, +- a prettier README by itself, +- our own demo bundles, +- our own CI using FixBundle, +- stars without real use, +- a one-off compliment, +- AI-generated root-cause prose with no external behavioral signal. + +## Immediate field case + +`tanbamboo/rusql#159` currently qualifies as User #1 candidate. + +Read-only evidence already narrowed the timeline from several suspected merges to the first verified `mysql-diff` red at the M63 merge, while two earlier overall-red runs actually had `mysql-diff` success and failed on Rust formatting. The harness also drops server child stdout/stderr, so panic-vs-lifecycle causality cannot be proven from the current CI artifact. + +That is the kind of contribution FixBundle must repeatedly produce before asking users to install it. + +## Decision after the gate + +Possible outcomes: + +### CONTINUE + +Repeat-use proves that portable cross-source compare is a real job. Only then choose the smallest next feature demanded by evidence. + +### NARROW + +Users like the evidence diff but want it embedded in CI/support tooling. Convert the value into a library, GitHub Action or protocol instead of growing a standalone CLI platform. + +### PIVOT + +The repeated pain is real, but the winning job is adjacent: support-bundle verification, temporal incident linking, or another evidence workflow discovered in field cases. + +### ARCHIVE + +The platform-native tools are good enough and users will not retain/compare artifacts. Stop spending engineering time. + +The product is allowed to die. That is a successful validation outcome if learned early. From 8343c86fe8d6e37a167b5ca3518d045e4f1e831a Mon Sep 17 00:00:00 2001 From: yaaertu Date: Wed, 2 Sep 2026 08:10:51 +0300 Subject: [PATCH 4/6] docs: replace vanity gates with activation and retention metrics --- docs/product/METRICS.md | 160 +++++++++++++++++++++++++++++++++++----- 1 file changed, 143 insertions(+), 17 deletions(-) diff --git a/docs/product/METRICS.md b/docs/product/METRICS.md index 60cf6a0..830ce0d 100644 --- a/docs/product/METRICS.md +++ b/docs/product/METRICS.md @@ -1,28 +1,154 @@ # Adoption scoreboard -FixBundle does not use vanity claims before there are users. This file records the public adoption state and the next evidence gates. +FixBundle does not use vanity claims before there are users. This file records the public adoption state and the evidence gates that can justify continuing the product. + +> Last checked: 2026-09-02 + +## Public baseline -## Baseline — 2026-09-02 - GitHub stars: 0 - forks: 0 -- open issues: 1 (maintainer roadmap issue) -- public release/tag: not yet published +- public repository: yes +- public v0.6.0 release: published +- open issues: 2 - PyPI: not yet published - unrelated-user feedback: none yet +- unrelated real captures: 0 +- retained external baselines: 0 +- repeat-use compares: 0 + +## Why the scoreboard changed + +The earlier scoreboard treated stars and install counts as early proof gates. The 2026-09-02 market re-audit found that major parts of the workflow already exist in DebugBundle, GitHub Agentic Workflows, TestSprite and support-bundle ecosystems. + +Therefore the primary question is no longer “can we attract attention?” It is: + +> **Does portable cross-source evidence comparison create behavior that platform-native tools do not already satisfy?** + +Stars remain useful distribution telemetry, but they are not product validation. + +## Stage A — Problem/evidence validation + +Run 10 qualified public field cases under `VALIDATION.md`. + +Record for each case: + +- did the evidence comparison add a material new fact, +- did it narrow the regression window, +- did it distinguish different failure identities hidden behind overall-red status, +- did it reveal environment/runtime/dependency drift, +- did it identify a concrete missing evidence class, +- was the finding already explicit in the issue. + +Working interpretation: + +- **STRONG:** ≥5/10 material findings +- **WEAK:** 3–4/10 +- **FAIL:** <3/10 + +These are product-learning thresholds, not statistical claims. + +## Stage B — External response + +Cap the first outreach set at 5 high-quality, evidence-first maintainer contacts. + +Measure: + +- useful response, +- correction accepted/disputed, +- investigation direction changed, +- maintainer asks how evidence was produced, +- maintainer attempts install/capture. + +Warning signal: 0 useful responses or capture attempts after 5 genuinely qualified contacts. + +## Stage C — Activation + +Activation requires an unrelated person generating a valid FixBundle from their own real failure. + +Target learning state: 3 unrelated real captures. + +For each activation record: -## First proof gates -1. **10 unrelated users** who run the tool or provide concrete feedback. -2. **25 GitHub stars** without paid/incentivized starring. -3. **3 real external bug handoffs** where users say what evidence was useful or missing. -4. **1 integration request** repeated by more than one unrelated user. -5. **1 second-use signal** from someone who uses FixBundle on another incident. +- install path, +- install → first valid bundle elapsed time if known, +- capture mode, +- command/token friction, +- privacy/redaction concern, +- artifact size, +- what evidence the user actually inspected. -## Growth gates -Only after first proof: -- publish PyPI package, -- submit to appropriate developer communities/directories, -- add agent/MCP integrations demanded by users, -- test hosted team features. +Do not count stars, clones, our fixtures or our own repositories as activation. + +## Stage D — Retention + +A retained baseline is stronger than an install. + +Count only when the unrelated user intentionally keeps or later references the artifact because they expect another incident/run to be comparable. + +Target: at least 1 real retained external baseline. + +## Stage E — Repeat-use compare + +North-star event: + +```text +external capture #1 + ↓ +artifact retained + ↓ +real incident #2 + ↓ +fixbundle compare + ↓ +concrete user-reported reduction in evidence reconstruction, +uncertainty or tool switching +``` + +Target: 1 unrelated same-user repeat-use compare before v0.7 feature work. + +## Competitive displacement metric + +For every activated user, capture the answer to one practical question: + +> Why was the native platform not enough for this incident? + +Candidate answers may involve: + +- cross-source comparison, +- historical source identity, +- local/offline workflow, +- portable handoff across tools, +- privacy/no instrumentation, +- integrity/checksum requirements. + +If users cannot name a concrete reason, the product boundary is weak. + +## Secondary distribution telemetry + +Track, but do not optimize ahead of retention: + +- GitHub stars, +- forks, +- release downloads when observable, +- issues/discussions from unrelated users, +- external links/mentions, +- repeat visitors or package installs if a trustworthy source becomes available. + +No paid/incentivized starring. No fabricated download numbers. + +## Kill signals + +- fewer than 3 material findings after 10 qualified field cases, +- 0 useful response/capture attempts after 5 evidence-first outreaches, +- no artifact retention after 3 unrelated real captures, +- native tools erase the cross-source/offline wedge, +- users consistently want this only as an embedded feature inside another tool. + +If a kill signal lands, do not answer with more features. Narrow, pivot or archive. ## 200k-star reality check -200k stars is an ambition, not a forecast or KPI we can manufacture. A developer debugging utility starts in a narrower market than a mass-market AI media generator. The route to exceptional reach is to become a broadly reusable **failure evidence standard**, integrate into existing workflows, and make the output useful across agents, CI systems, support teams and production observability tools. + +A huge star count is not a plan. FixBundle is in a narrower and increasingly competitive developer-tool market. Exceptional distribution would require a job that is both broadly repeated and dramatically clearer than platform-native alternatives. + +Until repeat-use exists, the only honest KPI is **learning rate per real external failure**. From ccc51dff9253901fdc758d7636b67f84570ffbec Mon Sep 17 00:00:00 2001 From: yaaertu Date: Wed, 2 Sep 2026 08:12:33 +0300 Subject: [PATCH 5/6] docs: sync audited repo-home metadata --- docs/product/REPO_HOME.md | 15 +++++++++++---- 1 file changed, 11 insertions(+), 4 deletions(-) diff --git a/docs/product/REPO_HOME.md b/docs/product/REPO_HOME.md index c44d3b2..368f76d 100644 --- a/docs/product/REPO_HOME.md +++ b/docs/product/REPO_HOME.md @@ -43,10 +43,17 @@ Use a 1280×640 visual that says: Do not include fabricated star/download counters. ## Last audited live state — 2026-09-02 -GitHub API previously confirmed: -- description: `Package a broken repo, failed command, or historical Git commit into a redacted AI-ready debugging bundle.` -- topics: `ai-coding-assistant`, `ai-debugging`, `bug-report`, `claude-code`, `codex`, `cursor`, `developer-tools`, `devtools`, `diagnostics`, `git`, `llm`, `production-debugging`, `reproducibility`, `support-bundle`, `temporal-debugging` +GitHub API confirmed: +- description: `Package failures into redacted evidence bundles and compare what changed across local, CI, and OpenTelemetry incidents.` - stars: 0 - forks: 0 +- live topics: `ai-coding-assistant`, `ai-debugging`, `bug-report`, `claude-code`, `codex`, `cursor`, `developer-tools`, `devtools`, `diagnostics`, `github-action`, `github-actions`, `incident-response`, `llm`, `observability`, `opentelemetry`, `production-debugging`, `regression-testing`, `reproducibility`, `support-bundle`, `temporal-debugging` -v0.5 added OpenTelemetry and v0.6 adds evidence comparison, so target metadata is ahead of the last confirmed live About state. The connected repository tools currently do not expose About/Topics mutation. Never claim target metadata is live until GitHub API confirms it after the maintainer updates the repository UI. +### Metadata delta +The About description is now synchronized with the v0.6 target. + +The live topic set is one slot away from the documented target: +- remove obsolete singular topic: `github-action` +- add missing target topic: `git` + +The connected repository tools currently do not expose About/Topics mutation. Never claim that topic delta is fixed until GitHub API confirms the live state after a maintainer UI update. From 6c1517a0c34bc3630331c6a2081926740ea1fdc7 Mon Sep 17 00:00:00 2001 From: yaaertu Date: Wed, 2 Sep 2026 08:12:45 +0300 Subject: [PATCH 6/6] docs: correct live issue and PR counts --- docs/product/METRICS.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/docs/product/METRICS.md b/docs/product/METRICS.md index 830ce0d..fb14ae0 100644 --- a/docs/product/METRICS.md +++ b/docs/product/METRICS.md @@ -10,7 +10,8 @@ FixBundle does not use vanity claims before there are users. This file records t - forks: 0 - public repository: yes - public v0.6.0 release: published -- open issues: 2 +- open non-PR issues: 1 (`#10`, maintainer distribution gate) +- open pull requests: 1 (`#11`, docs-only market validation gate) - PyPI: not yet published - unrelated-user feedback: none yet - unrelated real captures: 0