Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
136 changes: 119 additions & 17 deletions docs/product/LANDSCAPE.md
Original file line number Diff line number Diff line change
@@ -1,33 +1,135 @@
# Competitive landscape

FixBundle should not pretend the surrounding problems are unsolved. Several strong tools already cover adjacent pieces. The product boundary is deliberately narrower: **portable failure evidence**, not generic repository packing or generic agent memory.
> Last re-audited: 2026-09-02

## Repomix
Repository: https://github.com/yamadashy/repomix
FixBundle must not sell novelty that the market does not support. The surrounding problem is real, but several products now cover major pieces of the same workflow. The product only deserves to continue if users repeatedly choose the **portable, cross-source evidence-diff** job strongly enough to justify a standalone tool.

Repomix is the mature reference for **codebase → LLM-friendly context**. It packs local/remote repositories, counts tokens, supports compression and MCP integrations, and has a substantial user base.
## Market reality

FixBundle should not compete by becoming a worse repo packer.
The broad message **“turn failures into agent-ready debug bundles” is not unique**.

**Different job:** FixBundle packages a *failure event*: command output/exit state, exact Git identity, incident-vs-current revision, environment/stack evidence, selected source/config, redaction and checksums.
The broad message **“let an AI inspect a failed GitHub Actions run” is not unique**.

## temporal-debug-skill
Repository: https://github.com/MeherBhaskar/temporal-debug-skill
The broad message **“compare two runs and show what regressed” is not unique**.

This agent skill independently validates the temporal-debugging pain: production incidents may belong to an older commit, and temporary Git worktrees are safer than checking out over current work.
Therefore FixBundle must not position itself as a generic AI debugging platform, observability product, CI doctor, test platform, or repository packer.

FixBundle v0.3 overlaps on isolated worktree lifecycle. We should acknowledge that overlap instead of claiming novelty.
The only currently defensible wedge is narrower:

**Different job:** the skill teaches an agent *how to inspect history*. FixBundle creates a standalone, inspectable archive that can cross agent/vendor/support boundaries and carries logs, environment, source/config snapshots, redaction and integrity metadata.
> **Capture a known-good and a failing state as portable evidence artifacts, verify their integrity, and deterministically show which evidence changed across local, historical Git, GitHub Actions, and OpenTelemetry sources before anyone makes a causal claim.**

## GitHub Actions + Copilot
GitHub can explain failed workflow checks with Copilot and exposes downloadable workflow logs.
That wedge is still a hypothesis, not a validated market.

**Different job:** FixBundle v0.4 should not merely re-explain the same log. It should normalize a failed CI run into the same portable evidence schema used locally, tie it to commit/diff/config context, sanitize it, and make the archive usable outside GitHub Copilot.
## Direct and adjacent competitors

### DebugBundle

- Product: https://debugbundle.com/
- Repository: https://github.com/debugbundle/debugbundle
- Public repo observed 2026-09-02: created 2026-05-06, AGPL-3.0, TypeScript, 6 stars, 0 forks.
- Positioning: production errors → deterministic agent-ready bundles through SDKs, CLI, API, MCP, dashboard and hosted/self-hosted workflows.
- Capabilities include production event capture, redaction, incident grouping, reproduction artifacts, deploy metadata, probes, analytics and deploy-comparison analysis.

**Overlap:** deterministic bundles, redaction, agent-readable evidence, local-first workflow, production incident context.

**Current difference:** FixBundle does not require application SDK instrumentation or a hosted incident backend. It can package arbitrary local command failures, isolated historical Git commits, GitHub Actions runs and OTLP file exports into portable checksummed ZIPs, then compare two artifacts offline. That difference matters only if users value it enough to retain and reuse artifacts.

**Threat level: HIGH.** If DebugBundle or another product exposes equally low-friction arbitrary cross-source offline artifact comparison, FixBundle loses a major differentiation claim.

### GitHub Agentic Workflows / CI Failure Doctor

- Project: https://github.com/github/gh-aw
- CI failure investigation example: https://github.github.com/gh-aw/gallery/ci-failure-investigation/
- Audit/diff reference: https://github.github.com/gh-aw/reference/audit/

GitHub Agentic Workflows can start an agent after a failed workflow, preload failed jobs/logs/artifacts, correlate failures with repository changes and produce an actionable diagnosis. Its audit tooling can also compare workflow runs for behavioral drift.

**Overlap:** failed-run evidence collection, run comparison, agent consumption.

**Current difference:** GitHub's workflow is GitHub-native and agent-execution-oriented. FixBundle's claim is a portable, vendor-neutral evidence artifact that can leave GitHub and be compared with local or OTLP evidence without an AI being required.

**Threat level: HIGH for GitHub-only use cases.** FixBundle should never become a worse CI Doctor.

### TestSprite CLI

- Repository: https://github.com/TestSprite/testsprite-cli
- Product: https://www.testsprite.com/

TestSprite is a cloud testing platform. Its CLI exposes self-consistent failure bundles and `test diff <run-a> <run-b>`, including verdict changes, failure-kind changes, failed-step shifts, per-step status flips and code-version drift.

**Overlap:** failure bundles, durable run identity, before/after comparison, agent-friendly output.

**Current difference:** TestSprite owns the test execution platform and compares TestSprite test runs. FixBundle does not own execution; it imports evidence from arbitrary local, historical, CI and telemetry sources.

**Threat level: MEDIUM/HIGH.** It independently validates the “last green vs current red” job, but also proves that run-diff can be a feature inside a larger product rather than a standalone business.

### temporal-debug-skill

- Repository: https://github.com/MeherBhaskar/temporal-debug-skill

The skill addresses historical debugging by resolving an incident commit and inspecting it in an isolated Git worktree rather than changing the user's active workspace.

**Overlap:** historical revision safety and temporal debugging.

**Current difference:** the skill teaches an agent how to inspect history. FixBundle emits an inspectable portable artifact with command/log/environment/source evidence, redaction and integrity metadata.

**Threat level: MEDIUM.** Historical Git isolation by itself is not a product moat.

### Repomix

- Repository: https://github.com/yamadashy/repomix

Repomix is the mature reference for repository → LLM-friendly context packaging.

**Overlap:** portable context for AI tools.

**Current difference:** Repomix packages code. FixBundle packages a failure state and its evidence identity.

**Threat level: LOW if FixBundle stays failure-first. HIGH if FixBundle drifts into generic repository packing.**

### Support-bundle ecosystems

Examples include Replicated Troubleshoot and product-specific support bundles. These systems collect diagnostics, redact sensitive data and produce archives for support/debugging.

**Overlap:** bounded diagnostic collection, privacy, archive handoff.

**Current difference:** most support-bundle systems are product/platform specific and do not treat two arbitrary evidence archives as a cross-source temporal comparison contract.

**Threat level: MEDIUM.** “Make a redacted ZIP” is commodity functionality.

## What the market evidence says

Current public discussions repeatedly show engineers comparing last-green/current-red states manually, checking exact failed steps, recent changes, environment drift and historical failure signatures. This validates the problem shape.

It does **not** validate FixBundle as a standalone product. A real market test must show that users:

1. accept the capture friction,
2. intentionally retain an artifact,
3. later produce or obtain a second artifact,
4. run a comparison,
5. say the comparison reduced evidence reconstruction, uncertainty or tool switching,
6. choose this workflow over built-in GitHub/TestSprite/observability alternatives for a concrete reason.

## Product boundary
FixBundle wins only if this stays true:

> **Repomix packages code. Temporal Debug navigates time. Copilot explains a GitHub failure. FixBundle packages the evidence of a failure so any debugger can work from the same facts.**
FixBundle is not:

- an AI root-cause oracle,
- an observability SaaS,
- an AI SRE,
- a cloud testing platform,
- a generic repo packer,
- a replacement for GitHub Actions logs,
- a dashboard that happens to export JSON.

The candidate job is:

> **Prove what changed in failure evidence before anyone guesses why.**

## Kill rule

If external validation cannot prove that artifact portability + cross-source deterministic compare creates repeated value, do not solve the problem by adding more adapters, dashboards, MCP surfaces or AI summaries.

At that point FixBundle should be narrowed into a library/protocol/CI utility, pivoted to a stronger adjacent job, or archived.

If a new feature does not strengthen that result, it probably does not belong in the core.
See `docs/product/VALIDATION.md` for the falsification protocol.
161 changes: 144 additions & 17 deletions docs/product/METRICS.md
Original file line number Diff line number Diff line change
@@ -1,28 +1,155 @@
# Adoption scoreboard

FixBundle does not use vanity claims before there are users. This file records the public adoption state and the next evidence gates.
FixBundle does not use vanity claims before there are users. This file records the public adoption state and the evidence gates that can justify continuing the product.

> Last checked: 2026-09-02

## Public baseline

## Baseline — 2026-09-02
- GitHub stars: 0
- forks: 0
- open issues: 1 (maintainer roadmap issue)
- public release/tag: not yet published
- public repository: yes
- public v0.6.0 release: published
- open non-PR issues: 1 (`#10`, maintainer distribution gate)
- open pull requests: 1 (`#11`, docs-only market validation gate)
- PyPI: not yet published
- unrelated-user feedback: none yet
- unrelated real captures: 0
- retained external baselines: 0
- repeat-use compares: 0

## Why the scoreboard changed

The earlier scoreboard treated stars and install counts as early proof gates. The 2026-09-02 market re-audit found that major parts of the workflow already exist in DebugBundle, GitHub Agentic Workflows, TestSprite and support-bundle ecosystems.

Therefore the primary question is no longer “can we attract attention?” It is:

> **Does portable cross-source evidence comparison create behavior that platform-native tools do not already satisfy?**

Stars remain useful distribution telemetry, but they are not product validation.

## Stage A — Problem/evidence validation

Run 10 qualified public field cases under `VALIDATION.md`.

Record for each case:

- did the evidence comparison add a material new fact,
- did it narrow the regression window,
- did it distinguish different failure identities hidden behind overall-red status,
- did it reveal environment/runtime/dependency drift,
- did it identify a concrete missing evidence class,
- was the finding already explicit in the issue.

Working interpretation:

- **STRONG:** ≥5/10 material findings
- **WEAK:** 3–4/10
- **FAIL:** <3/10

These are product-learning thresholds, not statistical claims.

## Stage B — External response

Cap the first outreach set at 5 high-quality, evidence-first maintainer contacts.

Measure:

- useful response,
- correction accepted/disputed,
- investigation direction changed,
- maintainer asks how evidence was produced,
- maintainer attempts install/capture.

Warning signal: 0 useful responses or capture attempts after 5 genuinely qualified contacts.

## Stage C — Activation

Activation requires an unrelated person generating a valid FixBundle from their own real failure.

Target learning state: 3 unrelated real captures.

For each activation record:

## First proof gates
1. **10 unrelated users** who run the tool or provide concrete feedback.
2. **25 GitHub stars** without paid/incentivized starring.
3. **3 real external bug handoffs** where users say what evidence was useful or missing.
4. **1 integration request** repeated by more than one unrelated user.
5. **1 second-use signal** from someone who uses FixBundle on another incident.
- install path,
- install → first valid bundle elapsed time if known,
- capture mode,
- command/token friction,
- privacy/redaction concern,
- artifact size,
- what evidence the user actually inspected.

## Growth gates
Only after first proof:
- publish PyPI package,
- submit to appropriate developer communities/directories,
- add agent/MCP integrations demanded by users,
- test hosted team features.
Do not count stars, clones, our fixtures or our own repositories as activation.

## Stage D — Retention

A retained baseline is stronger than an install.

Count only when the unrelated user intentionally keeps or later references the artifact because they expect another incident/run to be comparable.

Target: at least 1 real retained external baseline.

## Stage E — Repeat-use compare

North-star event:

```text
external capture #1
↓
artifact retained
↓
real incident #2
↓
fixbundle compare
↓
concrete user-reported reduction in evidence reconstruction,
uncertainty or tool switching
```

Target: 1 unrelated same-user repeat-use compare before v0.7 feature work.

## Competitive displacement metric

For every activated user, capture the answer to one practical question:

> Why was the native platform not enough for this incident?

Candidate answers may involve:

- cross-source comparison,
- historical source identity,
- local/offline workflow,
- portable handoff across tools,
- privacy/no instrumentation,
- integrity/checksum requirements.

If users cannot name a concrete reason, the product boundary is weak.

## Secondary distribution telemetry

Track, but do not optimize ahead of retention:

- GitHub stars,
- forks,
- release downloads when observable,
- issues/discussions from unrelated users,
- external links/mentions,
- repeat visitors or package installs if a trustworthy source becomes available.

No paid/incentivized starring. No fabricated download numbers.

## Kill signals

- fewer than 3 material findings after 10 qualified field cases,
- 0 useful response/capture attempts after 5 evidence-first outreaches,
- no artifact retention after 3 unrelated real captures,
- native tools erase the cross-source/offline wedge,
- users consistently want this only as an embedded feature inside another tool.

If a kill signal lands, do not answer with more features. Narrow, pivot or archive.

## 200k-star reality check
200k stars is an ambition, not a forecast or KPI we can manufacture. A developer debugging utility starts in a narrower market than a mass-market AI media generator. The route to exceptional reach is to become a broadly reusable **failure evidence standard**, integrate into existing workflows, and make the output useful across agents, CI systems, support teams and production observability tools.

A huge star count is not a plan. FixBundle is in a narrower and increasingly competitive developer-tool market. Exceptional distribution would require a job that is both broadly repeated and dramatically clearer than platform-native alternatives.

Until repeat-use exists, the only honest KPI is **learning rate per real external failure**.
Loading
Loading