Skip to content

[2026 AI] Build an evaluation harness and published benchmarks for every shipped ML model #994

Description

@Nanle-code

Objective

Back every model accuracy claim with reproducible evaluation.

Target area

src/ml, src/lib/feePredictor.ts, phishingDetector.ts, liquidityModel.ts

Context

The README claims the fee predictor is '95% within 10% of actual fees' with '< 50ms' latency, but there's no benchmark script or dataset in CI to back it up. Several other models (phishing, liquidity, build, backup) have no stated evaluation.

Acceptance criteria

  • A pnpm ml:eval command runs each model on a held-out, versioned dataset and outputs metrics as JSON and markdown.
  • Metrics are compared to a naive baseline (e.g. median recent fee), and the model must beat it to ship.
  • The README and docs quote only numbers produced by the harness, with date and dataset version.
  • A scheduled CI job re-runs evaluation and posts the results as a workflow summary.
  • Automated tests cover the primary flow, at least one boundary case, and at least one failure case.
  • User-facing documentation or developer guidance is updated, including compatibility, security, or migration notes where applicable.

Pull request requirements

  • All continuous integration checks must pass before a pull request may be merged.
  • The branch must be free of merge conflicts with the target branch at the time of merge.
  • Do not request merge with failing workflows, skipped required checks, or unresolved conflicts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions