Skip to content

Build v0 living API-evolution + frontier freshness benchmark #24

Description

@in-c0

Owner direction

UpdAPI is evolving from a maintained public API-documentation index into a living API-evolution observatory and frontier coding-agent freshness benchmark.

The original index remains source-acquisition/provenance infrastructure. It is no longer the product thesis.

Implementation lead: Fable 5 Max / Claude Code. ChatGPT is product/research-methodology reviewer. Work should proceed autonomously until a genuine owner-only decision is required.

Do not use GitHub Actions. Every collector, validator, benchmark runner and publisher must have an explicit local/portable entry point so XUXI can later supervise the same semantics.

Current resume point — authoritative

Last completed implementation: PR #26 merged to main as d50bcdb (fix head bf4fc2a).

Completed:

  • direction reset + benchmark spec + no-Actions architecture (Direction reset: living API freshness observatory #25);
  • versioned schemas and 6 verified API-change events;
  • 3 executable cases with stale/current controls and discriminating deterministic validators;
  • local control harness;
  • one Claude Code adapter;
  • fresh out-of-repo workspaces, deterministic npm ci setup, provenance hashes, normalized artifacts;
  • attempt_status = scored | invalid, infrastructure-error normalization, symmetric validator-verdict contract, timeout semantics and dirty-tree provenance;
  • one Claude MCP smoke run, explicitly development-only / host-context, never a leaderboard datum.

Exact next action

Round 4: prove a standardized contamination-resistant Claude condition before invoking another benchmark agent.

Order:

  1. establish OS-enforced/fail-closed isolation (preferred on this Windows host: WSL2 + Claude sandbox; container/VM acceptable if empirically stronger);
  2. suppress/attribute inherited Claude context (CLAUDE.md, memory, hooks, plugins, skills, MCP, user config);
  3. automate a randomized forbidden-answer canary through both built-in Read and shell/subprocess channels, with allowed-read positive controls; pass twice with fresh nonces;
  4. if isolation is unavailable or a forbidden token is observable, stop — apparatus is not standardized;
  5. implement repeated-trial orchestration and immutable cohort-plan records;
  6. prove orchestration with controls;
  7. pre-register and execute exactly 5 scored Claude attempts on case-mcp-modern-era-negotiation-v2, retaining all invalid attempts and allowing at most 2 apparatus-invalid replacements;
  8. review the resulting cohort before adding adapter Required improvements in fast-scrapper.js #2 or expanding cases.

Do not invoke another real benchmark agent before steps 1–3 pass. If a clean WSL2 Claude environment requires a one-time authentication action that cannot be established safely, that is an owner-only escalation; do not weaken isolation to avoid it.

Read first

Core research question

When an external software interface changes, can a frontier coding agent produce a verified working implementation using the tools it normally has available — and how quickly after the change does it become reliable?

Remaining v0 sequence

  1. Standardized agent isolation + contamination canaries.
  2. Repeated-trial orchestration and normalized aggregation.
  3. Two additional materially different agent adapters.
  4. Expand to the first credible-release event/case/provider counts.
  5. Only then automate change discovery/ingestion.
  6. Do not build the public dashboard until credible real results exist.

Research / benchmark invariants

  • Outcome verification > prose judging.
  • Ground truth must be supported by executable behavior, versioned source/spec, or authoritative release evidence.
  • Discovery and ground-truth acceptance remain separate.
  • Published benchmark versions/results are immutable; corrections are explicit revisions.
  • Record exact agent/product/model/version/tool permissions/environment for every run.
  • Default leaderboard conditions should preserve normal useful agent tools while denying benchmark-answer leakage.
  • Apparatus invalid attempts are excluded from capability scores but remain visible; ordinary scored failures are never retried away.
  • Stale-API failure classification requires case-specific evidence, not generic failure inference.
  • Prefer recent temporal holdouts and preserve publication/observation/verification timestamps.
  • Multiple cases from one API change do not multiply that event's headline weight.
  • Do not optimize methodology to prove UpdAPI retrieval helps. A null result is valid.

First credible release gate

Before calling this a benchmark release:

  • = 30 verified change events;

  • = 5 provider/ecosystem families;

  • = 20 executable cases;

  • every executable case has positive + stale negative controls with demonstrated discrimination;
  • = 3 materially different frontier coding systems;

  • contamination-resistant standardized conditions;
  • normalized manifests/results + immutable cohort plans;
  • repeated-trial policy fixed before public results;
  • at least one controlled retrieval experiment;
  • machine-readable downloadable results;
  • public methodology + benchmark version history;
  • no GitHub Actions dependency.

Product direction after evidence exists

Publishing should resemble an independent measurement product (Artificial Analysis / Arena philosophy): leaderboard, newest change cohort, per-system/provider/event drilldowns, longitudinal adoption trends, methodology/version history and downloadable evidence.

The eventual headline metric should be a Balanced API Freshness view that aggregates attempts → cases → API change events → provider/ecosystem families, preventing one richly-authored release or provider from dominating the score. Raw success rates and category slices remain visible.

The durable asset is trust in the measurement.

Collaboration protocol

Fable implements aggressively and leaves evidence in commits/PRs/tests. ChatGPT challenges methodology, confounds, metric definitions, ground-truth quality, integrity boundaries and unnecessary complexity. Either side revises when the other's evidence is stronger.

Escalate to the owner only for genuine owner decisions, irreversible/high-risk actions, unavailable credentials/physical actions, or unresolved strategic tradeoffs that materially change the project thesis.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions