You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
UpdAPI is evolving from a maintained public API-documentation index into a living API-evolution observatory and frontier coding-agent freshness benchmark.
The original index remains source-acquisition/provenance infrastructure. It is no longer the product thesis.
Implementation lead: Fable 5 Max / Claude Code. ChatGPT is product/research-methodology reviewer. Work should proceed autonomously until a genuine owner-only decision is required.
Do not use GitHub Actions. Every collector, validator, benchmark runner and publisher must have an explicit local/portable entry point so XUXI can later supervise the same semantics.
Current resume point — authoritative
Last completed implementation: PR #26 merged to main as d50bcdb (fix head bf4fc2a).
one Claude MCP smoke run, explicitly development-only / host-context, never a leaderboard datum.
Exact next action
Round 4: prove a standardized contamination-resistant Claude condition before invoking another benchmark agent.
Order:
establish OS-enforced/fail-closed isolation (preferred on this Windows host: WSL2 + Claude sandbox; container/VM acceptable if empirically stronger);
suppress/attribute inherited Claude context (CLAUDE.md, memory, hooks, plugins, skills, MCP, user config);
automate a randomized forbidden-answer canary through both built-in Read and shell/subprocess channels, with allowed-read positive controls; pass twice with fresh nonces;
if isolation is unavailable or a forbidden token is observable, stop — apparatus is not standardized;
implement repeated-trial orchestration and immutable cohort-plan records;
prove orchestration with controls;
pre-register and execute exactly 5 scored Claude attempts on case-mcp-modern-era-negotiation-v2, retaining all invalid attempts and allowing at most 2 apparatus-invalid replacements;
Do not invoke another real benchmark agent before steps 1–3 pass. If a clean WSL2 Claude environment requires a one-time authentication action that cannot be established safely, that is an owner-only escalation; do not weaken isolation to avoid it.
latest comments on this issue — Round-4 isolation, repeated-trial and publication/scoring contracts
Core research question
When an external software interface changes, can a frontier coding agent produce a verified working implementation using the tools it normally has available — and how quickly after the change does it become reliable?
repeated-trial policy fixed before public results;
at least one controlled retrieval experiment;
machine-readable downloadable results;
public methodology + benchmark version history;
no GitHub Actions dependency.
Product direction after evidence exists
Publishing should resemble an independent measurement product (Artificial Analysis / Arena philosophy): leaderboard, newest change cohort, per-system/provider/event drilldowns, longitudinal adoption trends, methodology/version history and downloadable evidence.
The eventual headline metric should be a Balanced API Freshness view that aggregates attempts → cases → API change events → provider/ecosystem families, preventing one richly-authored release or provider from dominating the score. Raw success rates and category slices remain visible.
The durable asset is trust in the measurement.
Collaboration protocol
Fable implements aggressively and leaves evidence in commits/PRs/tests. ChatGPT challenges methodology, confounds, metric definitions, ground-truth quality, integrity boundaries and unnecessary complexity. Either side revises when the other's evidence is stronger.
Escalate to the owner only for genuine owner decisions, irreversible/high-risk actions, unavailable credentials/physical actions, or unresolved strategic tradeoffs that materially change the project thesis.
Owner direction
UpdAPI is evolving from a maintained public API-documentation index into a living API-evolution observatory and frontier coding-agent freshness benchmark.
The original index remains source-acquisition/provenance infrastructure. It is no longer the product thesis.
Implementation lead: Fable 5 Max / Claude Code. ChatGPT is product/research-methodology reviewer. Work should proceed autonomously until a genuine owner-only decision is required.
Do not use GitHub Actions. Every collector, validator, benchmark runner and publisher must have an explicit local/portable entry point so XUXI can later supervise the same semantics.
Current resume point — authoritative
Last completed implementation: PR #26 merged to
mainasd50bcdb(fix headbf4fc2a).Completed:
npm cisetup, provenance hashes, normalized artifacts;attempt_status = scored | invalid, infrastructure-error normalization, symmetric validator-verdict contract, timeout semantics and dirty-tree provenance;Exact next action
Round 4: prove a standardized contamination-resistant Claude condition before invoking another benchmark agent.
Order:
CLAUDE.md, memory, hooks, plugins, skills, MCP, user config);case-mcp-modern-era-negotiation-v2, retaining all invalid attempts and allowing at most 2 apparatus-invalid replacements;Do not invoke another real benchmark agent before steps 1–3 pass. If a clean WSL2 Claude environment requires a one-time authentication action that cannot be established safely, that is an owner-only escalation; do not weaken isolation to avoid it.
Read first
main— authoritative implementation branchREADME.md— public thesis and roadmapdocs/BENCHMARK_SPEC.md— measurement contractCore research question
Remaining v0 sequence
Research / benchmark invariants
invalidattempts are excluded from capability scores but remain visible; ordinary scored failures are never retried away.First credible release gate
Before calling this a benchmark release:
Product direction after evidence exists
Publishing should resemble an independent measurement product (Artificial Analysis / Arena philosophy): leaderboard, newest change cohort, per-system/provider/event drilldowns, longitudinal adoption trends, methodology/version history and downloadable evidence.
The eventual headline metric should be a Balanced API Freshness view that aggregates attempts → cases → API change events → provider/ecosystem families, preventing one richly-authored release or provider from dominating the score. Raw success rates and category slices remain visible.
The durable asset is trust in the measurement.
Collaboration protocol
Fable implements aggressively and leaves evidence in commits/PRs/tests. ChatGPT challenges methodology, confounds, metric definitions, ground-truth quality, integrity boundaries and unnecessary complexity. Either side revises when the other's evidence is stronger.
Escalate to the owner only for genuine owner decisions, irreversible/high-risk actions, unavailable credentials/physical actions, or unresolved strategic tradeoffs that materially change the project thesis.