docs(blogging): re-cut the blog queue around a broadcast filter - #120
Conversation
Flip `02` to Live (merged 2026-07-16 — the doc still advertised it as in review five weeks on), then reorder what ships next. Introduces an explicit selection filter — does the lesson transfer to a reader who will never use this harness? — and applies it. `03` (an internal failure taxonomy addressed by its own labels), `04` (re-argues `00`, harness-shaped framing) and `05` (U2 architecture-stance, and its gate was "earned by the empirics beneath it") are deferred rather than dropped; each carries its rationale in the status board. `04`'s ADR-0007 argument is reassigned to a section of `06`. The queue is now led by three full-kit backlog items plus the deterministic-grader recipe, all of which carry a lesson portable off this harness, on a two-week cadence through 2026-10-10. Also records that `increment-4-plan.md` — the flagship's gate — is gone from the repo root, and re-verifies every Evidence index path.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: cb99ca0981
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| **Snapshot:** 2 live · 1 in review · 4 stubs · 12 backlog. The stub dates for `03` and `04` | ||
| (2026-06-30, 2026-07-08) have already lapsed — they are re-based off `02`'s merge in the queue | ||
| below, not honoured as written. | ||
| **Snapshot (2026-08-23):** 3 live · 0 in review · 3 deferred · 1 stub (`06`, demo-gated) · 16 backlog. |
There was a problem hiding this comment.
Update this snapshot to 17 backlog items: the Backlog table contains 17 article rows, and the newly documented protocol explicitly says the four queued rows remain there until they receive directories. Leaving the count at 16 makes the repository's designated source of truth internally inconsistent.
AGENTS.md reference: AGENTS.md:L20-L20
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 3c86272 — you were right, and I had carried the stale figure forward instead of recounting (the pre-change doc said 12). The table has 17 rows; the snapshot now says 17 and states explicitly that the four queued rows stay in Backlog until they get directories.
| | # | Ship | Why here | Gate | Target | | ||
| | --- | --- | --- | --- | --- | | ||
| | **1** | `provider-reliability` → next free number | A provider hung ~5 min and returned a `\x00` body as **HTTP 200** — an error wearing a success costume — and it looked like a code regression until a serial control run (20/20, 0 NULs) exonerated the model. The rule (**classify by layer:** dead/empty = transport, malformed-but-present = model, slow-but-streaming = neither) applies to anyone calling an LLM API. Six-row run table already written. | none — 21KB kit | 2026-08-29 | | ||
| | **2** | `shell-syntax-boundary` → next free number | **The most broadcastable item in the backlog.** Every harness that avoids `shell=True` inherits this: `shlex.split` turns model-authored `&&` into argv words, so a declared 10-section verification chain **passed having verified 1 section** (`grep -q` exits 0 on first match while patterns 2–10 became unopenable filenames). Deterministic one-line repro; E5/P4/U4. | none — kit | 2026-09-12 | |
There was a problem hiding this comment.
Narrow the shell-boundary claim
Limit this claim to harnesses that pass model-authored shell strings through shlex.split without validating shell syntax. A harness can avoid shell=True while accepting structured argv or rejecting metacharacters, in which case it does not inherit this failure; the current “Every harness” wording turns one observed implementation pattern into an unsupported universal claim.
AGENTS.md reference: AGENTS.md:L36-L38
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 3c86272. Fair catch, and worse than a wording slip: the doc I was editing states as standing rule 1 that posts carry directional evidence rather than universal claims, and that line broke it. Narrowed to the actual pattern — a harness that hands model-authored command strings to shlex.split with shell=False and never checks them for shell syntax — and noted that structured argv or metacharacter rejection at the seam avoids it. Also added an instruction to hold the eventual post to that same scope, since the queue row is what the draft gets written from.
| ## Evidence index | ||
|
|
||
| Every path below is verified as of 2026-07-12. Research docs are date-prefixed (commit | ||
| Every path below re-verified 2026-08-23 (all resolve). Research docs are date-prefixed (commit |
There was a problem hiding this comment.
Exclude the missing plan from the all-resolve claim
Revise the assertion that every path resolves, because this same evidence index now records increment-4-plan.md as missing at line 229. As written, the verification summary contradicts the artifact inventory and gives readers a false signal that every receipt needed by the queue is available.
AGENTS.md reference: AGENTS.md:L20-L20
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in 3c86272. The check I ran only ever covered markdown-linked paths, so the summary was true of what it tested and misleading about what it implied. It now says every linked path resolves and names increment-4-plan.md as the exception — cited as a bare filename precisely because there is nothing to link to.
…the path check Addresses three P2 review comments on #120: - The snapshot said 16 backlog items; the table has 17. The stale figure was carried forward rather than recounted. Now counted, with a note that the four queued rows stay in the table until they get directories. - "Every harness that avoids shell=True inherits this" overclaimed from one observed implementation pattern — and violated standing rule 1 in the same doc that states it. Narrowed to harnesses that shlex.split model-authored command strings without checking shell syntax, with a note to hold the post to that scope. - The evidence index claimed every path resolves while its own rows record increment-4-plan.md as missing. The check only ever covered linked paths; the claim now says so and names the exception.
Reviewer quiz — blog queue re-cut4 questions on this change (answers inside)1. 2. Fill in the blank: 3. Which content survives the deferral of 4. The Evidence index now records one artifact as missing. Which one, and what does its absence block? Answers 1. (b) — 2. …no post directory yet. 3. ADR-0007's argument that a model which authors its own rubric sets it low up front ( 4. |
Description
Re-cuts the blog plan of action in
docs/blogging/blog-candidates.md: corrects02's status, introduces an explicit selection filter for what gets written at all, and reorders the queue around it. Docs only — no code, and nothing insarthak-blogis touched.Motivation
Two problems, one stale and one editorial.
Stale:
02(Don't judge an agent by its pass@1) merged tosarthak-blogon 2026-07-16, but this doc still advertised it as In review five weeks later. The doc's own Update protocol says status is read from git state, never from memory — so it was violating its own rule, and everything downstream re-dates off that merge.Editorial: reviewing the queue against the question does this lesson transfer to a reader who will never use this harness?, the three numbered stubs at the front fail it.
03is an internal failure taxonomy addressed by its own labels (A9,B3) — reference material for this repo, not broadcast material.04re-argues00, which already moved the authority for "done" out of the model, and frames it in harness-shaped terms.05is the lowest-uniqueness item in the series (U2) and pure architecture-stance, and its own gate was "earned by the empirical posts beneath it" — with03/04deferred there is less under it, not more. Meanwhile four full-kit backlog items that do carry portable lessons were sitting below them, unqueued.Changes
02→ Live, dated 2026-07-11, merged 2026-07-16 (PR feat(phase-3): async core + typed event bus + two-plane session (3.0 foundation) #7 insarthak-blog).Deferredstatus definition. The existingBacklogdefinition is "no post directory yet", and03/04/05have directories withdraft: true— moving them into that table would have broken the table's own rule.03,04,05→ Deferred, each carrying its rationale in the status board rather than being silently dropped.04's strongest content (ADR-0007's a model that authors its own rubric sets it low up front) is reassigned to a section of06, so it is not lost with the deferral.provider-reliability→shell-syntax-boundary→eval-probe-false-rejections→deterministic-grader→06, with target dates on a two-week cadence (2026-08-29 → 2026-10-10;06stays demo-gated).→ queued #N, so ordering lives in one place.increment-4-plan.mdis missing — it was untracked at the repo root and is gone — so the flagship's gate currently has no artifact behind it, and rebuilding that plan is named as the first task on its critical path.Testing
Docs-only change; no code paths touched, so no test suite applies.
02's live status directly fromsarthak-bloggit state (git log --first-parent main, merge commit9e1bf84, 2026-07-16,draft: false) rather than from the doc.../research/…,blog_kits/…) — all resolve.increment-4-plan.mdis genuinely absent from the repo root before recording it as missing.