Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,16 @@ shadow knowledge bases for any codebase.

---

## Unreleased

### Changed
- **More concise documentation** — consolidated README onboarding and workflow
guidance, with advanced operations linked to the skill references. Condensed
repeated guidance and examples in the core, Dream, Init, Meditate, Update,
and Viewer skills while preserving data formats, policy limits, and safety gates.

---

## 2026-09-22

### Added
Expand Down
610 changes: 172 additions & 438 deletions README.md

Large diffs are not rendered by default.

140 changes: 33 additions & 107 deletions skills/shadow-frog-dream/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -369,9 +369,6 @@ dream reports** in `_dreams/`.

### Build Exploration Coverage Map

File-level coverage breadth is the strongest predictor of dream success
(r²=0.63 vs bugs found), NOT dream count (r²=0.04).

Run the packet's `commands.coverage`. It already contains the pinned
interpreter/tool version and the original repository path.

Expand All @@ -391,8 +388,7 @@ For several, append repeated `--scope` argument pairs.
In broad mode, when using `--scope`, the per-category task quotas (Phase 2) still apply
but are interpreted against the scoped subset. Don't use scoped
exploration as the default — pick it only when there's a concrete reason
to concentrate effort. Unscoped diversity remains the strongest
predictor of useful discoveries.
to concentrate effort.

### Review Past Dreams (Required)

Expand Down Expand Up @@ -432,23 +428,14 @@ codebases (<30 source files), minimum 6 tasks across 4+ categories.

### The 6 Investigation Categories

| Category | What to look for | Priority signals |
|----------|-----------------|-----------------|
| **Investigation** | Under-explored files, shallow coverage, uncertain discoveries | Files with 0-2 discoveries, import chains not traced, `uncertain` entries |
| **Bug hunting** | Defects, edge cases, race conditions | Error-handling code, concurrency, unvalidated inputs |
| **Feature design** | New capabilities, missing functionality | TODOs, FIXMEs, user-facing gaps, integration opportunities |
| **Refactoring** | Structural improvements, duplication | God classes, copy-paste patterns, high-coupling files |
| **Optimization** | Algorithmic efficiency, performance | Hot paths, nested loops, repeated I/O, missing caches |
| **Security audit** | Vulnerabilities, unsafe patterns | Auth code, data handling, deserialization, user inputs |

| Category | What to experiment |
|----------|-------------------|
| **Investigation** | Write assertion-based tests proving/disproving behavior hypotheses |
| **Bug hunting** | Fuzz inputs, trigger error paths, reproduce race conditions |
| **Feature design** | Implement the feature, run it, evaluate integration |
| **Refactoring** | Do the refactor, run existing tests, measure complexity |
| **Optimization** | Benchmark, profile, implement optimization, measure before/after |
| **Security audit** | Craft adversarial inputs, test injection vectors (local only) |
| Category | Priority signals | Experiment |
|----------|------------------|------------|
| **Investigation** | Files with 0-2 discoveries, untraced import chains, `uncertain` entries | Test behavior hypotheses with assertions |
| **Bug hunting** | Error handling, concurrency, unvalidated inputs | Fuzz inputs, trigger error paths, reproduce races |
| **Feature design** | TODOs, FIXMEs, user-facing gaps, integration opportunities | Implement the feature, run it, evaluate integration |
| **Refactoring** | God classes, duplication, high coupling | Refactor, run existing tests, measure complexity |
| **Optimization** | Hot paths, nested loops, repeated I/O, missing caches | Profile, optimize, measure before/after |
| **Security audit** | Auth, data handling, deserialization, user inputs | Test adversarial inputs and injection vectors locally |

**Exception — user-directed focus**: If the user specifies a focus area
(e.g., "dream focus on security"), allocate ALL tasks to that category.
Expand All @@ -474,7 +461,8 @@ Tasks (by category):

### Diversity Rules

In broad mode, prevent fixation (exploring the same files while leaving most untouched):
In broad mode, prefer breadth over depth: more files with 2-3 discoveries over
one file with 20. Prevent fixation with these rules:

1. **Max 2 tasks per source file** (unless prior dream left concrete follow-up)
2. **≥30% of tasks on uncovered files** (from coverage map)
Expand All @@ -498,22 +486,8 @@ Also include the selected **mode**. Coherent tasks additionally carry their
own **goal** and **parent_connection**, not a shared sibling/tree objective.
Category still describes the experiment, but coherent mode has no category quota.

Good examples (one per category):
- **Investigation**: "Write assertion harness for request lifecycle —
instrument each layer to log entry/exit and reveal implicit contracts"
- **Bug hunting**: "Fuzz the CSV parser with malformed inputs — what
crashes or silently corrupts?"
- **Feature design**: "Implement retry logic with exponential backoff —
does it handle transient failures without masking permanent ones?"
- **Refactoring**: "Extract 5 duplicate auth checks into middleware —
run tests, measure if it simplifies without breaking special cases"
- **Optimization**: "Benchmark the hot path, implement LRU cache for
repeated lookups — measure before/after wall time"
- **Security audit**: "Craft SQL injection payloads for user-facing
endpoints — does parameterized query hold under nested quotes?"

Bad examples: "Look at the code", "Trace the flow", "Review error
handling", "Improve code quality"
For example, "Fuzz the CSV parser with malformed inputs to reproduce crashes
or silent corruption" is a concrete experiment; "Review error handling" is not.

### Feature Design: Motivation Required

Expand Down Expand Up @@ -561,14 +535,10 @@ but the branch persists on the remote.

### Reading Before Implementing

You must understand the code before changing it. For each task:

1. Read the source file(s) and their shadows (existing discoveries)
2. Read shadows of referenced/referencing files
3. Understand the current behavior, edge cases, and implicit contracts

Reading is *preparation*, not the deliverable. The deliverable is code
written, code run, results recorded.
Before implementing, read the target source and its shadow, plus shadows of
referenced/referencing files. Identify current behavior, edge cases, and
implicit contracts. Reading is preparation; the deliverable is code written
and run, with results recorded.

### Experiment Setup

Expand Down Expand Up @@ -623,10 +593,8 @@ or design, rather than being unrelated work parked on the same branch.

### Run

1. Implement the experiment — write real code, run tests/builds
2. Note what worked, broke, surprised
3. Debug if needed — the struggle produces the best discoveries
4. Record results as you go
Implement the experiment, run tests/builds, and debug as needed. Record what
worked, failed, or surprised you as you go.

### Write Shadow Discoveries

Expand Down Expand Up @@ -655,15 +623,7 @@ Find the `##`/`###` heading for the symbol, then:

After writing each discovery, evaluate whether it deserves any of the
five actionable labels from `/shadow-frog` (`bug`, `security`,
`performance`, `feature-gap`, `tech-debt`). Apply labels when:

| Label | Apply when the discovery describes... |
|-------|---------------------------------------|
| `bug` | A defect, silent failure, off-by-one, race, incorrect result, edge case that misbehaves, validate-then-use ordering hazard |
| `security` | Injection vector, unsafe default, missing auth/authz check, sensitive value logged, untrusted input reaching unsafe sink |
| `performance` | Measured bottleneck, O(N²) where N is large, repeated I/O that could batch, missing cache, blocking call on hot path |
| `feature-gap` | Missing capability the codebase clearly needs, asymmetric API (e.g., reads but no writes) |
| `tech-debt` | Duplication, dead code, leaky abstraction, vestigial parameter, inconsistent naming |
`performance`, `feature-gap`, `tech-debt`), using that skill's definitions.

Rules:
- Apply labels to BOTH the in-file discovery markdown AND the
Expand All @@ -677,7 +637,7 @@ Rules:
- Do not apply labels speculatively. The label says "an engineer should
act on this." If you wouldn't act on it, don't label it.

Examples:
Example:

```
- /api/upload accepts paths from request body without normalization,
Expand All @@ -686,13 +646,6 @@ Examples:
Dream report: `_dreams/20260518-161200Z-upload-traversal/`
```

```
- HttpClient.send retries 3x on transient failures.
_(verified, source: exploration)_
Dream report: `_dreams/20260518-163000Z-retry-audit/`
```
(No label — pure behavioral knowledge, no action implied.)

`dream-validate.py` emits non-blocking warnings when discovery text
contains label-signal keywords but no label is set. Treat those
warnings as a prompt to re-check the triage, not as a directive.
Expand All @@ -715,26 +668,14 @@ back-pointers in each file's `## Cross-References` section.
#### Discovery Quality

Discoveries must be **self-contained process knowledge** — understandable
with just the base codebase. Someone reading main's shadow should understand
the insight without checking out the dream branch. Capture **how to do it**,
**what you learned**, and **what to avoid** — not what was built. The branch
preserves the artifact; the shadow preserves the wisdom.

Good (behavioral insights about existing code):
- "functools.lru_cache is not thread-safe for initialization — two
threads can trigger duplicate expensive computations on first call."
- "agent.py's retry loop catches all exceptions including OOM, masking
fatal errors that should crash immediately."
- "To add a new eval metric, register in METRIC_MAP at metrics.py:25
and implement the Metric interface — missing either causes a silent
no-op in the pipeline."

Bad (descriptions of new code):
- "The implemented PluginFramework has PluginRegistry, PluginManager,
and 7 lifecycle hooks." — describes branch-only artifact.
- "Provides RewardShaper with 4 methods, Welford normalizer, and
GAE-lambda estimation." — feature spec, not behavioral insight.
- "Complete tested module with 58 passing tests." — verdict, not discovery.
from the base codebase without checking out the dream branch. Capture what
was learned, how to apply it, and what to avoid, not a description of the
experiment-only artifact.

- **Good:** "agent.py's retry loop catches all exceptions including OOM,
masking fatal errors that should crash immediately."
- **Bad:** "The implemented PluginFramework has PluginRegistry, PluginManager,
and 7 lifecycle hooks." This describes a branch-only artifact.

Per-file discoveries should reference the dream report:

Expand Down Expand Up @@ -920,11 +861,9 @@ bash "$SKILL_DIR/dream-cleanup.sh" "$WORKTREE_DIR" --repo-root "$REPO_ROOT"
```

`dream-cleanup.sh` does the equivalent of `git worktree remove --force`
followed by `git worktree prune`, but ALSO falls back to a safety-gated
`rm -rf` if `git worktree remove` silently fails — the failure mode that
leaked tens of dream worktrees per AFK session under the previous inline
snippet (see bug-worktree-leak.md). The rm fallback ONLY fires for paths
that match `${DREAM_WORKTREE_BASE:-/tmp/shadowfrog-dreams}/<ns>/dream-<slug>`
followed by `git worktree prune`, with a safety-gated `rm -rf` fallback if
worktree removal silently fails. The fallback only accepts paths matching
`${DREAM_WORKTREE_BASE:-/tmp/shadowfrog-dreams}/<ns>/dream-<slug>`
exactly; any other path is refused.

Remove as you go. If push failed, keep the worktree.
Expand Down Expand Up @@ -1145,9 +1084,7 @@ are four places they get cleaned up:
3. **`dream-gc.sh` (auto-triggered)** — `dream-setup.sh` invokes this
sweeper at the start of each new dream, throttled by a per-namespace
`.last-gc` tombstone to run at most once per `DREAM_GC_INTERVAL_MIN`
minutes (default 60). Catches orphans from crashed dreams, machine
reboots, OOM-killed agents — the long tail of cleanup failures that
accumulated GBs of leaked worktrees on long-running fleets.
minutes (default 60). Catches orphaned worktrees.

Env knobs (all optional, sensible defaults):
- `DREAM_GC_AUTO=0` — disable the auto-trigger entirely
Expand Down Expand Up @@ -1203,17 +1140,6 @@ git cherry-pick <tip_commit>
git show <tip_commit> | git apply
```

## Guidance

- **Always experiment.** Implementation reveals what reading cannot.
- **Small tasks, big lessons.** 30-minute experiment > 3 hours reading.
- **Fail forward.** "Tried X, broke because Y" is extremely valuable.
- **Broad mode: breadth over depth.** More files with 2-3 discoveries > one file with 20.
- **No descriptions.** "Catches all exceptions including OOM" yes.
"This function authenticates users" no.
- **Compound deliberately.** Read parent's report. Build on findings.
- **Branches are cheap, shadow is expensive.** Push freely, write carefully.

## Experiment Completion Criteria

A task is complete ONLY when ALL of these hold:
Expand Down
57 changes: 20 additions & 37 deletions skills/shadow-frog-init/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,19 +26,13 @@ python3 .github/skills/shadow-frog-init/shadow-init.py [options]
python3 .claude/skills/shadow-frog-init/shadow-init.py [options]
```

**IMPORTANT: Run from the repo/worktree root directory.** The script
auto-detects the root via `git rev-parse --show-toplevel`, which returns
the correct root for regular repos AND worktrees. If auto-detection fails
(common when python is routed through Docker or the `.git` file points to
an inaccessible path), pass `--root` explicitly:
**Run from the repo/worktree root.** The helper auto-detects it with
`git rev-parse --show-toplevel`. If Git cannot resolve or access the root
(for example, inside a container), use `--root` to bypass detection:

```bash
# If auto-detect fails, pass the root explicitly:
python3 .github/skills/shadow-frog-init/shadow-init.py --root "$(pwd)"

# In Docker wrapper scenarios (eval harness), git may not work inside
# the container. Use --root to bypass git detection:
python3 .github/skills/shadow-frog-init/shadow-init.py --root /testbed
```

### Options
Expand All @@ -51,32 +45,18 @@ python3 .github/skills/shadow-frog-init/shadow-init.py --root /testbed

### What it does

1. Discovers source files via `git ls-files`
2. Filters through `.shadow/.shadowignore` (gitignore syntax)
3. Extracts symbols from each file (classes, functions, methods)
4. Creates per-file shadow `.md` files with symbol headings
5. Creates `_index.md`, `_prefs.md`, `_meta/state.json`, `.shadowignore`
6. Reports: files, symbols, languages detected

### After running

Tell the user:
- "Edit `.shadow/.shadowignore` to exclude files that shouldn't be shadowed"
- "Run `/shadow-frog-dream` for autonomous exploration, or `/shadow-frog-update` after your next changes"
The helper discovers sources with `git ls-files`, applies
`.shadow/.shadowignore`, extracts symbols, and creates symbol-organized
per-file shadows plus `_index.md`, `_prefs.md`, `_meta/state.json`, and
`.shadowignore`. It reports file/symbol counts and detected languages.

Then decide the version-control mode (do not skip this — it is not handled
by the script). Ask the user: "Should `.shadow/` be **committed** (shared
with your team via git) or **gitignored** (local to your machine only)?"
Make the trade-off explicit before they choose — see
[Step 9: Handle .gitignore](#9-handle-gitignore) for the full committed vs
gitignored comparison. Key caveat: a gitignored `.shadow/` disables
`shadow-frog-dream` (dreams move `.shadow/` through git). If gitignored, add
`.shadow/` to `.gitignore`.
After creating `.shadow/`, complete [Post-init steps](#post-init-steps).
The helper does not choose the version-control mode.

## Fallback: Manual Init

If the Python script fails (wrong Python version, missing file, etc.),
follow these steps manually:
follow these steps manually, then complete [Post-init steps](#post-init-steps).

### 1. Check preconditions

Expand Down Expand Up @@ -156,10 +136,6 @@ out/
.claude/hooks/scripts/shadow-frog-*
```

Tell the user: "Edit `.shadow/.shadowignore` to exclude files or
folders that shouldn't be shadowed (e.g., vendored code, generated
files, tool configs)."

### 5. Create `_prefs.md`

```markdown
Expand Down Expand Up @@ -236,7 +212,11 @@ Rules:
| src/auth.py | Python | 5 (UserAuth, authenticate_user, ...) | 0 |
```

### 9. Handle .gitignore
## Post-init Steps

After creating `.shadow/` with either path:

### Version-Control Mode

Ask the user: "Should `.shadow/` be **committed** (shared with your team via
git) or **gitignored** (local to your machine only)?"
Expand All @@ -261,7 +241,10 @@ Before they decide, make the trade-off explicit:

- If gitignored: add `.shadow/` to `.gitignore`.

### 10. Report
### Report

Print: files discovered, languages detected, total symbols.
Suggest: `/shadow-frog-update` for deeper analysis, `/shadow-frog-dream` for autonomous exploration.
Tell the user to edit `.shadow/.shadowignore` to exclude unwanted paths,
such as vendored/generated files or tool configs. Suggest `/shadow-frog-update`
for deeper analysis or after code changes, and `/shadow-frog-dream` for
autonomous exploration if `.shadow/` is committed.
Loading
Loading