diff --git a/CHANGELOG.md b/CHANGELOG.md index 55ebf33..875a339 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,16 @@ shadow knowledge bases for any codebase. --- +## Unreleased + +### Changed +- **More concise documentation** — consolidated README onboarding and workflow + guidance, with advanced operations linked to the skill references. Condensed + repeated guidance and examples in the core, Dream, Init, Meditate, Update, + and Viewer skills while preserving data formats, policy limits, and safety gates. + +--- + ## 2026-09-22 ### Added diff --git a/README.md b/README.md index 96b572a..8b6b37f 100644 --- a/README.md +++ b/README.md @@ -4,12 +4,16 @@ ShadowFrog gives coding agents a **shadow knowledge base** for any codebase: a file-backed memory of tacit codebase knowledge learned from code reading, experiments, and conversations with you. -Most agent memory preserves what happened in past chats. ShadowFrog is built -for **tacit knowledge** that is hard to recover from chat history or source -alone: which refactor breaks downstream callers, which invariant the tests -never exercise, which "obvious" cleanup removes a production workaround, or -which cross-file edge case is easy to miss. The code tells you *what runs*. The -shadow tells future agents *what has been learned about how it behaves*. +A **shadow** mirrors your source tree under `.shadow/`, storing discoveries in +symbol-organized Markdown files. Lookup is **index-free**: agents follow source +paths and `file::symbol` references rather than a vector index or embedding +service. + +It records knowledge that is hard to recover from source alone: which refactor +breaks downstream callers, which invariant the tests never exercise, or which +"obvious" cleanup removes a production workaround. The code tells you *what +runs*. The shadow tells future agents *what has been learned about how it +behaves*. Read the launch blog post: [Shadow-Frog: Coding Agents that Dream and Discover](https://microsoft.github.io/debug-gym/blog/2026/06/shadow-frog/). @@ -22,15 +26,21 @@ Discover](https://microsoft.github.io/debug-gym/blog/2026/06/shadow-frog/). ## Quick Start +Install **per repository**, not globally. You need Git, Python 3, either GitHub +Copilot CLI or Claude Code, and a Git repository as the target project. +Hooks are limited to projects you explicitly opt into. + +### 1. Install + +From a checkout of ShadowFrog, choose one agent: + ```bash -# From your ShadowFrog checkout, install into your target repo. cd /path/to/ShadowFrog -# Choose one: ./install.sh --project /path/to/your-repo # Copilot CLI (default) # ./install.sh --agent claude --project /path/to/your-repo # Claude Code ``` -On **Windows**, use the PowerShell installer instead (no Bash shell needed): +On **Windows**, use the PowerShell installer, which does not require Bash: ```powershell cd C:\path\to\ShadowFrog @@ -38,7 +48,10 @@ cd C:\path\to\ShadowFrog # .\install.ps1 -Agent claude -Project C:\path\to\your-repo # Claude Code ``` -Commit the installed files in the target repo: +### 2. Share the setup + +For a shared setup, commit and push the installed files in the target repo. +The installer prints the exact staging paths for your selected components: ```bash cd /path/to/your-repo @@ -49,453 +62,178 @@ git commit -m "Add ShadowFrog skills, hooks, and context" git push ``` -Then open the target repo in your AI agent session: +Local-only use does not require a writable remote; Dream does. -| Command | Use | -|---------|-----| -| `/shadow-frog-init` | Create the shadow | -| `/shadow-frog-update` | Refresh after code changes | -| `/shadow-frog-dream` | Explore and experiment while you're away | -| `/shadow-frog-nap` | Generate grounded feature-task briefs with a bounded budget | -| `/shadow-frog-meditate` | Deduplicate and resolve conflicts | -| `/shadow-frog-viewer` | Browse what's in the shadow | - -If you plan to use `/shadow-frog-dream`, commit and push `.shadow/` after -init. Dream also needs permission to push `dream/...` branches to the repo's -remote. - -> See [Installation](#installation) for full details. - ---- +### 3. Initialize -## What is a Shadow? - -A shadow is a `.shadow/` directory that mirrors your source tree with markdown -files. It is not generated API documentation and it is not a transcript store. -It stores **discoveries**: behavioral facts, edge cases, implicit contracts, -warnings, and cross-file interactions that are useful to future agents. - -Shadow lookup is **index-free**: it follows the source tree instead of a -separate vector index. If an agent is editing `src/auth.py`, the corresponding -knowledge lives at `.shadow/src/auth.py.md`; if it is reasoning about -`src/auth.py::login`, the same shadow file contains the symbol-level section. -Cross-file discoveries use `file::symbol` back-pointers in `.shadow/_cross/`. -No embedding database or separate retrieval service is required for this -lookup path. - -``` -source code + conversations + dream experiments - | - v - shadow-frog skills and hooks - | - v - .shadow/ (per-file discoveries, cross-cutting notes, prefs, dreams) - | - v - future agent sessions (hooks, viewer, file reads, search) -``` - -The `.shadow/` layout mirrors the repo: +To start collecting shadow knowledge, open the target repo in your agent +session and run the following. Nap-only ideation can skip this step. ``` -your-repo/ - src/ - auth.py - db/models.py - .shadow/ - .shadowignore gitignore-style excludes - _index.md file list with discovery counts - _prefs.md project-wide user preferences - _cross/ cross-cutting discoveries (span multiple files) - token-expiry-config-split.md - _meta/ - state.json tracking state - _dreams/ experiment archive (reports + branch metadata) - _index.md table of all experiments (branch, parent, tip) - 20250115-143012Z-retry-logic/ one folder per experiment (mirrored from branch) - report.md structured report with YAML frontmatter - patch.diff code-only diff (excludes .shadow/) - manifest.json machine-readable discovery manifest - src/ - auth.py.md discoveries about auth.py - db/ - models.py.md discoveries about models.py -``` - -Discoveries come from three sources: - -**Agent exploration**: the agent reads code and runs experiments: -```markdown -- authenticate_user() silently returns None on expired tokens - instead of raising. 3 of 7 callers don't check the return value. - _(verified, source: exploration, labels: [bug])_ -``` - -**User knowledge**: things you tell the agent during conversation: -```markdown -- The retry logic here took 3 iterations to get right -- it handles - a subtle race condition during rolling deployments. Do not simplify. - _(verified, source: user)_ -``` - -**Collaborative work**: insights from debugging, refactoring, etc.: -```markdown -- While debugging issue #42, discovered that process_batch() silently - drops items exceeding 1MB -- logged at DEBUG level only. - _(verified, source: interaction)_ -``` - -Good discoveries are claims a future agent can act on, not summaries of what a -function is named. Prefer "silently returns `None` on expired tokens" over -"handles token expiration." - ---- - -## Skills - -| Skill | What it does | When to use | -|-------|-------------|-------------| -| **shadow-frog** | Reference docs for the shadow format and conventions | Use when working in a repo with `.shadow/` | -| **shadow-frog-init** | Creates `.shadow/` with structural templates for every file | Once per repo | -| **shadow-frog-update** | Refreshes shadows after code changes; captures knowledge from conversations | After commits, or when you share context | -| **shadow-frog-dream** | Autonomous exploration and experimentation while you're away | When you want the agent to explore on its own | -| **shadow-frog-nap** | Implementation-free proposal trees with independent judgments, selective probes, and task exports | When you need ideas or SWE task briefs rather than implemented features | -| **shadow-frog-meditate** | Deduplicates, merges, and resolves conflicting discoveries | Periodically, to keep the shadow clean | -| **shadow-frog-viewer** | Browse, search, inspect preferences and labels, render dream lineage, and check invariants | When you want to see what's in the shadow, or audit its integrity | - -**Design note:** skills are readable instructions backed by small helper -scripts for deterministic work: initializing shadows, managing dream -worktrees, validating artifacts, reconciling branches, repairing structure, -and rendering viewer outputs. The installer copies both into the target repo. - -### The Dream Skill - -Dream is ShadowFrog's active-discovery mode. Each dream is an **experiment**: -the agent implements real code, runs it, and persists the result as a **named -git branch** pushed to a writable remote. The experiment branch is live, -runnable code, but the primary output is the knowledge distilled into -`.shadow/`. - -- **Branch-based persistence**: each experiment becomes a `dream//` branch -- **Natural compounding**: future dreams branch from prior dream branches, - inheriting code + shadow from the ancestor chain -- **Shadow follows lineage**: each branch has its ancestor's shadow, not - sibling branches. The default branch accumulates all discoveries during - reconciliation. - -The experiment mode can surface tacit knowledge that was not already recorded -in the shadow. While you're away, the agent might try adding retry logic, -parallelizing a pipeline, or refactoring auth into middleware, then distill -what worked, what broke, and why into the shadow. - -**Compounding dreams**: Every experiment saves a report and branch to the -remote. Future dreams read past reports and can branch from prior experiment -branches, continuing partially useful work, avoiding dead ends, and chaining -discoveries across sessions. Dream #3 can branch from dream #1's code and pick -up where it left off. - -> The experiment code is not merged automatically. Turning a dream branch into -> a PR is a manual curation step; the dream skill includes guidance for deciding -> which experiments are worth proposing upstream. - -### Lightweight Ideation with Nap - -Nap reads focused source and optional shadow/dream evidence, refines a small -shortlist, and runs only probes of existing behavior that can change a decision. -It remains implementation-free: feature code and prototypes belong in Dream or -downstream implementation work. It does not require a remote, a -git-tracked shadow, or full shadow initialization. - -``` -/shadow-frog-nap -/shadow-frog-nap mode=coherent -``` - -Default limits are 7 recorded nodes, 2 probes, 2 judge batches, and 2 selected tasks. There is -no default depth cap; `max_depth` is an optional user limit (unset or `null` -otherwise). These are configurable ceilings, not quotas; rejected attempts and failed -probes count. The record validator does not enforce the host agent's actual -API spending. Store records outside `.shadow/`, or under an already initialized -`.shadow/_meta/naps/`, so proposals never become verified discoveries by accident. - -The bundled `nap.py` uses Python's standard library and Git, with no Bash -dependency. It manages a persistent proposal tree: code stays at one -pinned commit while children carry revised hypothetical design states. `@base` -selects the code root; it is not an implemented parent feature. - -```text -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --init --base DEFAULT_REF -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --context @base -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --add IDEAS.json --parent @base -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --add CHILDREN.json --parent n1 -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --review-packet n2 n3 -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --record-review JUDGMENT.json -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --select n2 -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --export TASKS.md -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --export HANDOFF.md --audience implementation -python .github/skills/shadow-frog-nap/nap.py RUN.json --repo REPO --mode coherent --trajectory n3 +/shadow-frog-init ``` -The host agent generates proposals and invokes a strong independent judge on -the shortlist; the Python helper does not call a model. Accepted judgments are -bound to the exact proposal, ancestor design state, mode and base commit. -Unreviewed or stale approvals cannot make tasks ready/exportable. Adding -unrelated siblings does not invalidate an existing approval. - -Updates allocate IDs, preserve parent proposal payloads, and use an exclusive -lock plus atomic replacement. The same record resumes across agent sessions. -Workers return submissions to one writer rather than editing the tree in -parallel. Verdicts are planning judgments, not verified implementations, and -recorded reviewer identities are not independently authenticated by the helper. - -The default export is a detailed **planning brief**. An implementation-audience -handoff presents the same active requirements more concisely, with binding -constraints separated from optional design suggestions and supporting evidence. -Both explicitly distinguish accepted planning review from implementation and -runtime validation, which Nap does not establish. Nonblocking implementation -risks can be recorded separately from questions that prevent planning approval. - -Selected-path review asks what outcome the path now describes and which steps -add capability, reduce uncertainty, or change a useful tradeoff. It does not -require one goal for the entire tree or implementation of superseded ancestors. - -For Claude Code use `.claude/skills/`; on Windows, `py -3` can be used in -place of `python`. See the [Nap skill](skills/shadow-frog-nap/SKILL.md) for -the canonical record and readiness requirements. - -### Coherent Parent-Child Exploration - -Both Dream and Nap support `mode=coherent`; the default remains `broad`. -Coherence applies to **each parent-child edge**, not a fixed tree-wide goal. -Children may extend, integrate, challenge, replace, simplify, or offer -alternatives to a parent's work. Ten children of one parent can pursue ten -different worthwhile directions; sibling diversity is encouraged, not forced -into a quota or a common feature. +This creates symbol-organized templates in `.shadow/`. Choose how to store it: -``` -/shadow-frog-dream mode=coherent -/shadow-frog-nap mode=coherent -``` +| Mode | Effect | +|------|--------| +| **Committed** | Team-shared knowledge; required for Dream | +| **Gitignored** | Local-only knowledge; update, meditate, viewer, and Nap remain available, but Dream is disabled | -Each coherent child records its own goal and an explicit parent connection. -Dream relaxes its breadth/category rules for this mode, while preserving real -execution and artifact requirements. Descendants wait for their parent; -siblings can run in parallel in separate worktrees. Dream validation uses -the selected mode, and reconciler cleanup retains coherent branches and their -canonical index ancestors until explicit curation so task baselines remain -available, including ancestors represented through supported lineage fallbacks. - -Before worktrees or branch switches, `dream-tools.py` pins the current helper -bundle and instructions into a new external run directory. Its returned command -arrays verify that snapshot before execution and supply the selected validation -mode. Continuing an older dream therefore cannot silently select its older -installed validator. The snapshot is host-local run state, not committed task -data, and remains available while children or resumed work use it. - -All automatic branch pruning uses the pinned Python reconciler, including broad -runs that recover pending coherent branches. If lineage metadata cannot be -read, cleanup exits nonzero with repair guidance before deleting any branches. - -Nap compounds ideas and evidence, not implemented APIs. A task must stand -alone at its pinned commit or be regenerated against a real implemented parent. -Export a root-to-leaf trajectory rather than stacking siblings. A final brief -contains the **active** requirements, not both a discarded design and its -replacement. Structural validation cannot establish semantic coherence or -feature feasibility; those still require agent review and appropriate evidence. +If you chose committed storage, commit and push `.shadow/` after initialization. +For installed paths and optional components, see [Installation options](#installation-options). --- -## Installation - -ShadowFrog supports both **GitHub Copilot CLI** (the default) and **Claude -Code**. The installer targets one agent's conventions at a time via -`--agent`: - -| Agent | Skills | Hooks | Context | -|-------|--------|-------|---------| -| `copilot` (default) | `.github/skills/` | `.github/hooks/hooks.json` | `.github/copilot-instructions.md` | -| `claude` | `.claude/skills/` | `.claude/settings.json` | `CLAUDE.md` | +## Everyday Workflow -### Install into your repo +Use the skill that matches your goal. Each link contains its full workflow, +helper commands, and format definitions. -ShadowFrog is installed **per repository**, not globally. This keeps its hooks -limited to projects you have explicitly opted in, and it lets shadow knowledge -and dream artifacts sync through that repository's normal git workflow. +| Command | When to use | +|---------|-------------| +| [`/shadow-frog`](skills/shadow-frog/SKILL.md) | Consult relevant knowledge before editing or investigating code | +| [`/shadow-frog-init`](skills/shadow-frog-init/SKILL.md) | Create the shadow once per repo | +| [`/shadow-frog-update`](skills/shadow-frog-update/SKILL.md) | Refresh after code changes and capture session insights | +| [`/shadow-frog-dream`](skills/shadow-frog-dream/SKILL.md) | Run autonomous experiments while you're away | +| [`/shadow-frog-nap`](skills/shadow-frog-nap/SKILL.md) | Generate reviewed feature-task briefs without implementing them | +| [`/shadow-frog-meditate`](skills/shadow-frog-meditate/SKILL.md) | Merge duplicates and resolve conflicting discoveries | +| [`/shadow-frog-viewer`](skills/shadow-frog-viewer/SKILL.md) | Browse, search, inspect lineage, and audit structural integrity | -Prerequisites: +As you work, the agent captures your code context as `source: user` and +collaborative findings as `source: interaction`. After commits, the pre-tool +hook can detect a shadow behind HEAD and remind the agent to run +`/shadow-frog-update`; the hook does not run the update itself. Meditate +consolidates accumulated knowledge and escalates unresolved conflicts to you. -- `git` and `python3` -- GitHub Copilot CLI or Claude Code -- A git repository as the target project -- To use `/shadow-frog-dream`: a pushable git remote where you can create - `dream/...` branches. After init, `.shadow/` must be tracked by git. +For example, use Viewer to find relevant knowledge or audit its structure: -```bash -cd /path/to/ShadowFrog -# Choose one: -./install.sh --project /path/to/your-repo # Copilot CLI (default) -# ./install.sh --agent claude --project /path/to/your-repo # Claude Code ``` - -On **Windows**, run the PowerShell equivalent (same flags, PowerShell style; -no Bash shell needed): - -```powershell -cd C:\path\to\ShadowFrog -.\install.ps1 -Project C:\path\to\your-repo # Copilot CLI (default) -# .\install.ps1 -Agent claude -Project C:\path\to\your-repo # Claude Code -``` - -This installs skills, hooks, and agent instructions all at once. -Use `--no-hooks` or `--no-context` (`-NoHooks` / `-NoContext` in PowerShell) to skip individual components. - -**After installing**, commit and push so future agent sessions find the skills. -The installer prints the exact `git add` paths for your chosen agent. For -Copilot CLI: - -```bash -cd your-repo -git add .github/skills/ .github/hooks/ .github/copilot-instructions.md -git commit -m "Add ShadowFrog skills, hooks, and context" -git push +/shadow-frog-viewer --search "auth" +/shadow-frog-viewer --top src/auth.py +/shadow-frog-viewer --check-invariants ``` -For Claude Code, stage `.claude/skills/`, `.claude/hooks/`, -`.claude/settings.json`, and `CLAUDE.md` instead. +The [Viewer reference](skills/shadow-frog-viewer/SKILL.md) also covers summaries, +recent discoveries, label filters, preferences, and interactive dream-lineage HTML. --- -## Usage - -### 1. Initialize - -Open your project in Copilot CLI or Claude Code and run: - -``` -/shadow-frog-init -``` - -This scans the codebase, extracts symbols (functions, classes, constants), and -creates `.shadow/` with a template for every file. +## Choose Dream or Nap -After init, choose how `.shadow/` should live: +| | Dream | Nap | +|---|---|---| +| Goal | Learn through implemented experiments | Develop source-grounded feature/task proposals | +| Output | Runnable experiment branches and discoveries | Reviewed task briefs and a persistent proposal tree | +| Continuation | Inherit code and shadow from an ancestor branch | Revise hypothetical designs over a pinned code baseline | +| Implementation | Write and run real code in isolated worktrees | No feature code or prototypes; optional probes inspect existing behavior | +| Prerequisites | Initialized, git-tracked shadow and writable remote | A Git repository with a commit; no remote or initialized shadow required | -| Mode | Use when | Tradeoff | -|------|----------|----------| -| **Committed** | You want team-shared memory and `/shadow-frog-dream` | Best for compounding knowledge; `.shadow/` travels through git | -| **Gitignored** | You want local-only notes | Update, meditate, and viewer still work; dream is disabled | +### Dream: learn by doing -If you plan to run `/shadow-frog-dream`, commit `.shadow/` after init. +Experiments persist as `dream//` branches. Future dreams can +continue a previous experiment, inheriting its code and shadow rather than +sibling branches. Reconciliation accumulates discoveries and experiment +reports on the default branch. -### 2. Work Normally +**Before running Dream, commit and push `.shadow/` and configure a remote +that permits pushing `dream/...` branches and reconciled shadow updates.** +Gitignored shadows cannot use Dream. Experiment code is **not merged +automatically**; adopting it into the project is a manual curation step. -As you code and talk to the agent, ShadowFrog gives the agent places to record -what would otherwise be lost: +See the [Dream workflow](skills/shadow-frog-dream/SKILL.md) for execution, +tooling snapshots, reconciliation, and safe cleanup. -- **You share context** ("don't touch the retry logic, it's subtle") → agent writes it as a `source: user` discovery -- **You debug together** → agent captures insights as `source: interaction` -- **You commit** → before the next mutating agent action, the hook can notice - the shadow is behind HEAD and remind the agent to run `/shadow-frog-update` +### Nap: plan without implementing -### 3. Update +Nap grows a resumable proposal tree and submits shortlisted paths to a strong +independent judge. Exported tasks require a current accepted judgment and +describe the complete change from a real code baseline, not assumed parent APIs. +**Planning approval is not runtime validation.** The host runs the judge; +the Python helper manages records and cannot authenticate reviewer identities. -After significant changes: +Default limits are **7 recorded nodes, 2 probes, 2 judge batches, and 2 selected +tasks**, with no default depth cap. These configurable ceilings are not quotas: +recorded rejections and failed probes count. They do not cap actual API spending. +Keep proposals outside `.shadow/`, or in an initialized `.shadow/_meta/naps/`; +they are not verified discoveries. -``` -/shadow-frog-update -``` +Exports can be detailed **planning briefs** or concise **implementation +handoffs**, with the same active requirements. Binding constraints are separate +from design suggestions; nonblocking implementation risks are separate from +questions that prevent planning approval. -Detects what changed via git diff, updates symbol headings, checks if existing -discoveries still hold, and captures any unrecorded session knowledge. +See the [Nap workflow and helper reference](skills/shadow-frog-nap/SKILL.md) +for tree operations, review receipts, and exports. -### 4. Dream +### Coherent parent-child exploration -Stepping away? Let the agent work while you're gone: +Dream and Nap default to `mode=broad`. Use `mode=coherent` to require a +meaningful connection along each parent-child edge: ``` -/shadow-frog-dream +/shadow-frog-dream mode=coherent +/shadow-frog-nap mode=coherent ``` -The agent explores uncovered code areas, runs experiments in isolated -worktrees, pushes results as persistent dream branches, and writes what it -learns into the shadow. When you return, the shadow is richer and experiment -code is accessible on named branches in the same remote you configured for -the repo. - -> **Requires a git-tracked `.shadow/`.** Dream moves shadow updates through -> git: experiment branches carry `.shadow/_dreams/` reports, manifests, and -> diffs, then reconciliation commits the accumulated `.shadow/` updates back -> to the default branch. If you chose "local only" (gitignored `.shadow/`) -> during init, dream is disabled. `dream-setup.sh` will tell you. The other -> skills (update, meditate, viewer) work either way. - -#### Remote requirements - -Dream needs a remote where it can push `dream/...` branches. For a repo you -already work on, use your existing checkout and its normal remote. -Dream branches are pushed to that remote under `dream//`, and -reconciliation commits shadow knowledge plus dream artifacts back to the -default branch. +Children can extend, integrate, challenge, replace, simplify, or offer an +alternative to their parent. Each has its own goal; siblings can pursue +different directions, including on the same files. There is no fixed +tree-wide goal or diversity quota. -If the repo's remote is not writable, configure a writable remote before -running dream. If you only want private local notes, gitignore `.shadow/` and -use update, meditate, and viewer. Dream is disabled because it depends on -git-tracked shadow artifacts. +Dream descendants wait for their implemented parent; coherent branches and +their ancestors are retained as reproducible baselines until explicit curation. +Nap compounds ideas, not implemented APIs. A combined task follows a +root-to-leaf path and preserves the final active requirements, rather than +stacking unrelated siblings or requiring both discarded and replacement designs. +Combining sibling work requires an explicit integration experiment or proposal. +Structural validation alone cannot establish semantic coherence or feasibility. -The onboarding flow is: - -1. Install ShadowFrog into the existing repo checkout. -2. Commit and push the installed skills, hooks, and context files. -3. Run `/shadow-frog-init`. -4. Commit and push `.shadow/`. -5. Run `/shadow-frog-dream` when you want autonomous exploration. - -After that, shadow knowledge and dream artifacts sync through normal git: -`.shadow/` lives on the default branch, while experiment code remains on -`dream/...` branches until you manually curate anything worth upstreaming. +--- -### 5. Meditate +## How Discoveries Work -Shadow getting noisy? Clean it up: +For example, knowledge about `src/auth.py` lives at `.shadow/src/auth.py.md`; +locations such as `src/auth.py::login` identify the relevant symbol. +Cross-file discoveries live once in `_cross/`, with links from the involved +per-file shadows. ``` -/shadow-frog-meditate +your-repo/ + src/auth.py + .shadow/ + src/auth.py.md file- and symbol-level discoveries + _cross/ cross-cutting discoveries + _prefs.md project-wide preferences + _dreams/ experiment reports, manifests, and patches + _index.md file inventory and counts + _meta/state.json update state + .shadowignore gitignore-style exclusions ``` -Scans for duplicate discoveries, merges near-duplicates, and resolves -conflicting claims. Escalates hard conflicts to you. - -### 6. Browse - -Want to see what's in the shadow? +Store **behavioral discoveries**, not API descriptions or chat transcripts. +For example: +**Agent exploration**: +```markdown +- authenticate_user() silently returns None on expired tokens + instead of raising. 3 of 7 callers don't check the return value. + _(verified, source: exploration, labels: [bug])_ ``` -/shadow-frog-viewer --summary # counts + per-file breakdown -/shadow-frog-viewer --search "auth" # keyword search across all discoveries -/shadow-frog-viewer --recent 5 # most recent discoveries -/shadow-frog-viewer --top src/auth.py # top actionable discoveries for one file -/shadow-frog-viewer --labels bug,security # filter actionable discoveries by label -/shadow-frog-viewer --prefs # all `_prefs.md` entries (project-wide) -/shadow-frog-viewer --check-invariants # audit structural integrity -``` - -Dream experiments are listed in `.shadow/_dreams/_index.md` (dream_id, -category, verdict, title, branch, parent, tip_commit). To render an -interactive view of the dream-branch tree (chains, fresh, and full-tree -tabs), run the bundled script directly: +**User knowledge**: +```markdown +- The retry logic here took 3 iterations to get right -- it handles + a subtle race condition during rolling deployments. Do not simplify. + _(verified, source: user)_ ``` -python3 .github/skills/shadow-frog-viewer/dream-lineage.py -o lineage.html -# (Claude Code: .claude/skills/shadow-frog-viewer/dream-lineage.py) -``` - ---- - -## How Discoveries Work -Each discovery is anchored to a file or symbol and has these properties: +**Collaborative work**: +```markdown +- While debugging issue #42, discovered that process_batch() silently + drops items exceeding 1MB -- logged at DEBUG level only. + _(verified, source: interaction)_ +``` | Property | Values | Meaning | |----------|--------|---------| @@ -513,52 +251,48 @@ Each discovery is anchored to a file or symbol and has these properties: | 4 | `uncertain` | Plausible but unconfirmed. | | 5 | `refuted` | Known wrong; skip. | -### Searching the Shadow - -```bash -cat .shadow/src/auth.py.md # read a file's shadow -cat .shadow/_prefs.md # project-wide preferences -grep -rl "src/auth.py::" .shadow/_cross/ # cross-cutting discoveries -grep -rl "error.handling\|exception" .shadow/ --include="*.md" # search by topic -grep -r "source: user" .shadow/ --include="*.md" # all user knowledge -``` +See the [worked coupon example](examples/coupon-demo/README.md) for a real +shadow, and the [core skill](skills/shadow-frog/SKILL.md) for the canonical +formats and reference rules. --- -## Repository Structure +## Installation Options -For contributors, the main directories are: +The installer copies readable skill instructions, their helper scripts, hooks, +and agent context. It targets one agent's conventions at a time: -| Path | Purpose | -|------|---------| -| `skills/` | The seven ShadowFrog skills and their helper scripts | -| `hook-templates/` | Copilot CLI and Claude Code hook configs plus shared hook scripts | -| `examples/coupon-demo/` | Tiny worked example with a real `.shadow/` | -| `eval/` | Evaluation methodology and results dashboard | -| `tests/` | Pytest suite for installer behavior, hooks, and skill helpers | +| Agent | Skills | Hooks | Context | +|-------|--------|-------|---------| +| `copilot` (default) | `.github/skills/` | `.github/hooks/hooks.json` | `.github/copilot-instructions.md` | +| `claude` | `.claude/skills/` | `.claude/settings.json` | `CLAUDE.md` | + +Use `--no-hooks` or `--no-context` (`-NoHooks` / `-NoContext` in PowerShell) to +skip individual components. On Windows, `py -3` can be used when invoking +Python helpers. Full helper usage is in the linked skill references above. --- -## Tests +## Contributing -ShadowFrog's test suite exercises the real Python scripts and shell hooks -against temporary shadow trees and git repositories, without mocked helper -layers. Nap and coherence coverage includes diverse siblings, parent cycles, -recorded budgets, final-contract exports, installed layouts, and operation -without Bash on PATH. +| Path | Purpose | +|------|---------| +| `skills/` | The seven skills and their helper scripts | +| `hook-templates/` | Agent hook configs and shared scripts | +| `examples/coupon-demo/` | Worked example with a real `.shadow/` | +| `eval/` | [Evaluation methodology](eval/README.md) and [results dashboard](eval/results_dashboard.html) | +| `tests/` | Installer, hook, and helper regression coverage | -Run the suite locally: +Install development dependencies and run the suite: ```bash -pip install -r requirements-dev.txt # pytest, pytest-cov, pathspec -python3 -m pytest # all tests -python3 -m pytest tests/skills/ # just the skill-script tests -python3 -m pytest -k viewer # everything matching "viewer" -python3 -m pytest --cov=skills # coverage report +pip install -r requirements-dev.txt +python3 -m pytest ``` -Test layout mirrors the source layout: `tests/skills/shadow_frog_viewer/` -tests `skills/shadow-frog-viewer/`, etc. +Tests use temporary shadow trees and Git repositories. Their layout mirrors +the source: `tests/skills/shadow_frog_viewer/` covers +`skills/shadow-frog-viewer/`, for example. --- diff --git a/skills/shadow-frog-dream/SKILL.md b/skills/shadow-frog-dream/SKILL.md index f409418..5183f90 100644 --- a/skills/shadow-frog-dream/SKILL.md +++ b/skills/shadow-frog-dream/SKILL.md @@ -369,9 +369,6 @@ dream reports** in `_dreams/`. ### Build Exploration Coverage Map -File-level coverage breadth is the strongest predictor of dream success -(r²=0.63 vs bugs found), NOT dream count (r²=0.04). - Run the packet's `commands.coverage`. It already contains the pinned interpreter/tool version and the original repository path. @@ -391,8 +388,7 @@ For several, append repeated `--scope` argument pairs. In broad mode, when using `--scope`, the per-category task quotas (Phase 2) still apply but are interpreted against the scoped subset. Don't use scoped exploration as the default — pick it only when there's a concrete reason -to concentrate effort. Unscoped diversity remains the strongest -predictor of useful discoveries. +to concentrate effort. ### Review Past Dreams (Required) @@ -432,23 +428,14 @@ codebases (<30 source files), minimum 6 tasks across 4+ categories. ### The 6 Investigation Categories -| Category | What to look for | Priority signals | -|----------|-----------------|-----------------| -| **Investigation** | Under-explored files, shallow coverage, uncertain discoveries | Files with 0-2 discoveries, import chains not traced, `uncertain` entries | -| **Bug hunting** | Defects, edge cases, race conditions | Error-handling code, concurrency, unvalidated inputs | -| **Feature design** | New capabilities, missing functionality | TODOs, FIXMEs, user-facing gaps, integration opportunities | -| **Refactoring** | Structural improvements, duplication | God classes, copy-paste patterns, high-coupling files | -| **Optimization** | Algorithmic efficiency, performance | Hot paths, nested loops, repeated I/O, missing caches | -| **Security audit** | Vulnerabilities, unsafe patterns | Auth code, data handling, deserialization, user inputs | - -| Category | What to experiment | -|----------|-------------------| -| **Investigation** | Write assertion-based tests proving/disproving behavior hypotheses | -| **Bug hunting** | Fuzz inputs, trigger error paths, reproduce race conditions | -| **Feature design** | Implement the feature, run it, evaluate integration | -| **Refactoring** | Do the refactor, run existing tests, measure complexity | -| **Optimization** | Benchmark, profile, implement optimization, measure before/after | -| **Security audit** | Craft adversarial inputs, test injection vectors (local only) | +| Category | Priority signals | Experiment | +|----------|------------------|------------| +| **Investigation** | Files with 0-2 discoveries, untraced import chains, `uncertain` entries | Test behavior hypotheses with assertions | +| **Bug hunting** | Error handling, concurrency, unvalidated inputs | Fuzz inputs, trigger error paths, reproduce races | +| **Feature design** | TODOs, FIXMEs, user-facing gaps, integration opportunities | Implement the feature, run it, evaluate integration | +| **Refactoring** | God classes, duplication, high coupling | Refactor, run existing tests, measure complexity | +| **Optimization** | Hot paths, nested loops, repeated I/O, missing caches | Profile, optimize, measure before/after | +| **Security audit** | Auth, data handling, deserialization, user inputs | Test adversarial inputs and injection vectors locally | **Exception — user-directed focus**: If the user specifies a focus area (e.g., "dream focus on security"), allocate ALL tasks to that category. @@ -474,7 +461,8 @@ Tasks (by category): ### Diversity Rules -In broad mode, prevent fixation (exploring the same files while leaving most untouched): +In broad mode, prefer breadth over depth: more files with 2-3 discoveries over +one file with 20. Prevent fixation with these rules: 1. **Max 2 tasks per source file** (unless prior dream left concrete follow-up) 2. **≥30% of tasks on uncovered files** (from coverage map) @@ -498,22 +486,8 @@ Also include the selected **mode**. Coherent tasks additionally carry their own **goal** and **parent_connection**, not a shared sibling/tree objective. Category still describes the experiment, but coherent mode has no category quota. -Good examples (one per category): -- **Investigation**: "Write assertion harness for request lifecycle — - instrument each layer to log entry/exit and reveal implicit contracts" -- **Bug hunting**: "Fuzz the CSV parser with malformed inputs — what - crashes or silently corrupts?" -- **Feature design**: "Implement retry logic with exponential backoff — - does it handle transient failures without masking permanent ones?" -- **Refactoring**: "Extract 5 duplicate auth checks into middleware — - run tests, measure if it simplifies without breaking special cases" -- **Optimization**: "Benchmark the hot path, implement LRU cache for - repeated lookups — measure before/after wall time" -- **Security audit**: "Craft SQL injection payloads for user-facing - endpoints — does parameterized query hold under nested quotes?" - -Bad examples: "Look at the code", "Trace the flow", "Review error -handling", "Improve code quality" +For example, "Fuzz the CSV parser with malformed inputs to reproduce crashes +or silent corruption" is a concrete experiment; "Review error handling" is not. ### Feature Design: Motivation Required @@ -561,14 +535,10 @@ but the branch persists on the remote. ### Reading Before Implementing -You must understand the code before changing it. For each task: - -1. Read the source file(s) and their shadows (existing discoveries) -2. Read shadows of referenced/referencing files -3. Understand the current behavior, edge cases, and implicit contracts - -Reading is *preparation*, not the deliverable. The deliverable is code -written, code run, results recorded. +Before implementing, read the target source and its shadow, plus shadows of +referenced/referencing files. Identify current behavior, edge cases, and +implicit contracts. Reading is preparation; the deliverable is code written +and run, with results recorded. ### Experiment Setup @@ -623,10 +593,8 @@ or design, rather than being unrelated work parked on the same branch. ### Run -1. Implement the experiment — write real code, run tests/builds -2. Note what worked, broke, surprised -3. Debug if needed — the struggle produces the best discoveries -4. Record results as you go +Implement the experiment, run tests/builds, and debug as needed. Record what +worked, failed, or surprised you as you go. ### Write Shadow Discoveries @@ -655,15 +623,7 @@ Find the `##`/`###` heading for the symbol, then: After writing each discovery, evaluate whether it deserves any of the five actionable labels from `/shadow-frog` (`bug`, `security`, -`performance`, `feature-gap`, `tech-debt`). Apply labels when: - -| Label | Apply when the discovery describes... | -|-------|---------------------------------------| -| `bug` | A defect, silent failure, off-by-one, race, incorrect result, edge case that misbehaves, validate-then-use ordering hazard | -| `security` | Injection vector, unsafe default, missing auth/authz check, sensitive value logged, untrusted input reaching unsafe sink | -| `performance` | Measured bottleneck, O(N²) where N is large, repeated I/O that could batch, missing cache, blocking call on hot path | -| `feature-gap` | Missing capability the codebase clearly needs, asymmetric API (e.g., reads but no writes) | -| `tech-debt` | Duplication, dead code, leaky abstraction, vestigial parameter, inconsistent naming | +`performance`, `feature-gap`, `tech-debt`), using that skill's definitions. Rules: - Apply labels to BOTH the in-file discovery markdown AND the @@ -677,7 +637,7 @@ Rules: - Do not apply labels speculatively. The label says "an engineer should act on this." If you wouldn't act on it, don't label it. -Examples: +Example: ``` - /api/upload accepts paths from request body without normalization, @@ -686,13 +646,6 @@ Examples: Dream report: `_dreams/20260518-161200Z-upload-traversal/` ``` -``` -- HttpClient.send retries 3x on transient failures. - _(verified, source: exploration)_ - Dream report: `_dreams/20260518-163000Z-retry-audit/` -``` -(No label — pure behavioral knowledge, no action implied.) - `dream-validate.py` emits non-blocking warnings when discovery text contains label-signal keywords but no label is set. Treat those warnings as a prompt to re-check the triage, not as a directive. @@ -715,26 +668,14 @@ back-pointers in each file's `## Cross-References` section. #### Discovery Quality Discoveries must be **self-contained process knowledge** — understandable -with just the base codebase. Someone reading main's shadow should understand -the insight without checking out the dream branch. Capture **how to do it**, -**what you learned**, and **what to avoid** — not what was built. The branch -preserves the artifact; the shadow preserves the wisdom. - -Good (behavioral insights about existing code): -- "functools.lru_cache is not thread-safe for initialization — two - threads can trigger duplicate expensive computations on first call." -- "agent.py's retry loop catches all exceptions including OOM, masking - fatal errors that should crash immediately." -- "To add a new eval metric, register in METRIC_MAP at metrics.py:25 - and implement the Metric interface — missing either causes a silent - no-op in the pipeline." - -Bad (descriptions of new code): -- "The implemented PluginFramework has PluginRegistry, PluginManager, - and 7 lifecycle hooks." — describes branch-only artifact. -- "Provides RewardShaper with 4 methods, Welford normalizer, and - GAE-lambda estimation." — feature spec, not behavioral insight. -- "Complete tested module with 58 passing tests." — verdict, not discovery. +from the base codebase without checking out the dream branch. Capture what +was learned, how to apply it, and what to avoid, not a description of the +experiment-only artifact. + +- **Good:** "agent.py's retry loop catches all exceptions including OOM, + masking fatal errors that should crash immediately." +- **Bad:** "The implemented PluginFramework has PluginRegistry, PluginManager, + and 7 lifecycle hooks." This describes a branch-only artifact. Per-file discoveries should reference the dream report: @@ -920,11 +861,9 @@ bash "$SKILL_DIR/dream-cleanup.sh" "$WORKTREE_DIR" --repo-root "$REPO_ROOT" ``` `dream-cleanup.sh` does the equivalent of `git worktree remove --force` -followed by `git worktree prune`, but ALSO falls back to a safety-gated -`rm -rf` if `git worktree remove` silently fails — the failure mode that -leaked tens of dream worktrees per AFK session under the previous inline -snippet (see bug-worktree-leak.md). The rm fallback ONLY fires for paths -that match `${DREAM_WORKTREE_BASE:-/tmp/shadowfrog-dreams}//dream-` +followed by `git worktree prune`, with a safety-gated `rm -rf` fallback if +worktree removal silently fails. The fallback only accepts paths matching +`${DREAM_WORKTREE_BASE:-/tmp/shadowfrog-dreams}//dream-` exactly; any other path is refused. Remove as you go. If push failed, keep the worktree. @@ -1145,9 +1084,7 @@ are four places they get cleaned up: 3. **`dream-gc.sh` (auto-triggered)** — `dream-setup.sh` invokes this sweeper at the start of each new dream, throttled by a per-namespace `.last-gc` tombstone to run at most once per `DREAM_GC_INTERVAL_MIN` - minutes (default 60). Catches orphans from crashed dreams, machine - reboots, OOM-killed agents — the long tail of cleanup failures that - accumulated GBs of leaked worktrees on long-running fleets. + minutes (default 60). Catches orphaned worktrees. Env knobs (all optional, sensible defaults): - `DREAM_GC_AUTO=0` — disable the auto-trigger entirely @@ -1203,17 +1140,6 @@ git cherry-pick git show | git apply ``` -## Guidance - -- **Always experiment.** Implementation reveals what reading cannot. -- **Small tasks, big lessons.** 30-minute experiment > 3 hours reading. -- **Fail forward.** "Tried X, broke because Y" is extremely valuable. -- **Broad mode: breadth over depth.** More files with 2-3 discoveries > one file with 20. -- **No descriptions.** "Catches all exceptions including OOM" yes. - "This function authenticates users" no. -- **Compound deliberately.** Read parent's report. Build on findings. -- **Branches are cheap, shadow is expensive.** Push freely, write carefully. - ## Experiment Completion Criteria A task is complete ONLY when ALL of these hold: diff --git a/skills/shadow-frog-init/SKILL.md b/skills/shadow-frog-init/SKILL.md index 6298466..df52a93 100644 --- a/skills/shadow-frog-init/SKILL.md +++ b/skills/shadow-frog-init/SKILL.md @@ -26,19 +26,13 @@ python3 .github/skills/shadow-frog-init/shadow-init.py [options] python3 .claude/skills/shadow-frog-init/shadow-init.py [options] ``` -**IMPORTANT: Run from the repo/worktree root directory.** The script -auto-detects the root via `git rev-parse --show-toplevel`, which returns -the correct root for regular repos AND worktrees. If auto-detection fails -(common when python is routed through Docker or the `.git` file points to -an inaccessible path), pass `--root` explicitly: +**Run from the repo/worktree root.** The helper auto-detects it with +`git rev-parse --show-toplevel`. If Git cannot resolve or access the root +(for example, inside a container), use `--root` to bypass detection: ```bash # If auto-detect fails, pass the root explicitly: python3 .github/skills/shadow-frog-init/shadow-init.py --root "$(pwd)" - -# In Docker wrapper scenarios (eval harness), git may not work inside -# the container. Use --root to bypass git detection: -python3 .github/skills/shadow-frog-init/shadow-init.py --root /testbed ``` ### Options @@ -51,32 +45,18 @@ python3 .github/skills/shadow-frog-init/shadow-init.py --root /testbed ### What it does -1. Discovers source files via `git ls-files` -2. Filters through `.shadow/.shadowignore` (gitignore syntax) -3. Extracts symbols from each file (classes, functions, methods) -4. Creates per-file shadow `.md` files with symbol headings -5. Creates `_index.md`, `_prefs.md`, `_meta/state.json`, `.shadowignore` -6. Reports: files, symbols, languages detected - -### After running - -Tell the user: -- "Edit `.shadow/.shadowignore` to exclude files that shouldn't be shadowed" -- "Run `/shadow-frog-dream` for autonomous exploration, or `/shadow-frog-update` after your next changes" +The helper discovers sources with `git ls-files`, applies +`.shadow/.shadowignore`, extracts symbols, and creates symbol-organized +per-file shadows plus `_index.md`, `_prefs.md`, `_meta/state.json`, and +`.shadowignore`. It reports file/symbol counts and detected languages. -Then decide the version-control mode (do not skip this — it is not handled -by the script). Ask the user: "Should `.shadow/` be **committed** (shared -with your team via git) or **gitignored** (local to your machine only)?" -Make the trade-off explicit before they choose — see -[Step 9: Handle .gitignore](#9-handle-gitignore) for the full committed vs -gitignored comparison. Key caveat: a gitignored `.shadow/` disables -`shadow-frog-dream` (dreams move `.shadow/` through git). If gitignored, add -`.shadow/` to `.gitignore`. +After creating `.shadow/`, complete [Post-init steps](#post-init-steps). +The helper does not choose the version-control mode. ## Fallback: Manual Init If the Python script fails (wrong Python version, missing file, etc.), -follow these steps manually: +follow these steps manually, then complete [Post-init steps](#post-init-steps). ### 1. Check preconditions @@ -156,10 +136,6 @@ out/ .claude/hooks/scripts/shadow-frog-* ``` -Tell the user: "Edit `.shadow/.shadowignore` to exclude files or -folders that shouldn't be shadowed (e.g., vendored code, generated -files, tool configs)." - ### 5. Create `_prefs.md` ```markdown @@ -236,7 +212,11 @@ Rules: | src/auth.py | Python | 5 (UserAuth, authenticate_user, ...) | 0 | ``` -### 9. Handle .gitignore +## Post-init Steps + +After creating `.shadow/` with either path: + +### Version-Control Mode Ask the user: "Should `.shadow/` be **committed** (shared with your team via git) or **gitignored** (local to your machine only)?" @@ -261,7 +241,10 @@ Before they decide, make the trade-off explicit: - If gitignored: add `.shadow/` to `.gitignore`. -### 10. Report +### Report Print: files discovered, languages detected, total symbols. -Suggest: `/shadow-frog-update` for deeper analysis, `/shadow-frog-dream` for autonomous exploration. +Tell the user to edit `.shadow/.shadowignore` to exclude unwanted paths, +such as vendored/generated files or tool configs. Suggest `/shadow-frog-update` +for deeper analysis or after code changes, and `/shadow-frog-dream` for +autonomous exploration if `.shadow/` is committed. diff --git a/skills/shadow-frog-meditate/SKILL.md b/skills/shadow-frog-meditate/SKILL.md index 0a6e81a..9a65ceb 100644 --- a/skills/shadow-frog-meditate/SKILL.md +++ b/skills/shadow-frog-meditate/SKILL.md @@ -17,19 +17,6 @@ Shadow hygiene — deduplicate, merge, and resolve conflicts across the entire `.shadow/` knowledge base. Prerequisite: `.shadow/` exists with discoveries. -## Why Meditate? - -Over time, shadows accumulate noise: -- **Duplicates**: the same insight written differently by different sessions -- **Near-duplicates**: one discovery is a subset of another -- **Conflicts**: two discoveries contradict each other (code may have changed, - or one was wrong) -- **Cross-scope duplicates**: a per-file discovery and a `_cross/` entry - saying the same thing - -This noise confuses downstream agents and dilutes signal. Meditate cleans -it up. - ## Phase 1: Scan Use parallel subagents to scan the shadow. Each subagent handles a batch @@ -44,21 +31,23 @@ Not every file needs scanning. To reduce cost: - **Always scan files with 5+ discoveries** — highest duplicate risk For the first meditate after a large dream run, most files will need -scanning. For incremental meditation after small updates, this can -reduce scope by 80%+. +scanning. ### Per-File Scan For each per-file shadow (e.g., `src/auth.py.md`): 1. Read all discoveries under each `## symbol` heading -2. For each pair of discoveries under the **same symbol**, classify: +2. Compare discoveries under the **same symbol** by behavioral meaning, + not wording: - **Duplicate**: same behavioral claim, different wording - **Near-duplicate**: one discovery is a subset/refinement of the other - **Conflict**: the two discoveries make contradicting claims - **Distinct**: genuinely different insights — no action needed 3. Record each finding as a structured action (see below) +Overlapping `Also involves:` refs can help identify related claims. + ### Scan Output Format Subagents must output findings as **one JSON object per line** so the @@ -97,18 +86,6 @@ After per-file scanning: or `_cross/` entry duplicates it 4. Record cross-scope findings the same way -### Scanning Guidelines - -- Compare claims semantically, not just textually. "Returns None on - expired tokens" and "Silently returns None when token expires" are - duplicates. -- Two discoveries about the same function but covering different - behaviors are **distinct**, not duplicates. E.g., "returns None on - expired tokens" vs "uses constant-time comparison" — these are - unrelated observations about the same function. -- Pay attention to `Also involves:` — two discoveries with overlapping - `Also involves:` refs are more likely related. - ## Phase 2: Resolve Process each finding by type. @@ -122,21 +99,6 @@ Combine into a single discovery: - Merge `Also involves:` refs (union of both) - Delete the weaker entry -Example: -``` -BEFORE (two entries under same symbol): -- authenticate_user() returns None on expired tokens. - _(verified, source: exploration)_ -- When the token is expired, authenticate_user silently returns None - instead of raising. 3 of 7 callers don't check. - _(verified, source: exploration)_ - -AFTER (merged): -- authenticate_user() silently returns None on expired tokens instead - of raising. 3 of 7 callers don't check the return value. - _(verified, source: exploration)_ -``` - ### Near-Duplicates → Absorb The broader discovery absorbs the narrower one: diff --git a/skills/shadow-frog-update/SKILL.md b/skills/shadow-frog-update/SKILL.md index a9764f3..05190e0 100644 --- a/skills/shadow-frog-update/SKILL.md +++ b/skills/shadow-frog-update/SKILL.md @@ -43,10 +43,14 @@ Categorize: modified, added, deleted, renamed. For each changed file, update its shadow at the symbol level: -- **Added symbols** → add new `##` section -- **Removed symbols** → mark section as `REMOVED`, keep discoveries for history -- **Renamed symbols** → update heading, preserve discoveries -- **Modified symbols** → check if discoveries still hold +| Symbol change | Action | +|---------------|--------| +| Added | Add the appropriate `##`/`###` heading | +| Removed | Mark the section `REMOVED`; keep discoveries for history | +| Renamed | Move discoveries to the updated heading | +| Modified | Check whether discoveries still hold | + +Only mark `source: user` discoveries stale if the symbol is completely removed. Lightweight update (auto/hook): re-extract symbols, update headings, flag stale. Deep update (manual/dream): read diffs, generate new discoveries, verify existing ones. @@ -56,16 +60,13 @@ Deep update (manual/dream): read diffs, generate new discoveries, verify existin When the user shares knowledge during the session, write it immediately. Do not batch for later. -Signals to capture: +Capture warnings, gotchas, design intent, history, deprecations, contracts, +and conventions. Representative examples: | Signal | Example | Category | |--------|---------|----------| | Warning | "Don't change the retry logic, it's subtle" | warning | | Design intent | "We use this pattern because the API is unreliable" | intent | -| History | "We tried caching here but it caused stale reads" | history | -| Gotcha | "This looks wrong but matches the tax authority spec" | warning | -| Deprecation | "This module is being replaced by v2/" | intent | -| Contract | "The 30s timeout matches our SLA" | contract | | Convention | "Always use the helper in utils.py, not raw SQL" | convention | Write as: @@ -192,10 +193,3 @@ update sessions: - If 3+ files involved → create in `_cross/.md` instead, add back-pointers - If project-wide preference with no file reference → write to `_prefs.md` - Slug naming: kebab-case derived from title (e.g., "Token expiry config split" → `token-expiry-config-split.md`) - -## Staleness Rules - -- Symbol modified → check if discovery still holds -- Symbol renamed → move discoveries to new heading -- Symbol removed → mark section `REMOVED`, keep discoveries -- `source: user` discoveries → only mark stale if symbol completely removed diff --git a/skills/shadow-frog-viewer/SKILL.md b/skills/shadow-frog-viewer/SKILL.md index 478f653..3c0e68e 100644 --- a/skills/shadow-frog-viewer/SKILL.md +++ b/skills/shadow-frog-viewer/SKILL.md @@ -52,35 +52,14 @@ No arguments defaults to `--summary`. ### Examples ```bash -# Quick overview -python3 shadow-viewer.py - -# Everything about auth (files, symbols, text, cross-cutting) -python3 shadow-viewer.py --search auth - -# Find token-related knowledge +# Search file names, symbols, discoveries, cross-cutting entries, and preferences python3 shadow-viewer.py --search "token expiry" -# 5 most recent discoveries -python3 shadow-viewer.py --recent 5 - -# All known bugs -python3 shadow-viewer.py --labels bug - # Security and performance issues python3 shadow-viewer.py --labels security,performance -# Top actionable discoveries for a single file (used by the preToolUse hook) -python3 shadow-viewer.py --top src/auth.py - -# Same, but broaden label filter and show up to 5 entries +# Broaden the per-file label filter and show up to 5 entries python3 shadow-viewer.py --top src/auth.py --top-labels bug,security,performance --top-limit 5 - -# Team preferences -python3 shadow-viewer.py --prefs - -# Structural audit (run before commits or after dream reconciliation) -python3 shadow-viewer.py --check-invariants ``` ## Dream Lineage Visualization @@ -102,14 +81,8 @@ python3 .github/skills/shadow-frog-viewer/dream-lineage.py -o my-lineage.html python3 .github/skills/shadow-frog-viewer/dream-lineage.py --shadow-dir /path/to/.shadow ``` -The HTML file has three tabs: -- **🌳 Chains** — compounding chains as tree cards, sorted by depth -- **📋 Fresh** — non-compounding experiments grouped by category -- **🗂️ Full Tree** — compact view of the entire lineage in one tree - -Each node shows the experiment's category icon, name, verdict, test count, -and discovery count. Click "▶ Show report" to expand the full experiment -report inline. +The HTML groups compounding chains and fresh experiments, includes a full +lineage tree, and supports expanding each experiment's report. ## Fallback: Shell One-Liners @@ -151,12 +124,7 @@ find .shadow -name '*.md' -not -path '*/_meta/*' -printf '%T@ %p\n' | sort -rn | ## Responding to the User -After running a view, present the results clearly: -- For `--summary`: show the output directly, highlight anything notable -- For `--search`: summarize key findings, group by relevance -- For `--recent`: present the discoveries conversationally -- For `--top`: typically called by the preToolUse hook before a file is - edited; output is intentionally short and pre-formatted. If invoked - manually, present as-is. +- Preserve `--top` output as-is; it is intentionally compact and pre-formatted + for the preToolUse hook. - If the shadow is empty or has no discoveries, suggest running `/shadow-frog-dream` to populate it diff --git a/skills/shadow-frog/SKILL.md b/skills/shadow-frog/SKILL.md index 8b0a840..9074571 100644 --- a/skills/shadow-frog/SKILL.md +++ b/skills/shadow-frog/SKILL.md @@ -24,30 +24,20 @@ to that code location. **Every time you work on code in a repo with `.shadow/`:** 1. **Read `_prefs.md` first** — it contains project-wide conventions, - user preferences, and things the user explicitly wants to avoid. - Violating a preference wastes the user's time. -2. **Read `_cross/` discoveries** — these are the highest-value findings, - spanning multiple files. List `_cross/` and read any files whose titles - relate to the area you're working in. Cross-cutting discoveries reveal - hidden contracts, interaction bugs, and design patterns that per-file - shadows alone cannot capture. -3. **Check `_dreams/` for experiment results** — `_dreams/_index.md` lists - autonomous exploration experiments. Read reports relevant to your task — - they contain verified bug analyses, attempted fixes, and architectural - insights. Dreams may contain knowledge not yet distilled into per-file - shadows, so always check when investigating a bug or unfamiliar area. -4. **Before editing any file**: read its shadow (`.shadow/.md`), - check `_cross/` for cross-cutting discoveries about it, and apply - what you learn. The shadow contains known bugs, edge cases, and - implicit contracts discovered by previous sessions. - **Note**: `_index.md` discovery counts may be stale — always check - per-file shadows and `_cross/` directly rather than relying solely on - the index summary. + user preferences, and things to avoid. +2. **Read relevant `_cross/` discoveries** — list `_cross/` and read entries + whose titles relate to the current area, including cross-file contracts + and interactions. +3. **Check `_dreams/_index.md`** and read relevant experiment reports, + especially when investigating bugs or unfamiliar code. They may contain + findings not yet distilled into per-file shadows. +4. **Before editing a file**, read its shadow (`.shadow/.md`) and + relevant `_cross/` entries, then apply the discoveries. + `_index.md` counts may be stale; inspect the actual shadows and `_cross/`. 5. **When the user explains something about code** (gotcha, design intent, warning, history): write a `source: user` discovery to the shadow - immediately. Do not ask where to put it — resolve the `file::symbol` - anchor yourself by searching `_index.md`, shadow files, and session - context (current file, recent edits). + immediately. Resolve its `file::symbol` anchor from `_index.md`, shadows, + and session context; do not ask where to put it. 6. **When the user states a preference or convention** (not tied to any specific file): write it to `_prefs.md` immediately. 7. **After code changes**: run `/shadow-frog-update` @@ -410,12 +400,10 @@ real code baseline separately from idea lineage. Export a root-to-leaf path for a trajectory, not stacked siblings; a final task uses its final active requirements rather than all superseded ancestor designs. -Nap's host agent generates proposals and delegates a strong independent -shortlist review. Version-2 tree records bind each judgment to the exact -proposal, ancestor design state, mode, and source commit. The helper manages -atomic append/review/select operations and rejects missing or stale approvals -for ready tasks, but it neither invokes an LLM nor authenticates that a recorded -review happened. Judgment approval is planning confidence, not execution proof. +Nap task export requires current acceptance from a strong independent judge; +see `/shadow-frog-nap` for the record and review workflow. The host runs the +judge; the helper checks approval bindings, not reviewer authenticity or +semantic truth. Approval is planning confidence, not execution proof. ## Related Skills