diff --git a/.github/workflows/tests.yml b/.github/workflows/tests.yml index 6d16cb1..0d089c3 100644 --- a/.github/workflows/tests.yml +++ b/.github/workflows/tests.yml @@ -8,6 +8,21 @@ permissions: contents: read jobs: + citation-compatibility: + name: Citation helpers (Python 3.9) + runs-on: ubuntu-latest + steps: + - name: Checkout + uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + - name: Set up Python + uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0 + with: + python-version: '3.9' + - name: Check citation and viewer entry points + run: | + python skills/shadow-frog/shadow-cite.py --help + python skills/shadow-frog-viewer/shadow-viewer.py --shadow-dir examples/coupon-demo/.shadow --check-invariants + pytest: name: pytest (${{ matrix.os }}, Python 3.12) runs-on: ${{ matrix.os }} diff --git a/CHANGELOG.md b/CHANGELOG.md index 773f964..7e2910c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,7 +7,14 @@ shadow knowledge bases for any codebase. --- -## Unreleased +## 2026-09-28 + +### Added +- **Visible knowledge citations** — `citation_score` lives alongside discovery + metadata in Markdown. Agents record consulted entries once per task after + normal file/symbol reads. A small locked increment helper preserves other + content; Viewer and reconciliation understand the same score field without + a database or additional retrieval workflow. ### Fixed - **Reconciliation path containment** — validate untrusted manifest destinations diff --git a/README.md b/README.md index 8b6b37f..960bd13 100644 --- a/README.md +++ b/README.md @@ -117,6 +117,13 @@ For example, use Viewer to find relevant knowledge or audit its structure: The [Viewer reference](skills/shadow-frog-viewer/SKILL.md) also covers summaries, recent discoveries, label filters, preferences, and interactive dream-lineage HTML. +Agents still read shadow files and symbols directly. Each entry carries a visible +`citation_score`, initially 0, as an approximate hint of how often agents revisit +it. After consulting an entry, the agent records it once per task with the small +core `shadow-cite.py` helper; its only job is to update the selected Markdown +counter safely. No database or special retrieval interface is required. Scores +never replace relevance, trust, or verification. + --- ## Choose Dream or Nap @@ -218,21 +225,21 @@ For example: ```markdown - authenticate_user() silently returns None on expired tokens instead of raising. 3 of 7 callers don't check the return value. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ ``` **User knowledge**: ```markdown - The retry logic here took 3 iterations to get right -- it handles a subtle race condition during rolling deployments. Do not simplify. - _(verified, source: user)_ + _(verified, source: user, citation_score: 0)_ ``` **Collaborative work**: ```markdown - While debugging issue #42, discovered that process_batch() silently drops items exceeding 1MB -- logged at DEBUG level only. - _(verified, source: interaction)_ + _(verified, source: interaction, citation_score: 0)_ ``` | Property | Values | Meaning | @@ -240,6 +247,7 @@ For example: | **Status** | `verified` / `uncertain` / `refuted` | Has the claim been confirmed? | | **Source** | `exploration` / `user` / `interaction` | Where did this knowledge come from? | | **Labels** | `bug`, `performance`, `security`, `feature-gap`, `tech-debt` | Optional; marks actionable discoveries | +| **Citation score** | Nonnegative integer; missing means `0` | Approximate agent-reported revisits, stored as `citation_score` in the entry | ### Trust Hierarchy diff --git a/agent-context.md b/agent-context.md index 11db20a..888d8b8 100644 --- a/agent-context.md +++ b/agent-context.md @@ -13,6 +13,16 @@ This project uses a `.shadow/` knowledge base with verified discoveries about no The shadow contains discoveries from code analysis and user conversations. Always consult it before making assumptions about code behavior. +Each entry's Markdown metadata includes `citation_score` (initially 0; omitted +also means 0). After deliberately consulting an entry, increment it once per +task using `shadow-cite.py` in the core `shadow-frog` skill, supplying the file, +symbol and exact claim text already read. Batch repeated `--text` arguments +for one file/section. Coordinate subagent updates through one writer. +Do not count every entry in an opened file or recount repeated reads. +The score is approximate revisit frequency, not confidence; relevance and trust +take precedence. `/shadow-frog-viewer` is for user-facing inspection, not a +required read path. There is no citation database or retrieval protocol. + ### Key directories - `.shadow/.md` — per-file shadows with symbol-level discoveries diff --git a/claude.md b/claude.md index 2d623e6..d6c911a 100644 --- a/claude.md +++ b/claude.md @@ -11,6 +11,8 @@ ShadowFrog/ skills/ shadow-frog/SKILL.md Main entrypoint (docs, reference system, search) shadow-frog/_coherence.py Shared structural parent-connection validation + shadow-frog/_citations.py Visible score metadata and safe per-file increments + shadow-frog/shadow-cite.py Record exact consulted claims after native reads shadow-frog-init/ First-time setup (create .shadow/) SKILL.md Init instructions + fallback steps shadow-init.py Python helper script @@ -84,7 +86,7 @@ Canonical formal spec: `/shadow-frog`. The shapes below are the minimum an agent Per-file discovery (anchored by `file::symbol` heading; labels and `Also involves:` are optional): ``` - - _(, source: [, labels: [bug, security]])_ + _(, source: [, labels: [bug, security]], citation_score: 0)_ Also involves: `file::symbol`, `file::symbol` ``` @@ -98,19 +100,32 @@ Cross-cutting (`_cross/.md`, slug = kebab-case from title, e.g. "DB connec **Discovery**: -_(, source: )_ +_(, source: , citation_score: 0)_ ``` Preference (`_prefs.md` — project-wide, no file/symbol anchor): ``` - - _(source: )_ + _(source: , citation_score: 0)_ ``` - Labels (lowercase, comma-separated): `bug`, `performance`, `security`, `feature-gap`, `tech-debt`. Only for actionable discoveries. - `Also involves:` always uses `file::symbol`, never bare file paths. - `Dream report: _dreams//` is optional — only for experiment-derived discoveries. +### Citation Scores + +- Keep one visible nonnegative integer `citation_score` in Markdown metadata, + after optional labels. New entries start at 0; omitted scores also mean 0. +- Agents read files/symbols directly and explicitly cite consulted entries once + per task. Use core `shadow-cite.py` for serialized exact-claim increments; no + database, opaque IDs, or required retrieval service. +- Score updates must not change discovery prose, provenance, references, or counts. + Coordinate ordinary edits with citation writes; only helper calls share its lock. +- Keep scores when rewording/moving the same claim, and use max rather than sum + when combining duplicates or inherited branch state. Scores are approximate, + not a global audited count or a trust/confidence value. + ### Verification - Observe-based: read source at `file::symbol`, trace logic, confirm claim. - Do-based: write and run a short test/script to confirm or refute. diff --git a/examples/coupon-demo/.shadow/_cross/coupon-case-normalization-mismatch.md b/examples/coupon-demo/.shadow/_cross/coupon-case-normalization-mismatch.md index 4538e31..1b0c04f 100644 --- a/examples/coupon-demo/.shadow/_cross/coupon-case-normalization-mismatch.md +++ b/examples/coupon-demo/.shadow/_cross/coupon-case-normalization-mismatch.md @@ -8,4 +8,4 @@ **Discovery**: validate_coupon normalizes coupon codes to uppercase via code.upper() before lookup, but calculate_total passes coupon_code directly to get_coupon without normalization. A user who validates "save20" (returns True) and then passes "save20" to calculate_total gets no discount — the coupon silently fails because load_coupon's keys are uppercase. This creates a validate-then-use inconsistency where validated codes don't work. -_(verified, source: exploration, labels: [bug])_ +_(verified, source: exploration, labels: [bug], citation_score: 0)_ diff --git a/examples/coupon-demo/.shadow/_cross/global-coupon-cache-side-effects.md b/examples/coupon-demo/.shadow/_cross/global-coupon-cache-side-effects.md index 3ee95a3..a939b78 100644 --- a/examples/coupon-demo/.shadow/_cross/global-coupon-cache-side-effects.md +++ b/examples/coupon-demo/.shadow/_cross/global-coupon-cache-side-effects.md @@ -9,4 +9,4 @@ **Discovery**: COUPON_CACHE is a module-level global dict shared by all importers. Any call to get_coupon (directly or via validate_coupon) permanently populates the cache, including caching None for invalid codes. In tests, cache entries from one test persist into the next — there is no reset mechanism. validate_coupon caches under the uppercased key, while calculate_total would cache under the original-case key, so a single logical coupon code can produce two separate cache entries ("SAVE20" and "save20") with different values. -_(verified, source: exploration, labels: [bug])_ +_(verified, source: exploration, labels: [bug], citation_score: 0)_ diff --git a/examples/coupon-demo/.shadow/_cross/mutation-through-discount-pipeline.md b/examples/coupon-demo/.shadow/_cross/mutation-through-discount-pipeline.md index 3dc557d..428f7b5 100644 --- a/examples/coupon-demo/.shadow/_cross/mutation-through-discount-pipeline.md +++ b/examples/coupon-demo/.shadow/_cross/mutation-through-discount-pipeline.md @@ -8,4 +8,4 @@ **Discovery**: apply_bulk_discount mutates item dicts in-place (modifying "price" keys) and returns the same list object. When piped into calculate_total, the mutation is invisible — calculate_total sees already-reduced prices. But any code holding a reference to the original items list now sees the discounted prices permanently. Calling apply_bulk_discount multiple times compounds discounts (0.90^N multiplier). The test_bulk_then_coupon test avoids this by creating fresh items, but real usage with shared item references would silently corrupt prices. -_(verified, source: exploration, labels: [bug])_ +_(verified, source: exploration, labels: [bug], citation_score: 0)_ diff --git a/examples/coupon-demo/.shadow/cart.py.md b/examples/coupon-demo/.shadow/cart.py.md index 63a62af..2238aec 100644 --- a/examples/coupon-demo/.shadow/cart.py.md +++ b/examples/coupon-demo/.shadow/cart.py.md @@ -5,53 +5,53 @@ ## File-Level - COUPON_CACHE is a module-level mutable global dict shared across all importers — any module that imports from cart.py shares the same cache instance, causing cross-module state pollution. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ Also involves: `inventory.py::validate_coupon` ## `COUPON_CACHE` - Never cleared or evicted — grows monotonically for the lifetime of the process. In a long-running server, every unique coupon code ever queried remains cached forever. - _(verified, source: exploration, labels: [performance])_ + _(verified, source: exploration, labels: [performance], citation_score: 0)_ - Caches None for invalid codes. Once an invalid code is looked up, the None result is permanently cached, preventing any future lookup even if the underlying data changes. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ - Case-variant lookups create duplicate cache entries for the same logical coupon. validate_coupon("save20") caches "SAVE20" → valid, then calculate_total("save20") caches "save20" → None. Cache grows 2× faster with mixed-case usage. - _(verified, source: exploration, labels: [bug, performance])_ + _(verified, source: exploration, labels: [bug, performance], citation_score: 0)_ Dream report: `_dreams/20260420-140000Z-cache-poison-sequence/` ## `load_coupon` - Case-sensitive lookup against uppercase keys ("SAVE20", "HALF"). Passing lowercase (e.g., "save20") returns None even though the coupon conceptually exists. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ - Returns None for unknown codes (via dict.get default), not an exception. Callers must handle None. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ ## `get_coupon` - The `if code not in COUPON_CACHE` guard means each code is loaded exactly once per process. But because None is a valid cached value, invalid codes are also "loaded once" and permanently considered invalid. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ Also involves: `cart.py::COUPON_CACHE`, `cart.py::load_coupon` - Accepts any hashable type as key (True, 42, lists-as-errors). Non-string keys permanently cache None entries that can never resolve to valid coupons — silent cache pollution. - _(verified, source: exploration, labels: [security])_ + _(verified, source: exploration, labels: [security], citation_score: 0)_ Dream report: `_dreams/20260420-142000Z-adversarial-inputs/` Also involves: `cart.py::COUPON_CACHE` ## `calculate_total` - Does NOT normalize coupon_code to uppercase before lookup. Lowercase codes silently produce no discount (coupon returns None from cache or load_coupon). This contradicts validate_coupon which does normalize. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ Also involves: `inventory.py::validate_coupon` - `if coupon_code:` is falsy for empty string "", None, 0, and False — all skip coupon lookup silently. No distinction between "no coupon" and "invalid coupon". - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ - Accepts negative prices and quantities without validation. Negative subtotals still have 8% tax applied, producing negative totals (e.g., price=-10, qty=1 → total=-10.80). - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ - Tax rate (0.08 = 8%) is hardcoded with no configuration mechanism. Changing tax requires editing source code. - _(verified, source: exploration, labels: [tech-debt])_ + _(verified, source: exploration, labels: [tech-debt], citation_score: 0)_ - Coupon min_total check uses the actual subtotal computed from items at call time. Since apply_bulk_discount mutates prices in-place before calculate_total runs, the min_total check sees post-bulk-discount prices. If bulk discount drops subtotal below min_total, the coupon is correctly rejected. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ Dream report: `_dreams/20260420-141000Z-bulk-min-total-interaction/` Also involves: `inventory.py::apply_bulk_discount` - Non-string coupon_code values (True, 42, etc.) pass the `if coupon_code:` truthiness check, reach get_coupon, cache None under the non-string key, and silently produce no discount. No type validation exists. - _(verified, source: exploration, labels: [security])_ + _(verified, source: exploration, labels: [security], citation_score: 0)_ Dream report: `_dreams/20260420-142000Z-adversarial-inputs/` Also involves: `cart.py::get_coupon`, `cart.py::COUPON_CACHE` diff --git a/examples/coupon-demo/.shadow/inventory.py.md b/examples/coupon-demo/.shadow/inventory.py.md index 2dea944..c6afe7a 100644 --- a/examples/coupon-demo/.shadow/inventory.py.md +++ b/examples/coupon-demo/.shadow/inventory.py.md @@ -5,36 +5,36 @@ ## File-Level - Imports get_coupon from cart, creating a dependency on cart's COUPON_CACHE. Any call to validate_coupon pollutes the shared cache as a side effect. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ Also involves: `cart.py::get_coupon`, `cart.py::COUPON_CACHE` ## `validate_coupon` - Normalizes code to uppercase via `code.upper()` before lookup, but calculate_total does NOT — so a code that validates successfully may still produce no discount when passed directly to calculate_total. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ Also involves: `cart.py::calculate_total`, `cart.py::load_coupon` - Side effect: populates COUPON_CACHE with the uppercased code. Calling validate_coupon("save20") caches under key "SAVE20", but a later calculate_total("save20") looks up lowercase "save20" — a cache miss that then caches None under "save20". - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ Also involves: `cart.py::COUPON_CACHE`, `cart.py::get_coupon` - Will crash with AttributeError if code is None (None has no .upper() method). No guard against non-string input. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ - Also crashes on any non-string type: int, list, dict all raise AttributeError on .upper(). Needs `isinstance(code, str)` guard. - _(verified, source: exploration, labels: [security])_ + _(verified, source: exploration, labels: [security], citation_score: 0)_ Dream report: `_dreams/20260420-142000Z-adversarial-inputs/` ## `apply_bulk_discount` - Mutates items in-place — modifies the original dict objects' "price" keys. The caller's list is permanently altered. Returns the same list object (not a copy). - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ - Crashes with KeyError if items lack 'qty' key, and TypeError if 'qty' is a string. No input validation. - _(verified, source: exploration, labels: [security])_ + _(verified, source: exploration, labels: [security], citation_score: 0)_ Dream report: `_dreams/20260420-142000Z-adversarial-inputs/` - Calling apply_bulk_discount twice on the same items compounds the discount: first call gives 0.90×, second gives 0.81×, third gives 0.729×. No idempotency guard. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ - Only applies discount to items with qty >= 5. Items with qty 4 or below are untouched, even if the total quantity across all items exceeds 5. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ - Uses round(price * 0.90, 2) which can produce floating-point artifacts on certain prices. For example, 33.33 * 0.90 = 29.997 → rounds to 30.0, not 29.997. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ ## Cross-References diff --git a/examples/coupon-demo/.shadow/test_cart.py.md b/examples/coupon-demo/.shadow/test_cart.py.md index ffcdbe4..1aff2ca 100644 --- a/examples/coupon-demo/.shadow/test_cart.py.md +++ b/examples/coupon-demo/.shadow/test_cart.py.md @@ -5,35 +5,35 @@ ## File-Level - Uses a manual `if __name__ == "__main__"` test runner, not pytest or unittest. Tests are plain functions with assert statements and no setup/teardown. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ - No test isolation: COUPON_CACHE is a shared global. test_coupon populates the cache with "SAVE20", and test_bulk_then_coupon populates it with "HALF". If test order changes or tests are rerun in the same process, cached values persist from earlier tests. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ Also involves: `cart.py::COUPON_CACHE` - No negative-path tests: no test for invalid coupon codes, negative prices, empty items, or the case-sensitivity mismatch between validate_coupon and calculate_total. - _(verified, source: exploration, labels: [feature-gap])_ + _(verified, source: exploration, labels: [feature-gap], citation_score: 0)_ ## `test_basic_total` - Verifies: 25.00 × 3 = 75.00 subtotal + 8% tax = 81.00. Tests the simplest happy path with no coupon. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ ## `test_coupon` - Verifies SAVE20 on 75.00 subtotal: 75.00 − 20% = 60.00 + 4.80 tax = 64.80. The 75.00 subtotal exceeds SAVE20's min_total of 50. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ - Only tests with uppercase coupon code "SAVE20". Does not test lowercase, revealing nothing about the case-normalization bug. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ Also involves: `cart.py::calculate_total` ## `test_bulk_then_coupon` - Tests the full pipeline: apply_bulk_discount mutates items (25.00 → 22.50), then calculate_total applies HALF coupon on 112.50 subtotal → 56.25 + 4.50 tax = 60.75. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ - Creates fresh items list, avoiding the mutation-persistence issue. But if this test's items object were reused in a subsequent test, prices would already be 22.50, not 25.00. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ Also involves: `inventory.py::apply_bulk_discount`, `cart.py::calculate_total` - Imports apply_bulk_discount inside the function body (lazy import), unlike the module-level import of calculate_total. This is inconsistent but functionally irrelevant. - _(verified, source: exploration, labels: [tech-debt])_ + _(verified, source: exploration, labels: [tech-debt], citation_score: 0)_ ## Cross-References diff --git a/skills/shadow-frog-dream/SKILL.md b/skills/shadow-frog-dream/SKILL.md index c643813..ce020a4 100644 --- a/skills/shadow-frog-dream/SKILL.md +++ b/skills/shadow-frog-dream/SKILL.md @@ -612,6 +612,15 @@ Follow the dedup and writing rules in `/shadow-frog`. Dream discoveries are typically `source: exploration`. Mark `verified` when confirmed by running code; `uncertain` if not fully testable. +New discoveries start with visible `citation_score: 0`. Keep that field in +the corresponding manifest entry as well. Existing knowledge that informed +the task can be cited once through the core increment helper or reported to +the coordinator for serialized updates. Do not count merely enumerated entries. +Scores are approximate within the relevant checkout; reconciliation preserves +the larger score when the same claim is supplied again, rather than summing +inherited counts. Counter-only branch edits are not imported unless represented +in a matching manifest discovery; they do not justify unsupported `op` values. + #### How to Append Find the `##`/`###` heading for the symbol, then: @@ -642,7 +651,7 @@ Example: ``` - /api/upload accepts paths from request body without normalization, allowing `../` traversal into /etc/. - _(verified, source: exploration, labels: [bug, security])_ + _(verified, source: exploration, labels: [bug, security], citation_score: 0)_ Dream report: `_dreams/20260518-161200Z-upload-traversal/` ``` @@ -682,7 +691,7 @@ Per-file discoveries should reference the dream report: ``` - Retrying with exponential backoff recovers from 99% of transient errors, but must exclude 4xx or it retries bad requests for 30s. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ Dream report: `_dreams/20250612-143012Z-retry-logic/` ``` @@ -707,6 +716,7 @@ After shadow writes, create `.shadow/_dreams/$DREAM_ID/manifest.json`: "status": "verified", "source": "exploration", "labels": ["bug"], + "citation_score": 0, "also_involves": ["src/parsers/utils.py::unescape"], "dream_report": "_dreams//" } @@ -731,6 +741,9 @@ reject any discovery whose `op` is not `add`. To revise or contradict an existing discovery, run a meditate session against main's `.shadow/` instead of trying to do it from a dream branch. +`citation_score` on per-file or cross-cutting manifest entries is a nonnegative +integer, defaulting to 0. It is a reuse hint, never a verification or trust signal. + **Hard gate — discoveries must be mirrored into per-file shadows.** The reconciler merges `manifest.json` entries into main directly (so discoveries are not lost at merge time), but the branch's per-file shadows must ALSO be @@ -904,7 +917,7 @@ rich area and the second half keeps digging there instead of spreading. 2. Each agent gets its own branch (inherently isolated) 3. Each agent writes its own manifest in its `$DREAM_ID/` directory 4. Do NOT write to main or shared files (`_index.md`, `state.json`) -5. Do NOT update metadata — reconciled post-dream by orchestrator +5. Do NOT update shared indexes or `state.json` — reconciled post-dream by orchestrator 6. Broad mode: fetch once and use the initial snapshot. Coherent mode: after a parent is pushed, the orchestrator refreshes that parent's ref/commit and branch map before launching its children. Siblings share the refreshed diff --git a/skills/shadow-frog-dream/dream-reconcile.py b/skills/shadow-frog-dream/dream-reconcile.py index 62253d3..3348df0 100755 --- a/skills/shadow-frog-dream/dream-reconcile.py +++ b/skills/shadow-frog-dream/dream-reconcile.py @@ -43,6 +43,20 @@ from datetime import datetime, timezone from pathlib import Path, PureWindowsPath +_bytecode = sys.dont_write_bytecode +sys.dont_write_bytecode = True +sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "shadow-frog")) +try: + from _citations import ( + CitationError, cross_metadata_line, metadata_score, set_metadata_score, + validate_score, + ) +except ImportError as exc: + raise SystemExit("ERROR: Missing core citation metadata parser; reinstall the full skill set") from exc +finally: + sys.path.pop(0) + sys.dont_write_bytecode = _bytecode + # Shared safety gate for `rm -rf `. Lives next to this script so # bash callers (dream-cleanup.sh, dream-gc.sh) and this module share ONE # source of truth for the "is this path safe to remove?" rules. Imported @@ -262,12 +276,14 @@ def _validate_manifest_paths(repo_root, dream_id, manifest): for filename in ('report.md', 'manifest.json', 'patch.diff'): _shadow_output_path(repo_root, f'_dreams/{dream_id}/{filename}', f"{label} {filename}") for index, disc in _manifest_entries(manifest, 'discoveries', dream_id): + validate_score(disc.get('citation_score', 0), f"{label} discoveries[{index}].citation_score") field = f"{label} discoveries[{index}].anchor" file_part = _anchor_file_part(disc.get('anchor', ''), field) if file_part is not None: _shadow_file_path(repo_root, file_part, field) for index, cross in _manifest_entries(manifest, 'cross_cutting', dream_id): field = f"{label} cross_cutting[{index}]" + validate_score(cross.get('citation_score', 0), f"{field}.citation_score") slug = cross.get('slug', '') if not slug: continue @@ -565,10 +581,12 @@ def _merge_meta(existing, new_status, new_source, new_labels): return status, source, labels, changed -def _format_meta_line(status, source, labels): +def _format_meta_line(status, source, labels, citation_score=0): + validate_score(citation_score) parts = [status, f'source: {source}'] if labels: parts.append(f"labels: [{', '.join(labels)}]") + parts.append(f"citation_score: {citation_score}") return f' _({", ".join(parts)})_\n' @@ -582,13 +600,11 @@ def merge_discovery_into_file(shadow_path, anchor_symbol, discovery, dream_id, * status = discovery.get('status', 'verified') source = discovery.get('source', 'exploration') labels = discovery.get('labels', []) + citation_score = validate_score(discovery.get('citation_score', 0)) also_involves = discovery.get('also_involves', []) # Build the discovery line - meta_parts = [status, f'source: {source}'] - if labels: - meta_parts.append(f"labels: [{', '.join(labels)}]") - meta_line = f' _({", ".join(meta_parts)})_' + meta_line = _format_meta_line(status, source, labels, citation_score).rstrip("\n") lines_to_add = [f'- {text}\n', f'{meta_line}\n'] if also_involves: @@ -639,9 +655,12 @@ def merge_discovery_into_file(shadow_path, anchor_symbol, discovery, dream_id, * return False merged = _merge_meta(existing_meta, status, source, labels) m_status, m_source, m_labels, changed = merged + previous_score = metadata_score(lines[meta_idx]) + merged_score = max(previous_score, citation_score) + changed = changed or merged_score != previous_score if not changed: return False - lines[meta_idx] = _format_meta_line(m_status, m_source, m_labels) + lines[meta_idx] = _format_meta_line(m_status, m_source, m_labels, merged_score) with open(shadow_path, 'w', encoding="utf-8") as f: f.writelines(lines) return True @@ -752,7 +771,7 @@ def add_cross_reference_backpointer(repo_root, file_part, slug, title, dream_id) return True -def _merge_refs_into_cross_file(cross_path, new_refs, *, repo_root): +def _merge_refs_into_cross_file(cross_path, new_refs, *, repo_root, citation_score=0): """Union new refs into an existing _cross/.md **Refs**: section. When two dreams use the same cross-cutting slug, the later one must not @@ -762,6 +781,7 @@ def _merge_refs_into_cross_file(cross_path, new_refs, *, repo_root): """ cross_path = _checked_shadow_destination(repo_root, cross_path, "cross-cutting destination") _validate_refs(repo_root, new_refs, "cross-cutting refs") + validate_score(citation_score) try: with open(cross_path, encoding="utf-8") as f: content = f.read() @@ -785,7 +805,16 @@ def _merge_refs_into_cross_file(cross_path, new_refs, *, repo_root): else: break to_add = [r for r in new_refs if r and r not in existing] - if not to_add: + score_changed = False + metadata_index = cross_metadata_line(lines) + if metadata_index is not None: + previous_score = metadata_score(lines[metadata_index]) + if citation_score > previous_score: + lines[metadata_index] = set_metadata_score(lines[metadata_index], citation_score) + score_changed = True + elif citation_score: + raise CitationError(f"{cross_path}: restore discovery metadata before merging citation scores") + if not to_add and not score_changed: return False lines[block_end:block_end] = [f'- `{r}`' for r in to_add] with open(cross_path, 'w', encoding="utf-8") as f: @@ -855,8 +884,10 @@ def merge_discoveries(repo_root, manifests, dry_run=False): f"**Category**: {cross.get('category', 'behavior')}\n" f"**Refs**:\n{refs_str}\n\n" f"**Discovery**: {cross.get('text', '')}\n\n" - f"_({cross.get('status', 'verified')}, " - f"source: {cross.get('source', 'exploration')})_\n" + + _format_meta_line( + cross.get('status', 'verified'), cross.get('source', 'exploration'), + cross.get('labels', []), cross.get('citation_score', 0), + ).lstrip() ) with open(cross_path, 'w', encoding="utf-8") as f: f.write(content) @@ -865,7 +896,10 @@ def merge_discoveries(repo_root, manifests, dry_run=False): # Cross file already exists (e.g. a prior dream used the same # slug). Union our refs into its **Refs**: block so it stays # consistent with the back-pointers added below. - if _merge_refs_into_cross_file(cross_path, refs, repo_root=repo_root): + if _merge_refs_into_cross_file( + cross_path, refs, repo_root=repo_root, + citation_score=cross.get('citation_score', 0), + ): merged_count += 1 else: skipped_count += 1 @@ -2080,6 +2114,6 @@ def main(): if __name__ == '__main__': try: main() - except (CoherentLineageError, UnsafeShadowPath) as exc: + except (CoherentLineageError, UnsafeShadowPath, CitationError) as exc: print(f"ERROR: {exc}", file=sys.stderr) sys.exit(1) diff --git a/skills/shadow-frog-dream/dream-tools.py b/skills/shadow-frog-dream/dream-tools.py index c29c41d..8e1da95 100644 --- a/skills/shadow-frog-dream/dream-tools.py +++ b/skills/shadow-frog-dream/dream-tools.py @@ -84,7 +84,7 @@ def _source_files() -> dict[str, Path]: raise ValueError(f"Tooling assets must be regular files: {path}") sources[f"{directory.name}/{path.name}"] = path required = { - "shadow-frog/SKILL.md", "shadow-frog/_coherence.py", + "shadow-frog/SKILL.md", "shadow-frog/_coherence.py", "shadow-frog/_citations.py", "shadow-frog-dream/SKILL.md", "shadow-frog-dream/dream-tools.py", "shadow-frog-dream/_worktree_safety.py", *(f"shadow-frog-dream/{name}" for name in TOOLS.values()), diff --git a/skills/shadow-frog-dream/dream-validate.py b/skills/shadow-frog-dream/dream-validate.py index 1b71823..09520b3 100644 --- a/skills/shadow-frog-dream/dream-validate.py +++ b/skills/shadow-frog-dream/dream-validate.py @@ -38,6 +38,7 @@ sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "shadow-frog")) try: from _coherence import MODES, validate_connection + from _citations import CitationError, validate_score except ImportError as exc: raise SystemExit("ERROR: Missing shared shadow-frog/_coherence.py; reinstall the full skill set") from exc finally: @@ -219,6 +220,10 @@ def main(): f'{type(disc).__name__}.' ) continue + try: + validate_score(disc.get('citation_score', 0), f"discoveries[{i}].citation_score") + except CitationError as exc: + errors.append(str(exc)) op = (disc.get('op') or 'add').lower() if op != 'add': errors.append( @@ -227,6 +232,13 @@ def main(): f'"add"), or split this into a meditate session.' ) + for i, cross in enumerate(manifest.get('cross_cutting', []) or []): + if isinstance(cross, dict): + try: + validate_score(cross.get('citation_score', 0), f"cross_cutting[{i}].citation_score") + except CitationError as exc: + errors.append(str(exc)) + # 10. Discoveries must be mirrored into per-file shadows on the dream # branch. The reconciler reads manifest entries directly when merging # into main, so the discoveries themselves are NOT lost — but the diff --git a/skills/shadow-frog-init/SKILL.md b/skills/shadow-frog-init/SKILL.md index df52a93..e8a5561 100644 --- a/skills/shadow-frog-init/SKILL.md +++ b/skills/shadow-frog-init/SKILL.md @@ -147,6 +147,8 @@ _No preferences recorded yet._ This file stores project-wide user preferences and conventions that are not tied to any specific file or symbol. It is populated by `/shadow-frog-update` when the user shares general directives. +New preference/discovery entries use `citation_score: 0`; empty placeholders +are not knowledge entries and do not have scores. ### 6. Create `_meta/state.json` diff --git a/skills/shadow-frog-meditate/SKILL.md b/skills/shadow-frog-meditate/SKILL.md index 9a65ceb..6ecf24d 100644 --- a/skills/shadow-frog-meditate/SKILL.md +++ b/skills/shadow-frog-meditate/SKILL.md @@ -55,7 +55,7 @@ orchestrator can auto-apply resolutions. This is critical for automation — prose recommendations require manual interpretation. ``` -{"action": "merge", "file": "src/auth.py.md", "symbol": "authenticate_user", "keep": "- silently returns None on expired tokens...", "remove": "- returns None when token expires...", "merged": "- authenticate_user() silently returns None on expired tokens instead of raising. 3 of 7 callers don't check.\n _(verified, source: exploration)_", "reason": "duplicate: same claim, different wording"} +{"action": "merge", "file": "src/auth.py.md", "symbol": "authenticate_user", "keep": "- silently returns None on expired tokens...", "remove": "- returns None when token expires...", "merged": "- authenticate_user() silently returns None on expired tokens instead of raising. 3 of 7 callers don't check.\n _(verified, source: exploration, citation_score: 0)_", "reason": "duplicate: same claim, different wording"} {"action": "merge", "file": "src/db.py.md", "symbol": "connect", "keep": "- connection pool exhaustion...", "remove": "- pool runs out...", "merged": "...", "reason": "near-duplicate: first extends second"} {"action": "conflict", "file": "src/auth.py.md", "symbol": "validate_token", "entry_a": "- raises ValueError...", "entry_b": "- returns False...", "resolution": "verified_a", "reason": "code inspection: line 42 raises ValueError"} {"action": "conflict", "file": "src/cache.py.md", "symbol": "invalidate", "entry_a": "...", "entry_b": "...", "resolution": "escalate", "reason": "both claims have evidence, needs user input"} @@ -97,6 +97,7 @@ Combine into a single discovery: - Keep the **stronger** trust: `source: user` > `source: interaction` > `source: exploration` - Keep the **stronger** status: `verified` > `uncertain` > `refuted` - Merge `Also involves:` refs (union of both) +- Preserve the larger `citation_score`; do not sum counts that may share history. - Delete the weaker entry ### Near-Duplicates → Absorb @@ -105,6 +106,7 @@ The broader discovery absorbs the narrower one: - Expand the broader entry to include any extra detail from the narrower - Delete the narrower entry - Preserve the stronger trust/status between the two +- Preserve the larger citation score for the merged knowledge. ### Conflicts → Investigate @@ -279,6 +281,8 @@ malformed discovery is worse than a duplicate; it breaks the viewer parser and downstream agents. Meditate-specific rules: +- Keep visible citation scores when rewording or moving the same knowledge; + a new behavioral claim starts at 0. Counts never override trust or correctness. - When merging, take the **union** of `labels: [...]` from both entries. - When merging, preserve every `Also involves: file::symbol` from both entries (union, not intersection). diff --git a/skills/shadow-frog-update/SKILL.md b/skills/shadow-frog-update/SKILL.md index 05190e0..941e954 100644 --- a/skills/shadow-frog-update/SKILL.md +++ b/skills/shadow-frog-update/SKILL.md @@ -72,13 +72,13 @@ and conventions. Representative examples: Write as: ```markdown - - _(verified, source: user)_ + _(verified, source: user, citation_score: 0)_ ``` For knowledge emerging from collaborative work (debugging, refactoring, test failures): ```markdown - - _(verified, source: interaction)_ + _(verified, source: interaction, citation_score: 0)_ ``` Anchor to the specific `file::symbol`. `source: user` and `source: interaction` @@ -187,6 +187,10 @@ See `/shadow-frog` § Discovery Format for the verbatim per-file, cross-cutting, and preference formats. Rules to keep in mind during update sessions: +- Start new claims/preferences with `citation_score: 0`; preserve the score + when updating the same knowledge. Citation-only edits do not add discoveries. +- Record deliberately consulted existing entries once per task using the core + `shadow-cite.py` helper, not a retrieval requirement or a second hidden store. - Be behavioral: "silently returns None on expired tokens" not "handles token expiration" - `source: user` and `source: interaction` → always `verified`, use user's own words - `source: exploration` → mark `uncertain` unless verified by code reading or tests diff --git a/skills/shadow-frog-viewer/SKILL.md b/skills/shadow-frog-viewer/SKILL.md index 3c0e68e..b5f9d85 100644 --- a/skills/shadow-frog-viewer/SKILL.md +++ b/skills/shadow-frog-viewer/SKILL.md @@ -1,7 +1,7 @@ --- name: shadow-frog-viewer description: >- - Browse and query the shadow knowledge base: overview, search for files + Help users browse and visualize the shadow knowledge base: overview, search for files or symbols or text, view preferences, or see recent discoveries. Invoke when the user wants to see what's in the shadow, get an overview, or find specific knowledge. @@ -12,7 +12,9 @@ scripts: # ShadowFrog Viewer -Query and browse `.shadow/` content. Prerequisite: `.shadow/` exists. +User-facing inspection of `.shadow/` content. Prerequisite: `.shadow/` exists. +Agents navigate the Markdown files/symbols directly; this viewer is not a +required retrieval interface. ## Primary: Python Helper Script @@ -40,6 +42,13 @@ python3 .claude/skills/shadow-frog-viewer/shadow-viewer.py [options] No arguments defaults to `--summary`. +Discovery views display the visible Markdown `citation_score` (missing means 0). +Search/label/top ordering uses it only after source trust and verification status; +refuted claims remain last. Reading, searching, and automatic previews do not +increment counts. The agent explicitly records deliberately consulted entries +with the core `shadow-cite.py` helper once per task. `--recent` is based on shadow +file modification time, which includes citation updates, not discovery creation time. + ### Options | Flag | Effect | diff --git a/skills/shadow-frog-viewer/shadow-viewer.py b/skills/shadow-frog-viewer/shadow-viewer.py index 3516b36..e379cdf 100755 --- a/skills/shadow-frog-viewer/shadow-viewer.py +++ b/skills/shadow-frog-viewer/shadow-viewer.py @@ -32,11 +32,16 @@ from pathlib import Path -_DISCOVERY_META_RE = re.compile( - r"_\((\w+),\s*source:\s*(\w+)" - r"(?:,\s*labels:\s*\[([^\]]*)\])?" - r"\)_" -) +_bytecode = sys.dont_write_bytecode +sys.dont_write_bytecode = True +sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "shadow-frog")) +try: + from _citations import DISCOVERY_META_RE as _DISCOVERY_META_RE, PREFERENCE_META_RE, CitationError, metadata_score +except ImportError as exc: + raise SystemExit("ERROR: Missing core citation metadata parser; reinstall the full skill set") from exc +finally: + sys.path.pop(0) + sys.dont_write_bytecode = _bytecode def warn(msg): @@ -71,7 +76,7 @@ def parse_discovery(line, continuation_lines=None): warn(f"parse_discovery: bad line input ({type(line).__name__}): {e}") return {"text": str(line) if line else ""} - meta = {} + meta = {"citation_score": 0} full_text = text if continuation_lines: @@ -83,6 +88,7 @@ def parse_discovery(line, continuation_lines=None): if m: meta["status"] = m.group(1) meta["source"] = m.group(2) + meta["citation_score"] = int(m.group(4) or 0) if m.group(3): meta["labels"] = [ l.strip() @@ -91,9 +97,12 @@ def parse_discovery(line, continuation_lines=None): ] else: # _(source: type)_ (preferences format) - m2 = re.match(r"_\(source:\s*(\w+)\)_", stripped) + m2 = PREFERENCE_META_RE.fullmatch(stripped) if m2: meta["source"] = m2.group(1) + meta["citation_score"] = int(m2.group(2) or 0) + elif stripped.startswith("_(") and "citation_score:" in stripped: + warn("Invalid citation_score metadata; repair it before recording citations") elif stripped.startswith("Also involves:"): refs = re.findall(r"`([^`]+)`", stripped) meta["also_involves"] = refs @@ -304,7 +313,7 @@ def parse_cross_cutting(shadow_dir): warn(f"Cannot read cross-cutting file {f}: {e}") continue - entry = {"slug": f.stem, "file": str(f.name)} + entry = {"slug": f.stem, "file": str(f.name), "citation_score": 0} try: # Title @@ -343,6 +352,7 @@ def parse_cross_cutting(shadow_dir): if m: entry["status"] = m.group(1) entry["source"] = m.group(2) + entry["citation_score"] = int(m.group(4) or 0) if m.group(3): entry["labels"] = [ l.strip() @@ -447,6 +457,17 @@ def collect_all_discoveries(shadow_dir): # --- View Functions --- +def _citation_rank(item): + if item.get("status") == "refuted": + trust = 5 + elif item.get("source") == "user": + trust = 0 + elif item.get("source") == "interaction": + trust = 1 + else: + trust = {"verified": 2, "uncertain": 3}.get(item.get("status"), 4) + return trust, -item.get("citation_score", 0) + def view_summary(shadow_dir): """Overview + detailed statistics. @@ -620,6 +641,7 @@ def view_search(shadow_dir, query): "text": text, "status": d.get("status", "?"), "source": d.get("source", "?"), + "citation_score": d.get("citation_score", 0), "also_involves": d.get("also_involves", []), "match": ( "file" if file_name_hit else @@ -680,11 +702,12 @@ def view_search(shadow_dir, query): for file, discs in sorted(by_file.items()): print(f"\n{file} ({len(discs)} matches)") print("-" * (len(file) + 15)) - for d in discs: + for d in sorted(discs, key=_citation_rank): sym = d["symbol"] print(f" {file}::{sym}") print(f" {d['text'][:120]}") print(f" ({d['status']}, source: {d['source']})") + print(f" citation_score: {d['citation_score']}") if d.get("also_involves"): print( f" Also involves: " @@ -698,13 +721,14 @@ def view_search(shadow_dir, query): try: print(f"\nCross-cutting ({len(cross_matches)} matches)") print("-" * 30) - for e in cross_matches: + for e in sorted(cross_matches, key=_citation_rank): title = e.get("title", e.get("slug", "?")) cat = e.get("category", "?") status = e.get("status", "?") source = e.get("source", "?") print(f"\n {title}") print(f" Category: {cat} | {status}, source: {source}") + print(f" citation_score: {e.get('citation_score', 0)}") print(f" Refs: {', '.join(e.get('refs', [])[:5])}") if e.get("discovery"): print(f" {e['discovery'][:120]}") @@ -716,8 +740,9 @@ def view_search(shadow_dir, query): try: print(f"\nPreferences ({len(pref_matches)} matches)") print("-" * 30) - for p in pref_matches: + for p in sorted(pref_matches, key=_citation_rank): print(f" [{p.get('source', '?')}] {p['text'][:120]}") + print(f" citation_score: {p.get('citation_score', 0)}") except Exception as e: warn(f"Search: failed to render preference results: {e}") @@ -736,10 +761,11 @@ def view_prefs(shadow_dir): print(f"Project Preferences ({len(prefs)} total)") print("=" * 40) - for p in prefs: + for p in sorted(prefs, key=_citation_rank): try: source = p.get("source", "?") print(f"\n [{source}] {p['text']}") + print(f" citation_score: {p.get('citation_score', 0)}") except Exception as e: warn(f"Failed to render preference: {e}") @@ -772,6 +798,7 @@ def view_labels(shadow_dir, label_filter): "status": entry.get("status", "?"), "source": entry.get("source", "?"), "labels": entry["labels"], + "citation_score": entry.get("citation_score", 0), }) except Exception as e: warn(f"Failed to include cross-cutting in label search: {e}") @@ -806,7 +833,7 @@ def view_labels(shadow_dir, label_filter): continue print(f"\n[{lbl}] ({len(discs)} discoveries)") print("-" * 30) - for d in discs: + for d in sorted(discs, key=_citation_rank): try: sym = d.get("symbol", "?") src_file = d.get("file", "?") @@ -814,6 +841,7 @@ def view_labels(shadow_dir, label_filter): print(f" {d['text'][:120]}") print(f" ({d.get('status', '?')}, " f"source: {d.get('source', '?')})") + print(f" citation_score: {d.get('citation_score', 0)}") all_labels = d.get("labels", []) other = [l for l in all_labels if l.lower() != lbl] if other: @@ -850,6 +878,7 @@ def view_recent(shadow_dir, count=10): "text": d.get("text", ""), "status": d.get("status", "?"), "source": d.get("source", "?"), + "citation_score": d.get("citation_score", 0), "mtime": mtime, }) except Exception as e: @@ -877,6 +906,7 @@ def view_recent(shadow_dir, count=10): "text": e.get("discovery", e.get("title", "")), "status": e.get("status", "?"), "source": e.get("source", "?"), + "citation_score": e.get("citation_score", 0), "mtime": mtime, }) except Exception as e: @@ -897,6 +927,7 @@ def view_recent(shadow_dir, count=10): "text": p.get("text", ""), "status": "-", "source": p.get("source", "?"), + "citation_score": p.get("citation_score", 0), "mtime": mtime, }) except Exception as e: @@ -928,6 +959,7 @@ def view_recent(shadow_dir, count=10): print(f" {item.get('text', '')[:120]}") print(f" ({item.get('status', '?')}, " f"source: {item.get('source', '?')})") + print(f" citation_score: {item.get('citation_score', 0)}") except Exception as e: warn(f"Recent: failed to render item: {e}") @@ -940,9 +972,9 @@ def view_top(shadow_dir, file_path, labels_filter, limit, max_chars): file. Pulls from both the per-file shadow and any _cross/ entries whose refs touch this file. - Ranking: verified > uncertain > refuted; within a tier, source - order is preserved. Output is hard-capped at max_chars (the trailing - "(...)" marker still fits). + Source trust/status precedes the visible citation score; equal scores + preserve document order. Reading does not increment citations. + Output is hard-capped at max_chars (the trailing "(...)" marker still fits). """ norm = file_path.strip() if norm.startswith("./"): @@ -964,6 +996,8 @@ def view_top(shadow_dir, file_path, labels_filter, limit, max_chars): "anchor": d.get("symbol") or "file-level", "text": d.get("text", "").strip(), "status": d.get("status", "?"), + "source": d.get("source", "?"), + "citation_score": d.get("citation_score", 0), "labels": sorted(disc_labels), }) except Exception as e: @@ -982,6 +1016,8 @@ def view_top(shadow_dir, file_path, labels_filter, limit, max_chars): "anchor": f"_cross/{entry.get('file', entry.get('slug', '?'))}", "text": (entry.get("discovery") or entry.get("title") or "").strip(), "status": entry.get("status", "?"), + "source": entry.get("source", "?"), + "citation_score": entry.get("citation_score", 0), "labels": sorted(cross_labels), }) except Exception as e: @@ -994,8 +1030,7 @@ def view_top(shadow_dir, file_path, labels_filter, limit, max_chars): ) return - tier = {"verified": 0, "uncertain": 1, "refuted": 2} - candidates.sort(key=lambda d: tier.get(d.get("status", "?"), 3)) + candidates.sort(key=_citation_rank) shown = candidates[:limit] header = ( @@ -1008,7 +1043,7 @@ def view_top(shadow_dir, file_path, labels_filter, limit, max_chars): anchor = d["anchor"] text = d["text"].replace("\n", " ").strip() lines.append( - f"- [{labels}] `{anchor}` ({d['status']}): {text}" + f"- [{labels}] `{anchor}` ({d['status']}, citation_score: {d['citation_score']}): {text}" ) out = "\n".join(lines) @@ -1069,6 +1104,13 @@ def v(path, line, kind, msg): also_involves_re = re.compile(r"^\s*Also involves:\s*(.+)$", re.I) file_sym_re = re.compile(r"`([^`]+::[^`]+)`") + def check_citation(path, line_number, line): + if line.strip().startswith("_(") and "citation_score:" in line: + try: + metadata_score(line) + except CitationError as exc: + v(path, line_number, "citation_score", str(exc)) + for shadow_path in get_all_shadow_files(shadow_dir): try: rel = shadow_path.relative_to(shadow_dir) @@ -1084,6 +1126,7 @@ def v(path, line, kind, msg): declared = set() for ln, raw in enumerate(text.split("\n"), 1): line = raw.rstrip() + check_citation(rel, ln, line) heading = md_heading_re.match(line) if heading: @@ -1170,6 +1213,8 @@ def v(path, line, kind, msg): continue rel_cf = cf.relative_to(shadow_dir) + for line_number, line in enumerate(text.splitlines(), 1): + check_citation(rel_cf, line_number, line) # Category enum check cat_m = re.search(r"\*\*Category\*\*:\s*(.+)", text) @@ -1235,6 +1280,14 @@ def v(path, line, kind, msg): f"Cross-References does not link back to " f"_cross/{slug}.md") + prefs_path = shadow_dir / "_prefs.md" + if prefs_path.is_file(): + try: + for line_number, line in enumerate(prefs_path.read_text(encoding="utf-8").splitlines(), 1): + check_citation("_prefs.md", line_number, line) + except (OSError, UnicodeError) as exc: + v("_prefs.md", 0, "unreadable", str(exc)) + # Output if not violations: print(f"✓ Invariants OK ({len(per_file_xref_targets)} per-file " diff --git a/skills/shadow-frog/SKILL.md b/skills/shadow-frog/SKILL.md index 9074571..1817ec6 100644 --- a/skills/shadow-frog/SKILL.md +++ b/skills/shadow-frog/SKILL.md @@ -10,7 +10,9 @@ description: >- to create it, shadow-frog-update to refresh it, shadow-frog-dream for autonomous experiments, shadow-frog-nap for lightweight feature-task ideation, shadow-frog-meditate for shadow hygiene, - or shadow-frog-viewer to browse it. + or shadow-frog-viewer for user-facing browsing and visualization. +scripts: + - shadow-cite.py --- # ShadowFrog @@ -42,6 +44,56 @@ to that code location. specific file): write it to `_prefs.md` immediately. 7. **After code changes**: run `/shadow-frog-update` +## Record Revisited Knowledge + +Continue navigating directly to shadow files and symbol headings. Each +discovery or preference has one visible `citation_score`: a nonnegative integer, +initially 0. Missing scores also mean 0. It is approximate, agent-reported +revisit frequency, not confidence or proof of usefulness. + +When updating an existing shadow, preserve all knowledge and existing scores: +missing `citation_score` fields mean `0` and are added on the first recorded +citation, so no reinitialization is required. + +After deliberately consulting an entry, record one citation for it **per task**. +Do not count every entry merely because its file was opened, repeat the count +on rereads, or count automatic previews that you did not use. Batch updates when +practical; the agent/coordinator tracks which entries it already counted. + +Use the small increment helper to avoid competing score edits. It does not +retrieve knowledge, create a database, or require opaque IDs: + +```text +python .github/skills/shadow-frog/shadow-cite.py .shadow/src/auth.py.md --symbol UserAuth.validate --text "Rejects expired tokens." +python .github/skills/shadow-frog/shadow-cite.py .shadow/_prefs.md --text "Keep public APIs stable." +``` + +For Claude Code, use `.claude/skills/`. Copy the exact claim text already read; +omit the bullet marker and metadata. Line wrapping is joined, but spaces inside +literals remain significant. +Use `--symbol File-Level` for file-wide entries; omit `--symbol` for preferences +and `_cross/` files. Repeat `--text` to update several entries in one section +atomically. `--shadow-dir` identifies an explicit nonstandard shadow root. +Unknown/ambiguous claims and invalid scores fail with corrective feedback. +Fenced examples and unrelated headings are not citation targets. + +The helper locks only its target file and publishes the score changes atomically. +Coordinate citation writes with ordinary knowledge edits; unrelated editors do +not participate in this lock. Subagents should report consulted entries to their +coordinator rather than race full-file rewrites. If a lock survives interruption, +confirm its writer stopped before removing that specific `.citation.lock` file. + +Scores stay with the Markdown and travel through Git. They are not exact global +counts across branches/clones: when merging the same knowledge, keep the larger +score rather than summing inherited counts. Keep scores when rewording/moving +the same claim; a genuinely new claim starts at 0. Citation-only edits do not +create discoveries or change discovery totals. + +Treat higher scores as a secondary hint after relevance and trust. Never omit +a relevant user constraint or new discovery because its score is low. Native +searches and targeted reads remain the normal tools; `/shadow-frog-viewer` +serves user browsing and does not automatically record citations. + ## Directory Layout ``` @@ -81,7 +133,7 @@ The symbol name is the stable anchor. ## File-Level - This module has no __all__ — all top-level names are public. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ ## `class UserAuth` @@ -89,12 +141,12 @@ The symbol name is the stable anchor. - Catches ALL exceptions and returns False — swallows connection errors, making network failures look like invalid tokens. - _(verified, source: exploration)_ + _(verified, source: exploration, citation_score: 0)_ ## `authenticate_user` - Silently returns None on expired tokens. Callers must check. - _(verified, source: exploration, labels: [bug])_ + _(verified, source: exploration, labels: [bug], citation_score: 0)_ Also involves: `src/middleware.py::require_auth` ## Cross-References @@ -134,7 +186,7 @@ Examples: **Discovery**: All database access goes through a connection pool that silently reconnects on failure. First request after DB restart is slow (~2s). -_(verified, source: exploration)_ +_(verified, source: exploration, citation_score: 0)_ ``` ## Preferences File (`_prefs.md`) @@ -146,19 +198,19 @@ specific file or symbol. These guide all agent work across the codebase. # Preferences - No backward compatibility — only keep the latest code, no shims or aliases. - _(source: user)_ + _(source: user, citation_score: 0)_ - Use snake_case for all Python function and variable names. - _(source: user)_ + _(source: user, citation_score: 0)_ - Prefer small, focused PRs over large sweeping changes. - _(source: interaction)_ + _(source: interaction, citation_score: 0)_ ``` Format: ``` - - _(source: )_ + _(source: , citation_score: 0)_ ``` Preferences are always trusted (same rank as `source: user`). They don't @@ -174,21 +226,21 @@ When to write to `_prefs.md` vs per-file shadow vs `_cross/`: Per-file discoveries (no IDs — anchored by their `file::symbol` heading): ``` - - _(, source: )_ + _(, source: , citation_score: 0)_ Also involves: `file::symbol`, `file::symbol` ``` With labels (optional — only when the discovery is actionable): ``` - - _(, source: , labels: [bug, security])_ + _(, source: , labels: [bug, security], citation_score: 0)_ Also involves: `file::symbol` ``` With dream report link (optional — only for experiment-derived discoveries): ``` - - _(, source: )_ + _(, source: , citation_score: 0)_ Dream report: `_dreams//` ``` @@ -202,7 +254,7 @@ Cross-cutting discoveries (one per `_cross/.md` file): **Discovery**: -_(, source: )_ +_(, source: , citation_score: 0)_ ``` Slug naming: use descriptive kebab-case derived from the title. @@ -227,7 +279,7 @@ Omit labels entirely for pure observational knowledge. Labels go in the metadata line: ``` -_(verified, source: exploration, labels: [bug])_ +_(verified, source: exploration, labels: [bug], citation_score: 0)_ ``` Cross-cutting discoveries can also have labels — add them to the metadata line. @@ -239,6 +291,7 @@ Cross-cutting discoveries can also have labels — add them to the metadata line - `source: user` — human stated it in conversation - `source: interaction` — emerged from collaborative work (debugging, refactoring) - `labels: [...]` — optional, actionable labels (see table above) +- `citation_score: N` — nonnegative integer after optional labels; new entries use 0 and omitted values mean 0 - `Also involves:` — `file::symbol` refs to other code locations (required if discovery touches other files) - `Dream report:` — optional, `_dreams//` link for experiment-derived discoveries - `Category` (cross-cutting only): pattern, behavior, edge-case, contract, performance, intent, warning, history, convention diff --git a/skills/shadow-frog/_citations.py b/skills/shadow-frog/_citations.py new file mode 100644 index 0000000..795b808 --- /dev/null +++ b/skills/shadow-frog/_citations.py @@ -0,0 +1,207 @@ +"""Visible Markdown citation metadata and serialized, exact-entry increments.""" + +from contextlib import contextmanager +import os +from pathlib import Path +import re +import stat +import tempfile +import time + + +DISCOVERY_META_RE = re.compile( + r"_\((\w+),\s*source:\s*(\w+)" + r"(?:,\s*labels:\s*\[([^\]]*)\])?" + r"(?:,\s*citation_score:\s*([0-9]+))?\)_" +) +PREFERENCE_META_RE = re.compile( + r"_\(source:\s*(\w+)(?:,\s*citation_score:\s*([0-9]+))?\)_" +) + + +class CitationError(ValueError): + """An entry cannot be safely identified or its score cannot be updated.""" + + +def validate_score(value, field="citation_score"): + if type(value) is not int or value < 0: + raise CitationError(f"{field} must be a nonnegative integer") + return value + + +def metadata_score(line): + stripped = line.strip() + discovery = DISCOVERY_META_RE.fullmatch(stripped) + preference = PREFERENCE_META_RE.fullmatch(stripped) + if discovery: + return int(discovery.group(4) or 0) + if preference: + return int(preference.group(2) or 0) + raise CitationError("Malformed metadata: use a nonnegative integer citation_score after optional labels") + + +def set_metadata_score(line, score): + """Change only the score field, preserving other text and line endings.""" + validate_score(score) + metadata_score(line) + if "citation_score:" in line: + return re.sub(r"(citation_score:\s*)[0-9]+", lambda match: match.group(1) + str(score), line, count=1) + closing = line.rfind(")_") + return line[:closing] + f", citation_score: {score}" + line[closing:] + + +def _symbol(value): + if value is None: + return None + value = value.strip().strip("`") + return re.sub(r"^(?:class|interface|enum|trait|struct|protocol|module) ", "", value) + + +def _claim(value): + # Join visual line wrapping, but do not collapse whitespace inside literals. + return re.sub(r"[ \t]*\r?\n[ \t]*", " ", value.strip()) + + +def _entries(lines, kind): + section = None + claim = None + fence = None + for index, line in enumerate(lines): + stripped = line.strip() + if fence is not None: + if claim is not None: + claim.append(stripped) + if re.fullmatch(re.escape(fence[0]) + "{" + str(len(fence)) + ",}", stripped): + fence = None + continue + opening = re.match(r"`{3,}|~{3,}", stripped) + if opening: + fence = opening.group(0) + if claim is not None: + claim.append(stripped) + continue + if re.match(r"#{1,6}\s", stripped): + heading = re.fullmatch(r"#{2,3}\s+(?:`(.+)`|(File-Level|Cross-References))", stripped) + section = _symbol(heading.group(1) or heading.group(2)) if heading else None + claim = None + continue + if kind == "cross" and stripped.startswith("**Discovery**:"): + claim = [stripped.partition(":")[2].strip()] + elif kind != "cross" and stripped.startswith("- "): + claim = [stripped[2:]] if kind == "preference" or section not in (None, "Cross-References") else None + elif claim is not None: + if stripped.startswith("_("): + try: + metadata_score(line) + except CitationError as exc: + raise CitationError(f"Line {index + 1}: {exc}") from exc + yield section if kind == "file" else None, _claim("\n".join(claim)), index + claim = None + elif stripped.startswith(("Also involves:", "Dream report:", "#")): + claim = None + elif stripped: + claim.append(stripped) + + +def cross_metadata_line(lines): + """Locate the single actual cross-cutting discovery's metadata, excluding examples.""" + entries = list(_entries(lines, "cross")) + if len(entries) > 1: + raise CitationError("Cross-cutting file contains multiple discoveries; resolve ambiguity before updating scores") + return entries[0][2] if entries else None + + +def _target(path, shadow_dir): + literal = Path(path).absolute() + if shadow_dir is None: + root = next((parent for parent in literal.parents if parent.name == ".shadow"), None) + if root is None: + raise CitationError("Provide a file under .shadow/ or an explicit --shadow-dir") + else: + root = Path(shadow_dir).absolute() + resolved_root = root.resolve() + if resolved_root != root.parent.resolve() / root.name: + raise CitationError("The shadow root is a filesystem alias; use a real shadow directory before recording citations") + resolved = literal.resolve() + if not resolved.is_relative_to(resolved_root) or not resolved.is_file(): + raise CitationError("Citation target must be an existing Markdown file inside the shadow directory") + relative = resolved.relative_to(resolved_root) + if relative.suffix != ".md" or relative.parts[0] in ("_meta", "_dreams", "_index.md"): + raise CitationError("Cite discovery or preference entries, not indexes or dream reports") + kind = "preference" if relative.as_posix() == "_prefs.md" else ( + "cross" if len(relative.parts) == 2 and relative.parts[0] == "_cross" else "file" + ) + return resolved, kind + + +@contextmanager +def _locked(path, timeout): + lock = path.with_name(path.name + ".citation.lock") + deadline = time.monotonic() + timeout + while True: + try: + descriptor = os.open(lock, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600) + break + except FileExistsError: + if time.monotonic() >= deadline: + raise CitationError( + f"Citation writer busy: {lock}. Retry later; after an interruption, " + "confirm its writer stopped before removing only that lock." + ) from None + time.sleep(min(0.02, max(0, deadline - time.monotonic()))) + try: + with os.fdopen(descriptor, "w", encoding="utf-8") as stream: + stream.write(f"pid={os.getpid()}\n") + yield + finally: + lock.unlink() + + +def record_citations(path, texts, *, symbol=None, shadow_dir=None, timeout=2.0): + """Increment each selected claim once; no retrieval, database, or hidden IDs.""" + path, kind = _target(path, shadow_dir) + if kind == "file" and not symbol: + raise CitationError("Per-file citations require --symbol (use File-Level for file-wide knowledge)") + if kind != "file" and symbol is not None: + raise CitationError("Preferences and cross-cutting discoveries do not use --symbol") + if not texts or any(not isinstance(text, str) or not text.strip() for text in texts): + raise CitationError("Supply --text with the exact discovery text already consulted") + targets = list(dict.fromkeys(_claim(text) for text in texts)) + selected_symbol = _symbol(symbol) + with _locked(path, timeout): + original = path.read_bytes() + lines = original.decode("utf-8").splitlines(keepends=True) + try: + entries = list(_entries(lines, kind)) + except CitationError as exc: + raise CitationError(f"{path}: {exc}") from exc + updates = [] + for text in targets: + matches = [ + index for entry_symbol, entry_text, index in entries + if entry_symbol == selected_symbol and entry_text == text + ] + if len(matches) != 1: + raise CitationError( + f"{path}: expected one matching entry for {symbol or kind} / {text!r}, " + f"found {len(matches)}. Re-read that entry; correct the text or resolve duplicates." + ) + index = matches[0] + before = metadata_score(lines[index]) + lines[index] = set_metadata_score(lines[index], before + 1) + updates.append({"text": text, "before": before, "after": before + 1}) + temporary = None + try: + with tempfile.NamedTemporaryFile(dir=path.parent, prefix=path.name + ".cite-", delete=False) as stream: + temporary = Path(stream.name) + stream.write("".join(lines).encode("utf-8")) + stream.flush() + os.fsync(stream.fileno()) + os.chmod(temporary, stat.S_IMODE(path.stat().st_mode)) + if path.read_bytes() != original: + raise CitationError("Shadow changed during citation update; coordinate writers and retry") + os.replace(temporary, path) + finally: + if temporary is not None and temporary.exists(): + temporary.unlink() + return updates diff --git a/skills/shadow-frog/shadow-cite.py b/skills/shadow-frog/shadow-cite.py new file mode 100755 index 0000000..114c970 --- /dev/null +++ b/skills/shadow-frog/shadow-cite.py @@ -0,0 +1,51 @@ +#!/usr/bin/env python3 +"""Record visits to existing Markdown discoveries after reading them normally. + +Example: + python shadow-cite.py .shadow/src/auth.py.md --symbol login --text "Rejects expired tokens." + +Repeat --text to cite several entries in the same file/section atomically. +Cite each deliberately consulted entry once per task. Repeated invocations +increment again; task-level deduplication belongs to the agent/coordinator. +""" + +import argparse +from pathlib import Path +import sys + + +_bytecode = sys.dont_write_bytecode +sys.dont_write_bytecode = True +sys.path.insert(0, str(Path(__file__).resolve().parent)) +try: + from _citations import record_citations +except ImportError as exc: + raise SystemExit("ERROR: Missing core citation helper; reinstall the full skill set") from exc +finally: + sys.path.pop(0) + sys.dont_write_bytecode = _bytecode + + +def main(): + for stream in (sys.stdout, sys.stderr): + if hasattr(stream, "reconfigure"): + stream.reconfigure(encoding="utf-8") + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("file", type=Path, help="The .shadow Markdown file already read") + parser.add_argument("--symbol", help="Exact source symbol, or File-Level (per-file shadows only)") + parser.add_argument("--text", action="append", required=True, help="Exact consulted claim; repeat for several") + parser.add_argument("--shadow-dir", type=Path, help="Explicit root for a nonstandard shadow location") + args = parser.parse_args() + try: + updates = record_citations(args.file, args.text, symbol=args.symbol, shadow_dir=args.shadow_dir) + except (ValueError, OSError, UnicodeError, RuntimeError) as exc: + print(f"ERROR: {exc}", file=sys.stderr) + return 1 + print(f"Recorded {len(updates)} citation(s) in {args.file}:") + for update in updates: + print(f" {update['before']} -> {update['after']}: {update['text']}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/conftest.py b/tests/conftest.py index 25cd182..d165c07 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -88,6 +88,11 @@ def coherence(repo_root): return _load_script(repo_root / "skills/shadow-frog/_coherence.py") +@pytest.fixture(scope="session") +def citations(repo_root): + return _load_script(repo_root / "skills/shadow-frog/_citations.py") + + @pytest.fixture(scope="session") def nap(repo_root): return _load_script(repo_root / "skills/shadow-frog-nap/nap.py") diff --git a/tests/skills/shadow_frog/test_citations.py b/tests/skills/shadow_frog/test_citations.py new file mode 100644 index 0000000..bad62ab --- /dev/null +++ b/tests/skills/shadow_frog/test_citations.py @@ -0,0 +1,281 @@ +"""Visible citation updates use real Markdown and per-file writer coordination.""" + +from pathlib import Path +import subprocess +import sys + +import pytest + +from tests.conftest import _load_script + + +CORE = Path(__file__).resolve().parents[3] / "skills/shadow-frog" + + +@pytest.fixture +def citations(): + return _load_script(CORE / "_citations.py") + + +def per_file(tmp_path, *, score="", newline="\n"): + path = tmp_path / ".shadow/src/auth.py.md" + path.parent.mkdir(parents=True) + text = ( + "# Shadow: src/auth.py\n\n## File-Level\n\n" + "- Importing opens no sockets.\n _(verified, source: exploration)_\n\n" + "## `class Auth`\n\n" + "- Construction leaves credentials untouched.\n _(verified, source: user)_\n\n" + "### `Auth.login`\n\n" + "- Key `a b` is distinct.\n _(verified, source: exploration, labels: [bug]" + score + ")_\n\n" + "- Rejects expired tokens.\n _(verified, source: interaction, citation_score: 4)_\n\n" + "## Cross-References\n\n- [auth](../_cross/auth.md)\n" + ) + path.write_bytes(text.replace("\n", newline).encode("utf-8")) + return path + + +def test_direct_read_then_increment_only_the_selected_visible_score(citations, tmp_path): + path = per_file(tmp_path) + before = path.read_text(encoding="utf-8") + assert "Key `a b`" in before + updates = citations.record_citations(path, ["Key `a b` is distinct."], symbol="Auth.login") + after = path.read_text(encoding="utf-8") + assert after == before.replace("labels: [bug])_", "labels: [bug], citation_score: 1)_") + assert updates[0]["before"] == 0 and updates[0]["after"] == 1 + assert sorted(p.name for p in path.parent.iterdir()) == ["auth.py.md"] + + +def test_batch_is_atomic_and_repeated_text_counts_once(citations, tmp_path): + path = per_file(tmp_path) + citations.record_citations( + path, ["Key `a b` is distinct.", "Rejects expired tokens.", "Rejects expired tokens."], + symbol="Auth.login", + ) + content = path.read_text(encoding="utf-8") + assert "citation_score: 1" in content and "citation_score: 5" in content + before = path.read_bytes() + with pytest.raises(citations.CitationError, match="found 0"): + citations.record_citations(path, ["Rejects expired tokens.", "Does not exist."], symbol="Auth.login") + assert path.read_bytes() == before + assert not list(path.parent.glob("*.citation.lock")) + + +def test_duplicate_claim_is_ambiguous_not_a_bulk_increment(citations, tmp_path): + path = per_file(tmp_path) + text = path.read_text(encoding="utf-8") + claim = "- Rejects expired tokens.\n _(verified, source: interaction, citation_score: 4)_\n" + path.write_text(text.replace(claim, claim + "\n" + claim), encoding="utf-8") + before = path.read_bytes() + with pytest.raises(citations.CitationError, match="found 2"): + citations.record_citations(path, ["Rejects expired tokens."], symbol="Auth.login") + assert path.read_bytes() == before + + +@pytest.mark.parametrize("symbol,text", [ + ("File-Level", "Importing opens no sockets."), + ("Auth", "Construction leaves credentials untouched."), +]) +def test_file_level_and_container_symbols(citations, tmp_path, symbol, text): + path = per_file(tmp_path) + updates = citations.record_citations(path, [text], symbol=symbol) + assert updates == [{"text": text, "before": 0, "after": 1}] + + +def test_literal_whitespace_is_not_normalized_away(citations, tmp_path): + path = per_file(tmp_path) + before = path.read_bytes() + with pytest.raises(citations.CitationError, match="found 0"): + citations.record_citations(path, ["Key `a b` is distinct."], symbol="Auth.login") + assert path.read_bytes() == before + + +@pytest.mark.parametrize("boundary", ["notes-heading", "fenced-example"]) +def test_non_discovery_sections_cannot_be_cited_as_previous_symbol(citations, tmp_path, boundary): + path = per_file(tmp_path) + sample = "- Example-only claim.\n _(verified, source: exploration, citation_score: 2)_\n" + if boundary == "notes-heading": + sample = "## Notes\n\n" + sample + else: + sample = "```markdown\n" + sample + "```\n" + path.write_text( + path.read_text(encoding="utf-8").replace("## Cross-References", sample + "\n## Cross-References"), + encoding="utf-8", + ) + before = path.read_bytes() + with pytest.raises(citations.CitationError, match="found 0"): + citations.record_citations(path, ["Example-only claim."], symbol="Auth.login") + assert path.read_bytes() == before + + +def test_discovery_after_fenced_example_keeps_its_actual_symbol(citations, tmp_path): + path = per_file(tmp_path) + fenced = ( + "```markdown\n## `WrongSymbol`\n\n- Example-only claim.\n" + " _(verified, source: exploration, citation_score: 2)_\n```\n\n" + "- Actual later claim.\n _(verified, source: user, citation_score: 0)_\n\n" + ) + path.write_text( + path.read_text(encoding="utf-8").replace("## Cross-References", fenced + "## Cross-References"), + encoding="utf-8", + ) + before = path.read_text(encoding="utf-8") + citations.record_citations(path, ["Actual later claim."], symbol="Auth.login") + assert path.read_text(encoding="utf-8") == before.replace( + "source: user, citation_score: 0", "source: user, citation_score: 1", + ) + + +def test_crlf_bom_and_existing_metadata_are_preserved(citations, tmp_path): + path = per_file(tmp_path, score=", citation_score: 7", newline="\r\n") + before = b"\xef\xbb\xbf" + path.read_bytes() + path.write_bytes(before) + citations.record_citations(path, ["Key `a b` is distinct."], symbol="Auth.login") + assert path.read_bytes() == before.replace(b"citation_score: 7", b"citation_score: 8") + + +def test_wrapped_claim_text_can_be_reported_without_copying_layout(citations, tmp_path): + path = per_file(tmp_path) + path.write_text( + path.read_text(encoding="utf-8").replace("Rejects expired tokens.", "Rejects expired\n tokens."), + encoding="utf-8", + ) + citations.record_citations(path, ["Rejects expired tokens."], symbol="Auth.login") + assert "citation_score: 5" in path.read_text(encoding="utf-8") + + +@pytest.mark.parametrize("kind", ["preference", "cross"]) +def test_preferences_and_cross_cutting_have_visible_scores(citations, tmp_path, kind): + path = tmp_path / ".shadow" / ("_prefs.md" if kind == "preference" else "_cross/shared.md") + path.parent.mkdir(parents=True) + if kind == "preference": + path.write_text("# Preferences\n\n- Keep the contract.\n _(source: user)_\n", encoding="utf-8") + else: + path.write_text( + "# Shared\n\n**Category**: contract\n**Refs**:\n- `src/a.py::run`\n\n" + "**Discovery**: Keep the contract.\n\n_(verified, source: exploration)_\n", + encoding="utf-8", + ) + citations.record_citations(path, ["Keep the contract."]) + assert "citation_score: 1" in path.read_text(encoding="utf-8") + + +def test_concurrent_processes_preserve_all_increments(citations, tmp_path): + path = per_file(tmp_path) + command = [ + sys.executable, str(CORE / "shadow-cite.py"), str(path), + "--symbol", "Auth.login", "--text", "Key `a b` is distinct.", + ] + processes = [subprocess.Popen(command, stdout=subprocess.PIPE, stderr=subprocess.PIPE) for _ in range(8)] + try: + for process in processes: + stdout, stderr = process.communicate(timeout=15) + assert process.returncode == 0, stderr.decode("utf-8") + finally: + for process in processes: + if process.poll() is None: + process.terminate() + process.communicate(timeout=10) + content = path.read_text(encoding="utf-8") + assert "labels: [bug], citation_score: 8" in content + assert "source: interaction, citation_score: 4" in content + assert sorted(p.name for p in path.parent.iterdir()) == ["auth.py.md"] + + +def test_existing_lock_fails_visibly_without_removing_it(citations, tmp_path): + path = per_file(tmp_path) + lock = path.with_name(path.name + ".citation.lock") + lock.write_text("other writer", encoding="utf-8") + before = path.read_bytes() + with pytest.raises(citations.CitationError, match="writer busy"): + citations.record_citations(path, ["Key `a b` is distinct."], symbol="Auth.login", timeout=0.02) + assert path.read_bytes() == before and lock.read_text() == "other writer" + + +def test_failed_publication_preserves_knowledge_and_cleans_temporary_files(citations, tmp_path, monkeypatch): + path = per_file(tmp_path) + before = path.read_bytes() + + def denied(*args): + raise PermissionError("simulated sharing violation") + + monkeypatch.setattr(citations.os, "replace", denied) + with pytest.raises(PermissionError, match="sharing violation"): + citations.record_citations(path, ["Rejects expired tokens."], symbol="Auth.login") + assert path.read_bytes() == before + assert sorted(p.name for p in path.parent.iterdir()) == ["auth.py.md"] + + +def test_detected_ordinary_editor_race_is_not_overwritten(citations, tmp_path, monkeypatch): + path = per_file(tmp_path) + before = path.read_bytes() + actual_fsync = citations.os.fsync + + def another_writer(fd): + path.write_bytes(before + b"\nAdditional knowledge from another writer.\n") + actual_fsync(fd) + + monkeypatch.setattr(citations.os, "fsync", another_writer) + with pytest.raises(citations.CitationError, match="changed during"): + citations.record_citations(path, ["Rejects expired tokens."], symbol="Auth.login") + assert path.read_bytes() == before + b"\nAdditional knowledge from another writer.\n" + assert sorted(p.name for p in path.parent.iterdir()) == ["auth.py.md"] + + +@pytest.mark.parametrize("score", [-1, 1.2, True, None, "3"]) +def test_json_score_requires_a_nonnegative_integer(citations, score): + with pytest.raises(citations.CitationError, match="nonnegative integer"): + citations.validate_score(score) + + +@pytest.mark.parametrize("raw", ["-1", "1.2", "true", "NaN", ""]) +def test_invalid_existing_score_is_not_reset(citations, tmp_path, raw): + path = per_file(tmp_path, score=f", citation_score: {raw}") + before = path.read_bytes() + with pytest.raises(citations.CitationError, match="Line .*Malformed metadata") as exc: + citations.record_citations(path, ["Key `a b` is distinct."], symbol="Auth.login") + assert str(path.resolve()) in str(exc.value) + assert path.read_bytes() == before + + +def test_cli_error_names_the_entry_and_preserves_file(tmp_path): + path = per_file(tmp_path) + before = path.read_bytes() + result = subprocess.run( + [sys.executable, str(CORE / "shadow-cite.py"), str(path), "--symbol", "Auth.login", "--text", "Wrong claim."], + capture_output=True, text=True, encoding="utf-8", + ) + assert result.returncode == 1 and "ERROR:" in result.stderr and "Re-read" in result.stderr + assert result.stdout == "" and path.read_bytes() == before + + +@pytest.mark.parametrize("relative", ["_index.md", "_dreams/dream/report.md", "_meta/notes.md"]) +def test_indexes_and_experiment_reports_are_not_citation_targets(citations, tmp_path, relative): + path = tmp_path / ".shadow" / relative + path.parent.mkdir(parents=True) + path.write_text("# Metadata\n", encoding="utf-8") + with pytest.raises(citations.CitationError, match="not indexes or dream reports"): + citations.record_citations(path, ["Metadata"], symbol="File-Level") + + +def test_outside_target_symlink_is_refused(citations, tmp_path, make_symlink): + path = per_file(tmp_path) + outside = tmp_path / "outside.md" + outside.write_bytes(path.read_bytes()) + path.unlink() + make_symlink(path, outside) + before = outside.read_bytes() + with pytest.raises(citations.CitationError, match="inside"): + citations.record_citations(path, ["Key `a b` is distinct."], symbol="Auth.login") + assert outside.read_bytes() == before + + +def test_aliased_shadow_root_is_refused(citations, tmp_path, make_symlink): + real_root = tmp_path / "real" + real_root.mkdir() + path = real_root / "file.md" + path.write_text("# Shadow: file\n\n## `run`\n\n- Claim.\n _(verified, source: exploration)_\n", encoding="utf-8") + make_symlink(tmp_path / ".shadow", real_root, target_is_directory=True) + before = path.read_bytes() + with pytest.raises(citations.CitationError, match="filesystem alias"): + citations.record_citations(tmp_path / ".shadow/file.md", ["Claim."], symbol="run") + assert path.read_bytes() == before diff --git a/tests/skills/shadow_frog_dream/test_citation_scores.py b/tests/skills/shadow_frog_dream/test_citation_scores.py new file mode 100644 index 0000000..ceee34d --- /dev/null +++ b/tests/skills/shadow_frog_dream/test_citation_scores.py @@ -0,0 +1,105 @@ +"""Dream artifacts preserve visible citation hints without summing shared ancestry.""" + +import pytest + + +def test_merge_keeps_larger_score_while_upgrading_metadata(dream_reconcile, tmp_path): + path = tmp_path / ".shadow/source.py.md" + path.parent.mkdir() + path.write_text( + "# Shadow: source.py\n\n## `run`\n\n- Claim.\n" + " _(uncertain, source: exploration, citation_score: 7)_\n\n## Cross-References\n", + encoding="utf-8", + ) + discovery = {"text": "Claim.", "status": "verified", "source": "user", "citation_score": 3} + assert dream_reconcile.merge_discovery_into_file( + str(path), "run", discovery, "dream", repo_root=str(tmp_path), + ) + assert "source: user, citation_score: 7" in path.read_text(encoding="utf-8") + discovery["citation_score"] = 9 + assert dream_reconcile.merge_discovery_into_file( + str(path), "run", discovery, "dream", repo_root=str(tmp_path), + ) + assert "citation_score: 9" in path.read_text(encoding="utf-8") + assert not dream_reconcile.merge_discovery_into_file( + str(path), "run", discovery, "dream", repo_root=str(tmp_path), + ) + + +def test_existing_cross_refs_and_scores_merge_independently(dream_reconcile, tmp_path): + path = tmp_path / ".shadow/_cross/contract.md" + path.parent.mkdir(parents=True) + path.write_text( + "# Contract\n\n**Refs**:\n- `a.py::run`\n\n**Discovery**: Shared claim.\n\n" + "_(verified, source: exploration, citation_score: 4)_\n", + encoding="utf-8", + ) + assert dream_reconcile._merge_refs_into_cross_file( + str(path), ["b.py::call"], repo_root=str(tmp_path), citation_score=2, + ) + assert "citation_score: 4" in path.read_text(encoding="utf-8") + assert dream_reconcile._merge_refs_into_cross_file( + str(path), ["a.py::run"], repo_root=str(tmp_path), citation_score=7, + ) + assert "citation_score: 7" in path.read_text(encoding="utf-8") + assert "`b.py::call`" in path.read_text(encoding="utf-8") + + +def test_cross_merge_does_not_edit_metadata_inside_examples(dream_reconcile, tmp_path): + path = tmp_path / ".shadow/_cross/contract.md" + path.parent.mkdir(parents=True) + original = ( + "# Contract\n\n**Refs**:\n- `a.py::run`\n\n" + "```markdown\n_(verified, source: exploration, citation_score: 2)_\n```\n\n" + "**Discovery**: Shared claim.\n\n" + "_(verified, source: exploration, citation_score: 4)_\n" + ) + path.write_text(original, encoding="utf-8") + assert dream_reconcile._merge_refs_into_cross_file( + str(path), ["a.py::run"], repo_root=str(tmp_path), citation_score=7, + ) + assert path.read_text(encoding="utf-8") == original.replace( + "citation_score: 4", "citation_score: 7", + ) + + +@pytest.mark.parametrize("score", [-1, 0.5, True, None, "4"]) +@pytest.mark.parametrize("kind", ["discoveries", "cross_cutting"]) +def test_invalid_manifest_score_fails_before_any_publication(dream_reconcile, tmp_path, score, kind): + invalid = {"anchor": "a.py::run", "text": "Claim.", "slug": "contract", "refs": [], "citation_score": score} + manifest = {"discoveries": [{"anchor": "valid.py::run", "text": "Valid claim."}]} + if kind == "discoveries": + manifest["discoveries"].append(invalid) + else: + manifest[kind] = [invalid] + with pytest.raises(dream_reconcile.CitationError, match="nonnegative integer"): + dream_reconcile.merge_discoveries(str(tmp_path), [("dream/p/id", "id", manifest)]) + assert list(tmp_path.iterdir()) == [] + + +@pytest.mark.slow +@pytest.mark.parametrize("score,expected", [(0, 0), (7, 0), (-1, 1), (False, 1), (0.5, 1), ("3", 1)]) +@pytest.mark.parametrize("kind", ["discoveries", "cross_cutting"]) +def test_validator_checks_manifest_citation_fields(tmp_git_repo, score, expected, kind): + from tests.skills.shadow_frog_dream.test_dream_validate import ( + _commit_base, _default_manifest, _default_report, _run_validate, _write_dream, + ) + + base = _commit_base(tmp_git_repo) + dream_id = "20260924-010000Z-citations" + entry = { + "anchor": "a.py::run", "slug": "contract", "refs": ["a.py::run"], + "text": "The function leaves its input unchanged.", "citation_score": score, + } + manifest = _default_manifest(dream_id, **{kind: [entry]}) + _write_dream(tmp_git_repo, dream_id, manifest=manifest, report=_default_report(dream_id, base)) + (tmp_git_repo / ".shadow/a.py.md").write_text( + "# Shadow: a.py\n\n## `run`\n\n- Leaves its input unchanged.\n" + " _(verified, source: exploration, citation_score: 0)_\n", + encoding="utf-8", + ) + result = _run_validate(dream_id, tmp_git_repo) + assert result.returncode == expected, result.stdout + result.stderr + if expected: + assert f"{kind}[0].citation_score" in result.stdout + assert "nonnegative integer" in result.stdout diff --git a/tests/skills/shadow_frog_dream/test_dream_reconcile.py b/tests/skills/shadow_frog_dream/test_dream_reconcile.py index 49b15d6..32ec085 100644 --- a/tests/skills/shadow_frog_dream/test_dream_reconcile.py +++ b/tests/skills/shadow_frog_dream/test_dream_reconcile.py @@ -375,7 +375,7 @@ def test_merge_discovery_creates_new_shadow(dream_reconcile, tmp_path): body = shadow.read_text(encoding="utf-8") assert "## `do_thing`" in body assert "- Returns None on empty input." in body - assert "_(verified, source: exploration)_" in body + assert "_(verified, source: exploration, citation_score: 0)_" in body assert "Dream report: `_dreams/20260101-000000Z-x/`" in body assert "## Cross-References" in body @@ -502,7 +502,7 @@ def test_merge_discovery_upgrades_uncertain_to_verified(dream_reconcile, tmp_pat ) assert written is True body = shadow.read_text(encoding="utf-8") - assert "_(verified, source: exploration)_" in body + assert "_(verified, source: exploration, citation_score: 0)_" in body def test_merge_discovery_never_downgrades_verified(dream_reconcile, tmp_path): @@ -1426,7 +1426,7 @@ def test_merge_discoveries_creates_cross_cutting_file_with_back_pointers( assert "**Category**: behavior" in body assert "`src/auth.py::login`" in body assert "`lib/session.py::Session`" in body - assert "_(verified, source: exploration)_" in body + assert "_(verified, source: exploration, citation_score: 0)_" in body # Each referenced per-file shadow must have a back-pointer. auth = (repo / ".shadow" / "src" / "auth.py.md").read_text(encoding="utf-8") diff --git a/tests/skills/shadow_frog_dream/test_dream_tools.py b/tests/skills/shadow_frog_dream/test_dream_tools.py index d071567..8c7e855 100644 --- a/tests/skills/shadow_frog_dream/test_dream_tools.py +++ b/tests/skills/shadow_frog_dream/test_dream_tools.py @@ -49,6 +49,7 @@ def test_pin_creates_complete_external_snapshot(tmp_git_repo, tmp_path): assert Path(packet["manifest"]).is_absolute() assert Path(packet["skill_dir"]) == output / "shadow-frog-dream" assert (output / "shadow-frog/_coherence.py").is_file() + assert (output / "shadow-frog/_citations.py").is_file() assert all(Path(path).is_file() for path in packet["instructions"]) metadata = json.loads(Path(packet["manifest"]).read_text(encoding="utf-8")) assert metadata["mode"] == "coherent" diff --git a/tests/skills/shadow_frog_viewer/test_citation_metadata.py b/tests/skills/shadow_frog_viewer/test_citation_metadata.py new file mode 100644 index 0000000..a332216 --- /dev/null +++ b/tests/skills/shadow_frog_viewer/test_citation_metadata.py @@ -0,0 +1,119 @@ +"""Viewer reads visible scores without recording implicit views or creating state.""" + +import os +import shutil +import subprocess +import sys + +import pytest + + +def test_metadata_score_is_not_discovery_text(shadow_viewer): + discovery = shadow_viewer.parse_discovery( + "- Rejects empty input.", + [" _(verified, source: user, labels: [bug], citation_score: 7)_"], + ) + assert discovery["text"] == "Rejects empty input." + assert discovery["citation_score"] == 7 + assert discovery["status"] == "verified" and discovery["labels"] == ["bug"] + assert shadow_viewer.parse_discovery("- Older claim.", [" _(verified, source: exploration)_"])["citation_score"] == 0 + assert shadow_viewer.parse_discovery("- Preference.", [" _(source: user, citation_score: 3)_"])["citation_score"] == 3 + + +def test_reading_and_ranking_scores_never_rewrites_shadow(shadow_viewer, tmp_path, capsys): + shadow = tmp_path / ".shadow" + shadow.mkdir() + path = shadow / "source.py.md" + path.write_text( + "# Shadow: source.py\n\n## `run`\n\n" + "- Rare finding.\n _(verified, source: exploration, labels: [bug], citation_score: 1)_\n\n" + "- Popular finding.\n _(verified, source: exploration, labels: [bug], citation_score: 9)_\n\n" + "- User constraint.\n _(verified, source: user, labels: [bug], citation_score: 0)_\n\n" + "## Cross-References\n", + encoding="utf-8", + ) + before = path.read_bytes() + shadow_viewer.view_top(shadow, "source.py", "bug", 3, 0) + output = capsys.readouterr().out + assert output.index("User constraint") < output.index("Popular finding") < output.index("Rare finding") + assert "citation_score: 9" in output + shadow_viewer.view_search(shadow, "finding") + output = capsys.readouterr().out + assert output.index("Popular finding") < output.index("Rare finding") + assert path.read_bytes() == before + assert sorted(p.name for p in shadow.iterdir()) == ["source.py.md"] + + +@pytest.mark.parametrize("location", ["source.py.md", "_prefs.md", "_cross/contract.md"]) +@pytest.mark.parametrize("score", ["-1", "1.5", "true"]) +def test_structural_audit_reports_invalid_visible_score(shadow_viewer, tmp_path, capsys, location, score): + shadow = tmp_path / ".shadow" + path = shadow / location + path.parent.mkdir(parents=True) + content = "# Shadow: source.py\n\n## `run`\n\n- Claim.\n" + metadata = f" _(verified, source: exploration, citation_score: {score})_\n" + if location == "_prefs.md": + content = "# Preferences\n\n- Claim.\n" + metadata = f" _(source: user, citation_score: {score})_\n" + elif location.startswith("_cross"): + content = "# Contract\n\n**Category**: contract\n**Refs**:\n- `source.py::run`\n\n**Discovery**: Claim.\n\n" + path.write_text(content + metadata, encoding="utf-8") + assert shadow_viewer.view_check_invariants(shadow) == 1 + assert "citation_score" in capsys.readouterr().out + + +@pytest.mark.slow +def test_direct_read_then_cite_then_user_viewer(repo_root, tmp_path): + shadow = tmp_path / ".shadow" + shadow.mkdir() + path = shadow / "source.py.md" + path.write_text( + "# Shadow: source.py\n\n## `run`\n\n- Rejects empty input.\n" + " _(verified, source: exploration, citation_score: 0)_\n", + encoding="utf-8", + ) + assert "citation_score: 0" in path.read_text(encoding="utf-8") + result = subprocess.run( + [sys.executable, str(repo_root / "skills/shadow-frog/shadow-cite.py"), str(path), + "--symbol", "run", "--text", "Rejects empty input."], + capture_output=True, text=True, encoding="utf-8", check=True, + ) + assert "0 -> 1" in result.stdout + before_view = path.read_bytes() + result = subprocess.run( + [sys.executable, str(repo_root / "skills/shadow-frog-viewer/shadow-viewer.py"), + "--shadow-dir", str(shadow), "--search", "empty"], + capture_output=True, text=True, encoding="utf-8", check=True, + ) + assert "citation_score: 1" in result.stdout and path.read_bytes() == before_view + + +@pytest.mark.parametrize("layout", [".github", ".claude"]) +def test_installed_counter_helper_has_no_hidden_store_or_bytecode(repo_root, tmp_path, layout): + installed = tmp_path / layout / "skills" + for name in ("shadow-frog", "shadow-frog-viewer"): + shutil.copytree( + repo_root / "skills" / name, installed / name, + ignore=shutil.ignore_patterns("__pycache__", "*.pyc"), + ) + shadow = tmp_path / ".shadow" + shadow.mkdir() + target = shadow / "sample.py.md" + target.write_text( + "# Shadow: sample.py\n\n## `run`\n\n- Keep the input unchanged.\n" + " _(verified, source: user, citation_score: 0)_\n", + encoding="utf-8", + ) + env = os.environ.copy() + env.pop("PYTHONDONTWRITEBYTECODE", None) + env.pop("PYTHONPYCACHEPREFIX", None) + result = subprocess.run( + [sys.executable, str(installed / "shadow-frog/shadow-cite.py"), str(target), + "--symbol", "run", "--text", "Keep the input unchanged."], + env=env, capture_output=True, text=True, encoding="utf-8", + ) + assert result.returncode == 0, result.stderr + assert "citation_score: 1" in target.read_text(encoding="utf-8") + assert not list(tmp_path.rglob("*.sqlite3")) + assert not list(installed.rglob("*.pyc")) + assert not list(shadow.rglob("*.citation.lock")) diff --git a/tests/test_smoke.py b/tests/test_smoke.py index 3d57fc4..b88b3be 100644 --- a/tests/test_smoke.py +++ b/tests/test_smoke.py @@ -13,7 +13,7 @@ def test_repo_root_resolves(repo_root): def test_all_script_fixtures_load( shadow_init, shadow_viewer, dream_reconcile, dream_validate, - dream_coverage, dream_lineage, meditate_repair, nap, coherence, dream_tools, + dream_coverage, dream_lineage, meditate_repair, nap, coherence, dream_tools, citations, ): for mod, expected_attr in [ (shadow_init, "main"), @@ -26,6 +26,7 @@ def test_all_script_fixtures_load( (meditate_repair, "main"), (nap, "main"), (coherence, "validate_connection"), + (citations, "record_citations"), ]: assert hasattr(mod, expected_attr), \ f"{mod.__name__} missing expected attribute {expected_attr!r}"