Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,21 @@ permissions:
contents: read

jobs:
citation-compatibility:
name: Citation helpers (Python 3.9)
runs-on: ubuntu-latest
steps:
- name: Checkout
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- name: Set up Python
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: '3.9'
- name: Check citation and viewer entry points
run: |
python skills/shadow-frog/shadow-cite.py --help
python skills/shadow-frog-viewer/shadow-viewer.py --shadow-dir examples/coupon-demo/.shadow --check-invariants

pytest:
name: pytest (${{ matrix.os }}, Python 3.12)
runs-on: ${{ matrix.os }}
Expand Down
9 changes: 8 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,14 @@ shadow knowledge bases for any codebase.

---

## Unreleased
## 2026-09-28

### Added
- **Visible knowledge citations** — `citation_score` lives alongside discovery
metadata in Markdown. Agents record consulted entries once per task after
normal file/symbol reads. A small locked increment helper preserves other
content; Viewer and reconciliation understand the same score field without
a database or additional retrieval workflow.

### Fixed
- **Reconciliation path containment** — validate untrusted manifest destinations
Expand Down
14 changes: 11 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,13 @@ For example, use Viewer to find relevant knowledge or audit its structure:
The [Viewer reference](skills/shadow-frog-viewer/SKILL.md) also covers summaries,
recent discoveries, label filters, preferences, and interactive dream-lineage HTML.

Agents still read shadow files and symbols directly. Each entry carries a visible
`citation_score`, initially 0, as an approximate hint of how often agents revisit
it. After consulting an entry, the agent records it once per task with the small
core `shadow-cite.py` helper; its only job is to update the selected Markdown
counter safely. No database or special retrieval interface is required. Scores
never replace relevance, trust, or verification.

---

## Choose Dream or Nap
Expand Down Expand Up @@ -218,28 +225,29 @@ For example:
```markdown
- authenticate_user() silently returns None on expired tokens
instead of raising. 3 of 7 callers don't check the return value.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
```

**User knowledge**:
```markdown
- The retry logic here took 3 iterations to get right -- it handles
a subtle race condition during rolling deployments. Do not simplify.
_(verified, source: user)_
_(verified, source: user, citation_score: 0)_
```

**Collaborative work**:
```markdown
- While debugging issue #42, discovered that process_batch() silently
drops items exceeding 1MB -- logged at DEBUG level only.
_(verified, source: interaction)_
_(verified, source: interaction, citation_score: 0)_
```

| Property | Values | Meaning |
|----------|--------|---------|
| **Status** | `verified` / `uncertain` / `refuted` | Has the claim been confirmed? |
| **Source** | `exploration` / `user` / `interaction` | Where did this knowledge come from? |
| **Labels** | `bug`, `performance`, `security`, `feature-gap`, `tech-debt` | Optional; marks actionable discoveries |
| **Citation score** | Nonnegative integer; missing means `0` | Approximate agent-reported revisits, stored as `citation_score` in the entry |

### Trust Hierarchy

Expand Down
10 changes: 10 additions & 0 deletions agent-context.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,16 @@ This project uses a `.shadow/` knowledge base with verified discoveries about no

The shadow contains discoveries from code analysis and user conversations. Always consult it before making assumptions about code behavior.

Each entry's Markdown metadata includes `citation_score` (initially 0; omitted
also means 0). After deliberately consulting an entry, increment it once per
task using `shadow-cite.py` in the core `shadow-frog` skill, supplying the file,
symbol and exact claim text already read. Batch repeated `--text` arguments
for one file/section. Coordinate subagent updates through one writer.
Do not count every entry in an opened file or recount repeated reads.
The score is approximate revisit frequency, not confidence; relevance and trust
take precedence. `/shadow-frog-viewer` is for user-facing inspection, not a
required read path. There is no citation database or retrieval protocol.

### Key directories

- `.shadow/<path>.md` — per-file shadows with symbol-level discoveries
Expand Down
21 changes: 18 additions & 3 deletions claude.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@ ShadowFrog/
skills/
shadow-frog/SKILL.md Main entrypoint (docs, reference system, search)
shadow-frog/_coherence.py Shared structural parent-connection validation
shadow-frog/_citations.py Visible score metadata and safe per-file increments
shadow-frog/shadow-cite.py Record exact consulted claims after native reads
shadow-frog-init/ First-time setup (create .shadow/)
SKILL.md Init instructions + fallback steps
shadow-init.py Python helper script
Expand Down Expand Up @@ -84,7 +86,7 @@ Canonical formal spec: `/shadow-frog`. The shapes below are the minimum an agent
Per-file discovery (anchored by `file::symbol` heading; labels and `Also involves:` are optional):
```
- <behavioral statement>
_(<verified|uncertain|refuted>, source: <exploration|user|interaction>[, labels: [bug, security]])_
_(<verified|uncertain|refuted>, source: <exploration|user|interaction>[, labels: [bug, security]], citation_score: 0)_
Also involves: `file::symbol`, `file::symbol`
```

Expand All @@ -98,19 +100,32 @@ Cross-cutting (`_cross/<slug>.md`, slug = kebab-case from title, e.g. "DB connec

**Discovery**: <behavioral statement>

_(<verified|uncertain|refuted>, source: <exploration|user|interaction>)_
_(<verified|uncertain|refuted>, source: <exploration|user|interaction>, citation_score: 0)_
```

Preference (`_prefs.md` — project-wide, no file/symbol anchor):
```
- <preference or convention>
_(source: <user|interaction>)_
_(source: <user|interaction>, citation_score: 0)_
```

- Labels (lowercase, comma-separated): `bug`, `performance`, `security`, `feature-gap`, `tech-debt`. Only for actionable discoveries.
- `Also involves:` always uses `file::symbol`, never bare file paths.
- `Dream report: _dreams/<dream-id>/` is optional — only for experiment-derived discoveries.

### Citation Scores

- Keep one visible nonnegative integer `citation_score` in Markdown metadata,
after optional labels. New entries start at 0; omitted scores also mean 0.
- Agents read files/symbols directly and explicitly cite consulted entries once
per task. Use core `shadow-cite.py` for serialized exact-claim increments; no
database, opaque IDs, or required retrieval service.
- Score updates must not change discovery prose, provenance, references, or counts.
Coordinate ordinary edits with citation writes; only helper calls share its lock.
- Keep scores when rewording/moving the same claim, and use max rather than sum
when combining duplicates or inherited branch state. Scores are approximate,
not a global audited count or a trust/confidence value.

### Verification
- Observe-based: read source at `file::symbol`, trace logic, confirm claim.
- Do-based: write and run a short test/script to confirm or refute.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -8,4 +8,4 @@

**Discovery**: validate_coupon normalizes coupon codes to uppercase via code.upper() before lookup, but calculate_total passes coupon_code directly to get_coupon without normalization. A user who validates "save20" (returns True) and then passes "save20" to calculate_total gets no discount — the coupon silently fails because load_coupon's keys are uppercase. This creates a validate-then-use inconsistency where validated codes don't work.

_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
Original file line number Diff line number Diff line change
Expand Up @@ -9,4 +9,4 @@

**Discovery**: COUPON_CACHE is a module-level global dict shared by all importers. Any call to get_coupon (directly or via validate_coupon) permanently populates the cache, including caching None for invalid codes. In tests, cache entries from one test persist into the next — there is no reset mechanism. validate_coupon caches under the uppercased key, while calculate_total would cache under the original-case key, so a single logical coupon code can produce two separate cache entries ("SAVE20" and "save20") with different values.

_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
Original file line number Diff line number Diff line change
Expand Up @@ -8,4 +8,4 @@

**Discovery**: apply_bulk_discount mutates item dicts in-place (modifying "price" keys) and returns the same list object. When piped into calculate_total, the mutation is invisible — calculate_total sees already-reduced prices. But any code holding a reference to the original items list now sees the discounted prices permanently. Calling apply_bulk_discount multiple times compounds discounts (0.90^N multiplier). The test_bulk_then_coupon test avoids this by creating fresh items, but real usage with shared item references would silently corrupt prices.

_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
28 changes: 14 additions & 14 deletions examples/coupon-demo/.shadow/cart.py.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,53 +5,53 @@
## File-Level

- COUPON_CACHE is a module-level mutable global dict shared across all importers — any module that imports from cart.py shares the same cache instance, causing cross-module state pollution.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
Also involves: `inventory.py::validate_coupon`

## `COUPON_CACHE`

- Never cleared or evicted — grows monotonically for the lifetime of the process. In a long-running server, every unique coupon code ever queried remains cached forever.
_(verified, source: exploration, labels: [performance])_
_(verified, source: exploration, labels: [performance], citation_score: 0)_
- Caches None for invalid codes. Once an invalid code is looked up, the None result is permanently cached, preventing any future lookup even if the underlying data changes.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
- Case-variant lookups create duplicate cache entries for the same logical coupon. validate_coupon("save20") caches "SAVE20" → valid, then calculate_total("save20") caches "save20" → None. Cache grows 2× faster with mixed-case usage.
_(verified, source: exploration, labels: [bug, performance])_
_(verified, source: exploration, labels: [bug, performance], citation_score: 0)_
Dream report: `_dreams/20260420-140000Z-cache-poison-sequence/`

## `load_coupon`

- Case-sensitive lookup against uppercase keys ("SAVE20", "HALF"). Passing lowercase (e.g., "save20") returns None even though the coupon conceptually exists.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
- Returns None for unknown codes (via dict.get default), not an exception. Callers must handle None.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_

## `get_coupon`

- The `if code not in COUPON_CACHE` guard means each code is loaded exactly once per process. But because None is a valid cached value, invalid codes are also "loaded once" and permanently considered invalid.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
Also involves: `cart.py::COUPON_CACHE`, `cart.py::load_coupon`
- Accepts any hashable type as key (True, 42, lists-as-errors). Non-string keys permanently cache None entries that can never resolve to valid coupons — silent cache pollution.
_(verified, source: exploration, labels: [security])_
_(verified, source: exploration, labels: [security], citation_score: 0)_
Dream report: `_dreams/20260420-142000Z-adversarial-inputs/`
Also involves: `cart.py::COUPON_CACHE`

## `calculate_total`

- Does NOT normalize coupon_code to uppercase before lookup. Lowercase codes silently produce no discount (coupon returns None from cache or load_coupon). This contradicts validate_coupon which does normalize.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
Also involves: `inventory.py::validate_coupon`
- `if coupon_code:` is falsy for empty string "", None, 0, and False — all skip coupon lookup silently. No distinction between "no coupon" and "invalid coupon".
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
- Accepts negative prices and quantities without validation. Negative subtotals still have 8% tax applied, producing negative totals (e.g., price=-10, qty=1 → total=-10.80).
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
- Tax rate (0.08 = 8%) is hardcoded with no configuration mechanism. Changing tax requires editing source code.
_(verified, source: exploration, labels: [tech-debt])_
_(verified, source: exploration, labels: [tech-debt], citation_score: 0)_
- Coupon min_total check uses the actual subtotal computed from items at call time. Since apply_bulk_discount mutates prices in-place before calculate_total runs, the min_total check sees post-bulk-discount prices. If bulk discount drops subtotal below min_total, the coupon is correctly rejected.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
Dream report: `_dreams/20260420-141000Z-bulk-min-total-interaction/`
Also involves: `inventory.py::apply_bulk_discount`
- Non-string coupon_code values (True, 42, etc.) pass the `if coupon_code:` truthiness check, reach get_coupon, cache None under the non-string key, and silently produce no discount. No type validation exists.
_(verified, source: exploration, labels: [security])_
_(verified, source: exploration, labels: [security], citation_score: 0)_
Dream report: `_dreams/20260420-142000Z-adversarial-inputs/`
Also involves: `cart.py::get_coupon`, `cart.py::COUPON_CACHE`

Expand Down
20 changes: 10 additions & 10 deletions examples/coupon-demo/.shadow/inventory.py.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,36 +5,36 @@
## File-Level

- Imports get_coupon from cart, creating a dependency on cart's COUPON_CACHE. Any call to validate_coupon pollutes the shared cache as a side effect.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
Also involves: `cart.py::get_coupon`, `cart.py::COUPON_CACHE`

## `validate_coupon`

- Normalizes code to uppercase via `code.upper()` before lookup, but calculate_total does NOT — so a code that validates successfully may still produce no discount when passed directly to calculate_total.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
Also involves: `cart.py::calculate_total`, `cart.py::load_coupon`
- Side effect: populates COUPON_CACHE with the uppercased code. Calling validate_coupon("save20") caches under key "SAVE20", but a later calculate_total("save20") looks up lowercase "save20" — a cache miss that then caches None under "save20".
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
Also involves: `cart.py::COUPON_CACHE`, `cart.py::get_coupon`
- Will crash with AttributeError if code is None (None has no .upper() method). No guard against non-string input.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
- Also crashes on any non-string type: int, list, dict all raise AttributeError on .upper(). Needs `isinstance(code, str)` guard.
_(verified, source: exploration, labels: [security])_
_(verified, source: exploration, labels: [security], citation_score: 0)_
Dream report: `_dreams/20260420-142000Z-adversarial-inputs/`

## `apply_bulk_discount`

- Mutates items in-place — modifies the original dict objects' "price" keys. The caller's list is permanently altered. Returns the same list object (not a copy).
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
- Crashes with KeyError if items lack 'qty' key, and TypeError if 'qty' is a string. No input validation.
_(verified, source: exploration, labels: [security])_
_(verified, source: exploration, labels: [security], citation_score: 0)_
Dream report: `_dreams/20260420-142000Z-adversarial-inputs/`
- Calling apply_bulk_discount twice on the same items compounds the discount: first call gives 0.90×, second gives 0.81×, third gives 0.729×. No idempotency guard.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
- Only applies discount to items with qty >= 5. Items with qty 4 or below are untouched, even if the total quantity across all items exceeds 5.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
- Uses round(price * 0.90, 2) which can produce floating-point artifacts on certain prices. For example, 33.33 * 0.90 = 29.997 → rounds to 30.0, not 29.997.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_

## Cross-References

Expand Down
18 changes: 9 additions & 9 deletions examples/coupon-demo/.shadow/test_cart.py.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,35 +5,35 @@
## File-Level

- Uses a manual `if __name__ == "__main__"` test runner, not pytest or unittest. Tests are plain functions with assert statements and no setup/teardown.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
- No test isolation: COUPON_CACHE is a shared global. test_coupon populates the cache with "SAVE20", and test_bulk_then_coupon populates it with "HALF". If test order changes or tests are rerun in the same process, cached values persist from earlier tests.
_(verified, source: exploration, labels: [bug])_
_(verified, source: exploration, labels: [bug], citation_score: 0)_
Also involves: `cart.py::COUPON_CACHE`
- No negative-path tests: no test for invalid coupon codes, negative prices, empty items, or the case-sensitivity mismatch between validate_coupon and calculate_total.
_(verified, source: exploration, labels: [feature-gap])_
_(verified, source: exploration, labels: [feature-gap], citation_score: 0)_

## `test_basic_total`

- Verifies: 25.00 × 3 = 75.00 subtotal + 8% tax = 81.00. Tests the simplest happy path with no coupon.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_

## `test_coupon`

- Verifies SAVE20 on 75.00 subtotal: 75.00 − 20% = 60.00 + 4.80 tax = 64.80. The 75.00 subtotal exceeds SAVE20's min_total of 50.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
- Only tests with uppercase coupon code "SAVE20". Does not test lowercase, revealing nothing about the case-normalization bug.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
Also involves: `cart.py::calculate_total`

## `test_bulk_then_coupon`

- Tests the full pipeline: apply_bulk_discount mutates items (25.00 → 22.50), then calculate_total applies HALF coupon on 112.50 subtotal → 56.25 + 4.50 tax = 60.75.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
- Creates fresh items list, avoiding the mutation-persistence issue. But if this test's items object were reused in a subsequent test, prices would already be 22.50, not 25.00.
_(verified, source: exploration)_
_(verified, source: exploration, citation_score: 0)_
Also involves: `inventory.py::apply_bulk_discount`, `cart.py::calculate_total`
- Imports apply_bulk_discount inside the function body (lazy import), unlike the module-level import of calculate_total. This is inconsistent but functionally irrelevant.
_(verified, source: exploration, labels: [tech-debt])_
_(verified, source: exploration, labels: [tech-debt], citation_score: 0)_

## Cross-References

Expand Down
Loading
Loading