Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions examples/01_quickstart.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
"""memorywire quickstart — end-to-end demo using the :class:`memorywire.Memory` facade.
"""memorywire quickstart — end-to-end demo using the :class:`memorywire.Memory` facade.

This script ingests 50 short facts, runs a few recalls, deletes one user's
records by filter, and prints final aggregate stats. It exists primarily
Expand All @@ -9,8 +9,8 @@
--------------
The default :class:`memorywire.store.sqlite_vec.SqliteVecStore` lazy-loads
``sentence-transformers/all-MiniLM-L6-v2`` on first embed call. To keep
this example runnable anywhere — CI, a fresh laptop, a Docker image
without ML wheels — we inject a tiny deterministic fake embedder
this example runnable anywhere — CI, a fresh laptop, a Docker image
without ML wheels — we inject a tiny deterministic fake embedder
(sha256-derived 384-d vectors). The embedder is *not* representative of
real recall quality; it exists to make the storage layer exercise its
ANN path without pulling sentence-transformers.
Expand All @@ -32,7 +32,7 @@
from memorywire.store.sqlite_vec import SqliteVecStore

# ---------------------------------------------------------------------------
# Fake embedder — sha256-derived 384-d deterministic vector
# Fake embedder — sha256-derived 384-d deterministic vector
# ---------------------------------------------------------------------------


Expand All @@ -51,7 +51,7 @@ def fake_embedder(text: str) -> list[float]:


# ---------------------------------------------------------------------------
# Seed data — 50 small facts
# Seed data — 50 small facts
# ---------------------------------------------------------------------------

# Each entry is (content, user_id) so we can demonstrate per-user filtering.
Expand Down Expand Up @@ -129,7 +129,7 @@ def _section(title: str) -> None:
async def main() -> None:
"""Run the full quickstart end-to-end."""
# Use a temp file path so the demo also exercises the on-disk path.
# An in-memory db works just as well — toggle the line below if needed.
# An in-memory db works just as well — toggle the line below if needed.
tmp_dir = tempfile.mkdtemp(prefix="amp-quickstart-")
db_path = os.path.join(tmp_dir, "quickstart.db")

Expand Down
6 changes: 3 additions & 3 deletions examples/03_procedural_fsm.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,10 @@

Runnable demo of the Phase-5 :mod:`memorywire.procedural` backend:

1. Build the canonical ``book-flight`` procedure from spec §7.
1. Build the canonical ``book-flight`` procedure from spec §7.
2. Statically validate the procedure.
3. Drive it through the happy path:
``found_options → picked → paid → receipt``.
``found_options → picked → paid → receipt``.
4. Demonstrate the ``"source": "*"`` wildcard idiom by ``cancel`` from a
mid-flow state.
5. Roundtrip through ``to_dict()`` / JSON / ``from_dict()`` and assert
Expand All @@ -24,7 +24,7 @@


def build_book_flight() -> Procedure:
"""Construct the spec §7 ``book-flight`` procedure."""
"""Construct the spec §7 ``book-flight`` procedure."""
return Procedure(
name="book-flight",
states=[
Expand Down
7 changes: 1 addition & 6 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -117,12 +117,7 @@ select = [
"SIM", # flake8-simplify
"RUF", # ruff-specific
]
ignore = [
"E501", # line length enforced by formatter
# Intentional typography in prose docstrings/comments/strings
# (em dashes, curly quotes, arrows) — deliberate, not defects.
"RUF001", "RUF002", "RUF003",
]
ignore = ["E501"] # line length enforced by formatter

[tool.ruff.format]
quote-style = "double"
Expand Down
2 changes: 1 addition & 1 deletion scripts/bump_version.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
manifest version.

Hatch-vcs derives the actual installed version from the latest ``v*`` git tag,
so the in-file values are *advisory* — they exist so editors and humans can see
so the in-file values are *advisory* — they exist so editors and humans can see
the intended version without running ``git describe``. release-please rewrites
all three on merge.

Expand Down
4 changes: 2 additions & 2 deletions scripts/extract_abstract.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@
clean = re.sub(r"\\ref\{[^}]+\}", "", clean)
clean = clean.replace(r"\&", "&").replace(r"\%", "%").replace(r"\#", "#").replace(r"\$", "$")
clean = re.sub(r"\\\\", " ", clean)
clean = clean.replace("---", "—").replace("--", "–")
clean = clean.replace("---", "—").replace("--", "–") # noqa: RUF001 -- deliberately emits em/en dashes
# Strip math-mode delimiters; arXiv's text field renders math as text.
clean = re.sub(r"\$([^$]+)\$", r"\1", clean)
# Replace common math symbols with plain-text equivalents.
Expand All @@ -44,7 +44,7 @@
clean = re.sub(r"\s+", " ", clean).strip()

print("=" * 72)
print(" PASTE THIS INTO arXiv's 'Abstract' field — plain text ready")
print(" PASTE THIS INTO arXiv's 'Abstract' field — plain text ready")
print("=" * 72)
print()
print(clean)
Expand Down
2 changes: 1 addition & 1 deletion scripts/inspect_pdf.py
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@
for label, needle in checks:
found = needle in all_text
mark = " OK " if found else " MISS"
# Last check is inverted — we want it NOT found
# Last check is inverted — we want it NOT found
if label.startswith("Old name placeholder absent"):
mark = " OK " if not found else " FAIL"
print(f" [{mark.strip()}] {label:38s} ({'present' if found else 'absent'})")
Expand Down
40 changes: 20 additions & 20 deletions scripts/lib/eval_common.py
Original file line number Diff line number Diff line change
Expand Up @@ -54,8 +54,8 @@ class EvalConfig:

Both ``run_longmemeval.py`` and ``run_locomo.py`` instantiate this
from argparse and pass it down through the eval loop. Keep it small
and JSON-serializable — the per-run JSON output embeds it for
reproducibility audit (paper §5 calls this out).
and JSON-serializable — the per-run JSON output embeds it for
reproducibility audit (paper §5 calls this out).
"""

stores: list[str] = field(default_factory=lambda: ["sqlite-vec://./eval.db"])
Expand All @@ -73,7 +73,7 @@ class EvalConfig:
out_csv_dir: Path | None = None

def to_jsonable(self) -> dict[str, Any]:
"""Return a JSON-serialisable dict (Path → str)."""
"""Return a JSON-serialisable dict (Path → str)."""
return {
"stores": list(self.stores),
"seeds": self.seeds,
Expand Down Expand Up @@ -120,7 +120,7 @@ def paired_bootstrap_ci(
seed:
RNG seed so the CI is reproducible across reruns.
alpha:
Two-sided significance level; defaults to 0.05 → 95% CI.
Two-sided significance level; defaults to 0.05 → 95% CI.

Returns
-------
Expand All @@ -130,13 +130,13 @@ def paired_bootstrap_ci(

Notes
-----
Pure-stdlib implementation — no NumPy dependency, because the eval
Pure-stdlib implementation — no NumPy dependency, because the eval
harness must run on a fresh ``pip install memorywire``
without numpy/scipy.

References
----------
Efron & Tibshirani (1993), "An Introduction to the Bootstrap" §16
Efron & Tibshirani (1993), "An Introduction to the Bootstrap" §16
(paired bootstrap).
"""
if len(a) != len(b):
Expand Down Expand Up @@ -269,8 +269,8 @@ class LLMGrader:

Cache invalidation
------------------
Swap models → key changes → cache miss → fresh grade. Swap prompt
template → key changes → cache miss → fresh grade. This is
Swap models → key changes → cache miss → fresh grade. Swap prompt
template → key changes → cache miss → fresh grade. This is
intentional: the paper's reproducibility claim hinges on the cache
capturing the *exact* grader prompt + model used to produce the
numbers, and on a different setup producing different numbers
Expand Down Expand Up @@ -358,7 +358,7 @@ def _ensure_client(self) -> Any:
def _call_openai(self, prompt: str) -> str:
client = self._ensure_client()
# Backoff is exponential with jitter; we don't retry on 4xx
# except 429 (rate limit). 429 and 5xx → retry.
# except 429 (rate limit). 429 and 5xx → retry.
last_exc: Exception | None = None
for attempt in range(self._max_retries):
try:
Expand All @@ -381,7 +381,7 @@ def _call_openai(self, prompt: str) -> str:
# Exponential backoff with jitter.
sleep_s = (2**attempt) + random.Random(attempt).random()
time.sleep(min(sleep_s, 30.0))
# Defensive — we only get here if max_retries == 0.
# Defensive — we only get here if max_retries == 0.
raise RuntimeError(f"grader call failed after {self._max_retries} retries: {last_exc}")

def grade_with_meta(
Expand Down Expand Up @@ -432,13 +432,13 @@ def _parse_grader_response(raw: str) -> tuple[float, str]:

Tolerant on purpose: trims markdown fences, swallows trailing text,
accepts ``score`` as either int or float, clamps to [0, 1]. Returns
``(0.0, raw)`` if the response cannot be parsed at all — better to
``(0.0, raw)`` if the response cannot be parsed at all — better to
score a malformed grader reply as zero than to crash a 1000-question
eval halfway through.
"""
stripped = raw.strip()
# Strip markdown fences if the grader wrapped JSON in them despite
# response_format=json_object — defensive, not expected.
# response_format=json_object — defensive, not expected.
if stripped.startswith("```"):
lines = stripped.splitlines()
stripped = "\n".join(line for line in lines if not line.startswith("```"))
Expand Down Expand Up @@ -483,7 +483,7 @@ def __init__(self, path: Path) -> None:
self._path = path
# writeback=False keeps memory bounded; we use the cache as a
# straight key-value store. The shelf is held for the grader's
# lifetime and closed via :meth:`close` — a ``with`` block here
# lifetime and closed via :meth:`close` — a ``with`` block here
# would close it before any grader call could use it.
self._shelf = shelve.open(str(path), writeback=False) # noqa: SIM115

Expand Down Expand Up @@ -634,7 +634,7 @@ def stage_dataset(

# Rough per-1k-token rates as of 2026-Q2. Conservative defaults; users
# can pass their own --grader-model and we'll fall back to a generic
# rate. We don't track *every* model — only the ones we recommend.
# rate. We don't track *every* model — only the ones we recommend.
GRADER_RATES_USD_PER_1K_TOKENS: dict[str, tuple[float, float]] = {
# model -> (input_per_1k, output_per_1k)
"gpt-4-turbo": (0.01, 0.03),
Expand All @@ -654,7 +654,7 @@ def estimate_grader_cost(
"""Cost estimate in USD for ``n_calls`` grader invocations.

If the model isn't in :data:`GRADER_RATES_USD_PER_1K_TOKENS` we use
the ``gpt-4-turbo`` rate as a pessimistic default — better the user
the ``gpt-4-turbo`` rate as a pessimistic default — better the user
overestimates and isn't surprised than the other way around.
"""
rate_in, rate_out = GRADER_RATES_USD_PER_1K_TOKENS.get(
Expand Down Expand Up @@ -689,7 +689,7 @@ def per_question_store_urls(
This helper sidesteps the issue by giving each (question, seed)
combination its own fresh SQLite file under ``workspace``. The path
encodes ``key`` so concurrent harness runs don't clobber each
other. Non-``sqlite-vec`` URLs pass through unchanged — Mem0,
other. Non-``sqlite-vec`` URLs pass through unchanged — Mem0,
Letta, etc. manage their own per-tenant isolation.

Parameters
Expand All @@ -707,8 +707,8 @@ def per_question_store_urls(
-------
``(rewritten_urls, owned_paths)``:

* ``rewritten_urls`` — the URL list to feed to ``Memory(stores=...)``.
* ``owned_paths`` — the files the helper created and the caller
* ``rewritten_urls`` — the URL list to feed to ``Memory(stores=...)``.
* ``owned_paths`` — the files the helper created and the caller
should delete after ``mem.close()``. Empty for non-sqlite-vec
URLs.
"""
Expand Down Expand Up @@ -739,7 +739,7 @@ def cleanup_question_dbs(paths: Iterable[Path]) -> None:

SQLite WAL/SHM journals are removed alongside the main DB so the
workspace stays bounded across a 1000-question run. Errors are
swallowed — losing a stale DB file is never worth aborting the
swallowed — losing a stale DB file is never worth aborting the
harness for.
"""
for db_path in paths:
Expand All @@ -753,7 +753,7 @@ def cleanup_question_dbs(paths: Iterable[Path]) -> None:
def build_grader_context(hits: Iterable[Any], *, max_chars: int = 4000) -> str:
"""Format a list of :class:`RecallHit` rows into a grader-facing context.

The grader doesn't see the raw recall output — it sees a flat string
The grader doesn't see the raw recall output — it sees a flat string
of the top-k passages, one per line, prefixed with ``[i]``. We cap
total length at ``max_chars`` so a runaway corpus can't blow the
grader's context window. Truncation happens at the hit boundary
Expand Down
2 changes: 1 addition & 1 deletion scripts/preflight_arxiv.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
"""arXiv pre-upload preflight — runs against the checklist arXiv displays
"""arXiv pre-upload preflight — runs against the checklist arXiv displays
just before file upload. Catches the issues that slow down announcement:

1. TeX source is present (not PDF-only).
Expand Down
Loading
Loading