Skip to content

Commit b32681b

Browse files
Ronald Tseronaldtse
authored andcommitted
contract identity: README describes the system + dependency directions; SPEC.md normative v1; runtime sample imports secryst
1 parent 8d1380e commit b32681b

3 files changed

Lines changed: 163 additions & 111 deletions

File tree

‎README.md‎

Lines changed: 56 additions & 110 deletions
Original file line numberDiff line numberDiff line change
@@ -1,125 +1,71 @@
1-
# interscript/ml-models
1+
# interscript-ml
22

3-
Unified training framework for Interscript ML-powered maps.
3+
The **contract** for Interscript's phonological layer — the normative
4+
definition of what a "hidden reading" model is, and the zoo that
5+
publishes models conforming to it.
46

5-
**Status:** skeleton (P0 in TODO.rababa/10). Framework abstractions
6-
implemented and tested. Task data modules + configs ready; full
7-
training requires GPU + dataset fetch.
7+
This repo owns three things and nothing else:
88

9-
## What this is
9+
1. **The `models.yaml` index** — the stable URL every runtime resolves
10+
model ids against (with per-artifact sha256s and split-part support).
11+
2. **The IMF v1 model-zip format** — the artifact contract:
12+
`metadata.yaml` + ONNX graphs + member sha256 manifest. The normative
13+
text is [SPEC.md](SPEC.md); the reference loader now lives in the
14+
[Python crystal](https://github.com/secryst/secryst-py).
15+
3. **The model zoo + publish pipeline** — teachers from
16+
[interscript-ml-train](https://github.com/interscript/interscript-ml-train)
17+
are distilled, gated (parity written into the artifact), and released
18+
as index entries here.
1019

11-
One training repo for every ML map in Interscript:
20+
## The system
1221

13-
- **rababa_arabic** — Arabic diacritization (adds harakat)
14-
- **rababa_hebrew** — Hebrew diacritization (adds nikud)
15-
- **secryst_thai_ipa** — Thai → IPA transliteration
16-
17-
Each task is a **config + data module**. The training framework is
18-
shared. Adding a new transliteration pair (Khmer → IPA, Japanese →
19-
Romaji) is one new directory under `src/tasks/` — zero edits to
20-
framework code.
21-
22-
## Architecture
23-
24-
```
25-
src/
26-
├── framework/ # SHARED abstractions (MECE)
27-
│ ├── config.py # TaskConfig loaded from YAML
28-
│ ├── registry.py # Plugin registry (OCP)
29-
│ ├── data.py # DataModule ABC
30-
│ ├── model.py # ModelModule ABC (teacher + student)
31-
│ ├── trainer.py # BaseTrainer + FineTune + Distill (DRY)
32-
│ ├── evaluator.py # BaseEvaluator + edit_distance + DER/PER utils
33-
│ ├── exporter.py # OnnxExporter ABC
34-
│ └── pipeline.py # TrainingPipeline orchestrator
35-
├── tasks/
36-
│ ├── rababa_arabic/ # config.yaml + data.py + student.py + metrics.py
37-
│ ├── rababa_hebrew/
38-
│ └── secryst_thai_ipa/
39-
└── cli.py # python -m src.cli train --task rababa_arabic
40-
```
41-
42-
## Design principles (project conventions)
43-
44-
- **OCP** — adding a task = one new directory. Adding a metric, model
45-
architecture, or data source = one new subclass + `@register_*`
46-
decorator. Framework code is never edited.
47-
- **MECE** — each module owns one concern. Data has no knowledge of
48-
model architecture. Model has no knowledge of trainer. Trainer has
49-
no knowledge of evaluator.
50-
- **DRY** — the epoch loop, edit-distance math, and ONNX export
51-
wrapper are written once.
52-
- **Model-driven, semantically-driven** — class names mirror domain
53-
concepts (`RababaArabicData`, `DEREvaluator`, `SecrystThaiIpaStudent`).
54-
- **Performance** — frozen dataclasses for config; lazy imports for
55-
torch so framework tests run without GPU deps.
56-
57-
## Quick start
58-
59-
```bash
60-
scripts/setup_env.sh # creates .venv, installs deps
61-
scripts/fetch_data.sh # fetch raw datasets (set env vars first)
62-
scripts/train.sh rababa_arabic # full training pipeline
63-
scripts/export.sh rababa_arabic # export student to ONNX
64-
scripts/publish.sh rababa_arabic # upload to HuggingFace Hub
6522
```
66-
67-
Or via the CLI directly:
68-
69-
```bash
70-
python -m src.cli list
71-
python -m src.cli train --task rababa_arabic --data-root data --out-root models
72-
python -m src.cli evaluate --task secryst_thai_ipa
73-
python -m src.cli export --task rababa_hebrew
23+
interscript deterministic transliteration maps + engines
24+
(ruby · js · py) │ maps that need vocalization dispatch to a
25+
│ crystal through stdlib adapters (optional)
26+
▼
27+
secryst crystals Ruby gem · pip install secryst · npm i secryst
28+
(secryst org) implement IMF v1 + models.yaml — nothing else
29+
│
30+
▼
31+
interscript-ml ◄────── models/zips resolve through this index
32+
(THIS repo) ──────► golden sets: crystals diffed against each other
33+
34+
interscript-ml-train teachers (arabic · persian · urdu + hebrew docs);
35+
(interscript org) students distilled here enter the zoo above
7436
```
7537

76-
## Adding a new task
77-
78-
1. Create `src/tasks/<name>/config.yaml` (copy from an existing task).
79-
2. Create `src/tasks/<name>/data.py` extending `DataModule`, decorated
80-
with `@register_data_module("<name>_data")`.
81-
3. Create `src/tasks/<name>/student.py` extending `ModelModule`,
82-
decorated with `@register_model_module("<name>_student")`.
83-
4. Create `src/tasks/<name>/metrics.py` extending `BaseEvaluator`,
84-
decorated with `@register_evaluator("<metric>")`.
85-
5. Run `python -m src.cli train --task <name>`.
86-
87-
That's it. No framework edits.
88-
89-
## Tests
90-
91-
```bash
92-
pytest -v
93-
```
94-
95-
Framework tests run without torch (CPU-only, fast). Training and ONNX
96-
export tests are gated behind `@pytest.mark.gpu` and require the
97-
`[train]` and `[export]` extras.
98-
99-
## Distribution
38+
Dependency directions, stated once:
10039

101-
Models ship as **IMF v1** zips (Interscript Model Format — spec in
102-
[`docs/imf-v1.md`](./docs/imf-v1.md)): byte-level tokenizer only, ONNX
103-
opset 14, sha256-verified graphs, metrics traceable to `RESULTS.md`
104-
anchors. Build/validate with `PYTHONPATH=src python -m imf pack|validate`.
40+
- **interscript-ml depends on nothing.** It is the contract: an index,
41+
a format, golden sets, and release tooling.
42+
- **Crystals depend only on the contract.** A crystal has zero
43+
interscript-core dependency — a TTS front-end can phonemize Khmer
44+
with `pip install secryst` and nothing else.
45+
- **Engines depend on crystals only optionally.** An engine without a
46+
crystal simply cannot execute maps that declare a vocalization step.
47+
- **Training depends on nothing downstream.** Teachers never import the
48+
contract; they're consumed by it (via the export gate).
10549

106-
Models reach end users through three channels (full plan in
107-
[`TODO.distribution/`](./TODO.distribution/)):
50+
## Repositories
10851

109-
| Channel | Audience | Why |
110-
|---|---|---|
111-
| **GitHub Releases** (primary) | All consumers | Versioned, immutable, checksums, tied to source tags |
112-
| **HuggingFace Hub** (canonical) | Researchers | Model cards, datasets, auto-conversion, inference API |
113-
| **jsdelivr CDN** (edge) | Browser | Edge-cached, CORS-friendly, no rate limits |
52+
| repo | role |
53+
|---|---|
54+
| [interscript/interscript-ml](https://github.com/interscript/interscript-ml) | this — contract + zoo |
55+
| [secryst/secryst](https://github.com/secryst/secryst) | Ruby crystal (the original, est. 2020) |
56+
| [secryst/secryst-py](https://github.com/secryst/secryst-py) | Python crystal — reference, owns golden generation |
57+
| [secryst/secryst-ts](https://github.com/secryst/secryst-ts) | TypeScript crystal (npm `secryst`) |
58+
| [secryst/secryst.github.io](https://www.secryst.org) | the crystals' documentation site |
59+
| [interscript/interscript-ml-train](https://github.com/interscript/interscript-ml-train) | training monorepo (arabic/persian/urdu) |
60+
| [interscript/rababa](https://github.com/interscript/rababa) · [rababa-farsi](https://github.com/interscript/rababa-farsi) · [rababa-urdu](https://github.com/interscript/rababa-urdu) | archived origins of the train monorepo (full history merged there) |
11461

115-
Per-task versioning: `rababa_arabic-v1.0.0`, `secryst_thai_ipa-v1.2.0`,
116-
etc. Each release ships fp32 + int8 + int4 variants with SHA256
117-
sidecars, SLSA provenance, and Sigstore signatures.
62+
`runtime/` in this repo is the **frozen origin** of the Python crystal —
63+
kept for provenance; live code and releases are in secryst-py.
11864

119-
Distribution phases (P2–P8) are tracked in `TODO.distribution/`. The
120-
first production release lands when phase P6 (first trained model)
121-
completes.
65+
## Environment (as implemented by every crystal)
12266

123-
## License
67+
`SECRYST_INDEX` (index URL or path; default: `models.yaml` on this
68+
repo's main) · `SECRYST_CACHE` (default `~/.cache/secryst`). Cache hits
69+
are re-verified against the index on every load.
12470

125-
BSD-3-Clause, for code and model weights alike (see `LICENSE`).
71+
License: BSD-3-Clause.

‎SPEC.md‎

Lines changed: 106 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,106 @@
1+
# interscript-ml contract — Specification (v1)
2+
3+
Normative. Key words MUST / MUST NOT / SHALL / SHOULD / MAY are to be
4+
interpreted as described in RFC 2119. This file is the canonical text;
5+
the [crystals' documentation site](https://www.secryst.org/spec.html)
6+
renders the same content for users.
7+
8+
## Conformance
9+
10+
An implementation conforms to interscript-ml v1 when it:
11+
12+
- **C1** — MUST resolve model ids against a `models.yaml` index of
13+
`version: 1`, honoring `SECRYST_INDEX`.
14+
- **C2** — MUST verify the whole-artifact sha256 before installation
15+
and re-verify every cache hit; a mismatch MUST fail loudly.
16+
- **C3** — MUST verify every `.onnx` member against the manifest
17+
sha256 map on load; members not covered by the manifest MUST NOT
18+
load.
19+
- **C4** — MUST implement the byte tokenizer exactly (§4); id
20+
sequences MUST NOT be treated as raw bytes.
21+
- **C5** — MUST produce byte-identical outputs to the reference
22+
crystal on the shared golden sets.
23+
- **C6** — SHOULD install artifacts atomically (temp file + rename).
24+
25+
## §1 The models.yaml index
26+
27+
version: 1
28+
models:
29+
<id>:
30+
filename: string # artifact file name
31+
url: string # single-file channel (http(s):// or file://)
32+
sha256: string # whole-artifact digest
33+
size: int
34+
precision: fp32 | fp16 # default fp32
35+
task: string
36+
parts: # OPTIONAL: split artifacts
37+
- url: string
38+
sha256: string # per-part digest, verified as it lands
39+
size: int
40+
41+
The `parts` mechanism exists for artifacts exceeding GitHub's 2 GiB
42+
per-asset cap: parts stream into one file in index order, each verified
43+
on arrival; the assembled file is then checked against the entry-level
44+
`sha256` exactly as a single-file model — the cache contract is
45+
identical.
46+
47+
## §2 Resolution algorithm
48+
49+
1. Resolve `<id>` in `models.models`. Unknown ids MUST raise an error
50+
enumerating known ids.
51+
2. If a cached copy exists at `<cache>/models/<id>/<filename>` whose
52+
whole-file sha256 matches, use it — cache hits are re-verified,
53+
never trusted blindly.
54+
3. Otherwise download (single URL, or parts in order) to a temporary
55+
file in the target directory, verifying digests as data lands.
56+
4. Verify the assembled artifact against the index sha256.
57+
5. Atomically rename into place, then load.
58+
59+
## §3 IMF v1 model zips
60+
61+
A zip containing at minimum `metadata.yaml`, `encoder.onnx`, and
62+
`decoder.onnx`. Zips MUST NOT rely on zip-level integrity; integrity
63+
is the manifest's job.
64+
65+
format: imf-v1
66+
tokenizer: bytes
67+
id: <model id>
68+
task: string
69+
decoder: plain | kv # kv iff decoder-kv.onnx is present
70+
precision: fp32
71+
opset: 14
72+
sha256: # every .onnx member MUST be covered
73+
encoder.onnx: <hex>
74+
decoder.onnx: <hex>
75+
76+
A conforming loader MUST reject: any other `format`, any `tokenizer`
77+
other than `bytes`, missing required members, and uncovered or
78+
mismatched member digests.
79+
80+
## §4 Byte tokenizer
81+
82+
| concept | rule |
83+
|---|---|
84+
| encoding a string | UTF-8 bytes `b` → ids `b + 3`, then one trailing `eos` |
85+
| decoding ids | stop at `eos`; skip `pad`/`unk`; `(id − 3) mod 256` per byte; reassemble as UTF-8 |
86+
| special ids | `pad = 0`, `eos = 1`, `unk = 2` |
87+
88+
Warning — the classic silent-garbage bug: ids are offset by 3 and carry
89+
a trailing EOS. Feeding `text.bytes` directly, or forgetting the EOS,
90+
produces plausible-but-wrong outputs that pass shape checks. Interop
91+
tests MUST cover both encode and decode round-trips, including
92+
multi-byte scripts.
93+
94+
## §5 Environment
95+
96+
| variable | meaning | default |
97+
|---|---|---|
98+
| `SECRYST_INDEX` | index URL or local path | `models.yaml` on this repo's main |
99+
| `SECRYST_CACHE` | artifact cache directory | `~/.cache/secryst` |
100+
101+
## Conformance kits
102+
103+
- Golden parity kit (deterministic decode-loop fixture + reference
104+
goldens): [secryst-py/parity](https://github.com/secryst/secryst-py/tree/main/parity).
105+
- Export gate (parity written into released artifacts): this repo's
106+
release pipeline (`src/imf/`, WO03).

‎runtime/README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -11,7 +11,7 @@ Home repo: https://github.com/secryst/secryst-py (this copy in
1111
ml-models/runtime is the frozen origin; the package now lives there).
1212

1313
```python
14-
from interscript_ml import Model
14+
from secryst import Model
1515

1616
model = Model.load("khm-latn-1.0") # id: index resolve -> download
1717
# -> sha256-verify -> cache -> load

0 commit comments

Comments
 (0)