Repository navigation
Breaking change: a boolean type (#1637, value-model release phase 1) - #1647
Merged
Merged
Conversation
The slot layer already carried TAG_BOOL; seven slot->Value bridges collapsed it to make_num(1/0). VAL_BOOL now exists as two immortal singletons (make_bool), the bridges keep it, and EQ/NE/LT/GT/LE/GE/NOT and the observer predicates push slot_from_bool. The NUM_REUSE branches in EQ/NE/NUM_CMP that wrote the 0/1 result into a heap-number operand are deleted (a bool cannot live there). Arithmetic on a bool raises (operators already refused non-numbers; sum/mean's flattener now refuses a bool element under strict), and ==/!= between a bool and a number raises via values_equal_op. Witness: tests/test_bool.eigs ([0fb]); 24 of 58 rows fail on v0.44.0. Plant: OP_LT pushing a number -> 4 rows red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TOK_TRUE/TOK_FALSE sit inside the TOK_IS..TOK_LOCAL run, so d.true still parses as a field (pinned in test_bool). AST_BOOL joins every AST_NULL case list in compiler/lint/lint_host/eigenlsp; OP_TRUE (95) and OP_FALSE (96) are appended after OP_SET_LOCAL_INTERNAL and pinned in test_opcode_abi.c. The LSP no longer documents true as 1. lib/eigen.eigs returns host bools for the literals, comparisons and not. Editor grammars highlight the two keywords. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The comparison emitter cmov's the false/true TAG_BOOL bits instead of 0.0/1.0; NOT flips a bool operand (xor 1) and produces bool bits for a number; JUMP_IF_* gets a bool arm (test $1,%al) ahead of the immediate-num arm, so a compiled 'if a < b' no longer bails. OP_TRUE/OP_FALSE join the inline immediate pushes and the last_imm peephole set. Measured (callgrind Ir, hot 'if i < 7' loop, 200k iterations): JIT 91,283,299 / interpreter 177,267,336; with the JUMP_IF bool arm removed the JIT run is 167,918,256 (bails each iteration). jit_smoke: new 'JIT bool' case pins the emitted slot bits and a run to the end of the chunk; both plants (no JUMP_IF bool arm; comparison emitting numbers) turn it red. [0fb] adds interpreter and forced-JIT rows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…item 4)
C builtins (every path, soft ARG_GUARD values included) now return bool:
predicates has_key contains starts_with ends_with list_contains
secure_equals try_parse channel_closed task_alive file_exists is_dir
is_file shared_has eigen_model_loaded; two-state statuses task_send
task_kill seed_random mkdir chdir rm rename remove_file write_text
stream_open stream_write stream_close tensor_save store_delete
store_update store_drop gfx_open audio_music_play audio_stream_push;
dict fields math_flags.{overflow,invalid,underflow},
trajectory.{observed,numeric}, gfx_poll shift/ctrl/alt, sandbox_run ok.
json_encode/store encode emit true/false; json_decode, the store decoder
and the db boolean column give bools.
lib: every predicate's literal 1/0 returns become true/false (85 functions
found by return-shape scan, pkg cmd_* exit codes excluded), and lib's own
'(pred of x) == 1/0' call sites become truthiness tests.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TRACE_FORMAT_VERSION 5 -> 6. The writers already render a bool as true/false; parse_value_p now reads them back as bools instead of 1/0. A v5 tape recorded a predicate's answer as 1/0 and is refused by the existing version check (exit 3) rather than misread. docs/TRACE.md records the bump and the compat decision. [0fb] gains a record/replay row (the probed file is deleted before replay, so the bool must come from the tape) and a v5 refusal row; planting make_num(1.0) back into the decoder turns the replay row red. test_tape_observer_config, test_trace_context and test_tsan stamp and refuse against v6/v5. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Band decision: a bool is one bit, entropy false 0.0 / true 1.0 (the values the old 1/0 comparison results had; no existing constant moves). A bool is not numeric, so the predicates read it on the entropy channel. The bool-slot skip in the observe sites (interpreter SET/name-post paths, the two JIT observe helpers, unobserved-block sampling, the SIGUSR1 query view, observer_entropy_now) is removed: a bool slot is observed through its singleton. Null remains the only unobserved value. Observer corpus (14 programs): no answer changed between the skip and the observed binary; vs the v0.44.0 goldens only printed predicate results change (0 -> false, 7 lines in two iLambdaAi programs), recaptured. [0fb] rows: a bool binding is observed, entropy 0/1, a constant flag converges, a flipping flag oscillates; restoring the skip turns 4 red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every in-repo site that compared a predicate or comparison with 1/0 now raises under #1637, so each is migrated (no shim): truthiness tests, == true / == false, or an explicit conversion where a count is wanted. tests/: the .eigs assertions, run_all_tests.sh goldens (T10-T13d, RA, HS, SB, WC, TP_), the child .sh/.py expectations (strict math, replay, read_bytes cap, fifo, loop halting, exe path, args import, handles MT fixtures, road goldens, non-finite buffer golden, tensor io limit), the C embed/sandbox tests (eigs_value_as_bool, VAL_BOOL ok), strict shape cases and the re-reviewed exempt-body pins (bool cases only; no argument-shape change). lib: the remaining mixed returns (handle_key, lab.is_stable, eigen _is_alnum). docs: SPEC's booleans section rewritten, every executed fence's output, COMPARISON/SYNTAX/llms.txt/README/BUILTINS/STDLIB/EMBEDDING prose. The embed API gains eigs_value_new_bool / eigs_value_as_bool; num of a bool raises (whether it should convert is undecided in #1637). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
The account paying for this security review has reached its Codex usage limits. The payer can check the Codex usage dashboard. For personal accounts, using credits requires enabling “Use credits for security reviews” in Code review settings. If you do not manage the paying account, contact this repository's admins. |
| /* #1637: the result is a bool immediate. The old NUM_REUSE | ||
| * branches wrote a 0/1 result INTO a heap-number operand; a bool | ||
| * cannot live there, so they are gone, not retagged. */ | ||
| int _r = (_ad == _bd); |
| /* #1637: the result is a bool immediate. The old NUM_REUSE | ||
| * branches wrote a 0/1 result INTO a heap-number operand; a bool | ||
| * cannot live there, so they are gone, not retagged. */ | ||
| int _r = (_ad != _bd); |
The freestanding profile's smoke expected 'print of converged' to print 1; with the boolean type it prints true. Missed by the item-7 migration (tools/, not tests/); caught by the CI freestanding lane. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
7 tasks
values_equal had a non-raising mode, and list_contains/list_index_of used it, so list_contains of [[1], 1 == 1] answered false: an old 1/0 membership check silently flipped from found to not-found. The non-raising mode is gone: values_equal_impl always raises on bool vs num, naming the operator or the builtin. Callers of the non-raising equality (git grep values_equal): exactly list_contains and list_index_of, both now raise. contains/index_of are string-only (wrong types are argument errors), sort/sort_by refuse bools already, match compiles to EQ, and no C builtin searches dict values. test_bool: 10 membership rows; 5 fail on 5df0b99, and restoring the exemption turns the same 5 red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d 2, item 2) EIGS_STRICT=0 kept the soft stand-ins for a bool: abs of true was 0 and sum of [true, true, 1] was 1 (bools counted as absent), so true read as 0. Now a bool where a number is wanted raises in both modes: - ARG_GUARD/_TAPED/_PRETAKE and STRICT_REQUIRE raise when the guard's want text names a number (eigs_want_numeric) and the argument, or a top-level element of it, is a bool; - tensor_total, tensor_to_flat and the byte-list validator refuse a bool element in both modes; - builtins that read a number with no typed guard (other wrong types answer null/[]) refuse a bool: usleep, state_at, num_copy, list_slice, fill's count, zeros, random_normal, softmax/log_softmax/relu/leaky_relu, reshape, nearest_in_range(_all), build_corpus. Witness: tests/test_bool_numeric.py ([0fb]) derives the numeric builtins by probing the binary (unary, list slots, and the numeric slots of every runnable strict_shape_cases row) and requires a type error for a bool in both modes: examined=309/309 unary=50 list=39 tuple_slots=134, violations=0; on 5df0b99 violations=191. Plant (guards strict-only again): 181 violations. test_bool also runs under EIGS_STRICT=0. Exempt shape-contract pins re-reviewed (bool refusal only) and re-pinned. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dicate (#1637 round 2, item 3) point_in_polygon toggled a 0/1 number; it now toggles a bool, and test_geometry asserts true/false instead of the laundered 1/0. tests/test_bool_lib_predicates.py ([0fb]) re-scans every public lib predicate by RUNNING it: the predicate-shaped names from --api (51) on a generic battery plus per-predicate true/false cases, recording type of each answer. Any non-bool answer fails, as does a predicate seen answering only one way unless it is listed with a reason (observer.is_* take a fresh parameter binding, observer_slots.is_diverging, check_openai, and json.json_has, whose always-true answer is a pre-existing bug, filed out of scope). Two pkg predicates are skipped: loading lib/pkg.eigs runs its CLI. Result: examined=51/51 nonbool=0; on 5df0b99's lib, point_in_polygon fails (non-bool num, answered only false). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…1637 round 2, item 4) lib/eigen.eigs evaluated a bare predicate with a 1/0 heuristic over an entropy it never updated (dH was always 0), so eigen_run of 'type of converged' was num and 1 where the VM says bool and false; the named form 'converged of x' did not parse at all (cannot call num). Predicates now go through the same value-only bridge report uses: the host predicate asked about a fresh parameter binding holding the value (the last observed value for the bare form, the operand for the named form, which the meta parser now accepts). That is always a bool and is the VM's answer for a binding without a full window. tests/test_bool_meta_diff.py ([0fb]): 43 snippets (comparisons, not, literals, predicate builtins, bare and named observer predicates, the bool-vs-number raises) through the VM and eigen_run; 0 disagree, 11 on 5df0b99. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…2, item 5) Editing a v6 tape's 'N 0 file_exists=true' to '=1' replayed as the number 1 at exit 0, and a malformed token was reported as 'malformed v5 ...' on a v6 tape. trace_replay_take now checks the parsed value against the kind its builtin returns (bool: file_exists, is_dir, is_file, mkdir; num: random, random_int, monotonic_ns/ms, clock_unix, heap_inuse) and refuses a bool/num mix with exit 3: trace: tape format v6 record 'N 0 file_exists=1': file_exists returns a bool, the tape holds a num; refusing to replay The malformed-value and malformed-stream-id messages (replay and the stepper) print the real format version. [0fb] adds both tamper directions; on 5df0b99 the file_exists tamper replays [1, "num"] with rc 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, item 6) Self-chosen extensions re-checked against the round-2 rule: deep [true] == [1] raises (kept, pinned); ordering two bools, or a bool against a number, raises (kept, now pinned in test_bool); the EIGS_STRICT=0 soft values for bools are gone (item 2). Re-audit of the item-7 migration: the lanes this box's release suite does not execute still asserted 1/0 for bools. Fixed, and run against a 'make server' build (http+net+model): test_http_server HS19/HS22 (shared_has), HS34 (contains), test_model_roundtrip MRT02/03b/06/07b, run_all_tests TR1/ TR6 (eigen_model_loaded) -> HTTP_SERVER 45/0, model roundtrip 0 FAIL; test_db DB21/DB24 (SQL booleans) migrated but not run (no PostgreSQL here). BUILTINS.md's db boolean row now says true/false. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ver raw reads (#1637 round 3, item 1) Critic r2 found a bool silently used as a number wherever C read a number behind `type == VAL_NUM ? x : default` or with no check at all: slice bounds, list_insert_at/list_remove_at indices, sgd_update lr, numerical_grad eps, screen_put/screen_render, set_observer_thresholds, buf_get, task_sleep. - eigs_num_arg / eigs_list_num / eigs_opt_num: the checked number reads. A bool or any other non-number raises a type error. 153 raw `items[N]->data.num` argument reads now go through them, and the check-then-default sites (sgd/numerical_grad lr/eps, eigen_generate, build_corpus, http_serve port, net timeouts, model token ids) raise instead of defaulting. - Builtin-call bool gate: every C-builtin call goes through eigs_call_builtin. A bool in a position its k_bool_policy entry does not declare raises before the builtin runs, in every strict mode. This closes `add of true` -> null, `exit of true` -> exit 0 and `pi of true`. - VM: a bool slice bound raises (an omitted bound is still the default). `x at <bool>` raises. - Guards (ARG_GUARD/STRICT_REQUIRE) raise for a bool two levels deep in every mode. - tensor_cells_numeric: numerical_grad_rows/_cols wrote `cell->data.num` into a bool cell, i.e. into the immortal true/false singleton. Cells are now checked before anything mutates. - classify of a trajectory read the bool `observed`/`numeric` flags as "not a number" and defaulted them. They are bools now. - nearest_in_range(_all): `active: false` skips the entity. - tools/num_read_check.sh (suite [0fb]) reports: `num-read: examined=320 allowlisted=16 violations=0`. It also prints `bool-gate: call_sites=9 ungated=0`. Its allowlist is tools/num_read_allowlist.txt; every entry has a reason and must be used. - strict_differential: list_insert_at/list_remove_at probes for their new guards; 8 exempt contract pins re-pinned. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
… round 3, item 2) The round-2 probe treated a slot as numeric only when a string there raised, so it never looked at slots that soft-accept a string. That is where the critic's slice-bound and lr/eps findings lived. tests/test_bool_fuzz.sh (bash) and its generator tests/bool_fuzz_gen.eigs (EigenScript) replace tests/test_bool_numeric.py. - Population: every builtin and extension name in `--api --json`. A name this build lacks is counted absent, never as a raise, and every core builtin must be examined. `report` is skipped by name (a reserved form). - Slots: the whole argument; elements 0 and 1 of a list argument; every argument of the reviewed call in strict_shape_cases.json; and the first element of each list-literal argument. - VM operands: index (list, string, buffer, dict), slice bounds, range forms, arithmetic, ordering, ==, list/string repetition, `at` line, `when` ordinal, index stores and for-in. - Each call runs under EIGS_STRICT=1 and 0. A call that returns is a violation unless its name|slot is on tests/bool_fuzz_anyvalue.txt. Every entry there has a reason, and a listed slot that never returns is stale and fails the run. - A crash, timeout or missing end marker counts as broken. Broken is a FAIL. Current binary: `BOOL_FUZZ: examined=308/349 absent=40 skipped=1 core=261 ... probes=4408 raised=4252 anyvalue=156 violations=0 broken=0 unused_anyvalue=0 PASS`. At 79e800a the same probe prints violations=1524 broken=4 FAIL. Those violations include slice-start/end, at-line, list_insert_at, buf_get, task_sleep and sgd_update lr. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…re theirs (#1637 round 3, item 3) Critic r2 showed replay_expected_kind covered 10 names. A tampered `N 0 random_normal=true` replayed as ["bool", true] at rc=0. - k_tape_kinds (src/trace.c) declares the return kinds of every taped builtin, all 37 that record an N value. Replay refuses (exit 3) a record whose kind its builtin cannot return. The message names the record, the builtin, its kinds and the tape version: `random_normal returns a list or null, the tape holds a bool`. It also refuses a name that is in no table. - tools/tape_kinds_check.sh (suite [0fb]) requires taped == declared. "Taped" means every TRACE_NONDET_*, ARG_GUARD_TAPED/PRETAKE, trace_replay_take or trace_nondet_value name in src/*.c. Output: `tape-kinds: taped=37 declared=37 missing=0 stale=0`. - Embedding API: eigs_trace_declare_kind(name, EIGS_KIND(EIGS_TYPE_X) | ...) declares a host-recorded name's kinds at registration. Replay refuses a record of another kind, and a host name that was never declared. The call refuses a runtime name, an empty name, FN/OTHER kinds and an empty set. - Tests: - [0fb] tamper row for random_normal=true. - test_trace_context (41 -> 44 checks) adds three checks: a declared host name refusing a bool record (forked child, exit 3 + message), an undeclared host name being refused, and the declare API refusals. - embed_smoke and test_trace_correspondence declare their host names. - embed_smoke: a #1417 check read a comparison result with eigs_value_as_num(r) == 1.0. A comparison is a bool now, so the check uses eigs_value_as_bool. `make embed-smoke` printed 8 failures before this change and `embed_smoke: OK` after it. - docs/TRACE.md, docs/EMBEDDING.md: the table, the declaration, the refusal. Fixed the stale "version 5 grammar" there. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…und 3, item 4) SPEC.md and llms.txt said "every numeric builtin given a bool raises". Critic r2 showed that was false: slice bounds, lr/eps and index slots accepted a bool silently. Items 1-2 made the rule true. The docs now state the rule exactly: - A bool raises in arithmetic, ordering, indexing, slice bounds, `range`, an `at` line and a `when` ordinal, in every strict mode. - A builtin raises on a bool anywhere it does not take one: as its argument, at a position of its argument list, or as an element of a list it reads as numbers. - The slots that do take a bool are reviewed and named; the full list is tests/bool_fuzz_anyvalue.txt, checked by tests/test_bool_fuzz.sh. Owner decision on #1637: keep deep =='s short-circuit. It is now documented: a deep comparison walks pairs in order (lists by index, dicts in the left operand's key order) and stops at the first unequal pair, so a bool/number pair raises only when the walk reaches it. This comes with an executed example (`[1, true] == [2, 1]` is false; `[2, true] == [2, 1]` raises). The SPEC bool example now also executes a bool slice bound and `sum of [1, true]`. At 79e800a that fence fails: `FAIL: docs/SPEC.md:532 (rc=0)`. SPEC population pin goes from 84/80 to 85/81. The changelog fragment is updated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…round 3, item 5)
The company is retiring its Python tooling, so the two remaining round-2
probes are rewritten. tests/test_bool_numeric.py was already replaced in
item 2.
- tests/test_bool_lib_predicates.sh replaces tests/test_bool_lib_predicates.py.
The population is the same: 51 predicate-shaped lib functions from
`--api`. The battery is the same: 24-value POOL tuples plus the reviewed
cases. The rules are unchanged: a return must be a bool, every
predicate must answer, and each must answer both true and false unless
it is one_sided with a reason. It still asserts
`examined=51/51`. Cases, one_sided and skip rows live in
tests/bool_lib_predicates_cases.txt.
Two holes in the old probe are closed:
- Old: a non-bool answer was only seen when it printed as `@type:value`
with no whitespace, so `return "not an email"` passed it.
Measured on a planted lib/validate.eigs, the .py printed
`nonbool=0 ... PASS`. The .sh prints the type on its own line and
reports `FAIL: validate.is_email: returned a non-bool (bool=10,str=38)`.
- Old: a zero-parameter predicate was called as `name of (n)`.
It is now called as `name of null`.
- tests/test_bool_meta_diff.sh replaces tests/test_bool_meta_diff.py.
It runs every snippet in tests/bool_meta_snippets.txt as a VM program
and through eigen_run, and asserts `snippets=N/N`. The 43 snippets are
now 47: the deep-== short-circuit pair, a bool index and `range of true`.
A plant that made eigen_run's `<` answer 1/0 printed `disagree=4 FAIL`.
- The three children print a PASS:/FAIL: marker for the runner's #988
vacuity rule. [0fb] runs them with bash.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…1637 round 3, item 3) The tapes hand-crafted a dict under `env_get`, which only returns a string. The kind table refuses that record now (exit 3), so [42a/47] failed with `dict replay (out='')`. The parser round-trip it tests is unchanged: the dict, and the list[dict, dict, buffer], now ride `ls`, which returns a list. [42a/47] is back to `RESULTS: 203/203 passed`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…round 4, item 1)
Critic r3 found `gather of [[true, 3], 0]` -> 0. The cause was a
type-check-then-0.0 soft default split over two lines, at
builtins_tensor.c:1523 and :1536. net_send at ext_net.c:362 had the same
shape. num_read_check.sh passed it because its soft-default test was
line-local text matching. That is the third round where the leak lived
inside the previous round's own heuristic, so the gate is now structural:
- Value's number member is `data.num_` and EigsSlot's double is `d_`.
`v->data.num` no longer compiles:
`error: 'union <anonymous>' has no member named 'num'`.
- Exactly two ways in:
- The checking accessors, which raise on a bool or any other non-number:
eigs_num_arg (now an inline fast path plus eigs_num_arg_slow),
eigs_list_num, eigs_opt_num, and the new eigs_elem_num. A bool element
raises in every mode; any other non-number raises under EIGS_STRICT
and reads as 0.0 under EIGS_STRICT=0.
- The raw macros VAL_NUM_RAW(v) / SLOT_NUM_RAW(s) / VAL_NUM_OFFSET (JIT),
for a type already proven.
All 180 raw reads in builtin files (builtins*.c, ext_*.c) now go through
eigs_num_arg. Raw tokens remain only in the core runtime.
- tools/num_read_check.sh is rewritten as token-exact:
- Every raw token in src/*.[ch] is attributed to its function (or
#define). Each (file, function) must be on tools/num_read_allowlist.txt
with its EXACT count and a reason.
- It also fails on: a missing row, a stale row, a row without a reason,
`data.num_`/`.d_` spelled out (a bypass), an ungated `->data.builtin(`
call, and examined == 0.
- Output: `num-read: examined=214 functions=69 allowlisted=69
violations=0 stale=0`.
- slot_as_num: same treatment. It had no callers, so it is deleted;
EigsSlot's `.d` became `d_` and its 90 reads are SLOT_NUM_RAW tokens,
counted by the same gate. Raw slot reads could leak the same way
(`slot_is_num(s) ? s.d : 0`).
- gather (2-D and 1-D cell read, and a bool where a row belongs) and
net_send's byte list read through eigs_elem_num. The descriptor
param_count of vm_run_bytecode/sandbox_run refuses a non-number instead
of reading it as 0.
- test_bool.eigs: two gather rows (depth 2 and 3). On the frozen r2 binary
they fail: `FAIL: gather of a bool cell: <no error, returned 0>`.
- strict_shape_contracts.json: 37 exempt helper pins re-pinned. The only
change in each is the mechanical accessor rename.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…ound 4, item 2) The fuzz never put a bool inside a nested argument, because first_elem_template gave up when the first element was a list. So gather's cell read (`[[2, 3]]`, depth 2) was never probed. - tests/bool_fuzz_gen.eigs builds templates from each reviewed case's own nesting. Every numeric literal at bracket depth 1-3 inside each argument (and in a dict value, keyed by its dict key) is replaced, one at a time, by true and then by false. Slot names are aN.dD[.key]. - Builtins without a case get `f of ([[v, 1], 1])` and `f of ([[[v, 1]], 1])` (e0.d2, e0.d3). - A depth-2/3 probe has a control: the reviewed call itself, or the same shape with the number 7. A probe whose result digest (`sha256 of str of result`) equals its control's is counted `unread`: the bool sat where the builtin does not look. Any other return is a violation unless the slot is on the any-value list. Depth-1 probes stay strict. exit and throw get no controls and no deep probes. - Probe count: calls 2468 -> 3384, probes 4408 -> 5992, plus 546 controls. - nearest_in_range(_all): a bool px/py raised nothing and the entity was skipped as "no position". It now raises in every mode; a string still gets the documented skip. `active: true/false` is listed as any-value. - Any-value additions, each with a reason: render/JSON at depth, dict and store record values, the active flag. Seven nullary nondeterministic builtins (clock_unix, random, ...) are also listed, because their controls cannot match. Results: - Current tree: `BOOL_FUZZ: examined=308/349 ... calls=3384 controls=546 probes=5992 raised=5124 anyvalue=276 unread=592 violations=0 broken=0 unused_anyvalue=0 PASS`. - 8532f2b (r3 head): `violations=28 ... FAIL`, with gather a0.d2 and nearest_in_range(_all) a0.d2.px/.py. - Plant (c), the gather fix reverted on this tree: `violations=4 FAIL` (`gather slot a0.d2 given true returned without raising`). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…ound 4, item 3) Critic r3 measured startup (0 iterations) at 1,345,804 -> 2,095,203 instructions, +56%. Every suite program pays startup. callgrind_annotate attributed 757,169 Ir (36% of the run) to bool_gate_build. That was eager work at registration, with three costs: - an O(n^2) alias search over ~350 builtins; - a strcmp of every builtin against all 26 policy rows; - a qsort of the table. Now: - The table is built on the first call that actually hands a builtin a bool (eigs_bool_gate_slow; double-checked under a mutex, published with release/acquire). The source is the running state's pristine builtin layer, limited to its first g_builtin_binding_count bindings, so host functions registered later stay ungated. - The build itself is cheaper: one qsort of the function pointers, aliases merged as neighbours, then one env_get + bsearch per policy row. Paired callgrind Ir, deterministic (two identical runs each), bench0 = the critic's 0-iteration program: - r2 (79e800a): 1,346,531 - r3 (8532f2b): 2,096,001 - this commit: 1,347,863 (+1,332 vs r2, +0.1%) The first bool to reach a builtin pays the build once: `print of 1` 1,165,065 vs `try: abs of true` 1,432,426, so +267k Ir once. The per-call cost is unchanged and was accepted: bench50k r2 368,783,895, r3 379,513,295, now 378,765,306. A bool first reached inside spawn workers (4 threads) and inside a task still raises type_mismatch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…ment (round 4, item 4) The fragment now lists every non-bool case that changed: - string `at` lines; - net timeouts and the http_serve port; - non-number bytecode bytes and param_count; - numerical_grad/sgd_update cells; - screen_put color and the write_bytes append flag; - gather under EIGS_STRICT. It also tells C embedders about the VAL_NUM_RAW / SLOT_NUM_RAW rename. Each case was run against the frozen r2 binary and this tree: `what is xs at "s"` r2 `RET null` -> `RAISE 'at' line must be a number, got str`. `sgd_update of [[1, "s"], [1, 1], 0.1]` r2 `RET [0.9, -0.1]` -> raises in both modes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…1637 fragment (round 5) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…#1637 round 5, item 7) The documented embedding example read its arguments with eigs_value_as_num, which answered 0.0 for a bool, so `host_add of [true, 1]` was 1. Owner decision: - eigs_value_as_num returns NaN for a bool, so the mistake propagates visibly in host C. Every other non-number keeps the documented 0.0. - EMBEDDING.md says so and points hosts at eigs_value_type / eigs_value_as_bool. Its host_add example (and embed_smoke's) now checks both operands are numbers and answers null otherwise. - embed_smoke gains 4 checks: - a bool reads as NaN; - eigs_value_as_bool reads it; - a string still reads 0.0; - `host_add of [true, 1]` is null. With the old eigs_value_as_num planted into the amalgamation: `FAIL src/embed_smoke.c:489 eigs_value_as_num of a bool is NaN` / `embed_smoke: 1 failure(s)`. Now `embed_smoke: OK`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
… type The previous pin (cf6ef5e) compares a bool with 0 in src/lcd.eigs:122, which raises under #1637. DMG#79 migrated it dual-compatibly (green on v0.44.0 too). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
… round 5b) tools/ui_surface_check.sh failed in CI's precheck: `UNTESTED UI FUNCTION: _ev_flag`, `UNDOCUMENTED UI FUNCTION: _ev_flag` and `_ev_stop`. The gate has no private-helper exemption: its self-test requires `_ui_private_probe` to be both asserted and documented. So both helpers get rows in docs/STDLIB.md's ui-surface-contracts region, next to _combo_matches, and _ev_flag gets four assert_eq rows (true, 1, 0, null). `UI surface: 381 definitions; 381 assertion-linked; 381 documented; 0 test gaps; 0 doc gaps`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Oct 6, 2026
…rted table (#1637 round 6) The merge lane's Ir gate (check_regression.sh --vs main 01bbf7e) failed: observed_loop and unobserved_loop were +7%. Reproduced locally against a 01bbf7e build in a separate worktree: `REGRESSION: unobserved_loop Ir 65341043 -> 70132863 (+7% > 5%)` (observed: 65333173 -> 70124860). Attribution: a callgrind function diff of bench/unobserved_loop.eigs, main vs r5. - malloc_consolidate +3,021,549 and unlink_chunk +1,343,766 account for 4.37M of the +4.79M. gdb on _int_free_maybe_consolidate puts the cause in eigs_bool_gate_slow -> qsort_r -> free. - The bench's only bool is the final `print of (x > 0)`. That first bool triggered round 4's lazy table build. Freeing qsort's merge buffer next to the top chunk made glibc consolidate every fastbin, and the fastbins held the loop's 60,000 freed number Values. - None of the other suspects shows up in the diff: per-call gate scan, NUM_REUSE, JUMP_IF on bool slots, checking accessors, observer. - On/off: removing the build removes the whole delta. Fix: no table, no allocation on the gate's slow path. - On the first bool, the ~30 k_bool_policy names are resolved to function pointers with env_get into a static array. A call's policy is a linear pointer scan. - Host functions are marked when eigs_register_function registers them (eigs_bool_gate_exempt). Compiled natives and the sandbox's blocked stub stay ungated. That is the same population the table's layer limit excluded; test_bool gains a row for the blocked stub. - The builtin's name is looked up only when raising. After (check_regression.sh --vs main): - observed_loop 65,550,857 (0%) - unobserved_loop 65,558,748 (0%) - every row `ok`: `no instruction-count regressions vs reference (threshold 5%)` - unobserved_loop function diff vs main: +217,500 total (+0.33%), all of it jit_helper_iter_next +165,000 and env_store_slot +44,992. The malloc rows are back to main's. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Oct 6, 2026
… shared_set takes any value (#1637 round 7, item 1) CI's server-db lane ran the bool fuzz on the full build and found 20 violations. The release build never compiles model_*.c or ext_http.c. - eigen_generate a0.d1, eigen_eval_loss a0.d1, native_train_step_builtin a0.d1/a1.d1: the token-id lists were read through eigs_num_arg (which raises), but only AFTER the "no model loaded" early return. Without a model, a bool id came back as the no-model answer. A bool anywhere in the arguments (two levels, BOOL_REFUSE) now raises first, in every mode, and for eigen_generate before the tape boundary. - shared_set a1 is a value store: the http shared store holds any JSON value, and a bool round-trips through shared_get as a bool. It is already BA_POS(1) in k_bool_policy, and now has a reasoned any-value row. - Audit of the other extension value-store slots: - shared_incr's delta is numeric (the gate refuses a bool); - net_send's byte list reads elements through eigs_elem_num (a bool raises); - db parameters take a bool as SQL boolean (round 5); - the http route/static slots take strings. - tools/num_read_check.sh scans src/*.c, model_*.c included; those files hold no raw number token. Local server-db build: `BOOL_FUZZ: examined=348/349 absent=0 skipped=1 core=261 ... probes=6768 raised=5740 anyvalue=280 unread=748 violations=0 broken=0 unused_anyvalue=0 PASS`. CI on 4c7941d showed `violations=20 FAIL`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
round 7, item 2) CI's db lane failed `Error line 148: ASSERT FAIL: DB17 message names the type rule`. Round 5 changed the db parameter error to "not a string, number or bool", and DB17 still expected "not a string or number". The assertion now matches the new message. New live row DB27 (runs only with the CI postgres service): - a true/false parameter inserted into a BOOLEAN column reads back as true/false; - a bool read from the db writes back; - a bool parameter compares with a boolean column. There is no PostgreSQL server here. Locally, the no-connection half passes on the server-db build (`PASS: all DB builtin checks (live-DB checks skipped: no connection)`). CI's postgres service is the confirming oracle for DB27 and for the typed-parameter binding. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…1637 round 7, items 3-4) On the full build, both model capability probes failed for the same reason: `eigen_model_loaded of null` answers `false` (a bool, as BUILTINS.md documents) and both probes still expected 0. - tools/strict_differential.sh's cap_contract asserted `result == 0`, which raised `cannot compare bool and num`. The script reported `CAPABILITY DID NOT MEASURE: model ... unexpected capability error` and exited 1 before a verdict ([99s]). It now asserts `result == false`. - [47i]'s probe compared the printed output with `0` and printed `false`: `FAIL: model capability probe (rc=0)`. It now expects `false`. Local server-db build, EIGS_SUITE_SECTIONS='0fb 47i 99s 46 47': all rows PASS ([99s] `PASS: every guard raises from its own`, [47i] the context-window rows). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
round 7) The server lane's `make ui-sdl-input-gfx` failed with `Error line 1: cannot compare bool and num with '=='` at `(gfx_open of [...]) == 1`. gfx_open returns a bool, and the event shift/ctrl/alt modifiers are bools. The C harness's eigen checks now compare with true/false. NOT RUN locally: this box has no SDL2 headers (SDL.h), so the harness cannot be built here. CI's server lane is the confirming oracle. tests/ui_native_input.py compares JSON true with 1 in Python, which holds (True == 1), so it needs no change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
…#1637 round 7, item 1) The shared_set any-value row is used on the server builds but unused on the release build, where ext_http is absent. The fuzz's stale-row check therefore failed the release lane (`unused_anyvalue=1 FAIL`). A row whose name was counted absent on this build is now skipped. An unused row for a PRESENT name still fails: a planted `abs|u` row prints `FAIL: unused any-value entry abs|u` / `unused_anyvalue=1 FAIL`. Results: release `... violations=0 broken=0 unused_anyvalue=0 PASS`; the same under bash 3.2; server-db `examined=348/349 ... unused_anyvalue=0 PASS`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Oct 6, 2026
…ls (#1637) Both now return a bool (round 1's status-code decision); their shape cases still asserted shape_result == 0, which raises under #1637. Only the macOS lane has SDL (audio/gfx), so Linux skipped these cases and the merge queue's macOS suite was the first to run them (4258/4260). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Oct 6, 2026
… audio_music_play (#1637) The previous commit fixed the valid_assert; the backend_absent.outputs goldens still expected the observed print to be 0 (macOS merge-queue lane: 'FAIL shape: audio_music_play — valid max/0: 0 false shape-complete'). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
github-merge-queue
Bot
removed this pull request from the merge queue due to failed status checks
Oct 6, 2026
… under a sanitizer build (#1637 round 8) CI's `asan + ubsan / shard 2/3 (http+model build)` was cancelled at its 55-minute limit; on main it takes 12-16 min. Shard 2/3 (section_plan.sh --emit-shard 2 3) carries [0fb]. Measured locally on `make asan-server` with CI's ASAN_OPTIONS=detect_leaks=1: - `SECTION_TIME: [0fb] 1097.62`, with every one of the 698 fuzz program runs and every lib-predicate probe `broken` at rc=134: `AddressSanitizer failed to allocate 0xdfff0001000 bytes ... ReserveShadowMemoryRange failed ... Perhaps you're using ulimit -v`. - tests/test_bool_fuzz.sh and tests/test_bool_lib_predicates.sh wrapped each probe in `ulimit -v 1500000` (the gfx-safety cap). ASan reserves ~15 TB of shadow address space at startup, so every probe aborted, and each abort cost ~1.5 s. - This was not a hang: no bool-gate mutex, publish ordering or tape path was involved. Fix: when the binary carries a sanitizer runtime (__asan_init / __tsan_init / __msan_init), the cap is lifted (VCAP=unlimited, with a printed NOTE). Video and audio stay on SDL's dummy drivers. Coverage is unchanged: the full probe set runs, no subset. Measured on asan-server after the fix: - tests/test_bool_fuzz.sh: 66 s, `examined=344/349 ... probes=6688 ... violations=0 broken=0 unused_anyvalue=0 PASS`; - [0fb]: `SECTION_TIME: [0fb] 105.86`, `RESULTS: 14/14 passed`; - shard 2/3: `RESULTS: 1717/1717 passed, 0 failed, 1 skipped`, wall 627 s. Main 01bbf7e's asan-server shard 2/3 on the same box: `1703/1703`, wall 861 s. Main's figure includes 350 s of first-time eigenlsp/eigsdap ASan builds in [88]/[126], which this tree already had. Excluding those, the rest is 494 s on main and 496 s here; [0fb] adds 106 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1637. This is phase 1 of the value-model release, #1644.
This is a breaking semantic change, not a refactor. Programs observe different output (
print of (1 < 2)→true), and some programs that used to run now raise.Semantics (owner decisions on #1637)
booltype:type of trueis"bool". Keywordstrue/false; new opcodesOP_TRUE/OP_FALSE(95/96, appended).not, the observer predicates, predicate builtins, andlib/predicates.and/orstill return an operand.print/strgivetrue/false, and JSON, the store and the db round-trip them.==/!=between a bool and a number, so every oldpred == 1fails loudly instead of flipping silently.Implementation
VAL_BOOLuses the slot layer's existing BOOL tag (0xFFF9), which was reserved but discarded until now. All 9 slot→Value collapse sites are fixed, and the NUM_REUSE result-in-operand branches are deleted.NOTpaths, plus a boolJUMP_IFfast path. With it a hotif i < 7loop costs 91M instructions; without it, 168M.changes/breaking/.Evidence
tests/test_bool.eigs: 24 of 58 rows fail on v0.44.0.OP_LTreturning a number;JUMP_IFfast path removed (instruction count back to 168M);trueas a number;jit-smokeandjit_diff(258 programs) OK; doc examples 193/193;test_boolleak-clean under ASan withdetect_leaks=1in all three tiers. Full ASan/TSan come from this PR's CI.Builder-chosen extensions, pending owner confirmation (see #1637)
bool == numraises at any depth ([true] == [1]).EIGS_STRICT=0, numeric builtins return their soft value for a bool, butsumandnum ofraise regardless.Follow-ups (lockstep or later)
🤖 Generated with Claude Code
https://claude.ai/code/session_01KC99CmwatKssQYggkBCgwF