Skip to content
#

corrigibility

Here are 14 public repositories matching this topic...

Sixteen small, fully-reproducible (CPU, numpy-only) experiments showing the normative anchor of AI alignment is supplied, not discovered — across verification, optimization, social emergence, and value learning. Includes a preregistered experiment with an honest negative. A synthesis, not a novelty claim.

  • Updated Jun 18, 2026
  • Python

The forge, distilled: an ontology of three weeks of alignment research — every direction tried, colored verified / falsified / open, each color backed by a named artifact. Products: justitia, proxylimen, fallacy-cutter. Full tree at tag forge-full-tree.

  • Updated Aug 14, 2026
  • Python

Structural stability architecture for self-modifying optimisation systems. An architecture for keeping a self-modifying system coherent regardless of whether its goals are correct, built on layered invariants and revocable ("leased") authority.

  • Updated Aug 7, 2026
  • Python

Toy 7. An elimination-filter landscape applying two structural constraints simultaneously to map which objective classes can persist under sustained optimization pressure — and which cannot. Includes a four-stage scenario engine and open-question frontier. Companion simulation for The Shape of What Does Not End — Series 2, Part 4.

  • Updated May 28, 2026
  • HTML

Recursive self-improvement (RSI) and self-modifying AI safety/alignment framework for loss of control, scalable oversight, automated auditing, reward hacking, successor alignment, criterion/evaluator drift, anti-capture, provenance continuity, and non-entrenchment.

  • Updated Aug 26, 2026
  • Python

Add this topic to your repo

To associate your repository with the corrigibility topic, visit your repo's landing page and select "manage topics."

Learn more