The framework is designed to distinguish between cognitive failure (reasoning degradation under task complexity) and system failure (artifacts arising from infrastructure limits such as token truncation or parsing errors).
-
Updated
Sep 21, 2026 - Jupyter Notebook