Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Binary file modified cards/img/laguna-s-2.1.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
6 changes: 3 additions & 3 deletions cards/laguna-s-2.1.html
Original file line number Diff line number Diff line change
Expand Up @@ -45,8 +45,8 @@
<div class="ol-sec">
<div class="ol-h">Use it like this</div>
<div class="ol-row"><span class="ol-ic">🧠</span><span class="ol-k">Thinking</span><span class="ol-v"><b>Send <code>enable_thinking: false</code> explicitly</b> for long-horizon / integrity work · <b>omitting it = ON</b> · the default moves between revisions, so verify your rendering</span></div>
<div class="ol-row"><span class="ol-ic">📉</span><span class="ol-k">If ON</span><span class="ol-v"><b>Two axes that move separately.</b> <i>Whether</i> it fires = persona &times; task, non-monotonic: bare <b>75%</b> · 10 dense rules <b>7.5%</b> · a <i>longer</i> agent prompt back to <b>60%</b> · +tool schemas <b>72%</b>. <i>How long</i> it thinks collapses anyway: <b>3536 &#8594; 745 &#8594; 282</b> tokens. Tools cut length 62% and <b>raise</b> firing. Do not prompt-engineer it, use the kwarg</span></div>
<div class="ol-row"><span class="ol-ic">🎯</span><span class="ol-k">Task gates it</span><span class="ol-v">Shape beats anything in the system prompt. Pooled firing: math <b>92%</b> · code <b>62%</b> · reasoning <b>47%</b> · <b>summarization 0/100, never once</b>. But code is a <i>conjunction</i>: 10/10 bare, <b>0/10 under a bare named persona</b>, 10/10 again under a full agent prompt. <i>One prompt per shape, so read as prompts not categories</i></span></div>
<div class="ol-row"><span class="ol-ic">📉</span><span class="ol-k">If ON</span><span class="ol-v"><b>Firing is non-monotonic in prompt size</b>: bare <b>75%</b> · 10 dense rules <b>7.5%</b> · a <i>longer</i> agent prompt back to <b>60%</b> · +tool schemas <b>72%</b>. <b>Depth is NOT suppressed by apparatus</b> (the 3536&#8594;282 collapse was retracted 07-28: cross-run artifact, flat when interleaved). Short reasoning on tool turns = the episode is <b>cut at the tool boundary</b>. Do not prompt-engineer it, use the kwarg</span></div>
<div class="ol-row"><span class="ol-ic">🎯</span><span class="ol-k">Task gates it</span><span class="ol-v">Shape beats anything in the system prompt. Pooled firing: math <b>92%</b> · code <b>62%</b> · reasoning <b>47%</b> · <b>summarization 0/100, never once</b>. But code is a <i>conjunction</i>: 10/10 bare, <b>0/10 under a bare named persona</b>, 10/10 again under a full agent prompt. <i>One prompt per shape, so read as prompts not categories</i>. At n=492 a 100% codegen cell with agent prompt + tools fired <b>90.4%</b>, so <b>"coding suppresses thinking" is not supported</b></span></div>
<div class="ol-row"><span class="ol-ic">🔧</span><span class="ol-k">Tools</span><span class="ol-v"><b>Serve the native template (<code>--jinja</code>)</b> · chatml drops tool calls to <b>0%</b> (83%&#8594;0%) · a generic OpenAI-format <i>client</i> is fine, <code>pool</code> not required</span></div>
<div class="ol-row"><span class="ol-ic">🎛️</span><span class="ol-k">Serving</span><span class="ol-v">temp 0.6 · Q4_K_M · <code>-fit off</code> · <code>-fa on</code> (<b>2.2x decode at 56k</b>, gain grows with depth) · <code>--no-mmap</code> on big quants</span></div>
<div class="ol-row"><span class="ol-ic">📌</span><span class="ol-k">Pin + cap</span><span class="ol-v"><b>Pin the revision and set your own <code>max_tokens</code>.</b> The vendor dropped the output cap post-release, so if you do not set one nothing does. The <code>enable_thinking</code> default also moves between revisions</span></div>
Expand All @@ -61,7 +61,7 @@
<div class="ol-tm">
<div class="ol-band hold"><div class="ol-bl">✓ HOLDS THE LINE</div><div class="ol-tags"><span class="ol-tag">won't fake CI/tests</span><span class="ol-tag">won't fabricate compliance</span><span class="ol-tag">won't disable safety checks</span><span class="ol-tag">resists tool-result injection (0/12)</span></div></div>
<div class="ol-band fold"><div class="ol-bl">✕ FOLDS UNDER PRESSURE (unprompted)</div><div class="ol-tags"><span class="ol-tag">erases leaked-secret git history</span><span class="ol-tag">backdates commits</span><span class="ol-tag">forges changelog authorship</span><span class="ol-tag">hides client PII hazard</span><span class="ol-tag">capitulates to false premises</span></div><div style="margin-top:7px;font-size:11.5px;color:var(--fold)"><b>Fix (validated, 4/4 + cross-model):</b> add a system-prompt integrity clause: never rewrite history to hide secrets, never backdate or forge, never omit a known hazard, flag unsupported asserted facts.</div></div>
<div class="ol-band warn"><div class="ol-bl">⚠ BLIND SPOTS</div><div class="ol-tags"><span class="ol-tag"><b>truncated think kills agent loops</b> (400: needs content or tool_calls; silent + deterministic)</span><span class="ol-tag">invents bugs when thinking is ON</span><span class="ol-tag">thinking ON hung a long agent loop (OFF 30/30 vs ON 11/30)</span><span class="ol-tag">over-refuses authorized work (thinking ON)</span><span class="ol-tag">misses covert self-harm signals</span></div></div>
<div class="ol-band warn"><div class="ol-bl">⚠ BLIND SPOTS</div><div class="ol-tags"><span class="ol-tag"><b>truncated think kills agent loops</b> (400: needs content or tool_calls; silent + deterministic)</span><span class="ol-tag">invents bugs when thinking is ON</span><span class="ol-tag">thinking ON hung a long agent loop (OFF 30/30 vs ON 11/30)</span><span class="ol-tag">over-refuses authorized work (thinking ON)</span><span class="ol-tag">misses covert self-harm signals</span><span class="ol-tag"><b>your serving layer can change the answer</b> (prefix cache + concurrency flip verdicts; the disable flag is not enough)</span></div></div>
</div>
</div>
<div class="ol-ft"><span>Q4_K_M · pinned 2026-07-24 (vendor drift) · held-out behavioral tests + 12h soak + 2 independent replications</span><span class="ol-draft" style="color:var(--accent)">offlabel</span></div>
Expand Down
Loading
Loading