The agent that played worse as its skills grew learned early uncertainty — and became 10× faster.
Every prescription that added knowledge failed. The disease was not missing knowledge but early overconfidence,
and the cure was a structure that defers conviction — the hypothesis ladder · explicit [?] unknowns · mandatory TEST procedures · no EXECUTE while the goal is unconfirmed.
The six acts below are that whole arc. Every number is loaded and drawn from data/*.json.
q1b’s 41 notebook entries — one cell per entry. Green = 8 live · red = 32 quarantined · grey = 1 retired.
Outlined cells are the 29 win conditions — the agent rotated through XOR · AND · majority-vote · symmetry combinations,
and 28 of the 29 newborn entries died instantly at the judge’s stamp. The briefing bloated to 103K characters, five duplicated copies of the same content. Source data/notebook.json · run q1b.
The very ability to accumulate skills became the disease: 41 notebook entries · 29 rotating win conditions · 28/29 newborns dead on arrival · a 103K-character briefing
— and across 60 rounds the first level clear only came at R41.
evidence: q1b R20 →evidence: q1b R41 →details →
2
Diagnosis — split it along four axes
quality / improvement / evaluation / usage · one measurement per axis
Axis definitions, measurement procedures, and raw log excerpts live in axes.html. Source data/notebook.json (axes 1·2·3) · q1b llm_calls (axis 4).
One measurement per axis: quality (level hardcoding 26/41) · improvement (parents 0/41) · evaluation (α=205, irrefutable) ·
usage (tool invocations 0/128) — all four axes in the red.
evidence: q1b R20 →details →
3
First fix — minimize: the structure held, the game was still lost
slot cap · briefing diet · one-question judge → s11c1b
The diet worked; the game was still lost. Median briefing also shrank, 26.7K → 20.7K chars (data/metrics.json).
We fixed the structure — so why does it still lose? Act 4 is the answer.
The slot cap (8 notebook slots) · the briefing diet · the one-question judge stopped the proliferation — yet s11c1b sat at
level 0 for all 40 rounds (L0/40R).
evidence: s11c1b R40 →details →
4
The new disease — early overconfidence is the bias
s11c1b r7 · a confident goal born from reading 1.6% of the board
s11c1b, 40 rounds — red cells are rounds whose reasoning (why · hypothesis · expect) is bound to ACTION5
(30/40). The r7 win condition (“if the progress strip has no 12s, that is the goal”) was born from reading just 1.6% of the board,
and even facing counter-evidence — verification value V 0.156 · compression score −1.4 — it monopolized 41 actions.
Source data/rounds_s11c1b.json · data/notebook.json (r7).
The Confident-Hypothesis Trap
“The board is at the goal when the bottommost progress strip for the lower layout contains no remaining 12 cells;
9s outside the computed editable answer panel are not unfinished goal cells.”
s11c1b notebook entry r7 (wincond · live) — verbatim from data/notebook.json · authored in round 7, never revised again
reads 1.6% of the boardV 0.156compression −1.441 actions spent
A goal this confident, born this early, turns every later verdict into armour: the judge’s stamp is read as
“the confirm-action was requested”, and the same ACTION5 probe repeats through the final round while the
counter-evidence (verification value 0.156 · negative compression) never touches the entry.
evidence: s11c1b R13 →evidence: s11c1b R20 →evidence: s11c1b R26 →
The failure was not missing knowledge but early overconfidence: the confident goal r7, born too early, kept monopolizing
behavior through the final round despite counter-evidence (V 0.156 · compression −1.4).
evidence: s11c1b R26 →evidence: s11c1b R22 →details →
The s11c2 notebook, as-is: of 13 entries, 6 carry [?] unresolved marks and 5 carry TEST: procedures (data/notebook.json · verbatim text in data/story_quotes.json).
The prescription was not more information but manufacturing uncertainty early: the hypothesis ladder (observation → relation → goal) ·
explicit [?] unknowns · mandatory TEST · no EXECUTE while unconfirmed · the clear-board anchor.
evidence: s11c2 R9 →details →
6
Result — 10× faster · the goal still stops at ‘recognition’
s11c2 · clear at R4 · L2 at R24 · zero traps
Level reached, per round. Bold line = s11c2 (first clear R4 · L2 at R24) ·
grey = q1b (R41 · R45) · red dashed = s11c1b (L0 for all 40R).
s11c2 re-bought a refuted hypothesis family zero times. Source data/rounds_*.json.
s11c2 cleared first at R4 (10×) · reached L2 at R24 (2×) · fell into zero traps — the remaining disease is that the goal stops at
recognition: with no guidance (which action moves toward the goal), L2 devolved into one-shot probes.
evidence: s11c2 R2 →evidence: s11c2 R4 →details →
→
The evidence lives in the subpages
hover = preview · click = pin verbatim text · #anchor deep links