← all weeks week of 2026‑08‑10 · TTSO · from early overconfidence to early uncertainty prev: 2026‑07‑30

test-time skill optimization · ARC-AGI-3 · weekly report · 2026-08-04 → 08-10 · q1b / s11c1b / s11c2

The agent that played worse as its skills grew learned early uncertainty — and became 10× faster.

Every prescription that added knowledge failed. The disease was not missing knowledge but early overconfidence, and the cure was a structure that defers conviction — the hypothesis ladder · explicit [?] unknowns · mandatory TEST procedures · no EXECUTE while the goal is unconfirmed. The six acts below are that whole arc. Every number is loaded and drawn from data/*.json.

R41 → R4first level clear · q1b → s11c2 · 10× R45 → R24L2 reached · q1b → s11c2 · 2× 381 → 105LLM calls · q1b → s11c2 41 → 8notebook entries · q1b total → s11c2 live
1

Problem — the more skills, the worse the play

q1b · 60R · self-intoxication
q1b’s 41 notebook entries — one cell per entry. Green = 8 live · red = 32 quarantined · grey = 1 retired. Outlined cells are the 29 win conditions — the agent rotated through XOR · AND · majority-vote · symmetry combinations, and 28 of the 29 newborn entries died instantly at the judge’s stamp. The briefing bloated to 103K characters, five duplicated copies of the same content. Source data/notebook.json · run q1b.

The very ability to accumulate skills became the disease: 41 notebook entries · 29 rotating win conditions · 28/29 newborns dead on arrival · a 103K-character briefing — and across 60 rounds the first level clear only came at R41. evidence: q1b R20 →evidence: q1b R41 →details →

2

Diagnosis — split it along four axes

quality / improvement / evaluation / usage · one measurement per axis
1  Quality — are the rules true? 26 of 41 entries hardcode level numbers (“Level 0” etc.) — instantly invalid on the next level 26 / 41 2  Improvement — do the rules get fixed? 41/41 entries have an empty parents lineage — every entry is a newborn, never an update; no improvement path exists 0 / 41 3  Evaluation — are verdicts refutable? one entry’s confidence coefficient α grew to 205 — no counter-evidence can move the posterior; irrefutable α = 205 4  Usage — do the rules drive action? skills invoked as tools 0 times across 128 LLM calls — hoarded, never used 0 / 128
Axis definitions, measurement procedures, and raw log excerpts live in axes.html. Source data/notebook.json (axes 1·2·3) · q1b llm_calls (axis 4).

One measurement per axis: quality (level hardcoding 26/41) · improvement (parents 0/41) · evaluation (α=205, irrefutable) · usage (tool invocations 0/128) — all four axes in the red. evidence: q1b R20 →details →

3

First fix — minimize: the structure held, the game was still lost

slot cap · briefing diet · one-question judge → s11c1b
structural metrics (q1b → s11c1b) notebook entries 41 8 (slot cap) win conditions 29 1 LLM calls 381 104 result L0 / 40R level 0 for all 40 rounds 0 clears · run s11c1b
The diet worked; the game was still lost. Median briefing also shrank, 26.7K → 20.7K chars (data/metrics.json). We fixed the structure — so why does it still lose? Act 4 is the answer.

The slot cap (8 notebook slots) · the briefing diet · the one-question judge stopped the proliferation — yet s11c1b sat at level 0 for all 40 rounds (L0/40R). evidence: s11c1b R40 →details →

4

The new disease — early overconfidence is the bias

s11c1b r7 · a confident goal born from reading 1.6% of the board
s11c1b, 40 rounds — red cells are rounds whose reasoning (why · hypothesis · expect) is bound to ACTION5 (30/40). The r7 win condition (“if the progress strip has no 12s, that is the goal”) was born from reading just 1.6% of the board, and even facing counter-evidence — verification value V 0.156 · compression score −1.4 — it monopolized 41 actions. Source data/rounds_s11c1b.json · data/notebook.json (r7).
The Confident-Hypothesis Trap
“The board is at the goal when the bottommost progress strip for the lower layout contains no remaining 12 cells; 9s outside the computed editable answer panel are not unfinished goal cells.”
s11c1b notebook entry r7 (wincond · live) — verbatim from data/notebook.json · authored in round 7, never revised again
reads 1.6% of the board V 0.156 compression −1.4 41 actions spent
A goal this confident, born this early, turns every later verdict into armour: the judge’s stamp is read as “the confirm-action was requested”, and the same ACTION5 probe repeats through the final round while the counter-evidence (verification value 0.156 · negative compression) never touches the entry. evidence: s11c1b R13 →evidence: s11c1b R20 →evidence: s11c1b R26 →

The failure was not missing knowledge but early overconfidence: the confident goal r7, born too early, kept monopolizing behavior through the final round despite counter-evidence (V 0.156 · compression −1.4). evidence: s11c1b R26 →evidence: s11c1b R22 →details →

5

Insight — manufacture uncertainty early

hypothesis ladder · [?] · TEST · EXECUTE gate → s11c2 design
1 Observation what is visible — coordinates · colors · panels 2 Relation what changes what 3 Goal when does the level end EXECUTE gate if the goal is unconfirmed ([?]), no execution — probes only • every entry marks its unresolved parts with [?] — conviction is not the default • every goal hypothesis carries a mandatory TEST: procedure — the actual clear board (level-up frame) is the anchor • the ladder climbs bottom-up only — no goal may be authored before observations and relations stand
The s11c2 notebook, as-is: of 13 entries, 6 carry [?] unresolved marks and 5 carry TEST: procedures (data/notebook.json · verbatim text in data/story_quotes.json).

The prescription was not more information but manufacturing uncertainty early: the hypothesis ladder (observation → relation → goal) · explicit [?] unknowns · mandatory TEST · no EXECUTE while unconfirmed · the clear-board anchor. evidence: s11c2 R9 →details →

6

Result — 10× faster · the goal still stops at ‘recognition’

s11c2 · clear at R4 · L2 at R24 · zero traps
Level reached, per round. Bold line = s11c2 (first clear R4 · L2 at R24) · grey = q1b (R41 · R45) · red dashed = s11c1b (L0 for all 40R). s11c2 re-bought a refuted hypothesis family zero times. Source data/rounds_*.json.

s11c2 cleared first at R4 (10×) · reached L2 at R24 (2×) · fell into zero traps — the remaining disease is that the goal stops at recognition: with no guidance (which action moves toward the goal), L2 devolved into one-shot probes. evidence: s11c2 R2 →evidence: s11c2 R4 →details →

The evidence lives in the subpages

hover = preview · click = pin verbatim text · #anchor deep links