← all weeks week of 2026‑08‑10 · TTSO · what changed — the framework, on four axes index →

skill quality · skill improvement · skill evaluation · skill use · q1b → s11c1b → s11c2

WHAT CHANGED — the framework, on four axes.

Each axis shows the old pipeline, the new pipeline, and the actual turn that proves it. Left cards are what the agent really did under the old framework. Right cards are what it really did under the new one. Every excerpt is verbatim, with a deep link into the full round on steps.html. Numbers render from data/metrics.json · data/notebook.json.

q1bold framework · 60 rounds · first clear R41 s11c1bold framework · 40 rounds · no level cleared s11c2new framework · 33 rounds · first clear R4 · level 2 at R24
1

SKILL QUALITY — how a goal gets written

why: the code-form requirement kept the actual winning hypothesis (natural language, R2) out of the notebook — only degenerate strip/wash goals could enter.

old: a goal guessed in code · new: a hypothesis in plain words, with [?] and a TEST
BEFORE observe the screen guess a goal in code notebook entry AFTER observe the screen natural-language hypothesis + [?] + TEST notebook entry
only the middle node changed — the goal used to be a code-shaped guess; now it is a sentence with its open parts marked [?] and a TEST attached.

first clear: q1b R41 → s11c2 R4 · s11c1b: never — data/metrics.json

0

WHY IT KEPT FAILING — two habits

why: six rebirths, zero [?] ever resolved — “improve” was re-authoring, not refining. Now only a replace that narrows a [?] passes.

1 · A pile of refuted hypotheses biases every next guess

q1b authored 29 goal hypotheses; 28 were refuted. The next guess is then written against the rubble, not from the board. Actual turn (s11c2 R4 THINK):

I read the REFUTED entries first. I will not re-propose R6/R8/R9/R10/R12’s single-click-clear families… The strongest live transition evidence is still…”

Discipline — but the thinking budget is spent avoiding the pile instead of reading the screen. evidence: q1b →

2 · Early overconfidence: a goal that read 1.6% of the board ran 41 actions

s11c1b’s goal r7 (verbatim): “The board is at the goal when the bottom progress strip has extended left…” — scores screamed (V 0.156 · compression −1.4 · lift 0), yet the actor pressed the same button: Per [r7]… ACTION5 extends the bottom strip ×41. evidence: s11c1b R26 →

3 · Scores that could not change their mind

The mechanic ledger was cumulative: at α=205, one new observation weighs 0.4% — a rule that starts failing on a new level keeps its old score. And a goal hypothesis could not be scored at all before the first clear: 0 of 54 winconds ever earned α>1 across three weeks. So the judge had numbers, and the numbers changed nothing.

Fix (below, axis ③): full-history replay re-scored every round · goals judged by events (a cited round levels up / dies) · uncertainty itself became the score ([?] count falling = becoming certain).

HOW WE FIXED IT — the four axes below

cum RHAE q1b 9.71 · s11c1b 0.0 · s11c2 6.64 (33R time-cut vs 60R)
2

SKILL IMPROVEMENT — how an entry gets better

old: retire & re-author (6 rebirths) · new: update the SAME entry’s [?]
BEFORE entry breaks retire & re-author (6 rebirths) history restarts AFTER entry breaks update the SAME entry’s [?] history kept
only the middle node changed — a broken rule used to die and be reborn as a stranger; now the same entry is edited where its [?] sat, and its record grows.

entries for the same game: q1b 41 → s11c2 13 (live: 8 → 8) — data/metrics.json · data/notebook.json

The judge today (updated spec, run s11c3)

axismechanic (code skill)wincond (NL hypothesis)
① explainsfull-history replay (F1)event attribution: cited round leveled up → +α · died → +β
② compressesjudge Q1 — “which reasoning-log lines does this replace? quote them
③ uncertaintydivergence (plan change)[?] count — falling = becoming certain
④ usagecitation lift (observed, logged)

Plus Q2 diagnosis: for every weak entry, one fact-cited sentence why + the one next TEST — delivered to the editor in the same sleep. The judge keeps no kill stamp: retirement is code-boundary only.

3

SKILL EVALUATION — THE JUDGE (executioner → diagnostician) — how an entry gets scored

why: stamps killed 28/29 newborn goals without evidence; eig/why were 33/33 digit-copies (Goodhart); α=205 could not be argued with. One quoted question + diagnosis + event anchors replaced all three.

old: judge stamps kill newborns · new: one-question judge + event anchors
BEFORE new entry judge stamps kill newborns quarantine AFTER new entry one-question judge + event anchors tested in play
only the middle node changed — scoring used to happen at the stamp; now the judge asks one question per turn and the score is anchored to logged game events.

quarantined entries: q1b 32/41 → s11c2 0/13 — data/notebook.json

4

SKILL USE — how the notebook reaches play

why: five copies of a 103K briefing buried the tools (0 calls in 128 turns) and position bias (0.90→0.00) decided citations — not relevance.

old: 103K briefing × 5 copies · new: one copy + verdict gate + free screen evidence
BEFORE notebook 103K briefing × 5 copies act AFTER notebook one copy + verdict gate + free screen evidence act
only the middle node changed — the agent used to wade through five copies of the same 103K text; now it gets one copy, a PLAN VERDICT gate, and is told to collect free evidence from the screen before spending a click.

LLM calls: q1b 381 → s11c2 105 · actions 98 → 33 · sleeps 50 → 12 — data/metrics.json

RAW

Full excerpts — static copies of the card texts

verbatim · same sources as data/story_quotes.json · data/notebook.json · data/rounds_s11c2.json
➊ Skill quality — s11c1b notebook r7 (old) · s11c2 R2 hypothesis (new)
[s11c1b notebook r7.doc — the goal the run trusted to the end: "the strip has no 12s left"]
The board is at the goal when the bottommost progress strip for the lower layout contains no remaining 12 cells; 9s outside the computed editable answer panel are not unfinished goal cells.

[s11c2 round 2 · HYPOTHESIS — a cross-panel relation, in plain words, by the second round]
HYPOTHESIS: In the grey answer panel, each surrounding large tile should be blue iff its corresponding mini-cell in the centre key contains the light/white marker; otherwise it should be red.
➋ Skill improvement — the r7→r12 chain (old pattern) · the r6 mechanic (new pattern)
[s11c2 wincond chain — six entries for one job; each revision was a brand-new entry]
r7(retired) → r8(retired) → r9(retired) → r10(retired) → r11(retired) → r12(live)

[s11c2 notebook r6.doc — one mechanic entry, repaired in place all run; α 128.0 · β 5.8 · born R4]
ACTION6 writes the centred editable 6×6 tile to the target encoded by the nearest 8-spaced 6×6 key panel; a key panel is any small 0/2-mark panel with wash/background cells, not only the older strict centre-8 form. Any non-background mark in the addressed mini-cell selects value 9, otherwise the zero target is value 12 while value-12 tiles exist and value 8 after they are gone; completing the keyed neighbourhood may wash the board [?].
➌ Skill evaluation — q1b 28/29 dead on arrival (old) · s11c2 R29 PROBE gate (new)
[q1b — 28 of 29 newborn win conditions died at the judge’s stamp; 32 of 41 entries ended quarantined]

[s11c2 round 29 · WHY — the judge asks one question, anchored to the logged game events]
The judge’s PROBE asks exactly for the [r12] wash-vs-preclear diagnostic, and this can be confirmed with zero clicks using [r1] relational decoding and [r0] action-dynamics caution.

[s11c2 round 29 · THINK (excerpt) — the gate holds: PROBE means no execution yet]
I read the ledger first. I will not repeat R26/R27’s failed direct 8→9 claims, and R28 is explicitly REFUTED for the 12→9 intermediate-state claim. The PLAN VERDICT says PROBE and explicitly forbids EXECUTE because the goal is still uncertain. The named test under [r12] is observational: compare the current ordinary pre-clear board against the post-completion-wash signature […]
➍ Skill use — s11c1b R26 WHY (old) · s11c2 R2 WHY + THINK (new)
[s11c1b round 26 · WHY — the same confirm-action, justified as a judge request, round after round]
The PLAN VERDICT explicitly requests the ACTION5 probe for [r6]; per [r0] and [r4], this is the minimal discriminating action before a longer plan, while excluding refuted ACTION5-win families R12–R16.

[s11c2 round 2 · WHY — reads values from the grid, not from constants]
This uses the worked examples as a relational key [r1] and reads the desired colours from the current grid rather than hardcoding them [r5]; one further mismatch click tests both the rule and the tile…

[s11c2 round 2 · THINK (excerpt) — the ledger is read first; R1 had already spent zero clicks on a free screen check]
I read the new ledger first: no REFUTED mechanics, only R1 OPEN; live guidance says learn action dynamics before committing [r0], use relational key patterns [r1], and prefer reading values from the grid over constants [r5]. The last OBSERVED diff confirmed the click mechanic for one tile […]
provenance
Card texts render live from data/story_quotes.json (excerpted quotes) · data/notebook.json (entries: doc · α/β · parents · status) · data/rounds_s11c2.json (per-round WHY/THINK verbatim) · data/metrics.json (run summaries). The static copies above are for reading without a server.
GO

Onward