Alex Leko
All content on this blog was fully or partially created using local AI (Apple MLX).

Running the Superpowers skills on a local model: Qwen3.6 35B A3B (part 3)

  • llm
  • local-models
  • qwen3
  • skills
  • l-coder

Part 3 of 4: The iterations

Series index: Part 1 · The Failure | Part 2 · The Method | Part 3 · The Iterations | Part 4 · The Lessons

You can find the final result here:


I ran the model six times, patching the skills after each run. In a sandbox repo (ldw.solutions.optiq-coder), it generated a spec and plan; I checked them against the skills' gates, made changes, and ran it again. Runs are named for the artifact dates supplied by the model. Those dates are wrong; Part 4 explains why.

The failures moved through several layers: execution bugs, plan coverage, spec coverage, constants, and finally falsifier grounding. Gates that passed continued to pass in later runs.

Run 1: 2026-08-28-knowledge-inject (skills v1)

Held: 11 assertion-shaped Acceptance Criteria; the plan directory structure, with a TOC-only index and a 165-line interfaces.md; the Iron Rule (tasks said "Create X with all functions from interfaces.md", and the implementation code appeared exactly once); RED gates with exact commands; the two-call dedup test at both levels (the pure closure and the injector instance); token-set scoring (D5 dead); ### topic nesting (D6 dead); missing-topic raises (D4 dead); dedupe state owned by a long-lived adapter object (D2 architecturally impossible).

Broke:

  • AC coverage holes: the real-directories criterion and the user-override criterion had no implementing task, and the discovery signature couldn't express three roots. The coverage-table self-review existed only as prose, so the model skipped it
  • No real-artifact count gate; every discovery test used tmp dirs (D1's killer absent)
  • An existence-only test for the wiring: assert callable(make_knowledge_injector)
  • A NameError in planned test code (an object used before assignment), plus a skipped RED step in the tests task
  • make_dedupe() -> callable had an invalid annotation under strict mypy
  • No **Spec:** path in index.md. This was a bug in my skill pack: the executor skill's Step 0 assumed it existed

Patched (v2): Spec Coverage became a required table in index.md, with STOP for any unmapped criterion. The REAL-ARTIFACT GATE became a template: run the real loader over the shipped paths, assert and print the expected count. The skills also required a **Spec:** path, banned existence-only tests, added an implementability STOP (each criterion's mechanism must appear in the Design), enforced a single header, and required the final task to run the full suite and mypy.

Run 2: 2025-07-18-knowledge-inject (v2)

Held: Spec Coverage table 12/12 rows, saved in index.md; **Spec:** line; no duplicated sections; the real-artifact gate, asserting len(entries) == 17 over the real l_coder/skills, matching the repo's actual 13 knowledge + 4 protocols; PASTE the output instructions propagated into the task steps (an emergent echo of the executor skill); two-call dedup test; RED gates.

Broke, at the root of the Iron Rule: interfaces.md contained six ... stubs, so "Copy from interfaces" was impossible. The tasks re-typed the implementations: 15 function definitions in the domain task and the entire Registry class in the adapter task. As the rule predicted, the copies drifted:

  • The MIN_SCORE_THRESHOLD = 2.0 vanished entirely from the code and AC; the adapter pre-filtered score > 0
  • Message content became a dict: {"customType": ..., "text": ...}. This was an invalid OpenAI content schema, and no criterion pinned it
  • A dead private cross-adapter import (_scan_dir), redundant double-scoring, a "lc-knowledge" magic string in two places defined nowhere
  • A coverage fidelity wobble: AC9 said "13 files ⇒ 13 entries"; the mapped test created 3
  • The NameError test bug recurred, and the single-header rule (then prose) was violated. The spec was dated 2025-07-18 (real date: 2026-08-28), a hallucination that persisted

Patched (v3): stubs were banned, with a ... grep added to self-review. A Constraint Coverage table mapped each constraint to a verifying AC and task. The Message Schema Rule required message content to be a string or typed parts; custom shapes needed Design justification and a pinning AC. Other changes made the single-header check mechanical, required named tests to reproduce their AC's stated input, added implementability and constraint-traceability STOPs, and restricted formatters to check mode. The scorecard grew to 16 defect rows, from D10 through D16.

Run 3: 2025-08-28-knowledge-inject (v3), the run that executed itself

Held: stub scan empty (140 lines of complete interfaces code); Constraint Coverage table present; the real-count rule fired mid-brainstorm: the model committed a spec correction (docs: fix entry count in knowledge injection specification) before planning; single-header 6/6; the content shape was a proper string with customType/display as sibling fields, pinned by an AC.

The surprise: after planning, the model started implementation on its own. It invoked executing-plans-small without asking, then executed the domain, port, and adapter tasks. The cause was in my skill: its Handoff section said "Execute with executing-plans-small," a directive rather than the user choice offered upstream. I changed it to a STOP gate: "Review the Spec Coverage and Constraint Coverage tables, then say go."

One useful result: the executed code matched the interfaces.md signatures exactly, with zero drift, in atomic per-task commits. The Iron Rule works when interfaces.md contains complete code.

Still broken: the threshold was present but unverified in the Design flow and planned code (with a comment!), but absent from every AC and from Global Constraints. It appeared only in flow prose. This became the project's deepest defect class.

Patched (v4): the spec template gained a ## Constants section (NAME = value # why, exercised by AC-n). The CONSTANT RULE says a constant absent from that section does not exist. Self-review now includes a CONSTANT SWEEP, the plan copies constants 1:1 into Global Constraints and Constraint Coverage, and the executor's Step 0 reads the Constants section.

Run 4: 2025-08-28-knowledge-inject (v4)

Held: the Constants section appeared (4 rows); Constraint Coverage matched 1:1; interfaces were complete (... grep clean); the single-header rule passed; self-review verified the real count (13, knowledge-only scoping); the model saved into index.md unprompted; dedupe state lived on LoopState (state._last_knowledge_block); message shape was plain; and the handoff STOP gate held, with spec and plan commits only.

Broke: all four constants' AC references were wrong (AC-1..AC-4 pointed at parse and discovery tests). AC-5 claimed score == 4.0 (two phrase matches), but the arithmetic is phrase 2.0 + word 1.0 = 3.0. AC-6 was self-contradictory: "selects first two (150+100=250 > 200, so only first one)". And the threshold exclusion test still didn't exist.

Found while fixing (a review false-negative): in my run-4 review, I had marked interfaces.md complete. During the fix pass, I found a re-typed select_entries in task-01. The cause was docstring-only stub bodies in interfaces.md; my round-3 grep missed them because there was no ... line. A function containing only a docstring silently returns None. The re-typed version also diverged semantically: it used continue on budget overflow (skip and try smaller) where the Design said stop. The two copies of the function already disagreed.

Patched (v5): FALSIFIABILITY RULE (a constant's criterion must fail if the value changes); ARITHMETIC CHECK (every numeric EXPECTED includes its computation); STUBS ARE FORBIDDEN in any disguise; and an AST-based STUB + RE-TYPE SCAN parses interfaces.md code blocks, flags docstring-only bodies, and catches interfaces names redefined in tasks. Constraint Coverage gained a Falsifier column. At the user's request, brainstorming questions adopted the grill-me frontier format: numbered questions with recommended answers, one round at a time. The change was about format only; it added no sub-agents or plan-file ending. I also fixed the experiment artifacts: AC-5 became 3.0, AC-6 was resolved, AC-16 (threshold exclusion) was added, constants were re-referenced to AC-16/14/6/8, the planned assert score == 4.0 was corrected, and the exclusion test was added to the task.

Run 5: 2025-08-28-knowledge-inject (v5)

Held: the Falsifier column, with 3 of 4 cells genuine ("Entry with cost 200 would exceed budget"; "Budget of 100 would allow entry with cost 150"); a new constant (TOKEN_ESTIMATE_DIVISOR = 3.5) with its computation inline, ceil(700/3.5) = 200, showing the ARITHMETIC CHECK in use; the real-artifact gate asserting 17 in the recursive scan, plus an unprompted smoke-test step (l-coder --preset minimal); the handoff STOP gate; and a post-plan commit where the model strengthened its own test, replacing an existence assertion with behavior assertions.

Broke: the threshold row failed for the third time. The Falsifier cell said the right thing ("Entry with score 1.5 would be excluded") but pointed to the scorer test, which contains no entry scored 1.5. The reference was false even though its wording was accurate. The AC list also lost its numbering, shifting every AC-n reference. ruff check. --fix (a mutating verify) recurred, and the date was 2025.

Patched (v6): the GROUNDING RULE requires the Falsifier cell to describe data that literally appears in the named test's code. If the numbers in the cell aren't in the test source, the row is false and the process must STOP. Grounding also applies to the spec-level sweep: the criterion's input must contain the falsifying data, so a happy-path criterion cannot ground a constant. The AC list must be numbered so AC-n references can be checked. I also numbered ACs 1 through 14, added AC-14 (the exclusion criterion) with a grounded test (scored = [(3.0, entry_hi), (1.5, entry_lo)] appears in the test body), and re-referenced all four constants with grounding notes.

Run 6: 2025-08-28-knowledge-inject (v6, final review)

Held, the cleanest run: STUBS: NONE. RE-TYPED: NONE. The AST scan produced its first fully clean report, confirmed twice. The run also had genuine falsifiers (3 of 4), inline arithmetic, the real-artifact gate (17, recursive scan), a smoke-test step, and a working handoff STOP. The model again strengthened a weak test on its own.

Still open: fresh regeneration omitted the manually added AC-14, so the threshold exclusion test was missing again. Manual fixes don't survive regeneration; only gates do. The model also attached the constant to the scorer test again. The GROUNDING rule should require the correct test on the next pass. Other remaining issues: ruff check . --fix in the verify task (twice now), the 2025 date in every run, and ambiguous requires_tools scoping.

The migration pattern

Run Skill version Where the failures lived
1 v1 plan coverage (unmapped criteria, missing gates)
2 v2 representation (... stubs → re-typing → drift)
3 v3 spec coverage (constants in prose, auto-start)
4 v4 constants grounding (false AC references, arithmetic)
5 v5 falsifier grounding (right sentence, wrong test)
6 v6 residuals only (threshold test pending, cosmetics)

Each round's fixes held in every later run. No defect regressed twice. The remaining failures moved one layer deeper with each patch.


Continue to Part 4 · The Lessons