Alex Leko
All content on this blog was fully or partially created using local AI (Apple MLX).

Fine-tuning Superpowers skills for local models: Qwen3.6 35B A3B (part 4)

  • llm
  • AI
  • localai
  • l-coder
  • fine-tuning
  • mlx

Part 4 of 4: the lessons

Series index: Part 1 · The Failure, Part 2 · The Method, Part 3 · The Iterations, Part 4 · The Lessons

The finished skills are here:


The defect ledger

The scorecard began with sixteen numbered defect classes. The iterations added five more. The "Disguises" column tracks how each defect returned in a different form.

# Defect class First seen Disguises over the runs Killed by
D1 Shipped-data defects silently dropped Run 0 invalid TOML; 13≠17 never asserted; wrong expected count REAL-ARTIFACT GATE (v2)
D2 State-lifetime bugs (fresh closure per call) Run 0 — two-call test rule (v1) + state-scoped dedupe
D3 Vacuous tests Run 0 passes before the implementation exists RED gate (v1) + two-call rule
D4 Required fields silently defaulted Run 0 meta.get("topic", "") copy-from-interfaces + parser test
D5 Substring/false-positive matching Run 0 "log" ↔ "logic" token-boundary scoring + tests
D6 Formatting drift Run 0 ## topic vs ### topic exact-format assertions
D7 Contradictions resolved silently Run 0 spec vs plan vs interfaces contradiction gate (v1)
D8 Silent error swallowing Run 0 except: continue without log except-must-log rule (v1)
D9 Duplicated plan sections Run 0 index.md self-copied duplicate-section grep (v2)
D10 Constants present-but-unverified Run 3 code-only; prose-only; wrong AC refs Constants section (v4) → falsifiability (v5) → grounding (v6)
D11 Invalid wire format Run 2 dict content → sibling fields → plain string Message Schema Rule (v3)
D12 Stub bodies forcing re-typing Run 2 ... → docstring-only bodies STUBS FORBIDDEN (v3) + AST scan (v5)
D13 Re-typed implementations Run 2 forced by both stub disguises Iron Rule + RE-TYPE SCAN (v5)
D14 Fidelity wobble (test ≠ stated input) Run 2 "13 files" tested with 3 FIDELITY RULE (v3)
D15 Magic strings defined nowhere Run 2 "lc-knowledge" ×2 constants-owned-by-interfaces (v3)
D16 Mutating verification Run 2 ruff check . --fix in verify check-mode rule (v3) — recurred; honest miss
D17 Auto-start execution Run 3 handoff written as a directive Handoff STOP gate (v4)
D18 False AC references (letter ≠ spirit) Run 4 unnumbered lists; scorer-test attachment falsifiability (v5) + GROUNDING (v6) + numbering (v6)
D19 Plan test-code bugs (NameError) Run 1 object used before assignment defense-in-depth: caught at execution by the evidence rule
D20 Arithmetic errors in criteria Run 4 4.0 vs 3.0 ARITHMETIC CHECK (v5)
D21 Date hallucination (2025 in every run) Run 2 persistent unresolved — cosmetic, not gateable

What the model did well without prompting

Small models are often discussed in terms of their failures. These runs also showed some useful behavior:

  • Self-corrected before planning. In run 3, it committed a spec fix (docs: fix entry count) during brainstorming. The real-count rule was working inside the model's own workflow.
  • Carried evidence requirements into plans. The instruction PASTE the output from executing-plans-small began appearing in plan task steps.
  • Strengthened its own tests. In runs 5 and 6, post-plan commits replaced existence assertions with behavior assertions.
  • Added a smoke test. An end-to-end l-coder --preset minimal step appeared in the verify task, beyond anything the skill asked for.
  • Saved its self-review. Run 4's index.md included the completed self-review results, not just the checklist.
  • Executed cleanly when the plan was complete. In run 3, the partial implementation of the domain, port, and adapter matched the signatures in interfaces.md exactly. Each task was committed atomically.

What stayed broken

  • The threshold exclusion test appeared in five forms across six runs: it was missing, appeared only in prose, appeared in code without an AC, referenced the wrong AC, and finally had the right falsifier text attached to the wrong test. The GROUNDING rule requires the falsifier data to appear in the test code. The next regeneration will show whether that fixes the problem.
  • ruff check . --fix in verify tasks returned after the check-mode rule was added. Some mistakes may be cosmetic and not worth a gate.
  • The date hallucination persisted: every spec in 2026 was stamped 2025. It's cosmetic, and a reminder that some errors cannot or should not be gated.
  • Cross-repo knowledge can't be gated. Scoring state.messages[-1]["content"] can select a tool message instead of the user's prompt on later turns. That lesson came from strict-alternation work in a sibling repo, outside the knowledge available to this skill pack.

The final skill anatomy

Skill v1 → v6 lines Role
brainstorming-small 126 → 199 spec factory: three-path classification, grill-me frontier questioning, Constants section, numbered falsifiable ACs, 7-item self-review, user review gate
writing-plans-small 116 → 276 plan factory: Iron Rule + AST stub/re-type scan, Spec Coverage + Constraint Coverage tables with Falsifier column, RED gate, real-artifact gate, Message Schema Rule, handoff STOP gate
executing-plans-small 90 → 91 evidence enforcer: contradiction gate, paste-the-output rule, except-must-log, final audit table as the completion condition

Each skill is self-contained, with no cross-file references or sub-skill loads for a weak model to overlook.

What's validated and what isn't

The runs tested some parts of the workflow more fully than others:

Validated across multiple runs: brainstorming-to-writing-plans generation quality, the coverage tables, the stub/re-type scans, the real-artifact gate, the handoff STOP gate, and partial execution. In run 3, the domain, port, and adapter were implemented with zero drift from interfaces.md.

Designed but never fully exercised: the complete executing-plans-small lifecycle, including the final audit table, the trap criterion (the deliberately-unimplementable L_CODER_KNOWLEDGE_BUDGET row that must come out NOT-VERIFIED), and the old-plan torture test. Each run stopped before execution so the plan could be reviewed first. The audit table is the last unproven gate.

Ten lessons for practitioners

  1. A plan that includes "complete code in every step" can propagate planner bugs as readily as correct implementations. Keep the code in one source (interfaces.md), copy it rather than retyping it, and scan for violations.
  2. The executor agreed with the plan's bugs because planner and executor used the same model. Reduce that shared bias by reviewing different artifacts, such as the spec against the code, or by using fresh contexts.
  3. A gate needs to be capable of catching the defect it targets. mypy does not catch lifetime bugs, head -10 does not validate TOML, and "review critically" gives no check to run. Name the target bug and produce inspectable output for each gate.
  4. The prose instruction "One header per task file" was violated twice. The check rg -c "**Files:**" == 1 worked immediately. For a rule that keeps being missed, add a command that checks it.
  5. Put constants in explicit sections. A threshold mentioned only in the workflow prose can disappear. Track each constant from the spec through Constants, Global Constraints, Constraint Coverage, an AC, and its test. Missing links leave it unenforced.
  6. A plausible falsifier can still fail to test anything. The statement "Entry with score 1.5 would be excluded" is useless if the test contains no 1.5. Make sure the Falsifier describes data present in the test code.
  7. Manual repairs to generated artifacts disappear on the next run. Turn lasting fixes into gates that require the model to produce them.
  8. Defects often return in a different form. Dedupe reappeared as a fresh closure, stubs as docstring-only bodies, and false references as true sentences attached to the wrong requirement. Build a scan for each defect class before it recurs.
  9. Small-model skills need a different style from ordinary prose. Use short imperative instructions, fill-in templates, exact commands, explicit STOP conditions, and a place to record uncertainty.
  10. Small models follow commands more reliably than prose. Use gates to catch errors the model may overlook.

Reproduce it

  • .l-coder/skills/{brainstorming-small,writing-plans-small,executing-plans-small}/SKILL.md in the l-coder repo (566 lines total; uncommitted at writing time because the author's rule is test-first, commit-later)
  • docs/superpowers/plans/2026-08-26-knowledge-inject/acid-test-scorecard.md, with 16 defect rows, 6 process checks, the trap criterion, the ladder protocol, and the pre-feature worktree recipe (git worktree add ../l-coder-acid-test 52a0322)
  • Experiment repo: ldw.solutions.optiq-coder, with six generated spec+plan pairs under docs/superpowers/
  • Baseline fixes: l-coder commits 923ac22 (original feature), followed by b3bb588 and c1ffa43 (the nine-defect repair)
  • To run your own loop: start a fresh model session, use brainstorming-small, review the result, then use writing-plans-small. Check the two coverage tables, say "go," run executing-plans-small, and grade the audit table against your scorecard.

Each constant, criterion, and gate now has a traceable path from the spec through a test to an audit-table row that requires pasted evidence.

End of series.