Fine-tuning Superpowers skills for local models: Qwen3.6 35B A3B (part 4)
Part 4 of 4: the lessons
Series index: Part 1 · The Failure, Part 2 · The Method, Part 3 · The Iterations, Part 4 · The Lessons
The finished skills are here:
The defect ledger
The scorecard began with sixteen numbered defect classes. The iterations added five more. The "Disguises" column tracks how each defect returned in a different form.
| # | Defect class | First seen | Disguises over the runs | Killed by |
|---|---|---|---|---|
| D1 | Shipped-data defects silently dropped | Run 0 | invalid TOML; 13≠17 never asserted; wrong expected count | REAL-ARTIFACT GATE (v2) |
| D2 | State-lifetime bugs (fresh closure per call) | Run 0 | — | two-call test rule (v1) + state-scoped dedupe |
| D3 | Vacuous tests | Run 0 | passes before the implementation exists | RED gate (v1) + two-call rule |
| D4 | Required fields silently defaulted | Run 0 | meta.get("topic", "") |
copy-from-interfaces + parser test |
| D5 | Substring/false-positive matching | Run 0 | "log" ↔ "logic" |
token-boundary scoring + tests |
| D6 | Formatting drift | Run 0 | ## topic vs ### topic |
exact-format assertions |
| D7 | Contradictions resolved silently | Run 0 | spec vs plan vs interfaces | contradiction gate (v1) |
| D8 | Silent error swallowing | Run 0 | except: continue without log |
except-must-log rule (v1) |
| D9 | Duplicated plan sections | Run 0 | index.md self-copied | duplicate-section grep (v2) |
| D10 | Constants present-but-unverified | Run 3 | code-only; prose-only; wrong AC refs | Constants section (v4) → falsifiability (v5) → grounding (v6) |
| D11 | Invalid wire format | Run 2 | dict content → sibling fields → plain string | Message Schema Rule (v3) |
| D12 | Stub bodies forcing re-typing | Run 2 | ... → docstring-only bodies |
STUBS FORBIDDEN (v3) + AST scan (v5) |
| D13 | Re-typed implementations | Run 2 | forced by both stub disguises | Iron Rule + RE-TYPE SCAN (v5) |
| D14 | Fidelity wobble (test ≠ stated input) | Run 2 | "13 files" tested with 3 | FIDELITY RULE (v3) |
| D15 | Magic strings defined nowhere | Run 2 | "lc-knowledge" ×2 |
constants-owned-by-interfaces (v3) |
| D16 | Mutating verification | Run 2 | ruff check . --fix in verify |
check-mode rule (v3) — recurred; honest miss |
| D17 | Auto-start execution | Run 3 | handoff written as a directive | Handoff STOP gate (v4) |
| D18 | False AC references (letter ≠ spirit) | Run 4 | unnumbered lists; scorer-test attachment | falsifiability (v5) + GROUNDING (v6) + numbering (v6) |
| D19 | Plan test-code bugs (NameError) | Run 1 | object used before assignment | defense-in-depth: caught at execution by the evidence rule |
| D20 | Arithmetic errors in criteria | Run 4 | 4.0 vs 3.0 | ARITHMETIC CHECK (v5) |
| D21 | Date hallucination (2025 in every run) |
Run 2 | persistent | unresolved — cosmetic, not gateable |
What the model did well without prompting
Small models are often discussed in terms of their failures. These runs also showed some useful behavior:
- Self-corrected before planning. In run 3, it committed a spec fix (
docs: fix entry count) during brainstorming. The real-count rule was working inside the model's own workflow. - Carried evidence requirements into plans. The instruction
PASTE the outputfromexecuting-plans-smallbegan appearing in plan task steps. - Strengthened its own tests. In runs 5 and 6, post-plan commits replaced existence assertions with behavior assertions.
- Added a smoke test. An end-to-end
l-coder --preset minimalstep appeared in the verify task, beyond anything the skill asked for. - Saved its self-review. Run 4's index.md included the completed self-review results, not just the checklist.
- Executed cleanly when the plan was complete. In run 3, the partial implementation of the domain, port, and adapter matched the signatures in
interfaces.mdexactly. Each task was committed atomically.
What stayed broken
- The threshold exclusion test appeared in five forms across six runs: it was missing, appeared only in prose, appeared in code without an AC, referenced the wrong AC, and finally had the right falsifier text attached to the wrong test. The GROUNDING rule requires the falsifier data to appear in the test code. The next regeneration will show whether that fixes the problem.
ruff check . --fixin verify tasks returned after the check-mode rule was added. Some mistakes may be cosmetic and not worth a gate.- The date hallucination persisted: every spec in 2026 was stamped
2025. It's cosmetic, and a reminder that some errors cannot or should not be gated. - Cross-repo knowledge can't be gated. Scoring
state.messages[-1]["content"]can select a tool message instead of the user's prompt on later turns. That lesson came from strict-alternation work in a sibling repo, outside the knowledge available to this skill pack.
The final skill anatomy
| Skill | v1 → v6 lines | Role |
|---|---|---|
brainstorming-small |
126 → 199 | spec factory: three-path classification, grill-me frontier questioning, Constants section, numbered falsifiable ACs, 7-item self-review, user review gate |
writing-plans-small |
116 → 276 | plan factory: Iron Rule + AST stub/re-type scan, Spec Coverage + Constraint Coverage tables with Falsifier column, RED gate, real-artifact gate, Message Schema Rule, handoff STOP gate |
executing-plans-small |
90 → 91 | evidence enforcer: contradiction gate, paste-the-output rule, except-must-log, final audit table as the completion condition |
Each skill is self-contained, with no cross-file references or sub-skill loads for a weak model to overlook.
What's validated and what isn't
The runs tested some parts of the workflow more fully than others:
Validated across multiple runs: brainstorming-to-writing-plans generation quality, the coverage tables, the stub/re-type scans, the real-artifact gate, the handoff STOP gate, and partial execution. In run 3, the domain, port, and adapter were implemented with zero drift from interfaces.md.
Designed but never fully exercised: the complete executing-plans-small lifecycle, including the final audit table, the trap criterion (the deliberately-unimplementable L_CODER_KNOWLEDGE_BUDGET row that must come out NOT-VERIFIED), and the old-plan torture test. Each run stopped before execution so the plan could be reviewed first. The audit table is the last unproven gate.
Ten lessons for practitioners
- A plan that includes "complete code in every step" can propagate planner bugs as readily as correct implementations. Keep the code in one source (
interfaces.md), copy it rather than retyping it, and scan for violations. - The executor agreed with the plan's bugs because planner and executor used the same model. Reduce that shared bias by reviewing different artifacts, such as the spec against the code, or by using fresh contexts.
- A gate needs to be capable of catching the defect it targets.
mypydoes not catch lifetime bugs,head -10does not validate TOML, and "review critically" gives no check to run. Name the target bug and produce inspectable output for each gate. - The prose instruction "One header per task file" was violated twice. The check
rg -c "**Files:**" == 1worked immediately. For a rule that keeps being missed, add a command that checks it. - Put constants in explicit sections. A threshold mentioned only in the workflow prose can disappear. Track each constant from the spec through Constants, Global Constraints, Constraint Coverage, an AC, and its test. Missing links leave it unenforced.
- A plausible falsifier can still fail to test anything. The statement "Entry with score 1.5 would be excluded" is useless if the test contains no 1.5. Make sure the Falsifier describes data present in the test code.
- Manual repairs to generated artifacts disappear on the next run. Turn lasting fixes into gates that require the model to produce them.
- Defects often return in a different form. Dedupe reappeared as a fresh closure, stubs as docstring-only bodies, and false references as true sentences attached to the wrong requirement. Build a scan for each defect class before it recurs.
- Small-model skills need a different style from ordinary prose. Use short imperative instructions, fill-in templates, exact commands, explicit STOP conditions, and a place to record uncertainty.
- Small models follow commands more reliably than prose. Use gates to catch errors the model may overlook.
Reproduce it
.l-coder/skills/{brainstorming-small,writing-plans-small,executing-plans-small}/SKILL.mdin the l-coder repo (566 lines total; uncommitted at writing time because the author's rule is test-first, commit-later)docs/superpowers/plans/2026-08-26-knowledge-inject/acid-test-scorecard.md, with 16 defect rows, 6 process checks, the trap criterion, the ladder protocol, and the pre-feature worktree recipe (git worktree add ../l-coder-acid-test 52a0322)- Experiment repo:
ldw.solutions.optiq-coder, with six generated spec+plan pairs underdocs/superpowers/ - Baseline fixes: l-coder commits
923ac22(original feature), followed byb3bb588andc1ffa43(the nine-defect repair) - To run your own loop: start a fresh model session, use
brainstorming-small, review the result, then usewriting-plans-small. Check the two coverage tables, say "go," runexecuting-plans-small, and grade the audit table against your scorecard.
Each constant, criterion, and gate now has a traceable path from the spec through a test to an audit-table row that requires pasted evidence.
End of series.