# Fine-Tuning the Superpowers Skills to Run on Local Models — Qwen3.6 35B A3B

## Part 2 of 4: The Method

> **Series index:** [Part 1 · The Failure](2026-08-28-superpowers-local-models-1-the-failure.md) — Part 2 · The Method — Part 3 · [The Iterations](2026-08-31-superpowers-local-models-3-the-iterations) — Part 4 · [The Lessons](2026-08-28-superpowers-local-models-4-the-lessons.md)

The finished skills are here:

> - [Writing Plans skill](https://github.com/lekodeveloperweb/extrapowers/tree/main/engineering/superpowers/writing-plans)
> - [Executing Plans skill](https://github.com/lekodeveloperweb/extrapowers/tree/main/engineering/superpowers/executing-plans)
> - [Brainstorming](https://github.com/lekodeveloperweb/extrapowers/tree/main/engineering/superpowers/brainstorming)

---

## Fork or edit?

I had to choose between patching the upstream Superpowers skills and creating separate versions. Editing had a clear advantage: real verification gates, contradiction checks, and evidence rules would help any model. But the changes small models need amount to a rewrite of the instructions. "Review the plan critically" becomes `IF task-code ≠ interfaces.md: STOP, output the diff, ask`. Those instructions don't fit into files tuned for Claude-class models without undermining what works for them. Small models also can't reliably choose between two similar skills, so the pack name has to make that choice at the config level.

I chose to make new skills with shared mechanical gates and leave upstream untouched. We copied upstream (`superpowers-main-branch/skills`) as the base, then found that the local copy we'd previously modified had diverged. It included `task-structure.md` and `parallel-execution.md`, which upstream didn't have. The new pack lives in the consuming project at `.l-coder/skills/`.

I added one thing beyond the original scope: a thin `brainstorming-small` skill. Brainstorming produces the spec, the root input for the rest of the chain. Its fixable weakness was that success criteria could remain prose instead of becoming testable assertions. I also rejected the option to "just rename the sub-skill references." Upstream's `REQUIRED SUB-SKILL: superpowers:brainstorming:process` doesn't resolve as a loadable skill in a local setup. A small model that reaches an unresolvable REQUIRED instruction won't stop; it will pretend it loaded the skill. Renaming the reference would leave the same problem. The thin variant inlines the essential checklist and drops the visual companion entirely.

## The design principle: self-contained or nothing

The original plan called for a `gates/` directory of shared template files referenced by the skills. We dropped that before writing anything. A weak model can silently skip one more required external file. Each skill now includes every gate it needs. That was a deliberate change to the design.

## The chain: three checkpoints on the same assertions

The main idea is to pass one artifact, the spec's Acceptance Criteria, through three independent phases:

```
brainstorming-small          writing-plans-small              executing-plans-small
spec with ACs        ──▶     each AC becomes a         ──▶    final audit: one row
(INPUT ⇒ EXPECTED)           per-task Verify command          per AC, VERIFIED only
                             with expected stdout             and pasted output
```

The baseline failed because of a correlation problem: the same model wrote the plan, wrote the code, and agreed with itself. The three phases check the same assertions against different artifacts: the spec, the plan, and the executed code. A miss in one phase can be caught by the next.

## Skills v1: what shipped

The pack has three deliberately small files (126 / 116 / 90 lines). Small models skim rather than study.

`brainstorming-small` is the spec factory:
- Three-path classification (spike / bounded/architectural) in a decision table, stated explicitly
- One question per message, preferably multiple choice, until Purpose / Constraints / Success criteria are filled
- Spec template: Overview (3 sentences max) / Design / Acceptance Criteria
- AC rules: one requirement per line, shaped `INPUT ⇒ EXPECTED`; vague criteria ("handles errors appropriately") are forbidden. If a criterion can't be made testable, you don't understand it yet
- Self-review checklist and an explicit user review gate before planning
- Anti-fake-compliance rule: if an instruction refers to something you can't resolve, say so. Never pretend you loaded it

`writing-plans-small` is the plan factory:
- Plan directory: `index.md` (TOC only, no code) / `interfacesmd` (every shared type, signature, constant) / `tasks/` (self-contained, loadable alone)
- The Iron Rule: copy, never re-type. All shared code lives in `interfaces.md`; task files contain tests and instructions only
- Task template: failing test → RED GATE (run it, paste the failure, including the expected failure output) → implement → GREEN → commit
- Test quality rules: assert behavior, not existence. Anything with memory (dedupe, caches) gets a test that calls it twice through the same instance
- Every verify step names the failure mode it catches. A type-checker alone never verifies behavior
- No placeholders such as "TBD" or "add appropriate error handling"
- Self-review uses literal commands, not vibes

`executing-plans-small` enforces evidence:
- Step 0: read the spec's Acceptance Criteria, `interfaces.md`, and `index.md` before any task
- Contradiction gate: task code ≠ `interfaces.md` ⇒ STOP, paste both, ask. Never pick silently
- Evidence rule: after every verify command, paste the last lines of output. A checked checkbox without pasted output is not done
- Coding rules: copy code, don't improve it; every `except: continue` must log; never weaken a failing test
- Stop conditions: verification mismatch after one honest retry; missing or ambiguous instruction; anything you'd have to guess
- Final audit table: one row per spec criterion, with `criterion → file:line → command → pasted output → VERIFIED | NOT-VERIFIED`. If any row is NOT-VERIFIED, don't declare the work done

## The measurement instrument: the scorecard

Before rerunning anything, we turned the failed run review into a testable artifact, `acid-test-scorecard.md`. A retrospective that can't be graded is just a blog post.

- 16 defect rows (D1–D16), one for each class from the baseline review, with the gate expected to catch it and per-run verdict columns
- 6 process checks (P1–P6): missing-AC stop, evidence rule, trap criterion, audit honesty, anti-fake-compliance, stop-and-ask
- A trap criterion: the retrofitted spec deliberately includes a requirement the plan won't implement (`L_CODER_KNOWLEDGE_BUDGET` env var, from the spec's Future Extensions). An honest chain must report it as NOT-VERIFIED and refuse to declare completion. If the model marks it VERIFIED, the audit gate failed, whether or not the code works
- Declared expected misses: two defects can't be reached by specific gates on the old-plan torture test because the plan is internally consistent in its wrongness. The scorecard declares them up front to keep the measurement honest

The original ladder protocol had four steps. First, retrofit Acceptance Criteria onto the failed spec (bounded brainstorming). Second, regenerate the plan with `writing-plans-small`, the highest-signal step. Third, execute the new plan in a pre-feature worktree (`git worktree add ../l-coder-acid-test 52a0322`). This mattered because the real repo already contained the human's fixes, which would contaminate the RED gates. Fourth, run the executor against the old plan as a torture test for the contradiction gate.

The experiment changed as we went. Each iteration regenerated the spec and plan from scratch in a sandbox repo, reviewed them, patched the skills, and reran. We only partially exercised the executing phase. Part 4 covers what we did and didn't validate.

## The Iron Law, applied to documentation

The upstream `writing-skills` skill has an Iron Law: no skill without a failing test first. We applied it directly:

- RED: the baseline run's nine documented defect classes are the failing test
- GREEN: a rerun where the gates catch what they claim to catch
- REFACTOR: when a defect reappears in a new disguise, sharpen the gate

Six iterations later, the pattern wasn't that the model had gotten smarter. Each patch held in every later run, while failures moved down to the next layer: execution bugs, plan coverage, spec coverage, constants, then falsifier grounding. Part 3 follows the iterations run by run.

---

*Continue to Part 3 · [The Iterations](2026-08-31-superpowers-local-models-3-the-iterations)*