Running the Superpowers skills on a local model: Qwen3.6 35B A3B (part 2)
Part 2 of 4: The Method
Series index: Part 1 · The Failure — Part 2 · The Method — Part 3 · The Iterations — Part 4 · The Lessons
The finished skills are here:
Fork or edit?
I had to choose between patching the upstream Superpowers skills and creating separate versions. Editing had a clear advantage: real verification gates, contradiction checks, and evidence rules would help any model. But the changes small models need amount to a rewrite of the instructions. "Review the plan critically" becomes IF task-code ≠ interfaces.md: STOP, output the diff, ask. Those instructions don't fit into files tuned for Claude-class models without undermining what works for them. Small models also can't reliably choose between two similar skills, so the pack name has to make that choice at the config level.
I chose to make new skills with shared mechanical gates and leave upstream untouched. We copied upstream (superpowers-main-branch/skills) as the base, then found that the local copy we'd previously modified had diverged. It included task-structure.md and parallel-execution.md, which upstream didn't have. The new pack lives in the consuming project at .l-coder/skills/.
I added one thing beyond the original scope: a thin brainstorming-small skill. Brainstorming produces the spec, the root input for the rest of the chain. Its fixable weakness was that success criteria could remain prose instead of becoming testable assertions. I also rejected the option to "just rename the sub-skill references." Upstream's REQUIRED SUB-SKILL: superpowers:brainstorming:process doesn't resolve as a loadable skill in a local setup. A small model that reaches an unresolvable REQUIRED instruction won't stop; it will pretend it loaded the skill. Renaming the reference would leave the same problem. The thin variant inlines the essential checklist and drops the visual companion entirely.
The design principle: self-contained or nothing
The original plan called for a gates/ directory of shared template files referenced by the skills. We dropped that before writing anything. A weak model can silently skip one more required external file. Each skill now includes every gate it needs. That was a deliberate change to the design.
The chain: three checkpoints on the same assertions
The main idea is to pass one artifact, the spec's Acceptance Criteria, through three independent phases:
brainstorming-small writing-plans-small executing-plans-small
spec with ACs ──▶ each AC becomes a ──▶ final audit: one row
(INPUT ⇒ EXPECTED) per-task Verify command per AC, VERIFIED only
with expected stdout and pasted output
The baseline failed because of a correlation problem: the same model wrote the plan, wrote the code, and agreed with itself. The three phases check the same assertions against different artifacts: the spec, the plan, and the executed code. A miss in one phase can be caught by the next.
Skills v1: what shipped
The pack has three deliberately small files (126 / 116 / 90 lines). Small models skim rather than study.
brainstorming-small is the spec factory:
- Three-path classification (spike / bounded/architectural) in a decision table, stated explicitly
- One question per message, preferably multiple choice, until Purpose / Constraints / Success criteria are filled
- Spec template: Overview (3 sentences max) / Design / Acceptance Criteria
- AC rules: one requirement per line, shaped
INPUT ⇒ EXPECTED; vague criteria ("handles errors appropriately") are forbidden. If a criterion can't be made testable, you don't understand it yet - Self-review checklist and an explicit user review gate before planning
- Anti-fake-compliance rule: if an instruction refers to something you can't resolve, say so. Never pretend you loaded it
writing-plans-small is the plan factory:
- Plan directory:
index.md(TOC only, no code) /interfacesmd(every shared type, signature, constant) /tasks/(self-contained, loadable alone) - The Iron Rule: copy, never re-type. All shared code lives in
interfaces.md; task files contain tests and instructions only - Task template: failing test → RED GATE (run it, paste the failure, including the expected failure output) → implement → GREEN → commit
- Test quality rules: assert behavior, not existence. Anything with memory (dedupe, caches) gets a test that calls it twice through the same instance
- Every verify step names the failure mode it catches. A type-checker alone never verifies behavior
- No placeholders such as "TBD" or "add appropriate error handling"
- Self-review uses literal commands, not vibes
executing-plans-small enforces evidence:
- Step 0: read the spec's Acceptance Criteria,
interfaces.md, andindex.mdbefore any task - Contradiction gate: task code ≠
interfaces.md⇒ STOP, paste both, ask. Never pick silently - Evidence rule: after every verify command, paste the last lines of output. A checked checkbox without pasted output is not done
- Coding rules: copy code, don't improve it; every
except: continuemust log; never weaken a failing test - Stop conditions: verification mismatch after one honest retry; missing or ambiguous instruction; anything you'd have to guess
- Final audit table: one row per spec criterion, with
criterion → file:line → command → pasted output → VERIFIED | NOT-VERIFIED. If any row is NOT-VERIFIED, don't declare the work done
The measurement instrument: the scorecard
Before rerunning anything, we turned the failed run review into a testable artifact, acid-test-scorecard.md. A retrospective that can't be graded is just a blog post.
- 16 defect rows (D1–D16), one for each class from the baseline review, with the gate expected to catch it and per-run verdict columns
- 6 process checks (P1–P6): missing-AC stop, evidence rule, trap criterion, audit honesty, anti-fake-compliance, stop-and-ask
- A trap criterion: the retrofitted spec deliberately includes a requirement the plan won't implement (
L_CODER_KNOWLEDGE_BUDGETenv var, from the spec's Future Extensions). An honest chain must report it as NOT-VERIFIED and refuse to declare completion. If the model marks it VERIFIED, the audit gate failed, whether or not the code works - Declared expected misses: two defects can't be reached by specific gates on the old-plan torture test because the plan is internally consistent in its wrongness. The scorecard declares them up front to keep the measurement honest
The original ladder protocol had four steps. First, retrofit Acceptance Criteria onto the failed spec (bounded brainstorming). Second, regenerate the plan with writing-plans-small, the highest-signal step. Third, execute the new plan in a pre-feature worktree (git worktree add ../l-coder-acid-test 52a0322). This mattered because the real repo already contained the human's fixes, which would contaminate the RED gates. Fourth, run the executor against the old plan as a torture test for the contradiction gate.
The experiment changed as we went. Each iteration regenerated the spec and plan from scratch in a sandbox repo, reviewed them, patched the skills, and reran. We only partially exercised the executing phase. Part 4 covers what we did and didn't validate.
The Iron Law, applied to documentation
The upstream writing-skills skill has an Iron Law: no skill without a failing test first. We applied it directly:
- RED: the baseline run's nine documented defect classes are the failing test
- GREEN: a rerun where the gates catch what they claim to catch
- REFACTOR: when a defect reappears in a new disguise, sharpen the gate
Six iterations later, the pattern wasn't that the model had gotten smarter. Each patch held in every later run, while failures moved down to the next layer: execution bugs, plan coverage, spec coverage, constants, then falsifier grounding. Part 3 follows the iterations run by run.
Continue to Part 3 · The Iterations