# Fine-tuning the Superpowers skills to run on local models: Qwen3.6 35B A3B

## Part 1 of 4: The failure

> **Series index:** Part 1 · The Failure; Part 2 · [The Method](2026-08-29-superpowers-local-models-2-the-method); Part 3 · [The Iterations](2026-08-31-superpowers-local-models-3-the-iterations); Part 4 · [The Lessons](2026-09-01-superpowers-local-models-4-the-lessons.md)

*"Fine-tuning" here means iterating on process documentation rather than model weights. Nobody touched a GPU; we tuned the scaffolding.*

The final skills are here:

> - [Writing Plans skill](https://github.com/lekodeveloperweb/extrapowers/tree/main/engineering/superpowers/writing-plans)
> - [Executing Plans skill](https://github.com/lekodeveloperweb/extrapowers/tree/main/engineering/superpowers/executing-plans)
> - [Brainstorming](https://github.com/lekodeveloperweb/extrapowers/tree/main/engineering/superpowers/brainstorming)
>

---

## The experiment

**l-coder** is a terminal coding agent that runs locally served LLMs on Apple Silicon through an OpenAI-compatible endpoint. Its **knowledge-inject** feature reads Markdown files under `l_coder/skills/knowledge/` and `l_coder/skills/protocols/`, each with TOML frontmatter (`topic`, `token_cost`, `keywords`). On each turn, the agent scores the user's prompt against entry keywords (single word +1.0, phrase +2.0), drops scores below 2.0, fits the remaining entries into a 200-token budget, and adds them to the conversation in an `## Algorithm Reference` block. The feature ships with 17 files: 13 knowledge entries and 4 protocols.

The model was **Qwen3.6 35B A3B**, a Mixture-of-Experts model with roughly 3 billion active parameters per token. It runs locally, quickly, and cheaply. I used **Superpowers** (brainstorming → writing-plans → executing-plans), a skill system originally written for Claude-class models, from a locally vendored copy.

The model's task was to implement knowledge-inject end to end by following those skills.

From the outside, the run looked good. The model announced its skills, wrote a spec, produced a multi-file plan with interfaces and small tasks, executed it, and committed the changes. The test suite passed, and `mypy` and `ruff` were clean.

## The review

I compared the implementation with its spec and found nine classes of defects, all of which had shipped:

### D1: All four protocol files were invalid TOML and silently dropped

Every protocol file's frontmatter contained doubled quotes:

```toml
triggers = [""/cite""]
```

That's not valid TOML. The parser raises `TOMLDecodeError` on line 3. The discovery adapter caught the error and continued:

```python
except (ValueError, OSError):
    continue
```

Result: **zero of four protocol entries ever loaded.** 13 of 17 files worked, and nothing anywhere asserted the count. The doubled quotes came from the migration script in the plan, which applied `repr()` to list items that already contained quotes. A manual "quote normalization" pass followed, but no step parsed the real files.

### D2: Dedupe did nothing, so context grew without bound

The spec required dedupe state to live on `LoopState`:

```python
dedupe_state: Callable[[str], bool] = field(default_factory=make_dedupe)
```

The implementation instead created a fresh closure on every call, inside the injection function:

```python
dedupe = make_dedupe()   # fresh every turn — always accepts
```

The ~200-token knowledge block was appended again on *every* turn, to the latest user message. Its keywords then became part of the next turn's extracted "prompt", so scoring selected the same entries again. Context growth was inevitable.

### D3: The dedup test was vacuous

```python
def test_inject_knowledge_dedup(...):
    entry = KnowledgeEntry(topic="DP", ..., keywords=["dp", "memoize"])
    ...
```

The entry scores 1.0 (one keyword hit), below the threshold of 2.0. **Nothing was injected in either call.** The equality assertion held trivially. The test never asserted that the first injection happened. It verified nothing, and it passed.

### D4: `topic` was optional

The spec said `topic` is required, `ValueError` on missing. The parser shipped `topic = meta.get("topic", "")` and proceeded.

### D5: Scoring semantics drifted

The spec assigned +1.0 to a single-word keyword and +2.0 to a *multi-word keyword phrase*, with word-boundary matching. The implementation matched every keyword as a raw substring (+1.0, so `"log"` matches `"logic"` and `"login"`). It also added an undocumented +2.0 score for a multi-word *topic* phrase.

### D6: Heading collision

The spec called for `### {topic}\n{body}` sections under a `## Algorithm Reference` header. The implementation emitted `## {topic}`, making topics siblings of the section header instead of children.

### D7: A spec contradiction went unresolved

The design doc said "inject as a tail user message"; interfaces.md said "append to the latest user message". These architectures have different failure modes. For example, a bare user turn after tool results breaks strict-alternation templates like Mistral/Devstral. The executor chose one without flagging the disagreement.

### D8: Nothing was observable

No log line on injection (the spec's hook logged `knowledge: +N entries`). Combined with D1's silent skip, the feature's two most important failure modes were invisible.

### D9: The plan duplicated itself

The plan's `index.md` contained its entire Goal / Architecture / Global Constraints section **twice**. For a small model, duplicated but slightly diverging instructions are attention poison.

## The fixes

I fixed all nine defects in the real codebase across two commits: `b3bb588` handled the semantics (frontmatter, turn-scoped dedupe, spec-conformant scoring, required topic, heading levels, logging, and a regression test that parses every shipped skill file and checks for a topic). `c1ffa43` normalized formatting across 20 files with `ruff format`. The suite, now at 429 tests, stayed green throughout.

The review ran into tooling problems too. The Read tool reflows long lines, which made pattern-based edits fail. I switched to byte-exact Python splices. Running `ruff format` also normalized formatting drift in 16 files I had not touched. That is worth accounting for when reviewing AI-generated code professionally.

## Why a model with a 392-test safety net shipped this

The plan files the model wrote for itself explained most of it:

**1. The plan contained the bugs verbatim.** `task-05b-injection-hook.md`, which specified the injection function, put `dedupe = make_dedupe()` inside the function body as "complete code". That was *the exact bug*. `task-01a-domain-model.md` included `topic = meta.get("topic", "")` and the invented topic-phrase scoring. The spec was right, but the plan contradicted it. The executor was told to *"follow plan steps exactly"*. The plan had effectively compiled the bugs in.

**2. The errors were correlated.** The same model interpreted the spec in the plan and then wrote the code. When the executor saw the buggy plan, it found code it would have written itself. Nothing challenged the mistakes. Two passes from one model don't make two reviewers.

**3. The verification gates couldn't catch these bugs.** The injection task used `mypy` as its verification step, but a type-checker can't detect a state-lifetime bug. The migration task used `head -10`, which showed nothing wrong with the invalid TOML. Tests ran against synthetic fixtures, and no test parsed the 17 real files. A gate that can't fail is decoration.

**4. There was no red phase.** TDD appeared as prose in the skill's "Remember" section, not as a gate. The model wrote a test that passed before the implementation existed. It was vacuous and failed to expose the bug it was meant to catch.

**5. Attention economics.** With ~3B active parameters and the spec, interfaces, and task files all in context, the model made *locally plausible* simplifications that strayed from the spec: substring matching instead of word boundaries, a default instead of a raise, and a missing log line. Each choice made sense on its own, but each was a deviation. The duplicated sections in the plan added to the problem.

## The honest scoreboard

| What shipped | What the gates caught |
|---|---|
| 9 defect classes | none |
| 17 data files, 4 broken | no count was ever asserted |
| A test suite that passed | by testing the implementation, not the spec |

Small models can do this, but the workflow's gates were prose, and prose doesn't execute. The skills told the model to "review the plan critically" and "don't skip verifications." It needed commands, templates, and outputs it could use directly.

The next post covers the fix: a `-small` fork of the Superpowers skills, built around mechanical gates, and the measurement instrument we designed before rerunning the experiment.

---

*Continue to Part 2 · [The Method](2026-08-29-superpowers-local-models-2-the-method)*