Alex Leko
All content on this blog was fully or partially created using local AI (Apple MLX).

Running the Superpowers skills on a local model: Qwen3.6 35B A3B (part 1)

  • llm
  • local-models
  • qwen3
  • skills
  • l-coder

Part 1 of 4: The failure

Series index: Part 1 · The Failure; Part 2 · The Method; Part 3 · The Iterations; Part 4 · The Lessons

"Fine-tuning" here means iterating on process documentation rather than model weights. Nobody touched a GPU; we tuned the scaffolding.

The final skills are here:


The experiment

l-coder is a terminal coding agent that runs locally served LLMs on Apple Silicon through an OpenAI-compatible endpoint. Its knowledge-inject feature reads Markdown files under l_coder/skills/knowledge/ and l_coder/skills/protocols/, each with TOML frontmatter (topic, token_cost, keywords). On each turn, the agent scores the user's prompt against entry keywords (single word +1.0, phrase +2.0), drops scores below 2.0, fits the remaining entries into a 200-token budget, and adds them to the conversation in an ## Algorithm Reference block. The feature ships with 17 files: 13 knowledge entries and 4 protocols.

The model was Qwen3.6 35B A3B, a Mixture-of-Experts model with roughly 3 billion active parameters per token. It runs locally, quickly, and cheaply. I used Superpowers (brainstorming → writing-plans → executing-plans), a skill system originally written for Claude-class models, from a locally vendored copy.

The model's task was to implement knowledge-inject end to end by following those skills.

From the outside, the run looked good. The model announced its skills, wrote a spec, produced a multi-file plan with interfaces and small tasks, executed it, and committed the changes. The test suite passed, and mypy and ruff were clean.

The review

I compared the implementation with its spec and found nine classes of defects, all of which had shipped:

D1: All four protocol files were invalid TOML and silently dropped

Every protocol file's frontmatter contained doubled quotes:

triggers = [""/cite""]

That's not valid TOML. The parser raises TOMLDecodeError on line 3. The discovery adapter caught the error and continued:

except (ValueError, OSError):
    continue

Result: zero of four protocol entries ever loaded. 13 of 17 files worked, and nothing anywhere asserted the count. The doubled quotes came from the migration script in the plan, which applied repr() to list items that already contained quotes. A manual "quote normalization" pass followed, but no step parsed the real files.

D2: Dedupe did nothing, so context grew without bound

The spec required dedupe state to live on LoopState:

dedupe_state: Callable[[str], bool] = field(default_factory=make_dedupe)

The implementation instead created a fresh closure on every call, inside the injection function:

dedupe = make_dedupe()   # fresh every turn — always accepts

The ~200-token knowledge block was appended again on every turn, to the latest user message. Its keywords then became part of the next turn's extracted "prompt", so scoring selected the same entries again. Context growth was inevitable.

D3: The dedup test was vacuous

def test_inject_knowledge_dedup(...):
    entry = KnowledgeEntry(topic="DP", ..., keywords=["dp", "memoize"])
    ...

The entry scores 1.0 (one keyword hit), below the threshold of 2.0. Nothing was injected in either call. The equality assertion held trivially. The test never asserted that the first injection happened. It verified nothing, and it passed.

D4: topic was optional

The spec said topic is required, ValueError on missing. The parser shipped topic = meta.get("topic", "") and proceeded.

D5: Scoring semantics drifted

The spec assigned +1.0 to a single-word keyword and +2.0 to a multi-word keyword phrase, with word-boundary matching. The implementation matched every keyword as a raw substring (+1.0, so "log" matches "logic" and "login"). It also added an undocumented +2.0 score for a multi-word topic phrase.

D6: Heading collision

The spec called for ### {topic}\n{body} sections under a ## Algorithm Reference header. The implementation emitted ## {topic}, making topics siblings of the section header instead of children.

D7: A spec contradiction went unresolved

The design doc said "inject as a tail user message"; interfaces.md said "append to the latest user message". These architectures have different failure modes. For example, a bare user turn after tool results breaks strict-alternation templates like Mistral/Devstral. The executor chose one without flagging the disagreement.

D8: Nothing was observable

No log line on injection (the spec's hook logged knowledge: +N entries). Combined with D1's silent skip, the feature's two most important failure modes were invisible.

D9: The plan duplicated itself

The plan's index.md contained its entire Goal / Architecture / Global Constraints section twice. For a small model, duplicated but slightly diverging instructions are attention poison.

The fixes

I fixed all nine defects in the real codebase across two commits: b3bb588 handled the semantics (frontmatter, turn-scoped dedupe, spec-conformant scoring, required topic, heading levels, logging, and a regression test that parses every shipped skill file and checks for a topic). c1ffa43 normalized formatting across 20 files with ruff format. The suite, now at 429 tests, stayed green throughout.

The review ran into tooling problems too. The Read tool reflows long lines, which made pattern-based edits fail. I switched to byte-exact Python splices. Running ruff format also normalized formatting drift in 16 files I had not touched. That is worth accounting for when reviewing AI-generated code professionally.

Why a model with a 392-test safety net shipped this

The plan files the model wrote for itself explained most of it:

1. The plan contained the bugs verbatim. task-05b-injection-hook.md, which specified the injection function, put dedupe = make_dedupe() inside the function body as "complete code". That was the exact bug. task-01a-domain-model.md included topic = meta.get("topic", "") and the invented topic-phrase scoring. The spec was right, but the plan contradicted it. The executor was told to "follow plan steps exactly". The plan had effectively compiled the bugs in.

2. The errors were correlated. The same model interpreted the spec in the plan and then wrote the code. When the executor saw the buggy plan, it found code it would have written itself. Nothing challenged the mistakes. Two passes from one model don't make two reviewers.

3. The verification gates couldn't catch these bugs. The injection task used mypy as its verification step, but a type-checker can't detect a state-lifetime bug. The migration task used head -10, which showed nothing wrong with the invalid TOML. Tests ran against synthetic fixtures, and no test parsed the 17 real files. A gate that can't fail is decoration.

4. There was no red phase. TDD appeared as prose in the skill's "Remember" section, not as a gate. The model wrote a test that passed before the implementation existed. It was vacuous and failed to expose the bug it was meant to catch.

5. Attention economics. With ~3B active parameters and the spec, interfaces, and task files all in context, the model made locally plausible simplifications that strayed from the spec: substring matching instead of word boundaries, a default instead of a raise, and a missing log line. Each choice made sense on its own, but each was a deviation. The duplicated sections in the plan added to the problem.

The honest scoreboard

What shipped What the gates caught
9 defect classes none
17 data files, 4 broken no count was ever asserted
A test suite that passed by testing the implementation, not the spec

Small models can do this, but the workflow's gates were prose, and prose doesn't execute. The skills told the model to "review the plan critically" and "don't skip verifications." It needed commands, templates, and outputs it could use directly.

The next post covers the fix: a -small fork of the Superpowers skills, built around mechanical gates, and the measurement instrument we designed before rerunning the experiment.


Continue to Part 2 · The Method