Alex Leko
All content on this blog was created fully or partially using local AI (Apple MLX).

Recent Posts

Modern identity, workload security, and agent authentication

This article explains how workload identity and token exchange can reduce reliance on static credentials for AI agents. It covers OAuth 2.0, OIDC, SPIFFE/SPIRE, and identity services from AWS, Microsoft, and Google Cloud, including authentication between agents, tools, and APIs.

Fine-tuning Superpowers skills for local models: Qwen3.6 35B A3B (part 4)

Tracks 21 defect classes found while adapting the Superpowers skills for Qwen3.6 35B A3B, including silent TOML failures, docstring-only stubs, and defects that returned in new forms. Reviews what the model handled well, what remained unresolved, which parts of the workflow were validated, and how to reproduce the runs.

Running the Superpowers skills on a local model: Qwen3.6 35B A3B (part 3)

Across six runs and six skill versions, gates that passed continued to hold as failures shifted from execution bugs to plan and spec coverage, constants, and falsifier grounding. This post tracks the changes and how the model's output eventually helped strengthen its tests.

Running the Superpowers skills on a local model: Qwen3.6 35B A3B (part 2)

This post describes a `-small` skill pack with `brainstorming-small`, `writing-plans-small`, and `executing-plans-small`. Its instructions use explicit checks, including a rule to copy code rather than retype it, a stop when task code conflicts with the interfaces, and a final audit that requires output for each requirement. The scorecard also includes a trap criterion for tasks the model has not implemented.

Running the Superpowers skills on a local model: Qwen3.6 35B A3B (part 1)

An implementation of the Superpowers skills for Qwen3.6 35B A3B passed its tests but had nine defect classes, including invalid TOML, ineffective deduplication, and tests that never exercised injection. This post examines how the plan and verification process let those defects through.

Local LLM benchmarks on an M1 Max: MTP, quantization, and speed

Benchmarks of four model families on an M1 Max MacBook Pro compare OptiQ and oQ4e quantization with several MTP draft settings. Qwen3.6 35B finished faster with oQ4e and Draft 2, while MTP slowed Gemma 4. KAT Coder and Qwen3.6 reached about 53-54 tokens per second; Qwen3.8 took more than six minutes per run.