(Modified:
2026-10-05)
— Written by
SimeonOnSecurity— 11 min read
A practical October 2026 guide to choosing local AI models for coding agents. Compare GPU memory, KV-cache growth, context size, quantization quality, prompt processing, reasoning settings, and cloud rental costs.
(Modified:
2026-10-03)
— Written by
SimeonOnSecurity— 11 min read
A practical October 2026 guide to running a Qwen 27B model with 32GB of usable accelerator memory. Compare single GPUs, two-card builds, unified memory, used data-center cards, rented compute, software support, and full-context limits.
(Modified:
2026-09-22)
— Written by
SimeonOnSecurity— 12 min read
Ollama connects Claude Desktop and Claude Code to local models through an Anthropic-compatible API. Here is the correct setup for each, the environment variable most guides omit, which models to use in September 2026, and which features stay cloud-only.
(Modified:
2026-09-22)
— Written by
SimeonOnSecurity— 16 min read
Qwen3.8 27B and its ternary repack Bonsai 2 27B now outscore Claude Sonnet 4.6 on the aggregate intelligence index while running on one consumer GPU. What changed, what the benchmarks show, and why memory pricing makes waiting the most expensive option.