(Modified:
2026-10-05)
— Written by
SimeonOnSecurity— 11 min read
A practical October 2026 guide to choosing local AI models for coding agents. Compare GPU memory, KV-cache growth, context size, quantization quality, prompt processing, reasoning settings, and cloud rental costs.
(Modified:
2026-10-03)
— Written by
SimeonOnSecurity— 11 min read
A practical October 2026 guide to running a Qwen 27B model with 32GB of usable accelerator memory. Compare single GPUs, two-card builds, unified memory, used data-center cards, rented compute, software support, and full-context limits.
(Modified:
2026-10-02)
— Written by
SimeonOnSecurity— 8 min read
NVIDIA’s 64GB DGX Spark tier shows how memory supply pressure is shaping local AI hardware. This guide explains the capacity cut, unified-memory tradeoffs, and why DRAM relief is unlikely before late 2027.
(Modified:
2026-10-02)
— Written by
SimeonOnSecurity— 16 min read
Measured Ollama benchmarks of Llama 3.1 8B and Qwen3.8 27B across leading Vast.ai GPUs. Compare decode speed, long-context behavior, rental cost, self-hosting tradeoffs, API credits, and subscriptions.