(Modified:
2026-10-03)
— Written by
SimeonOnSecurity— 11 min read
A practical October 2026 guide to running a Qwen 27B model with 32GB of usable accelerator memory. Compare single GPUs, two-card builds, unified memory, used data-center cards, rented compute, software support, and full-context limits.
(Modified:
2026-10-02)
— Written by
SimeonOnSecurity— 8 min read
NVIDIA’s 64GB DGX Spark tier shows how memory supply pressure is shaping local AI hardware. This guide explains the capacity cut, unified-memory tradeoffs, and why DRAM relief is unlikely before late 2027.
(Modified:
2026-09-22)
— Written by
SimeonOnSecurity— 16 min read
Qwen3.8 27B and its ternary repack Bonsai 2 27B now outscore Claude Sonnet 4.6 on the aggregate intelligence index while running on one consumer GPU. What changed, what the benchmarks show, and why memory pricing makes waiting the most expensive option.
(Modified:
2026-09-20)
— Written by
SimeonOnSecurity— 19 min read
Use measured Raspberry Pi 4 and 5 results to choose local language models by memory, speed, and workload, from tiny LFM2.5 models to 3B and larger options.
(Modified:
2026-09-12)
— Written by
SimeonOnSecurity— 7 min read
A practical account of building a local AI assistant for Meshtastic, with private configuration, queued TCP access, retrieval tools, and a safer dashboard.