(Modified:
2026-10-06)
— Written by
SimeonOnSecurity— 10 min read
Compare local AI with ChatGPT and Claude through benchmark scores, memory estimates, serving speed, and agent verification. Includes RTX 5090, DGX Spark, and Mac Studio scenarios.
(Modified:
2026-10-03)
— Written by
SimeonOnSecurity— 11 min read
A practical October 2026 guide to running a Qwen 27B model with 32GB of usable accelerator memory. Compare single GPUs, two-card builds, unified memory, used data-center cards, rented compute, software support, and full-context limits.
(Modified:
2026-10-02)
— Written by
SimeonOnSecurity— 16 min read
Measured Ollama benchmarks of Llama 3.1 8B and Qwen3.8 27B across leading Vast.ai GPUs. Compare decode speed, long-context behavior, rental cost, self-hosting tradeoffs, API credits, and subscriptions.
(Modified:
2026-09-22)
— Written by
SimeonOnSecurity— 16 min read
Qwen3.8 27B and its ternary repack Bonsai 2 27B now outscore Claude Sonnet 4.6 on the aggregate intelligence index while running on one consumer GPU. What changed, what the benchmarks show, and why memory pricing makes waiting the most expensive option.