Table of Contents

Local AI is a practical option for bounded work with clear checks. Replacing the work you give ChatGPT or Claude requires three things to align: model capability, usable memory, and an agent equipped to inspect and verify its output.

A model which fits your workstation still needs suitable tools and enough speed for repeated attempts. This article separates benchmark results, memory estimates, and completed-task evidence so you evaluate each on its own terms.

Key Takeaways

  • Capability: reference scores describe specific evaluation settings, not your local quantized build.
  • Memory: budget for weights, context cache, and runtime overhead together.
  • Speed: replaying an agent conversation measures serving performance, not correctness.
  • Verification: an agent needs checks matched to the requested result.
  • Selection: compare repeated tasks, repair time, and total cost before buying hardware.

Read Scores in Context

The Artificial Analysis Intelligence Index aggregates several evaluations into a reference score. The model leaderboard and local hardware results show the following selected entries as checked on October 6, 2026.

Model and settingIndex scoreDeployment in this comparison
Qwen3.8 27B, xhigh34Downloadable weights
GLM-5.3, max45Downloadable weights
GPT-6 Astra, max53Hosted service
Claude Opus 5.5, max with fallback58Hosted service

Index points are not percentages of intelligence or task success. A 13-point gap between GLM and Claude does not establish a 13% difference in your coding results. Reasoning settings also matter when comparing entries for the same model.

The published evaluation methodology describes the test setup. Some evaluations include tools and agent infrastructure. Treat the score as a result inside those conditions. Do not transfer it directly to a compressed local copy running through another application.

ChatGPT and Claude are products, while this table compares selected underlying models. Your subscription, selected model, available tools, and task context introduce additional differences. Start with a task you repeat, then test the complete setup.

Budget the Whole Request

Memory fit starts with three allocations. Model weights store learned parameters. The key-value cache, or KV cache, retains attention data used during inference. Runtime buffers and other software consume the remaining space.

Required memory = resident weights + context cache + runtime allowance
Available memory = physical capacity - operating system and application reserve

Quantization reduces the storage precision of model weights. Lower precision reduces memory requirements, with quality dependent on the model and quantized build. Cache precision is a separate setting. A Q4 weight file does not establish a four-bit cache.

For a rough weights-only example, 27 billion parameters at 16 bits require about 50.3 GiB. Eight bits gives about 25.1 GiB, and four bits gives about 12.6 GiB. Published file sizes also reflect metadata, packing, and mixed precision, so use the exact files for deployment planning.

The Qwen3.8-27B model card and GLM-5.3 model card describe different architectures. A mixture-of-experts model activates only part of its parameters for each token, but the remaining weights still require storage. Offloading changes placement and latency, rather than eliminating those weights.

Start With Published Files

Download size gives a more concrete starting point than parameter count. Unsloth’s published Qwen and GLM builds illustrate how much the storage requirement changes with the chosen quantization.

Quantized buildPublished size, decimal GBApproximate GiB
Qwen3.8 27B Q4_016.115.0
Qwen3.8 27B Q8_029.027.0
GLM-5.3 UD-Q4_K_XL467434.9
GLM-5.3 UD-IQ2_M239222.6

File sources: Qwen Q4_0 , Qwen Q8_0 , GLM UD-Q4_K_XL shards , and GLM UD-IQ2_M shards . Values are rounded listings checked on October 6, 2026. GiB conversions divide decimal bytes by 2³⁰.

These are weight-file sizes, not total resident-memory measurements. Loading behavior, caches, temporary buffers, and additional model components affect the running process. Do not compare a 29 GB download directly against a 30 GiB allocation without converting units.

Context Changes the Fit

Qwen’s attention cache provides a worked example. Its published configuration specifies 64 layers, full attention every fourth layer, four key-value heads, and a head dimension of 256.

Full-attention KV bytes per token:
16 layers × 4 KV heads × 256 dimensions × 2 (K and V) × 2 bytes
= 65,536 bytes

8,192 tokens  = 0.5 GiB
32,768 tokens = 2.0 GiB

This calculated cache component assumes 16-bit keys and values for one sequence. It excludes linear-attention state, runtime buffers, allocator overhead, and optional vision or speculative-decoding components. Those allocations still need a reserve.

For an illustrative planning budget, add 3–5 GiB for those remaining allocations to the rounded weight sizes. This allowance is an assumption to replace with measurements from your selected runtime.

Build and contextCalculated planning range
Qwen Q4_0, 8K18.5–20.5 GiB
Qwen Q8_0, 8K30.5–32.5 GiB
Qwen Q8_0, 32K32.0–34.0 GiB

A 30 GiB usable allocation leaves substantial room for the Q4 example, while the Q8 scenarios exceed it under these assumptions. A smaller measured reserve changes the boundary. This is why a successful short prompt does not prove sufficient capacity for a long coding session.

Concurrent requests add another dimension. Ollama documents context-memory growth with parallel requests, alongside separate cache-precision controls. Its Q8 cache uses approximately half the F16 cache memory, while Q4 uses approximately one-quarter, with model-dependent quality tradeoffs. See the Ollama runtime FAQ .

Match Hardware to Allocation

An RTX 5090 offers 32 GB of dedicated graphics memory. A DGX Spark with 128 GB uses shared memory. Artificial Analysis documents both configurations in its hardware results. More capacity permits larger allocations, but capacity alone does not establish serving speed.

Hardware scenarioPractical implication
RTX 5090, 32 GBQwen Q4 leaves more context headroom than Q8
DGX Spark, 128 GBRoom for either Qwen build, with runtime overhead still required
Mac Studio, 256 GBLarger weight sets are candidates, subject to runtime and allocation limits

GLM UD-Q4_K_XL exceeds all three capacities before adding a cache. The approximately 222.6 GiB IQ2 weight set also exceeds a 128 GB system. On a 256 GB Mac, its feasibility depends on the actual GPU-accessible allocation and remaining overhead. Apple’s specifications establish hardware options, not a guaranteed inference allocation.

A hypothetical 240 GiB usable budget leaves about 17.4 GiB after those IQ2 weights. A 192 GiB budget fails at the weights alone. Neither budget establishes an operating-system default, compressed-cache support, useful throughput, or acceptable IQ2 quality. Require a demonstrated runtime configuration before purchasing for this workload.

Speed Is a Separate Result

Artificial Analysis’ local inference benchmark replays a recorded workload of 168 model turns. Its laptop and workstation results list the following times for the tested Qwen3.8 27B configurations.

SystemServing replay time
DGX Spark, 128 GB24.2 minutes
RTX 50904.9 minutes
Mac Studio, 256 GBNo result in this comparison

The RTX result takes roughly one-fifth of the Spark time. Serving configurations matter, so this is not a universal hardware ratio. The tested serving builds also differ from the GGUF planning examples above.

Replay timing excludes tool execution and forces responses to recorded lengths. It does not grade whether the generated answers solve the original task. A MacBook measurement also does not establish Mac Studio performance.

Give the Agent Feedback

An agent execution system supplies tools, context, an action loop, and verification. A model writes or chooses actions inside this system. For a CSV export, useful capabilities include reading project files, changing code, running tests, and receiving the resulting errors.

Consider an illustrative export task, rather than a measured experiment. The agent writes a download feature, but its output uses the wrong date format. A visual review misses the problem. A test opens the exported file and compares its dates against the required format. The agent receives the failure, changes the formatter, and reruns the check.

Missing componentLikely failure
Relevant contextEdits an unrelated file
Execution toolsDescribes a fix without applying it
Verification stepStops after plausible code
Error feedbackRepeats a failed approach

LangChain reports a change from 52.8% to 66.5% on Terminal Bench 2.0 while holding GPT-5.2-Codex fixed. Its agent engineering report describes verification guidance, environment context, repeated-edit detection, and reasoning-budget changes.

This is vendor-reported evidence, not proof of a small local model matching a frontier model. It supports a narrower conclusion: model selection alone does not explain agent outcomes. LangChain also reports worse results with maximum reasoning throughout, due to timeouts.

Match Checks to Work

A passing check establishes only what the check covers. A CSV test validating column names leaves date formatting, quoting, Unicode, and access control untested. Define the requested result before deciding how to verify it.

TaskUseful completion evidence
Code changeRelevant tests plus inspection of the resulting behavior
Research answerRetrieved sources supporting individual claims
Booking or refundCorrect stored state and compliance with applicable policy
Document exportParsed output matching the requested fields and formats

The τ-bench study evaluates agents interacting with tools, users, and domain rules. Its research paper checks final database state against expected outcomes. This illustrates why fluent confirmation text is an insufficient completion check for a transaction.

Local Does Not Mean Offline

Local inference controls where the model runs. The surrounding application still determines where documents, search queries, traces, and tool results travel. A local coding model connected to remote search or cloud tools remains a networked system.

Ollama states it does not receive prompts or answers for local execution and documents disabling its cloud features. This is a runtime-specific statement, not a privacy guarantee for every connected agent. Verify both the model endpoint and each integration against the runtime’s documentation .

Data pathWhat to check
Inference endpointLocal process, remote server, or automatic fallback
Search and retrievalQueries and document fragments sent externally
Tool connectionsFiles and records exposed to each service
Logs and tracesStorage location, retained content, and access

A hybrid workflow needs an explicit handoff rule. For example, keep private document extraction local, then send only approved aggregate results for hosted analysis. Review the exact outgoing material. A summary still contains sensitive information if it preserves names, customer details, or confidential findings.

Test Before You Buy

Choose a repeated task with a clear finish condition. Compare the local setup against the hosted product you use, including its tools. Give both the same inputs and acceptance criteria, then repeat the task with several representative examples.

  1. Record the configuration: exact model file, quantization, backend version, context limit, and reasoning setting.
  2. Define success: expected artifact, required behavior, and forbidden side effects.
  3. Measure completion: elapsed time, passed checks, failed attempts, and manual repairs.
  4. Include ownership cost: hardware, electricity, API fees, maintenance, and your time.
  5. Classify failures: reasoning errors, memory limits, latency, missing context, or missing verification.

Local inference suits work whose evaluated quality and latency meet your requirements. A hosted model remains useful when its additional capability reduces failures or review effort. A hybrid workflow assigns different tasks to each after defining which data is allowed to leave the machine.

Count Cost per Accepted Result

Human repair time often changes the economics. A fast local response which needs ten minutes of correction costs more working time than a slower response which passes review. Count accepted results alongside inference expense.

Cost per accepted task =
(hardware allocation + electricity + service fees + maintenance + review time)
÷ accepted tasks

Illustrative review cost: at an assumed $30 per hour, eight minutes of correction costs $4 per task. One hundred such tasks consume $400 of review time. These figures are arithmetic examples, not measured local-model failure rates.

Compare equivalent outcomes. Include rejected attempts in elapsed time and cost. For owned hardware, allocate purchase cost over a realistic service period and task volume. For a hosted service, include the applicable subscription or API expense and the review work it still requires.

For hardware detail, continue with the local model, GPU, and context guide . For capacity and pricing considerations, read the DGX Spark memory comparison . Use your task results to choose the smallest setup which meets your quality, speed, and data requirements.

Reference Video

References