Table of Contents

The best local coding model is the one which keeps enough of your project in memory and reads it fast enough to stay useful. Model size still matters, but a model which fits only after spilling into system RAM often feels worse than a smaller model with a stable context window.

Choose the model and the memory budget together. Do not choose a model from a tier list, then force it onto unsuitable hardware.

This guide focuses on coding agents, not short autocomplete prompts. An agent reads files, tool output, compiler errors, and test results before it writes a fix. Those inputs consume context, and context changes the hardware decision.

The Video Is a Reference

The following video compares local models across several GPU memory tiers. It is useful for its practical model-selection observations and reported hardware results. This article takes a separate path by organizing the decision around memory behavior, agent context, prompt processing, and total ownership cost.

Why Tier Lists Fail

A model tier list usually maps parameter count to a GPU memory number. The shortcut is useful for a first estimate. It stops being useful once a coding agent starts reading a real repository.

The model weights are the first allocation. The KV cache is the growing allocation. It stores attention keys and values for the active context so the runtime does not recompute the entire prompt after every generated token.

Memory consumerWhat it holdsWhy it matters
Model weightsThe quantized parametersA one-time loading requirement
KV cacheNotes from the active contextGrows as the prompt and conversation grow
Runtime buffersTemporary computation spaceVaries by backend and batch size
Agent instructionsSystem prompts and tool definitionsUses context before project files arrive

A model file which fits the card is not proof of a usable agent. The runtime needs room for the cache, temporary buffers, tool definitions, and the next response.

Context Is the Real Budget

Coding agents spend context on more than source files. The budget also includes system instructions, tool schemas, directory listings, shell output, compiler messages, test results, and prior conversation turns.

Three MCP connections with verbose tool definitions often consume several thousand tokens before the agent opens a project file. A small default context then leaves little room for code. The agent still responds, but it loses the working set needed for multi-step repairs.

A Simple Context Record

Before comparing GPUs, record these values:

  • Project size: files and approximate source lines in a normal task.
  • Tool overhead: system prompt, MCP schemas, shell tools, and editor instructions.
  • Error payload: normal compiler and test output length.
  • Target context: the largest prompt you want to keep without truncation.
  • Response allowance: space reserved for the planned patch and explanation.

Use the result as a workload specification. A developer who edits one file at a time has a different memory requirement from a developer who asks an agent to trace an API across a monorepo.

KV Cache Changes the Ranking

Two models with similar parameter counts often have different context costs. A dense model with every layer contributing to the growing cache often needs more memory than a model using a hybrid attention design.

One 27B coding model discussed in the reference material uses a limited set of full attention layers while other layers use a fixed-size summary. The reported 128K context cost stays below 9GB for the cache. A similarly sized model with full growing attention across every layer reportedly needs more than 21GB at the same context length.

These figures describe specific model architectures and runtime settings. Treat them as a reason to inspect the architecture, not as a universal memory formula.

Model behaviorContext effectHardware implication
Full attention at every layerLarge cache growthNeeds more memory for long prompts
Hybrid or sliding attentionLower cache growth in some layersMore context headroom at the same model size
Mixture of expertsFewer active parameters per tokenLower compute per token, but total weights still need storage
Long-context extensionLarger working windowMore cache memory and more prompt-processing work

Read the model architecture before buying memory. Parameter count alone hides the cost of a long coding session.

Choose by GPU Memory Tier

Four to Eight GB

Small dense models remain the practical choice. They fit the card, respond quickly, and work well for autocomplete, short explanations, and narrow file edits.

A larger model with CPU offload might produce text, but response time often becomes the limiting factor. A coding agent needs repeated file reads and tool calls. A four-token-per-second setup turns each repair into a long wait, even when the model technically runs.

Mixture-of-experts models offer another path. A small active portion reduces compute pressure, while the full weight set still lives partly in system memory. This approach rewards a system with ample RAM and fast transfer paths.

WorkloadSuggested direction
AutocompleteSmall dense model with a short prompt
Single-file editsSmall instruct model with tool support
Repository-wide agentRent a larger GPU or increase system memory first
Private code under an NDAUse a local model, but accept a smaller scope or slower runs

At this tier, buy system RAM before chasing a large model. A stable small model beats a large model which spends most of its time moving data across the bus.

Twelve to Sixteen GB

This tier opens the door to a 27B coding model, but quantization choice becomes central. A standard 4-bit build near 17GB does not fit on a 16GB card once runtime overhead enters the picture.

A 3-bit build near 13GB leaves more room for context. A 12GB card pushes the choice toward a 2-bit build or a smaller mixture-of-experts model. Quality depends on the specific quantizer and calibration process, so two uploads with the same bit depth often produce different coding results.

Check the following before downloading a quantized model:

  • Quantizer author and release notes
  • Calibration data and evaluation results
  • Tokenizer compatibility
  • Tool-calling tests
  • Context length at the chosen quantization
  • Runtime support in Ollama, llama.cpp, or the selected front end

Do not treat “2-bit” or “3-bit” as a complete quality description. The packing method and calibration record matter.

Twenty-four to Thirty-two GB

This is the most flexible range for a 27B coding model. A 24GB card often fits a 4-bit build with a useful context, but a full 128K window might exceed the remaining memory. A 32GB card gives the runtime more room for the cache and temporary allocations.

This range also makes ownership easier to justify. One GPU handles the workload without the complexity of a two-card split. The system uses less power than a multi-GPU build, and software support is easier to test.

CapacityPractical position
24GBStrong 4-bit model with context limits to measure
32GB4-bit or 6-bit 27B model with broader context headroom
48GBHigher precision or larger cache, without changing the core model

The 24GB to 32GB range is the sensible ownership zone for frequent private coding work. It avoids the weakest offload behavior without forcing data-center hardware into a desktop case.

Forty-eight GB and Above

More memory does not automatically mean a new model. The same 27B model might run at 8-bit precision on a 48GB card, with a larger cache and fewer compromises. The improvement is consistency, context room, and output quality rather than a new level of reasoning.

At 128GB, the decision changes. A much larger mixture-of-experts model becomes possible, but prompt processing becomes a serious concern. A model often generates quickly while taking a long time to ingest a large repository or a fresh tool result.

For large models, measure prefill speed separately from decode speed. A fast answer after a slow prompt read still feels slow in an agent workflow.

Decode Speed Is Only Half the Test

Decode speed measures generated tokens per second. It answers, “How fast does the model write?” Prefill speed measures prompt processing. It answers, “How fast does the model read?”

An agent spends much of its time reading. Every tool call adds new text. A long source file, stack trace, or test log enters the prompt before the next response begins.

MetricUser experience
Decode tokens per secondHow quickly the answer appears after processing
Prompt tokens per secondHow long the agent waits before the answer starts
Time to first tokenThe combined delay from prompt processing and setup
Context retentionHow much project state remains available during the task

Benchmark your own prompt sizes. A short synthetic prompt hides the cost which matters most during repository work.

Free Runtime Settings First

Hardware upgrades are not the first performance step. Test runtime settings before opening a shopping page.

Multi-Token Prediction

Some model and backend combinations support predicting several future tokens, then verifying them in one pass. The feature is often exposed through a runtime flag or a compatible draft setup.

Reported tests in the reference material show large gains on some high-end cards. Results vary by model file, backend, driver, and card. Apple Metal paths might not preserve the required feature during conversion.

Reasoning Level

A model shipped with maximum reasoning enabled spends more time on internal work before returning a result. Medium reasoning often gives a better balance for code repair, especially when the task already includes a clear error message and a narrow file target.

Use a simple test matrix:

  1. Run the same bug-fix task with low, medium, and high reasoning.
  2. Record time to first token, total time, patch success, and test result.
  3. Repeat with a short prompt and a repository-sized prompt.
  4. Keep the setting which produces the best completed task, not the highest token rate.

Medium reasoning with a compatible multi-token path is a strong starting point. Verify quality on your own code before making it the default.

Local Hardware or Rental GPU?

Rental compute wins for occasional work. You pay for active sessions instead of buying, cooling, updating, and powering a card all year.

Owned hardware wins when the workload is frequent, private, or offline. It also removes queue time and gives you a stable environment for repeatable tests.

SituationBetter first move
A few sessions per monthRent a GPU or use an API
Daily private codingBuy a supported 24GB to 32GB system
Large repository with frequent reloadsRent first and measure prefill speed
No code leaves the buildingOwn the smallest system which meets the context target
Experimenting with a new modelRent before purchasing hardware

Calculate break-even with active hours, not calendar hours. Include electricity, storage, cooling, maintenance, and the time required to keep the runtime working.

A high-end GPU bought for occasional experimentation is a hobby expense. A 24GB to 32GB system used daily for private work has a stronger economic case.

A Better Buying Checklist

Use this order when comparing a model and a GPU:

  1. Define the task. Autocomplete, single-file repair, repository agent, or long-context analysis.
  2. Measure the prompt. Count normal system instructions, tool schemas, files, and test output.
  3. Inspect cache behavior. Look for architecture notes and context-memory measurements.
  4. Choose a quantization. Check quality results from the specific upload, not only its bit count.
  5. Test prefill and decode. Use prompts from your own repository.
  6. Tune reasoning. Compare completed task time at several reasoning levels.
  7. Test privacy and maintenance. Confirm where source code travels and who maintains the backend.
  8. Compare rental cost. Use expected active hours and include power in the owned-system estimate.

Final Recommendation

Under 12GB, run a smaller model or a mixture-of-experts model with enough system RAM. Do not force a 27B dense model into a setup which spends most of its time offloading.

From 12GB to 16GB, focus on quantization quality and a controlled context target. A well-tested 2-bit or 3-bit build with tool support is more useful than a 4-bit file which never fits cleanly.

From 24GB to 32GB, a 27B coding model becomes the practical default for private daily work. Measure context usage and prompt speed before assuming the full advertised window is available.

At 48GB and above, spend the extra memory on precision, cache room, and stable sessions before moving to a larger model. Once the model crosses into 128GB territory, prompt-reading speed and rental economics deserve more attention than raw capacity.

The model name is only the starting point. The useful question is how much project context remains after the model, runtime, tools, and cache take their share.