Table of Contents

32GB on the product label does not mean 32GB available to a Qwen 27B workload. The operating system, runtime, KV cache, model format, driver stack, and card layout all change the result. A cheap card with enough capacity might still lose badly on prompt processing or setup time.

For October 2026, compare usable memory and completed work per dollar, not the VRAM number by itself.

This guide looks at the practical ways to run a Qwen 27B model at 6-bit quality with a serious context window. It separates published specifications from reported test results, because one person’s tokens-per-second result is not a universal property of a GPU.

Why 32GB Is the Target

The memory requirement starts with the weights. A dense 27B model reads the full parameter set during generation. At 6-bit quantization, the weights occupy roughly 20.5GB before runtime allocations and context memory enter the calculation.

A 128K-token context adds another large allocation. The exact KV-cache size changes with the model architecture, cache precision, batch settings, and backend. A practical estimate near 8GB puts the combined weight and cache requirement around 28.5GB before the operating system and inference runtime take their share.

ConfigurationWhat it supports
24GB discrete GPU4-bit weights with reduced context, or partial offload
32GB discrete GPU6-bit weights with a useful 128K context target
Two 16GB GPUsA split model with less headroom for runtime overhead
64GB unified memoryLarger context and model options, subject to the GPU memory allocation limit

The 32GB target is a sizing point, not a guarantee. A model may load while leaving too little room for a long prompt, a larger batch, multimodal inputs, or a second process.

The Video Is One Useful Field Reference

The following video presents a hands-on comparison of the main ways to reach 32GB for local AI. It is useful as a field reference for reported prices, setup friction, and runtime results. This article makes a separate comparison using hardware specifications, model memory behavior, software support, and total ownership cost.

Usable Memory Beats Box Memory

Apple Unified Memory

Apple silicon uses a shared pool for the CPU and GPU. Apple lists 32GB unified-memory Mac mini configurations, but the full pool is not a dedicated graphics allocation. macOS, the application, and the inference runtime all need space.

Some llama.cpp sessions report roughly 22,906MB available on a 32GB system. This is about 21.3GiB, leaving little room for a 6-bit 27B model plus a large KV cache. The exact ceiling depends on the operating system, model runner, memory pressure, and launch settings.

Apple silicon remains attractive for quiet systems. llama.cpp treats Metal as a first-class backend, and Apple avoids the driver and power issues found in many used data-center cards. The tradeoff is lower throughput than a high-end discrete GPU and less predictable memory headroom.

Two 16GB Cards

Two 16GB cards provide 32GB of physical VRAM, but the model runtime still needs per-device buffers and communication space. One card might report 13.4GB in use while the other reports 14.4GB. The remaining capacity is not a clean 4GB reserve for context.

The split method also matters. Layer splitting assigns different model layers to different cards. It expands capacity, but the cards take turns through the model. Tensor splitting places work from the same layer across both cards, which increases communication but lets both devices contribute more directly.

Split methodMain benefitMain cost
Layer splitSimple capacity expansionOne card often waits while the other works
Tensor splitBetter parallel compute in some workloadsMore tuning and more interconnect traffic
CPU plus GPU offloadFits models above total VRAMLarge speed penalty once the CPU handles frequent layers

A two-card plan needs a tested backend before purchase. PCIe generation, slot spacing, power delivery, driver support, and the exact quantization all affect the result.

Single-Card Options

RTX 5090

NVIDIA lists the RTX 5090 with 32GB of GDDR7 and 1,792GB/s of memory bandwidth. It has broad CUDA support, mature tooling, and many tested front ends. The current CUDA GPU table lists the RTX 5090 at compute capability 12.0.

The price is the problem. A new 5090 gives a clean path to 32GB, but the purchase cost competes with months of rented accelerator time. It also draws up to 575W total graphics power, so the system needs a serious power supply and cooling plan.

Radeon AI PRO R9700

The R9700 is a 32GB-class AMD option with high memory bandwidth and a lower entry price than the 5090 in many listings. Its value depends heavily on the runtime. A default llama.cpp path might produce a usable result, while a faster community backend could produce a much higher result on the same card.

This is the clearest example of why a GPU comparison needs two speed columns. Decode speed measures generated tokens. Prefill speed measures how quickly the backend reads the prompt before the first generated token. A 50,000-token codebase exposes a weak prefill path even when decode speed looks acceptable.

Intel Arc Pro B70

The Arc Pro B70 targets professional workloads with 32GB of memory. It is an interesting capacity-per-dollar option, but the software path needs more inspection than CUDA hardware. Standard GGUF support, patched builds, draft-token settings, and backend maturity all affect the result.

A lower purchase price does not compensate for hours spent finding a working build. For a tinkerer, this work is part of the project. For a workstation used every day, driver and application support carry a real value.

Used Data-Center Cards

Tesla V100

Used 32GB Tesla V100 cards attract attention because the card price looks low. The full build costs more. A passive data-center card needs airflow, a suitable chassis or shroud, power adapters, and a host with enough PCIe space.

The other issue is software age. The current CUDA GPU table focuses on newer architectures and does not list the V100 in its current architecture table. A V100 owner needs to check the exact CUDA release, driver branch, backend, and community fork before buying.

Budget for a working V100 system, not a bare card. A card listed near $680 might become a roughly $1,000 project after cooling, adapters, and the host platform. Used prices also rise when a popular local-AI workload sends buyers toward the same retired accelerator.

AMD Instinct MI50

The MI50 offers an attractive used price per gigabyte, yet its current software support is a serious constraint. A card that needs an older ROCm stack or a community-maintained fork is not equivalent to a current consumer card with mainstream application support.

The reported prompt-processing gap is more important than the sticker price. One comparison measured about 128 tokens per second on an MI50 against roughly 2,652 tokens per second on a newer 32GB card for a prompt-processing workload. A 50,000-token codebase would take minutes on the slower path and seconds on the faster path before generation begins.

MI50 fits a narrow use case. It suits a builder who values capacity above interaction speed and accepts driver maintenance. It is a poor choice for rapid codebase analysis when prefill time matters.

Software Is Part of the GPU

llama.cpp supports CUDA, HIP, Metal, Vulkan, SYCL, and other backends. It also supports CPU plus GPU hybrid inference and multi-GPU use. Those features do not produce identical performance across vendors or model formats.

The same physical card often shows large differences between:

  • A default application build with conservative settings.
  • A tuned llama.cpp build with backend-specific kernels.
  • A community fork with speculative or multi-token prediction.
  • A patched model format designed for one accelerator path.

Reported results in the supplied comparison include roughly 27 tokens per second for a default 32GB AMD path, about 46 tokens per second with multi-token prediction, and much higher results on a specialized backend. The result is useful as evidence of software headroom. It is not a promise for a fresh installation.

Record the backend, commit, model file, quantization, context length, prompt length, batch size, and power state with every benchmark. Without those fields, tokens per second is an advertisement number.

Renting Instead of Buying

Renting a fast GPU changes the calculation. A reported RTX 5090 rental price near $0.69 per hour makes a $4,400 purchase equal to about 6,377 rental hours before electricity, host hardware, maintenance, and resale value enter the comparison.

That is close to nine months of continuous rental or about two years at eight hours per day. A two-card desktop purchase near $1,560 equals about 2,261 hours at the same rental rate. The rented 5090 also avoids local heat, noise, driver setup, and card failure.

Usage patternBetter first test
A few hours per weekHosted model or rented GPU
A short projectRent a 5090-class card and measure the real workload
Daily interactive useBuy only after testing the backend and context target
Agents running overnightOwned hardware becomes easier to justify
Private data with steady loadOwned hardware with network isolation

Rent before buying when the model, backend, or context length is still uncertain. One weekend of measured use is cheaper than a wrong workstation purchase.

Buying Recommendations

Two RTX 5060 Ti 16GB Cards

Two 16GB cards provide a practical CUDA path for buyers who need 32GB and want common tutorials, drivers, and front ends. Use tensor splitting only after verifying the model format and backend. Layer splitting is simpler, but it often leaves performance on the table.

This route also needs a motherboard with two usable slots, enough power, and airflow between the cards. A pair of cards is not a small workstation build once the platform cost is included.

One 32GB Card

Choose a single 32GB card when reliability, warranty coverage, and simple deployment matter more than the lowest memory price. The R9700 and RTX 5090 represent different versions of this choice. The AMD path emphasizes capacity and price. The NVIDIA path emphasizes software breadth and mature CUDA support.

A Used V100 or MI50

Choose used data-center hardware only when the project itself includes driver testing, cooling work, and backend maintenance. The card price is the first line of the budget, not the final line.

Apple Silicon

Choose a 32GB Apple silicon system when silence, low idle power, and a compact desktop matter more than maximum throughput. Check the memory allocation reported by the exact runner before assuming a full 32GB is available to the model.

Cloud Rental

Choose rental when the workload is occasional, the model changes often, or a fast prompt response matters more than local ownership. Shut the instance down after the test. Idle time removes the price advantage quickly.

A Better Comparison Worksheet

Before buying, record these values for every candidate:

  • Usable model memory after runtime overhead.
  • Weight format and expected model size.
  • KV-cache size at the target context length.
  • Prefill speed for a representative prompt.
  • Decode speed after the context is full.
  • Backend and driver version required for the result.
  • Idle and loaded power at your local electricity rate.
  • Host cost, cooling, adapters, storage, and warranty.
  • Rental equivalent for the number of hours you expect to use the system.

The most useful test is a prompt from your real workload. For coding, use a representative codebase. For research, use a long document with the normal retrieval or paste workflow. Measure time to first token, prompt processing time, generation speed, memory use, and output quality.

Final Recommendation

Buy two RTX 5060 Ti 16GB cards for the lowest-friction 32GB CUDA build. Use a tensor-split configuration only after a small test confirms the backend works as expected.

Buy a single 32GB card when you need a cleaner workstation. Prefer the card with the stronger software path for your chosen runner, not the card with the lowest dollars-per-gigabyte figure.

Use a V100 or MI50 only when you accept the maintenance project. A cheap accelerator with weak prefill performance is not a bargain for long codebase prompts.

Rent a 5090-class GPU when use is irregular. The rental test gives you real memory, speed, and context data before a large purchase.

The 32GB question is therefore not “Which card has the cheapest gigabyte?” It is “Which system leaves enough usable memory, reads my workload quickly, works with my software, and earns its purchase price through regular use?”

References