Table of Contents

The H100 SXM was the fastest card in both Ollama model tests, but it was not the best value. The RTX PRO 5000 and RTX 6000 Ada delivered a stronger speed-to-rental-cost balance, while the CMP 170HX offered the lowest hourly rate among the completed Qwen3.8 27B runs.

These results measure one model, one Ollama wrapper, one default quantization, and one serving pattern. Treat them as sizing data for this workload, not as a universal GPU ranking.

This article compares Llama 3.1:8b, an Ollama Q4_K_M build at 8.03B parameters, with Qwen3.8:27b, an Ollama Q4_K_M build at roughly 27.3B parameters. Both were tested across GPU instances rented from Vast.ai. The tests focused on decode throughput at actual prompt lengths up to the supported context limit.

The Short Answer

Rent an RTX PRO 5000 or RTX 6000 Ada for regular single-user work. The RTX PRO 5000 led the value comparison at short and medium context lengths. The RTX 6000 Ada cost less and kept a useful lead over the A100 cards. Rent an H100 SXM when response speed matters more than hourly cost or when long prompts are part of the workload.

GoalRecommendationReason
Lowest test costCMP 170HX$0.420 per hour and enough VRAM for the tested Q4_K_M models
Best value under $1.25/hrRTX PRO 5000 or RTX 6000 AdaStrong decode speed without H100 rental cost
Fastest repliesH100 SXMHighest decode speed in both model tests
Long-context Qwen workH100 SXM or RTX PRO 6000 Max-QBoth completed the approximately 131k-token Qwen run with substantial VRAM headroom
Popular 8B baselineLlama 3.1:8bA practical 8B comparison with a 128K context limit
Small recurring workloadAPI credits or a chat subscriptionNo idle GPU charge, maintenance, or deployment work
Private and frequent workloadOwned GPU workstationHardware cost becomes easier to justify with sustained utilization
Abliterated or uncensored experimentsRent first, then buy only after sustained useModel behavior, licensing, and operational risk deserve a separate test before hardware spend

What Was Tested

The benchmark used Ollama with the default Q4_K_M model build. No special quantization, speculative decoding, custom sampler, or serving stack was added. The goal was a simple, locally reproducible comparison of the GPU instances.

Vast.ai describes its service as a GPU cloud marketplace connecting compute providers with users. Its documentation lists on-demand, reserved, interruptible, and serverless options, with search filters for GPU model, memory, price, and availability. The hourly rates below came from the selected instances during this benchmark session, not from a permanent Vast.ai price list.

The measured output was decode throughput, expressed as generated tokens per second. Prefill throughput is listed separately because it describes prompt processing, not the speed at which the answer appears. A long prompt often takes substantial time to process even when decode speed remains high.

The benchmark used two valid trials per context and recorded the median decode result. The completed prompt levels were based on actual prompt_eval_count values:

  • Approximately 16.4k prompt tokens
  • Approximately 32.8k prompt tokens
  • Approximately 65.5k prompt tokens
  • Approximately 131k prompt tokens

The requested 256k-token run was not completed. The 40 GB A100 cards and several smaller cards were still processing or failed to finish at the target context length.

Benchmark limit: The figures describe two Ollama Q4_K_M workloads under one serving pattern. They do not predict training speed, multi-user throughput, image generation speed, vLLM performance, or another model’s result.

Llama 3.1 8B Rerun

Llama 3.1:8b provides a useful practical baseline for the same GPU pool. Ollama lists the model at roughly 120 million downloads, which makes it a reasonable popularity-based choice for a second benchmark. The official Ollama model page lists a 4.9 GB Q4_K_M build, 8.03B parameters, and a 128K context window.

The rerun used a model-specific benchmark runner and calibrated the prompt generator against Ollama’s tokenizer. The labels below use actual prompt counts rather than requested context settings:

  • Approximately 32,676 actual prompt tokens for the 32k target
  • Approximately 65,342 actual prompt tokens for the 64k target
  • Approximately 130,671 actual prompt tokens for the 128k target

The 256k level is explicitly unsupported for Llama 3.1 8B because the model’s listed context limit is 128K. The runner added retry handling and status rows, so a timeout remains visible instead of disappearing from the table.

Llama 3.1 8B Results

GPU~32k decode tok/s~64k decode tok/s~128k decode tok/s
H100 SXM163.5124.082.4
H100 SXM158.6121.280.6
RTX PRO 5000110.179.849.9
CMP 170HX 64GB91.971.050.1
RTX 6000 Ada88.353.929.3
RTX PRO 6000 Max-Q66.753.040.3
RTX PRO 6000 S60.342.533.9
A100 SXM4 40GB2.040.638.5
A100 PCIe 40GB1.220.320.3
2x RTX 5060 Ti1.326.8Timeout

The two H100 rows came from separate valid instances of the same GPU class. Both were clearly ahead of the other cards. Their 32k results were 44% to 48% faster than the RTX PRO 5000, and their 128k results were about 61% to 65% faster.

The A100 and RTX 5060 Ti 32k results are too low to represent comparable GPU execution for this workload. The RTX 5060 Ti host also timed out during the second 128k trial. Treat the host as unsuitable for valid GPU throughput comparisons rather than as a genuine ranking result.

Context targetStatus
32kCompleted on all 10 hosts
64kCompleted on all 10 hosts
128kCompleted on 9 hosts. The 2x RTX 5060 Ti host timed out during its second trial
256kUnsupported for all hosts because Llama 3.1 8B supports up to 128K context

Comparing the Two Models

The smaller Llama model was faster on every comparable GPU, but model size and quality are not interchangeable. Llama 3.1 8B is useful for a popular, lower-memory baseline. Qwen3.8 27B needs more memory and delivers a different capability profile, so the numbers should guide hardware selection rather than declare one model universally better.

Use caseBetter benchmark referenceWhy
Fast 8B assistantLlama 3.1:8bHigher decode speed and lower model memory demand
27B local model sizingQwen3.8:27bMore representative of a larger single-GPU workload
128K context ceilingLlama 3.1:8bThe model has an explicit 128K limit, so 256k is not a valid target
Testing VRAM headroomQwen3.8:27bThe larger model exposes memory and long-context limits sooner
Cross-provider baselineLlama 3.1:8bPopularity and model size make the result easier to reproduce

The earlier Qwen rerun also corrected its actual prompt counts to approximately 32,770, 65,538, and 131,074 tokens on the supported context levels. Those corrected labels replace requested num_ctx values when comparing Qwen runs.

Approximately 16k Tokens

The H100 SXM led the short-context run at 134.4 decode tokens per second. The RTX PRO 5000 followed at 113.9 tokens per second while costing less than half as much per hour.

RankGPUVRAMDecode tok/sPrefill tok/sRental rate
1H100 SXM79.6 GB134.489,974$2.693/hr
2RTX PRO 500047.8 GB113.968,726$1.005/hr
3RTX 6000 Ada48 GB89.195,856$0.813/hr
4RTX PRO 6000 Max-Q95.6 GB84.172,235$1.507/hr
5A100 SXM440 GB60.653,048$0.680/hr
6CMP 170HX64 GB51.551,122$0.420/hr
7RTX PRO 6000 S95.6 GB49.352,773$1.756/hr
8A100 PCIe40 GB38.438,886$0.565/hr

At this prompt length, the RTX PRO 5000 produced about 84.7% of the H100’s decode speed at 37.3% of the hourly rental rate. The RTX 6000 Ada produced about 66.3% of the H100’s speed at 30.2% of the hourly rate.

Approximately 32.8k Tokens

The ranking stayed stable through the medium-context run. The H100 SXM fell to 120.0 tokens per second, while the RTX PRO 5000 reached 105.0 and the RTX 6000 Ada reached 75.7.

RankGPUVRAMDecode tok/sPrefill tok/sRental rate
1H100 SXM79.6 GB120.0139,748$2.693/hr
2RTX PRO 500047.8 GB105.0103,707$1.005/hr
3RTX 6000 Ada48 GB75.7141,678$0.813/hr
4RTX PRO 6000 Max-Q95.6 GB66.9120,537$1.507/hr
5RTX PRO 6000 S95.6 GB65.9109,280$1.756/hr
6A100 SXM440 GB51.181,929$0.680/hr
7CMP 170HX64 GB46.277,984$0.420/hr
8A100 PCIe40 GB36.173,377$0.565/hr

The RTX PRO 5000 remained the most attractive rental in this group. It was only 12.5% slower than the H100 while costing 62.7% less per hour.

Approximately 65.5k Tokens

Longer prompts exposed a wider gap between cards with large memory capacity and cards with less headroom. The H100 stayed first at 112.0 tokens per second. The RTX PRO 6000 Max-Q moved ahead of the RTX 6000 Ada in decode speed.

RankGPUVRAMDecode tok/sPrefill tok/sRental rate
1H100 SXM79.6 GB112.0195,252$2.693/hr
2RTX PRO 500047.8 GB96.9148,897$1.005/hr
3RTX PRO 6000 Max-Q95.6 GB75.7156,443$1.507/hr
4RTX 6000 Ada48 GB70.8208,400$0.813/hr
5A100 SXM440 GB49.3147,699$0.680/hr
6CMP 170HX64 GB45.9118,657$0.420/hr
7RTX PRO 6000 S95.6 GB37.0124,927$1.756/hr
8A100 PCIe40 GBN/AN/A$0.565/hr

The RTX PRO 5000 was still the best balance in the completed table. The RTX 6000 Ada was slower than the RTX PRO 5000, but its lower rental rate made it attractive for budgets focused on availability over peak speed.

Approximately 131k Tokens

The H100 was the only card above 100 decode tokens per second at this context length. The RTX PRO 6000 Max-Q completed the run at 73.4 tokens per second, followed by the RTX 6000 Ada at 61.7.

RankGPUVRAMDecode tok/sPrefill tok/sRental rate
1H100 SXM79.6 GB108.2242,093$2.693/hr
2RTX PRO 6000 Max-Q95.6 GB73.4197,208$1.507/hr
3RTX 6000 Ada48 GB61.7278,283$0.813/hr
4RTX PRO 6000 S95.6 GB39.4124,653$1.756/hr
5RTX PRO 500047.8 GBN/AN/A$1.005/hr
6A100 SXM440 GBN/AN/A$0.680/hr
7CMP 170HX64 GBN/AN/A$0.420/hr
8A100 PCIe40 GBN/AN/A$0.565/hr

The RTX PRO 6000 Max-Q became the practical alternative to the H100 for long prompts. It delivered 67.8% of the H100’s decode speed at 56.0% of the hourly rate. The RTX 6000 Ada cost less, but its 48 GB of VRAM leaves less room for long prompts, larger batches, or another model configuration.

Overall Decode Comparison

GPU~16k~32.8k~65.5k~131k
H100 SXM134.4120.0112.0108.2
RTX PRO 5000113.9105.096.9N/A
RTX 6000 Ada89.175.770.861.7
RTX PRO 6000 Max-Q84.166.975.773.4
RTX PRO 6000 S49.365.937.039.4
A100 SXM460.651.149.3N/A
CMP 170HX51.546.245.9N/A
A100 PCIe38.436.1N/AN/A

The H100 led every completed context length. Its advantage narrowed from 20.5 tokens per second over the RTX PRO 5000 at roughly 16k tokens to 12.0 at roughly 32.8k tokens. The value ranking depends on whether you need the fastest first response or the lowest cost for a long-running session.

Cost per Million Decode Tokens

Hourly price alone hides the useful comparison. Dividing the supplied hourly rate by decode throughput gives a rough rental cost for one million generated tokens. This ignores prompt processing, idle time, storage, data transfer, taxes, discounts, and failed or interrupted runs.

GPUShort-context rateApprox. cost per 1M decode tokens
CMP 170HX$0.420/hr$2.27
RTX PRO 5000$1.005/hr$2.45
RTX 6000 Ada$0.813/hr$2.53
A100 SXM4$0.680/hr$3.12
A100 PCIe$0.565/hr$4.09
RTX PRO 6000 Max-Q$1.507/hr$4.98
H100 SXM$2.693/hr$5.57
RTX PRO 6000 S$1.756/hr$9.89

The CMP 170HX won this narrow arithmetic comparison. The RTX PRO 5000 and RTX 6000 Ada were close behind, and both provided more throughput than the CMP 170HX. For interactive use, the extra waiting time often matters more than the lowest token cost.

Rent, Buy Credits, Subscribe, or Buy Hardware?

The right option depends on utilization and control requirements. A GPU rental is a variable compute bill. API credits are a variable model bill. A chat subscription is a fixed access fee with usage limits. Owned hardware is a capital purchase followed by electricity, cooling, maintenance, and depreciation.

Monthly budget and workloadRecommended pathWhyMain compromise
Under $10Buy API credits or use a free local modelProves the workflow without idle GPU costLess control over model hosting and rate limits
$10 to $30API credits for occasional work, or a chat subscription for daily interactive useA subscription suits human chat, while credits suit scripts and measured callsNeither option gives a private 27B endpoint by default
$30 to $100Rent an RTX 6000 Ada or RTX PRO 5000 only during active sessionsThis budget covers short research bursts and avoids ownership overheadThe instance still needs setup, storage, monitoring, and shutdown discipline
$100 to $250Rent a fast card for scheduled work, or test an owned workstation planSustained monthly usage starts to make utilization visibleRental availability and hourly rates change
$250 to $600Compare a month of RTX PRO 5000 rental against a used or new local GPUThis is the decision point for frequent private inferenceHardware adds power, heat, support, and resale risk
Above $600Buy hardware only with high utilization, or reserve H100-class rental for burst workA mixed strategy avoids owning peak hardware for occasional demandCapital remains tied up in a single configuration

These bands are decision guides, not provider quotes. The benchmark rates show the key break-even issue: an H100 at $2.693 per hour costs about $64.63 for a continuous 24-hour day and about $1,941.00 for 30 days before storage or other charges. The RTX PRO 5000 at $1.005 per hour costs about $24.12 per day and $723.60 for 30 continuous days. A GPU rented only for active sessions has a different total from a GPU left running all month.

When API Credits Win

API credits win when usage is irregular, automation matters, and the task does not require a specific local model. Credits avoid GPU provisioning, model downloads, driver problems, and idle time. They also let you compare several hosted models before committing to hardware.

API pricing differs by model and provider. OpenRouter’s model catalog is designed for comparing model context sizes, providers, and token pricing. Use the provider’s live calculator before comparing an API bill with this benchmark. A local GPU produces tokens without a per-token provider charge, but its rental hour continues while the instance waits.

When a Subscription Wins

A subscription suits a person who uses a hosted assistant every day. Claude’s current individual plans, for example, separate free, Pro, and Max access levels with different usage limits and features. A subscription is not the same as API access. It often gives a polished interface and bundled tools, while API credits give programmatic control.

Choose a subscription when the work is mostly interactive writing, research, coding, or file review. Choose API credits when a script, application, or agent needs a metered endpoint.

When Vast.ai Wins

Vast.ai wins when you need an open model, private runtime control, or a short burst of high throughput. You select the GPU, launch an instance, run Ollama, test a model, and shut the instance down when finished. This is a strong path for Qwen3.8 27B, abliterated variants, and other model files unable to fit a consumer laptop.

The cost is operational work. You need to confirm the model is using the GPU, watch VRAM, protect the exposed service, preserve any required model files, and stop the instance after the run. A low hourly rate does not fix a forgotten instance.

When Buying a GPU Wins

Buying hardware wins after sustained utilization, stable model requirements, and a need for local control. The purchase removes rental availability risk and gives you a persistent environment. It also gives you the option to run private data without sending prompts to a hosted provider.

The purchase price is only one part of the bill. Add the host system, power supply, storage, cooling, electricity, replacement risk, and the value of your setup time. A high-end card also sits idle during periods when API credits or a subscription would cost less.

NVIDIA’s RTX 4090 reference specifications list 24 GB of memory, a 450 W total graphics power rating, and an 850 W minimum system power recommendation. This makes it a useful local-inference reference point, but its 24 GB capacity is below the 27B Q4_K_M configurations tested here. Check actual model memory use rather than assuming a card fits from the parameter count alone.

Abliterated and Uncensored Models

Abliterated and uncensored model variants change the buying decision. Their appeal is often fewer refusal behaviors or fewer built-in restrictions. Their costs include weaker safety behavior, uncertain provenance, inconsistent prompt formatting, higher review needs, and possible license or acceptable-use limits.

For these models, rent before buying. A one-day RTX PRO 5000 test at the supplied rate costs about $24.12 before other charges. This is enough to confirm model fit, GPU use, response quality, and operational risk.

Do not expose an uncensored model directly to the public internet. Put authentication, rate limits, logging controls, and network isolation in front of it. Avoid sending personal, confidential, or regulated data to a model source you have not reviewed.

If the model becomes a daily tool, compare the monthly rental total with the full cost of an owned workstation. If usage is occasional, the rental remains easier to stop and replace.

Benchmark Caveats

The missing results are part of the result. An N/A value means the run did not produce a valid measurement. It does not mean the GPU is slow by the amount implied by another card’s score.

  • The second H100 instance produced no valid benchmark set because the model download completed too late.
  • The 2x RTX 5060 Ti instance was excluded because Ollama loaded the model on the CPU instead of using the GPUs.
  • The Qwen3.8 27B 256k-token prompt was not completed. Its highest valid level was approximately 131k actual prompt tokens. Llama 3.1 8B marks 256k as unsupported because its context limit is 128K.
  • The requested num_ctx was not used as the ranking label. The final tables use actual prompt_eval_count values.
  • Two valid trials per context were summarized with the median decode result. This reduces the effect of one slow run, but it does not replace a larger test set.
  • The wrapper was simple by design. A tuned vLLM or llama.cpp deployment will produce a different ranking.

Final Recommendation

Start with an RTX 6000 Ada or RTX PRO 5000 rental. The RTX 6000 Ada is the budget choice when its $0.813 hourly rate is available. The RTX PRO 5000 is the stronger speed choice below H100 pricing. Use the H100 SXM for long prompts, demanding interactive latency, or a short benchmark window where time matters more than rental cost. For an 8B model, the CMP 170HX is worth testing when its lower hourly rate matters more than peak speed.

Do not buy a GPU after one successful test. Rent for a month of real workloads, record active hours, prompt lengths, concurrency, and model changes, then compare the rental total with the complete ownership bill. Buy hardware when the workload is frequent enough to keep the card busy and stable enough to justify the loss of flexibility.

For occasional work, API credits or a chat subscription are simpler. For private local inference, open model experimentation, or abliterated model testing, Vast.ai provides a lower-commitment path than buying a workstation immediately. Use Llama 3.1 8B as the lower-memory baseline and Qwen3.8 27B when the larger model is the actual deployment target.

References

  1. Vast.ai GPU cloud marketplace
  2. Vast.ai documentation: marketplace, instances, pricing, and rental types
  3. Ollama Qwen3 model library
  4. Ollama Llama 3.1 8B model page
  5. OpenRouter model catalog and pricing comparison
  6. Claude pricing and individual subscription plans
  7. NVIDIA GeForce RTX 4090 specifications
  8. Local AI in 2026: Qwen3.8 27B and self-hosted model hardware
  9. Ollama model testing on Raspberry Pi hardware