Local LLM Hardware

Local LLM on AMD Radeon RX 9070 XT

16GB VRAM · GDDR6 · 640 GB/s memory bandwidth · 304W TDP · released 2025

At a glance

VRAM16 GB
MemoryGDDR6
Memory bandwidth640 GB/s
TDP304 W
Released2025
2026 price (used / new)~$600
Tiermid-range AMD 2025
Best fornewest AMD option with better ROCm support
Reported throughput classQ4 7B: 25-35, Q4 13B: 16-22

Is the AMD Radeon RX 9070 XT the right card for this?

16 GB is enough for the 7B-to-14B range with room for context, and not enough for 32B at Q4. That single fact decides most of what this card is for. 640 GB/s of GDDR6 sets the ceiling on tokens per second once a model does fit. At about $600 it costs roughly $38 per gigabyte of VRAM and pulls up to 304 W. Catalogue throughput class: Q4 7B: 25-35, Q4 13B: 16-22.

4 other cards here carry 16 GB, so all of them load the same weights and differ only in how fast they read them. At 640 GB/s this is not the quickest of that group, so its argument is price rather than speed.

At 2025 this is among the newest entries here, which cuts both ways. The memory technology is current, and the runtime support is the least settled of anything in this guide: check that the backend you intend to use has shipped a release naming this generation before buying for it.

What the price gap buys

The cheapest 16 GB option in this guide is the NVIDIA RTX 4060 Ti 16GB at about $450. This card costs $150 more and reads memory 122% faster, which is roughly the gain in tokens per second. It buys no additional model: both load exactly the same weights, so the question is only whether you are waiting on the output.

At 1 year old this is still a current retail part, so the price above is a shop price rather than a listing price and stock is the usual constraint rather than condition. Buying new also means the 304 W figure is the number your power supply has to meet from day one.

Where the AMD Radeon RX 9070 XT sits in this guide

Across the 20 cards catalogued here, the AMD Radeon RX 9070 XT ranks 14th on memory bandwidth at 640 GB/s, 7th cheapest at about $600, and 19th on bandwidth per watt at 2.1 GB/s per watt. Those three positions, not the capacity figure, are what separate it from other 16 GB parts.

GDDR6 is the conservative choice on this card, and at 640 GB/s it sets the ceiling on tokens per second more firmly than the core clock does. When comparing this against a card with the same 16 GB but newer memory, the capacity is identical and the throughput is not.

Running this card means running ROCm, which our catalogue records as the supported path for AMD hardware. Ollama and llama.cpp both have ROCm builds, and llama.cpp additionally has a Vulkan backend that works when ROCm does not. Check that your distribution and kernel are on the supported list before buying: that, rather than the silicon, is where AMD builds usually fail.

Which catalogue models fit on AMD Radeon RX 9070 XT?

Of the 46 models in our catalogue with per-size memory figures, 32 have at least one size that fits in 16 GB at Q4_K_M with room for context. The largest of each are below, biggest first.

ModelLargest size that fitsWeights at Q4_K_M
GPT-OSS 20B 14 GB
StarCoder2 15B 10 GB
LLaVA 13B 10 GB
DeepSeek Coder V2 16B 10 GB
Qwen 2.5 14B 9 GB
Qwen 2.5-Coder 14B 9 GB
Qwen 3 14B 9 GB
DeepSeek R1 14B 9 GB

A further 24 smaller models also fit; see the model catalogue.

The next sizes up are out of reach without offloading: Mistral Small 24B at 15 GB, Magistral 24B at 15 GB, Devstral 24B at 15 GB.

Expected tokens per second on AMD Radeon RX 9070 XT

Memory bandwidth, not compute, is the bottleneck for single-stream token generation. The AMD Radeon RX 9070 XT's 640 GB/s puts the following ceiling on a batch of one at 2048 tokens of context.

ModelQuantizationStatusApprox tokens/sec
Llama 3.2 1B Q4_K_M Fits in VRAM ~115 tok/s
Llama 3.1 8B Q4_K_M Fits in VRAM ~38 tok/s
Llama 3.1 8B Q8_0 Fits in VRAM ~29 tok/s
Mistral Nemo 12B Q4_K_M Fits in VRAM ~26 tok/s
Qwen 2.5 14B Q4_K_M Fits in VRAM ~22 tok/s
Qwen 3 32B Q4_K_M Needs CPU offload ~3 tok/s

These come from a bandwidth heuristic (tokens/sec ≈ bandwidth in GB/s × a per-model efficiency factor), not from a test rig in our office. Real numbers move with model architecture, batch size and KV-cache size. Treat them as an order of magnitude, and treat the "reported throughput class" row in the table above as the community-measured figure.

Build recommendations for AMD Radeon RX 9070 XT

Best model to download first

Llama 3.1 8B or Mistral Nemo 12B. Start with ollama pull llama3.1:8b.

Recommended inference backend

Use Ollama for general use or Mullama if you need a drop-in Ollama alternative with native bindings for 6 languages. For production serving on multi-GPU setups, see vLLM.

Sources

[1] amd.com/rx-9070-xt · VRAM figures for catalogue models from src/data/models.json (2026-06-29); card specifications from src/data/gpus.json (2026-06-29).