Local LLM Hardware

Local LLM on AMD Radeon RX 7900 XTX

24GB VRAM · GDDR6 · 960 GB/s memory bandwidth · 355W TDP · released 2022

At a glance

VRAM24 GB
MemoryGDDR6
Memory bandwidth960 GB/s
TDP355 W
Released2022
2026 price (used / new)~$900
Tierhigh-end AMD consumer
Best forAMD alternative to 3090/4090 with ROCm support
Reported throughput classQ4 7B: 30-45, Q4 32B: 10-15

Is the AMD Radeon RX 7900 XTX the right card for this?

24 GB is the capacity most local-LLM tooling is written around, and this card sits in that bracket at 960 GB/s. Released in 2022, it is 4 years old in 2026 and sells for about $900, or roughly $38 per gigabyte of VRAM. It draws up to 355 W, so the power supply and the case airflow are part of the budget, not an afterthought. Catalogue throughput class: Q4 7B: 30-45, Q4 32B: 10-15.

2 other cards here carry 24 GB, so all of them load the same weights and differ only in how fast they read them. At 960 GB/s this is not the quickest of that group, so its argument is price rather than speed.

Released in 2022, this is one of the older entries in the guide, and age here is mostly an advantage: the current generation of local-AI tooling was written while this hardware was already widespread, so backend support is settled and the failure modes are documented. What it cannot have is any of the memory technology that arrived after it, and memory is what sets generation speed.

What the price gap buys

The cheapest 24 GB option in this guide is the NVIDIA RTX 3090 at about $700. This card costs $200 more and reads memory 3% faster, which is roughly the gain in tokens per second. It buys no additional model: both load exactly the same weights, so the question is only whether you are waiting on the output.

Released in 2022, this sits between the shelf and the second-hand market: new stock still turns up, and used examples are common enough that the price above spans both. That makes it worth checking the two markets against each other before committing, because the gap can be wider than the gap to the next card up.

Where the AMD Radeon RX 7900 XTX sits in this guide

Across the 20 cards catalogued here, the AMD Radeon RX 7900 XTX ranks 8th on memory bandwidth at 960 GB/s, 11th cheapest at about $900, and 13th on bandwidth per watt at 2.7 GB/s per watt. Those three positions, not the capacity figure, are what separate it from other 24 GB parts.

GDDR6 is the conservative choice on this card, and at 960 GB/s it sets the ceiling on tokens per second more firmly than the core clock does. When comparing this against a card with the same 24 GB but newer memory, the capacity is identical and the throughput is not.

Running this card means running ROCm, which our catalogue records as the supported path for AMD hardware. Ollama and llama.cpp both have ROCm builds, and llama.cpp additionally has a Vulkan backend that works when ROCm does not. Check that your distribution and kernel are on the supported list before buying: that, rather than the silicon, is where AMD builds usually fail.

Which catalogue models fit on AMD Radeon RX 7900 XTX?

Of the 46 models in our catalogue with per-size memory figures, 38 have at least one size that fits in 24 GB at Q4_K_M with room for context. The largest of each are below, biggest first.

ModelLargest size that fitsWeights at Q4_K_M
Qwen 2.5 32B 20 GB
Qwen 2.5-Coder 32B 20 GB
Qwen 3 32B 20 GB
DeepSeek R1 32B 20 GB
DeepSeek Coder 33B 20 GB
Gemma 4 31b 20 GB
Code Llama 34B 20 GB
QwQ 32B 20 GB

A further 30 smaller models also fit; see the model catalogue.

The next sizes up are out of reach without offloading: LLaVA 34B at 22 GB, Command R 35B at 22 GB, Llama 3.1 70B at 42 GB.

Expected tokens per second on AMD Radeon RX 7900 XTX

Memory bandwidth, not compute, is the bottleneck for single-stream token generation. The AMD Radeon RX 7900 XTX's 960 GB/s puts the following ceiling on a batch of one at 2048 tokens of context.

ModelQuantizationStatusApprox tokens/sec
Llama 3.2 1B Q4_K_M Fits in VRAM ~173 tok/s
Llama 3.1 8B Q4_K_M Fits in VRAM ~58 tok/s
Llama 3.1 8B Q8_0 Fits in VRAM ~43 tok/s
Mistral Nemo 12B Q4_K_M Fits in VRAM ~38 tok/s
Qwen 2.5 14B Q4_K_M Fits in VRAM ~34 tok/s
Qwen 3 32B Q4_K_M Fits in VRAM ~17 tok/s
DeepSeek R1 distilled 32B Q4_K_M Fits in VRAM ~17 tok/s
Llama 3.1 70B Q4_K_M Needs CPU offload ~3 tok/s

These come from a bandwidth heuristic (tokens/sec ≈ bandwidth in GB/s × a per-model efficiency factor), not from a test rig in our office. Real numbers move with model architecture, batch size and KV-cache size. Treat them as an order of magnitude, and treat the "reported throughput class" row in the table above as the community-measured figure.

Build recommendations for AMD Radeon RX 7900 XTX

Best model to download first

Llama 3.1 8B for everyday chat and coding; a 32B at Q4_K_M for top quality within 24 GB. Start with ollama pull llama3.1:8b.

Recommended inference backend

Use Ollama for general use or Mullama if you need a drop-in Ollama alternative with native bindings for 6 languages. For production serving on multi-GPU setups, see vLLM.

Sources

[1] amd.com/rx-7900xtx · [2] ROCm benchmarks · VRAM figures for catalogue models from src/data/models.json (2026-06-29); card specifications from src/data/gpus.json (2026-06-29).