Local LLM Hardware
Local LLM on NVIDIA RTX 5070 Ti
16GB VRAM · GDDR7 · 896 GB/s memory bandwidth · 300W TDP · released 2025
At a glance
| VRAM | 16 GB |
|---|---|
| Memory | GDDR7 |
| Memory bandwidth | 896 GB/s |
| TDP | 300 W |
| Released | 2025 |
| 2026 price (used / new) | ~$750 |
| Tier | upper-mid consumer 2025 |
| Best for | 16GB with GDDR7 bandwidth, future-proof |
| Reported throughput class | Q4 7B: 40-55, Q4 13B: 25-35 |
Is the NVIDIA RTX 5070 Ti the right card for this?
16 GB is enough for the 7B-to-14B range with room for context, and not enough for 32B at Q4. That single fact decides most of what this card is for. 896 GB/s of GDDR7 sets the ceiling on tokens per second once a model does fit. At about $750 it costs roughly $47 per gigabyte of VRAM and pulls up to 300 W. Catalogue throughput class: Q4 7B: 40-55, Q4 13B: 25-35.
4 other cards here carry 16 GB, so all of them load the same weights and differ only in how fast they read them. At 896 GB/s this is the quickest of that group, so what you gain over the others is tokens per second, not a model you could not otherwise open.
At 2025 this is among the newest entries here, which cuts both ways. The memory technology is current, and the runtime support is the least settled of anything in this guide: check that the backend you intend to use has shipped a release naming this generation before buying for it.
What the price gap buys
The cheapest 16 GB option in this guide is the NVIDIA RTX 4060 Ti 16GB at about $450. This card costs $300 more and reads memory 211% faster, which is roughly the gain in tokens per second. It buys no additional model: both load exactly the same weights, so the question is only whether you are waiting on the output.
At 1 year old this is still a current retail part, so the price above is a shop price rather than a listing price and stock is the usual constraint rather than condition. Buying new also means the 300 W figure is the number your power supply has to meet from day one.
Where the NVIDIA RTX 5070 Ti sits in this guide
Across the 20 cards catalogued here, the NVIDIA RTX 5070 Ti ranks 10th on memory bandwidth at 896 GB/s, 9th cheapest at about $750, and 12th on bandwidth per watt at 3.0 GB/s per watt. Those three positions, not the capacity figure, are what separate it from other 16 GB parts.
GDDR7 is the current generation of board-mounted graphics memory, and it is the reason this card reaches 896 GB/s on a 16 GB bus where the previous generation needed a wider one. For token generation, which reads every weight once per token, that bandwidth translates almost linearly into speed.
This is a CUDA card, which in practice means every local runner in our tool directory supports it first. Ollama and llama.cpp pick it up without configuration; vLLM and TensorRT-LLM target it specifically. Driver and CUDA toolkit versions are the usual source of trouble, not the card.
Which catalogue models fit on NVIDIA RTX 5070 Ti?
Of the 46 models in our catalogue with per-size memory figures, 32 have at least one size that fits in 16 GB at Q4_K_M with room for context. The largest of each are below, biggest first.
| Model | Largest size that fits | Weights at Q4_K_M |
|---|---|---|
| GPT-OSS | 20B | 14 GB |
| StarCoder2 | 15B | 10 GB |
| LLaVA | 13B | 10 GB |
| DeepSeek Coder V2 | 16B | 10 GB |
| Qwen 2.5 | 14B | 9 GB |
| Qwen 2.5-Coder | 14B | 9 GB |
| Qwen 3 | 14B | 9 GB |
| DeepSeek R1 | 14B | 9 GB |
A further 24 smaller models also fit; see the model catalogue.
The next sizes up are out of reach without offloading: Mistral Small 24B at 15 GB, Magistral 24B at 15 GB, Devstral 24B at 15 GB.
Expected tokens per second on NVIDIA RTX 5070 Ti
Memory bandwidth, not compute, is the bottleneck for single-stream token generation. The NVIDIA RTX 5070 Ti's 896 GB/s puts the following ceiling on a batch of one at 2048 tokens of context.
| Model | Quantization | Status | Approx tokens/sec |
|---|---|---|---|
| Llama 3.2 1B | Q4_K_M | Fits in VRAM | ~161 tok/s |
| Llama 3.1 8B | Q4_K_M | Fits in VRAM | ~54 tok/s |
| Llama 3.1 8B | Q8_0 | Fits in VRAM | ~40 tok/s |
| Mistral Nemo 12B | Q4_K_M | Fits in VRAM | ~36 tok/s |
| Qwen 2.5 14B | Q4_K_M | Fits in VRAM | ~31 tok/s |
| Qwen 3 32B | Q4_K_M | Needs CPU offload | ~5 tok/s |
These come from a bandwidth heuristic (tokens/sec ≈ bandwidth in GB/s × a per-model efficiency factor), not from a test rig in our office. Real numbers move with model architecture, batch size and KV-cache size. Treat them as an order of magnitude, and treat the "reported throughput class" row in the table above as the community-measured figure.
Build recommendations for NVIDIA RTX 5070 Ti
Best model to download first
Llama 3.1 8B or Mistral Nemo 12B. Start with ollama pull llama3.1:8b.
Recommended inference backend
Use Ollama for general use or Mullama if you need a drop-in Ollama alternative with native bindings for 6 languages. For production serving on multi-GPU setups, see vLLM.
Sources
[1] nvidia.com/rtx-5070-ti · VRAM figures for catalogue models from src/data/models.json (2026-06-29); card specifications from src/data/gpus.json (2026-06-29).