Archetypes
Local AI PC builds
Historical example configurations, not tested shopping lists. GPU launch MSRP is not a current full-PC quote; model-size targets are illustrative estimates, not measurements or fit guarantees.
Entry 16GB VRAM workstation
Q4–Q5 7B–8B GQA models at 8–16K without offload
1× GeForce RTX 4060 Ti 16GB- Illustrative target (not measured)
- 7B–8B GQA at Q4_K_M / Q5_K_M in VRAM. Phi-4 14B Q4 at 8K is tight; 16K needs offload.
- System RAM
- 32 GB
- PSU
- 650 W
- GPU launch MSRP observation
- $499 on 2023-07-18
Limitations: 70B Q4 will offload. Bandwidth is modest (288 GB/s).
High-bandwidth 16GB build
Higher-bandwidth 7B–8B Q4–Q5 on a single 16 GB Blackwell card
1× GeForce RTX 5080- Illustrative target (not measured)
- 7B–8B Q4–Q5 in VRAM. 32B Q4 weights already exceed 16 GB; 70B requires offload.
- System RAM
- 64 GB
- PSU
- 850 W
- GPU launch MSRP observation
- $999 on 2025-01-07
Limitations: 16 GB cannot hold 32B Q4 or 70B Q4 in-VRAM. MSRP is launch price, not street.
24–32GB high-performance build
32B Q4–Q6 in VRAM at moderate context; 70B only with offload
1× GeForce RTX 5090- Illustrative target (not measured)
- 32B Q4–Q6 in VRAM. 70B Q4_K_M weights (~40 GB) and 32B FP16 (~61 GB) do not fit 32 GB.
- System RAM
- 64 GB
- PSU
- 1000 W
- GPU launch MSRP observation
- $1999 on 2025-01-07
Limitations: 575 W TDP. DeepSeek-V3 671B weights still do not fit.
48GB workstation
Larger dense models and longer context without consumer power spikes
1× RTX 6000 Ada Generation- Illustrative target (not measured)
- 70B Q4_K_M at 8K is tight on 48 GB. 32B Q8 fits; 32B FP16 and 70B Q8 do not.
- System RAM
- 128 GB
- PSU
- 850 W
- GPU launch MSRP observation
- $6800 on 2022-12-06
Limitations: Launch workstation MSRP is not a current street quote.
Multi-GPU workstation
vLLM/SGLang tensor parallel for models that exceed one GPU
2× GeForce RTX 5090- Illustrative target (not measured)
- Two 32 GB cards still cannot hold 70B FP16 (~131 GB). Use Q4/Q5 70B with a serving runtime, or MoE with expert offload.
- System RAM
- 128 GB
- PSU
- 1600 W
- GPU launch MSRP observation
- $3998 on 2025-01-07 (GPU MSRP × count)
Limitations: Consumer GeForce cards are not NVLink. Multi-GPU needs a serving runtime, not Ollama-as-default.