Archetypes

Local AI PC builds

Historical example configurations, not tested shopping lists. GPU launch MSRP is not a current full-PC quote; model-size targets are illustrative estimates, not measurements or fit guarantees.

Catalogue review recorded · individual claims have separate datesSource synthesisMethodology

Entry 16GB VRAM workstation

Q4–Q5 7B–8B GQA models at 8–16K without offload

1× GeForce RTX 4060 Ti 16GB
Illustrative target (not measured)
7B–8B GQA at Q4_K_M / Q5_K_M in VRAM. Phi-4 14B Q4 at 8K is tight; 16K needs offload.
System RAM
32 GB
PSU
650 W
GPU launch MSRP observation
$499 on 2023-07-18

Limitations: 70B Q4 will offload. Bandwidth is modest (288 GB/s).

High-bandwidth 16GB build

Higher-bandwidth 7B–8B Q4–Q5 on a single 16 GB Blackwell card

1× GeForce RTX 5080
Illustrative target (not measured)
7B–8B Q4–Q5 in VRAM. 32B Q4 weights already exceed 16 GB; 70B requires offload.
System RAM
64 GB
PSU
850 W
GPU launch MSRP observation
$999 on 2025-01-07

Limitations: 16 GB cannot hold 32B Q4 or 70B Q4 in-VRAM. MSRP is launch price, not street.

24–32GB high-performance build

32B Q4–Q6 in VRAM at moderate context; 70B only with offload

1× GeForce RTX 5090
Illustrative target (not measured)
32B Q4–Q6 in VRAM. 70B Q4_K_M weights (~40 GB) and 32B FP16 (~61 GB) do not fit 32 GB.
System RAM
64 GB
PSU
1000 W
GPU launch MSRP observation
$1999 on 2025-01-07

Limitations: 575 W TDP. DeepSeek-V3 671B weights still do not fit.

48GB workstation

Larger dense models and longer context without consumer power spikes

1× RTX 6000 Ada Generation
Illustrative target (not measured)
70B Q4_K_M at 8K is tight on 48 GB. 32B Q8 fits; 32B FP16 and 70B Q8 do not.
System RAM
128 GB
PSU
850 W
GPU launch MSRP observation
$6800 on 2022-12-06

Limitations: Launch workstation MSRP is not a current street quote.

Multi-GPU workstation

vLLM/SGLang tensor parallel for models that exceed one GPU

2× GeForce RTX 5090
Illustrative target (not measured)
Two 32 GB cards still cannot hold 70B FP16 (~131 GB). Use Q4/Q5 70B with a serving runtime, or MoE with expert offload.
System RAM
128 GB
PSU
1600 W
GPU launch MSRP observation
$3998 on 2025-01-07 (GPU MSRP × count)

Limitations: Consumer GeForce cards are not NVLink. Multi-GPU needs a serving runtime, not Ollama-as-default.