The best local-AI machine is the least expensive complete system that runs your exact model, quantization, context, and workload at an acceptable speed—not the box with the largest parameter claim.
Last fully reviewed: August 14, 2026 at 12:31:17 PM PDT
Next scheduled review: August 21, 2026 at 12:31:17 PM PDT
Market basis: United States, USD, complete usable systems before tax; prices and availability accessed August 14, 2026.
Opening verdict
Buy a used RTX 3090 workstation at about $2,400 if 24GB is enough. It remains the strongest entry point because it combines mature CUDA support, 936GB/s of VRAM bandwidth, and replaceable PC parts. Its first pairing in this guide is the newly released Qwen3.8-27B at Unsloth’s downloadable Q4_K_M GGUF, but that recommendation has an unusually important warning: the model and quant arrived on the day of this review, so fit is calculated and no reliable exact RTX 3090 performance benchmark exists yet.
At roughly $4,000, choose capacity over peak speed: the 128GB Framework Desktop can run models that a 24GB GPU cannot, and Framework reports 38 tokens/s for gpt-oss-120b MXFP4. At $4,699, DGX Spark is the compact CUDA appliance to buy when software compatibility and 128GB of coherent memory matter more than decode latency. At $6,699, the 128GB M5 Max MacBook Pro is the best portable large-model system and has the strongest single-box bandwidth/capacity balance here. At $9,449, two Sparks become a defensible but expert-only distributed system.
Do not force a $12,000 purchase. Z.ai announced GLM-5.3 on the day of this review, but said downloadable weights would follow after a two-week safety and hardening period. The proposed three-Spark build therefore has no released model card, license, 3-bit DQ artifact, backend integration, or benchmark to verify. The current downloadable GLM-5.2 MLX four-bit/MXFP4 artifacts also exceed three nodes’ nominal memory before runtime headroom, and three current Sparks cost more than $14,000 once the topology is completed. Above the dual-Spark tier, rent before buying unless a measured workload proves otherwise.
The short answer
Performance in this table belongs only to the exact model and artifact named in the same row. It is not a generic speed rating for the machine.
| Budget | Recommended complete system | Current system cost | Accelerator and usable memory | Bandwidth or interconnect | Best-matched model and exact quantization | Practical tested context | Verified decode speed | Power considerations | Main compromise | Evidence class | Confidence | Last verified |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ~$2.4K | Refurbished Dell Precision 5820, 64GB/1TB/950W + used RTX 3090 | $2,376–$2,406 conditional market estimate | RTX 3090, 24GB GDDR6X | 936GB/s | Qwen3.8-27B, Unsloth Q4_K_M GGUF (15.93GiB) + F16 vision projector (0.86GiB); llama.cpp qwen35 support exists in principle, exact artifact load/stability unverified |
Calculated: 32K with FP16 KV; 64K is an edge case | No reliable benchmark found | 350W GPU; 950W PSU; used-card thermals vary | Same-day model/runtime maturity; no pinned GPU SKU; fit, connectors, and warranty vary | Calculated fit + officially specified hardware | Medium | 2026-08-14 |
| ~$4.0K | Framework Desktop, Ryzen AI Max+ 395, 128GB, 2TB, Linux | $3,978 | Radeon 8060S; 128GB shared. Windows preset allows up to 96GB dedicated VRAM; Linux can configure higher, subject to OS/runtime headroom | 256GB/s shared | gpt-oss-120b MXFP4 in LM Studio |
Vendor did not disclose benchmark context | 38 tok/s (Framework test) | 120W sustained/140W boost processor envelope; 400W PSU | Soldered memory; lower bandwidth; exact backend support varies | Officially specified/vendor benchmark | Medium-high | 2026-08-14 |
| ~$4.7K | NVIDIA DGX Spark, 128GB/4TB | $4,699 | GB10; 128GB coherent unified memory | 273GB/s | Stable fit: Mistral Small 4 119B official NVFP4 (70.847GB). Experimental: June DeepSeek-V4-Flash custom ~81GiB IQ2XXS/Q2 GGUF in DS4 | 2K, 32K, 65K, 131K, 262K tested; peak runtime ~115GiB | 13.2 tok/s at 2K; 11.4 at 65K; 7.94 at 262K; no reliable Mistral result | 140W GB10 TDP; 240W PSU | Official/current DeepSeek-0731 does not fit; custom engine and extreme quant | Independently benchmarked custom artifact + official hardware | Medium-high for exact custom run; low for retained quality | 2026-08-14 |
| ~$6.7K | 14-inch MacBook Pro, M5 Max 40-core GPU, 128GB/2TB | $6,699 | 128GB unified; exact custom run used 114.5–115.1GB | 614GB/s | Measured speed path: June DeepSeek-V4-Flash custom ~81GB IQ2XXS/Q2 in DS4. Current 0731 option: MLX OptiQ 2-bit, 92.487GB/2.43 achieved bpw, SSD-streamed and unbenchmarked on M5 Max | DS4 configured 300K; actual measured prompts were short. Separate preview DQ sweep reached 195K with incomplete harness | 34.1 tok/s mean across three DS4 prompts; ~10% thermal decline | 96W adapter included; GPU reached 98°C in cited sustained test; watts unreported | Fast evidence targets the superseded June preview; current 0731 quality/speed unresolved; no CUDA | Independently benchmarked custom artifact + official hardware | Medium-high for exact preview run; low for current-checkpoint/quality claims | 2026-08-14 |
| ~$9.45K | NVIDIA dual-DGX-Spark bundle with direct cable | $9,449 | 2 × 128GB local pools; not one transparent 256GB pool | 273GB/s per node; 200GbE/RoCE cross-node | Two distinct June/native DeepSeek custom paths: Tony recipe uses TP=2/PP=1, MTP=2, FP8 KV without expert parallel; Elsung uses vLLM 0.21.1, TP=2, expert parallel, MTP=2, FP8 KV. Neither is official NVFP4-on-GB10 | Both set/report up to 500K; Tony omits a complete raw harness | Tony: project-reported ~40 tok/s single stream, ~92 aggregate at c=8. Elsung: 40/38/32 at 8K/200K/500K and ~351 aggregate at 256K/c=32. Artifact lineage remains incompletely sealed | 2 × 140W GB10 TDP; 2 × 240W PSU capacity | Custom mutable images, incomplete benchmark metadata, multi-node failure modes; current 0731 equivalence unresolved | Community-reported demonstrations | Medium for feasibility; low-medium for exact speed | 2026-08-14 |
| ~$12K | Buy nothing by default; keep the dual-Spark tier or rent | $0 incremental | — | — | GLM-5.3 announced, but weights/3-bit DQ/backend/run were not released; GLM-5.2 is the current downloadable fallback | — | — | — | The next defensible hardware tier moved above this budget | Official announcement + verified rejection | High | 2026-08-14 |
Three labels matter in that table. Officially specified means a manufacturer or model publisher states the fact. Calculated means the arithmetic is shown but the pairing was not tested. Community-reported means a public result has useful configuration detail but is not an independent lab result. None of the machines in this review was physically tested by this publication, so nothing is labeled “Verified hands-on.”
How to interpret “can run”
“It loads” is the weakest useful claim. A system can hold the weights and still be miserable because the context cache exhausts memory, prompt processing takes minutes, generation is too slow, the backend lacks optimized kernels, or a distributed link turns every token into a synchronization exercise.
This guide uses these experience bands for batch size 1:
- Comfortable interactive: at least 20 decode tokens/s and acceptable time to first token at the stated context.
- Usable with compromises: 8–20 decode tokens/s, or fast decode paired with noticeable prompt delay.
- Slow but practical for queued work: 2–8 decode tokens/s.
- Fits but is not usefully interactive: below 2 decode tokens/s.
- Experimental: an immature model conversion, backend, hardware port, or topology makes reproducibility the main risk regardless of speed.
- Does not fit: weights plus runtime allocations and the target cache cannot coexist.
- Unverified: the necessary artifact, backend, memory trace, or benchmark evidence is missing.
Prefill and decode are separate. A dual-Spark DeepSeek run can decode around 32 tokens/s at 500K context yet still take about six minutes to process that prompt. For an interactive user, that is a queued workload despite the healthy decode number. Concurrency is separate again: the same dual-Spark result exposes roughly 1.1 million KV-cache tokens shared by all live requests, not 1.1 million per user.
Detailed recommendations by budget
About $2,400: used RTX 3090 workstation — best value, conditional buy
The purchase. The conditional market estimate is a $1,149 refurbished Precision 5820 with Xeon W-2155, 64GB RAM, 1TB NVMe, Windows 11 Pro, and a 950W PSU, plus the current $1,227 used-market RTX 3090 median and a $0–$30 lead/adapter allowance. That totals $2,376–$2,406, but it is not a reproducible pinned-SKU BOM: RTX 3090 AIB cards vary in length, thickness, and use two or three power connectors. NVIDIA specifies 24GB GDDR6X, 936GB/s, and 350W for the card. Confirm the exact card’s clearance and required two or three independent feeds before ordering. See the live refurb listing, used-price index, and NVIDIA specifications.
The model pairing. Qwen3.8-27B is a 27B dense, multimodal, Apache-2.0-licensed open-weight model with 262,144 native context and official 1M YaRN guidance. The exact Unsloth Q4_K_M GGUF is 17.107GB (15.932GiB); the F16 multimodal projector is 0.928GB (0.864GiB). Both were downloadable when checked. Unsloth calls this a Dynamic V3 preview, so it is third-party and same-day, not an official Qwen quant.
The architecture has 16 full-attention layers, four KV heads, head dimension 256, and FP16 K and V. The cache is therefore:
16 layers × 4 KV heads × 256 dimensions × 2 (K + V) × 2 bytes
= 65,536 bytes/token = 64KiB/token
That is 0.5GiB at 8K, 2GiB at 32K, 4GiB at 64K, 8GiB at 128K, and 16GiB at the native 262K limit. Q4_K_M weights + projector + 32K FP16 KV consume about 18.80GiB before graph buffers, allocator overhead, OS/display use, and fragmentation. A 24GB 3090 therefore has a credible 32K target. 64K FP16 is an edge case; quantized KV may make it fit, but it needs a real test. Full 262K does not fit on this card.
Why it wins. CUDA and llama.cpp-family tooling are mature, VRAM bandwidth is high, and the rest of the PC is replaceable. What it cannot do: it cannot hold high-quality frontier-scale weights or Qwen’s full context entirely in VRAM. Planning horizon: two to three years for 20–30B low-bit models, subject to used-hardware health—not a warranty promise. Upgrade path: more system RAM/SSD, then a higher-memory GPU or a new platform; a second 3090 is not a casual addition to this chassis. Setup: medium-to-high. Buy it for low-latency single-user chat, coding, and experimentation. Skip it if you need a warranty-backed appliance, quiet portability, or models above roughly the 24GB class.
Before the return deadline, stress all 24GB of VRAM, log memory-junction and hotspot temperatures, verify the PCIe link, and inspect power connectors. A former mining card with failing rear-side GDDR6X is not a bargain.
About $4,000: Framework Desktop 128GB — best capacity per current complete-system dollar
The current complete Linux configuration is $3,978: $3,449 for the Ryzen AI Max+ 395/128GB system, $505 for Framework’s 2TB SSD, $19 for the required fan, and $5 for the power cord. Framework specifies 128GB LPDDR5X-8000 and lists a 96GB maximum dedicated-VRAM setting for Windows; its official ML page explicitly says Linux can override the VRAM setting higher. Total model-usable memory still remains below 128GB after the OS, runtime, cache, and buffers. AMD specifies 256GB/s shared bandwidth, and the processor runs at a 120W sustained, 140W boost envelope. Memory is soldered. Pricing and configuration are from the official configurator, with machine-learning results and AMD bandwidth specifications.
Framework reports 38 decode tokens/s for OpenAI’s gpt-oss-120b MXFP4 in LM Studio on Fedora 42. This is a vendor benchmark, and its context/prompt length, LM Studio build, flags, and power mode were not disclosed; treat the speed as promising, not lab-grade. The 117B-total/5.1B-active MoE is a better pairing than a dense model of similar stored size because all weights must fit but only a small fraction participates in each token’s expert computation.
Why it wins. It crosses from 24GB to genuinely large shared memory without distributed inference. What it cannot do: 256GB/s is far below the 3090’s dedicated VRAM bandwidth, and 128GB installed is not 128GB free for weights even when Linux configures more than 96GB for graphics. Planning horizon: three to four years for large quantized models if ROCm/Vulkan support continues improving. Upgrade path: SSD, fan, PSU, and case parts are serviceable; memory is not. Setup: medium; Linux and exact runtime selection matter. Buy it for large-model capacity, an x86 host, and quiet compact use. Skip it if a 24GB model covers your needs and latency matters more than capacity, or if your workload requires CUDA-only libraries.
$4,699: DGX Spark — best compact CUDA appliance, not the fastest decoder
NVIDIA’s current listing includes GB10, 128GB coherent unified memory, 4TB NVMe, a 240W PSU, and a power cord for $4,699. The hardware guide specifies 273GB/s memory bandwidth, a 140W GB10 TDP, a 20-core Arm CPU, 10GbE, and two ConnectX-7 ports. NVIDIA’s “1 PFLOP FP4” and “up to 200B parameters” are peak/capacity positioning, not measured LLM throughput.
The conservative pairing is Mistral’s official Mistral Small 4 119B NVFP4, a 70.847GB compressed-tensors artifact derived from the Apache-2.0 base release. The MoE has roughly 6–6.5B routed-active parameters, or about 8B when embeddings/output are counted. The artifact leaves ample nominal capacity for runtime and cache, but no reliable exact GB10 decode benchmark was found.
The frontier experiment is DeepSeek-V4-Flash. The current publisher checkpoint is DeepSeek-V4-Flash-0731: 284B core parameters/13B active, 1M context, an attached speculative module in the packaged checkpoint, and 166.899GB of mixed FP4/FP8 files. It does not fit one 128GB Spark. The best one-Spark evidence instead targets the superseded June preview through a community IQ2XXS/Q2 GGUF of roughly 81GiB and the dedicated DS4 engine—not generic llama.cpp.
An independent exact Spark test pinned DS4 commit a97e7a3, Ubuntu 24.04, CUDA 13.0, driver 580.142, and generated 128 tokens through 65K contexts, plus a 64-token doubling test to 262K. Decode fell from 13.2 tokens/s at 2K to 11.8 at 32K, 11.4 at 65K, 9.95 at 131K, and 7.94 at 262K; peak runtime memory was about 115GiB. Reported prefill was 64.9 tokens/s at 2K, 162.6 at 32K, and 247.1 in the one-shot 65K run; the separate incremental test reported 165.2 at 131K and 109.3 at 262K. Those methods differ, so the values are not one monotonic context curve. No power measurement was supplied. This is an aggressive, non-official conversion with higher-precision attention/shared/output pieces and unknown quality loss. It is not “current DeepSeek at 2.5-bit” in a quality-neutral sense.
Why it wins. It is a compact, warrantied CUDA development appliance with far more memory than consumer NVIDIA cards. What it cannot do: it cannot load the current official DeepSeek-V4-Flash-0731 artifact, and its 273GB/s bandwidth limits single-stream decode. Planning horizon: three to four years as a CUDA development node, conditional on Arm64 container support. Upgrade path: external/network storage and a second Spark; memory is soldered. Setup: low for supported single-node containers, high for custom quants. Buy it for compact CUDA capacity and NVIDIA’s software ecosystem. Skip it if you only run sub-24GB models or expect desktop-GPU latency.
$6,699: 128GB M5 Max MacBook Pro — best portable large-model system
Apple’s exact 14-inch M5 Max 40-core GPU, 128GB, 2TB configuration is $6,699 with its display, battery, keyboard, 96W adapter, and macOS included. Apple specifies 614GB/s memory bandwidth and up to 128GB unified memory in the M5 Max announcement. Memory and storage are not upgradeable.
There are three distinct DeepSeek paths; do not blend them:
- Best exact speed evidence, superseded preview: the same custom ~81GB IQ2XXS/Q2 artifact used on Spark, running in DS4. A detailed M5 Max setup and three-prompt test pinned DS4 commit
560662d, configured a 300K context ceiling, used temperature 0/seed 1, and reported 36.3, 33.6, and 32.5 decode tokens/s—34.1 average—with 114.5–115.1GB runtime memory. A thermal follow-up observed roughly a 10% decode decline as GPU temperature rose from 51°C to 98°C. The actual measured prompts were short; a configured ceiling is not a filled 300K prompt. - Better context sweep, still preview:
mlx-community/DeepSeek-V4-Flash-2bit-DQis a 96.531GB third-party dynamic quant, about 2.72 effective stored bits per 284B parameter, with selected tensors at four, six, or eight bits. The oMLX benchmark reports 38.1 decode/560.7 prefill tokens/s at 4K and 34.1/266.0 at 16K with about 91.4GB peak memory, but the backend audit found its output length, flags, and warm/cold state incompletely disclosed. Use directionally, not as lab-grade evidence. - Current 0731 checkpoint:
DeepSeek-V4-Flash-0731-OptiQ-2bitis 92.487GB, claims 2.43 achieved bits/weight, and streams routed experts from SSD viamlx-optiq. Its producer reports about 2.5 tokens/s on an M3 Max and publishes no capability-retention score. No exact M5 Max result was found.
Why it wins. It pairs high bandwidth and large memory with a complete portable computer, and Metal/MLX has a strong Apple-native inference path. What it cannot do: it cannot run CUDA software; the fast evidence is for a superseded preview and custom engine; the current 0731 quant’s capability loss and M5 Max speed are not established. Full 1M context was not tested. Planning horizon: three to five years for portable local inference, bounded by soldered capacity and evolving MLX support. Upgrade path: none internally; use external storage or replace the machine. Setup: low-to-medium. Buy it if portability and single-box capacity matter. Skip it if CUDA compatibility, repairability, or guaranteed current-checkpoint behavior dominates.
$9,449: two DGX Sparks — best compact frontier experiment
The official bundle contains two 128GB/4TB units and a direct cable for $9,449. The link is Ethernet/RoCE over ConnectX-7 at up to 200Gb/s—not cross-box NVLink. Two boxes provide two local 128GB pools and 8TB aggregate storage; a distributed backend must shard the model. NVIDIA documents the direct topology in its clustering guide.
Two community branches establish feasibility but not an official product profile. The maintained TP2 500K recipe uses a custom arm64 vLLM image, tensor parallelism 2, pipeline parallelism 1, FP8 KV, 500K maximum context, eight sequence slots, MTP draft length 2, and an MP distributed executor. It reports roughly 40 tokens/s single stream and 92 aggregate at concurrency eight, but omits an immutable image digest, exact prompts/output lengths, raw result log, TTFT, prefill, power, and test date.
A second dual-Spark repository reports a roughly 149GB FP8 path with vLLM 0.21.1, tensor parallelism 2, expert parallelism, MTP 2, and FP8 KV. It publishes the more detailed figures: roughly 40 decode tokens/s at 8K, 38 at 200K, and 32 at 500K; TTFT rose from four seconds to roughly six minutes, while a 256K profile reached about 351 aggregate tokens/s at concurrency 32 with a shared ~1.1M-token KV pool. Its authors explicitly call the documentation AI-written and advise caution. Neither repo proves NVIDIA’s 168.305GB NVFP4 artifact is officially supported on GB10; NVIDIA’s card validates B200/GB300-class hardware. Neither result is sealed tightly enough to establish current DeepSeek-V4-Flash-0731 equivalence.
Why it wins. It is the smallest current system with public evidence of interactive DeepSeek-V4-class serving at high precision and meaningful concurrency. What it cannot do: it does not make distributed inference turnkey, and it does not pool memory transparently. Planning horizon: two to four years as an experimental serving cluster; distributed software risk dominates hardware age. Upgrade path: more nodes are theoretically possible, but the cited recipe validates only two. Setup: high; expect RoCE/NCCL, container, kernel, and failure-recovery work. Buy it only if the exact DeepSeek-quality/capacity or multi-user workload justifies the operational burden. Skip it for a personal chat box, first local-AI system, or anything that already fits one GPU.
Around $12,000: stop, measure, and rent
Three current Sparks are not a $12K system. The official two-node bundle plus a third $4,699 unit is already $14,148 before networking. NVIDIA’s direct three-Spark ring requires three QSFP links total, so a bundle that supplies one link still needs two additional approved cables; this guide did not capture a sufficiently evidenced cable price and therefore does not publish a false complete total. No exact GLM runtime was validated on that physical topology. NVIDIA’s RTX PRO 6000 Blackwell Workstation Edition is currently $16,000 for the GPU alone and unavailable on the manufacturer marketplace.
The model premise is premature. GLM-5.3 was officially announced on August 14, but Z.ai said weights would follow after a two-week safety/hardening period. At cutoff there was no downloadable official model card/license, architecture tree, exact 3-bit DQ artifact, or backend run. The current downloadable fallback is GLM-5.2, roughly 744B/40B active by its published description (753.33B including packaged tensors), with a 1M context claim. The checked MLX four-bit and MXFP4 repositories occupy roughly 418GB and 395GB, respectively, before runtime/cache, already beyond three nodes’ nominal 384GB. Unsloth’s existing UD-IQ3_S GGUF is 308.641GB and may pass a raw aggregate-capacity screen, but it is not “3-bit DQ” and has no validated three-Spark loader, context margin, or throughput result. Community pruning does not rescue the recommendation either.
Buy nothing by default. Use the dual-Spark system, rent a larger GPU instance for the frontier model, or wait for a verified 192GB AMD system/next workstation tier. Reopen the capital decision only after a representative week of prompts records required quality, context, concurrency, latency, and utilization.
Model-to-hardware compatibility matrix
“Good” means a downloadable artifact and credible runtime path exist with useful headroom. “Possible” means fit is plausible but the exact combination lacks a strong benchmark. “Experimental” signals conversion/backend risk. “No” means the recommended artifact does not fit or the hardware path is inappropriate. Context entries are publisher maxima, not promises of practical local use.
License language is deliberate. Open weights means the weights can be downloaded under stated terms; it does not prove that training data, code, or the full pipeline is open. Apache-2.0 and MIT weight releases are identified as permissively licensed open weights. Source-available is used for public weights/code under custom or restrictive terms. MiniMax-M2.7 is non-commercial without separate authorization; NVIDIA Nemotron and Kimi K3 use custom agreements. None is flattened into the blanket label “open source.”
| Model | Type; total / active parameters | Publisher context | Weight-license status | Practical local artifact | RTX 3090 24GB | Framework 128GB | 1× Spark | M5 Max 128GB | 2× Spark |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.8-27B | Dense multimodal; 27.78B parsed | 262K native; 1M YaRN | Apache-2.0 open weights | Unsloth Q4_K_M GGUF 17.107GB + 0.928GB projector |
Possible/in principle at calculated ~32K; llama.cpp qwen35 code exists, but no exact same-day load or benchmark |
Possible; exact backend speed unverified | Possible; exact speed unverified | Possible via GGUF/Metal; no exact run | Wasteful |
| Qwen3.6-27B | Dense multimodal; 27.78B parsed | 262K native; ~1.01M extended | Apache-2.0 open weights | ggml-org Q4_K_M 19.1GB; MLX 4-bit 16.1GB |
Good; detailed 3090 predecessor run exists | Good | Good | Good | Wasteful |
| Qwen3.6-35B-A3B | MoE multimodal; 35.95B / ~3B | 262K native; ~1.01M extended | Apache-2.0 open weights | ggml-org Q4_K_M 20.420GB + optional 1.060GB MTP; MLX 20.429GB |
Fits, but cache margin tight | Good | Good | Good | Wasteful |
| DeepSeek-V4-Flash-0731 | MoE reasoning/coding; 284B core / 13B active; packaged checkpoint parses 304.18B with speculative module | 1M | MIT open weights | Official mixed tree 166.899GB; MLX OptiQ 92.487GB/2.43 bpw with SSD expert streaming | No | No verified Framework/Linux path; OptiQ is Apple-specific | No | Experimental; exact M5 Max speed/quality absent | Possible capacity, but no sealed 0731 TP2 run |
| DeepSeek-V4-Flash (June preview) | MoE reasoning/coding; 284B / 13B | 1M | MIT open weights; superseded checkpoint | NVIDIA NVFP4 168.305GB; custom ~81GiB DS4 GGUF; MLX DQ 96.531GB | No | No verified DS4/Vulkan path | Experimental; independently tested custom DS4 | Experimental; independently tested DS4 and community oMLX paths | Community TP2 demonstrated; not official NVFP4-on-GB10 |
| GLM-4.7-Flash | MoE assistant/coding; 31.22B / ~3B | 202,752 | MIT open weights | Unsloth Q4_K_M 18.3GB; MLX 4-bit 16.9GB |
Good | Good | Good | Good | Wasteful |
| GLM-5.2 | Frontier MoE; ~744B / ~40B (753.33B packaged) | 1M | MIT open weights | Unsloth UD-IQ3_S 308.641GB; patched MLX 2.56-bpw ~246.9GB |
No | No | No | No | No; still exceeds two-node capacity/headroom |
gpt-oss-20b |
MoE reasoning/tool use; 21B / 3.6B | 131K | Apache-2.0 open weights | Official MXFP4 | Good | Good; vendor reports 58 tok/s | Good | Good | Wasteful |
gpt-oss-120b |
MoE reasoning/tool use; 117B / 5.1B | 131K | Apache-2.0 open weights | Official MXFP4 | No | Good; vendor reports 38 tok/s | Good; llama.cpp reports 42.76 tok/s at 32K | Good | Unnecessary |
| Gemma 4 26B-A4B-it | MoE multimodal; 25.2B / 3.8B | 256K | Apache-2.0 open weights | Google QAT Q4_0 GGUF 14.439GB; 15.634GB with projector | Good | Good | Good | Good | Wasteful |
| Mistral Small 4 119B | MoE multimodal/coding; 119.4B / ~6–8B | 256K | Apache-2.0 open weights | Official NVFP4 70.847GB; Unsloth Q4_K_M 73.763GB | No | Possible via GGUF/Vulkan; exact run absent | Good NVFP4 fit; exact speed unverified | Possible via GGUF/Metal; exact run absent | Good but unnecessary |
| Kimi-Linear-48B-A3B-Instruct | Hybrid linear-attention MoE; 48B / 3B | 1M | MIT open weights | MLX 4-bit 27.644GB; community AWQ 4-bit | No | Good via suitable backend | Possible; exact run absent | Good MLX fit | Unnecessary |
| MiniMax-M2.7 | MoE; 228.69B; active total not disclosed for this release | 204,800 | Non-commercial/source-available; commercial authorization required | MLX 3-bit 100.103GB; community GGUF | No | Experimental; MLX artifact is Apple-specific and a GGUF/Vulkan path needs separate headroom testing | Experimental community low-bit only | Possible MLX fit; near memory edge | Possible; license and quality caveats |
| Nemotron 3 Nano Omni 30B-A3B Reasoning | Hybrid multimodal MoE; 31B / ~3B | 256K | Custom NVIDIA Open Model Agreement; not Apache/MIT/OSI | Official NVFP4 22.432GB | Fits by bytes but no optimized Ampere NVFP4 path verified | Possible via conversion | Good NVFP4 fit | Possible only after conversion | Wasteful |
| IBM Granite 4.0 H Small | Hybrid Mamba/attention MoE; 32B / 9B | 128K | Apache-2.0 open weights | IBM official Q4_K_M GGUF 19.477GB |
Good | Good | Good | Good | Wasteful |
| Kimi K3 | Frontier multimodal MoE; 2.8T / 104B | 1M | Custom Kimi K3 license; open-weight/source-available, not OSI open source | Official QAT MXFP4/MXFP8 repository ~1.561TB | No | No | No | No | No |
The matrix intentionally includes some “No” rows. A current-model guide should show where local ownership stops being rational, not imply that every headline release belongs on a desk.
Memory and quantization methodology
Start with the artifact, not a slogan. The rough equation
raw weight memory = parameter count × effective bits per weight ÷ 8
is only a screening tool. Actual files also contain scales, zero points, metadata, tensors deliberately kept at higher precision, embeddings, vision projectors, and sometimes an MTP draft head. A “2-bit” filename therefore need not occupy two bits per published parameter. The preview-based DeepSeek 96.531GB MLX artifact divided by 284B parameters is roughly 2.72 bits/parameter; the custom ~81GiB GGUF is roughly 2.44. The current 0731 OptiQ card reports 2.43 achieved bits/weight, but streams experts from SSD. None of those figures measures quality.
Then add, separately:
- runtime graph, kernel, allocator, and temporary buffers;
- KV cache at the actual cache precision and architecture;
- target context and batch/concurrency;
- OS, display, and shared-memory reservations;
- fragmentation and a safety margin;
- a multimodal projector and image-token cache where relevant.
For MoE, total stored parameters determine capacity; active parameters per token influence compute. gpt-oss-120b still needs space for roughly 117B parameters even though about 5.1B are active per token. Conversely, a 27B dense model touches its full dense stack every token. Active count is not a license to discard the inactive experts.
Formats are not interchangeable:
- GGUF is the dominant llama.cpp-family container for CPU, CUDA, Metal, and Vulkan offload.
Q4_K_Mis a mixed four-bit family, not exactly four bits for every tensor. - MLX artifacts target Apple Silicon’s unified-memory/Metal stack. A conversion being small enough does not make it runnable elsewhere.
- AWQ, GPTQ, and EXL2 are weight-only families used mainly on discrete GPUs; kernel/model support and calibration differ.
- NVFP4 is NVIDIA’s four-bit floating-point format with scaling metadata and hardware-specific fast paths. An NVFP4 repository is not evidence that a pre-Blackwell GPU will execute optimized NVFP4 kernels.
- DQ means dynamic or mixed quantization in the cited community artifacts; inspect the tensor map rather than reading the filename as a uniform bit rate.
Quality evidence is the weakest part of this market. None of the aggressive DeepSeek community quants in the recommendations has a strong public task-level comparison against the publisher checkpoint. Treat speed and fit as proven only where stated; treat retained intelligence as unresolved.
NVIDIA versus Apple Silicon versus compact clustered systems
| Platform | What it does best | What the spec hides | Best fit in this guide |
|---|---|---|---|
| RTX 3090 / discrete NVIDIA | Low-latency decode for models that fit; broad CUDA and llama.cpp support; replaceable hardware | VRAM is separate from host RAM; 24GB is a hard capacity wall for full GPU residence; used thermals/warranty | 9B–27B low-bit models, coding/chat, fastest value tier |
| Ryzen AI Max / shared x86 APU | Large single-address-space capacity at ~$4K; compact, serviceable storage/platform | CPU and GPU share 256GB/s; Windows offers a 96GB dedicated-VRAM maximum preset while Linux can configure higher; total usable remains below 128GB; backend maturity varies | 70B–120B-class low-bit MoE/dense capacity when CUDA is not required |
| DGX Spark / GB10 | Compact 128GB CUDA appliance; supported NVIDIA OS/container path; built-in high-speed NICs | 273GB/s constrains decode; Arm64 compatibility; “1 PFLOP” says little about tokens/s | Large official NVFP4 models and CUDA development; custom frontier experiments |
| M5 Max / Apple unified memory | 614GB/s, 128GB, battery/display, strong MLX/Metal path | No CUDA, all core components soldered, OS uses the same pool | Portable large-model inference and exact MLX artifacts |
| Dual Spark | Distributed capacity, high aggregate throughput, direct 200GbE/RoCE | Two local pools, not one; link is 25GB/s wire rate before overhead; software complexity dominates | One measured TP=2 workload that cannot run acceptably on one node |
No platform wins every axis. Apple provides the best portable capacity/bandwidth balance; a used 3090 provides the best low-budget decode ecosystem; Framework offers the cheapest clean 128GB capacity; Spark buys NVIDIA software compatibility at a bandwidth penalty; multiple Sparks buy a distributed systems project.
Multi-GPU and multi-node realities
Two accelerators do not automatically provide twice the usable memory or performance. The runtime must implement tensor, pipeline, expert, or layer parallelism for the exact model and quant. Every token can require cross-device collectives, so interconnect latency and bandwidth matter alongside local memory bandwidth.
For two Sparks:
- each node has 128GB local coherent memory and 273GB/s local bandwidth;
- the direct ConnectX-7 path is up to 200Gb/s Ethernet/RoCE, or 25GB/s wire rate per direction before overhead;
- the official bundle supplies the cable, but the model server still needs RoCE/NCCL and a compatible container;
- the Tony DeepSeek recipe uses TP2/PP1 without expert parallelism; the separate Elsung setup uses TP2 plus expert parallelism; neither creates magical unified pooling;
- a failure takes down the sharded service unless the operator builds a restart/fallback plan;
- 2 × 240W PSU capacity is not the same as 480W measured wall draw.
The public dual-Spark result is useful because it reports topology, runtime, context, concurrency, and long-context TTFT. It still remains community evidence with a custom image and a previously diagnosed stability problem. The repository says a prefix-cache host-memory leak was fixed in its production-v2 image and reports a subsequent stable run; an editorial lab should reproduce that before calling the system production-ready.
Ordinary multi-GPU PCs have similar caveats. PCIe peer-to-peer support, slot spacing, bifurcation, NUMA placement, power delivery, and cooling can matter more than aggregate VRAM on a shopping list. NVLink exists only on specific older/professional cards and does not, by itself, create a flat model-memory pool.
Power, noise, portability, and ownership cost
The published power figures are not directly comparable: 350W is a GPU board rating, 140W is the GB10 TDP, and 96W is an included laptop adapter. None is a measured whole-system inference draw.
- Used 3090 workstation: potentially the loudest and hottest option; the 350W card sits in an older workstation with a 950W PSU. Electricity, fan noise, and used GDDR6X health belong in the purchase decision.
- Framework Desktop: 120W sustained/140W boost processor envelope in a 4.5L case, with a 400W PSU. Compact and serviceable, but the small PSU fan and shared-memory load need an independent noise test.
- DGX Spark: 140W GB10 TDP and 240W PSU in a very small enclosure. One box is desk-friendly; two add network, cabling, and service complexity.
- M5 Max MacBook Pro: the only complete portable system. Apple includes a 96W adapter for the 14-inch SKU, but long inference can consume battery quickly and the smaller chassis may throttle differently from the 16-inch model. No comparable wall-power/noise test was found.
Storage matters. Model collections can consume terabytes; Spark includes 4TB, while the comparison Mac and Framework include 2TB and the refurb includes 1TB. Back up model manifests, adapters, prompts, and environment files; cached weights are replaceable, experimental conversions may not be.
Ownership cost also includes failures and time. A week spent stabilizing a two-node container can cost more than months of cloud inference. Used hardware needs a return window and stress testing. Soldered-memory machines need enough capacity on day one because there is no DIMM upgrade later.
Recommendations by use case
- Best-value private chat and coding: used RTX 3090 workstation. Target 9B–27B low-bit models and 8K–32K contexts. Choose Qwen3.8 only if you accept same-day software maturity; keep a known-stable model installed.
- Largest clean model capacity near $4K: Framework Desktop 128GB. Start with
gpt-oss-120bMXFP4 because an exact vendor result exists; validate context and power yourself. - CUDA development with more than 24GB: DGX Spark. Prefer official NVFP4 artifacts such as Mistral Small 4 over extreme community quants for production-like evaluation.
- Portable research and large local models: M5 Max 128GB. MLX/oMLX makes the strongest case; the DeepSeek 2bit-DQ result is fast but requires a quality evaluation.
- Multi-user frontier serving in a lab: dual Spark only when the exact vLLM/DeepSeek path is the workload. At 256K, concurrency 32 achieved high aggregate throughput, but each live request cannot simultaneously consume the full context.
- Long-document single-user analysis: favor enough memory first, then examine prefill. A 500K context that takes six minutes before generation is a batch job, not chat.
- Fine-tuning: none of the inference recommendations implies comfortable full-model training. LoRA/QLoRA feasibility depends on model, optimizer state, activation memory, sequence length, and framework support; buy against a measured training recipe.
- CPU-only: use it for small models, testing, or overnight batch jobs on hardware already owned. Buying a new CPU-only system for these model classes is poor value.
What not to buy
- A “$2,000 RTX 3090 PC” without a pinned card. Current used-card and host pricing puts the conditional market estimate around $2.4K, but exact AIB dimensions, connectors, condition, and warranty determine whether it is actually complete.
- An RTX 5090 for capacity. It is extremely fast when 32GB is enough, but current market pricing is around the cost of a 128GB Spark. Buy it for speed, gaming, or training kernels—not because it solves large-model fit.
- A DGX Spark on the “1 PFLOP” line alone. Sparse FP4 peak compute is not decode speed. Small dense models can be slower than on a used 3090 because Spark has only 273GB/s memory bandwidth.
- One Spark for the current DeepSeek-V4-Flash-0731 checkpoint. Its official tree is 166.899GB; one node has 128GB shared with the OS/runtime. Only much more aggressive third-party variants of the older preview have a measured one-Spark path.
- Two Sparks described as “256GB unified memory.” They are two 128GB nodes connected by Ethernet/RoCE and require explicit distributed execution.
- Three Sparks for “GLM-5.3 3-bit.” The model was announced but weights were explicitly delayed; the proposed artifact/backend/benchmark is absent; mainstream current GLM-5.2 files do not fit, while its smaller community GGUF lacks a three-node loader/run; the complete topology exceeds the budget.
- A GPU based on announced memory alone. AMD’s 192GB Ryzen AI Max+ PRO systems and NVIDIA’s RTX Spark Windows PCs were announced but lacked a verified shipping US configuration and price at cutoff. Upcoming is not purchasable.
- An aggressive low-bit frontier quant without a quality gate. A file that decodes quickly can still erase the capability you bought the larger model to obtain.
FAQ
Is local AI cheaper than an API?
Only at sustained utilization or when privacy/offline control has independent value. Divide complete purchase price, electricity, failures, and setup time by useful tokens—not theoretical peak. For occasional frontier-model use, renting usually wins above the dual-Spark tier.
How much memory headroom should I leave?
There is no universal percentage. Reserve the measured OS/runtime baseline, calculate the architecture-specific cache at the target context and precision, add graph/scratch allocations, then leave fragmentation margin. As a shopping screen, 10–20% beyond weights is safer than zero, but it does not replace an exact load test.
Does a 128GB unified-memory computer provide 128GB for the model?
No. The OS, runtime, display, cache, and buffers use the same pool. Framework’s Windows setting tops out at 96GB dedicated VRAM, while Linux can configure higher; neither makes all 128GB available to model weights. macOS and DGX OS likewise require headroom.
Does active MoE parameter count determine memory use?
No. Total parameters normally determine stored weight capacity; active parameters influence per-token compute. DeepSeek-V4-Flash’s core is 284B total and 13B active, not a 13B download; the 0731 packaged checkpoint also carries a speculative module.
Is “2-bit DQ” exactly two bits per weight?
No. The checked 96.531GB preview artifact implies roughly 2.72 bits per published parameter and keeps selected tensors at higher precision. The current 0731 OptiQ artifact separately reports 2.43 achieved bits/weight and uses SSD expert streaming. File naming describes a method/family, not a uniform physical rate or a quality guarantee.
Can the RTX 3090 run Qwen3.8-27B at its full 262K context?
Not entirely in 24GB with the cited quant and FP16 KV. The cache alone is 16GiB by the model’s architecture. A 32K target is calculated to fit with useful headroom; exact same-day performance remains unverified.
Is the dual-Spark 500K result interactive?
Generation is interactive at roughly 32 tokens/s once it starts, but the cited 500K prompt took about six minutes to prefill. That makes it a long-running analysis job.
Should I buy Apple, NVIDIA, or AMD?
Buy the software path. Choose NVIDIA for CUDA breadth and discrete-GPU latency, Apple for portable high-bandwidth unified memory and MLX, and Ryzen AI Max for the lowest current cost of clean 128GB x86 capacity. Exact model/backend support beats brand allegiance.
What about model quality?
This is a hardware/fit guide, not a leaderboard. Bigger parameter counts and higher prices do not guarantee better results for your tasks. Run a private evaluation set on the exact quant before purchasing around it.
Can I train these models locally?
Inference fit does not imply training fit. Full training requires weights, gradients, optimizer state, and activations; even adapters can exceed these systems at long sequence lengths. Follow an exact measured training recipe.
Research methodology
This is a BUILD run because no current article or prior ledger was supplied. Research began at 2026-08-14 12:05:57 PDT. Every recommendation was checked against current US pricing, an exact model card, an existing downloadable artifact, a credible backend path, memory headroom, and any configuration-rich benchmark available. Sources were accessed on August 14, 2026 unless a row says otherwise.
The hierarchy was: model publisher and license; hardware manufacturer price/specification; runtime documentation; reproducible benchmark repository; independent test; current retail/used listing. Search snippets were used only to find a page, never as the final evidence when a direct source was available. Retail and used prices are snapshots, not promises.
No hardware was physically tested by this publication. “Officially specified” is not relabeled as hands-on. Vendor benchmarks remain vendor benchmarks. Community results remain community results even when the repository is detailed. Calculated fit is not a speed claim.
The review tested the five starting hypotheses directly:
| Starting hypothesis | Verdict | Evidence-based replacement |
|---|---|---|
| ~$2K RTX 3090 + Qwen3.8-27B | Revise/support with caveats | Complete system is ~$2.4K. Exact Q4_K_M fit at 32K is calculated; model and quant are same-day; speed unverified. |
| ~$4K one Spark + DeepSeek-V4-Flash “2.5-bit” | Reject as stated | Spark is $4,699. Current 0731 official artifact does not fit. A preview-based custom ~2.44 effective-bpw GGUF fits and independently measured 13.2 tok/s at 2K, falling to 7.94 at 262K, with unresolved quality. |
| ~$7K M5 Max 128GB + DeepSeek “2.5-bit” | Support hardware; split the model verdict | Exact $6,699 system. The fast 34.1 tok/s mean is a preview-based custom DS4 artifact using ~115GB. Current 0731 has a 92.487GB/2.43-bpw SSD-streamed OptiQ artifact but no M5 Max speed or quality score. |
| ~$8K dual Spark using NVFP4/6.5-bit DQ | Reject price and quant premise | Current official bundle is $9,449. Best evidence is a community TP=2 FP8 path at ~40 tok/s, not a verified 6.5-bit DQ recommendation. |
| ~$12K three Spark + GLM-5.3 3-bit DQ | Reject for now | GLM-5.3 was announced today, but weights are delayed and no 3-bit DQ/backend/run exists. GLM-5.2 does not validate the claim; complete hardware starts above $14.3K. |
Sources
Models, artifacts, and licenses
- Qwen: Qwen3.8-27B official model card; Unsloth Qwen3.8-27B GGUF; Qwen3.6-35B-A3B.
- DeepSeek: current DeepSeek-V4-Flash-0731 card/license; official V4 collection; technical report; June preview card; custom GGUF; preview MLX 2bit-DQ; current 0731 OptiQ; NVIDIA preview NVFP4.
- GLM: GLM-5.3 announcement and delayed-weight statement; GLM-5.2 official card; GLM-5.2 UD-IQ3_S artifact; experimental GLM-spark project.
- OpenAI:
gpt-oss-20b;gpt-oss-120b. - Mistral: Mistral Small 4 119B; official NVFP4 artifact.
- NVIDIA: Nemotron 3 Nano Omni 30B-A3B BF16; NVFP4 artifact; Open Model Agreement.
- Google: Gemma 4 26B-A4B-it; Google QAT Q4_0 GGUF.
- Moonshot AI: Kimi Linear 48B-A3B; Kimi K3 card and custom license.
- MiniMax/IBM: MiniMax-M2.7 and its non-commercial license; IBM Granite 4.0 H Small; official Granite GGUF.
Hardware and current prices
- NVIDIA: RTX 3090 specifications; DGX Spark store; DGX Spark hardware guide; Spark clustering guide; two-unit bundle; RTX PRO 6000 marketplace.
- Used workstation: Precision 5820 listing; RTX 3090 used-price index.
- Framework/AMD: Framework configurator; Framework AI results; Ryzen AI Max+ 395 specifications.
- Apple: exact M5 Max configuration; M5 Max specifications/release; 14-inch technical specifications.
Benchmarks and runtimes
- Independent DS4 DeepSeek-V4-Flash on one DGX Spark, with commit, contexts, and runtime memory.
- Detailed M5 Max DS4 run and thermal follow-up.
- oMLX DeepSeek-V4-Flash preview 2bit-DQ on M5 Max, a useful but incompletely disclosed context/batch sweep.
- Dual/single DGX Spark DeepSeek-V4-Flash repository, including dual-node configuration, cross-machine result, and long-context result.
- Framework’s LM Studio/Fedora benchmark, treated as a vendor result.
- llama.cpp repository, the relevant GGUF runtime family.
What changed in this review
This is the first BUILD, so there is no prior recommendation set to diff. The initial baseline:
- replaces the requested round-number ladder with live complete-system costs of about $2.4K, $3.98K, $4.70K, $6.70K, and $9.45K;
- adds same-day Qwen3.8-27B and its exact downloadable
Q4_K_Martifact while refusing to invent RTX 3090 speed; - separates current DeepSeek-V4-Flash-0731 from the superseded preview and every preview-based aggressive GGUF/MLX/NVFP4 artifact;
- supports the M5 Max hardware while separating the fast, preview-based DS4 result from the current 0731 SSD-streaming artifact with no M5 Max benchmark;
- moves dual Spark from $8K to $9,449 and labels its best result as a custom community stack over 200GbE/RoCE;
- records GLM-5.3 as announced-but-unreleased, rejects the current $12K three-node claim, and rejects any suggestion that two nodes create transparent unified memory;
- creates an evidence ledger and a weekly review state so the next pass can update only claims that changed.
