AI News

Local AI Hardware Guide 2026: What Models Run at Every Budget

The best local-AI machine is the least expensive complete system that runs your exact model, quantization, context, and workload at an acceptable speed—not the box with the largest parameter claim.

Last fully reviewed: September 4, 2026 at 12:36:16 PM PDT
Next scheduled review: September 11, 2026 at 12:36:16 PM PDT
Market basis: United States, USD, complete usable systems before tax; prices and availability accessed September 4, 2026.

For the wider decision map, use the AI Hardware hub to compare tested systems, fit guides, buying options and local runtimes.

Opening verdict

Buy a used RTX 3090 workstation at about $2,450 if 24GB is enough. It remains the strongest entry point because it combines mature CUDA support, 936GB/s of VRAM bandwidth, and replaceable PC parts. Its first pairing in this guide is Qwen3.8-27B at ggml-org’s 18.974GB Q4_K_M GGUF. A pinned five-run llama.cpp package measures 37.39 decode tokens/s without speculative decoding and 60.33 tokens/s with the matching Q4_0 MTP draft on one RTX 3090 at a configured 64K context. Those are synthetic 512-token medians, not a promise for every prompt or runtime.

At roughly $4,000, choose capacity over peak speed: the 128GB Framework Desktop can run models that a 24GB GPU cannot. Framework now publishes an exact current-checkpoint DeepSeek-V4-Flash-0731 lane at 14.7–14.9 tokens/s without speculative decoding and 21.73–24.64 with it, alongside its older 38-tokens/s gpt-oss-120b result. At $4,699, DGX Spark is the compact CUDA appliance to buy when software compatibility and 128GB of coherent memory matter more than decode latency. At $6,699, the 128GB M5 Max MacBook Pro remains the best portable large-model system. Apple has also opened preorders for a 128GB M5 Max Mac Studio at $5,899, shipping September 22; it is a serious desktop-value challenger, but it has no exact local-model benchmark yet. At $9,449, two Sparks are an expert-only distributed system, and the official bundle was out of stock at this review.

Do not force a $12,000 purchase. GLM-5.3 now has a public four-DGX-Spark run at 27.2 single-stream decode tokens/s, but four units cost $18,796 before interconnect and the result still lacks an independent retained-quality evaluation. No validated three-Spark run was found. Apple’s newly announced 256GB M5 Ultra Mac Studio is $11,299 and ships September 22, making it the first serious single-box in-budget candidate for this tier—but it has no exact public local-model run yet. The default remains wait, measure, and rent rather than preorder on specifications alone.

The short answer

Performance in this table belongs only to the exact model and artifact named in the same row. It is not a generic speed rating for the machine.

Budget Recommended complete system Current system cost Accelerator and usable memory Bandwidth or interconnect Best-matched model and exact quantization Practical tested context Verified decode speed Power considerations Main compromise Evidence class Confidence Last verified
~$2.5K Refurbished Dell Precision 5820, 64GB/1TB/950W + used RTX 3090 $2,433–$2,463 conditional market estimate RTX 3090, 24GB GDDR6X 936GB/s Qwen3.8-27B, ggml-org Q4_K_M GGUF (17.671GiB) + BF16 vision projector (0.867GiB); optional Q4_0 MTP draft (1.565GiB) Measured: 64K with Q4 KV; calculated conservative target: 32K with FP16 KV 37.39 tok/s baseline; 60.33 with MTP (five-run 512-token medians) 350W GPU; 950W PSU; used-card thermals vary No pinned GPU SKU; fit, connectors, quant quality, and warranty vary Pinned community benchmark + officially specified hardware Medium-high for exact measured path; medium for purchase 2026-09-04
~$4.0K Framework Desktop, Ryzen AI Max+ 395, 128GB, 2TB, Linux $3,978 Radeon 8060S; 128GB shared. Windows preset allows up to 96GB dedicated VRAM; Linux can configure higher, subject to OS/runtime headroom 256GB/s shared Current primary lane: DeepSeek-V4-Flash-0731, Unsloth UD-IQ2_XXS (90.9GB); stable alternative: gpt-oss-120b MXFP4 DeepSeek vendor test: 2,048-token prompt + 128 generated tokens 14.7–14.9 tok/s baseline; 21.73–24.64 with DSpark; older gpt-oss-120b result: 38 tok/s 120W sustained/140W boost processor envelope; 400W PSU Soldered memory; lower bandwidth; exact backend support varies; vendor results only Officially specified/vendor benchmark Medium-high for the exact reported run; lower for transferability and quality 2026-09-04
~$4.7K NVIDIA DGX Spark, 128GB/4TB $4,699 GB10; 128GB coherent unified memory 273GB/s Stable fit: Mistral Small 4 119B official NVFP4 (70.847GB). Experimental: GLM-5.3-Flash UD-Q2_K_XL (101.253GiB) in a custom llama.cpp branch GLM: 32K measured; 256K with Q8 KV served; 1M Q8 KV caused hard OOM 17.9–33.8 tok/s across four 32K workload shapes with MTP; 17.5 without MTP 140W GB10 TDP; 240W PSU Aggressive quant and custom branches; capability retention and independent reproduction remain unresolved Community-reported exact run + officially specified hardware Medium for exact run; low for retained quality and transferability 2026-09-04
~$6.7K 14-inch MacBook Pro, M5 Max 40-core GPU, 128GB/2TB $6,699 128GB unified; exact custom run used 114.5–115.1GB 614GB/s Measured speed path: June DeepSeek-V4-Flash custom ~81GB IQ2XXS/Q2 in DS4. Current 0731 option: MLX OptiQ 2-bit, 92.487GB/2.43 achieved bpw, SSD-streamed and unbenchmarked on M5 Max DS4 configured 300K; actual measured prompts were short. No independently verified long-context result is established here. 34.1 tok/s mean across three DS4 prompts; ~10% thermal decline 96W adapter included; GPU reached 98°C in cited sustained test; watts unreported Fast evidence targets the superseded June preview; current 0731 quality/speed unresolved; no CUDA Independently benchmarked custom artifact + official hardware Medium-high for exact preview run; low for current-checkpoint/quality claims 2026-09-04
~$9.45K NVIDIA dual-DGX-Spark bundle with direct cable $9,449; out of stock at review 2 × 128GB local pools; not one transparent 256GB pool 273GB/s per node; 200GbE/RoCE cross-node Two distinct June/native DeepSeek custom paths: Tony recipe uses TP=2/PP=1, MTP=2, FP8 KV without expert parallel; Elsung uses vLLM 0.21.1, TP=2, expert parallel, MTP=2, FP8 KV. Neither is official NVFP4-on-GB10 Both set/report up to 500K; Tony omits a complete raw harness Tony: project-reported ~40 tok/s single stream, ~92 aggregate at c=8. Elsung: 40/38/32 at 8K/200K/500K and ~351 aggregate at 256K/c=32. Artifact lineage remains incompletely sealed 2 × 140W GB10 TDP; 2 × 240W PSU capacity Custom mutable images, incomplete benchmark metadata, multi-node failure modes; current 0731 equivalence unresolved Community-reported demonstrations Medium for feasibility; low-medium for exact speed 2026-09-04
~$12K Buy nothing by default; wait for exact tests $0 incremental Candidate: M5 Ultra Mac Studio 256GB/2TB at $11,299, shipping September 22 1.2TB/s specified Serious single-box candidate, but no exact public local-model benchmark; four-Spark GLM-5.3 path costs $18,796 before interconnect — — Unmeasured for the new Mac Studio configuration Preorder hardware is not tested hardware; no validated three-Spark GLM path exists Official pricing/specifications + community four-Spark run + verified rejection High 2026-09-04

Three labels matter in that table. Officially specified means a manufacturer or model publisher states the fact. Calculated means the arithmetic is shown but the pairing was not tested. Community-reported means a public result has useful configuration detail but is not an independent lab result. None of the machines in this review was physically tested by this publication, so nothing is labeled “Verified hands-on.”

How to interpret “can run”

“It loads” is the weakest useful claim. A system can hold the weights and still be miserable because the context cache exhausts memory, prompt processing takes minutes, generation is too slow, the backend lacks optimized kernels, or a distributed link turns every token into a synchronization exercise.

This guide uses these experience bands for batch size 1:

  • Comfortable interactive: at least 20 decode tokens/s and acceptable time to first token at the stated context.
  • Usable with compromises: 8–20 decode tokens/s, or fast decode paired with noticeable prompt delay.
  • Slow but practical for queued work: 2–8 decode tokens/s.
  • Fits but is not usefully interactive: below 2 decode tokens/s.
  • Experimental: an immature model conversion, backend, hardware port, or topology makes reproducibility the main risk regardless of speed.
  • Does not fit: weights plus runtime allocations and the target cache cannot coexist.
  • Unverified: the necessary artifact, backend, memory trace, or benchmark evidence is missing.

Prefill and decode are separate. A dual-Spark DeepSeek run can decode around 32 tokens/s at 500K context yet still take about six minutes to process that prompt. For an interactive user, that is a queued workload despite the healthy decode number. Concurrency is separate again: the same dual-Spark result exposes roughly 1.1 million KV-cache tokens shared by all live requests, not 1.1 million per user.

Detailed recommendations by budget

About $2,500: used RTX 3090 workstation — best value, conditional buy

The purchase. The conditional market estimate is a $1,134 refurbished Precision 5820 with Xeon W-2155, 64GB RAM, 1TB NVMe, Windows 11 Pro, and a 950W PSU, plus the current $1,299 used-market RTX 3090 median and a $0–$30 lead/adapter allowance. That totals $2,433–$2,463, but it is not a reproducible pinned-SKU BOM: the index is based on eight qualifying recent sales, and RTX 3090 AIB cards vary in length, thickness, and use two or three power connectors. NVIDIA specifies 24GB GDDR6X, 936GB/s, and 350W for the card. Confirm the exact card’s clearance and required two or three independent feeds before ordering. See the live refurb listing, used-price index, and NVIDIA specifications.

The model pairing. Qwen3.8-27B is a 27B dense, multimodal, Apache-2.0-licensed open-weight model with 262,144 native context and official 1M YaRN guidance. The reproducible speed path uses ggml-org’s Q4_K_M GGUF, 18.974GB (17.671GiB), plus a 0.931GB (0.867GiB) BF16 multimodal projector. Its optional Q4_0 MTP draft is 1.680GB (1.565GiB). These are third-party conversions of the official BF16 release, not publisher quants.

The exact Unsloth Dynamic v3 UD-Q4_K_M is a smaller alternative at 16.464GB (15.334GiB), with a 0.928GB (0.864GiB) F16 projector. Unsloth replaced its initial release during the week, so the old 17.107GB byte claim is obsolete. No benchmark in this guide transfers the ggml-org result to that different quant.

The architecture has 16 full-attention layers, four KV heads, head dimension 256, and FP16 K and V. The cache is therefore:

16 layers × 4 KV heads × 256 dimensions × 2 (K + V) × 2 bytes
= 65,536 bytes/token = 64KiB/token

That is 0.5GiB at 8K, 2GiB at 32K, 4GiB at 64K, 8GiB at 128K, and 16GiB at the native 262K limit. The recommended ggml-org weights + projector + 32K FP16 KV consume about 20.54GiB before graph buffers, allocator overhead, OS/display use, and fragmentation. A 24GB 3090 therefore has a credible 32K FP16-KV target. A pinned run demonstrated 64K with Q4_0 K/V cache and the MTP draft; full 262K still does not fit on this card in the recommended configuration.

The reproduction package pins llama.cpp commit 9b05354e, the ggml-org repository revision and model hashes, exact profiles, prompts, and repeated results. On one RTX 3090, five 512-token runs produced a 37.39 tok/s median without speculative decoding and 60.33 tok/s with the Q4_0 MTP draft at a configured 65,536-token context, Q4_0 KV, flash attention, and a 2,048-token batch. The test is unusually reproducible for community evidence, but it is synthetic decode throughput, not a quality score, long-prompt prefill test, or whole-system power/thermal measurement.

Why it wins. CUDA and llama.cpp-family tooling are mature, VRAM bandwidth is high, and the rest of the PC is replaceable. What it cannot do: it cannot hold high-quality frontier-scale weights or Qwen’s full context entirely in VRAM. Planning horizon: two to three years for 20–30B low-bit models, subject to used-hardware health—not a warranty promise. Upgrade path: more system RAM/SSD, then a higher-memory GPU or a new platform; a second 3090 is not a casual addition to this chassis. Setup: medium-to-high. Buy it for low-latency single-user chat, coding, and experimentation. Skip it if you need a warranty-backed appliance, quiet portability, or models above roughly the 24GB class.

Before the return deadline, stress all 24GB of VRAM, log memory-junction and hotspot temperatures, verify the PCIe link, and inspect power connectors. A former mining card with failing rear-side GDDR6X is not a bargain.

About $4,000: Framework Desktop 128GB — best capacity per current complete-system dollar

The current complete Linux configuration is $3,978: $3,449 for the Ryzen AI Max+ 395/128GB system, $505 for Framework’s 2TB SSD, $19 for the required fan, and $5 for the power cord. Framework specifies 128GB LPDDR5X-8000 and lists a 96GB maximum dedicated-VRAM setting for Windows; its official ML page explicitly says Linux can override the VRAM setting higher. Total model-usable memory still remains below 128GB after the OS, runtime, cache, and buffers. AMD specifies 256GB/s shared bandwidth, and the processor runs at a 120W sustained, 140W boost envelope. Memory is soldered. Pricing and configuration are from the official configurator, with machine-learning results and AMD bandwidth specifications. Framework now labels a 192GB Desktop as “coming soon,” but does not publish a buyable US configuration or complete price; it is a watch item, not a replacement recommendation.

Framework’s current primary 128GB pairing is DeepSeek-V4-Flash-0731 through Unsloth’s 90.9GB UD-IQ2_XXS quant. In Framework’s September 2 vendor test, a 2,048-token prompt and 128 generated tokens produced 14.90 tokens/s in llama.cpp ROCm and 14.70 in llama.cpp Vulkan/RADV without speculative decoding. With the DSpark draft technique, the same paths reached 21.73 and 24.64 tokens/s. This is the first exact current-checkpoint Framework result in the guide, but it is still vendor evidence: capability retention, sustained thermals, power, long-context behavior, and independent reproduction remain unresolved. Framework’s published gpt-oss-120b MXFP4 result of 38 decode tokens/s remains a simpler stable alternative, though its context, prompt/output lengths, LM Studio build, flags, and power mode were not disclosed.

Why it wins. It crosses from 24GB to genuinely large shared memory without distributed inference. What it cannot do: 256GB/s is far below the 3090’s dedicated VRAM bandwidth, and 128GB installed is not 128GB free for weights even when Linux configures more than 96GB for graphics. Planning horizon: three to four years for large quantized models if ROCm/Vulkan support continues improving. Upgrade path: SSD, fan, PSU, and case parts are serviceable; memory is not. Setup: medium; Linux and exact runtime selection matter. Buy it for large-model capacity, an x86 host, and quiet compact use. Skip it if a 24GB model covers your needs and latency matters more than capacity, or if your workload requires CUDA-only libraries.

$4,699: DGX Spark — best compact CUDA appliance, not the fastest decoder

NVIDIA’s current listing includes GB10, 128GB coherent unified memory, 4TB NVMe, a 240W PSU, and a power cord for $4,699. The direct single-unit listing showed Add to Cart when checked September 4; stock can change independently of list price. The hardware guide specifies 273GB/s memory bandwidth, a 140W GB10 TDP, a 20-core Arm CPU, 10GbE, and two ConnectX-7 ports. NVIDIA’s “1 PFLOP FP4” and “up to 200B parameters” are peak/capacity positioning, not measured LLM throughput.

The conservative pairing is Mistral’s official Mistral Small 4 119B NVFP4, a 70.847GB compressed-tensors artifact derived from the Apache-2.0 base release. The MoE has roughly 6–6.5B routed-active parameters, or about 8B when embeddings/output are counted. The artifact leaves ample nominal capacity for runtime and cache, but no reliable exact GB10 decode benchmark was found.

An experimental pairing is GLM-5.3-Flash, a separate 320B-total/18B-active MIT-licensed model—not the 753B GLM-5.3 flagship. Unsloth’s UD-Q2_K_XL totals 108.720GB (101.253GiB). A one-Spark community run used a custom llama.cpp branch with multi-token prediction and reported 17.9 tokens/s for prose, 25.0 for code, 26.6 for JSON, and 33.8 for structured output at a 32K configuration; the same code workload produced 17.5 tokens/s without MTP. It also served at 256K with Q8 KV, while 1M Q8 KV caused a hard out-of-memory failure. A newer single-Spark EXL3 recipe reports roughly 15.7–16.5 tokens/s with MTP at 8K and a patched 258K prefill, while a dual-Spark NVFP4 recipe reports 26.5 tokens/s and a 200K retrieval run. These repositories expand feasibility evidence but use different artifacts, custom patches, and community harnesses; they do not establish retained quality or a portable product-level speed. Treat GLM-5.3-Flash as a test lane, not a reason by itself to buy the hardware.

The frontier experiment is DeepSeek-V4-Flash. The current publisher checkpoint is DeepSeek-V4-Flash-0731: 284B core parameters/13B active, 1M context, an attached speculative module in the packaged checkpoint, and 166.899GB of mixed FP4/FP8 files. It does not fit one 128GB Spark. The best one-Spark evidence instead targets the superseded June preview through a community IQ2XXS/Q2 GGUF of roughly 81GiB and the dedicated DS4 engine—not generic llama.cpp.

An independent exact Spark test pinned DS4 commit a97e7a3, Ubuntu 24.04, CUDA 13.0, driver 580.142, and generated 128 tokens through 65K contexts, plus a 64-token doubling test to 262K. Decode fell from 13.2 tokens/s at 2K to 11.8 at 32K, 11.4 at 65K, 9.95 at 131K, and 7.94 at 262K; peak runtime memory was about 115GiB. Reported prefill was 64.9 tokens/s at 2K, 162.6 at 32K, and 247.1 in the one-shot 65K run; the separate incremental test reported 165.2 at 131K and 109.3 at 262K. Those methods differ, so the values are not one monotonic context curve. No power measurement was supplied. This is an aggressive, non-official conversion with higher-precision attention/shared/output pieces and unknown quality loss. It is not “current DeepSeek at 2.5-bit” in a quality-neutral sense.

Why it wins. It is a compact, warrantied CUDA development appliance with far more memory than consumer NVIDIA cards. What it cannot do: it cannot load the current official DeepSeek-V4-Flash-0731 artifact, and its 273GB/s bandwidth limits single-stream decode. Planning horizon: three to four years as a CUDA development node, conditional on Arm64 container support. Upgrade path: external/network storage and a second Spark; memory is soldered. Setup: low for supported single-node containers, high for custom quants. Buy it for compact CUDA capacity and NVIDIA’s software ecosystem. Skip it if you only run sub-24GB models or expect desktop-GPU latency.

$6,699: 128GB M5 Max MacBook Pro — best portable large-model system

Apple’s exact 14-inch M5 Max 40-core GPU, 128GB, 2TB configuration is $6,699 with its display, battery, keyboard, 96W adapter, and macOS included. Apple specifies 614GB/s memory bandwidth and up to 128GB unified memory in the M5 Max announcement. Memory and storage are not upgradeable.

There are three distinct DeepSeek paths; do not blend them:

  1. Best exact speed evidence, superseded preview: the same custom ~81GB IQ2XXS/Q2 artifact used on Spark, running in DS4. A detailed M5 Max setup and three-prompt test pinned DS4 commit 560662d, configured a 300K context ceiling, used temperature 0/seed 1, and reported 36.3, 33.6, and 32.5 decode tokens/s—34.1 average—with 114.5–115.1GB runtime memory. A thermal follow-up observed roughly a 10% decode decline as GPU temperature rose from 51°C to 98°C. The actual measured prompts were short; a configured ceiling is not a filled 300K prompt.
  2. Preview quantization; benchmark temporarily unavailable: mlx-community/DeepSeek-V4-Flash-2bit-DQ is a 96.531GB third-party dynamic quant, about 2.72 effective stored bits per 284B parameter, with selected tensors at four, six, or eight bits. The oMLX community benchmark browser is temporarily unavailable during a database redesign. Its earlier throughput and context reports are excluded from this recommendation; the independently documented DS4 results above remain separate evidence.
  3. Current 0731 checkpoint: DeepSeek-V4-Flash-0731-OptiQ-2bit is 92.487GB, claims 2.43 achieved bits/weight, and streams routed experts from SSD via mlx-optiq. Its producer reports about 2.5 tokens/s on an M3 Max and publishes no capability-retention score. No exact M5 Max result was found.

Why it wins. It pairs high bandwidth and large memory with a complete portable computer, and Metal/MLX has a strong Apple-native inference path. What it cannot do: it cannot run CUDA software; the fast evidence is for a superseded preview and custom engine; the current 0731 quant’s capability loss and M5 Max speed are not established. Full 1M context was not tested. Planning horizon: three to five years for portable local inference, bounded by soldered capacity and evolving MLX support. Upgrade path: none internally; use external storage or replace the machine. Setup: low-to-medium. Buy it if portability and single-box capacity matter. Skip it if CUDA compatibility, repairability, or guaranteed current-checkpoint behavior dominates.

$5,899 preorder watch: 128GB M5 Max Mac Studio — likely desktop-value challenger, not yet tested

Apple announced a new M5 Max and M5 Ultra Mac Studio on August 25, with delivery beginning September 22. The exact M5 Max 40-core GPU, 128GB, 2TB configuration is $5,899, $800 below the like-memory MacBook Pro. Apple specifies the same 614GB/s memory bandwidth for this M5 Max tier. The desktop chassis may sustain long inference better than the 14-inch laptop, but that is an inference from form factor—not a measured result.

Do not preorder it solely for local AI. No exact current-model test, wall-power result, acoustic measurement, or sustained thermal run exists for this shipping configuration yet. If September 22 testing confirms the expected sustained behavior, it could replace the MacBook as the Apple desktop recommendation while the MacBook remains the portable winner.

$9,449: two DGX Sparks — best compact frontier experiment

The official bundle contains two 128GB/4TB units and a direct cable for $9,449, but NVIDIA’s page showed Out of Stock on September 4. The link is Ethernet/RoCE over ConnectX-7 at up to 200Gb/s—not cross-box NVLink. Two boxes provide two local 128GB pools and 8TB aggregate storage; a distributed backend must shard the model. NVIDIA documents the direct topology in its clustering guide.

Two community branches establish feasibility but not an official product profile. The maintained TP2 500K recipe uses a custom arm64 vLLM image, tensor parallelism 2, pipeline parallelism 1, FP8 KV, 500K maximum context, eight sequence slots, MTP draft length 2, and an MP distributed executor. It reports roughly 40 tokens/s single stream and 92 aggregate at concurrency eight, but omits an immutable image digest, exact prompts/output lengths, raw result log, TTFT, prefill, power, and test date.

A second dual-Spark repository reports a roughly 149GB FP8 path with vLLM 0.21.1, tensor parallelism 2, expert parallelism, MTP 2, and FP8 KV. It publishes the more detailed figures: roughly 40 decode tokens/s at 8K, 38 at 200K, and 32 at 500K; TTFT rose from four seconds to roughly six minutes, while a 256K profile reached about 351 aggregate tokens/s at concurrency 32 with a shared ~1.1M-token KV pool. Its authors explicitly call the documentation AI-written and advise caution. Neither repo proves NVIDIA’s 168.305GB NVFP4 artifact is officially supported on GB10; NVIDIA’s card validates B200/GB300-class hardware. Neither result is sealed tightly enough to establish current DeepSeek-V4-Flash-0731 equivalence.

Why it wins. It is the smallest current system with public evidence of interactive DeepSeek-V4-class serving at high precision and meaningful concurrency. What it cannot do: it does not make distributed inference turnkey, and it does not pool memory transparently. Planning horizon: two to four years as an experimental serving cluster; distributed software risk dominates hardware age. Upgrade path: more nodes are theoretically possible, but the cited recipe validates only two. Setup: high; expect RoCE/NCCL, container, kernel, and failure-recovery work. Buy it only if the exact DeepSeek-quality/capacity or multi-user workload justifies the operational burden. Skip it for a personal chat box, first local-AI system, or anything that already fits one GPU.

Around $12,000: stop, measure, and rent

Three current Sparks are not a $12K system. Three single units cost $14,097 before networking. NVIDIA’s direct three-Spark ring requires three QSFP links total; this guide did not capture a sufficiently evidenced cable price and therefore does not publish a false complete total. No exact GLM runtime was validated on that physical topology. Four Sparks now have a public GLM-5.3 community run, but the hardware alone costs $18,796 before interconnect. NVIDIA’s RTX PRO 6000 Blackwell Workstation Edition is currently $16,000 for the GPU alone and unavailable on the manufacturer marketplace.

The model premise is now testable but still unproven. GLM-5.3 shipped August 28 as a 753B-total MoE. Its official mixed/F8 safetensors total 755.632GB and its BF16 repository totals 1.507TB. The weights use the custom GLM-5.3 License, not Apache-2.0 or MIT; the terms are broadly permissive but include a security-review condition for very large model-as-a-service operators. Official runtime support names SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth.

Unsloth’s UD-Q3_K_XL GGUF totals 342.966GB (319.41GiB). Three nominal 128GB nodes provide 384GB (about 357.6GiB), leaving only about 38.2GiB before three operating systems, runtime allocations, KV cache, buffers, collectives, and fragmentation. A four-Spark mixed-precision community repository now reports 27.2 decode tokens/s at concurrency one, 59.1 aggregate at four, 86.9 at eight, and a 200,064-token KV pool. It uses 377.4GiB of weights across four ranks and reports a ten-sample 91.4% quality battery—too small to establish capability retention. This validates one four-node experimental path, not the proposed three-node system. No three-Spark loader, long-context memory trace, throughput result, failure-recovery test, or capability-retention comparison was found.

Apple’s exact M5 Ultra 36-core CPU, 80-core GPU, 256GB, 2TB Mac Studio is $11,299 and specifies 1.2TB/s memory bandwidth. It ships September 22. On paper, that makes it the first serious single-box candidate inside this tier, but no exact local-model artifact, context, speed, power, or thermal result exists. A preorder page is not a benchmark.

DeepSeek also published DeepSeek-V4-Pro-0813 during this review cycle. Its official repository parses to about 1.65T parameters and 892.762GB of files, and the publisher’s example uses four GB300 GPUs. It is MIT-licensed but does not fit any system in this ladder, so it does not displace the local Flash-0731 analysis.

Buy nothing by default. Use the dual-Spark system, rent a larger GPU instance for the frontier model, or wait for post-launch M5 Ultra and 192GB AMD evidence. Reopen the capital decision only after exact tests record required quality, context, concurrency, latency, power, and sustained behavior.

Model-to-hardware compatibility matrix

“Good” means a downloadable artifact and credible runtime path exist with useful headroom. “Possible” means fit is plausible but the exact combination lacks a strong benchmark. “Experimental” signals conversion/backend risk. “No” means the recommended artifact does not fit or the hardware path is inappropriate. Context entries are publisher maxima, not promises of practical local use.

License language is deliberate. Open weights means the weights can be downloaded under stated terms; it does not prove that training data, code, or the full pipeline is open. Apache-2.0 and MIT weight releases are identified as permissively licensed open weights. Source-available is used for public weights/code under custom or restrictive terms. MiniMax-M2.7 is non-commercial without separate authorization; NVIDIA Nemotron and Kimi K3 use custom agreements. None is flattened into the blanket label “open source.”

Model Type; total / active parameters Publisher context Weight-license status Practical local artifact RTX 3090 24GB Framework 128GB 1× Spark M5 Max 128GB 2× Spark
Qwen3.8-27B Dense multimodal; 27.78B parsed 262K native; 1M YaRN Apache-2.0 open weights Measured path: ggml-org Q4_K_M 18.974GB + 0.931GB projector + optional 1.680GB Q4_0 MTP; smaller Unsloth Dynamic v3 alternative 16.464GB Good at 64K with Q4 KV; pinned llama.cpp result is 37.39 tok/s baseline, 60.33 with MTP Possible; exact backend speed unverified Possible; exact speed unverified Possible via GGUF/Metal; no exact run Wasteful
Qwen3.8-Flash-Next Hybrid MoE; 125B main / 6B active plus n-gram table and MTP modules; ~176B stored total 262K native Custom Qwen Community License; source-available open weights Official safetensors 360.000GB; Unsloth UD-Q2_K_XL 78.869GB Experimental; no pinned run Experimental; reported decode cliff after ~1K context Experimental; native 262K CUDA abort reported Possible by bytes; exact MLX run absent Unnecessary until runtime stabilizes
Qwen3.6-27B Dense multimodal; 27.78B parsed 262K native; ~1.01M extended Apache-2.0 open weights ggml-org Q4_K_M 19.1GB; MLX 4-bit 16.1GB Good; detailed 3090 predecessor run exists Good Good Good Wasteful
Qwen3.6-35B-A3B MoE multimodal; 35.95B / ~3B 262K native; ~1.01M extended Apache-2.0 open weights ggml-org Q4_K_M 20.420GB + optional 1.060GB MTP; MLX 20.429GB Fits, but cache margin tight Good Good Good Wasteful
DeepSeek-V4-Flash-0731 MoE reasoning/coding; 284B core / 13B active; packaged checkpoint parses 304.18B with speculative module 1M MIT open weights Official mixed tree 166.899GB; Framework-tested Unsloth UD-IQ2_XXS 90.9GB; MLX OptiQ 92.487GB/2.43 bpw No Experimental; vendor-measured 14.7–14.9 tok/s baseline, 21.73–24.64 with DSpark at 2K prompt No Experimental; exact M5 Max speed/quality absent Possible capacity, but no sealed 0731 TP2 run
DeepSeek-V4-Flash (June preview) MoE reasoning/coding; 284B / 13B 1M MIT open weights; superseded checkpoint NVIDIA NVFP4 168.305GB; custom ~81GiB DS4 GGUF; MLX DQ 96.531GB No No verified DS4/Vulkan path Experimental; independently tested custom DS4 Experimental; independently tested DS4 path; oMLX benchmark browser unavailable Community TP2 demonstrated; not official NVFP4-on-GB10
DeepSeek-V4-Pro-0813 Frontier MoE; 1.65T parsed 1M MIT open weights Official mixed repository 892.762GB No No No No No
GLM-4.7-Flash MoE assistant/coding; 31.22B / ~3B 202,752 MIT open weights Unsloth Q4_K_M 18.3GB; MLX 4-bit 16.9GB Good Good Good Good Wasteful
GLM-5.2 Frontier MoE; ~744B / ~40B (753.33B packaged) 1M MIT open weights Unsloth UD-IQ3_S 308.641GB; patched MLX 2.56-bpw ~246.9GB No No No No No; still exceeds two-node capacity/headroom
GLM-5.3-Flash Hybrid multimodal MoE; 320B / 18B Up to 1M in publisher evaluations MIT open weights Official mixed/F8 328.337GB; Unsloth UD-Q2_K_XL 108.720GB No No verified path Experimental; one community 32K run and 256K load Possible by bytes; exact MLX/run absent Possible but unnecessary for the cited quant
GLM-5.3 Frontier MoE; 753B total 1M in publisher evaluations Custom GLM-5.3 License; source-available open weights Official mixed/F8 755.632GB; Unsloth UD-Q3_K_XL 342.966GB No No No No No; three-node aggregate fit is plausible but unvalidated
gpt-oss-20b MoE reasoning/tool use; 21B / 3.6B 131K Apache-2.0 open weights Official MXFP4 Good Good; vendor reports 58 tok/s Good Good Wasteful
gpt-oss-120b MoE reasoning/tool use; 117B / 5.1B 131K Apache-2.0 open weights Official MXFP4 No Good; vendor reports 38 tok/s Good; llama.cpp reports 42.76 tok/s at 32K Good Unnecessary
Gemma 4 26B-A4B-it MoE multimodal; 25.2B / 3.8B 256K Apache-2.0 open weights Google QAT Q4_0 GGUF 14.439GB; 15.634GB with projector Good Good Good Good Wasteful
Mistral Small 4 119B MoE multimodal/coding; 119.4B / ~6–8B 256K Apache-2.0 open weights Official NVFP4 70.847GB; Unsloth Q4_K_M 73.763GB No Possible via GGUF/Vulkan; exact run absent Good NVFP4 fit; exact speed unverified Possible via GGUF/Metal; exact run absent Good but unnecessary
Kimi-Linear-48B-A3B-Instruct Hybrid linear-attention MoE; 48B / 3B 1M MIT open weights MLX 4-bit 27.644GB; community AWQ 4-bit No Good via suitable backend Possible; exact run absent Good MLX fit Unnecessary
MiniMax-M2.7 MoE; 228.69B; active total not disclosed for this release 204,800 Non-commercial/source-available; commercial authorization required MLX 3-bit 100.103GB; community GGUF No Experimental; MLX artifact is Apple-specific and a GGUF/Vulkan path needs separate headroom testing Experimental community low-bit only Possible MLX fit; near memory edge Possible; license and quality caveats
Nemotron 3 Nano Omni 30B-A3B Reasoning Hybrid multimodal MoE; 31B / ~3B 256K Custom NVIDIA Open Model Agreement; not Apache/MIT/OSI Official NVFP4 22.432GB Fits by bytes but no optimized Ampere NVFP4 path verified Possible via conversion Good NVFP4 fit Possible only after conversion Wasteful
IBM Granite 4.0 H Small Hybrid Mamba/attention MoE; 32B / 9B 128K Apache-2.0 open weights IBM official Q4_K_M GGUF 19.477GB Good Good Good Good Wasteful
Kimi K3 Frontier multimodal MoE; 2.8T / 104B 1M Custom Kimi K3 license; open-weight/source-available, not OSI open source Official QAT MXFP4/MXFP8 repository ~1.561TB No No No No No

The matrix intentionally includes some “No” rows. A current-model guide should show where local ownership stops being rational, not imply that every headline release belongs on a desk.

Memory and quantization methodology

Start with the artifact, not a slogan. The rough equation

raw weight memory = parameter count × effective bits per weight ÷ 8

is only a screening tool. Actual files also contain scales, zero points, metadata, tensors deliberately kept at higher precision, embeddings, vision projectors, and sometimes an MTP draft head. A “2-bit” filename therefore need not occupy two bits per published parameter. The preview-based DeepSeek 96.531GB MLX artifact divided by 284B parameters is roughly 2.72 bits/parameter; the custom ~81GiB GGUF is roughly 2.44. The current 0731 OptiQ card reports 2.43 achieved bits/weight, but streams experts from SSD. None of those figures measures quality.

Then add, separately:

  1. runtime graph, kernel, allocator, and temporary buffers;
  2. KV cache at the actual cache precision and architecture;
  3. target context and batch/concurrency;
  4. OS, display, and shared-memory reservations;
  5. fragmentation and a safety margin;
  6. a multimodal projector and image-token cache where relevant.

For MoE, total stored parameters determine capacity; active parameters per token influence compute. gpt-oss-120b still needs space for roughly 117B parameters even though about 5.1B are active per token. Conversely, a 27B dense model touches its full dense stack every token. Active count is not a license to discard the inactive experts.

Formats are not interchangeable:

  • GGUF is the dominant llama.cpp-family container for CPU, CUDA, Metal, and Vulkan offload. Q4_K_M is a mixed four-bit family, not exactly four bits for every tensor.
  • MLX artifacts target Apple Silicon’s unified-memory/Metal stack. A conversion being small enough does not make it runnable elsewhere.
  • AWQ, GPTQ, and EXL2 are weight-only families used mainly on discrete GPUs; kernel/model support and calibration differ.
  • NVFP4 is NVIDIA’s four-bit floating-point format with scaling metadata and hardware-specific fast paths. An NVFP4 repository is not evidence that a pre-Blackwell GPU will execute optimized NVFP4 kernels.
  • DQ means dynamic or mixed quantization in the cited community artifacts; inspect the tensor map rather than reading the filename as a uniform bit rate.

Quality evidence is the weakest part of this market. None of the aggressive DeepSeek or new GLM community quants in the recommendations has a strong public task-level comparison against the publisher checkpoint. Treat speed and fit as proven only where stated; treat retained intelligence as unresolved.

NVIDIA versus Apple Silicon versus compact clustered systems

Platform What it does best What the spec hides Best fit in this guide
RTX 3090 / discrete NVIDIA Low-latency decode for models that fit; broad CUDA and llama.cpp support; replaceable hardware VRAM is separate from host RAM; 24GB is a hard capacity wall for full GPU residence; used thermals/warranty 9B–27B low-bit models, coding/chat, fastest value tier
Ryzen AI Max / shared x86 APU Large single-address-space capacity at ~$4K; compact, serviceable storage/platform CPU and GPU share 256GB/s; Windows offers a 96GB dedicated-VRAM maximum preset while Linux can configure higher; total usable remains below 128GB; backend maturity varies 70B–120B-class low-bit MoE/dense capacity when CUDA is not required
DGX Spark / GB10 Compact 128GB CUDA appliance; supported NVIDIA OS/container path; built-in high-speed NICs 273GB/s constrains decode; Arm64 compatibility; “1 PFLOP” says little about tokens/s Large official NVFP4 models and CUDA development; custom frontier experiments
M5 Max / Apple unified memory 614GB/s, 128GB, battery/display, strong MLX/Metal path No CUDA, all core components soldered, OS uses the same pool Portable large-model inference and exact MLX artifacts
Dual Spark Distributed capacity, high aggregate throughput, direct 200GbE/RoCE Two local pools, not one; link is 25GB/s wire rate before overhead; software complexity dominates One measured TP=2 workload that cannot run acceptably on one node

No platform wins every axis. Apple provides the best portable capacity/bandwidth balance; a used 3090 provides the best low-budget decode ecosystem; Framework offers the cheapest clean 128GB capacity; Spark buys NVIDIA software compatibility at a bandwidth penalty; multiple Sparks buy a distributed systems project.

Multi-GPU and multi-node realities

Two accelerators do not automatically provide twice the usable memory or performance. The runtime must implement tensor, pipeline, expert, or layer parallelism for the exact model and quant. Every token can require cross-device collectives, so interconnect latency and bandwidth matter alongside local memory bandwidth.

For two Sparks:

  • each node has 128GB local coherent memory and 273GB/s local bandwidth;
  • the direct ConnectX-7 path is up to 200Gb/s Ethernet/RoCE, or 25GB/s wire rate per direction before overhead;
  • the official bundle supplies the cable, but the model server still needs RoCE/NCCL and a compatible container;
  • the Tony DeepSeek recipe uses TP2/PP1 without expert parallelism; the separate Elsung setup uses TP2 plus expert parallelism; neither creates magical unified pooling;
  • a failure takes down the sharded service unless the operator builds a restart/fallback plan;
  • 2 × 240W PSU capacity is not the same as 480W measured wall draw.

The public dual-Spark result is useful because it reports topology, runtime, context, concurrency, and long-context TTFT. It still remains community evidence with a custom image and a previously diagnosed stability problem. The repository says a prefix-cache host-memory leak was fixed in its production-v2 image and reports a subsequent stable run; an editorial lab should reproduce that before calling the system production-ready.

Ordinary multi-GPU PCs have similar caveats. PCIe peer-to-peer support, slot spacing, bifurcation, NUMA placement, power delivery, and cooling can matter more than aggregate VRAM on a shopping list. NVLink exists only on specific older/professional cards and does not, by itself, create a flat model-memory pool.

Power, noise, portability, and ownership cost

The published power figures are not directly comparable: 350W is a GPU board rating, 140W is the GB10 TDP, and 96W is an included laptop adapter. None is a measured whole-system inference draw.

  • Used 3090 workstation: potentially the loudest and hottest option; the 350W card sits in an older workstation with a 950W PSU. Electricity, fan noise, and used GDDR6X health belong in the purchase decision.
  • Framework Desktop: 120W sustained/140W boost processor envelope in a 4.5L case, with a 400W PSU. Compact and serviceable, but the small PSU fan and shared-memory load need an independent noise test.
  • DGX Spark: 140W GB10 TDP and 240W PSU in a very small enclosure. One box is desk-friendly; two add network, cabling, and service complexity.
  • M5 Max MacBook Pro: the only complete portable system. Apple includes a 96W adapter for the 14-inch SKU, but long inference can consume battery quickly and the smaller chassis may throttle differently from the 16-inch model. No comparable wall-power/noise test was found.

Storage matters. Model collections can consume terabytes; Spark includes 4TB, while the comparison Mac and Framework include 2TB and the refurb includes 1TB. Back up model manifests, adapters, prompts, and environment files; cached weights are replaceable, experimental conversions may not be.

Ownership cost also includes failures and time. A week spent stabilizing a two-node container can cost more than months of cloud inference. Used hardware needs a return window and stress testing. Soldered-memory machines need enough capacity on day one because there is no DIMM upgrade later.

Recommendations by use case

  • Best-value private chat and coding: used RTX 3090 workstation. Target 9B–27B low-bit models and 8K–64K contexts. Qwen3.8 now has a pinned llama.cpp path, but keep a known-stable fallback while its quants and speculative runtimes are still moving quickly.
  • Largest clean model capacity near $4K: Framework Desktop 128GB. The current-checkpoint DeepSeek-0731 lane is now exact enough to test, while gpt-oss-120b MXFP4 remains the simpler stable starting point; validate quality, context, and power yourself.
  • CUDA development with more than 24GB: DGX Spark. Prefer official NVFP4 artifacts such as Mistral Small 4 for production-like evaluation; treat the new GLM-5.3-Flash Q2 result as an experimental test lane until quality and reproduction improve.
  • Portable research and large local models: M5 Max 128GB MacBook Pro. The independently documented DS4 run supports feasibility, but exact artifact quality and current-checkpoint performance still require evaluation. For a desktop purchase, wait for September 22 M5 Max Mac Studio tests.
  • Multi-user frontier serving in a lab: dual Spark only when the exact vLLM/DeepSeek path is the workload. At 256K, concurrency 32 achieved high aggregate throughput, but each live request cannot simultaneously consume the full context.
  • Long-document single-user analysis: favor enough memory first, then examine prefill. A 500K context that takes six minutes before generation is a batch job, not chat.
  • Fine-tuning: none of the inference recommendations implies comfortable full-model training. LoRA/QLoRA feasibility depends on model, optimizer state, activation memory, sequence length, and framework support; buy against a measured training recipe.
  • CPU-only: use it for small models, testing, or overnight batch jobs on hardware already owned. Buying a new CPU-only system for these model classes is poor value.

What not to buy

  1. A “$2,000 RTX 3090 PC” without a pinned card. Current used-card and host pricing puts the conditional market estimate around $2,450, but exact AIB dimensions, connectors, condition, and warranty determine whether it is actually complete.
  2. An RTX 5090 for capacity. It is extremely fast when 32GB is enough, but current market pricing is around the cost of a 128GB Spark. Buy it for speed, gaming, or training kernels—not because it solves large-model fit.
  3. A DGX Spark on the “1 PFLOP” line alone. Sparse FP4 peak compute is not decode speed. Small dense models can be slower than on a used 3090 because Spark has only 273GB/s memory bandwidth.
  4. One Spark for the current DeepSeek-V4-Flash-0731 checkpoint. Its official tree is 166.899GB; one node has 128GB shared with the OS/runtime. Only much more aggressive third-party variants of the older preview have a measured one-Spark path.
  5. Two Sparks described as “256GB unified memory.” They are two 128GB nodes connected by Ethernet/RoCE and require explicit distributed execution.
  6. Three Sparks for “GLM-5.3 3-bit.” The model and a 342.966GB Unsloth Q3 artifact now exist, but raw aggregate fit is not a validated three-node run: OS/runtime/cache headroom is thin, no loader/throughput/quality result exists, and the complete topology exceeds the budget.
  7. A system based on announced memory alone. Apple’s new Mac Studio has exact preorder prices but no shipping-machine local-AI test; AMD’s 192GB Ryzen AI Max+ PRO systems and NVIDIA’s RTX Spark Windows PCs still lack a verified shipping US configuration, complete price, and exact model run. Upcoming is not tested.
  8. An aggressive low-bit frontier quant without a quality gate. A file that decodes quickly can still erase the capability you bought the larger model to obtain.

FAQ

Is local AI cheaper than an API?

Only at sustained utilization or when privacy/offline control has independent value. Divide complete purchase price, electricity, failures, and setup time by useful tokens—not theoretical peak. For occasional frontier-model use, renting usually wins above the dual-Spark tier.

How much memory headroom should I leave?

There is no universal percentage. Reserve the measured OS/runtime baseline, calculate the architecture-specific cache at the target context and precision, add graph/scratch allocations, then leave fragmentation margin. As a shopping screen, 10–20% beyond weights is safer than zero, but it does not replace an exact load test.

Does a 128GB unified-memory computer provide 128GB for the model?

No. The OS, runtime, display, cache, and buffers use the same pool. Framework’s Windows setting tops out at 96GB dedicated VRAM, while Linux can configure higher; neither makes all 128GB available to model weights. macOS and DGX OS likewise require headroom.

Does active MoE parameter count determine memory use?

No. Total parameters normally determine stored weight capacity; active parameters influence per-token compute. DeepSeek-V4-Flash’s core is 284B total and 13B active, not a 13B download; the 0731 packaged checkpoint also carries a speculative module.

Is “2-bit DQ” exactly two bits per weight?

No. The checked 96.531GB preview artifact implies roughly 2.72 bits per published parameter and keeps selected tensors at higher precision. The current 0731 OptiQ artifact separately reports 2.43 achieved bits/weight and uses SSD expert streaming. File naming describes a method/family, not a uniform physical rate or a quality guarantee.

Can the RTX 3090 run Qwen3.8-27B at its full 262K context?

Not entirely in 24GB with the cited quant. The FP16 cache alone would be 16GiB at the full limit. A 32K FP16-KV target is calculated to fit with useful headroom, while a pinned run measured 64K using Q4_0 KV and MTP. Full 262K remains outside the recommended 24GB configuration.

Is the dual-Spark 500K result interactive?

Generation is interactive at roughly 32 tokens/s once it starts, but the cited 500K prompt took about six minutes to prefill. That makes it a long-running analysis job.

Should I buy Apple, NVIDIA, or AMD?

Buy the software path. Choose NVIDIA for CUDA breadth and discrete-GPU latency, Apple for high-bandwidth unified memory and MLX, and Ryzen AI Max for the lowest current cost of clean 128GB x86 capacity. The MacBook is the proven portable choice; the September 22 Mac Studio needs shipping-machine evidence. Exact model/backend support beats brand allegiance.

What about model quality?

This is a hardware/fit guide, not a leaderboard. Bigger parameter counts and higher prices do not guarantee better results for your tasks. Run a private evaluation set on the exact quant before purchasing around it.

Can I train these models locally?

Inference fit does not imply training fit. Full training requires weights, gradients, optimizer state, and activations; even adapters can exceed these systems at long sequence lengths. Follow an exact measured training recipe.

Research methodology

This is a weekly UPDATE run against the published source-audited baseline. Research began at 2026-09-04 12:30:00 PDT. Every recommendation was rechecked against current US pricing, exact model cards and repository revisions, downloadable artifacts, licenses, backend paths, memory headroom, and configuration-rich benchmarks. Sources were accessed on September 4, 2026 unless a row says otherwise.

The hierarchy was: model publisher and license; hardware manufacturer price/specification; runtime documentation; reproducible benchmark repository; independent test; current retail/used listing. Search snippets were used only to find a page, never as the final evidence when a direct source was available. Retail and used prices are snapshots, not promises.

No hardware was physically tested by this publication. “Officially specified” is not relabeled as hands-on. Vendor benchmarks remain vendor benchmarks. Community results remain community results even when the repository is detailed. Calculated fit is not a speed claim.

The review tested the five starting hypotheses directly:

Starting hypothesis Verdict Evidence-based replacement
~$2K RTX 3090 + Qwen3.8-27B Revise/support with caveats Complete system is now ~$2.45K. A pinned ggml-org Q4_K_M/llama.cpp path measures 37.39 tok/s baseline and 60.33 with MTP at a configured 64K context; used-hardware and quant-quality caveats remain.
~$4K one Spark + DeepSeek-V4-Flash “2.5-bit” Reject as stated Spark is $4,699. Current 0731 official artifact does not fit. A preview-based custom ~2.44 effective-bpw GGUF fits and independently measured 13.2 tok/s at 2K, falling to 7.94 at 262K, with unresolved quality.
~$7K M5 Max 128GB + DeepSeek “2.5-bit” Support hardware; split the model verdict Exact $6,699 system. The fast 34.1 tok/s mean is a preview-based custom DS4 artifact using ~115GB. Current 0731 has a 92.487GB/2.43-bpw SSD-streamed OptiQ artifact but no M5 Max speed or quality score.
~$8K dual Spark using NVFP4/6.5-bit DQ Reject price and quant premise Current official bundle is $9,449. Best evidence is a community TP=2 FP8 path at ~40 tok/s, not a verified 6.5-bit DQ recommendation.
~$12K three Spark + GLM-5.3 3-bit DQ Reject for now Four-Spark GLM-5.3 is now demonstrated, but costs $18,796 before interconnect; no validated three-Spark run exists. The $11,299 M5 Ultra Mac Studio is a serious September 22 candidate, not yet an evidence-backed buy.

Sources

Models, artifacts, and licenses

Hardware and current prices

Benchmarks and runtimes

What changed in this review

This weekly UPDATE preserves the original recommendation ladder while changing only evidence that moved:

  • moves the conditional used 3090 complete-system estimate to $2,433–$2,463 after the refurb host rose to $1,134 and the used-card median fell to $1,299 on eight qualifying sales;
  • makes DeepSeek-V4-Flash-0731 the Framework test lane after an exact September 2 vendor run measured 14.7–14.9 tokens/s baseline and 21.73–24.64 with DSpark;
  • records Apple’s September 22 Mac Studio launch: the $5,899 128GB M5 Max is a desktop-value watch item, while the $11,299 256GB M5 Ultra is the first serious single-box candidate inside the $12K tier; neither is recommended before exact tests;
  • records a four-Spark GLM-5.3 run at 27.2 single-stream tokens/s, while preserving the rejection of an unvalidated three-Spark path and noting the four-box hardware cost of $18,796 before interconnect;
  • expands GLM-5.3-Flash feasibility evidence across one and two Sparks without promoting the custom artifacts to a quality-proven buying case;
  • adds Qwen3.8-Flash-Next as an experimental watch item because its exact artifacts exist but current llama.cpp reports include a Framework decode cliff and a DGX Spark native-context abort;
  • records that the $4,699 single-Spark listing showed Add to Cart while the $9,449 dual-Spark bundle showed Out of Stock;
  • confirms the checked Qwen3.8-27B, DeepSeek, Mistral, gpt-oss, GLM-5.3, and GLM-5.3-Flash identities and material artifact sizes remain stable.