NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
NVIDIA announced a 64GB unified-memory configuration of DGX Spark, available October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI, starting at $4,999. It keeps the GB10 Grace Blackwell Superchip, DGX OS and the full NVIDIA AI software stack, and runs models of up to 100 billion parameters fully on device. Two units can be clustered through NVIDIA Sync Cluster Assistant to pool 128GB of memory and support up to 200 billion parameters. In NVIDIA's Qwen 3.8 27B test, two clustered 64GB systems delivered up to 1.7x the performance of one system.
On October 2, 2026, NVIDIA announced on its blog that DGX Spark will gain a 64GB unified-memory configuration. The new version goes on sale Friday, October 23, through six manufacturer partners (Acer, ASUS, Dell, Gigabyte, HP and MSI), starting at $4,999. The same post introduces NVIDIA Sync Cluster Assistant, a tool for joining two DGX Spark units into a cluster. NVIDIA's framing is simple: as AI agents move from experiments into everyday development, increasingly capable open models are shrinking to fit on more devices, so builders have more to run locally. What was announced DGX Spark combines NVIDIA Grace Blackwell compute, unified memory, NVIDIA ConnectX-7 networking and a CUDA-accelerated AI software stack in one system. NVIDIA positions it as a complete local platform for agents, inference, fine-tuning, data science and edge development. The new 64GB configuration is available exclusively from manufacturer partners. It keeps the GB10 Grace Blackwell Superchip, DGX OS and the full NVIDIA AI software stack, the same as the 128GB model. NVIDIA says it supports models of up to 100 billion parameters, plus the agentic applications built on them, fully on device and without cloud dependency.
Out-of-the-box software is the other selling point. The NVIDIA Agent Toolkit, CUDA-X AI libraries, Nemotron open models, and popular runtimes such as Ollama, vLLM and PyTorch with CUDA are supported from day one. NVIDIA says developers can go from power-on to running models in minutes. Blender is among the first major creator application providers to support the platform, with a prebuilt, downloadable installer coming soon. How the two-unit cluster works Every DGX Spark ships with a built-in ConnectX-7 NIC. Two units can connect directly with a QSFP cable. Their memory pools to 128GB, model support rises to up to 200 billion parameters, memory bandwidth doubles, and performance rises by up to 1.7x. The post also describes the link as a 200 GbE fabric. The feature that lowers the barrier is Cluster Assistant inside the NVIDIA Sync app. It detects connected units, validates device configuration and configures the ConnectX-7 network, so developers do not have to set up networking and nodes by hand. Every node runs the same NVIDIA software stack, so the software environment does not need reconfiguring when you move from one unit to two. One caution: the 1.7x figure comes from NVIDIA's own Qwen 3.8 27B test and is stated as an upper bound. It sits below the theoretical 2x, which suggests some cost for communication between nodes. That reading is our interpretation, not a conclusion stated in the post.
Workflow changes: Model Launcher and three use cases The post also previews NVIDIA Sync Model Launcher, due at the end of the month. With a few clicks, developers can download and launch Qwen3.8 27B on a single DGX Spark or on a cluster. NVIDIA Sync configures the model to run across the connected devices and makes it reachable from the user's laptop. The launcher also sets up OpenCode to use the model, so developers can start coding in their browser. NVIDIA gives three workflow examples. First, run an AI agent around the clock: keep a coding or research agent running on DGX Spark, ready to review code, analyze documents or carry out multistep tasks, with a cluster adding capacity for larger models, longer context windows or several agents at once. Second, power AI apps on your everyday PC: run a language or image-generation model on DGX Spark while an agent or creative application runs on a laptop or desktop, so inference does not tie up the PC. Third, scale when the work grows: when a task outgrows one unit, two 64GB systems connected over the 200 GbE fabric pool their memory to 128GB, and the same workflow runs without reconfiguring the software environment. What it means for developers and the industry The following is our analysis, not NVIDIA's wording. First, price. A $4,999 starting point gives individual developers, researchers and small teams an easier entry, and they can buy one unit now and a second later. The post gives no price for the 128GB model, so the price gap between the two cannot be read from the source. Second, privacy and cost. Running locally means models and private data need not leave the device, and there is no cloud instance to rent for every task. That appeals to teams that handle sensitive code and documents. Third, ecosystem pull. Shipping with NVIDIA's agent toolkit, CUDA-X libraries and Nemotron models, while also supporting open runtimes such as Ollama and vLLM, lowers the cost of getting started and reinforces the CUDA ecosystem at the same time. Several cautions apply. Memory capacity decides how large a model and how long a context one unit can hold. The post says up to 100 billion parameters but does not state the quantization level or context length behind that claim. The 1.7x result covers a single 27B model, and other models and workloads may scale differently. Model Launcher and the Blender installer are both still listed as coming soon, so real-world experience after launch remains to be checked. Getting started and outlook NVIDIA's starting steps are short. Download a supported inference framework: llama.cpp, Ollama, vLLM or LM Studio. Download the recommended local model for your workflow. To scale to two units, connect them through their ConnectX-7 ports and launch NVIDIA Sync Cluster Assistant, which configures the network and routes workloads automatically. For agentic AI playbooks, the NemoClaw, OpenClaw, Hermes Agent and OpenShell pages on build.nvidia.com are the places to look.
The signal from this launch is clear. Local AI is being packaged less as a hobbyist toy and more as a development platform with a defined growth path, where single-unit capacity plus a two-unit cluster lets users start small and grow. The next things to watch are real pricing and supply after the October 23 launch, independent third-party benchmarks, and whether Model Launcher ships on time at the end of the month. If those pieces hold up, the 64GB DGX Spark could become a common starting point for local agent development.