Large-model platform guide

Ryzen AI Max+ 395 PCs for local LLMs

Systems with 64 or 128 GB of unified memory can offer more model capacity than typical consumer GPUs. The crucial questions are how much memory the GPU can actually use and whether your runtime supports the Radeon 8060S.

Short answer

When does this platform make sense?

64 GB: strong for 20B to 32B

A useful capacity tier for quantized 32B models, longer contexts, and local coding or RAG workflows. The operating system and runtime consume part of the shared pool.

128 GB: room for 70B

The larger configuration creates enough room for many quantized 70B models. Whether a specific model and context fit depends on quantization, the UMA limit, and KV cache.

Capacity before peak speed

The Radeon 8060S shares memory with the CPU. That enables large models in a compact machine, but it does not replace a high-end discrete GPU when maximum token throughput is the priority.

128 GB of unified memory is not 128 GB of free VRAM. Always reserve memory for the operating system, runtime, context, and other applications.

Complete systems

Three 128 GB configurations for different priorities

These systems use Ryzen AI Max+ 395 or its PRO variant with Radeon 8060S. Storage, networking, ports, physical layout, warranty, and final price are therefore the deciding differences because the memory is not user-upgradeable.

Best for: large models in a compact PC

GMKtec EVO-X2

Ryzen AI Max+ 395 with 128 GB LPDDR5X unified memory and a 2 TB SSD. Read the configuration guide.

128 GB

Good match for

Quantized 70B models, larger local knowledge systems, and users who value model capacity more than maximum discrete-GPU speed.

Check before buying

Select exactly 128 GB / 2 TB, confirm the maximum UMA allocation, and verify the operating-system and runtime combination you plan to use. The soldered memory cannot be upgraded later.

CPU / graphics
Ryzen AI Max+ 395 / Radeon 8060S
Memory
128 GB LPDDR5X, soldered
Storage expansion
2× M.2 PCIe 4.0; 2 TB included
Networking
2.5GbE, Wi-Fi 7
Ports
2× USB4, HDMI, DisplayPort, SD reader
Size / power
193 × 186 × 77 mm; external 230 W adapter

Best for: a fast home-network model server

Beelink GTR9 Pro

Ryzen AI Max+ 395 with 128 GB unified memory, a 2 TB SSD, and two 10GbE ports. Read the configuration guide.

128 GB

Good match for

Quantized 70B models, a compact home model server, and workflows that move large model files or serve several local clients over 10GbE.

Check before buying

Confirm exactly 128 GB / 2 TB and your intended UMA allocation. Memory is soldered. Linux users should also verify the current Ethernet driver and Radeon runtime path before planning unattended operation.

CPU / graphics
Ryzen AI Max+ 395 / Radeon 8060S
Memory
128 GB LPDDR5X-8000, soldered
Storage expansion
2× M.2 PCIe 4.0; 2 TB included
Networking
2× 10GbE, Wi-Fi 7
Ports
2× USB4, HDMI, DisplayPort, SD reader
Size / power
compact chassis; external power adapter

Best for: professional support and Windows

HP Z2 Mini G1a

Ryzen AI Max+ PRO 395 with 128 GB unified memory, a 1 TB SSD, and a three-year workstation warranty. Read the configuration guide.

128 GB

Good match for

Quantized 70B models, professional local AI work, and buyers who value Thunderbolt 4, an internal power supply, and on-site support more than maximum included storage.

Check before buying

Confirm BN8E8UA#ABA with 128 GB / 1 TB, the intended UMA allocation, and your Radeon runtime. One terabyte is a limited starting point for a large model library, and the workstation package can carry a substantial price premium.

CPU / graphics
Ryzen AI Max+ PRO 395 / Radeon 8060S
Memory
128 GB LPDDR5X-8533, soldered
Storage expansion
2× M.2 2280; 1 TB included
Networking
2.5GbE, Wi-Fi 7
Ports
2× Thunderbolt 4, 2× mini-DP, USB-C/A
Warranty
3 years parts, labor, and on-site repair

* Paid link. We may earn a commission if you buy; your price is unchanged. As an Amazon Associate we earn from qualifying purchases. Recheck configuration, seller, and delivery details on the destination page.

Memory decision

64 or 128 GB?

Question64 GB128 GB
Typical focusquantized 20B–32B32B with ample reserve; many 70B quantizations
Long contextscalculate reserve carefullymore headroom, still model-dependent
Multiple local servicespossible, but pressure rises quicklybetter for RAG, embeddings, and parallel components
Later RAM upgradenot possiblenot possible
Recommendationwhen 70B is not essentialwhen large models drive the purchase

Before ordering

Seven checks that matter

1

RAM configuration

LPDDR5X memory is soldered. A 64 GB unit cannot be upgraded to 128 GB later.

2

UMA/VRAM allocation

Check the BIOS version and maximum GPU allocation. Advertised total memory is not automatically all available to the model.

3

Runtime and backend

llama.cpp lists HIP and Vulkan GPU backends. AMD's HIP documentation includes Ryzen AI Max+ 395, but your exact runtime release still needs to match.

4

SSD capacity

Model files, multiple quantizations, and embedding models can consume hundreds of gigabytes. A 2 TB SSD is the more comfortable starting point.

5

Sustained-load profile

Quiet, balanced, and performance modes use different power limits. Keep the vents clear for long responses or batch jobs; a quieter mode may be preferable in a living or sleeping area.

6

Ports and networking

Fast Ethernet and enough USB4/display outputs are useful for a NAS, a local model server, or multiple clients.

7

Returns and warranty

Confirm seller, included power adapter, delivery scope, and support conditions for the exact configuration.

Next comparison

Capacity or CUDA speed?

If 32 GB of VRAM is enough for your models, an RTX 5090 tower may be the better fit.

Open RTX 5090 guide