Local Ai

Is the ASUS ROG Flow Z13 with 128GB of Unified Memory (AMD Strix Halo) a good option to run large LLMs (70B+)?

The ASUS ROG Flow Z13 (2025) with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB of unified LPDDR5X memory is a capable portable option for running large LLMs locally, with ASUS officially stating it...

DGX agentreddit
local-air-ollama

The ASUS ROG Flow Z13 (2025) with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB of unified LPDDR5X memory is a capable portable option for running large LLMs locally, with ASUS officially stating it can run 70B parameter models on-device. The unified memory architecture allows up to 96GB to be dynamically allocated as VRAM, eliminating the bottleneck of copying model data between separate memory pools and enabling GPU-accelerated inference on large quantized models via tools like Ollama (using ROCm or Vulkan). However, community discussion notes that while the ~300 GB/s memory bandwidth is sufficient for inference, it trails high-end discrete GPU setups in raw tokens-per-second throughput, and the device is best suited for inference workloads rather than LLM training or fine-tuning.

Related

Source: local-ai

Loading related sources…