Local Ai
Is the ASUS ROG Flow Z13 with 128GB of Unified Memory (AMD Strix Halo) a good option to run large LLMs (70B+)?
The ASUS ROG Flow Z13 (2025) with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB of unified LPDDR5X memory is a capable portable option for running large LLMs locally, with ASUS officially stating it...
The ASUS ROG Flow Z13 (2025) with AMD Ryzen AI Max+ 395 (Strix Halo) and 128GB of unified LPDDR5X memory is a capable portable option for running large LLMs locally, with ASUS officially stating it can run 70B parameter models on-device. The unified memory architecture allows up to 96GB to be dynamically allocated as VRAM, eliminating the bottleneck of copying model data between separate memory pools and enabling GPU-accelerated inference on large quantized models via tools like Ollama (using ROCm or Vulkan). However, community discussion notes that while the ~300 GB/s memory bandwidth is sufficient for inference, it trails high-end discrete GPU setups in raw tokens-per-second throughput, and the device is best suited for inference workloads rather than LLM training or fine-tuning.
Related
- Mac mini M4 48GB
- Recommended Model for a 4060ti 8gb and 16gb ram
- How do I know if an AI model could work locally on my computer?
- Listening to @alexocheema from @exolabs talking about running LLMs locally at @aiDotEngineer @swyx 18 months before we get a SOTA LLM at hom…
- Use the Same Model Across Ollama, LM Studio, Jan, and your Favorite Local AI Apps
Source: local-ai