Skip to main content

Blog

Guides to running large language models on unified-memory machines: Mac Studio with MLX, NVIDIA DGX Spark with CUDA, memory sizing and the economics of renting.

  • DGX Spark

    Mac Studio vs DGX Spark: MLX or CUDA?

    5 min read

    Two desktop machines with 128 GB of unified memory, built on very different bets. How the Mac Studio M5 Max and the NVIDIA DGX Spark compare on bandwidth, compute, software and tooling, and when to pick which.

    • Mac Studio
    • DGX Spark
    • MLX
    • CUDA
  • DGX Spark

    NVIDIA DGX Spark and the GB10, explained

    5 min read

    What's inside NVIDIA's desktop AI computer: the GB10 Grace Blackwell superchip, 128 GB of LPDDR5x at 273 GB/s, up to 1 PFLOP of FP4, the CUDA software stack, and which models fit.

    • DGX Spark
    • GB10
    • CUDA
    • Unified Memory
  • Mac Studio

    Rent or buy? The economics of a Mac Studio by the day

    5 min read

    A 128 GB Mac Studio is a big purchase. Here's the break-even against renting one at €20 a day, the costs people forget, and the jobs where renting simply makes more sense.

    • Mac Studio
    • Economics
    • Rent vs Buy
  • Mac Studio

    M5 Max vs M5 Ultra: sizing memory for local LLMs

    5 min read

    How to work out whether a model fits in 96, 128 or 256 GB of unified memory — weights, KV cache and headroom — and what the M5 Ultra's 1.2 TB/s buys you over the M5 Max's 614 GB/s.

    • Mac Studio
    • M5 Max
    • M5 Ultra
    • Memory Sizing
  • Mac Studio

    Running Qwen3 and Llama with MLX on a Mac Studio M5 Max

    5 min read

    From a fresh Mac Studio to a working model server: installing mlx-lm, choosing a model and quantisation, generating, chatting, serving an OpenAI-compatible API, and what speed to expect.

    • Mac Studio
    • MLX
    • Qwen3
    • Llama