How do I run an LLM locally on my Mac?

Install Ollama, pull a model sized for your Mac's memory, and run it: three commands, and nothing you type has to leave the machine. Ollama (MIT) installs from its website or with Homebrew, runs models through llama.cpp, and answers on a local API. As a rule of thumb from Jan's published requirements, 8 GB of memory handles models of about 3B parameters, 16 GB about 7B and 32 GB about 13B. A MacBook or a Mac mini follows the same rule, because memory is what decides the size. If you prefer a window to a terminal, LM Studio does the same job in a desktop app and can use Apple's MLX. Once a model runs you can chat with it in the terminal, point other apps at it, or install Grux OS (MIT), which finds Ollama by itself and uses the model to act on your mail, calendar, files and shell.

Grux OS 3.0.0 · last checked 2026-09-30 · generated from the shipping release

The three steps

  1. Install Ollama. Download the Mac app from ollama.com, or run brew install ollama.
  2. Pick a model that fits your memory. Use the table below, or let the which LLM can my Mac run page work it out from your chip and RAM.
  3. Run it. ollama run downloads the model the first time and opens a chat in the terminal.
brew install ollama
ollama run llama3.2:3b
# optional: let the model act on your Mac
brew install --cask dotcomjack/tap/grux

How big a model your Mac can run

MemoryModel size, roughly
8 GBAbout 3B parameters
16 GBAbout 7B parameters
32 GBAbout 13B parameters

These are the minimums Jan publishes for a decent experience on macOS. Quantized models fit in less, and Macs with more memory run larger ones.

Other ways to run an LLM on a Mac

Where to check this

ClaimSource
Ollama installs on macOS from a script or a downloadable app, and ollama run chats with a modelOllama README
8 GB for 3B models, 16 GB for 7B, 32 GB for 13B on macOSJan README, System Requirements
llama.cpp downloads and runs a model straight from Hugging Facellama.cpp README
MLX LM installs with pip and generates text on Apple siliconMLX LM README
LM Studio runs GGUF and MLX modelsLM Studio docs
Grux OS installs with HomebrewGrux README, Install

Questions

How much RAM do I need to run an LLM on a Mac?
As a rough guide, 8 GB runs models of about 3B parameters, 16 GB about 7B and 32 GB about 13B. More memory lets you run larger and better models.
Can a Mac mini run an LLM locally?
Yes. A Mac mini with Apple silicon runs local models the same way a MacBook does; its memory decides how large a model it can hold.
Do I need the internet to run an LLM on my Mac?
Only to download the model. After that it runs on your Mac with no connection, and nothing you type leaves the machine.
Download Grux 3.0.0 Read the source

Free, MIT licensed. macOS 14 or later, Apple silicon. 23.7 MB, signed and notarized by Apple. No account, no server, no subscription.