Install Ollama, pull a model sized for your Mac's memory, and run it: three commands, and nothing you type has to leave the machine. Ollama (MIT) installs from its website or with Homebrew, runs models through llama.cpp, and answers on a local API. As a rule of thumb from Jan's published requirements, 8 GB of memory handles models of about 3B parameters, 16 GB about 7B and 32 GB about 13B. A MacBook or a Mac mini follows the same rule, because memory is what decides the size. If you prefer a window to a terminal, LM Studio does the same job in a desktop app and can use Apple's MLX. Once a model runs you can chat with it in the terminal, point other apps at it, or install Grux OS (MIT), which finds Ollama by itself and uses the model to act on your mail, calendar, files and shell.
brew install ollama.ollama run downloads the model the first time and opens a chat in the terminal.brew install ollama
ollama run llama3.2:3b
# optional: let the model act on your Mac
brew install --cask dotcomjack/tap/grux
| Memory | Model size, roughly |
|---|---|
| 8 GB | About 3B parameters |
| 16 GB | About 7B parameters |
| 32 GB | About 13B parameters |
These are the minimums Jan publishes for a decent experience on macOS. Quantized models fit in less, and Macs with more memory run larger ones.
llama cli -hf with a model name downloads and runs it straight from Hugging Face.pip install mlx-lm, then mlx_lm.generate --prompt hello.| Claim | Source |
|---|---|
Ollama installs on macOS from a script or a downloadable app, and ollama run chats with a model | Ollama README |
| 8 GB for 3B models, 16 GB for 7B, 32 GB for 13B on macOS | Jan README, System Requirements |
| llama.cpp downloads and runs a model straight from Hugging Face | llama.cpp README |
| MLX LM installs with pip and generates text on Apple silicon | MLX LM README |
| LM Studio runs GGUF and MLX models | LM Studio docs |
| Grux OS installs with Homebrew | Grux README, Install |