Install Ollama and run ollama run qwen3.5, which pulls the 9b model, a 6.6GB to 7.6GB download that reads text and images with a 256K context window. Pick the size by your Mac's memory: the 2b (2.7GB to 3.1GB) suits 8 GB, the 9b suits 16 GB and 24 GB, the newest Qwen 3.8 27b (18GB) is tight on 32 GB and comfortable from 48 GB, the Qwen 3.5 35b (22GB) fits 64 GB, and the 122b (81GB) needs 128 GB. Qwen 3.6 27b, tuned for agentic coding, has the same footprint as 3.8. All of them are Apache 2.0 licensed, and Ollama's library also lists MLX builds, tagged -mlx. Once it runs, point any app that speaks Ollama at it, or install Grux OS (MIT), whose model cookbook lists Qwen 3.5 and 3.8 by memory and can use them to act on your mail, calendar, files and shell.
brew install ollama
ollama run qwen3.5 # the 9b model, 6.6GB to 7.6GB
ollama run qwen3.5:2b # 2.7GB to 3.1GB, for 8 GB Macs
ollama run qwen3.8:27b # 18GB, for 48 GB or more
| Your Mac's memory | Qwen model | Download |
|---|---|---|
| 8 GB | Qwen 3.5 2b (the 4b is tight) | 2.7GB to 3.1GB |
| 16 GB | Qwen 3.5 9b | 6.6GB to 7.6GB |
| 24 GB | Qwen 3.5 9b comfortably | 6.6GB to 7.6GB |
| 32 GB | Qwen 3.8 27b is tight; 9b is easy | 18GB |
| 48 GB | Qwen 3.8 27b or Qwen 3.6 27b | 18GB to 19GB |
| 64 GB | Qwen 3.5 35b | 22GB |
| 128 GB | Qwen 3.5 122b, tight | 81GB |
The table applies the rule on the which LLM can my Mac run page: the GPU may use about three quarters of unified memory, and the model file should stay under about two thirds of that, leaving room for context. Sizes are from Ollama's Qwen 3.5, Qwen 3.6 and Qwen 3.8 pages, read on 5 October 2026.
Qwen 3.5 is the family with every size, from 0.8b to 122b, and the place to start. Qwen 3.6 comes as a 27b and a 35b and is tuned for agentic coding. Qwen 3.8 comes as a 27b only, the newest, with thinking on by default. All three read images as well as text. For a coding model sized to your Mac, see the best local LLM for coding page.
| Claim | Source |
|---|---|
| Qwen 3.5 sizes from 0.8b to 122b, 256K context, text and image, MLX builds | Ollama library, qwen3.5 |
| Qwen 3.8 27b at 18GB, thinking on by default | Ollama library, qwen3.8 |
| Qwen 3.6 27b and 35b, agentic coding | Ollama library, qwen3.6 |
| Apache 2.0 license | Qwen3.5-9B on Hugging Face |
| Grux's cookbook lists Qwen 3.5 2b, 4b, 9b, 35b and Qwen 3.8 27b by memory | Cookbook.swift |
ollama run qwen3.5 pulls by default. It is a 6.6GB to 7.6GB download and leaves room for its context.