Open-weight models run entirely on your Mac — reading code, editing files, running commands. Expert offload fits 35B-class models on a 16 GB MacBook, and nothing you write ever leaves the machine.
macOS Sequoia or newer · Apple Silicon · free
The open-weight models worth running are mixture-of-experts — and even quantized to 4-bit, they're ~21 GB of weights, which a resident runtime refuses to load on a 16 GB Mac. Moonlet keeps the experts a token actually needs in RAM and leaves the rest on the SSD — Qwen 3.6 35B runs in 10.4 GB, at 36 tok/s.
And there is no server side at all: no account, no API key, no token cost. Your code stays on the machine.
Every one is tuned and tested inside Moonlet before it ships. One-click download in the app — no Hugging Face account required.
One download. Pick a model on first launch, point it at a folder, start working.
Download for macOSBy downloading or using Moonlet you agree to the End-User License Agreement. See the third-party notices.
Bug reports, model requests, rough edges — we read all of it.