Moonlet app icon

Your local coding agent.

Open-weight models run entirely on your Mac — reading code, editing files, running commands. Expert offload fits 35B-class models on a 16 GB MacBook, and nothing you write ever leaves the machine.

macOS Sequoia or newer  ·  Apple Silicon  ·  free

MoE offload

Run the 21 GB model anyway.

The open-weight models worth running are mixture-of-experts — and even quantized to 4-bit, they're ~21 GB of weights, which a resident runtime refuses to load on a 16 GB Mac. Moonlet keeps the experts a token actually needs in RAM and leaves the rest on the SSD — Qwen 3.6 35B runs in 10.4 GB, at 36 tok/s.

And there is no server side at all: no account, no API key, no token cost. Your code stays on the machine.

Weights in memory — measured on the shipped configs
16 GB Mac
Qwen 3.6 35B · resident20.7 GB
out of memory — won't load
Qwen 3.6 35B · Moonlet offload10.4 GB
36 tok/s
Gemma 4 26B · Moonlet offload8.8 GB
40 tok/s
Peak RAM during decode, 4-bit quants, Apple Silicon.

Supported models

Every one is tuned and tested inside Moonlet before it ships. One-click download in the app — no Hugging Face account required.

Qwen 3.5 / 3.6 Alibaba Best agentic quality of the five — the default.
Gemma 4 Google Strong generalist for code and conversation.
Nemotron 3 NVIDIA Dense and predictable, with disciplined tool calls.
Mellum 2 JetBrains Code-focused; a custom architecture Moonlet supports natively.
LFM 2.5 Liquid AI Small and quick, with RAM to spare.

Get Moonlet

One download. Pick a model on first launch, point it at a folder, start working.

Download for macOS
Apple Silicon Macs Apple Silicon (M1+) 16 GB RAM recommended
Feedback

Tell us what broke. Or what didn't.

Bug reports, model requests, rough edges — we read all of it.