Back
How to set up a local coding agent on macOS with Gemma 4 and llama.cpp
SiTech AI Team3 წთ. საკითხავი

How to set up a local coding agent on macOS with Gemma 4 and llama.cpp

Kyle Howells documents a local coding-agent stack for macOS: llama.cpp with Metal, Gemma 4 26B-A4B, an MTP speculative draft model and Pi. Generation speed rose from 58.2 to 72.2 tokens per second.

Why run a coding agent locally

Kyle Howells lost his internet connection a few times and found himself stranded without a coding agent, so when he saw the “Gemma 4 now runs 2x faster with MTP” update for Multi-Token Prediction, he decided to run one on his own Mac.

His requirements were practical: fast enough to actually use, an OpenAI-compatible API so other tools could talk to it, and preferably support for screenshots and images, so he could feed the agent a picture of what it had just built.

The stack he settled on

The final setup: llama.cpp built with Metal on macOS, Gemma 4 26B-A4B in GGUF format, a Q8 MTP draft model for speculative decoding, the Gemma 4 multimodal projector, and Pi as the terminal coding agent. It was tested on an Apple M1 Max with 64 GB of unified memory running macOS 15.7.7.

The main model is gemma-4-26B-A4B-it-UD-Q4_K_XL.gguf from Unsloth’s Hugging Face repository, about 16 GB; with the draft head and projector the folder grows to roughly 17 GB. Benchmarks used one prompt asking for a compact Python function that parses a unified diff, generating about 128 tokens per run.

MTP lifts generation speed by about 24%

Baseline llama.cpp with Metal acceleration reached 298.0 prompt tokens/second and 58.2 generation tokens/second — usable, but slow for an agent making many tool calls. Adding the Q8 MTP draft model as a speculative decoder pushed generation to 72.2 tokens/second. Howells swept the draft length from 1 to 6: three was fastest at 72.2, two was nearly identical at 72.0, and larger values dropped to 63.7 and 61.2. Prompt processing stayed flat at about 296 tokens/second, giving a 1.24x speedup overall.

He also compared MLX-LM, expecting it to win on Mac hardware: the best MLX run reached 45.8 tokens/second, and community 4-bit variants were slower still at 43.9 and 38.1. A test of Gemma 4 MTP through gemma-4-swift-mlx failed because the 26B 4-bit checkpoints did not match the loader’s expected weight keys.

Images, setup and the Qwen alternative

For screenshots the llama.cpp server needs the Gemma 4 multimodal projector, since only the 12B model is natively multimodal; with it loaded the server advertises multimodal support and Pi can pass images through. Pi’s model entry also had to list text and image as inputs, otherwise image tool output was not sent to the model. Re-running the text benchmark with the projector showed no slowdown.

The build is standard: install cmake, git, tmux and Python 3.11 with Homebrew, compile llama.cpp with Metal and Accelerate enabled, download the model, the MTP draft and mmproj-BF16.gguf, then start llama-server on 127.0.0.1:8080 with --spec-type draft-mtp, --spec-draft-n-max 3 and a 65,536-token context. That provides an OpenAI-compatible endpoint at /v1, which Pi points to from its models.json.

His conclusion: the MTP draft model is worth using, taking Gemma 4 from 58.2 to 72.2 tokens per second while keeping the setup simple. He notes a common suggestion to use Qwen3.6 35B-A3B instead: the benchmarks he found suggest Qwen is a better coding agent, but it is slower — about 55 tokens/second in the same configuration against 72 for Gemma 4.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.