Back
Running local models is good now: a developer's setup on a 2022 Mac
SiTech AI Team3 წთ. საკითხავი

Running local models is good now: a developer's setup on a 2022 Mac

Engineer Vicki Boykis says local models have finally become good enough for agentic coding: on a 2022 M2 Mac with 64 GB of RAM they reach roughly 75% of frontier-model accuracy and speed, with every agent session sandboxed in Docker.

Engineer Vicki Boykis has worked with local models since they came out, and in a post published on 15 June 2026 she sums up the change: they are finally surprisingly good. Her test machine is a 2022 M2 Mac with 64 GB of RAM and 1 TB of storage.

She has run Mistral 7B, Gemma 3, OpenAI OSS-20B, Qwen 3 MoE and other Qwen variants such as Qwen 2.5 Coder across several setups: raw llama.cpp with Open WebUI, llama-cpp-python, Ollama, llamafiles and LM Studio.

Where local models stand now

Boykis marks the release of GPT-OSS as the turning point — the first model she stopped double-checking against an API model so often. Her personal metric is simple: do I have to verify this against an API model?

With the most recent releases in the Gemma 4 family she finally started doing agentic coding locally, with loops running at roughly 75% of the accuracy and speed of frontier models. Her default local model is gemma-4-26b-a4b in LM Studio. She has used it to refactor a notebook script into a repository of five or six modules, fix type hints for generics, proofread blog posts, write unit tests and bootstrap a two-tower recommendation model repo from a blank slate. The K-V cache grows to 64 GB of RAM, and these tasks would have been impossible for local models as recently as six months ago.

Running local agents today

Three pieces are needed: a local inference engine, an agentic harness and the model artifact, with the harness pointed at the local endpoint. Boykis uses Pi as the harness and LM Studio as the inference server, noting that going straight to llama.cpp would likely be faster.

For safety, every Pi session runs in a Docker container with permissions only for bash, so it cannot run Python code or browse the web. For the model she picked gemma-4-12b-qat — more recent, smaller and faster, without much sacrifice in accuracy.

What still does not work

Inference can be slow, context windows are small and limited by your own hardware, and early releases suffer from prompt template mismatches, though those are usually patched quickly. The ecosystem is much easier thanks to LM Studio and HuggingFace's Use This Model button.

In her view this is not ready for production software development quite yet. The benefit is transparency: you can watch token inference live, change the local context window and watch performance shift, adjust the system prompt and quantizations, and pit models against each other.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.