Back
Needle 2: a 14 MB agentic LLM for phones, wearables and robots
SiTech Team2 წთ. საკითხავი

Needle 2: a 14 MB agentic LLM for phones, wearables and robots

Cactus Compute has released Needle 2, an open 45-million-parameter model that fits in a single 14 MB binary, runs in about 28 MB of RAM and focuses on tool calling and structured extraction.

Cactus Compute has released Needle 2, an open model with 45 million parameters for tool calling, device use and structured extraction. The entire model ships as a single 14 MB binary that runs in about 28 MB of RAM; Cactus says it was built on the company's Simple Attention Network, compressed to 2-bit CQ2 quantisation with Cactus Quants and packaged with its own inference engine.

What it is for

Rather than open-ended chat, Needle 2 is aimed at mapping a user's sentence onto functions exposed by an application — deciding which function to call and filling in the arguments — and at turning messy text into typed fields. The company's argument is that this framing requires no world knowledge and no prose generation, which is why, it says, 45 million parameters are enough where general chat needs models billions of parameters large. Cactus positions the model for smart homes, robots, phones, wearables, AR glasses and cars: hardware with no GPU or NPU and only a few dozen megabytes of RAM.

Benchmarks

On tool-call and mobile device-use benchmarks, Cactus says Needle 2 trades wins with much larger small models — FunctionGemma 270M, LFM2.5 230M and Apple FM — despite being 5 to 70 times smaller and running at 2 bits against their 16-bit precision. The comparison uses Google's Mobile-Actions evaluation split (961 rows), with Needle 2 measured end-to-end through the shipped binary.

Performance figures listed by the company: about 500 tokens per second decoding on a Raspberry Pi 5, with more than 800 tokens per second prefill; 400 to 1,500 tokens per second on VR headsets such as Meta Quest 3S and Apple Vision Pro; and 300 to 700 tokens per second on sub-$200 phones like Samsung's A-series. With peak session RAM around 28 MB, the model also runs on newer microcontrollers such as the ESP32-S3.

Why cheap hardware

Cactus frames its bet as a bet on cheap devices: the company notes there are more than 21 billion IoT devices against roughly 1.5 billion PCs, that most phones in emerging markets ship under $200, and that roughly four in five edge devices cost less than $200. Needle 2 is an open release, and the announcement walks through example flows such as smart-home commands, multi-step robot actions, form filling, extraction from documents and routing across a set of tools.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.