Back
Seven ESP32-S3 boards run a 1.58-bit language model as one cluster
SiTech AI Team3 წთ. საკითხავი

Seven ESP32-S3 boards run a 1.58-bit language model as one cluster

A GitHub project splits a quantized Qwen2-0.5B into 24 layers across six ESP32-S3 compute nodes and one master, chained over SPI. Ternary weights keep each node's firmware inside 16 MB of flash.

A project published on GitHub as ESP32s3-LLM-Cluster shows how seven ESP32-S3 boards serve one small language model together. Six of them are compute nodes and the seventh is the master, which runs the BPE tokenizer, the embedding lookup, the final normalization and the LM head. The base model is Qwen2-0.5B, pruned and quantized for microcontrollers.

The ESP32-S3 boards of the cluster

Each compute node holds four of the model's 24 transformer blocks. The boards do not talk to each other directly: they are wired into a daisy chain, so a hidden state leaves the master, passes through every node in order and returns for the last layer.

1.58 bits per weight

Every linear layer uses BitNet-style ternary quantization: a weight can only be -1, 0 or 1. Three values fit in two bits, so four weights are packed into a single byte, and that is what lets the model fit on such small chips. The embedding table is kept separately in INT4.

According to the README, one layer takes about 3.82 MB and the four layers on a node about 15.3 MB, inside a 16 MB flash partition. The master needs roughly 14 MB for INT4 embeddings plus 64 KB for the final norm. The context window is 512 tokens, with the KV cache kept in each node's PSRAM.

An SPI daisy chain

Each board carries two SPI channels: one transmits (CS 4, MOSI 5, CLK 7) and one receives (CS 15, MOSI 16, CLK 18). Node 1's transmit channel feeds node 2's receive channel, and the last node sends the result back to the master. Separate lines handle reset and readiness: the master's GPIO 1 resets the nodes, and GPIO 3 brings back a signal from the end of the chain that all boards are up. A status LED sits on GPIO 8.

Power and limits

The author reports about 1.17 W at idle (5 V, 0.23 A) and roughly 1.53 W while inferencing. Each node needs about 1.3 seconds for its own part of the computation, so latency grows linearly as boards are added: one ESP32-S3 covers four layers.

The project is candid about its weak points. The quantization-aware training script was only partially trained, the loss settles near 8.0, and the model often emits random tokens or repeats one word under greedy sampling. The README also warns that if the 2-bit packing order in the Python tooling does not match the unpacking code in C and assembly, every weight is silently corrupted. The code is MIT-licensed, and the credits name Microsoft's BitNet and two earlier ESP32 projects as inspiration.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.