Back
MicroLLM Lab: seven tiny LLMs run inside the browser
SiTech AI Team3 წთ. საკითხავი

MicroLLM Lab: seven tiny LLMs run inside the browser

A new browser experiment, MicroLLM Lab, runs, benchmarks and compares seven small language models from 26M to 360M parameters on-device through WebGPU, with no servers, no accounts and no data leaving the machine.

A lab that runs in the browser

MicroLLM Lab, published at stateofutopia.com/experiments/microllmlab, is a browser experiment for running, benchmarking and comparing Small Language Models (SLMs) directly on the user's hardware. The formula is short: WebGPU, zero server, Q4 quantization. Inference is not offloaded anywhere, no account is needed, and no prompt or user data leaves the device. The project is a WebGPU port, Q4 packer and multi-model harness inspired by yangqi0's PetitGPT research.

The page argues that frontier models such as GPT-4 or Claude cost millions to serve and add network latency, while compact models between 25M and 360M parameters act as a fast edge layer, classifying queries and extracting intent in milliseconds. Because the work runs on the client GPU, concurrency carries no API bill, and first-token latency is under 10 ms.

Seven models, all quantized to Q4

The catalog holds seven checkpoints. PetitGPT research-v1 (124.6M parameters, Apache-2.0, by yangqi0) is the reference baseline, trained on a single RTX 4090 with roughly 13B tokens. SmolLM2 135M Instruct and SmolLM2 360M Instruct come from Hugging Face's Smol Models Research and were pretrained on about 2 trillion tokens before instruction tuning. L20-Edu 135M (AliceYin) follows the same Llama-style recipe but was trained on one NVIDIA L20 with about 13B tokens. Two MiniMind models from jingyaogong, at 104M and 26M parameters, plus GPT-2 124M in its nanoGPT shape complete the set.

Every file is Q4 quantized: weights shrink from 16-bit floats to 4 bits per parameter, cutting memory use by about 75%. On-device weights range from roughly 16 MB for the 26M MiniMind2-Small to about 226 MB for the 360M SmolLM2, and models above 100M parameters fit in 50-84 MB of browser memory. WebGPU, the W3C standard that runs compute shaders on Apple Metal, DirectX 12 or Vulkan, does the work and falls back to WASM and then JavaScript. Loading a model caches it in the browser's private IndexedDB.

Benchmarks and certificates

The benchmark mode runs objective, deterministic checks, regular expressions and exact-token tests rather than writing-quality scores. A 135M model is allowed to fail, the page notes, because that is the measurement. The suite can run on the active model or on every loaded model, and a sustained 256-token speed test is available too.

A Compare & Certificate tab turns the figures into a shareable PNG certificate. A custom-eval panel lets users write their own benchmark in JavaScript: the code is eval'd in the page's origin and each check runs against the model's decoded text. A runtime panel reports backend, tokens per second, JS heap, GPU buffers and IndexedDB usage. The project can also be downloaded as a 589 MB zip, but it must be served over HTTP.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.