A 1-bit LLM now runs on hardware you can buy for pocket money: a hobbyist BitNet ESP32 cluster splits a pruned Qwen2-0.5B model across seven ESP32-S3 boards, while Cactus Compute’s 2-bit “Needle” tool-calling model, shipped at 8–29 MB, climbed GitHub’s trending list. Here is what ternary weights are, and what they really buy embedded engineers.
The news: a 1-bit LLM on seven ESP32-S3 boards
On 26 September 2026, developer Low Zi Hong published the ESP32s3-LLM-Cluster repository: a distributed pipeline inference engine running a “1.58-bit (BitNet)” language model on seven ESP32-S3 boards linked by an SPI daisy-chain. One board is the master; it runs the BPE tokenizer, an INT4 token embedding (about 14 MB in flash), the final RMSNorm, the LM head and greedy sampling. Six compute nodes each run four transformer blocks, covering all 24 layers of a vocabulary-pruned Qwen2-0.5B.
According to a detailed write-up by Shamyl Bin Mansoor, the whole cluster draws about 1.5 W during inference and 1.15 W idle, and each node takes roughly 1.3 seconds per inference step. The ESP-IDF node firmware includes hand-written assembly (bitlinear_forward.S) for ternary multiply-accumulate.
| Parameter | ESP32-S3 BitNet cluster (as reported) |
|---|---|
| Base model | Qwen2-0.5B, vocabulary pruned from 151K to 32K tokens |
| Weights | 1.58-bit ternary linear layers, INT4 embeddings |
| Packing | 4 ternary weights per byte (2 bits each) |
| Per layer / per node | ~3.82 MB per layer, ~15.3 MB per node (fits 16 MB flash) |
| Context | 512 tokens, KV cache in PSRAM |
| Power | ~1.5 W inference, ~1.15 W idle (all 7 boards) |
| Licence | MIT |
The honest caveat comes from the project itself: the quantization-aware training script only partially trains the model, and the output is semi-coherent and can get stuck repeating a token. It is a proof of concept, not an assistant.
How BitNet ternary weights work
A normal LLM stores each weight as a 16-bit float. BitNet b1.58, introduced by Microsoft Research in February 2024, restricts every weight in the linear layers to one of three values: −1, 0 or +1. Three states carry log2(3) ≈ 1.58 bits of information, hence the name.

The practical consequences for a microcontroller are big:
- Memory shrinks about 8–10×. In practice you pack ternary values into 2 bits, four per byte, so a 0.5B model drops from roughly 1 GB at FP16 to around 100 MB.
- Multiplications disappear. Multiplying an activation by −1, 0 or +1 is a subtract, a skip or an add. The inner loop of a matrix-vector product becomes additions plus a per-tensor scale at the end.
- Training matters more than conversion. BitNet models are trained with ternary weights from the start (quantization-aware).
Microsoft’s BitNet b1.58 2B4T, released in April 2025, was the first open native 1-bit LLM at 2 billion parameters.
Needle: a 2-bit model that only does tool calls
The second story points in a more practical direction. Cactus Compute’s Needle repository describes an “automation foundation model for tiny devices”: 2-bit weights, a single 8–29 MB binary, built for tool calls, structured extraction and embeddings rather than open chat. The repo shows about 13.4k stars and 907 forks at the time of writing, and The AI Vibe reported it topping GitHub’s trending feed on 4 October 2026.
A few details from the repo and the Cactus devices guide are worth noting for engineers:
- Its quantisation format, “Cactus Quants”, runs at 2.125 bits per weight.
- Every depth from 2 to 20 layers is a trained, deployable model. The full 20-layer file is 29 MB; the 4-layer rung is about 8 MB.
- A byte-level grammar compiled from your tool schemas constrains decoding, so the output always parses, and each response carries a confidence score. An off-topic request returns an empty call list instead of a guess.
- Cactus reports decode speeds on a Raspberry Pi 5 from about 400 tokens/s (20 layers) to 4,000 tokens/s (2 layers). These are vendor numbers, not independent benchmarks.
One important reality check: despite the “microcontrollers” wording, the published platform folders are Linux (x86-64, ARM64, ARMv7, RISC-V, MIPS), Windows, Android, Apple platforms and WebAssembly. There is no ESP-IDF or bare-metal MCU build listed today. On embedded hardware, think Linux-class boards such as a Raspberry Pi or a MIPS camera SoC, not an ESP32.
Low-bit quantization beyond BitNet ESP32 hacks
Research is pushing the same direction. The EdgeRazor paper (arXiv:2605.04062, revised May 2026) describes mixed-precision quantization-aware distillation for small LLMs. Its authors report that a 1.88-bit Qwen3-0.6B beats state-of-the-art 2-bit baselines by 11.27 points, and that a 1.58-bit version cuts storage from 1.11 GB to 0.19 GB while decoding 15.16× faster than the 16-bit baseline. These are self-reported results.
At the truly tiny end, the open-source BitNetMCU project shows what low-bit weights do for classic TinyML. Using quantization-aware training, it passes 99% test accuracy on a 16×16 MNIST task on a CH32V003 RISC-V microcontroller with no multiply instructions, in 2 KB of RAM and 16 KB of flash. It added ternary (1.58-bit) inference export in January 2026.
| Project | Precision | Size | Target | Status |
|---|---|---|---|---|
| ESP32-S3 LLM Cluster | 1.58-bit + INT4 embeddings | ~15 MB per node, 7 boards | ESP32-S3 (ESP-IDF) | Proof of concept, semi-coherent output |
| Cactus Needle 3 | 2-bit (2.125 bpw) | 8–29 MB | Linux, mobile, WASM | Released, vendor benchmarks |
| EdgeRazor Qwen3-0.6B | 1.58–1.88-bit | 0.19 GB at 1.58-bit | CPU (llama.cpp) | Research paper |
| BitNetMCU | 1–8-bit incl. ternary | 16 KB flash | CH32V003 and similar MCUs | Working, MNIST-class models |
Practical limits of a 1-bit LLM on microcontrollers
Ternary weights solve the storage problem first. They do not solve everything:
- Latency. Pipeline parallelism means each token must pass through all six compute nodes in series. With about 1.3 s per node step, a token takes several seconds. Adding nodes adds layers but also adds time.
- Activations and KV cache. Weights shrink, but hidden states are still sent as FP32 vectors between boards and the KV cache needs PSRAM. SRAM, not flash, is often the real limit.
- Training cost. Ternary models only work well when trained or distilled with quantization in the loop. The cluster’s weak output is a training-compute problem, not a firmware one.
- Power integrity. Seven boards on one supply is a classic recipe for brownouts. If your nodes reset under load, read our guide on why ESP32 boards randomly reboot.
What this means for embedded engineers
The useful lesson is not “put ChatGPT on an ESP32”. It is that narrow, low-bit models are becoming a normal part of the embedded toolbox. A practical path to get ahead:
- Learn quantization-aware training on something small. Clone BitNetMCU, train its MNIST model, export the C header, and compare 8-bit, 4-bit and ternary accuracy against flash and RAM use.
- Write a ternary dot product yourself. Pack four weights per byte, decode with a 256-entry look-up table or bit masks, and accumulate with add/subtract only. Benchmark it against an INT8 loop on an ESP32-S3 with ESP-IDF.
- Use a tool-calling model where it fits. For voice or text commands on a Linux gateway, Needle’s pattern—fixed tool schemas, grammar-constrained JSON, confidence gating—is the right shape. Keep the MCU for sensing and actuation.
- Measure, do not trust headlines. Log tokens per second, peak RAM and energy per answer on your own hardware, and test accuracy on your own commands. Our post on why high accuracy can hide a bad model explains what to look for.
For a wider roadmap of skills, see Edge AI on microcontrollers: what embedded engineers should learn next.
Key takeaways
- BitNet’s ternary weights (−1, 0, +1) cut model storage roughly 8–10× and replace multiplies with adds.
- A seven-board ESP32-S3 cluster proves a 0.5B-class transformer can be split across $5-class MCUs, but output quality and speed are not product-ready.
- Cactus Needle shows the more practical trend: tiny 2-bit models built for one job, such as tool calls, running on Linux-class edge devices.
- Quantization-aware training is the core skill to learn next.
FAQ
What is a 1-bit LLM?
It is a language model whose linear-layer weights use one to two bits each. BitNet b1.58 uses ternary weights (−1, 0, +1), about 1.58 bits of information, usually stored as 2 bits per weight.
Can an ESP32 run an LLM on its own?
Not a useful general-purpose one. A ternary 0.5B model is about 100 MB, far beyond one ESP32-S3’s 16 MB flash, which is why the cluster splits it across six compute boards. A single ESP32-S3 is better suited to keyword spotting and small classifiers.
Does Needle run on microcontrollers like the ESP32?
Not today, based on Cactus’s published platform list, which covers Linux, Windows, Android, Apple platforms and WebAssembly. Its smallest builds target Linux-class boards such as a Raspberry Pi or MIPS camera SoC.
Is 2-bit quantization worse than 1.58-bit?
Not necessarily. Ternary weights are stored in 2 bits anyway, and true 2-bit schemes have four levels instead of three. Accuracy depends far more on whether the model was trained or distilled for low precision.
Want to build these skills properly? Explore the embedded systems, ESP32 and AI courses from Educational Engineering Team at eduengteam.com/.
