LLM on a microcontroller: a small dev board generating text and images without a GPU or cloudLLM on a microcontroller: a small dev board generating text and images without a GPU or cloud

Generative AI has reached bare silicon: an open-source project now runs a 28.9-million-parameter LLM on a microcontroller, an ESP32-S3, at 9.88 tokens per second, while a second one draws faces with a diffusion model on a $1 RP2350. Neither chip runs Linux, has a GPU or talks to a cloud API.

What actually happened (and a correction on “289M”)

In late September 2026, several outlets reported that “a 289-million-parameter LLM” was running on an NXP FRDM-MCXN947 and that a diffusion model was producing 64 × 64 grayscale images on an STM32N6570-DK. We followed those reports back to their sources, and the details do not hold up as written.

The Open Source For You report (28 September) uses the same headline wording and image as a roundup The AI Vibe published on 24 September. That roundup describes two different projects: a 28.9M-parameter LLM on an ESP32-S3 (slvDev’s esp32-ai) and a latent diffusion transformer on a Raspberry Pi RP2350 (Tim’s pico-faces). The “289M” figure looks like 28.9M with the decimal point dropped. We found no public repository for an LLM on the MCXN947 or a diffusion model on the STM32N6570-DK. The OSFY piece also gives the STM32N6 an Arm Ethos-U55 NPU, but ST’s own product page says the STM32N6 uses ST’s in-house Neural-ART accelerator.

The verified projects are impressive enough on their own. Here is how they work.

How an LLM on a microcontroller fits: memory tiering, not magic

The esp32-ai README is clear about the problem. The ESP32-S3 used has 512 KB of SRAM, 8 MB of PSRAM and 16 MB of flash, and the model is 14.9 MB at 4-bit, far bigger than the fast memory. The trick is to place each part of the model in the memory tier that matches how often it is read.

Diagram of memory tiering for an LLM on a microcontroller: activations in SRAM, transformer core in PSRAM, embedding table in flash
  • SRAM (fast, tiny): activations and norm weights, which are touched many times per token.
  • PSRAM (medium): the dense transformer core and output head, read once per position.
  • Flash (large, slow): a 25-million-parameter embedding table. The model looks rows up in it instead of computing with it. About six rows, roughly 450 bytes, are read per token.

This is Google’s Per-Layer Embeddings (PLE) idea from Gemma 3n, applied to a microcontroller’s memory map. Of the 28.9M stored parameters, 25M sit in that flash lookup table, so most of the model is never loaded into RAM at all. The author reports 9.88 tokens/s end to end and 94.9 ms of compute per token.

The README is just as honest about the limits. The model was trained on TinyStories, a dataset of short, simple stories. It writes simple stories and does not answer questions, follow instructions, write code or know facts. The memory trick lets a big parameter count fit on the chip. It does not make the small reasoning core any smarter.

How a diffusion model generates images on an RP2350

In his write-up, Tim describes a generative image model on the RP2350’s dual Cortex-M33 with 520 KB of RAM. The model and inference code take less than 4 MB of flash. It generates 128 × 128 RGB faces in 10–20 seconds each and can show them on a VGA monitor or send them over USB. Several design choices make this possible:

  1. Latent diffusion. A VAE compresses the 128×128×3 image into a 16×16×8 latent, a 24× reduction. Only the small decoder runs on the chip. The large encoder is used only during training.
  2. Flow matching with 8 steps. The model predicts a velocity toward the clean image and takes small steps along it, eight times.
  3. A tiny diffusion transformer. Two variants have 2.9M and 1.7M parameters and support five classes (gender × smile, plus unconditional). The latent is split into 64 tokens of dimension 128.
  4. Lookup tables instead of compute. The model only ever sees 8 timesteps and 5 classes, so the AdaLN conditioning is stored as precomputed tables.
  5. INT8 quantization. Weights are quantized to INT8 after training, and quantization-aware self-distillation repairs some of the accuracy damage.
  6. Streaming weights with DMA. Weights are streamed from flash into a ping-pong buffer while the previous layer computes. Each weight is reused across all 64 tokens, so flash bandwidth is rarely the bottleneck.

Tim also used SMLAD SIMD intrinsics, ran both cores at 300 MHz (overclocked from the stock 150 MHz), and used a ReLU² activation that increases sparsity, which the inference engine exploits for about 15% faster inference. The VGA framebuffer’s SRAM doubles as scratch space during inference. The fast model takes about 5 s per image. The larger model with classifier-free guidance takes about 20 s.

Side-by-side specs

Item esp32-ai (LLM) pico-faces (diffusion)
Chip ESP32-S3 RP2350 (dual Cortex-M33)
Fast RAM 512 KB SRAM + 8 MB PSRAM 520 KB SRAM
Parameters 28.9M (25M in flash table) 2.9M or 1.7M + small VAE decoder
Weight format 4-bit, 14.9 MB INT8, under 4 MB with engine
Key trick Per-Layer Embeddings in flash Latent DiT + DMA weight streaming
Output 9.88 tokens/s 128×128 RGB face in ~5–20 s
Training data TinyStories FFHQ faces

Limits of an LLM on a microcontroller

None of this replaces a GPU, and neither author claims it does. The LLM writes toy stories. The image model makes one class of small faces. Both rely on hardware tricks such as large external flash, PSRAM, overclocking and DMA, and you have to plan for them in your board design. For scale: MLCommons describes TinyML models as typically under 2M weights, and its MLPerf Tiny benchmark still focuses on keyword spotting, visual wake words, image classification and anomaly detection. Generative models on MCUs are hobbyist and research demos so far, not benchmarked product workloads.

The hardware is moving quickly, though. NXP rates the MCX N94x’s eIQ Neutron NPU at 4.8G INT8 operations per second. ST specifies up to 600 GOPS for the STM32N6’s Neural-ART NPU, with 4.2 MB of contiguous RAM. Once someone shows that a model fits a microcontroller’s memory, NPUs like these are what will make it fast.

What this means for embedded engineers

The lesson is less about the models and more about systems thinking. Both projects succeed because the authors designed the model and the memory map together. If you want to try it yourself:

  1. Reproduce first. esp32-ai ships two commands per model, scripts/fetch_model.sh barista and scripts/deploy.sh barista. The fetch script checks SHA-256 hashes before it installs anything. pico-faces is on GitHub for any RP2350 board.
  2. Write down your memory budget. For each tensor, note its size, how often it is read per token or step, and which tier it lives in (SRAM, PSRAM, internal flash, QSPI/XIP). This table decides your architecture more than any accuracy chart.
  3. Prefer lookups over compute where you can. Embedding tables and precomputed conditioning tables move work from MACs to flash reads.
  4. Learn quantization properly. Start with INT8 post-training quantization. Measure the damage, then try quantization-aware training or distillation. Choose per-channel scales deliberately.
  5. Use the DMA. Double-buffered weight streaming hides flash latency. Check cache coherency and alignment, which are classic sources of C bugs that become hardware failures.
  6. Try the vendor NPU flow. On the FRDM-MCXN947, NXP’s documented path takes a TFLite model through neutron-converter --target mcxn94x and deploys it in MCUXpresso. On the STM32N6, the equivalent is ST Edge AI.

If you are still choosing hardware, our guide to microcontroller vs microprocessor vs SoC walks through the trade-offs. For the wider skill map, see what embedded engineers should learn next for edge AI on microcontrollers.

Key takeaways

  • The verified demos are a 28.9M-parameter LLM on an ESP32-S3 and a 2.9M/1.7M-parameter diffusion transformer on an RP2350. We could not verify the widely repeated “289M on NXP” claim.
  • Memory placement made both possible: weights live in flash and are streamed or looked up, and SRAM is kept for activations.
  • 4-bit and INT8 quantization, latent spaces and lookup tables did as much work as raw compute.
  • The outputs are toy-grade. The techniques are production-relevant.

FAQ

Can you really run an LLM on a microcontroller?

Yes, a small one. esp32-ai runs a 28.9M-parameter model on an ESP32-S3 at 9.88 tokens/s. It writes simple TinyStories-style text and cannot follow instructions or answer factual questions.

Did someone run a 289M-parameter LLM on the NXP FRDM-MCXN947?

We found no public project that shows this. The reports appear to trace back to a roundup about a 28.9M model on an ESP32-S3, so treat the 289M figure as unverified.

How does a diffusion model fit in 520 KB of RAM?

It generates a compressed 16×16×8 latent instead of pixels, stores INT8 weights in flash, streams them in with DMA layer by layer, and reuses the framebuffer SRAM as scratch space.

Which skills matter most for generative AI on MCUs?

Quantization, memory-hierarchy planning, DMA and cache handling, SIMD intrinsics on Cortex-M, and vendor NPU toolchains such as NXP eIQ and ST Edge AI.

Ready to build these skills on real hardware? Explore the embedded systems, IoT and AI courses from the Educational Engineering Team at https://eduengteam.com/wp-content/uploads/2026/10/meta-muse-gadgets-esp32-ai-agent-diagram.jpg.

Leave a Reply